Image generation method and device, equipment and medium

By using user-provided design assistance to separate foreground and background layers with shared attention mechanisms, the method addresses inefficiencies in promotional image generation, improving speed and quality while aligning with user intent.

CN120318092APending Publication Date: 2025-07-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480429.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, generating promotional images requires a lot of manual design and modification, resulting in waste of resources and the generation effect is not in line with the user's wishes.

Method used

By obtaining image design auxiliary information, generating description information of the foreground layer and background layer, using the model to generate image, using the shared attention mechanism for denoising, and combining foreground and background layers.

Benefits of technology

It improves image generation efficiency and effect, reduces waste of artificial resources, and the generated images are more in line with user wishes and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318092A_ABST
    Figure CN120318092A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, equipment and a medium. A specific implementation mode of the method comprises the steps of obtaining image design auxiliary information; the image design auxiliary information at least comprises text guide information; based on the image design auxiliary information, generating condition information; the condition information comprises description information of a foreground image layer corresponding to a to-be-generated target image and a background image layer corresponding to the target image; and generating the target image based on the condition information by using a first model. According to the embodiment, waste of manual resources is avoided, the image generation efficiency is improved, the image generation effect is improved, the generated image better conforms to the willingness of a user, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to methods, apparatuses, devices, and media for image generation. Background Art

[0002] A promotional image is a visual media tool that can be used for publicity and information dissemination. It is usually a flat image, and the content of a promotional image can include patterns, pictures, texts, colors, etc. Through promotional images, advertising content can be displayed to the public, warning information can also be publicized, and cultural dissemination can be carried out, etc. Promotional images can be used in different usage scenarios, such as product posters, social media content, business cards, invitations, movie promotion covers, book covers, advertising promotion pictures, etc. Therefore, promotional images play a very important role in aspects such as information transmission, brand promotion, and emotional communication. Currently, a method for generating promotional images is needed. Summary of the Invention

[0003] Embodiments of the present disclosure describe a method, apparatus, device, and medium for image generation.

[0004] According to a first aspect, a method for image generation is provided. The method includes: obtaining image design auxiliary information; the image design auxiliary information at least includes text guiding information; based on the image design auxiliary information, generating conditional information; the conditional information includes description information of a foreground layer corresponding to a target image to be generated and a background layer corresponding to the target image; using a first model, based on the conditional information, generating the target image.

[0005] According to a second aspect, an apparatus for image generation is provided. The apparatus includes: an obtaining unit configured to obtain image design auxiliary information; the image design auxiliary information at least includes text guiding information; a determining unit configured to generate conditional information based on the image design auxiliary information; the conditional information includes description information of a foreground layer corresponding to a target image to be generated and a background layer corresponding to the target image; a generating unit configured to use a first model, based on the conditional information, generate the target image.

[0006] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is caused to execute the method according to any one of the first aspect.

[0007] According to a fourth aspect, an electronic device is provided, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of the first aspect is implemented.

[0008] According to a solution for image generation provided by an embodiment of the present disclosure, by obtaining image design auxiliary information provided by a user, based on the image design auxiliary information, condition information is generated, and the condition information includes description information of a background layer corresponding to a target image to be generated and a foreground layer corresponding to the target image. Using a first model, based on the condition information, the target image is generated. Thereby, the waste of human resources is avoided, not only the efficiency of image generation is improved, but also the effect of image generation is improved, making the generated image more in line with the user's wishes and enhancing the user experience.

[0009] Since in this embodiment, description information corresponding to the foreground layer and the background layer is respectively generated based on the image design auxiliary information, thereby decoupling the generation of the foreground part and the background part, facilitating the user to modify the foreground part or the background part respectively, further improving the efficiency of image generation, and contributing to enhancing the user experience.

[0010] Since in this embodiment, the generation of a text layer including text content and a main image layer including a material image is decoupled, further improving the effect of image generation, and also providing the user with a richer way of image design, contributing to enhancing the user experience.

[0011] Since in this embodiment, the background layer can be generated according to the foreground layer and the description information for the background layer, considering the degree of fusion between the layout of the content in the foreground layer and the background layer, so that after the background layer and the foreground layer are merged, the picture is more harmonious, further improving the effect of image generation.

[0012] Since in this embodiment, a shared attention mechanism is adopted to perform denoising processing on the image to be processed, enabling the first model to better predict the noise to be removed according to the shared features, thereby further improving the picture effect of the background layer and making the degree of fusion of the picture after the background layer and the foreground layer are merged higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a schematic diagram of an application scenario for image generation shown according to an exemplary embodiment of the present disclosure;

[0014] Figure 2 is a schematic diagram of an exemplary system architecture applying an embodiment of the present disclosure;

[0015] Figure 3 is a flowchart of a method for image generation shown according to an exemplary embodiment of the present disclosure;

[0016] Figure 4 is another schematic diagram of an application scenario for image generation shown according to an exemplary embodiment of the present disclosure;

[0017] Figure 5It is a schematic diagram of another application scenario of image generation shown according to an exemplary embodiment of the present disclosure;

[0018] Figure 6 It is a block diagram of a device for image generation shown according to an exemplary embodiment of the present disclosure;

[0019] Figure 7 It is a schematic block diagram of an electronic device provided in some embodiments of the present disclosure. Detailed implementation manners

[0020] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0021] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0022] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0023] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0024] The technical solutions provided by the present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and do not limit the invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0025] A promotional image is a visual media tool that can be used for promotion and information dissemination. It is usually a flat image and can include content such as patterns, pictures, text, and colors. Through promotional images, advertising content can be displayed to the public, warning information can be promoted, and cultural dissemination can also be carried out. Promotional images can be used in different scenarios, such as product posters, social media content, business cards, invitations, movie promotion covers, book covers, advertising promotion pictures, etc. Therefore, promotional images play a very important role in information transmission, brand promotion, and emotional communication. In the related art, usually, designers manually typeset according to the requirements and materials provided by users and manually design promotional images. However, in order to meet the needs of users, designers may need to continuously modify the designed promotional images, resulting in a great waste of human resources.

[0026] A scheme for image generation provided by the present disclosure generates conditional information based on the image design auxiliary information obtained from the user. The conditional information includes the description information of the background layer corresponding to the target image to be generated and the foreground layer corresponding to the target image. Using the first model, based on the conditional information, the target image is generated. Thereby, the waste of human resources is avoided, not only the efficiency of image generation is improved, but also the effect of image generation is improved, making the generated image more in line with the user's wishes and enhancing the user experience.

[0027] See Figure 1 , which is a schematic diagram of an application scenario of image generation shown according to an exemplary embodiment.

[0028] As Figure 1 shown, for example, if a user wants to generate an advertising poster, the user can input the design auxiliary information for the advertising poster to the image generation client through a terminal device. The design auxiliary information can include text guidance information 101 and a material image 102. Among them, the specific content included in the text guidance information 101 can be "Design a cool promotional poster with a bold and avant-garde aesthetic style, featuring eye-catching graphics, bright colors, and dynamic fonts, emphasizing the main offers, and including product pictures, store logos, and space for a prominent call-to-action button. This design is intended to attract customers' attention."

[0029] The image generation client can generate layout information for the foreground layer 103 and description information 105 for the background layer 104 based on the text guidance information 101 and the material image 102 through the model M1. Among them, the layout information for the foreground layer 103 can include the text content in the foreground layer 103 and the text layout information of the text content in the foreground layer 103, and also include the material layout information of the material image 102 in the foreground layer 103. The description information 105 can specifically be "black background, green texture strokes, white torn edges, gray ellipses, and green neon shapes".

[0030] Next, the image generation client can use the renderer R to render a text layer based on the text content in the foreground layer 103 and the text layout information of the text content in the foreground layer 103. Based on the material layout information of the material image 102 in the foreground layer 103, a main image layer is rendered, and then the text layer and the main image layer are merged to obtain the foreground layer 103.

[0031] Then, the image generation client can input the foreground layer 103 and the description information 105 for the background layer 104 as conditional information 106, together with a random noise image 107, into the diffusion model M2. The diffusion model M2 can perform multi-step denoising processing on the noise image 107 based on the foreground layer 103 and the description information 105 for the background layer 104 to obtain the background layer 104. Finally, the image generation client can perform a merging operation on the foreground layer 103 and the background layer 104 to obtain the target image 108, which is an advertising poster for the user's needs.

[0032] It should be noted that Figure 1 The embodiment is described by taking the image generation client directly generating the target image as an example. In other embodiments, the image generation client can also transmit the text guidance information 101 and the material image 102 to an image generation server deployed on the service platform through the network. The image generation server can generate the target image 108 based on the text guidance information 101 and the material image 102, and transmit the target image 108 to the image generation client through the network to provide the target image 108 to the user. For details, see Figure 2 the embodiment.

[0033] Figure 2 It is a schematic diagram of an exemplary system architecture for applying the embodiments of the present disclosure.

[0034] As Figure 2 shown, the system architecture 200 can include a terminal device 202, a network 203, and a server 204. It should be understood that Figure 2The number or type of terminal devices, networks, and servers therein is merely illustrative. According to implementation requirements, there can be any number or type of terminal devices, networks, and servers.

[0035] Network 203 is used to provide a medium for communication links between terminal devices and servers. Network 203 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0036] An image generation client is installed in the terminal device 202. The terminal device 202 can interact with the server through the network 203 to receive or send requests, information, etc. The terminal device 202 can be various electronic devices, including but not limited to smartphones, tablets, laptop portable computers, desktop computers, and smart wearable devices, etc.

[0037] An image generation server is deployed in the server 204. The server 204 can perform processing such as storing and analyzing the received data, and can also send control commands or requests to the terminal device or other servers. The server can provide an image processing service in response to a user's service request. It can be understood that one server can provide one or more services, and the same service can also be provided by multiple servers.

[0038] Based on Figure 2 The system architecture shown, in the embodiments of the present disclosure, the user 201 can input image design auxiliary information for generating a target image through the terminal device 202. The image design auxiliary information can include but not limited to text guiding information and material images. The terminal device 202 can transmit the image design auxiliary information to the server 204 through the network 203. After receiving the image design auxiliary information, the server 204 can generate a target image based on the image design auxiliary information. Finally, the server 204 can return the target image to the terminal device 202 through the network 203, enabling the user 201 to view and save the target image through the terminal device 202.

[0039] The following will describe the present disclosure in detail in conjunction with specific embodiments.

[0040] Figure 3 It is a flowchart of a method for image generation shown according to an exemplary embodiment. This method can be applied to an image generation client or an image generation server. In this embodiment, the image generation client is installed in a terminal device, and the terminal device can include but not limited to mobile terminal devices such as smartphones, smart wearable devices, tablets, laptop computers, and desktop computers, etc. The image generation server is deployed in a service platform, and the service platform can be implemented as any device, server, or device cluster with computing and processing capabilities. The method can include the following steps:

[0041] As Figure 3 shown, in step 301, image design auxiliary information is obtained.

[0042] In this embodiment, the image design auxiliary information may be information for assisting in generating a target image, and the target image may be a promotional image such as an advertising poster, a movie promotion cover, an invitation, etc. Generally speaking, a promotional image usually may include text content, a foreground main object, and a background picture, etc. Therefore, the image design auxiliary information may include information capable of indicating the text content, the foreground main object, and the background picture in the target image, etc. Specifically, the image design auxiliary information may at least include text guiding information, and the text guiding information may include but is not limited to the text content to be displayed in the target image and the user's design intention information, etc., where the user's design intention information may include, for example, the image style and atmosphere that the user wants to design. Optionally, the image design auxiliary information may further include at least one material image, and the material image may be an image including the foreground main object. The foreground main object may be an object that the user wants to prominently display in the target image. For example, in a movie promotion cover, the foreground main object may be the main character of the movie, and the material image may be a still photo of an important scene in the movie. Another example is that in a product advertising promotion picture, the foreground main object may be the product to be promoted, and the material image may be a photo of the product.

[0043] In one implementation, the image design auxiliary information may only include text guiding information. Refer to Figure 4 As Figure 4 shown, in the image design auxiliary information 401 input by the user, only the text guiding information 402 is included. In another implementation, in addition to including text guiding information, the image design auxiliary information may further include a material image. Refer to Figure 5 As Figure 5 shown, in the image design auxiliary information 501 input by the user, the text guiding information 502 and the material image 503 may be included. It should be noted that the user may provide one material image or multiple material images.

[0044] In step 302, condition information is generated based on the image design auxiliary information.

[0045] In this embodiment, based on the image design auxiliary information, a foreground layer and description information for the background layer may be generated as condition information. The foreground layer may be a layer including the foreground picture in the target image, and the background layer may be a layer including the background picture in the target image. The description information of the background layer may be a description of the content information and texture information of the background picture in the background layer, and the description information may be information in text form.

[0046] For example, a generative model can be utilized to directly generate and output the foreground layer and the description information for the background layer according to the image design assistance information. Another example is that the image design assistance information can be input into a second model, enabling the second model to generate the description information for the background layer and the layout information for the foreground layer based on the image design assistance information. Then, a rendering operation is performed based on the layout information to obtain the foreground layer. Among them, the second model can be, for example, a multi-modal large language model. Since in this embodiment, the description information corresponding to the foreground layer and the background layer is generated respectively based on the image design assistance information, the generation of the foreground part and the background part is decoupled, facilitating the user to modify the foreground part or the background part separately, further improving the efficiency of image generation and contributing to enhancing the user experience.

[0047] In one implementation manner, the image design assistance information may only include text guidance information. Then, the layout information for the foreground layer may include the text content in the foreground layer and the text layout information of the text content in the foreground layer. Among them, the text layout information may include, but is not limited to, information such as the font type, font color, alignment method, font position, and layout method of the font corresponding to the text content in the foreground layer. The second model can be used to obtain the description information for the background layer, the text content in the foreground layer, and the text layout information of the text content in the foreground layer based on the image design assistance information. Then, a rendering operation is performed based on the text layout information to obtain the text layer, and the text layer is used as the foreground layer.

[0048] In another implementation manner, the image design assistance information may include text guidance information and a material image. Then, the layout information for the foreground layer may include the text content in the foreground layer, the text layout information of the text content in the foreground layer, and the material layout information of the material image in the foreground layer. Among them, the material layout information may include information such as the position, size, and layout method of the material image in the foreground layer. The second model can be used to obtain the description information for the background layer, the text content in the foreground layer, the text layout information of the text content in the foreground layer, and the material layout information of the material image in the foreground layer based on the image design assistance information. Then, a rendering operation is performed based on the text layout information to obtain the text layer, a rendering operation is performed based on the material layout information to obtain the main image layer, and the text layer and the main image layer are merged to obtain the foreground layer.

[0049] Since in this embodiment, the generation of the text layer including the text content and the main image layer including the material image is decoupled, the effect of image generation is further improved, and it also provides the user with a more abundant image design approach, contributing to enhancing the user experience.

[0050] In step 303, a target image is generated based on the conditional information using the first model.

[0051] In this embodiment, a target image can be generated based on the conditional information using the first model. Among them, the first model can be, for example, a diffusion model. Specifically, first, a noise image can be obtained. The noise image can be an image randomly generated using Gaussian white noise. This embodiment places no limitation on the specific generation method of the noise image. Then, using the first model, based on the foreground layer and the description information for the background layer, multi-step denoising is performed on the noise image to obtain the background layer, and then the background layer and the foreground layer are merged to obtain the target image.

[0052] Since this embodiment can generate the background layer according to the foreground layer and the description information for the background layer, taking into account the fusion degree between the layout of the content in the foreground layer and the background layer, the picture becomes more harmonious after the background layer and the foreground layer are merged, further improving the effect of image generation.

[0053] In one implementation, feature extraction can be performed on the foreground layer and the description information for the background layer to obtain a conditional vector. Then, the conditional vector and the noise image are input into the first model, and the first model directly performs multi-step denoising processing on the noise image based on the conditional vector to obtain the background layer.

[0054] In another implementation, feature extraction can also be performed on the foreground layer and the description information for the background layer to obtain a conditional vector, and the conditional vector and the noise image are input into the first model. The first model can perform multi-step denoising processing on the noise image based on the conditional vector through a shared attention mechanism to obtain the background layer. Among them, for any step of the multi-step denoising processing, the following steps can be included: First, obtain the image to be processed. If it is the first time to perform denoising processing, obtain the noise image as the image to be processed. If it is not the first time to perform denoising processing, obtain the intermediate image obtained from the previous denoising processing as the image to be processed. Then, through the shared attention mechanism, based on the conditional vector and the image to be processed, calculate the shared feature. Specifically, the hidden layer feature obtained based on the conditional vector and the hidden layer feature obtained based on the image to be processed can be input into the shared attention network to calculate the shared feature. Finally, based on the shared feature, perform denoising processing on the image to be processed.

[0055] Since this embodiment uses a shared attention mechanism to perform denoising processing on the image to be processed, the first model can better predict the noise to be removed according to the shared feature, thereby further improving the picture effect of the background layer and making the fusion degree of the picture higher after the background layer and the foreground layer are merged.

[0056] A method for image generation provided by the present disclosure obtains image design assistance information provided by a user, generates conditional information based on the image design assistance information, where the conditional information includes description information of a foreground layer corresponding to a target image to be generated and a background layer corresponding to the target image, and uses a first model to generate the target image based on the conditional information. Thereby, the waste of human resources is avoided, not only the efficiency of image generation is improved, but also the effect of image generation is improved, making the generated image more in line with the user's wishes and enhancing the user experience.

[0057] Next, a complete and specific application example is used to illustrate the solution and effect of the present disclosure schematically.

[0058] First, refer to Figure 4 , as Figure 4 shown, a user can input image design assistance information 401 to an image generation client through a terminal device. The image design assistance information 401 includes text guidance information 402, and the specific content of the text guidance information 402 can include "Place the text 'week end graze' in an arc above the product, and the background is dark green to highlight the product. The light projection on the product gives a natural and peaceful feeling."

[0059] The image generation client can generate conditional information 403 based on the image design assistance information 401. The conditional information 403 can include a foreground layer 404 and description information 405 for the background layer. The specific content of the description information 405 can include "What is shown in the photo is a piece of cake roll. The outer layer is soft and golden brown, the inner layer is soft and white, placed on a wooden board, and the background is dark green. The cake roll is cut open to reveal fresh fruits and powdered sugar inside, forming a contrast between the textured outer layer and the smooth inner layer." The image generation client can generate a background layer 406 according to the conditional information 403, and merge the foreground layer 404 and the background layer 406 to obtain a target image 407.

[0060] Next, another complete and specific application example is used to illustrate the solution and effect of the present disclosure schematically.

[0061] First, refer to Figure 5 , as Figure 5As shown, the user can input image design assistance information 501 into the image generation client through the terminal device. The image design assistance information 501 includes text guidance information 502 and material images 503. The specific content of the text guidance information 502 can include "Design a compelling banner advertisement for 'Golden glow serum, Golden glow serum promotion'. Use bold text below the product name. The design should adopt gold and yellow tones and include elements that glow or have a glowing effect. Add a button that displays '70% off' to increase an additional level of interaction. The slogan 'Cherish the natural beauty of your skin and use our golden washable skin care brand' can be added to emphasize the natural skin beautifying effect of the product."

[0062] The image generation client can generate condition information 504 based on the image design assistance information 501. The condition information 504 can include a foreground layer 505 and description information 506 for the background layer. The specific content of the description information 506 can include "golden tone, liquid splash, arch, smooth texture, geometric emphasis, balanced composition". The image generation client can generate a background layer 507 according to the condition information 504, and merge the foreground layer 505 and the background layer 507 to obtain a target image 508.

[0063] Corresponding to the foregoing method embodiments of image generation, the present disclosure also provides an embodiment of an image generation device.

[0064] As Figure 6 shown, Figure 6 is a block diagram of an image generation device shown by the present disclosure according to an exemplary embodiment. The device can include: an acquisition module 601, a determination unit 602, and a generation unit 603.

[0065] Among them, the acquisition unit 601 is configured to acquire image design assistance information, and the image design assistance information includes at least text guidance information.

[0066] The determination unit 602 is configured to generate condition information based on the image design assistance information. The condition information includes a foreground layer corresponding to the target image to be generated and description information for the background layer corresponding to the target image.

[0067] The generation unit 603 is configured to use a first model to generate a target image based on the condition information.

[0068] In some embodiments, the determination unit 602 is configured to: input image design assistance information into a second mold to obtain description information output by the second model and layout information for a foreground layer, and perform a rendering operation based on the layout information to obtain the foreground layer.

[0069] In other embodiments, the layout information includes the text content in the foreground layer and the text layout information of the text content in the foreground layer.

[0070] In other embodiments, the image design assistance information further includes at least one material image for generating the target image, and the layout information further includes the material layout information of the material image in the foreground layer.

[0071] In other embodiments, the determination unit 602 performs a rendering operation based on the layout information in the following manner to obtain the foreground layer: perform a rendering operation based on the text layout information to obtain a text layer, perform a rendering operation based on the material layout information to obtain a main image layer, and merge the text layer and the main image layer to obtain the foreground layer.

[0072] In other embodiments, the generation unit 603 is configured to: obtain a noise image, use a first model to perform multi-step denoising processing on the noise image based on condition information to obtain a background layer, and merge the background layer and the foreground layer to obtain the target image.

[0073] In other embodiments, for any step of the multi-step denoising processing, the generation unit performs the following steps: obtain an image to be processed, where if it is the first time to perform denoising processing, obtain the noise image as the image to be processed. If it is not the first time to perform denoising processing, obtain the intermediate image obtained from the previous denoising processing as the image to be processed, calculate a shared feature based on the condition information and the image to be processed through a shared attention mechanism. Based on the shared feature, perform denoising processing on the image to be processed.

[0074] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure. Those of ordinary skill in the art can understand and implement it without creative work.

[0075] Next, refer to Figure 7 , Figure 7A schematic block diagram of an electronic device provided by some embodiments of the present disclosure. The electronic device 920 is, for example, suitable for implementing the method for image generation provided by the embodiments of the present disclosure. The electronic device 920 may be a terminal device or the like, and may be used to implement a client or a server. The electronic device 920 may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that Figure 7 The electronic device 920 shown is merely an example and will not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0076] As Figure 7 shown, the electronic device 920 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 921, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 922 or the program loaded from the storage device 928 into the random access memory (RAM) 923. In the RAM 923, various programs and data required for the operation of the electronic device 920 are also stored. The processing device 921, the ROM 922, and the RAM 923 are connected to each other through a bus 924. The input / output (I / O) interface 925 is also connected to the bus 924.

[0077] Generally, the following devices may be connected to the I / O interface 925: an input device 926 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 927 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 928 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 929. The communication device 929 may allow the electronic device 920 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 920 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 920 may alternatively implement or have more or fewer devices. Figure 7 Each block shown in

[0078] According to an embodiment of the present disclosure, the above method for image generation can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the above method for image generation. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by a processing device 921, the functions defined in the method for image generation provided by the embodiments of the present disclosure can be implemented.

[0079] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method provided by the present disclosure.

[0080] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0081] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0082] The various embodiments in the present disclosure are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the storage medium and the computing device, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0083] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present disclosure can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0084] The above-described specific embodiments further elaborate on the objectives, technical solutions, and beneficial effects of the embodiments of the present disclosure. It should be understood that the above is only the specific embodiments of the embodiments of the present disclosure and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present disclosure shall be included in the protection scope of the present invention.

Claims

1. A method for image generation, the method comprising: Obtaining image design auxiliary information; The image design auxiliary information at least includes text guidance information; Based on the image design auxiliary information, generating conditional information; The conditional information includes description information of a foreground layer corresponding to a target image to be generated and a background layer corresponding to the target image; Using a first model, based on the conditional information, generating the target image.

2. The method according to claim 1, wherein The generating conditional information based on the image design auxiliary information includes: Inputting the image design auxiliary information into a second model to obtain the description information output by the second model and layout information for the foreground layer; Performing a rendering operation based on the layout information to obtain the foreground layer.

3. The method according to claim 2, wherein, The layout information includes text content in the foreground layer and text layout information of the text content in the foreground layer.

4. The method according to claim 3, wherein, The image design auxiliary information further includes at least one material image for generating the target image; the layout information further includes material layout information of the material image in the foreground layer.

5. The method according to claim 4, wherein, The performing a rendering operation based on the layout information to obtain the foreground layer includes: Performing a rendering operation based on the text layout information to obtain a text layer; Performing a rendering operation based on the material layout information to obtain a main image layer; Merging the text layer and the main image layer to obtain the foreground layer.

6. The method according to claim 1, wherein The using a first model, based on the conditional information, generating the target image includes: Obtaining a noise image; Using the first model, performing multi-step denoising processing on the noise image based on the conditional information to obtain the background layer; Merging the background layer and the foreground layer to obtain the target image.

7. The method according to claim 6, wherein Any one of the multi-step denoising processing includes the following steps: Obtaining an image to be processed; wherein, if performing denoising processing for the first time, obtaining the noise image as the image to be processed; if not performing denoising processing for the first time, obtaining the intermediate image obtained from the previous denoising processing as the image to be processed; Calculating shared features based on the conditional information and the image to be processed through a shared attention mechanism; Based on the shared features, performing denoising processing on the image to be processed.

8. An image generation apparatus, the apparatus comprising: An obtaining unit configured to obtain image design auxiliary information; The image design auxiliary information at least includes text guidance information; A determining unit configured to generate conditional information based on the image design auxiliary information; The conditional information includes description information of a foreground layer corresponding to a target image to be generated and a background layer corresponding to the target image; A generating unit configured to use a first model, based on the conditional information, generate the target image.

9. A computer program product, comprising a computer program, where when the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1-7.

11. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-7 is implemented.