Image generation method and device and electronic equipment

By adjusting the layout template and controlling the denoising processing of the diffusion model, the problem of insufficient rationality of image layout in the prior art is solved, and high-quality image layout and visual effects are achieved.

CN120014115APending Publication Date: 2025-05-16NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411996904.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The images generated by the existing diffusion model have shortcomings in the rationality of layout structure, and it is difficult to meet the high-quality requirements of posters, advertisements and other images in layout.

Method used

By obtaining the target text and the initial image, the preset layout template is adjusted to generate layout adjustment information, and then the diffusion model is controlled for denoising, and the final image is generated. This method ensures that the target object and the attached object are laid out according to the layout parameters.

Benefits of technology

It realizes that users accurately control image layout through text and images, improves the rationality and quality of image layout, and enhances visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014115A_ABST
    Figure CN120014115A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device and electronic equipment. The method comprises the following steps: acquiring a target text and an initial image; wherein the initial image comprises a target object; adjusting a preset layout template based on the target text and the initial image to obtain layout adjustment information; the layout adjustment information comprises at least part of layout elements in the layout template and layout parameters after the at least part of layout elements are adjusted; the layout element comprises a target object and an affiliated object related to the target object; generating a control condition based on the layout adjustment information and the initial image; and controlling a preset diffusion model to carry out denoising processing through the control condition, and generating a final image. Wherein the final image comprises the target object and an affiliated object associated with the target object, and the target object and the affiliated object are arranged according to the layout parameters. According to the mode, the reasonability of image layout is improved, and the quality and the visual effect of the image are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to an image generating method, device and electronic equipment. Background Art

[0002] The diffusion model can generate images based on the user's input data, which includes text descriptions, reference images, etc. For image generation tasks such as poster images and advertising images, there are high requirements for the layout of images, but the images generated by the diffusion model often have poor structural layout rationality, resulting in image quality and visual effects that are difficult to meet user needs. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide an image generation method, device and electronic device, so that users can accurately control the image layout through input text and images, improve the rationality of the image layout, and further improve the quality and visual effect of the image.

[0004] In a first aspect, an embodiment of the present invention provides an image generation method, the method comprising: acquiring a target text and an initial image; wherein the initial image comprises a target object; based on the target text and the initial image, adjusting a preset layout template to obtain layout adjustment information; wherein the layout template comprises: layout elements and default parameters of the layout elements; the layout adjustment information comprises: at least some of the layout elements in the layout template, and layout parameters after at least some of the layout elements are adjusted; the layout elements comprise a target object and subsidiary objects related to the target object; based on the layout adjustment information and the initial image, generating control conditions; controlling a preset diffusion model through the control conditions to perform denoising processing to generate a final image; wherein the final image comprises the target object and subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0005] In a second aspect, an embodiment of the present invention provides an image generating device, the device comprising: an acquisition module, used to acquire a target text and an initial image; wherein the initial image comprises a target object; a layout module, used to adjust a preset layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout template comprises: layout elements and default parameters of the layout elements; the layout adjustment information comprises: at least some of the layout elements in the layout template, and layout parameters after at least some of the layout elements are adjusted; the layout elements comprise a target object and subsidiary objects related to the target object; a control module, used to generate control conditions based on the layout adjustment information and the initial image; a denoising module, used to control a preset diffusion model to perform denoising processing through the control conditions to generate a final image; wherein the final image comprises the target object and subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0006] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned image generation method.

[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned image generation method.

[0008] The embodiments of the present invention bring the following beneficial effects:

[0009] The above-mentioned image generation method, device and electronic device obtain target text and initial image; wherein the initial image includes a target object; based on the target text and the initial image, a preset layout template is adjusted to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters after at least some of the layout elements are adjusted; the layout elements include the target object and subsidiary objects related to the target object; based on the layout adjustment information and the initial image, a control condition is generated; a preset diffusion model is controlled by the control condition to perform denoising processing to generate a final image; wherein the final image includes the target object and the subsidiary objects related to the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0010] In this method, the user inputs the target text and the initial image, and the layout adjustment information is generated based on the target text and the initial image. The denoising process of the image is then controlled through the layout adjustment information so that the image layout of the final image matches the data input by the user. The user can accurately control the image layout through the input text and image, which improves the rationality of the image layout and further improves the image quality and visual effect.

[0011] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0012] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0014] Figure 1 A flowchart of an image generation method provided by an embodiment of the present invention;

[0015] Figure 2 A schematic diagram of the structure of a layout decoder provided by an embodiment of the present invention;

[0016] Figure 3 A schematic diagram of generating a second layout feature based on layout adjustment information provided by an embodiment of the present invention;

[0017] Figure 4 An example flowchart of an image generation method provided by an embodiment of the present invention;

[0018] Figure 5 An example diagram of layout adjustment information, text and final image provided by an embodiment of the present invention;

[0019] Figure 6 A schematic diagram of an image generating device provided by an embodiment of the present invention;

[0020] Figure 7 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0022] Diffusion models, such as the Stable Diffusion model, can generate images based on user input data, such as text descriptions, reference images, and other data. In practical applications, image generation for specific tasks often requires more sophisticated control and higher quality requirements. For example, in the generation of product posters and advertising images, it is necessary not only to generate attractive images, but also to ensure that the products in the images are accurately presented, the text information is clear and readable, and the layout is reasonable and meets aesthetic standards. Traditional image generation methods may have shortcomings in dealing with these complex requirements and it is difficult to take multiple factors into account at the same time.

[0023] In the related art, the layout of images can be achieved in two ways.

[0024] One is to rely on heuristic algorithms and templates to generate layouts based on fixed patterns and rules, which is difficult to cope with diverse product features, text content and design requirements. For example, for different types of products, such as electronic products, fashion products, food, etc., the key points that need to be highlighted in the poster and the appropriate layout method may be different, but heuristic algorithms and templates are often unable to be flexibly adjusted according to specific circumstances, resulting in the generated layout may be stereotyped and cannot attract the attention of the target audience well. In actual applications, each brand or product may have its own unique brand image and promotional strategy, requiring personalized poster layouts to convey specific information and emotions. However, fixed patterns and rules cannot meet such personalized requirements.

[0025] Therefore, although the approach of relying on heuristic algorithms and templates can generate layouts to a certain extent, it lacks flexibility and adaptability due to its fixed patterns and rules, and is difficult to meet diverse needs.

[0026] Another layout method is the content-aware method based on deep learning, which predicts the layout based on the content of the pre-generated image. However, in practical applications, it has high requirements on the quality and content integrity of the input image. If the input image has problems such as noise, blur or unclear content, it may affect the accuracy and reliability of layout generation. For example, when the background of the product image is complex or the features of the product itself are not obvious, it may be difficult for the model to accurately identify the position and size of the product, resulting in an unreasonable generated layout.

[0027] Although the content-aware approach based on deep learning can learn some layout patterns to a certain extent, it is still insufficient in terms of the diversity of generated layouts. It may tend to generate some common layout patterns, but it is difficult to effectively generate some innovative or special layouts. At the same time, in terms of the controllability of the layout, it may be difficult for users to directly intervene and adjust the generation process in detail to meet specific design requirements.

[0028] Therefore, although the content-aware method based on deep learning can predict layout based on image content, it has strict requirements on image content and there is still room for improvement in the diversity and controllability of generated layouts.

[0029] Based on the above problems, an image generation method, device and electronic device provided in the embodiments of the present invention can be applied to the generation of various types of images, such as posters, advertisements and other images.

[0030] See also Figure 1An image generation method is shown, the method comprising the following steps:

[0031] Step S102, obtaining a target text and an initial image; wherein the initial image includes a target object;

[0032] This embodiment is used to generate a poster image, an advertisement image, a cover image, etc. containing a target object. The target object may be various objects such as cosmetics, electronic devices, and food; the initial image contains the target object, for example, the initial image may be an image obtained by photographing the target object, or the initial image may only include the target object without including background, text, and other information.

[0033] This embodiment aims to generate a final image containing a target object based on an initial image. The final image contains the target object and may also contain target text, image background, etc.

[0034] The target text is used to describe the target object. The target text may include the text reflected in the final image, and may also include descriptions of the target object, image background, and overall image style, such as a blue background, a fresh style, and the like.

[0035] Step S104, adjusting the preset layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters corresponding to at least some of the layout elements; the layout elements include the target object and subsidiary objects related to the target object;

[0036] The layout template includes multiple layout elements, each layout element has multiple specified attributes, and each specified attribute has one or more attribute values. The layout elements include: target object, image background, image text, etc. For each layout element, the specified attributes such as category, center point, size, etc. are set.

[0037] In a specific implementation, the above default parameters include: a specified attribute, and at least one attribute value preset for the specified attribute; wherein the attached object includes an image background and / or image text; the specified attribute includes: at least one of an object category, a center position, and an object size. wherein the object category may include: a main object, a background, text, etc.

[0038] The subsidiary object may include both the image background and the image text, or may include only one of the image background and the image text. The target object and the subsidiary object are both layout elements in the layout template. For the target object or each subsidiary object, the default parameters of the object include the specified attributes and the corresponding attribute values.

[0039] In one example, the specified attributes may include five attributes, namely, object category, center position, object width, and object height; wherein the center position may be represented by coordinates. For each specified attribute of each layout element, an attribute value range may be preset, and the attribute value range may be at least one attribute value preset for the aforementioned specified attribute. When adjusting the layout template, the attribute value is adjusted within the attribute value range.

[0040] In actual implementation, the layout template can be adjusted through self-attention, cross-attention and other algorithms. For example, the target text, initial image and layout template are input into the algorithm, the adjustment probability of each attribute value is output, and then the corresponding attribute value is adjusted according to the probability.

[0041] The image generation process will undergo denoising processing at multiple time steps. For each time step, further adjustments can be made based on the layout adjustment information of the previous time step to obtain the layout adjustment information of the time step.

[0042] Among them, the layout adjustment information is the final layout after adjustment based on the layout template. In the process of adjusting the layout template, all or part of the layout elements in the layout template may be adjusted. For the layout elements that need to be adjusted, the attribute values ​​of all or part of the specified attributes of the layout elements may be adjusted.

[0043] The layout adjustment information includes layout parameters of the target object and each subordinate object, for example, attribute values ​​of attributes such as object category, center position, object width, and object height.

[0044] In addition, during the adjustment of the layout template, the initial image can also be adjusted based on the layout parameters for the target object in the layout adjustment information, for example, the position and size of the target object in the initial image can be adjusted. Specifically, the layout of the target object in the initial image is adjusted based on the layout adjustment information to generate an adjusted initial image; for example, the layout adjustment information includes information such as the position and size of the target object. When the layout of the target object in the initial image does not match the layout adjustment information, the layout of the target object needs to be adjusted based on the layout adjustment information to obtain an adjusted initial image; if the initial image is adjusted, in subsequent steps, control conditions are generated based on the layout adjustment information and the adjusted initial image.

[0045] Step S106, generating control conditions based on the layout adjustment information and the initial image;

[0046] The control condition is used to control the image layout when the final image is subsequently generated, for example, the position and size of each layout element. For example, the layout adjustment information and the initial image are converted into feature vectors, and the feature vectors are input into the control network to control the diffusion model.

[0047] Step S108, controlling the preset diffusion model through the control condition to perform denoising processing to generate a final image; wherein the final image includes the target object and the subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0048] The diffusion model may specifically be a denoising diffusion probability model (DDPM for short), a diffusion model without classifier guidance (such as a GLIDE model), a distillation diffusion model, and the like.

[0049] The diffusion model performs multi-time-step denoising processing based on a random noise image to generate a final image. In this embodiment, the control condition is generated based on the layout adjustment information and the adjusted initial image. Therefore, the layout adjustment information and the adjusted initial image will guide the denoising process so that the target object and the subsidiary objects in the final image are laid out according to the layout parameters.

[0050] The above-mentioned image generation method obtains a target text and an initial image; wherein the initial image includes a target object; based on the target text and the initial image, a preset layout template is adjusted to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters after at least some of the layout elements are adjusted; the layout elements include the target object and subsidiary objects related to the target object; based on the layout adjustment information and the initial image, a control condition is generated; a preset diffusion model is controlled by the control condition to perform denoising processing to generate a final image; wherein the final image includes the target object and the subsidiary objects related to the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0051] In this method, the user inputs the target text and the initial image, and the layout adjustment information is generated based on the target text and the initial image. The denoising process of the image is then controlled through the layout adjustment information so that the image layout of the final image matches the data input by the user. The user can accurately control the image layout through the input text and image, which improves the rationality of the image layout and further improves the image quality and visual effect.

[0052] In a specific implementation method, a first layout feature corresponding to a layout template, a text feature corresponding to a target text, and a first image feature corresponding to an initial image are generated; a cross-attention process is performed on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result; and layout adjustment information is generated based on the processing result.

[0053] First, the layout template, target text and initial image are converted into the feature space to obtain the corresponding first layout features, text features and first image features. These first layout features, text features and first image features can be embedded features, that is, the feature form of embedding, or other feature forms.

[0054] The extended model needs to perform a denoising process of multiple time steps, and accordingly, in this embodiment, a layout adjustment process of multiple time steps is required. At each time step, a cross-attention process is performed on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result.

[0055] Specifically, the current time step and the first layout feature are normalized to obtain a first intermediate result; the first intermediate result is self-attention processed to obtain a second intermediate result; the second intermediate result, the text feature and the first image feature are cross-attention processed to obtain a third intermediate result; the third intermediate result and the first intermediate result are processed through a feedforward function to obtain a processing result.

[0056] The first intermediate result can be expressed by the following formula:

[0057] h t =AdaLN(e t ,t)

[0058] Among them, h t represents the first intermediate result at time step t, AdaLN represents the adaptive normalization function, e t Represents the first layout feature at time step t.

[0059] The above second intermediate result can be expressed by the following formula:

[0060] a t =h t +SA(h t )

[0061] Among them, a t represents the second intermediate result at time step t, and SA represents the self-attention function. The self-attention function allows the model to pay more attention to the relationship between different parts in the feature vector of the second intermediate result, and to perform weighted updates on the vector by calculating the correlation weight, so as to better capture the intrinsic connection and dependency between the elements in the vector, and to more effectively learn and process layout and other information.

[0062] The above processing results can be expressed by the following formula:

[0063] e t =FF(a t +CA(a t,CAT(e T ,e I )))

[0064] Among them, e t represents the processing result of time step t, FF represents the feedforward function, CA represents the cross attention function, CAT represents the connection operation, and e T Represents text features, e I represents the first image feature; CAT is used to concatenate the feature vectors of the text feature and the first image feature; CA(a t ,CAT(e T ,e I )) represents the aforementioned third intermediate result.

[0065] The cross attention function is used to combine text features and image images, and interact with the feature vector in the second intermediate result, adjust the representation of each element in the feature vector according to the characteristics of the text and image, so that the generated layout or image is more in line with the visual characteristics and text semantic requirements of the target object, and establish the connection between information such as text, image and layout in the feature space.

[0066] The feedforward function is used to further transform and integrate the feature vectors after self-attention and cross-attention processing. The feedforward function is usually composed of a fully connected layer, which can perform nonlinear transformations on the feature vectors and extract higher-level feature representations. The feedforward function can further process and fuse the relationship information about layout, image, and text learned by self-attention and cross-attention, so that the feature vectors can better adapt to subsequent decoding and image generation tasks. By effectively extracting and transforming the feature vectors, the feedforward function helps improve the accuracy and quality of the layout and image generated by the model.

[0067] In actual implementation, at the first time step, the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing result of the first time step; in each time step, a cross-attention processing process is performed, and the specific process can refer to the aforementioned embodiment, including multi-step processing methods such as self-attention, cross-attention, and feedforward function.

[0068] In a subsequent time step of the first time step, the first layout feature is updated based on the processing result of the previous time step; the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing result of the subsequent time step.

[0069] The layout features between multiple time steps are iterated cyclically. Based on this, for time step t-1, the previous time step is t; the processing result of the previous time step t updates the first layout feature of time step t-1, and then based on the updated first layout feature, cross-attention processing is performed to obtain the processing result corresponding to time step t-1. In each subsequent time step, a cross-attention processing process is performed. The specific process can refer to the aforementioned embodiment, including multi-step processing methods such as self-attention, cross-attention, and feedforward function.

[0070] See also Figure 2 In the first time step, the layout template is input to the first fully connected layer, that is, the initial value of the layout adjustment information is the layout target. In the subsequent time steps, the layout adjustment information output in the previous time step is input to the first fully connected layer.

[0071] Specifically, based on the processing result of the previous time step, the conversion probability of the first layout feature is determined; based on the conversion probability, the first layout feature is converted to obtain an updated first layout feature.

[0072] The transition probability can be expressed by the following matrix:

[0073]

[0074] Matrix Q t It expresses the transition probability of any specified attribute in any layout element between different time steps, where α t Indicates the probability of maintaining the current attribute value at time step t; β t represents the probability of changing the attribute value at time step t; γ t Represents the probability that the attribute value remains unchanged at time step t.

[0075] As the time step changes, the attribute value will change according to the aforementioned conversion probability until all attribute values ​​no longer change. By converting the probability, the image layout can have a certain degree of randomness and uncertainty, making the image layout output by the diffusion model rich and varied, and improving the diversity and flexibility of the layout form.

[0076] The following formula expresses the total transition probability over multiple time steps for any given attribute:

[0077] q(z t |z0)=v T (z t )·Q t Q t-1 ·…·Q1·v(z0)

[0078] Among them, v T (zt ) is the first layout feature z at time step t t The transposed vector of the one-hot column vector, Q t is the transition probability at time step t, and v(z0) is the one-hot column vector of the first layout feature z0 at time step 0.

[0079] For a specific implementation, see Figure 2 , input the layout template into the first fully connected layer, generate the first layout feature, input the current time step, the first layout feature, the text feature and the first image feature corresponding to the initial image into the preset cross-attention processing network; wherein the cross-attention processing network includes: multiple series-connected conversion modules; the conversion module is also called a Transformer module.

[0080] In each conversion module, the target feature data corresponding to the conversion module is determined, and cross-attention processing is performed on the current time step, the target feature data, the text feature and the first image feature, and an intermediate result is output; wherein, for the first conversion module, the target feature data is the first layout feature, and for the subsequent conversion modules of the first conversion module, the target feature is the intermediate result output by the previous conversion module; the intermediate result output by the last conversion module is input into the second fully connected layer to generate layout adjustment information.

[0081] The first fully connected layer is connected to the first conversion module in the aforementioned cross-attention processing network, and the last conversion module in the cross-attention processing network is connected to the second fully connected layer.

[0082] First, the first layout feature, the text feature corresponding to the target text and the first image feature corresponding to the initial image are input into the first conversion module to obtain the intermediate result output by the first conversion module; wherein, in the first conversion module, the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed; in each conversion module, an adaptive normalization function, a self-attention function, a cross-attention function and a feedforward function are included.

[0083] Then, the current time step, the intermediate result output by the first conversion module, the text features corresponding to the target text, and the first image features corresponding to the initial image are input to the next conversion module to obtain the intermediate result output by the next conversion module, until the intermediate result output by the last conversion module; wherein, in the next conversion module, the current time step, the intermediate result output by the first conversion module, the text features, and the first image features are subjected to cross-attention processing; it should be noted that the structures within the N conversion modules are the same, and in other subsequent conversion modules except the first conversion module, the intermediate result output by the previous conversion model needs to be used as the updated first layout feature, and then the adaptive normalization function processing is performed. At the same time, the text features and the first image features need to be input to each conversion module to participate in the cross-attention function processing in the conversion module.

[0084] Specifically, each conversion module includes: adaptive normalization function, self-attention function, cross-attention function and feedforward function;

[0085] The current time step and the target feature data are normalized by the adaptive normalization function to obtain a first intermediate result; and the first intermediate result is self-attentioned by the self-attention function to obtain a second intermediate result.

[0086] Through the self-attention function, the model can pay more attention to the relationship between different parts in the feature vector of the second intermediate result, and perform weighted updates on the vector by calculating the correlation weights, so as to better capture the intrinsic connections and dependencies between the elements in the vector, and more effectively learn and process information such as layout.

[0087] Through the cross-attention function, the second intermediate result, the text feature and the first image feature are cross-attention processed to obtain a third intermediate result; through the feedforward function, the third intermediate result and the first intermediate result are processed by the feedforward function to obtain a processing result.

[0088] The cross attention function is used to combine text features and image images, and interact with the feature vector in the second intermediate result, adjust the representation of each element in the feature vector according to the characteristics of the text and image, so that the generated layout or image is more in line with the visual characteristics and text semantic requirements of the target object, and establish the connection between information such as text, image and layout in the feature space.

[0089] The feedforward function is used to further transform and integrate the feature vectors after self-attention and cross-attention processing. The feedforward function is usually composed of a fully connected layer, which can perform nonlinear transformations on the feature vectors and extract higher-level feature representations. The feedforward function can further process and fuse the relationship information about layout, image, and text learned by self-attention and cross-attention, so that the feature vectors can better adapt to subsequent decoding and image generation tasks. By effectively extracting and transforming the feature vectors, the feedforward function helps improve the accuracy and quality of the layout and image generated by the model.

[0090] Finally, the intermediate result output by the last conversion module is input to the second fully connected layer to generate layout adjustment information, and then the layout of the target object in the initial image is adjusted based on the layout adjustment information to generate the adjusted initial image. The second fully connected layer is connected to the last conversion module and is used to convert the intermediate result in feature form into layout adjustment information.

[0091] The layout adjustment information includes layout parameters of the target object, such as the attribute values ​​of specified attributes of the target object, such as the object category, center position, and object size; based on the layout parameters, the target object in the initial image is adjusted, for example, the size and position of the target object are adjusted, so as to obtain an adjusted initial image.

[0092] After the denoising process of the last time step is completed, the final layout adjustment information Z0 is obtained. Based on the layout adjustment information, the posterior probability of the denoising process can be calculated by the following formula:

[0093]

[0094] Among them, z t represents the processing result at time step t, z t-1 represents the processing result of time step t-1; q(z t-1 ∣z t ,z0) means that when z0 is known, t , convert to z t-1 The probability of q(z t |z0) means the conversion from z0 to z t The probability of q(z t ∣z t-1 ,z0) means that when z0 is known, t-1 , convert to z t The probability of q(z t-1 |z0) means the transition from z0 to z t-1 probability.

[0095] In the known z tand z0, the reverse diffusion process can be calculated, that is, z in the denoising process t-1 The posterior probability of z is used as conditional information to guide the diffusion model from z t Derived z t-1 , the diffusion model needs to learn in the process of training given z t and z0, we can deduce z t-1 .

[0096] In a specific implementation, a second layout feature corresponding to the layout adjustment information is generated, and a second image feature corresponding to the initial image is generated; and a control condition is generated based on the second layout feature, the second image feature and the noise feature of the current time step.

[0097] The layout adjustment information can be converted into a mask image. For example, each layout element corresponds to a mask image, which records the area occupied by the layout element. In the mask image, the pixel value in the area occupied by the layout element is 1, and the pixel value outside the area is 0.

[0098] In actual implementation, the target object corresponds to a mask image, and each subsidiary object corresponds to a mask image. However, when the subsidiary object includes multiple image texts, each image text can correspond to a mask image.

[0099] The layout adjustment information can be encoded into a second layout feature through a fully connected layer, a convolutional layer or other networks. The second layout feature can be a feature vector in an embedding form or a feature vector in other forms.

[0100] For details, see Figure 3 , input the layout adjustment information into the first convolutional network, and output the intermediate layout features; fuse the object features between the target object and the subordinate objects in the intermediate layout features to obtain fused features; process the fused features with a multi-head attention mechanism to generate global features; input the global features into the second convolutional network, and output the second layout features.

[0101] The layout adjustment information includes the attribute value of the specified attribute corresponding to each layout element. According to the information such as the center position and the object size, a mask image corresponding to the layout element can be generated. Figure 3 The M here means there are M layout elements in total, that is, the target object and its subsidiary objects together have M objects in total.

[0102] The first convolutional network can be specifically a three-layer convolutional network, through which the mask image is encoded to obtain a feature map LM' corresponding to the mask image. The first convolutional network can gradually extract the features of the mask image, reduce the feature dimension and capture the layout space information contained in the mask image. Each layer of convolution performs a convolution operation by sliding the convolution kernel on the mask image to extract features at different levels. For example, the first layer of convolution may capture some basic edge and shape information, and as the number of layers increases, more advanced spatial structure information is gradually extracted.

[0103] Then, the feature fusion module is used to fuse the layout features of different layout elements to learn the spatial relationship between the layout features of different layout elements. It can be specifically expressed by the following formula:

[0104]

[0105] in, is the fusion feature, is an aggregate token added to the input, 1,j ,…,l M,j are the layout features of different layout elements; j represents the pixel position; CAT is a connection operation that concatenates the aggregation token and the layout features of each layout element.

[0106] In the above feature fusion module, the aggregation token is a marker. After the aggregation token is input into the feature fusion module, it is used to summarize and transmit global information when fusing the features of different layout elements. The aggregation token plays a bridging role in fusing local and global information, which helps to generate more representative and comprehensive layout embeddings, thereby providing more accurate layout guidance for subsequent image generation.

[0107] Furthermore, a specified image operation is performed on the initial image so that the target object satisfies the layout parameters for the target object in the layout adjustment information; the initial image after the specified image operation is input into the third convolutional network, and the second image feature is output.

[0108] The specified image operation includes operations such as scaling, cropping, and translation. The initial image is further processed through the layout parameters of the layout adjustment information to ensure that the initial image can be accurately placed at the position indicated by the layout adjustment information and the size of the target object matches the layout adjustment information.

[0109] The third convolutional network can be specifically a six-layer convolutional network, which can gradually extract the visual features of the target object in the initial image, including color, texture, shape and other features. Each layer of convolution abstracts and extracts the features of the target object at different levels, from low-level pixel features to high-level semantic features.

[0110] Through the processing of the third convolutional network, the second image feature related to the target object is finally obtained, that is, the visual feature vector Z V The visual feature vector is a low-dimensional vector representation that contains the key visual features of the target object so that the visual information of the target object can be accurately incorporated into the layout when generating the final image. For example, when combined with ControlNet to generate the final image, Z V It can provide visual feature information of the target object, so that the target object in the generated final image can more realistically present the original characteristics of the target object.

[0111] In a specific implementation method, control conditions are input into a control network to output control features; noise features and control features are input into a diffusion model to generate an intermediate image; wherein the intermediate image includes a target object and an image background; and image text is rendered in the intermediate image to generate a final image.

[0112] See also Figure 4 The flow chart of the image generation method shown in the figure, the layout decoder is used to generate layout adjustment information and an initial image based on a layout template, a first image feature and a text feature; the layout adjustment information is converted into a second layout feature, and the initial image is converted into a second image feature; the second layout feature, the second image feature and the noise feature are superimposed to generate a control condition, the control condition is input into a control network to generate a control feature, the control feature and the noise feature are input into a diffusion model together to finally generate an intermediate image.

[0113] The encoder is used to encode the original data, such as the initial image, target text, etc., into the feature space, that is, the latent space, to obtain the latent space vector. Different types of encoders, such as encoders based on convolutional neural networks and encoders based on transformer structures, may have different effects on the quality and feature representation of the latent space vector. Corresponding to the encoder, the decoder is used to decode the latent space vector back to the original data, such as image data. The decoder converts the features and information learned in the latent space into visible output by learning the mapping relationship between the latent space vector and the output data.

[0114] The aforementioned diffusion model can specifically be a Unet network, which is a model structure in the diffusion model Stable Diffusion, and is mainly used for encoding and decoding images. U-Net can process image features in feature space. For example, in the process of predicting noise or generating images, U-Net can extract features and perform upsampling or downsampling operations on feature vectors through its unique convolution and pooling structure to achieve the step-by-step generation of images or the prediction of noise.

[0115] The second layout feature and the second image feature are used as control conditions of the control network ControlNet, and the control network is combined with the diffusion model Stable Diffusion to jointly generate an image. The control network plays a guiding and constraining role in the image generation process. The control network adjusts the stable diffusion generation process according to the input control conditions so that the generated image can meet the characteristics of the target object. For example, through the second layout feature, the control network can understand the structure of the layout and the distribution of layout elements, thereby guiding the diffusion model to generate appropriate backgrounds and elements at the corresponding positions; through the second image feature, it can be ensured that the target object in the generated image has the correct visual characteristics.

[0116] The control condition can be expressed by the following formula:

[0117] Z ′ =Z t +Z L +Z V

[0118] Among them, Z ′ is the control condition, Z t is the noise characteristic at time step t, Z L is the second layout feature, Z V Second image feature. By combining this feature with the generated control condition, the diffusion model takes into account the visual information of the layout and the object when generating the image, thereby generating an image that matches the layout and contains the target object.

[0119] The control network adjusts the image generation process of the Stable Diffusion diffusion model according to the input control conditions, so that the generated image can meet the given layout and visual characteristics of the target object, and finely controls the generation and fusion of elements such as the layout and target object during the image generation process.

[0120] The noise feature is a given random noise. On the basis of the noise feature, denoising processing is performed under the control of the aforementioned control conditions to generate an intermediate image; the intermediate image includes the target object and the image background, but the image text is not rendered; therefore, it is necessary to further render the image text on the basis of the intermediate image to obtain the final image.

[0121] It should be noted that the diffusion model and control network in this embodiment need to be trained. During the training stage of the diffusion model and the control network, the prediction noise of the diffusion model is generated based on the model parameters of the diffusion model, the training noise characteristics, the training time step, the text characteristics of the training text and the training control conditions; the noise distance between the prediction noise and the random noise is determined; the loss function value is determined based on the noise distance, and the parameters of the diffusion model and the control network are adjusted according to the loss function value.

[0122] During the training phase, it is necessary to calculate the relevant loss function, which is used to measure the difference between the generated image of the diffusion model and the sample image, so as to adjust the parameters of the diffusion model and the control network through the optimization algorithm to make the generated image closer to the sample image.

[0123] First, the training text can be encoded through a pre-trained CLIP (Contrastive Language-Image Pre-Training) model or a text encoder of a diffusion model to obtain text features of the training text, which can be in the form of embedding features.

[0124] At training time step t in the training phase, random noise ∈ is added to the image features of the sample image, thereby generating the training noise feature Z t , then, the diffusion model outputs the predicted noise ∈ θ .

[0125] Loss Function It can be expressed by the following formula:

[0126]

[0127] in, Represents the expected value of the distance between the predicted noise and the random noise when the random noise changes randomly in the range of (0, 1). is the prediction noise output by the diffusion model, θ is the model parameter of the diffusion model, Z t is the training noise feature, t is the training time step, τ θ (τ) represents the text features of the training text, τ θ represents a text encoder, To predict the noise under the training control condition, the control network is trained according to the training control condition Z ′ Calculated.

[0128] Training control condition Z ′ , which is generated based on the layout features and image features, and is used to guide the image generation process. stands for L2 norm, which is used to calculate the distance between the predicted noise and the random noise.

[0129] In the diffusion model, the diffusion process is a core operation flow, including two stages: forward diffusion and backward diffusion. Forward diffusion is to gradually add noise to the feature vector in the model, that is, the latent space vector, so that the feature vector gradually becomes blurred and random. This process can be regarded as a disturbance to the data, gradually converting the original data, such as the initial image or layout template, into noise data in the feature space.

[0130] In the process of image generation, reverse diffusion gradually recovers image data from noise data. Through the gradual denoising operation of the latent space vector, the diffusion model learns the distribution and characteristics of the data, thereby generating an image that meets the requirements. In this process, the latent space vector is constantly adjusted as the time step changes. The model achieves the goal of promoting the diffusion process and generating images by operating and controlling the latent space vector.

[0131] The denoising process is part of the back diffusion in the diffusion model, and its main purpose is to recover the image data from the latent space vector with noise added. Each step in the denoising process involves the calculation and update of the latent space vector. By continuously optimizing these operations, the model can learn how to effectively remove noise and generate accurate layouts. Among them, the latent space vector is the main operation object and information carrier of the denoising process.

[0132] After the intermediate image is output, text rendering is required for the intermediate image. Specifically, the text color is determined based on the image parameters of the background area in the intermediate image; the text font is determined based on the object theme of the target object, the layout adjustment information, and the text content of the target text; the text placement is determined based on the layout adjustment information; based on the text color, text font, and text placement, the target text is rendered in the intermediate image to generate the final image.

[0133] The purpose of text rendering is to add text information to the intermediate image and ensure that the text information is both clear and readable in the final image and coordinated with the overall style of the image.

[0134] The target object and the background are already included in the aforementioned intermediate image. The image parameters of the background area may include image parameters such as the color and brightness of the background area; specifically, the background area of ​​the intermediate image may be sampled to obtain pixel data of a plurality of sampled pixels, and the image parameters such as the color and brightness of the background area may be calculated based on the pixel data.

[0135] In another method, the background area can be divided into multiple sub-areas. In each sub-area, the pixel data of one or more pixels are sampled, the color and brightness of each sub-area are calculated, and then the color and brightness of each sub-area are averaged to obtain the overall color and brightness of the background area.

[0136] In a specific implementation, if the background brightness of the background area is high, such as the brightness value is greater than a preset brightness threshold, in order to ensure the readability of the text, a darker color can be used as the text color. For example, the text color can be selected from a preset dark color set, such as black, dark gray, etc. This is because dark text can form a sharp contrast on a background with higher brightness and is easier to be recognized.

[0137] If the background color is complex or the color distribution is uneven, more complex strategies can be used. For example, the main color of the background color can be calculated through color clustering and other methods, and then a color that is complementary to or contrasts with the main color can be selected as the text color. Since complementary colors are opposite to each other on the color wheel, they can produce a strong visual contrast, such as the main color is red and the text color is green, or the main color is blue and the text color is yellow. This can make the text stand out against a complex background while maintaining a certain degree of coordination with the overall picture.

[0138] Furthermore, when determining the text font, the subject of the target object is first classified. For example, if the target object is a fashion product, the subject of the object is "fashion"; if the target object is a technology product, the subject of the object is "technology", and so on.

[0139] Different object themes correspond to different text fonts. For example, for an image with an object theme of "fashion", you can choose a font with smooth lines, a strong sense of modernity, and a sense of design from the font library as the text font, such as a thin sans serif font or a handwritten font with unique decorative features. These fonts can convey a sense of fashion and trend, matching the "fashion" theme. For another example, for an image with an object theme of "technology", you can choose a simple and regular font as the text font, such as a sans serif font with obvious geometric shapes, to reflect the simplicity, efficiency, and modernity of the technology theme.

[0140] In addition, when determining the text font, you also need to consider the text content and layout: for example, if it is the title text of a poster, you may need to choose a more eye-catching, larger font as the text font; and for some auxiliary explanatory text, you can choose a relatively small, concise font as the text font to avoid being too complicated and affecting the overall visual effect.

[0141] At the same time, the text font can also be adjusted according to the layout of the text in the image. If the text is located in the center or important area of ​​the image, the text font is adjusted to a more prominent font style; if the text is located at the edge or secondary area of ​​the image, the font prominence can be appropriately reduced, for example, the font size, thickness, etc. can be reduced to maintain overall balance and coordination.

[0142] Furthermore, in the aforementioned layout adjustment information, the text display area has been indicated, and the text display area indicated by the layout adjustment information can be further adjusted so that the text display has visual balance and the final text placement position is obtained.

[0143] Specifically, the text placement is determined based on the layout adjustment information of the image and the principle of visual balance. For example, if the text display area indicated by the layout adjustment information is at the edge of the image, or the text display area is close to or even overlaps with the target object, further adjustments are required to avoid displaying the text too close to the edge of the image or overlapping with the product or other layout elements, so as not to affect readability and visual effects.

[0144] For example, the text can be placed in the blank area above or below the image, or around the target object or in a relatively open area. At the same time, the spatial relationship between the text and the image elements should be considered, and the text and the target object should be kept at an appropriate distance to make the whole picture look comfortable and harmonious.

[0145] After determining the text color, text font, and text placement, you can start text rendering. In actual implementation, you can use a graphics processing library or related image editing tools to add the text with the determined color and font to the intermediate image with appropriate transparency, size, and angle.

[0146] Using image processing libraries in Python, such as the Pillow library, you can draw text onto the intermediate image by specifying parameters such as the text content, font file path, text color, and coordinates of the text placement position. During the drawing process, you can also perform some special effects on the text, such as shadows and strokes, to enhance the visual effect of the text.

[0147] Through text rendering, clear, beautiful and overall style-coordinated text information can be added to the final generated image, making the image more complete and attractive, and meeting the actual needs of posters, advertisements and other images.

[0148] The image generation method provided in this embodiment mainly includes two main parts: layout inference and rendering generation. The layout inference part is used to generate layout adjustment information, and the rendering generation part is used to render the image, realizing the generation process from the target object and text information to the final image.

[0149] In the layout push part, the layout elements are represented by five attributes, and the state transition matrix is ​​defined to describe the transition probability of the attributes during the diffusion process. In the denoising process, the layout decoder is used to convert the layout template into layout features. The layout decoder consists of two fully connected layers and multiple conversion modules. Through the projection of the layout features at time step t, conversion module processing, and cross-attention calculation, it gradually learns the relationship between elements, images, and text features, thereby generating reasonable layout adjustment information.

[0150] The image encoder and text encoder are fine-tuned based on the ALBEF (Align before Fuse) model, and the layout decoder training is set based on the discrete diffusion model. During the training process, relevant parameters such as the number of transformation modules, the number of attention heads, and the feature dimension are adjusted, and the AdamW optimizer is used for optimization to ensure that the layout inference network can accurately generate layout adjustment information.

[0151] In the rendering generation part, the layout adjustment information is converted into a mask image, which is encoded through a three-layer convolutional network. Then, the layout spatial relationship is learned through a feature fusion module to obtain a unified layout representation, and finally the second layout feature is obtained through a convolutional network. Then, the initial image containing the target object is adjusted according to the layout adjustment information, and the second image feature is extracted from the adjusted initial image using a six-layer convolutional network.

[0152] The second layout feature and the second image feature are used as the control conditions of ControlNet to guide the stable diffusion to generate images. During training, loss functions such as image reconstruction loss and perceptual loss are calculated, and the optimization algorithm is used to adjust the parameters of the diffusion model and ControlNet to generate high-quality images.

[0153] After the diffusion model generates the image, the text color and font are determined based on heuristic rules. For example, the appropriate text color is selected based on the background color and brightness, and the appropriate font is selected based on the poster style and theme. The text is then added to the image with the appropriate position and effect to generate the final image.

[0154] In addition, during the training phase of the model, in order to obtain appropriate parameters for some structures or weights in the model, the model needs to be trained with a large number of specific types of images, such as poster images, advertising images, etc. Taking product poster images as an example, the training images need to remove images containing portraits to ensure that the images are mainly focused on product display. At the same time, remove images with unsightly backgrounds to improve the overall quality of images in the training dataset, making them more suitable for product poster generation tasks and ensuring that the generated posters are visually attractive.

[0155] The image encoder and text encoder can be fine-tuned based on ALBEF. The image encoder can use the 12-layer visual transformer ViT-B / 16, and the text encoder can use the first 6 layers of the RoBERTa pre-trained model. Fine-tuning is performed according to the original training objectives of ALBEF, including tasks such as image-text contrastive learning, mask language model, and image-text matching. Through fine-tuning of these tasks, the encoder is able to better extract features of images and texts, as well as learn the associations between them. In addition, an additional training objective is added, which is to predict the category of the product based on the image and text. This objective allows the layout decoder to pay more attention to the relationship between the features of the product and the text description during the learning process, so as to better adapt to the needs of accurate product identification and layout planning in the product poster generation task.

[0156] The layout decoder sets relevant parameters during training, such as the number of conversion modules is 4, the number of attention heads is 8, the feature dimension is 512, the hidden dimension is 2048, and the dropout rate is 0.1. The AdamW optimizer is used for training, and the learning rate is 5×10 -4 , T p By continuously adjusting the model parameters and reducing the value of the loss function, the layout decoder can gradually learn the ability to accurately generate layouts.

[0157] Figure 5 An example of a final image generated by the image generation method of this embodiment is shown, wherein the target object in this example is a sofa, and there are three target texts. Figure 5 Three layout adjustment information based on the target object are output in the image. In different layout adjustment information, the position of the target text in the image is different, and the font, color, etc. of the target text are also different. Figure 5 A final image generated based on each layout adjustment information is also shown.

[0158] The image generation method provided in this embodiment can more accurately learn the relationship between the target object, text and layout, and the generated layout is more reasonable and meets the design principles and aesthetic standards of images such as posters and advertisements. Compared with the methods in the related art, the problems of unreasonable layout and chaotic element distribution are reduced, and the overall quality and visual effect of the image are improved. For example, in the arrangement of the positions of the target object and text, the layout can be more scientifically arranged according to their importance and mutual relationship, so that the information conveyed by the image is clearer and more effective.

[0159] The image generation method provided in this embodiment makes the generated image more realistic and rich in details, texture, color, etc., and can better integrate the target object, background and text information, avoiding problems such as mismatch between the background and the target object style, image blur or distortion. In addition, through the learning of a large amount of data and the optimization of the model, a variety of images can be generated to meet the needs and creativity of different users. For example, for the same target object and the same text description, a variety of images with different styles and layouts can be generated, providing users with more choices and inspiration.

[0160] The image generation method provided in this embodiment can process data and generate images more efficiently; compared with traditional manual design or complex image processing processes, it greatly shortens the generation time and improves production efficiency. Through model training and parameter adjustment, users can more effectively control the image generation results by adjusting the input initial image, text content and related parameters. For example, users can flexibly adjust the font, color and position of the text, as well as the display method of the target object in the image according to the characteristics and promotional focus of the target object, thereby achieving more personalized and accurate image generation.

[0161] See also Figure 6 A schematic diagram of an image generating device shown in FIG. 1 , the device comprising:

[0162] An acquisition module 60 is used to acquire a target text and an initial image; wherein the initial image includes a target object;

[0163] The layout module 62 is used to adjust the preset layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters of at least some of the layout elements after adjustment; the layout elements include the target object and subsidiary objects related to the target object;

[0164] A control module 64, for generating control conditions based on the layout adjustment information and the initial image;

[0165] The denoising module 66 is used to control the preset diffusion model to perform denoising processing through control conditions to generate a final image; wherein the final image includes the target object and the subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0166] The above-mentioned image generation device obtains a target text and an initial image; wherein the initial image includes a target object; based on the target text and the initial image, a preset layout template is adjusted to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters after at least some of the layout elements are adjusted; the layout elements include the target object and subsidiary objects related to the target object; based on the layout adjustment information and the initial image, a control condition is generated; a preset diffusion model is controlled by the control condition to perform denoising processing to generate a final image; wherein the final image includes the target object and the subsidiary objects related to the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0167] In this method, the user inputs the target text and the initial image, and the layout adjustment information is generated based on the target text and the initial image. The denoising process of the image is then controlled through the layout adjustment information so that the image layout of the final image matches the data input by the user. The user can accurately control the image layout through the input text and image, which improves the rationality of the image layout and further improves the image quality and visual effect.

[0168] The above-mentioned default parameters include: a specified attribute, and at least one attribute value preset for the specified attribute; the specified attribute includes: at least one of an object category, a center position, and an object size.

[0169] The above-mentioned layout module is used to: generate a first layout feature corresponding to a layout template, a text feature corresponding to a target text, and a first image feature corresponding to an initial image; perform cross-attention processing on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result; and generate layout adjustment information based on the processing result.

[0170] The above-mentioned layout module is used to: normalize the current time step and the first layout feature to obtain a first intermediate result; perform self-attention processing on the first intermediate result to obtain a second intermediate result; perform cross-attention processing on the second intermediate result, text features and first image features to obtain a third intermediate result; process the third intermediate result and the first intermediate result through a feedforward function to obtain a processing result.

[0171] The above-mentioned layout module is used to: at the first time step, perform cross-attention processing on the current time step, the first layout feature, the text feature and the first image feature to obtain the processing result of the first time step; at the subsequent time step of the first time step, update the first layout feature based on the processing result of the previous time step; perform cross-attention processing on the current time step, the first layout feature, the text feature and the first image feature to obtain the processing results of the subsequent time steps.

[0172] The above-mentioned layout module is used to: determine the conversion probability of the first layout feature based on the processing result of the previous time step; and convert the first layout feature based on the conversion probability to obtain an updated first layout feature.

[0173] The above-mentioned layout module is used to: input the layout template into the first fully connected layer to generate the first layout feature; input the current time step, the first layout feature, the text feature and the first image feature corresponding to the initial image into the preset cross-attention processing network; wherein the cross-attention processing network includes: multiple series-connected conversion modules; in each conversion module, determine the target feature data corresponding to the conversion module, perform cross-attention processing on the current time step, the target feature data, the text feature and the first image feature, and output an intermediate result; wherein, for the first conversion module, the target feature data is the first layout feature, and for the subsequent conversion modules of the first conversion module, the target feature is the intermediate result output by the previous conversion module; the intermediate result output by the last conversion module is input into the second fully connected layer to generate layout adjustment information.

[0174] The above-mentioned conversion module includes: an adaptive normalization function, a self-attention function, a cross-attention function and a feedforward function; the above-mentioned layout module is used to: normalize the current time step and the target feature data through the adaptive normalization function to obtain a first intermediate result; perform self-attention processing on the first intermediate result through the self-attention function to obtain a second intermediate result; perform cross-attention processing on the second intermediate result, text features and first image features through the cross-attention function to obtain a third intermediate result; and process the third intermediate result and the first intermediate result through the feedforward function to obtain a processing result.

[0175] The device also includes an image adjustment module for adjusting the layout of the target object in the initial image based on the layout adjustment information to generate an adjusted initial image; the control module is used to generate control conditions based on the layout adjustment information and the adjusted initial image.

[0176] The control module is used to: generate a second layout feature corresponding to the layout adjustment information, and generate a second image feature corresponding to the initial image; and generate a control condition based on the second layout feature, the second image feature and the noise feature of the current time step.

[0177] The above-mentioned control module is used to: input the layout adjustment information into the first convolutional network and output the intermediate layout features; fuse the object features between the target object and the subordinate objects in the intermediate layout features to obtain fused features; process the fused features with a multi-head attention mechanism to generate global features; input the global features into the second convolutional network and output the second layout features.

[0178] The above-mentioned control module is used to: perform a specified image operation on the initial image so that the target object satisfies the layout parameters for the target object in the layout adjustment information; input the initial image after the specified image operation into the third convolutional network, and output the second image feature.

[0179] The above-mentioned denoising module is used to: input control conditions into the control network and output control features; input noise features and control features into the diffusion model to generate an intermediate image; wherein the intermediate image includes the target object and the image background; based on the target text, render the image text in the intermediate image to generate the final image.

[0180] The above-mentioned device also includes a training module, which is used to: generate the prediction noise of the diffusion model based on the model parameters of the diffusion model, training noise characteristics, training time steps, text characteristics of the training text and training control conditions during the training stage of the diffusion model and the control network; determine the noise distance between the prediction noise and the random noise; determine the loss function value based on the noise distance, and adjust the parameters of the diffusion model and the control network through the loss function value.

[0181] The above-mentioned denoising module is used to: determine the text color based on the image parameters of the background area in the intermediate image; determine the text font based on the object subject, layout adjustment information and text content of the target object; determine the text placement position based on the layout adjustment information; render the target text in the intermediate image based on the text color, text font and text placement position to generate a final image.

[0182] This embodiment also provides an electronic device, including a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the above-mentioned image generation method. The electronic device can be a server or a terminal device.

[0183] See also Figure 7 As shown, the electronic device includes a processor 100 and a memory 101 , wherein the memory 101 stores computer executable instructions that can be executed by the processor 100 , and the processor 100 executes the computer executable instructions to implement the above-mentioned image generating method.

[0184] Further, Figure 7 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .

[0185] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0186] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 100. The above processor 100 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and completes the steps of the method of the above embodiment in combination with its hardware.

[0187] The processor in the above electronic device can implement the following operations in the above image generation method by executing computer executable instructions:

[0188] An image generation method, comprising: obtaining a target text and an initial image; wherein the initial image includes a target object; adjusting a layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout adjustment information includes: layout elements and layout parameters corresponding to the layout elements; the layout elements include the target object and subsidiary objects related to the target object; generating control conditions based on the layout adjustment information and the initial image; performing denoising processing by controlling a diffusion model through the control conditions to generate a final image; wherein the final image includes the target object and subsidiary objects related to the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0189] The default parameters include: a specified attribute, and at least one attribute value preset for the specified attribute; the specified attribute includes: at least one of an object category, a center position, and an object size.

[0190] Generate a first layout feature corresponding to the layout template, a text feature corresponding to the target text, and a first image feature corresponding to the initial image; perform cross-attention processing on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result; and generate layout adjustment information based on the processing result.

[0191] The current time step and the first layout feature are normalized to obtain a first intermediate result; the first intermediate result is self-attention processed to obtain a second intermediate result; the second intermediate result, the text feature and the first image feature are cross-attention processed to obtain a third intermediate result; the third intermediate result and the first intermediate result are processed through a feedforward function to obtain a processing result.

[0192] At the first time step, the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing result of the first time step; at the subsequent time step of the first time step, the first layout feature is updated based on the processing result of the previous time step; the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing results of the subsequent time steps.

[0193] Based on the processing result of the previous time step, the conversion probability of the first layout feature is determined; based on the conversion probability, the first layout feature is converted to obtain an updated first layout feature.

[0194] The layout template is input into the first fully connected layer to generate the first layout feature; the current time step, the first layout feature, the text feature and the first image feature corresponding to the initial image are input into the preset cross-attention processing network; wherein the cross-attention processing network includes: a plurality of conversion modules connected in series; in each conversion module, the target feature data corresponding to the conversion module is determined, the current time step, the target feature data, the text feature and the first image feature are cross-attention processed, and an intermediate result is output; wherein, for the first conversion module, the target feature data is the first layout feature, and for the subsequent conversion modules of the first conversion module, the target feature is the intermediate result output by the previous conversion module; the intermediate result output by the last conversion module is input into the second fully connected layer to generate layout adjustment information.

[0195] The conversion module includes: an adaptive normalization function, a self-attention function, a cross-attention function and a feedforward function; the current time step and the target feature data are normalized by the adaptive normalization function to obtain a first intermediate result; the first intermediate result is subjected to self-attention processing by the self-attention function to obtain a second intermediate result; the second intermediate result, the text feature and the first image feature are subjected to cross-attention processing by the cross-attention function to obtain a third intermediate result; the third intermediate result and the first intermediate result are processed by the feedforward function to obtain a processing result.

[0196] The layout of the target object in the initial image is adjusted based on the layout adjustment information to generate an adjusted initial image; and the control condition is generated based on the layout adjustment information and the adjusted initial image.

[0197] Generate a second layout feature corresponding to the layout adjustment information, and generate a second image feature corresponding to the initial image; and generate a control condition based on the second layout feature, the second image feature and the noise feature of the current time step.

[0198] The layout adjustment information is input into the first convolutional network, and the intermediate layout features are output; the object features between the target object and the subordinate objects in the intermediate layout features are fused to obtain fused features; the fused features are processed by a multi-head attention mechanism to generate global features; the global features are input into the second convolutional network, and the second layout features are output.

[0199] Perform a specified image operation on the initial image so that the target object satisfies the layout parameters for the target object in the layout adjustment information; input the initial image after the specified image operation into the third convolutional network, and output the second image feature.

[0200] The control conditions are input into the control network to output the control features; the noise features and the control features are input into the diffusion model to generate an intermediate image; wherein the intermediate image includes the target object and the image background; based on the target text, the image text is rendered in the intermediate image to generate the final image.

[0201] During the training phase of the diffusion model and the control network, the prediction noise of the diffusion model is generated based on the model parameters of the diffusion model, the training noise characteristics, the training time step, the text characteristics of the training text, and the training control conditions; the noise distance between the prediction noise and the random noise is determined; the loss function value is determined based on the noise distance, and the parameters of the diffusion model and the control network are adjusted through the loss function value.

[0202] Determine the text color based on the image parameters of the background area in the intermediate image; determine the text font based on the object theme of the target object, layout adjustment information and text content of the target text; determine the text placement position based on the layout adjustment information; render the target text in the intermediate image based on the text color, text font and text placement position to generate a final image.

[0203] In this method, the user inputs the target text and the initial image, and the layout adjustment information is generated based on the target text and the initial image. The denoising process of the image is then controlled through the layout adjustment information so that the image layout of the final image matches the data input by the user. The user can accurately control the image layout through the input text and image, which improves the rationality of the image layout and further improves the image quality and visual effect.

[0204] This embodiment also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned image generation method.

[0205] The computer executable instructions stored in the computer readable storage medium can implement the following operations in the image generation method by executing the computer executable instructions:

[0206] An image generation method, comprising: obtaining a target text and an initial image; wherein the initial image includes a target object; adjusting a layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout adjustment information includes: layout elements and layout parameters corresponding to the layout elements; the layout elements include the target object and subsidiary objects related to the target object; generating control conditions based on the layout adjustment information and the initial image; performing denoising processing by controlling a diffusion model through the control conditions to generate a final image; wherein the final image includes the target object and subsidiary objects related to the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

[0207] The default parameters include: a specified attribute, and at least one attribute value preset for the specified attribute; the specified attribute includes: at least one of an object category, a center position, and an object size.

[0208] Generate a first layout feature corresponding to the layout template, a text feature corresponding to the target text, and a first image feature corresponding to the initial image; perform cross-attention processing on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result; and generate layout adjustment information based on the processing result.

[0209] The current time step and the first layout feature are normalized to obtain a first intermediate result; the first intermediate result is self-attention processed to obtain a second intermediate result; the second intermediate result, the text feature and the first image feature are cross-attention processed to obtain a third intermediate result; the third intermediate result and the first intermediate result are processed through a feedforward function to obtain a processing result.

[0210] At the first time step, the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing result of the first time step; at the subsequent time step of the first time step, the first layout feature is updated based on the processing result of the previous time step; the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing results of the subsequent time steps.

[0211] Based on the processing result of the previous time step, the conversion probability of the first layout feature is determined; based on the conversion probability, the first layout feature is converted to obtain an updated first layout feature.

[0212] The layout template is input into the first fully connected layer to generate the first layout feature; the current time step, the first layout feature, the text feature and the first image feature corresponding to the initial image are input into the preset cross-attention processing network; wherein the cross-attention processing network includes: a plurality of conversion modules connected in series; in each conversion module, the target feature data corresponding to the conversion module is determined, the current time step, the target feature data, the text feature and the first image feature are cross-attention processed, and an intermediate result is output; wherein, for the first conversion module, the target feature data is the first layout feature, and for the subsequent conversion modules of the first conversion module, the target feature is the intermediate result output by the previous conversion module; the intermediate result output by the last conversion module is input into the second fully connected layer to generate layout adjustment information.

[0213] The conversion module includes: an adaptive normalization function, a self-attention function, a cross-attention function and a feedforward function; the current time step and the target feature data are normalized by the adaptive normalization function to obtain a first intermediate result; the first intermediate result is subjected to self-attention processing by the self-attention function to obtain a second intermediate result; the second intermediate result, the text feature and the first image feature are subjected to cross-attention processing by the cross-attention function to obtain a third intermediate result; the third intermediate result and the first intermediate result are processed by the feedforward function to obtain a processing result.

[0214] The layout of the target object in the initial image is adjusted based on the layout adjustment information to generate an adjusted initial image; and the control condition is generated based on the layout adjustment information and the adjusted initial image.

[0215] Generate a second layout feature corresponding to the layout adjustment information, and generate a second image feature corresponding to the initial image; and generate a control condition based on the second layout feature, the second image feature and the noise feature of the current time step.

[0216] The layout adjustment information is input into the first convolutional network, and the intermediate layout features are output; the object features between the target object and the subordinate objects in the intermediate layout features are fused to obtain fused features; the fused features are processed by a multi-head attention mechanism to generate global features; the global features are input into the second convolutional network, and the second layout features are output.

[0217] Perform a specified image operation on the initial image so that the target object satisfies the layout parameters for the target object in the layout adjustment information; input the initial image after the specified image operation into the third convolutional network, and output the second image feature.

[0218] The control conditions are input into the control network to output the control features; the noise features and the control features are input into the diffusion model to generate an intermediate image; wherein the intermediate image includes the target object and the image background; based on the target text, the image text is rendered in the intermediate image to generate the final image.

[0219] During the training phase of the diffusion model and the control network, the prediction noise of the diffusion model is generated based on the model parameters of the diffusion model, the training noise characteristics, the training time step, the text characteristics of the training text, and the training control conditions; the noise distance between the prediction noise and the random noise is determined; the loss function value is determined based on the noise distance, and the parameters of the diffusion model and the control network are adjusted through the loss function value.

[0220] Determine the text color based on the image parameters of the background area in the intermediate image; determine the text font based on the object theme of the target object, layout adjustment information and text content of the target text; determine the text placement position based on the layout adjustment information; render the target text in the intermediate image based on the text color, text font and text placement position to generate a final image.

[0221] In this method, the user inputs the target text and the initial image, and the layout adjustment information is generated based on the target text and the initial image. The denoising process of the image is then controlled through the layout adjustment information so that the image layout of the final image matches the data input by the user. The user can accurately control the image layout through the input text and image, which improves the rationality of the image layout and further improves the image quality and visual effect.

[0222] The computer program products of the image generation method, device and electronic device provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments and will not be repeated here.

[0223] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0224] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0225] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0226] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0227] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire a target text and an initial image; wherein the initial image includes a target object; Based on the target text and the initial image, a preset layout template is adjusted to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters of at least some of the layout elements after adjustment; the layout elements include the target object and subsidiary objects related to the target object; generating a control condition based on the layout adjustment information and the initial image; The control condition controls the preset diffusion model to perform denoising processing to generate a final image; wherein the final image includes the target object and the subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

2. The method according to claim 1, characterized in that The default parameters include: a specified attribute, and at least one attribute value preset for the specified attribute; the specified attribute includes: at least one of an object category, a center position, and an object size.

3. The method according to claim 1, characterized in that The step of adjusting a preset layout template based on the target text and the initial image to obtain layout adjustment information includes: Generate a first layout feature corresponding to the layout template, a text feature corresponding to the target text, and a first image feature corresponding to the initial image; Performing cross attention processing on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result; Layout adjustment information is generated based on the processing result.

4. The method according to claim 3, characterized in that The step of performing cross attention processing on the current time step, the first layout feature, the text feature and the first image feature to obtain a processing result comprises: Normalizing the current time step and the first layout feature to obtain a first intermediate result; Performing self-attention processing on the first intermediate result to obtain a second intermediate result; Performing cross attention processing on the second intermediate result, the text feature, and the first image feature to obtain a third intermediate result; The third intermediate result and the first intermediate result are processed by a feedforward function to obtain a processing result.

5. The method according to claim 3, characterized in that: The step of performing cross attention processing on the current time step, the first layout feature, the text feature and the first image feature to obtain a processing result comprises: At a first time step, cross-attention processing is performed on the current time step, the first layout feature, the text feature, and the first image feature to obtain a processing result of the first time step; At a subsequent time step of the first time step, the first layout feature is updated based on the processing result of the previous time step; and the current time step, the first layout feature, the text feature and the first image feature are cross-attention processed to obtain the processing result of the subsequent time step.

6. The method according to claim 5, characterized in that The step of updating the first layout feature based on the processing result of the previous time step includes: Determining a conversion probability of the first layout feature based on a processing result of a previous time step; Based on the conversion probability, the first layout feature is converted to obtain an updated first layout feature.

7. The method according to claim 1, characterized in that The step of adjusting a preset layout template based on the target text and the initial image to obtain layout adjustment information includes: Inputting the layout template into the first fully connected layer to generate a first layout feature; Inputting the current time step, the first layout feature, the text feature and the first image feature corresponding to the initial image into a preset cross-attention processing network; wherein the cross-attention processing network includes: a plurality of serially connected conversion modules; In each of the conversion modules, the target feature data corresponding to the conversion module is determined, the current time step, the target feature data, the text feature and the first image feature are cross-attention processed, and an intermediate result is output; wherein, for the first conversion module, the target feature data is the first layout feature, and for the subsequent conversion modules of the first conversion module, the target feature is the intermediate result output by the previous conversion module; The intermediate result output by the last conversion module is input into the second fully connected layer to generate layout adjustment information.

8. The method according to claim 7, characterized in that The conversion module includes: an adaptive normalization function, a self-attention function, a cross-attention function and a feedforward function; The step of performing cross attention processing on the current time step, the target feature data, the text feature and the first image feature and outputting an intermediate result comprises: The current time step and the target feature data are normalized by the adaptive normalization function to obtain a first intermediate result; Performing self-attention processing on the first intermediate result through the self-attention function to obtain a second intermediate result; Performing cross-attention processing on the second intermediate result, the text feature, and the first image feature through the cross-attention function to obtain a third intermediate result; The third intermediate result and the first intermediate result are processed by the feedforward function to obtain a processing result.

9. The method according to claim 1, characterized in that: After the step of adjusting the preset layout template based on the target text and the initial image to obtain layout adjustment information, the method further includes: adjusting the layout of the target object in the initial image based on the layout adjustment information to generate the adjusted initial image; The step of generating control conditions based on the layout adjustment information and the initial image includes: generating control conditions based on the layout adjustment information and the adjusted initial image.

10. The method according to claim 1, characterized in that The step of generating a control condition based on the layout adjustment information and the initial image comprises: generating a second layout feature corresponding to the layout adjustment information, and generating a second image feature corresponding to the initial image; A control condition is generated based on the second layout feature, the second image feature and the noise feature at the current time step.

11. The method according to claim 10, characterized in that The step of generating a second layout feature corresponding to the layout adjustment information includes: Inputting the layout adjustment information into a first convolutional network and outputting intermediate layout features; Fusing the object features between the target object and the subordinate objects in the intermediate layout features to obtain fused features; Processing the fused features with a multi-head attention mechanism to generate global features; The global features are input into a second convolutional network, and a second layout feature is output.

12. The method according to claim 10, characterized in that The step of generating a second image feature corresponding to the initial image comprises: Performing a specified image operation on the initial image so that the target object satisfies the layout parameters for the target object in the layout adjustment information; The initial image after the specified image operation is input into a third convolutional network, and a second image feature is output.

13. The method according to claim 1, characterized in that The step of controlling the preset diffusion model to perform denoising processing by the control condition to generate a final image includes: Inputting the control conditions into a control network and outputting control characteristics; Inputting the noise feature and the control feature into a preset diffusion model to generate an intermediate image; wherein the intermediate image includes the target object and an image background; Based on the target text, the image text is rendered in the intermediate image to generate a final image.

14. The method according to claim 13, characterized in that The method further comprises: In the training phase of the diffusion model and the control network, based on the model parameters of the diffusion model, the training noise characteristics, the training time step, the text characteristics of the training text and the training control conditions, the prediction noise of the diffusion model is generated; determining a noise distance between the predicted noise and the random noise; A loss function value is determined based on the noise distance, and parameters of the diffusion model and the control network are adjusted according to the loss function value.

15. The method according to claim 13, characterized in that The step of rendering the image text in the intermediate image based on the target text to generate a final image comprises: Determining text color based on image parameters of a background area in the intermediate image; determining a text font based on an object theme of the target object, the layout adjustment information, and text content of the target text; Determining a text placement position based on the layout adjustment information; The target text is rendered in the intermediate image based on the text color, the text font and the text placement position to generate a final image.

16. An image generating device, characterized in that: The device comprises: An acquisition module, used to acquire a target text and an initial image; wherein the initial image includes a target object; A layout module, configured to adjust a preset layout template based on the target text and the initial image to obtain layout adjustment information; wherein the layout template includes: layout elements and default parameters of the layout elements; the layout adjustment information includes: at least some of the layout elements in the layout template, and layout parameters of at least some of the layout elements after adjustment; the layout elements include the target object and subsidiary objects related to the target object; A control module, configured to generate a control condition based on the layout adjustment information and the initial image; A denoising module is used to control a preset diffusion model to perform denoising processing through the control conditions to generate a final image; wherein the final image includes the target object and the subsidiary objects associated with the target object, and the target object and the subsidiary objects are laid out according to the layout parameters.

17. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the image generating method according to any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the image generation method according to any one of claims 1 to 15.

Citation Information

Cited By

  • Video generation method and device for picture story, medium and computer program product

    CN121099083A

  • Advertisement material generation method and device based on multi-modal large model, equipment and storage medium

    CN121564139A

  • Image generation method and apparatus, and electronic device

    WO2026144662A1