An image generation method, apparatus, electronic device, and storage medium

By selecting the optimal generation size of the image generation model and optimizing the prompt words, combined with size transformation operations, the problem of fixed size limitation of the image generation model is solved, and the generation and automated processing of images of diverse sizes are realized.

CN122176079BActive Publication Date: 2026-08-04HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
Filing Date
2026-05-13
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing image generation models can only support a limited number of fixed output sizes, which cannot meet users' needs for diverse image sizes. Existing solutions result in the subject being truncated, compositional imbalance, or resolution loss when images are cropped or scaled, and it is difficult to achieve automated batch processing.

Method used

By obtaining the target size and prompt words, the optimal generation size supported by the image generation model is selected, and the prompt words are optimized. After generating the image of the optimal size, size transformation is performed, including operations such as scaling, semantic cropping, and canvas expansion, to ensure that the image matches the target size.

Benefits of technology

It enables the generation of images of any custom size without modifying the image generation model, with reasonable composition, complete subject, and lossless image quality, making it suitable for large-scale mass production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176079B_ABST
    Figure CN122176079B_ABST
Patent Text Reader

Abstract

This invention provides an image generation method, apparatus, electronic device, and storage medium, applied in the field of image processing. The method comprises: obtaining a target size and a first prompt word; selecting the optimal generation size from multiple image generation sizes supported by an image generation model according to the target size; optimizing the first prompt word to obtain a second prompt word; inputting the optimal generation size and the second prompt word into the image generation model to obtain a first image generated by the image generation model; and performing a size transformation operation on the first image to obtain a second image matching the target size. This method can generate images of various sizes, meeting the diverse needs of users for different image sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] In recent years, AIGC image generation technology based on diffusion models has developed rapidly, and image generation models are now able to generate high-quality image content based on text descriptions.

[0003] Due to limitations in the network architecture design of image generation models and the resolution distribution of training data, image generation models can only support a limited number of fixed output sizes. However, in real-world applications, users have highly diverse needs for image sizes, and image generation models can no longer meet these diverse requirements. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an image generation method, apparatus, electronic device, and storage medium to meet users' diverse needs for various image sizes.

[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0006] The first aspect of this invention discloses an image generation method, the method comprising:

[0007] Get the target size and the first prompt word;

[0008] According to the target size, the optimal generation size is selected from multiple image generation sizes supported by the image generation model;

[0009] The first prompt word is optimized to obtain the second prompt word;

[0010] The optimal generated size and the second prompt word are input into the image generation model to obtain a first image generated by the image generation model.

[0011] The first image is resized to obtain a second image that matches the target size.

[0012] Preferably, according to the target size, the optimal generation size is selected from multiple image generation sizes supported by the image generation model, including:

[0013] Based on the target size and the multiple image generation sizes supported by the image generation model, calculate the comprehensive score for each of the image generation sizes;

[0014] The image generation size with the highest overall score is selected as the optimal generation size.

[0015] Preferably, according to the target size and multiple image generation sizes supported by the image generation model, a comprehensive score is calculated for each of the image generation sizes, including:

[0016] Using the first aspect ratio of the target size and the second aspect ratio of each image generation size supported by the image generation model, calculate the proportional deviation value corresponding to the image generation size;

[0017] Calculate the area adequacy corresponding to the image generation size based on the target size and the image generation size;

[0018] A comprehensive score for the generated image size is calculated based on the area adequacy and the proportional deviation value.

[0019] Preferably, the first prompt word is optimized to obtain the second prompt word, including:

[0020] Obtain target composition guide words that match the first aspect ratio of the target size;

[0021] If it is determined that the user selects an application scenario template, the compositional semantic constraints corresponding to the application scenario template selected by the user are obtained, and the target compositional guide word and the compositional semantic constraints are added to the first prompt word to obtain the second prompt word;

[0022] If it is determined that the user has not selected an application scenario template, the target image guidance word is added to the first prompt word to obtain the second prompt word.

[0023] Preferably, performing a size transformation operation on the first image to obtain a second image matching the target size includes:

[0024] An adaptation strategy is selected using specified data, wherein the specified data includes at least: the target size, the optimal generated size, the pixel area corresponding to the target size, the pixel area corresponding to the optimal generated size, the scale deviation value corresponding to the optimal generated size, and the area sufficiency corresponding to the optimal generated size; the adaptation strategy includes one or more of scaling processing, semantic cropping processing, and canvas expansion processing.

[0025] The first image is resized using the adaptation strategy to obtain a second image that matches the target size.

[0026] Preferably, the optimal generated size and the second prompt word are input into the image generation model to obtain a first image generated by the image generation model, including:

[0027] The optimal generated size and the second prompt word are input into the image generation model to obtain multiple candidate images generated by the image generation model;

[0028] The candidate images are scored to obtain image scores for the candidate images;

[0029] The candidate image with the highest image score is selected as the first image.

[0030] A second aspect of this invention discloses an image generation apparatus, the apparatus comprising:

[0031] The acquisition unit is used to acquire the target size and the first prompt word;

[0032] The selection unit is used to select the optimal generation size from multiple image generation sizes supported by the image generation model according to the target size;

[0033] An optimization unit is used to optimize the first prompt word to obtain a second prompt word;

[0034] A generation unit is used to input the optimal generation size and the second prompt word into the image generation model to obtain a first image generated by the image generation model;

[0035] The transformation unit is used to perform a size transformation operation on the first image to obtain a second image that matches the target size.

[0036] Preferably, the selection unit includes:

[0037] The calculation module is used to calculate the comprehensive score for each of the image generation sizes according to the target size and the multiple image generation sizes supported by the image generation model.

[0038] The selection module is used to select the image generation size with the highest overall score as the optimal generation size.

[0039] A third aspect of the present invention discloses an electronic device, comprising: a processor and a memory, the processor and the memory being connected via a bus; wherein the processor is configured to call and execute a program stored in the memory; the memory is configured to store the program, the program being configured to implement the image generation method disclosed in the first aspect of the present invention.

[0040] A fourth aspect of the present invention discloses a storage medium storing computer-executable instructions for executing the image generation method disclosed in the first aspect of the present invention.

[0041] Based on the above embodiments of the present invention, an image generation method, apparatus, electronic device, and storage medium are provided. The method comprises: obtaining a target size and a first prompt word; selecting the optimal generation size from multiple image generation sizes supported by an image generation model according to the target size; optimizing the first prompt word to obtain a second prompt word; inputting the optimal generation size and the second prompt word into the image generation model to obtain a first image generated by the image generation model; and performing a size transformation operation on the first image to obtain a second image matching the target size. In this scheme, the optimal generation size is selected from multiple image generation sizes supported by the image generation model according to the target size given by the user, and the first prompt word given by the user is optimized to obtain the second prompt word. The optimal generation size and the second prompt word are input into the image generation model to generate the first image, and the first image is subjected to a size transformation operation to obtain a second image matching the target size. This method can generate images of various sizes to meet the diverse needs of users for various image sizes. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0043] Figure 1 A flowchart of an image generation method provided in an embodiment of the present invention;

[0044] Figure 2 A flowchart for selecting the optimal generation size provided in an embodiment of the present invention;

[0045] Figure 3 A flowchart for obtaining a second image provided in an embodiment of the present invention;

[0046] Figure 4 An example diagram illustrating the execution order and interrelationships of functional modules provided in embodiments of the present invention;

[0047] Figure 5 This is an overall flowchart of an image generation method provided in an embodiment of the present invention;

[0048] Figure 6 This is a structural block diagram of an image generation device provided in an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0051] In recent years, AIGC image generation technology based on diffusion models has developed rapidly. The image generation model can now generate high-quality image content based on text descriptions and has been widely used in creative design, e-commerce, social media operation, advertising production and other fields.

[0052] However, existing image generation models have significant limitations in terms of output size. Specifically, due to the limitations of the network architecture design of the image generation model and the resolution distribution of the training data, the image generation model can only support a limited number of fixed output sizes, and some image generation models also require the width and height of the output image to be integers of 64 or 8.

[0053] Research has revealed that, in real-world applications, user requirements for image sizes are highly diverse. For example, in e-commerce, software A requires product images to be 800*800 pixels, software B requires 350*350 pixels, and software C requires 750*352 pixels. In social media, software D requires cover images to be 900*383 pixels, and software E recommends 1080*1440 pixels for note images. The size specifications in advertising are even more complex, with dozens of size specifications available for in-feed ads and splash screen ads on different media platforms.

[0054] Faced with the contradiction between the fixed output size of image generation models and the diverse practical needs, the common practice in the industry is "generate first, then crop and scale," that is, first generate images at the size supported by the image generation model, and then adapt them to the target size through simple cropping or scaling operations. However, this approach has the following problems:

[0055] 1) Direct cropping may result in the main subject being cut off or the composition becoming unbalanced;

[0056] 2) Forced scaling can cause image distortion or resolution loss;

[0057] 3) The above operations usually require manual intervention to determine the cropping position and scaling ratio, making it difficult to achieve automated batch processing.

[0058] To address this, embodiments of the present invention propose an image generation method, apparatus, electronic device, and storage medium. The method selects the optimal generation size from multiple image generation sizes supported by an image generation model based on a user-provided target size, and optimizes a first prompt word provided by the user to obtain a second prompt word. The optimal generation size and the second prompt word are input into the image generation model to generate a first image, and the first image is then resized to obtain a second image matching the target size. This method can generate images of various sizes, satisfying diverse user needs for different image sizes.

[0059] After applying this solution, users can input any custom target size (any aspect ratio and any resolution) to obtain images with reasonable composition, complete subject, and lossless image quality without modifying the image generation model (any AIGC image generation model).

[0060] See Figure 1 The diagram illustrates a flowchart of an image generation method provided by an embodiment of the present invention, the image generation method comprising:

[0061] Step S101: Obtain the target size and the first prompt word.

[0062] In the specific implementation step S101, the target size (W_t, H_t) and the first prompt word input by the user are received, where W_t is the width of the target size and H_t is the height of the target size.

[0063] Step S102: Select the optimal generation size from multiple image generation sizes supported by the image generation model according to the target size.

[0064] It should be noted that the image generation model can be any AIGC model with image generation capabilities. This image generation model supports a list of generation sizes L, L={(W_1,H_1),(W_2,H_2),..., (W_n,H_n)}, where (W_1,H_1) to (W_n,H_n) are different image generation sizes. The image generation sizes in the list of generation sizes supported by the image generation model can be called candidate sizes.

[0065] In the specific implementation step S102, according to the target size, a certain image generation size is selected from the image generation model as the optimal generation size (W_g, H_g). That is, the optimal generation size (W_g, H_g) is any image generation size in the list of generation sizes supported by the image generation model.

[0066] Step S103: Optimize the first prompt word to obtain the second prompt word.

[0067] It should be noted that the first aspect ratio of the target dimensions (W_t, H_t) is denoted as R_t, and it is calculated as R_t = W_t / H_t. A built-in proportion-composition mapping knowledge base maps common aspect ratio ranges to corresponding composition guidance words.

[0068] In the specific implementation step S103, the first prompt word is optimized based on the first aspect ratio and the "proportion-composition mapping knowledge base" to obtain the second prompt word.

[0069] Specifically, based on the "proportion-composition mapping knowledge base", the target composition guide word that matches the first aspect ratio of the target size is obtained. The target composition guide word is the composition guide word corresponding to the first aspect ratio.

[0070] For example: when R_t is greater than 2.0 (extremely wide banner), the panoramic composition guide word is determined as the target composition guide word. When R_t is less than 0.6 (narrow vertical banner), the vertical composition guide word is determined as the target composition guide word.

[0071] If it is determined that the user selects an application scenario template, the compositional semantic constraints corresponding to the selected application scenario template are obtained. The target compositional guide word and the compositional semantic constraints are added to the first prompt word to obtain the second prompt word. In this way, the enhanced prompt word can guide the image generation model to generate an image that is more suitable for the target size in terms of spatial layout (i.e., the first image mentioned later) while retaining the user's original semantic intent.

[0072] If it is determined that the user has not selected an application scenario template, add the target composition guide word to the first prompt word to obtain the second prompt word.

[0073] Step S104: Input the optimal generation size and the second prompt word into the image generation model to obtain the first image generated by the image generation model.

[0074] In the specific implementation step S104, the optimal generation size (W_g, H_g) and the second prompt word are input into the image generation model to obtain multiple candidate images generated by the image generation model.

[0075] Specifically, the optimal generated size (W_g, H_g) and the second prompt word are used as input parameters to call the API interface of the image generation model to generate multiple candidate images.

[0076] The candidate images are scored to obtain image scores. Specifically, each candidate image is scored using an aesthetic scoring model and a composition scoring model to obtain an image score for each candidate image.

[0077] Multiple candidate images are sorted according to their image scores, and the candidate image with the highest image score is selected as the first image.

[0078] For example, using the optimal generation size (W_g, H_g) and the second prompt word as input parameters, the API interface of the image generation model is called to generate 2 to 4 candidate images. Each candidate image is scored by the aesthetic scoring model and the composition scoring model, and the candidate image with the highest score is selected as the first image.

[0079] In practical applications, the optimal generation size and the second prompt word can also be input into the image generation model, and the image generation model will generate only one image, which is the first image.

[0080] It should be noted that the size of the first image obtained through the above method is the optimal generated size.

[0081] Step S105: Perform a size transformation operation on the first image to obtain a second image that matches the target size.

[0082] In the specific implementation step S105, a size transformation operation is performed on the first image with the optimal generated size to obtain a second image with the target size.

[0083] In this embodiment of the invention, the optimal image generation size is selected from multiple image generation sizes supported by the image generation model according to the target size given by the user. The first prompt word given by the user is then optimized to obtain a second prompt word. The optimal generation size and the second prompt word are input into the image generation model to generate a first image. The first image is then resized to obtain a second image that matches the target size. This method can generate images of various sizes, satisfying users' diverse needs for different image sizes.

[0084] Regarding the above embodiments of the present invention Figure 1 For the optimal generation size involved in step S102, see [link to documentation]. Figure 2 The flowchart illustrating the selection of the optimal generation size provided by an embodiment of the present invention includes the following steps:

[0085] Step S201: Calculate the comprehensive score for each image generation size according to the target size and the multiple image generation sizes supported by the image generation model.

[0086] It should be noted that (W_i, H_i) is the i-th image generation size supported by the image generation model, where i is greater than or equal to 1 and less than or equal to n.

[0087] In the specific implementation step S201, the proportional deviation value corresponding to the image generation size is calculated by using the first aspect ratio of the target size and the second aspect ratio of each image generation size supported by the image generation model.

[0088] Specifically, the first aspect ratio of the target size (W_t, H_t) is calculated as R_t = W_t / H_t, and the second aspect ratio of the i-th generated image size (W_i, H_i) is calculated as R_i = W_i / H_i. Using the first aspect ratio R_t of the target size and the second aspect ratio R_i of the i-th generated image size (W_i, H_i), the scaling deviation value D_i = |R_i - R_t| corresponding to the i-th generated image size (W_i, H_i) is calculated.

[0089] By calculating the scaling deviation value corresponding to each image generation size supported by the image generation model using the above method, we can obtain D_1, D_2, ..., D_n. Let D_max be the maximum value among the scaling deviation values ​​corresponding to each image generation size, i.e., D_max = max(D_1, D_2, ..., D_n).

[0090] Calculate the area adequacy corresponding to the image generation size based on the target size and the image generation size. Specifically, calculate the area adequacy A_i corresponding to the i-th image generation size (W_i, H_i) based on the target size (W_t, H_t) and the i-th image generation size (W_i, H_i).

[0091] Here, the area adequacy A_i corresponding to the i-th image generation size (W_i, H_i) is the ratio between the area of ​​the i-th image generation size (W_i, H_i) and the area of ​​the target size (W_t, H_t), that is, A_i = (W_i × H_i) / (W_t × H_t). The area adequacy corresponding to each image generation size supported by the image generation model can be calculated in this way, i.e., A_1, A_2…A_n can be calculated.

[0092] The overall score for the generated image size is calculated based on the area adequacy and scale deviation values ​​corresponding to the generated image size.

[0093] Specifically, for the i-th image generation size (W_i, H_i), the comprehensive score Score_i of the i-th image generation size is calculated as follows: Score_i = α × (1 - D_i / D_max) + β × min(A_i, 1).

[0094] Where α and β are configurable weighting coefficients.

[0095] Using the methods described above, the comprehensive score for each image generation size supported by the image generation model can be calculated.

[0096] Step S202: Select the image generation size with the highest comprehensive score as the optimal generation size.

[0097] In the specific implementation step S202, the image generation size with the highest comprehensive score is selected as the optimal generation size (W_g, H_g).

[0098] The above embodiments of the present invention Figure 2 This is an explanation of how to select the optimal generation size.

[0099] Regarding the above embodiments of the present invention Figure 1 The second image involved in step S105, see [link / reference]. Figure 3 The flowchart illustrating the process of obtaining a second image according to an embodiment of the present invention includes the following steps:

[0100] Step S301: Select an adaptation strategy using specified data.

[0101] In the specific implementation step S301, the adaptation strategy is selected by specifying data.

[0102] The specified data includes at least: target size, optimal generated size, pixel area (W_t×H_t) corresponding to the target size, pixel area (W_g×H_g) corresponding to the optimal generated size, scale deviation value corresponding to the optimal generated size, and area sufficiency corresponding to the optimal generated size.

[0103] The selected adaptation strategy includes one or more of scaling, semantic cropping, and canvas expansion.

[0104] It should be noted that the proportional deviation value corresponding to the optimal generated size (W_g, H_g) is denoted as D_g. The calculation method of D_g can be found in the above embodiment of the present invention. Figure 2 The details regarding D_i are omitted here. The area adequacy corresponding to the optimal generated size (W_g, H_g) is denoted as A_g. The calculation method for A_g can be found in the above-described embodiments of the invention. Figure 2 The details regarding A_i will not be elaborated upon here.

[0105] In some embodiments, the way to select an adaptation strategy is as follows: If the "proportional deviation value D_g corresponding to the optimal generated size" is less than the threshold e1 and the "area abundance A_g corresponding to the optimal generated size" is close to 1, the selected adaptation strategy is scaling processing (equivalent to a pure scaling path), where "close to 1" means the difference from 1 is within a preset range.

[0106] The pixel area (W_g × H_g) corresponding to the optimal generated size is the generated area, and the pixel area (W_t × H_t) corresponding to the target size is the target area. If the generated area is greater than the target area and the proportions are different, the selected adaptation strategy is semantic cropping processing (equivalent to a semantic cropping path).

[0107] If the generated area is not sufficient to cover the target area, the selected adaptation strategy is canvas expansion processing (equivalent to a canvas expansion path).

[0108] It should be noted that "the generated area is not sufficient to cover the target area" specifically means that: in the width and height dimensions of the optimal generated size, at least one dimension is smaller than the target size, which will result in being unable to obtain the target size by cropping based on the optimal generated size.

[0109] Furthermore, if W_g < W_t or H_g < H_t, it can be determined that the generated area is not sufficient to cover the target area. In practical applications, if A_g < 1 and the proportions are inconsistent, it can also be determined that the generated area is not sufficient to cover the target area.

[0110] If the judgment conditions for any one of the above-mentioned scaling processing, semantic cropping processing, and canvas expansion processing are not met, the selected adaptation strategy includes scaling processing, semantic cropping processing, and canvas expansion processing, where semantic cropping processing, scaling processing, and canvas expansion processing can be performed sequentially, or canvas expansion processing, semantic cropping processing, and scaling processing can be performed in series.

[0111] Step S302: Perform a size transformation operation on the first image through the adaptation strategy to obtain a second image that matches the target size.

[0112] In the process of specifically implementing step S302, perform a size transformation operation on the first image through the selected adaptation strategy, so as to obtain a second image that precisely matches the target size.

[0113] It should be noted that the selected adaptation strategy includes one or more of scaling processing, semantic cropping processing, and canvas expansion processing. When executing the adaptation strategy, functional modules such as "subject detection and semantic segmentation", "semantic-aware cropping", "canvas expansion", "super-resolution scaling", and "edge harmonization processing" will be called. The following will first explain these functional modules separately.

[0114] Subject detection and semantic segmentation: Perform object detection and semantic segmentation on the first image to obtain the bounding box coordinates and pixel-level separation results of the foreground / background of the main object in the first image. The obtained "bounding box coordinates and pixel-level separation results of foreground / background" are used in subsequent semantic-aware cropping and canvas expansion for subject protection.

[0115] In other words, object detection is performed on the first image to obtain the bounding box coordinates of the main object. These bounding box coordinates can represent the position and occupied area of ​​the main object. Semantic segmentation is performed on the first image to obtain pixel-level separation results of the foreground / background. The pixel-level separation results of the foreground / background are used to represent whether each pixel belongs to the main object or the background.

[0116] Semantic-aware cropping: When the adaptation strategy includes semantic cropping, the cropping box is designed to ensure it does not cut into the main subject area based on the bounding box coordinates; only background or unimportant content is cropped. Specifically, the size of the cropping box is calculated using the center point of the main subject's bounding box as the anchor point and the first aspect ratio R_t. The cropping box must satisfy the subject integrity constraint, meaning it must completely contain the subject's bounding box while maintaining a certain safety margin. If the subject is too large to be completely contained while maintaining the target aspect ratio, the cropping area is appropriately enlarged and processed in subsequent scaling steps.

[0117] Canvas Expansion: When the adaptation strategy includes canvas expansion processing, the position and orientation of the subject are determined based on the pixel-level separation results of the foreground / background, deciding in which direction to expand the canvas to avoid the subject shifting to the edge of the image after expansion. Simultaneously, the expanded filling content also references the characteristics of the background area to maintain consistency. Specifically, an Outpainting operation is performed on the first image in the direction requiring extension. Using the original image content as a constraint, the image expansion interface (or Inpainting interface) of the image generation model is called to generate the extended area while maintaining consistency with the style of the original image content. The expansion operation employs an overlapping block strategy, preserving a certain width of overlap between adjacent blocks, and eliminating seam traces within the overlapping area through weighted fusion.

[0118] Super-resolution scaling: Adjust to the accurate target resolution W_t×H_t using a high-quality scaling algorithm (such as bicubic interpolation or a deep learning-based super-resolution model).

[0119] Edge harmonization processing: Global harmonization post-processing is performed on the boundary traces that may be generated by the operation of the above functional modules, including tone unification, brightness balance and texture smoothing, to ensure that the final output image is visually natural and coherent.

[0120] The above is an explanation of the functional modules such as "Subject Detection and Semantic Segmentation", "Semantic Aware Cropping", "Canvas Expansion", "Super-Resolution Scaling", and "Edge Harmonization".

[0121] The execution order and interrelationships of functional modules such as "Subject Detection and Semantic Segmentation", "Semantic Aware Cropping", "Canvas Expansion", "Super-Resolution Scaling", and "Edge Harmonization" are as follows: Figure 4 As shown, after obtaining the first image and selecting the adaptation strategy, the "subject detection and semantic segmentation" function module is executed. Then, according to the adaptation strategy, the "semantic-aware cropping", "canvas expansion", "super-resolution scaling" and "edge harmonization" function modules are selected and executed.

[0122] The scaling process requires calling the "Super-resolution Scaling" and "Edge Harmonization" modules. Since the aspect ratio is already matched, only the resolution needs to be adjusted, and the image quality can be guaranteed by scaling and then performing edge harmonization.

[0123] Semantic cropping primarily involves calling the "Subject Detection and Semantic Segmentation" and "Semantic-Aware Cropping" modules. First, subject detection identifies key content areas in the image. Then, based on semantic information, the position of the cropping box is determined to ensure that unimportant parts are cropped. If further size adjustments are needed after cropping, the "Super-Resolution Scaling" and "Edge Harmonization" modules may be called.

[0124] Canvas expansion primarily requires calling the "Canvas Expansion" and "Edge Harmonization" modules. Outpainting adds pixels outwards, and edge harmonization eliminates seams and color differences between the expanded and original areas. Similarly, before expansion, the "Subject Detection and Semantic Segmentation" module may be called to determine the most appropriate expansion direction (avoiding pushing the subject to the edge). The expanded image may also require scaling using the "Super-Resolution Scaling" module.

[0125] Understandably, the "super-resolution scaling" module receives an image that needs to be scaled as input.

[0126] During scaling, the "super-resolution scaling" module receives the "first image" as input and adjusts the first image of size "W_g×H_g" to the second image of size "W_t×H_t" through the "super-resolution scaling" module.

[0127] When performing semantic cropping, the "super-resolution scaling" module receives the "cropped image" as input.

[0128] When performing canvas expansion processing, the "Super-resolution Scaling" module receives the "expanded image" as input.

[0129] When performing scaling, semantic cropping, and canvas expansion in combination, the "super-resolution scaling" module receives the "combined image" as input.

[0130] The above embodiments of the present invention Figure 3 This is an explanation of how to obtain the second image.

[0131] In some embodiments, after obtaining a second image that matches the target size, the quality of the second image is evaluated to obtain evaluation metrics for the second image, which include, but are not limited to, subject integrity, compositional aesthetics score, edge coherence, etc.

[0132] When the evaluation metric of the second image is lower than the quality threshold, the parameters are automatically adjusted (such as adjusting the cropping anchor point offset, increasing the Outpainting overlap width, etc.), and step S103 or step S301 is re-executed until the evaluation metric reaches the quality threshold or the maximum number of re-executions is reached.

[0133] To better understand this solution, through Figure 5 The following is an example of an overall flowchart of an image generation method, including the following steps:

[0134] Step S501: Analyze the target size and calculate the first aspect ratio R_t.

[0135] Step S502: Iterate through the image generation sizes supported by the image generation model and calculate the proportional deviation value of each image generation size.

[0136] Step S503: Select the optimal generation size.

[0137] Step S504: Select an adaptation strategy.

[0138] Step S505: Call the image generation model to generate a first image, and perform a size transformation operation on the first image through an adaptation strategy to obtain a second image that matches the target size.

[0139] It should be noted that the execution principle of steps S501 to S505 can be found in the contents of the above embodiments, and will not be repeated here.

[0140] To further demonstrate the practical application of this solution, examples will be given below from three backgrounds: multi-platform adaptation of e-commerce product main images, multi-scenario content production on social media, and precise size output of printed materials.

[0141] E-commerce product main image multi-platform adaptation background: An e-commerce operations team needs to generate main images for the same product that are adapted to three platforms: Software A (target size 800×800), Software B (target size 350x350), and Software C (target size 750×352). The first suggestion for users entering the product description is "white minimalist ceramic vase, desktop placement, light gray background, commercial photography style," and they specify three target sizes in sequence.

[0142] For a target size of 800×800 (R_t=1.0), the optimal generation size is found to be 1024×1024 (R_i=1.0) among the image generation sizes supported by the image generation model. The proportions are perfectly matched, and the chosen adaptation strategy is scaling. A symmetrical composition guide word is appended to the first prompt word at a 1:1 ratio to obtain the second prompt word. After generating the first image at 1024×1024, the image generation model performs high-quality downsampling and scaling to 800×800 to obtain the second image.

[0143] For the target size of 350×350 (R_t=1.0), the processing logic is the same as the relevant content of the "target size of 800×800" described above. After generating the first image at 1024×1024, the first image is downsampled and scaled to 350×3500 with high quality. During scaling, a downsampling algorithm with sharpening is used to maintain the clarity of details, thereby obtaining the second image.

[0144] For the target size of 750×352 (R_t=2.13), the proportion of this target size deviates significantly from the image generation size supported by the image generation model. The closest image generation size is 1024×768 (R_i=1.33). The selected adaptation strategies include scaling, semantic cropping, and canvas expansion. A wide horizontal composition guide word is appended to the first cue word to obtain the second cue word. After the image generation model generates the first image at 1024×768, it first detects the position of the vase subject, then calculates the cropping box with the subject as the center at R_t=2.13. Due to insufficient horizontal width, outpainting expansion is performed on the left and right sides to generate background extension content consistent with the style of the original image. Then, it is cropped and scaled to 750×352 to obtain the second image.

[0145] Social media content production context: A self-media creator needs to generate a WeChat official account cover image (target size 900×383, R_t=2.35) and a software-generated image (target size 1080×1440, R_t=0.75) for the same article. The first suggested phrase entered by the user is "a cafe terrace under cherry blossom trees in spring, warm light, Japanese aesthetics".

[0146] For the target size of 900×383, 1024×768 was chosen as the optimal generation size. The selected adaptation strategies included scaling, semantic cropping, and canvas expansion. An ultra-wide banner-style guiding text was added to the first prompt. After the image generation model generated the first image, the subject was detected and located in the café terrace area. The image was then cropped vertically to a size close to the target. For any insufficient areas on either side, outpainting was used to extend the cherry blossom trees and sky. Finally, the image was scaled to 900×383 to obtain the second image.

[0147] For the target size of 1080×1440, 768×1024 was chosen as the optimal generation size, with a perfect aspect ratio. The selected adaptation strategy included scaling. A vertical composition guide word was added to the first prompt. After the image generation model generated the first image of 768×1024, it was enlarged to 1080×1440 using a super-resolution model to obtain the second image.

[0148] Precise Dimension Output Background for Printed Materials: A design company needs to create a roll-up banner background for an exhibition. The physical dimensions are 80cm × 200cm, and the printing resolution requirement is 150 DPI, corresponding to a pixel size of 4724 × 11811 (R_t = 0.40).

[0149] The physical size plus DPI is converted to pixel size. The target image is identified as having an extremely narrow vertical format and extremely high resolution. 768×1024 is selected as the optimal generation size. The chosen adaptation strategies include scaling, semantic cropping, and canvas expansion. A guiding phrase for the extremely narrow vertical format is added to the first prompt.

[0150] After generating a first image of 768×1024 using the image generation model, outpainting is first performed in multiple steps in the vertical direction to approximate a scale of R_t=0.40. The expansion amount in each step is controlled within 30% of the original image size to ensure stylistic consistency. Semantic consistency constraints are applied during the expansion process. After expansion, the image is cropped according to the target scale. Finally, a super-resolution model is used to progressively enlarge the image to the target resolution of 4724×11811, thus obtaining the second image. Throughout the process, edge coherence is checked after each expansion step; if a stylistic break occurs, the expansion step is automatically retried.

[0151] The above are examples illustrating the practical application scenarios of this solution. In summary, this solution offers the following beneficial effects:

[0152] Model independence: The entire solution treats the underlying AIGC model as a black-box API call, which does not depend on the internal architecture of any specific model. Therefore, it can be seamlessly compatible with various current and future image generation models. When the underlying model is upgraded, the system can immediately benefit without modification.

[0153] Arbitrary size support: Users can input any pixel-level target size, without being limited by the model's natively supported size list, fundamentally solving the contradiction between fixed-size output and diverse application needs.

[0154] Adaptive composition: Through the synergy of size strategy planning and cue word composition enhancement, the generated image composition naturally fits the target size.

[0155] Subject integrity protection: The semantically aware cropping mechanism can automatically identify and protect the subject of the image, ensuring that the subject is not truncated or obscured during any size transformation.

[0156] Automation and mass production capability: The entire production line operates automatically without the need for manual intervention in determining the cropping position or scaling ratio, making it suitable for large-scale, batch content production scenarios.

[0157] Corresponding to the image generation method provided in the above embodiments of the present invention, see also... Figure 6 The present invention also provides a structural block diagram of an image generation device, which includes: an acquisition unit 601, a selection unit 602, an optimization unit 603, a generation unit 604, and a transformation unit 605.

[0158] The acquisition unit 601 is used to acquire the target size and the first prompt word.

[0159] Selection unit 602 is used to select the optimal generation size from multiple image generation sizes supported by the image generation model according to the target size.

[0160] The optimization unit 603 is used to optimize the first prompt word to obtain the second prompt word.

[0161] In some embodiments, the optimization unit 603 is specifically used to: obtain a target composition guide word that matches a first aspect ratio of the target size; if it is determined that the user selects an application scenario template, obtain the composition semantic constraint corresponding to the application scenario template selected by the user, and add the target composition guide word and the composition semantic constraint to the first prompt word to obtain a second prompt word; if it is determined that the user does not select an application scenario template, add the target composition guide word to the first prompt word to obtain a second prompt word.

[0162] The generation unit 604 is used to input the optimal generation size and the second prompt word into the image generation model to obtain the first image generated by the image generation model.

[0163] In some embodiments, the generation unit 604 is specifically used to: input the optimal generation size and the second prompt word into the image generation model to obtain multiple candidate images generated by the image generation model; score the candidate images to obtain image scores for the candidate images; and select the candidate image with the highest image score as the first image.

[0164] The transformation unit 605 is used to perform a size transformation operation on the first image to obtain a second image that matches the target size.

[0165] In this embodiment of the invention, the optimal image generation size is selected from multiple image generation sizes supported by the image generation model according to the target size given by the user. The first prompt word given by the user is then optimized to obtain a second prompt word. The optimal generation size and the second prompt word are input into the image generation model to generate a first image. The first image is then resized to obtain a second image that matches the target size. This method can generate images of various sizes, satisfying users' diverse needs for different image sizes.

[0166] Preferred, combined Figure 6 The selection unit 602, as shown, includes a calculation module and a selection module. The execution principle of each module is as follows:

[0167] The calculation module is used to calculate the comprehensive score for each image generation size according to the target size and multiple image generation sizes supported by the image generation model.

[0168] In specific implementation, the calculation module is used to: calculate the proportional deviation value corresponding to the image generation size using the first aspect ratio of the target size and the second aspect ratio of each image generation size supported by the image generation model; calculate the area adequacy corresponding to the image generation size according to the target size and the image generation size; and calculate the comprehensive score of the image generation size based on the area adequacy and the proportional deviation value.

[0169] The selection module is used to select the image generation size with the highest overall score as the optimal generation size.

[0170] Preferred, combined Figure 6 The transformation unit 605, as shown, includes a selection module and a transformation module. The execution principle of each module is as follows:

[0171] The selection module is used to select an adaptation strategy using specified data, wherein the specified data includes at least: target size, optimal generated size, pixel area corresponding to the target size, pixel area corresponding to the optimal generated size, scale deviation value corresponding to the optimal generated size, and area adequacy corresponding to the optimal generated size; the adaptation strategy includes one or more of scaling processing, semantic cropping processing, and canvas expansion processing.

[0172] The transformation module is used to perform a size transformation operation on the first image through an adaptation strategy to obtain a second image that matches the target size.

[0173] Preferably, the present invention also provides an electronic device, including: a processor and a memory, the processor and the memory being connected via a bus; wherein, the processor is used to call and execute a program stored in the memory; the memory is used to store the program, the program being used to implement the image generation method provided in the above method embodiments.

[0174] Preferably, the present invention also provides a storage medium storing computer-executable instructions for executing the image generation method provided in the above-described method embodiments.

[0175] In summary, embodiments of the present invention provide an image generation method, apparatus, electronic device, and storage medium. The method selects the optimal generation size from multiple image generation sizes supported by an image generation model according to a target size provided by the user, and optimizes a first prompt word provided by the user to obtain a second prompt word. The optimal generation size and the second prompt word are input into the image generation model to generate a first image, and the first image is then resized to obtain a second image matching the target size. This method can generate images of various sizes, satisfying diverse user needs for different image sizes.

[0176] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0177] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0178] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image generation method characterized by, The method includes: Get the target size and the first prompt word; According to the target size, the optimal generation size is selected from multiple image generation sizes supported by the image generation model; The first prompt word is optimized to obtain the second prompt word; The optimal generated size and the second prompt word are input into the image generation model to obtain a first image generated by the image generation model. The first image is resized to obtain a second image that matches the target size. Performing a size transformation operation on the first image to obtain a second image that matches the target size includes: An adaptation strategy is selected using specified data, wherein the specified data includes at least: the target size, the optimal generated size, the pixel area corresponding to the target size, the pixel area corresponding to the optimal generated size, the ratio deviation value corresponding to the optimal generated size, and the area sufficiency corresponding to the optimal generated size; the adaptation strategy includes one or more of scaling processing, semantic cropping processing, and canvas expansion processing; wherein the ratio deviation value is calculated by the first aspect ratio of the target size and the second aspect ratio of each image generated size supported by the image generation model, and the area sufficiency is calculated by the target size and the image generated size; The first image is resized using the adaptation strategy to obtain a second image that matches the target size.

2. The method of claim 1, wherein, Based on the target size, the optimal generation size is selected from multiple image generation sizes supported by the image generation model, including: Based on the target size and the multiple image generation sizes supported by the image generation model, calculate the comprehensive score for each of the image generation sizes; The image generation size with the highest overall score is selected as the optimal generation size.

3. The method of claim 2, wherein, Based on the target size and the multiple image generation sizes supported by the image generation model, a comprehensive score is calculated for each of the image generation sizes, including: Using the first aspect ratio of the target size and the second aspect ratio of each image generation size supported by the image generation model, calculate the proportional deviation value corresponding to the image generation size; Calculate the area adequacy corresponding to the image generation size based on the target size and the image generation size; A comprehensive score for the generated image size is calculated based on the area adequacy and the proportional deviation value.

4. The method of claim 1, wherein, The first prompt word is optimized to obtain the second prompt word, including: Obtain target composition guide words that match the first aspect ratio of the target size; If it is determined that the user selects an application scenario template, the compositional semantic constraints corresponding to the application scenario template selected by the user are obtained, and the target compositional guide word and the compositional semantic constraints are added to the first prompt word to obtain the second prompt word; If it is determined that the user has not selected an application scenario template, the target image guidance word is added to the first prompt word to obtain the second prompt word.

5. The method of claim 1, wherein, The optimal generated size and the second prompt word are input into the image generation model to obtain a first image generated by the image generation model, including: The optimal generated size and the second prompt word are input into the image generation model to obtain multiple candidate images generated by the image generation model; The candidate images are scored to obtain image scores for the candidate images; The candidate image with the highest image score is selected as the first image.

6. An image generation apparatus, characterized in that, The device includes: The acquisition unit is used to acquire the target size and the first prompt word; The selection unit is used to select the optimal generation size from multiple image generation sizes supported by the image generation model according to the target size; An optimization unit is used to optimize the first prompt word to obtain a second prompt word; A generation unit is used to input the optimal generation size and the second prompt word into the image generation model to obtain a first image generated by the image generation model; A transformation unit is used to perform a size transformation operation on the first image to obtain a second image that matches the target size; The transformation unit includes: The selection module is used to select an adaptation strategy using specified data, wherein the specified data includes at least: the target size, the optimal generated size, the pixel area corresponding to the target size, the pixel area corresponding to the optimal generated size, the ratio deviation value corresponding to the optimal generated size, and the area adequacy value corresponding to the optimal generated size; the adaptation strategy includes one or more of scaling processing, semantic cropping processing, and canvas expansion processing; wherein the ratio deviation value is calculated by the first aspect ratio of the target size and the second aspect ratio of each image generated size supported by the image generation model, and the area adequacy value is calculated by the target size and the image generated size; The transformation module is used to perform a size transformation operation on the first image through the adaptation strategy to obtain a second image that matches the target size.

7. The apparatus according to claim 6, characterized in that, The selection unit includes: The calculation module is used to calculate the comprehensive score of each of the image generation sizes according to the target size and the multiple image generation sizes supported by the image generation model. The selection module is used to select the image generation size with the highest overall score as the optimal generation size.

8. An electronic device, characterized in that, include: A processor and a memory are connected via a bus; wherein the processor is used to call and execute a program stored in the memory; The memory is used to store a program for implementing the image generation method as described in any one of claims 1-5.

9. A storage medium, characterized in that, The storage medium stores computer-executable instructions for performing the image generation method as described in any one of claims 1-5.