Image generation method and device, electronic equipment and storage medium

By determining the target ratio in AIGC technology and performing padding, cropping, and scaling on the initial image, the distortion and misalignment problems during the generation of irregularly shaped images are solved, thus improving the image generation quality.

CN122115628APending Publication Date: 2026-05-29BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing AIGC technology, due to the probabilistic generation characteristics of the model, is prone to image distortion and element misalignment when expanding the size of irregularly shaped images, resulting in poor generation quality.

Method used

By determining the initial image and prompt words at the target ratio, the initial image is filled based on the standard ratio to generate the first image at the standard ratio. Then, the image is generated using an image generation model, followed by cropping and scaling to obtain the target image at the target ratio.

Benefits of technology

It effectively avoids the problems of image distortion and element misalignment that occur during the generation of irregularly shaped images, thus improving the quality and presentation of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115628A_ABST
    Figure CN122115628A_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image generation method, device, electronic equipment and storage medium. The method comprises: determining an initial image of a target scale and a prompt word according to image generation requirement data, the prompt word being used to describe picture content requirements of the image; in the case that the target scale meets a preset scale requirement, determining a standard scale according to the target scale, and performing padding processing on the initial image based on the standard scale to obtain a first image corresponding to the standard scale; inputting the first image and the prompt word into an image generation model to perform image generation processing, and outputting a second image corresponding to the standard scale; and performing cropping and scaling processing on the second image to obtain a target image corresponding to the target scale. The present method can improve the generation quality of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] Traditional advertising poster production is time-consuming and labor-intensive, requiring a significant investment of time and energy in shooting, retouching, and layout. The application of AIGC (Artificial Intelligence Generated Content) technology has greatly optimized this process. Designers can quickly obtain high-quality images and make fine adjustments simply by inputting text prompts, shortening the creative implementation cycle, reducing costs, and meeting personalized needs.

[0003] The current mainstream solutions for generating multi-size images in the AIGC field have uncontrollable generation effects. Due to the probabilistic generation characteristics of the model, when the size is greatly expanded to meet the irregular shape standard, problems such as image distortion are likely to occur, resulting in poor image generation quality in the final product. Summary of the Invention

[0004] This disclosure provides an image generation method, apparatus, and electronic device to at least solve the problem of poor image generation quality caused by image distortion when the size is significantly expanded due to the probabilistic generation characteristics of the model. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, an image generation method is provided, comprising:

[0006] Based on the raw image requirement data, determine the initial image and prompt words with the target ratio, wherein the prompt words are used to describe the image's content requirements;

[0007] If the target ratio meets the preset ratio requirement, a standard ratio is determined according to the target ratio, and the initial image is filled based on the standard ratio to obtain a first image corresponding to the standard ratio.

[0008] The first image and the prompt word are input into an image generation model for image generation processing, and a second image corresponding to the standard ratio is output.

[0009] The second image is cropped and scaled to obtain a target image corresponding to the target ratio.

[0010] In one embodiment, the step of filling the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio includes:

[0011] Based on the size information of the initial image and the standard ratio, determine the horizontal fill size and the vertical fill size;

[0012] Fill areas are added to the left and right sides of the initial image according to the horizontal fill size, and fill areas are added to the top and bottom sides of the initial image according to the vertical fill size, to obtain a first image corresponding to the standard ratio.

[0013] In one embodiment, cropping and scaling the second image to obtain a target image corresponding to the target ratio includes:

[0014] Based on the target size of the first image, the resolution of the second image is transformed to obtain a third image corresponding to the target size;

[0015] The third image is cropped based on the horizontal fill size and the vertical fill size to obtain a target image corresponding to the target ratio.

[0016] In one embodiment, determining the initial image and prompt words of the target ratio based on the raw image requirement data includes:

[0017] Obtain the target ratio and prompt words from the raw image demand data;

[0018] Based on the target ratio and the prompt words, an initial image corresponding to the target ratio is generated.

[0019] In one embodiment, generating an initial image corresponding to the target ratio based on the target ratio and the prompt word includes:

[0020] Determine the standard scale based on the target scale, and generate a canvas corresponding to the standard scale;

[0021] The canvas and the prompt words are input into the image generation model for image generation processing, and a fourth image corresponding to the standard ratio is output.

[0022] The fourth image is filled according to the target ratio to obtain an initial image corresponding to the target ratio.

[0023] In one embodiment, the preset ratio requirement includes a first preset ratio requirement and a second preset ratio requirement; wherein, the first preset ratio requirement includes a target ratio greater than a first ratio, and the second preset ratio requirement includes a target ratio less than a second ratio.

[0024] In one embodiment, determining the standard ratio based on the target ratio includes:

[0025] If the target ratio meets the first preset ratio requirement, the standard ratio is determined as the first standard ratio, and the first standard ratio corresponds to the horizontal image.

[0026] If the target ratio meets the second preset ratio requirement, the standard ratio is determined as the second standard ratio, and the second standard ratio corresponds to the vertical image.

[0027] In one embodiment, the filled region of the first image includes any one of the following: a solid color region, a Gaussian blurred image region, or a noisy image region.

[0028] According to a second aspect of the present disclosure, an image generation apparatus is provided, comprising:

[0029] The determining unit is configured to execute the determination of an initial image and prompt words based on the raw image requirement data, wherein the prompt words are used to describe the image's content requirements.

[0030] The filling unit is configured to perform the following actions when the target ratio meets the preset ratio requirement: determine a standard ratio based on the target ratio, and fill the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio.

[0031] The image generation unit is configured to perform image generation processing on the first image and the prompt word input to the image generation model, and output a second image corresponding to the standard ratio.

[0032] The image processing unit is configured to perform cropping and scaling processing on the second image to obtain a target image corresponding to the target ratio.

[0033] In one embodiment, the step of filling the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio includes:

[0034] Based on the size information of the initial image and the standard ratio, determine the horizontal fill size and the vertical fill size;

[0035] Fill areas are added to the left and right sides of the initial image according to the horizontal fill size, and fill areas are added to the top and bottom sides of the initial image according to the vertical fill size, to obtain a first image corresponding to the standard ratio.

[0036] In one embodiment, cropping and scaling the second image to obtain a target image corresponding to the target ratio includes:

[0037] Based on the target size of the first image, the resolution of the second image is transformed to obtain a third image corresponding to the target size;

[0038] The third image is cropped based on the horizontal fill size and the vertical fill size to obtain a target image corresponding to the target ratio.

[0039] In one embodiment, determining the initial image and prompt words of the target ratio based on the raw image requirement data includes:

[0040] Obtain the target ratio and prompt words from the raw image demand data;

[0041] Based on the target ratio and the prompt words, an initial image corresponding to the target ratio is generated.

[0042] In one embodiment, generating an initial image corresponding to the target ratio based on the target ratio and the prompt word includes:

[0043] Determine the standard scale based on the target scale, and generate a canvas corresponding to the standard scale;

[0044] The canvas and the prompt words are input into the image generation model for image generation processing, and a fourth image corresponding to the standard ratio is output.

[0045] The fourth image is filled according to the target ratio to obtain an initial image corresponding to the target ratio.

[0046] In one embodiment, the preset ratio requirement includes a first preset ratio requirement and a second preset ratio requirement; wherein, the first preset ratio requirement includes a target ratio greater than a first ratio, and the second preset ratio requirement includes a target ratio less than a second ratio.

[0047] In one embodiment, determining the standard ratio based on the target ratio includes:

[0048] If the target ratio meets the first preset ratio requirement, the standard ratio is determined as the first standard ratio, and the first standard ratio corresponds to the horizontal image.

[0049] If the target ratio meets the second preset ratio requirement, the standard ratio is determined as the second standard ratio, and the second standard ratio corresponds to the vertical image.

[0050] In one embodiment, the filled region of the first image includes any one of the following: a solid color region, a Gaussian blurred image region, or a noisy image region.

[0051] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement any of the image generation methods provided in the first aspect.

[0052] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the image generation methods provided in the first aspect.

[0053] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including instructions that, when executed by a processor of an electronic device, enable the electronic device to perform any of the image generation methods provided in the first aspect.

[0054] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0055] The image generation method, apparatus, electronic device, and storage medium provided in this disclosure can determine an initial image with a target ratio and prompts describing the image content requirements based on the image generation requirement data. If the target ratio meets the preset ratio requirements, a standard ratio is determined based on the target ratio, and the initial image is filled based on the standard ratio to obtain a first image with the corresponding standard ratio. The first image and prompts are input into an image generation model for image generation processing, and a second image with the corresponding standard ratio is output. Finally, the second image is cropped and scaled to obtain a target image with the corresponding target ratio. By employing the image generation method, apparatus, electronic device, and storage medium provided in this disclosure, when the target ratio meets the preset ratio requirements, it can be determined that the current initial image is an image with an irregular ratio, that is, the aspect ratio of the image is not the standard ratio, and the final generated image is also an image with an irregular ratio. At this time, the initial image with an irregular ratio is first filled to the standard ratio adapted by the model. Relying on the model's high-precision generation capability of standard ratio images, it is ensured that the image content is complete in detail and has a regular composition when it is generated. Then, it is restored to the target ratio and target size by cropping and scaling. This can avoid problems such as image distortion and element misalignment when the size is greatly expanded or the ratio deviates from the standard, thereby greatly improving the generation quality and presentation effect of the target image.

[0056] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0057] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0058] Figure 1 This is a flowchart illustrating an image generation method according to an exemplary embodiment.

[0059] Figure 2 This is a detailed flowchart illustrating step 104 according to an exemplary embodiment.

[0060] Figure 3 This is a detailed flowchart illustrating step 108 according to an exemplary embodiment.

[0061] Figure 4 This is a detailed flowchart illustrating step 102 according to an exemplary embodiment.

[0062] Figure 5 This is a detailed flowchart illustrating step 404 according to an exemplary embodiment.

[0063] Figure 6 This is a schematic diagram illustrating an image generation method according to an exemplary embodiment.

[0064] Figure 7a This is a schematic diagram of an initial image shown according to an exemplary embodiment.

[0065] Figure 7b This is a schematic diagram of a first image shown according to an exemplary embodiment.

[0066] Figure 7c This is a schematic diagram of a second image shown according to an exemplary embodiment.

[0067] Figure 7d This is a schematic diagram of a target image according to an exemplary embodiment.

[0068] Figure 8 This is a block diagram illustrating an image generation apparatus according to an exemplary embodiment.

[0069] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0070] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0071] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0072] Figure 1 This is a flowchart illustrating an image generation method according to an exemplary embodiment. This embodiment uses the application of the method to a terminal as an example for illustration. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes steps 102 to 108, wherein:

[0073] Step 102: Determine the initial image and prompts for the target scale based on the raw image requirement data. The prompts are used to describe the content requirements of the image.

[0074] In this embodiment, the raw image requirement data refers to a set of relevant data initiated by the user or system to define the image generation target. This data may include various parameters such as image size requirements, visual style, core elements, and application scenarios (e.g., advertising posters, promotional images). The target aspect ratio is the aspect ratio of the final image expected by the user. It can be a standard aspect ratio (e.g., 1:1, 4:3, 16:9) or an irregular aspect ratio (e.g., 1:5, 7:2). The initial image with the target aspect ratio is a preliminary image carrier determined based on the raw image requirement data. Its aspect ratio is consistent with the target aspect ratio and can serve as a reference image for the subsequent image generation model during the image generation process. The initial image can be directly included in the raw image requirement data, generated from user-uploaded basic materials, or automatically generated by the system based on a blank template according to the target aspect ratio. This embodiment does not specifically limit the specific generation method or initial image quality of the initial image.

[0075] Prompts are textual information used to describe the content requirements of an image. They may include the main subject, color scheme, lighting effects, details and textures, style (such as realistic, cartoon, retro, etc.), and background atmosphere. They are the primary basis for the image generation model to understand the intended image and generate a visually appealing image. For example, in an advertising poster scenario, the prompts could be set as "Red background, skincare product in the center, soft lighting, delicate texture, simple and sophisticated style, background adorned with scattered light spots, no unnecessary clutter." Prompts can be manually entered by the user or automatically generated by the system based on the image requirements data. This embodiment does not specifically limit the length or format of the prompts; those skilled in the art can flexibly set them according to the image accuracy requirements.

[0076] Step 104: If the target ratio meets the preset ratio requirements, determine the standard ratio according to the target ratio, and fill the initial image based on the standard ratio to obtain the first image with the corresponding standard ratio.

[0077] In this embodiment, the preset aspect ratio requirement is a standard used to determine whether the target aspect ratio is an irregular aspect ratio. It can be preset based on the adaptability of the image generation model and industry standards, or a specific range of aspect ratios can be customized as the preset aspect ratio requirement. By using the preset aspect ratio requirement, irregular aspect ratio scenes that need to be filled can be filtered out, avoiding the image distortion problem caused by directly generating images with irregular aspect ratios.

[0078] For example, if the target ratio meets the preset ratio requirement, it means that the current target ratio is an irregular ratio (i.e., not the standard ratio adapted by the image generation model). In this case, it is necessary to first determine the standard ratio of the adapted image generation model based on the target ratio, and then fill the initial image to make it reach the standard ratio. The image obtained after filling is used as the first image. The first image after filling corresponds to the standard ratio, and the corresponding size information is the target size. Here, filling refers to supplementing the edge area of ​​the initial image with adapted image content based on the standard ratio, so that the aspect ratio of the initial image is adjusted to the standard ratio, and finally the first image is obtained. For example, the content to be filled can fit the style and core elements of the picture described by the prompt. The system can automatically generate adapted background content based on the prompt to avoid the filled part from being abrupt with the main body of the initial image. For example, if the initial image is a vertical figure and the target ratio is 2:7, gradient background or texture elements consistent with the picture style can be added above and below the figure during filling.

[0079] In one exemplary embodiment, the preset ratio requirement includes a first preset ratio requirement and a second preset ratio requirement; wherein, the first preset ratio requirement includes a target ratio greater than a first ratio value, and the second preset ratio requirement includes a target ratio less than a second ratio value.

[0080] In this embodiment of the disclosure, the preset ratio requirement may specifically include a first preset ratio requirement and a second preset ratio requirement. The first preset ratio requirement is that the target ratio (i.e., aspect ratio) is greater than a first ratio value, and the second preset ratio requirement is that the target ratio is less than a second ratio value. The first and second ratio values ​​respectively correspond to extreme horizontal and vertical irregular aspect ratio scenarios. The first and second ratio values ​​can be flexibly set according to industry standards and the adaptation performance of the image generation model. In one example, the first ratio value can be set to 3 (i.e., aspect ratio > 3), and the second ratio value can be set to 1 / 3 (i.e., aspect ratio < 1 / 3). It should be noted that this embodiment of the disclosure does not limit the specific values ​​of the first and second ratio values; those skilled in the art can set them according to their needs.

[0081] By using the preset ratio requirement, irregularly shaped scenes that require filling can be filtered out, avoiding the image distortion problem caused by directly generating irregularly shaped scenes. In this embodiment, the first and second preset ratio requirements can achieve bidirectional threshold definition of the irregularly shaped scene range, improving the accuracy and flexibility of irregularly shaped scene determination.

[0082] In one exemplary embodiment, determining the standard ratio based on the target ratio includes:

[0083] If the target ratio meets the first preset ratio requirement, the standard ratio is determined as the first standard ratio, which corresponds to the horizontal image; if the target ratio meets the second preset ratio requirement, the standard ratio is determined as the second standard ratio, which corresponds to the vertical image.

[0084] In this embodiment of the disclosure, if the target ratio meets the preset ratio requirement, it indicates that the current target ratio is an irregular ratio. In this case, it is necessary to determine the standard ratio for the image generation model based on the specific type of the preset ratio requirement. The standard ratio is the aspect ratio at which the image generation model can stably output high-quality images, which is usually a common industry ratio. It is divided into a first standard ratio and a second standard ratio, which are adapted to different types of irregular ratio scenarios. Among them, the first standard ratio is a common standard ratio adapted to horizontal images, and the second standard ratio is a common standard ratio adapted to vertical images.

[0085] For example, if the target ratio meets the first preset ratio requirement (width-to-height ratio is greater than the first ratio value, such as width-to-height ratio is greater than 3), the standard ratio is determined as the first standard ratio (such as the first standard ratio is 3:1); if the target ratio meets the second preset ratio requirement (width-to-height ratio is less than the second ratio value, such as width-to-height ratio is less than 1:3), the standard ratio is determined as the second standard ratio (such as the first standard ratio is 1:3).

[0086] In this embodiment of the disclosure, the selection of the standard ratio needs to take into account both the adaptability of the image generation model and the convenience of subsequent cropping and restoration. Therefore, it is preferable to select the corresponding type of standard ratio with a smaller difference from the target ratio, which can reduce the impact of subsequent cropping operations on the core content of the image, thereby improving the image quality of the final target image.

[0087] Step 106: Input the first image and the prompt word into the image generation model for image generation processing, and output the second image with the corresponding standard ratio.

[0088] In this embodiment, the image generation model is capable of generating high-quality images based on text prompts and a base image. Its architecture can employ existing AIGC image generation architectures such as diffusion models or Transformer-based generation models. This embodiment does not limit the specific architecture of the model. This image generation model can accurately identify the content requirements described by the prompt words and, combined with the basic framework and style of the first image, generate images with complete details, regular composition, and conforming to standard proportions.

[0089] For example, training an image generation model includes the following steps: First, a training dataset is constructed, which contains a large number of text-image data pairs. The text data is used to describe the content, style, and other information of the corresponding images, consistent with the prompt word expression logic in the aforementioned embodiments. The image data are all high-quality images of standard proportions (including the first standard proportion and the second standard proportion), used to provide learning samples for the model. The training dataset can be further divided into a training set and a validation set, used for model parameter training and training effect verification, respectively.

[0090] Next, the image generation model parameters are initialized. Random initialization is used to assign values ​​to the network parameters of each layer, determining the optimizer, loss function (such as cross-entropy loss, mean squared error loss, etc.), batch size, number of iterations, and other hyperparameters. These hyperparameters can be flexibly adjusted according to the model architecture and computational resources. Then, the iterative training process begins: text data and corresponding standard-scale image data from the training set are input into the initial image generation model. The model generates image prediction results based on text semantics, calculates the differences between the predicted image and the sample image (including dimensions such as image detail, semantic alignment, and image quality) using the loss function, and updates the model network parameters using the backpropagation algorithm based on the loss value to minimize the prediction error.

[0091] After each iteration, the model's current performance can be evaluated using a validation set to determine whether the standard-ratio images generated by the model meet the image quality requirements and whether the text and image semantics are accurately aligned. If the evaluation result does not reach the preset performance threshold, training continues iteratively; if the preset performance threshold is reached or the preset number of iterations is completed, training stops, resulting in a trained image generation model. Furthermore, regularization and data augmentation strategies can be introduced during training to improve the model's generalization ability, avoid overfitting, and ensure that the model can accurately respond to prompts of different styles and content, stably outputting high-quality standard-ratio images.

[0092] The first image and the prompt are used as input to the image generation model. The model collaboratively parses the first image and the prompt, transforming the textual semantics in the prompt into visual elements. Simultaneously, leveraging the model's high-precision generation capability for standard-ratio images, it optimizes image details and avoids distortion issues. For example, if the first image is a filled 3:1 standard-ratio image, and the prompt is described as "blue sky background, white clouds, green grass and wildflowers below, a fresh and healing style," the image generation model will generate a 3:1 image—the second image—based on the 3:1 frame of the first image, that matches the style of the prompt and is rich in detail.

[0093] It should be noted that the embodiments disclosed herein do not impose specific limitations on the training method, generation speed, or resolution of the output image of the image generation model. These limitations can be flexibly adjusted according to actual computing resources and image accuracy requirements, as long as the image generation model can adapt to standard-ratio image generation and accurately respond to prompts.

[0094] Step 108: Crop and scale the second image to obtain a target image with the corresponding target ratio.

[0095] In this embodiment of the disclosure, the cropping and scaling process refers to adjusting the aspect ratio of a second image with a standard ratio to the target ratio by cropping the filling area and scaling the image size, so as to finally obtain a target image that meets the user's needs.

[0096] For example, the cropping and scaling process prioritizes adjusting the resolution of the second image to match the size of the first image, ensuring that the image size and resolution meet the requirements while maintaining the standard aspect ratio of the second image to avoid blurring or distortion during scaling. The cropping process is performed after scaling, prioritizing the preservation of core image elements and removing redundant areas or non-critical content created by padding. The aspect ratio of the image is adjusted to the target ratio, ensuring the cropped image has a complete subject and reasonable composition. The resulting target image has the target aspect ratio, and both size and resolution meet the requirements. The cropping and scaling process can be implemented using existing image processing algorithms. This embodiment does not limit the specific algorithm, as long as it achieves accurate aspect ratio conversion and ensures image quality. Furthermore, if the target ratio is a standard ratio, the second image can be directly used as the target image, or only simple scaling adjustments can be made to match the size requirements.

[0097] The image generation method provided in this disclosure can determine an initial image with a target ratio and prompts describing the image content requirements based on the image generation requirements data. If the target ratio meets the preset ratio requirements, a standard ratio is determined based on the target ratio, and the initial image is filled based on the standard ratio to obtain a first image with the corresponding standard ratio. The first image and prompts are then input into an image generation model for image generation processing, outputting a second image with the corresponding standard ratio. Finally, the second image is cropped and scaled to obtain a target image with the corresponding target ratio. Using the image generation method provided in this disclosure, if the target ratio meets the preset ratio requirements, it can be determined that the current initial image is an irregularly shaped image, meaning the aspect ratio is not standard, and the final generated image will also be an irregularly shaped image. In this case, the irregularly shaped initial image is first filled to the standard ratio adapted by the model. Relying on the model's high-precision generation capability for standard ratio images, it ensures complete details and regular composition when generating image content. Then, by cropping and scaling, it is restored to the target ratio and target size. This can fundamentally avoid problems such as image distortion and element misalignment when the size is greatly expanded or the ratio deviates from the standard, thereby significantly improving the generation quality and presentation effect of the target image.

[0098] In one exemplary embodiment, reference is made to Figure 2 As shown, in step 104, the initial image is filled based on a standard ratio to obtain a first image with the corresponding standard ratio. This may include steps 202 to 204, wherein:

[0099] Step 202: Based on the size information and standard ratio of the initial image, determine the horizontal fill size and the vertical fill size;

[0100] Step 204: Add fill areas to the left and right sides of the initial image according to the horizontal fill size, and add fill areas to the top and bottom sides of the initial image according to the vertical fill size, to obtain the first image with the corresponding standard ratio.

[0101] In this embodiment of the disclosure, the size information of the initial image refers to the actual width and height dimensions of the initial image (such as the width and height values ​​in pixels), which can be obtained by reading the attribute information of the initial image through the system. The horizontal fill size refers to the total width that needs to be added to the left and right sides of the initial image, and the vertical fill size refers to the total height that needs to be added to the top and bottom sides of the initial image. For example, based on the size information of the initial image and the standard ratio, the difference in width dimensions (i.e., the horizontal fill size) and the difference in height dimensions (i.e., the vertical fill size) required to achieve the standard ratio can be derived to ensure that the aspect ratio of the filled image perfectly matches the standard ratio.

[0102] For example, only one dimension can be filled, whether horizontal or vertical. Taking vertical filling as an example, assuming the initial image size is 700×200 pixels (aspect ratio 3.5, meeting the first preset ratio requirement), and the determined standard ratio is 3 (corresponding to the horizontal image). Since the initial horizontal image size is fixed at 700 pixels, adapting to the 3:1 standard ratio requires calculating the corresponding height as 700÷3≈233.33 pixels. Therefore, the vertical fill size is approximately 33.33 pixels, and the horizontal fill size is 0 pixels, meaning only vertical single-sided filling is performed.

[0103] It should be noted that the horizontal and vertical fill dimensions can both be positive, both be zero, or one of them can be zero, depending on the adaptation of the initial image's dimensions to the standard ratio. If a certain dimension of the initial image already meets the requirements of the corresponding dimension in the standard ratio, then the fill dimension for that dimension is zero, and only the other dimension is filled. If neither the width nor the height dimension of the initial image adapts to the standard ratio, then both the horizontal and vertical fill dimensions are positive, meaning that both horizontal and vertical fill are performed simultaneously to ensure that the aspect ratio of the filled image perfectly matches the standard ratio. This embodiment does not limit the specific range of fill dimensions; the standard ratio is the condition that the filled image meets the requirements.

[0104] In another example, both horizontal and vertical fill can be applied. Assume the initial image size is 200×700 pixels (aspect ratio ≈ 0.29, meeting the second preset ratio requirement), and the determined standard ratio is 1 / 3 (corresponding to the vertical image). Since the initial image's aspect ratio does not match the standard ratio, both width and height dimensions need to be filled simultaneously. Adhering to the principle of maintaining the original image's main body proportion, the height of the first image after filling is set to 750 pixels. Based on the 1:3 standard ratio, the corresponding width is calculated to be 250 pixels. Therefore, the horizontal fill size can be determined to be 50 pixels, and the vertical fill size to be 50 pixels, thus completing bidirectional filling in both width and height. This not only ensures that the aspect ratio of the filled image conforms to the standard ratio but also guarantees that the main body is not stretched or distorted.

[0105] The fill area refers to the image region added to the edges of the initial image based on the horizontal and / or vertical fill dimensions. The content of the fill area can be adapted to the visual style and core elements described by the prompt. It can be automatically generated by the system based on the prompt, or it can be generated using solid color fill, gradient fill, etc., to ensure the visual harmony of the first image after filling. Specifically, horizontal fill can distribute the calculated horizontal fill size evenly or according to a preset ratio to the left and right sides of the initial image, while vertical fill can distribute the vertical fill size evenly or according to a preset ratio to the top and bottom of the initial image. After filling is completed, the aspect ratio of the initial image will be adjusted to the standard ratio, forming a first image with the corresponding standard ratio.

[0106] The image generation method provided in this disclosure calculates and determines the horizontal and vertical fill dimensions by using the size information of the initial image and the standard ratio, and performs image filling based on the horizontal and vertical fill dimensions to ensure that the first image obtained after filling strictly conforms to the standard ratio. This allows the model to utilize its high-precision generation capability for standard ratio images, thereby avoiding problems such as image distortion and element misalignment that may occur in subsequent image generation processes from the source.

[0107] In one exemplary embodiment, the filled region of the first image includes any one of the following: a solid color region, a Gaussian blurred image region, or a noisy image region.

[0108] Among them, solid color areas refer to areas where the edges of the initial image are filled with a single color; Gaussian blurred image areas refer to images processed using a Gaussian blur algorithm as fill content, usually based on the edges of the main subject of the initial image, so that the fill area transitions naturally with the main subject of the initial image, suitable for realistic style and complex background image generation scenarios. This type of fill area can effectively weaken the fill boundary, avoid obvious fill marks, and improve the overall visual harmony of the first image; noisy image areas refer to images generated with random noise as fill content. Noisy image areas can enrich image details and at the same time cover up fill marks.

[0109] The image generation method provided in this embodiment can clearly define the boundary between the core generation area and the edge filling area, so that the core generation area is concentrated in the middle, which helps the image generation model to clarify the focus of the generated image and avoid the generation content offset or spatial layout chaos; and all types of filling areas can achieve natural connection with the subject of the initial image through reasonable content design, weaken the filling boundary, avoid obvious filling traces, and at the same time not interfere with the generation quality of the core area, ensuring the overall visual integrity of the generated first image.

[0110] In one exemplary embodiment, reference is made to Figure 3 As shown, in step 108, the second image is cropped and scaled to obtain a target image with the corresponding target ratio. This may include steps 302 to 304, wherein:

[0111] Step 302: Based on the target size of the first image, perform resolution transformation on the second image to obtain a third image corresponding to the target size;

[0112] Step 304: Crop the third image based on the horizontal and vertical fill dimensions to obtain a target image with the corresponding target ratio.

[0113] In this embodiment, the target size of the first image refers to the pixel size corresponding to the first image when adapted to a standard aspect ratio. Resolution transformation refers to adjusting the pixel size of the second image to match the pixel size requirement of the first image, while maintaining the aspect ratio of the second image unchanged to avoid distortion problems such as stretching and compression, ultimately generating a third image. The aspect ratio of the third image is the same as that of the second image, both being standard proportions, with only the pixel size adapted to the resolution requirements, which is the target size of the first image. It should be noted that resolution transformation can be implemented using existing image scaling algorithms. This embodiment does not limit the specific algorithm; for example, traditional bilinear interpolation or AI-based super-resolution algorithms can be used, as long as image quality and standard aspect ratio are maintained while adjusting the resolution.

[0114] The process of cropping the third image based on the horizontal and vertical fill sizes can be understood as the reverse of the aforementioned fill processing. If the horizontal fill size is allocated to the left and right sides of the initial image and the vertical fill size is allocated to the top and bottom sides during the fill processing, then the corresponding area equal to the horizontal fill size is cropped from the left and right sides of the third image, and the area equal to the vertical fill size is cropped from the top and bottom sides, resulting in the target image. The aspect ratio of the target image is consistent with the target ratio at this time, and the core generation area is completely preserved, avoiding the cropping of core image elements.

[0115] For example, taking an initial image with dimensions of 700×200 pixels as an example, the horizontal fill size is 0 pixels, and the vertical fill size is approximately 33.33 pixels (evenly distributed on the top and bottom sides, approximately 16.67 pixels each); the second image is 700×233.33 pixels (aspect ratio 3:1, standard proportion). The cropping process can then be based on the vertical fill size of 33.33 pixels, cropping approximately 16.67 pixels of redundant fill area from the top and bottom sides of the second image, resulting in a 700×200 pixel third image whose aspect ratio perfectly matches the target proportion.

[0116] The image generation method provided in this disclosure relies on the high-precision generation capability of the image generation model for standard-ratio images. After generating a high-quality standard-ratio image, the redundant filled areas are removed, and the standard-ratio image is directly restored to the target ratio. This effectively avoids the core generation area being mistakenly cropped, ensuring the integrity of the main subject in the third image and the final target image, thereby solving the problems of subject offset and missing parts that are prone to occur during the restoration of irregular ratios. At the same time, the target ratio remains unchanged during the resolution transformation process, which can not only meet the high-precision requirements of the data for generating images, but also retain image details to the maximum extent and avoid problems such as scaling distortion and image quality degradation, thereby greatly improving the image generation quality.

[0117] In one exemplary embodiment, reference is made to Figure 4 As shown, in step 102, determining the initial image and prompt words of the target ratio based on the raw image requirement data may include steps 402 to 404, wherein:

[0118] Step 402: Obtain the target ratio and prompt words from the raw image requirement data;

[0119] Step 404: Generate an initial image corresponding to the target scale based on the target scale and the prompt words.

[0120] In this embodiment, the raw image requirement data may only include text content and not image data. The target aspect ratio can then be obtained from the raw image requirement data. If the target aspect ratio meets a preset aspect ratio condition, it indicates that the target aspect ratio is an irregular aspect ratio. A high-quality initial image corresponding to the target aspect ratio can then be generated based on the target aspect ratio and the prompt text. That is, the initial image generation uses the target aspect ratio as a size constraint and the prompt text as a style and content basis. It is a basic image carrier with an aspect ratio consistent with the target aspect ratio and initially conforms to the visual requirements of the prompt text.

[0121] The image generation method provided in this embodiment analyzes and extracts the raw image requirement data to obtain the target ratio and prompt words. The initial image generated based on the target ratio and prompt word style provides a style and ratio reference for the image generation model. This avoids subsequent repeated adjustments caused by ratio deviations and reduces style conflicts between the initial image and the final generated image. No additional style calibration is required, reducing adaptation costs and workload.

[0122] In one exemplary embodiment, reference is made to Figure 5 As shown, in step 404, generating an initial image corresponding to the target ratio based on the target ratio and the prompt words may include the following steps 502 to 506, wherein:

[0123] Step 502: Determine the standard scale based on the target scale and generate a canvas with the corresponding standard scale.

[0124] Step 504: Input the canvas and prompt words into the image generation model for image generation processing, and output the fourth image with the corresponding standard ratio;

[0125] Step 506: Fill the fourth image according to the target ratio to obtain the initial image corresponding to the target ratio.

[0126] In this embodiment, the process of determining the standard ratio is the same as in the foregoing embodiments, and will not be repeated here. After determining the standard ratio, a canvas with the corresponding standard ratio can be generated. This canvas is the basic carrier for carrying the image generation content. For example, if the target ratio is 3.5 (meeting the first preset ratio requirement), the standard ratio can be determined to be 3. Based on this standard ratio, canvases with aspect ratios of 700×233.33 pixels and 1400×466.67 pixels, etc., conforming to a 3:1 ratio, can be generated.

[0127] Using the canvas and prompts as input to the image generation model, the model strictly adheres to the standard canvas proportions for composition, ensuring that the aspect ratio of the generated fourth image matches the standard ratio. It also faithfully reproduces the subject matter, color scheme, lighting effects, and stylistic tone described in the prompts, ensuring a high degree of match between the content, style, and intended image of the fourth image. Compared to directly generating images with irregular proportions, generating a fourth image using a standard-proportion canvas effectively reduces issues such as subject distortion, loss of detail, and chaotic composition, thereby improving image generation quality.

[0128] Since the fourth image is a standard-ratio image, it can be adjusted to the target ratio through fill processing, ultimately generating an initial image at the target ratio. For example, the fill processing can first calculate the required horizontal and vertical fill dimensions based on the size of the fourth image and the target ratio. The horizontal fill dimension is the total width that needs to be added to adapt the fourth image to the target ratio, and the vertical fill dimension is the total height that needs to be added. By deriving the difference between the target ratio and the size information of the fourth image, the horizontal and vertical fill dimensions that ensure the aspect ratio of the filled image strictly matches the target ratio can be obtained. The fill area can be allocated symmetrically or with a preset ratio to prioritize centering the core content of the fourth image and avoid subject offset. The fill area can also be a solid color area, a Gaussian blurred image area, or a noisy image area to avoid visual abruptness between the filled part and the main body of the fourth image, ensuring a unified style and harmonious composition in the generated initial image.

[0129] Using the image generation method provided in this disclosure, when the target ratio is an irregular ratio, a canvas adapted to the standard ratio can be generated based on the standard ratio matched by the irregular ratio. By relying on the high-precision generation capability of the image generation model for standard ratio images, problems such as subject distortion, lack of detail, and chaotic composition caused by directly generating images with irregular ratios can be avoided, resulting in a high-quality fourth image. Based on this fourth image, an initial image of the target ratio is constructed by filling. Based on this high-quality initial image and prompt words, image generation can be performed, which can greatly improve the generation quality and presentation effect of the target image.

[0130] To enable those skilled in the art to better understand the embodiments of this disclosure, the embodiments of this disclosure are described below through specific examples.

[0131] For example, refer to Figure 6 The diagram illustrates a data processing flow according to an embodiment of this disclosure, including:

[0132] Step S1: Obtain the generation request and parameter parsing. The system first receives the user's raw image requirement data. The raw image requirement data includes: target width and height (size, such as Wtarget (width) × Htarget (height)) and prompt (text used to describe the content of the image), where the target width and height can be calculated from the initial image carried in the raw image requirement data.

[0133] Step S2: Calculate the target ratio Ratio = Wtarget / Htarget. If Ratio > 3 (extra-wide horizontal image) or Ratio < 1 / 3 (extra-long vertical image), then the target ratio is determined to be an irregular ratio.

[0134] Step S3: Calculate and map the standard ratio. If the target ratio is greater than 3, determine the standard ratio as 3:1; if the target ratio is less than 1 / 3, determine the standard ratio as 1:3.

[0135] Step S4: Dimension Alignment and Fill Size Calculation. The fill size can be used to scale the target scale to the same width scale (for vertical irregularities) or the same height scale (for horizontal irregularities), effectively adjusting the target scale to the standard scale through fill. Taking width alignment as an example, the system calculates the difference between the standard scale's required height Hstd and the target scale's current height Htarget at this width, denoted as the vertical fill size h. The calculation process for the horizontal fill size is similar.

[0136] Step S5: Padding. Top edge padding: Fill the area above the initial image at the target scale with a height of 1 / 2 × h; Bottom edge padding: Fill the area below the initial image at the target scale with a height of 1 / 2 × h. At this point, the aspect ratio of the constructed first image reverts to the standard proportions familiar to the model. The padding area can be a solid color, a Gaussian blur, or a noise map to help the model understand spatial location, but the core generated area is concentrated in the center.

[0137] Step S6: AI Generation (Inpainting / Generation). Input the processed first image and the prompt word into an AIGC model (such as Stable Diffusion). The model generates a second image based on the prompt word and the first image at a standard scale. Because the input scale is normal, the model can generate a structurally complete and perspective-correct subject image.

[0138] Step S7: Image scaling and cropping (inverse transformation). Image scaling: Remove the area of ​​the upper edge 1 / 2×h and the lower edge 1 / 2×h. Scale the generated high-resolution second image back to the width and height corresponding to the first image. Edge cropping: Crop the edges according to the fill size h. After cropping, the remaining middle part is the high-quality image that conforms to the target proportions and whose main compositional elements are not stretched or distorted.

[0139] In one example, suppose an application needs an image with dimensions of 584*160 pixels (ratio 3.65:1, which is a non-standard aspect ratio, with an aspect ratio greater than 3). The target ratio is 3.65:1, while the established standard ratio is 3:1.

[0140] At 2k pixels, the generated image at a 3:1 aspect ratio is 3552*1184 pixels. First, generate an image of this size using an AIGC model and prompts, referring to... Figure 7a As shown; a target aspect ratio of 3.65:1 indicates that the image is wider, so adjustments are needed. Figure 7aThe left and right sides, as well as the top and bottom sides, of the image shown are filled to transform a non-standard size problem into a standard size generation problem. The filling result is referenced. Figure 7b As shown. The filled image and prompts are input into the AIGC model, and the generated image is referenced. Figure 7c As shown, the image output by the AIGC model is cropped and resized to obtain a target image with a target size of 584*160px. (Refer to...) Figure 7d As shown.

[0141] The image generation method provided in this disclosure can flexibly generate advertising materials of any proportion based on the same benchmark model, solving the limitation of traditional models that can only generate fixed proportions and greatly improving the adaptation efficiency of materials when they are placed in multiple locations. By calculating the height difference h -> filling the top and bottom edges -> generating -> cropping, the model is forced to perform inference under a standard canvas ratio. This avoids the spatial distortion and illusion of multiple heads and feet caused by the model when dealing with extreme aspect ratios, ensuring the integrity of the main structure of the generated content. This method does not require collecting thousands of irregularly shaped images for model fine-tuning; it can be implemented by directly reusing existing general-purpose SD (Stable Diffusion) or MJ (MidJourney) models. This not only saves the expensive costs of data annotation and training but also allows the solution to be quickly transferred to different base models (such as from a two-dimensional model to a realistic model).

[0142] It should be understood that, although Figures 1-6 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 1-6 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0143] It is understood that the same / similar parts between the various embodiments of the methods described above in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments, and relevant parts can be referred to the description of other method embodiments.

[0144] Figure 8 This is a block diagram 800 of an image generation apparatus according to an exemplary embodiment. (Refer to...) Figure 8The device includes a determining unit 802, a filling unit 804, an image generating unit 806, and an image processing unit 808, wherein:

[0145] The determining unit 802 is configured to execute the determination of the initial image and prompt words based on the raw image requirement data, wherein the prompt words are used to describe the image content requirements;

[0146] The filling unit 804 is configured to perform the following operations when the target ratio meets the preset ratio requirement: determine a standard ratio based on the target ratio, and fill the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio.

[0147] The image generation unit 806 is configured to perform image generation processing on the first image and the prompt word input to the image generation model, and output a second image corresponding to the standard ratio.

[0148] The image processing unit 808 is configured to perform cropping and scaling processing on the second image to obtain a target image corresponding to the target ratio.

[0149] Using the image generation apparatus provided in this embodiment, when the target ratio meets the preset ratio requirements, it can be determined that the current initial image is an image with an irregular ratio, that is, the aspect ratio of the image is not the standard ratio, and the final generated image is also an image with an irregular ratio. At this time, the initial image with an irregular ratio is first filled to the standard ratio adapted by the model. Relying on the model's high-precision generation capability of standard ratio images, it is ensured that the image content is complete in detail and the composition is regular when it is generated. Then, it is restored to the target ratio and target size by cropping and scaling. This can avoid problems such as image distortion and element misalignment when the size is greatly expanded or the ratio deviates from the standard, thereby greatly improving the generation quality and presentation effect of the target image.

[0150] In one embodiment, the step of filling the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio includes:

[0151] Based on the size information of the initial image and the standard ratio, determine the horizontal fill size and the vertical fill size;

[0152] Fill areas are added to the left and right sides of the initial image according to the horizontal fill size, and fill areas are added to the top and bottom sides of the initial image according to the vertical fill size, to obtain a first image corresponding to the standard ratio.

[0153] In one embodiment, cropping and scaling the second image to obtain a target image corresponding to the target ratio includes:

[0154] Based on the target size of the first image, the resolution of the second image is transformed to obtain a third image corresponding to the target size;

[0155] The third image is cropped based on the horizontal fill size and the vertical fill size to obtain a target image corresponding to the target ratio.

[0156] In one embodiment, determining the initial image and prompt words of the target ratio based on the raw image requirement data includes:

[0157] Obtain the target ratio and prompt words from the raw image demand data;

[0158] Based on the target ratio and the prompt words, an initial image corresponding to the target ratio is generated.

[0159] In one embodiment, generating an initial image corresponding to the target ratio based on the target ratio and the prompt word includes:

[0160] Determine the standard scale based on the target scale, and generate a canvas corresponding to the standard scale;

[0161] The canvas and the prompt words are input into the image generation model for image generation processing, and a fourth image corresponding to the standard ratio is output.

[0162] The fourth image is filled according to the target ratio to obtain an initial image corresponding to the target ratio.

[0163] In one embodiment, the preset ratio requirement includes a first preset ratio requirement and a second preset ratio requirement; wherein, the first preset ratio requirement includes a target ratio greater than a first ratio, and the second preset ratio requirement includes a target ratio less than a second ratio.

[0164] In one embodiment, determining the standard ratio based on the target ratio includes:

[0165] If the target ratio meets the first preset ratio requirement, the standard ratio is determined as the first standard ratio, and the first standard ratio corresponds to the horizontal image.

[0166] If the target ratio meets the second preset ratio requirement, the standard ratio is determined as the second standard ratio, and the second standard ratio corresponds to the vertical image.

[0167] In one embodiment, the filled region of the first image includes any one of the following: a solid color region, a Gaussian blurred image region, or a noisy image region.

[0168] Figure 9 This is a block diagram illustrating an electronic device 900 for an image generation method according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0169] Reference Figure 9 The electronic device 900 may include one or more of the following components: processing component 902, memory 904, power supply component 906, multimedia component 908, audio component 910, input / output (I / O) interface 912, sensor component 914, and communication component 916.

[0170] Processing component 902 typically controls the overall operation of electronic device 900, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 902 may include one or more modules to facilitate interaction between processing component 902 and other components. For example, processing component 902 may include a multimedia module to facilitate interaction between multimedia component 908 and processing component 902.

[0171] Memory 904 is configured to store various types of data to support the operation of electronic device 900. Examples of such data include instructions for any application or method operating on electronic device 900, contact data, phonebook data, messages, pictures, videos, etc. Memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, optical disk, or graphene storage.

[0172] Power supply component 906 provides power to various components of electronic device 900. Power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 900.

[0173] Multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 908 includes a front-facing camera and / or a rear-facing camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0174] Audio component 910 is configured to output and / or input audio signals. For example, audio component 910 includes a microphone (MIC) configured to receive external audio signals when electronic device 900 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 904 or transmitted via communication component 916. In some embodiments, audio component 910 also includes a speaker for outputting audio signals.

[0175] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0176] Sensor assembly 914 includes one or more sensors for providing state assessments of various aspects of electronic device 900. For example, sensor assembly 914 can detect the on / off state of electronic device 900, the relative positioning of components such as the display and keypad of electronic device 900, changes in position of electronic device 900 or its components, the presence or absence of user contact with electronic device 900, orientation or acceleration / deceleration of device 900, and temperature changes of electronic device 900. Sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 914 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0177] Communication component 916 is configured to facilitate wired or wireless communication between electronic device 900 and other devices. Electronic device 900 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 6G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 916 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 916 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0178] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0179] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, which can be executed by a processor 920 of an electronic device 900 to perform the above-described method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0180] In an exemplary embodiment, a computer program product is also provided, the computer program product including instructions that can be executed by a processor 920 of an electronic device 900 to perform the above-described method.

[0181] It should be noted that the above-mentioned apparatus, electronic equipment, computer-readable storage medium, computer program product, etc., may also include other implementation methods according to the description of the method embodiments. For specific implementation methods, please refer to the description of the relevant method embodiments, which will not be elaborated here.

[0182] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0183] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An image generation method, characterized in that, include: Based on the raw image requirement data, determine the initial image and prompt words with the target ratio, wherein the prompt words are used to describe the image's content requirements; If the target ratio meets the preset ratio requirement, a standard ratio is determined according to the target ratio, and the initial image is filled based on the standard ratio to obtain a first image corresponding to the standard ratio. The first image and the prompt word are input into an image generation model for image generation processing, and a second image corresponding to the standard ratio is output. The second image is cropped and scaled to obtain a target image corresponding to the target ratio.

2. The method according to claim 1, characterized in that, The step of filling the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio includes: Based on the size information of the initial image and the standard ratio, determine the horizontal fill size and the vertical fill size; Fill areas are added to the left and right sides of the initial image according to the horizontal fill size, and fill areas are added to the top and bottom sides of the initial image according to the vertical fill size, to obtain a first image corresponding to the standard ratio.

3. The method according to claim 2, characterized in that, The step of cropping and scaling the second image to obtain a target image corresponding to the target ratio includes: Based on the target size of the first image, the resolution of the second image is transformed to obtain a third image corresponding to the target size; The third image is cropped based on the horizontal fill size and the vertical fill size to obtain a target image corresponding to the target ratio.

4. The method according to any one of claims 1 to 3, characterized in that, The initial image and prompts for determining the target aspect ratio based on the raw image requirement data include: Obtain the target ratio and prompt words from the raw image demand data; Based on the target ratio and the prompt words, an initial image corresponding to the target ratio is generated.

5. The method according to claim 4, characterized in that, The step of generating an initial image corresponding to the target ratio based on the target ratio and the prompt word includes: Determine the standard scale based on the target scale, and generate a canvas corresponding to the standard scale; The canvas and the prompt words are input into the image generation model for image generation processing, and a fourth image corresponding to the standard ratio is output. The fourth image is filled according to the target ratio to obtain an initial image corresponding to the target ratio.

6. The method according to any one of claims 1 to 3, characterized in that, The preset ratio requirement includes a first preset ratio requirement and a second preset ratio requirement; wherein, the first preset ratio requirement includes a target ratio greater than a first ratio, and the second preset ratio requirement includes a target ratio less than a second ratio.

7. The method according to claim 6, characterized in that, Determining the standard ratio based on the target ratio includes: If the target ratio meets the first preset ratio requirement, the standard ratio is determined as the first standard ratio, and the first standard ratio corresponds to the horizontal image. If the target ratio meets the second preset ratio requirement, the standard ratio is determined as the second standard ratio, and the second standard ratio corresponds to the vertical image.

8. The method according to claim 2, characterized in that, The filled area of ​​the first image includes any one of the following: a solid color area, a Gaussian blurred image area, or a noisy image area.

9. An image generation apparatus, characterized in that, include: The determining unit is configured to execute the determination of an initial image and prompt words based on the raw image requirement data, wherein the prompt words are used to describe the image's content requirements. The filling unit is configured to perform the following actions when the target ratio meets the preset ratio requirement: determine a standard ratio based on the target ratio, and fill the initial image based on the standard ratio to obtain a first image corresponding to the standard ratio. The image generation unit is configured to perform image generation processing on the first image and the prompt word input to the image generation model, and output a second image corresponding to the standard ratio. The image processing unit is configured to perform cropping and scaling processing on the second image to obtain a target image corresponding to the target ratio.

10. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the image generation method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the image generation method as described in any one of claims 1 to 8.