A material expansion cutting beautifying method and system

By employing a step-by-step strategy involving the image expansion model ProOut, the layout generation model PosterO, the cropping model S2CNet, and the refinement network Refiner, the problem of semantic inconsistency in the composition of multimodal large models is solved. This achieves automatic adjustment of visual unity and compositional aesthetics at any size, making it suitable for advertising design and digital marketing.

CN121962348BActive Publication Date: 2026-07-31GUANGZHOU TAIDONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU TAIDONG TECH CO LTD
Filing Date
2026-01-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

When using multimodal large models for composition, existing technologies rely on semantic understanding limited to local matching, which fails to achieve semantic coordination of the overall design logic. This results in mechanical or uninspired generated results, making it difficult to maintain visual consistency and stylistic uniformity at any size.

Method used

By employing the image expansion model ProOut, the layout generation model PosterO, the cropping model S2CNet, and the refinement network Refiner, the layer content ratio and layout are adjusted through a step-by-step strategy. Combined with control signals and a composition planning module, a structured layout tree is generated to optimize the layout results and achieve visual unity and compositional aesthetics.

Benefits of technology

Automatically adjusts the proportion and layout of layer content at any size, reducing manual design costs and achieving visual unity and compositional aesthetics, suitable for scenarios such as advertising design and digital marketing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962348B_ABST
    Figure CN121962348B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image processing technology, and more particularly to a method and system for expanding, cropping, and beautifying source material. The method includes: inputting a background image into a preset image expansion model to obtain a target background image; inputting the target background image and the main element information of a subject image into a preset layout generation model to obtain a first layout result; combining the subject image and the target background image according to the first layout result to obtain a combined image; cropping the combined image according to a target size to obtain a cropped image; inputting the cropped image and the decorative element information of a decorative layer into the layout generation model to obtain a second layout result; optimizing the second layout result using a preset refined network based on the cropped image to obtain a third layout result; and combining the decorative layer and the cropped image according to the third layout result to generate a target image. The method of this invention ensures semantic consistency and aesthetic appeal in the composition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and system for expanding, cropping and beautifying materials. Background Technology

[0002] With the rapid development of digital advertising, e-commerce, and multi-platform media, the same creative material often needs to be adapted to display scenarios of different sizes and proportions. In the existing design process, the composition and visual aesthetics of the creative material rely heavily on manual adjustments, which not only consumes a lot of time and manpower but also makes it difficult to maintain visual consistency and stylistic unity in multi-size adaptation tasks.

[0003] Existing technologies primarily employ automatic layout planning schemes based on multimodal large models. These schemes abstract elements such as images, text, and icons into semantic tags or layout sequences, utilizing multimodal large models to understand the semantic relationships between content and text, and generating hierarchical layout structures within a given canvas size. However, the semantic understanding of these schemes is often limited to local matching, lacking a grasp of the overall design logic and failing to truly achieve semantic coordination between content and composition. The generated results often appear mechanical or lack design flair.

[0004] Therefore, ensuring semantic consistency and aesthetic appeal of the composition at any size is a pressing technical problem that needs to be solved. Summary of the Invention

[0005] To address the aforementioned technical problem of semantic inconsistency when using multimodal large models for graph construction, this invention provides solutions in the following aspects.

[0006] In a first aspect, the present invention provides a method for expanding, cropping, and beautifying source material, comprising: obtaining layers corresponding to a target size and a target image source material; the layers including a background image, a main image, and a decorative layer; inputting the background image into a preset image expansion model for expansion filling to obtain a target background image; inputting the main element information of the target background image and the main image into a preset layout generation model to obtain a first layout result; combining the main image and the target background image according to the first layout result to obtain a combined image; cropping the combined image according to the target size to obtain a cropped image; inputting the cropped image and the decorative element information of the decorative layer into the layout generation model to obtain a second layout result; optimizing the second layout result using a preset refined network based on the cropped image to obtain a third layout result; and combining the decorative layer and the cropped image according to the third layout result to generate a target image.

[0007] Furthermore, after obtaining the cropped image, the method further includes: using a preset image coordination model to soften the cropped image.

[0008] Furthermore, the combined image is cropped according to the target size, including cropping the combined image using S2CNet based on the target size.

[0009] Furthermore, the image expansion model is a ProOut model, and the background image is input into the preset image expansion model for expansion filling, including: generating control signals by the composition planning module; The control signal is injected into the intermediate feature layer of the latent space diffusion network to guide the generation of local expansion and obtain the local expansion result of the current iteration. The local expansion result generated in the current iteration step is pasted back to the corresponding position to update the global intermediate expansion image; the global intermediate expansion image is generated step by step through multiple progressive expansion iterations with the background image as the initial state.

[0010] Furthermore, the calculation expression for the control signal is as follows:

[0011] In the formula, For control signals, This represents the intermediate feature mapping of the current diffusion network U-Net. Zero convolutional layer, To control the encoder, Features that provide compositional hints For global semantic conditions, , These are all learnable parameters in the control encoder. For time steps.

[0012] Furthermore, the calculation expression for composition cue features is as follows:

[0013] In the formula, For target masking, For reference image, It is a shallow convolutional layer. This is for splicing operations.

[0014] Furthermore, the refinement network is the Refiner network in SEGA; the second layout result is optimized using a preset refinement network, including: rendering the elements corresponding to the decoration layer onto the clipping image based on the second layout result to generate a visual cue image; inputting the visual cue image, the geometrically normalized second layout result, and the corresponding task instructions into the Refiner network to obtain the third layout result and the layout quality score.

[0015] Further, the main element information of the target background image and the main image is input into a preset layout generation model to obtain the first layout result, including: extracting design intent information from the target background image; retrieving similar example layouts from a preset database; combining the main element information, design intent information, and example layouts into a Prompt; inputting the Prompt into an LLM model to generate a structured layout tree containing element spatial attributes and hierarchical relationships to obtain the first layout result.

[0016] Furthermore, the layout generation model is PosterO.

[0017] In a second aspect, the present invention provides a material expansion, cropping, and enhancement system, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the material expansion, cropping, and enhancement method described in the first aspect is implemented.

[0018] The beneficial effects of this invention are as follows: by adopting a step-by-step strategy of first expanding the background image, then laying out the main image, then cropping the combined image, and finally laying out the decorative layer, the proportion and layout of the content of each layer can be automatically adjusted under any given target size. This reduces manual design costs while ensuring visual unity and compositional aesthetics, and provides an efficient and intelligent re-creation solution for advertising design, digital marketing and other scenarios. Attached Figure Description

[0019] Figure 1 The process of material expansion, cropping and beautification method in the embodiments of the present invention Figure 1 ; Figure 2 The process of material expansion, cropping and beautification method in the embodiments of the present invention Figure 2 ; Figure 3 This is a structural block diagram of the material expansion, cropping, and beautification system in an embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0022] Figure 1 The process of material expansion, cropping and beautification method in the embodiments of the present invention Figure 1 .

[0023] In a first aspect, the present invention provides a method for expanding, cropping, and beautifying materials. Specifically, in conjunction with... Figure 1 and Figure 2 As shown, the method of the present invention includes the following steps.

[0024] S1. Obtain the target size and the layer corresponding to the target image material.

[0025] In this embodiment, the layers corresponding to the target image material include a background image, a main image, and decorative layers. The decorative layers include text, icons, and other elements. It can be understood that each layer corresponds to a specific element; for example, the main image corresponds to a person or object, and the text layer corresponds to text.

[0026] Specifically, the target size is input by the user, and the layer corresponding to the target image material is stored in a preset database, which is selected and entered by the user.

[0027] S2. Input the background image into the preset image expansion model for expansion filling to obtain the target background image.

[0028] In this embodiment, the image expansion model adopts the ProOut model. The ProOut model is an expansion painting model that can achieve progressive expansion at any size. The ProOut model adopts an architecture that combines a stable diffusion-based branch based on latent space with a composition planning module (CPM). The stable diffusion branch is responsible for filling the blank parts of the local target window, while the composition planning module provides global information to the stable diffusion branch to achieve natural expansion.

[0029] Specifically, set For the first Subglobal intermediate extended image, from The target window area is cropped to obtain the first... Input image for this operation And create a mask for the region to be filled. At the same time, for Scaling is performed to obtain the reference image. And mark the area of ​​the target window to obtain the target mask corresponding to the reference image. .

[0030] In each operation At that time, the composition planning module is based on and Extract global information and pass it to the stable diffusion branch, enabling the stable diffusion branch to utilize the global information. , Generate local expansion results with other conditions (such as text prompts). Then Composite into the global window to obtain the next complete global intermediate extended image. Repeat the above process until all outer areas are filled, and the final target background image is obtained.

[0031] In the stable diffusion branch, the input image I containing holes (i.e., the region to be filled) is first compressed into latent space features by a variational autoencoder. At the same time, the mask M of the region to be filled is marked. Then, the mask of the region to be filled is adjusted to the same spatial dimension as the latent space features by a downsampling operation. The latent space features and the downsampled mask are concatenated to obtain the noise latent variable. The noise latent variable is used as the input of the denoising network.

[0032] In one embodiment, the process of obtaining the noise latent variable is represented as follows:

[0033] In the formula, For time step The noise latent variable, This indicates a channel-level concatenation operation. This indicates a downsampling operation. This is the mask for the region to be filled. Features of latent space.

[0034] Corresponding optimization objective / loss function It can be:

[0035] In the formula, This represents the mathematical expectation function operation. Features of the initial latent space For time steps, These are constraints used to guide model generation / editing, such as text prompts and composition constraints. To follow a standard normal distribution random noise, For time steps noise latent variables, For noise reduction networks, It is an L2 norm.

[0036] This loss function guides the model to progressively denoise within the latent space, generating locally expanded results. and the results of local expansion Update global intermediate expansion image This yields a new global intermediate extended image. It is understandable that the global intermediate expansion image is generated gradually through multiple progressive expansions, starting from the background image.

[0037] In one embodiment, to maintain semantic coherence and compositional alignment between the extended region and the original image, a ControlNet-based CPM is employed. The ControlNet-based CPM generates control signals for compositional planning via a control encoder, injecting global information into the stable diffusion branch. The generation process of the control signals can be represented as follows:

[0038] In the formula, For control signals, This represents the intermediate feature mapping of the current diffusion network U-Net. Zero-convolutional layers are used to implement feature mapping and weight resetting along the channel dimension. To control the encoder, These are compositional cue features, i.e., local compositional features. These are global semantic conditions, containing global information about the entire image, describing the image's overall style, tone, and visual orientation. , These are all learnable parameters in the control encoder. For time steps.

[0039] The control signal is injected into the middle layer of the latent space diffusion network and jointly updated with the noise prediction features, so that the ProOut model maintains the continuity of local details and follows the overall composition constraints when expanding outward, thereby achieving global coordination in spatial distribution, color style and semantic rhythm.

[0040] In one embodiment, composition cue features The extraction process can be represented as:

[0041] In the formula, For target masking, For reference image, It is a shallow convolutional layer. This is for splicing operations.

[0042] In one embodiment, the ProOut model is trained with a small number of samples to meet the compositional characteristics and background style requirements of advertising and commercial visual design scenarios, so that it can maintain the global composition. Figure 1 While maintaining consistency, it can more accurately learn common color matching rules and visual layers in advertising materials, thereby achieving background expansion generation that is stylistically consistent, semantically harmonious, and has commercial aesthetic characteristics.

[0043] Compared to simple stretching or filling, by introducing the ProOut model, it is possible to generate an expanded background with natural texture and consistent style with the original background. Therefore, no matter how the target size changes, the system can pre-build a perfect background image, effectively solving the problem of stretching or cramped composition caused by insufficient background when the size is changed.

[0044] S3. Input the main element information of the target background image and the main image into the preset layout generation model to obtain the first layout result.

[0045] In this embodiment, the main element information refers to the information of the element graph, such as its length, width, and shape. This main element information is pre-stored in a database and input into the layout generation model as a prompt. The layout generation model used is PosterO. PosterO models the layout as a structured layout tree, where nodes represent elements and edges represent the hierarchy and subordination relationships between elements.

[0046] Specifically, the U-net model extracts placeable areas / design intent information from the target background image; retrieves similar example layouts from a pre-set database (which stores the layout tree corresponding to each training data); and combines the design intent information, example layouts, and main element information into a Prompt in the form of structured text. This Prompt is then input into the LLM model to generate a structured layout tree of element spatial attributes and hierarchical relationships, which is used as the first layout result.

[0047] When constructing a structured layout tree, nodes that are geometrically aligned and have similar boundaries are matched using a preset condition function. With nodes They are merged into the same subtree, thus constructing a layout tree structure that simultaneously encodes element shapes, relative geometric relationships, and design intent.

[0048] In one embodiment, the condition function It can be:

[0049] In the formula, and These are two nodes in the layout tree, representing either elements or design intent areas. For nodes The horizontal coordinate of the top-left corner of the bounding box of the corresponding element. For nodes The y-axis coordinate of the top-left corner of the bounding box of the corresponding element. For nodes The width of the corresponding element. For nodes The height of the corresponding element, For nodes The horizontal coordinate of the top-left corner of the bounding box of the corresponding element. For nodes The y-axis coordinate of the top-left corner of the bounding box of the corresponding element. For nodes The width of the corresponding element. For nodes The height of the corresponding element, A preset tolerance threshold is used to determine whether the boundaries of two nodes approximately coincide in space. Represents the logical AND operation, when The node will only be joined if all sub-conditions of the connection are true. With nodes Merge them into the same subtree.

[0050] By constructing a structured layout tree, we can understand the semantic relationships and functional priorities between different layers, ensuring the rationality and professional aesthetics of the composition design.

[0051] S4. Based on the first layout result, combine the main image with the target background image to obtain a combined image.

[0052] In this embodiment, the first layout result includes the coordinates, category, and size of each element within the canvas / layout. Based on this first layout result, the elements corresponding to the main image are rendered onto the target background image to obtain a combined image. Specifically, the combination process employs an image synthesis technique based on the alpha channel. First, an affine transformation matrix is ​​constructed based on the coordinate and size parameters (and possibly rotation angle parameters) in the first layout result. Using this matrix and a bilinear interpolation algorithm, the main image and its corresponding alpha transparency channel are resampled and aligned, mapping them to the coordinate system of the target background image. Subsequently, according to the formula... The transformed subject image and the target background image are then superimposed using pixel-level weighted summaries. The normalized transparency value. This represents the RGB color value of the transformed main image at this pixel location. The RGB color value of the target background image at that pixel location is used to obtain a combined image containing the subject and background.

[0053] In other alternative embodiments, those skilled in the art may also employ other existing combinations of methods, which are not limited here.

[0054] S5. Crop the combined images according to the target size to obtain the cropped image.

[0055] In this embodiment, S2CNet (Spatial-Semantic Collaborative Cropping Network) is used to crop the combined image to obtain a cropped image of the target size.

[0056] Specifically, for the input image (i.e., the composite image), multiple cropping candidate regions are first generated based on the target size. A pre-trained Faster R-CNN is then used to mine the target region of the input image, obtaining several potential visual targets. Subsequently, a convolutional neural network is used to extract features from the input image, and feature representations of each target region and cropping candidate region are obtained through region feature alignment. The cropping candidate regions and target regions are modeled as nodes in a graph structure. Semantic relationships are constructed based on region feature similarity, and spatial relationships are established by combining regional spatial location information, thus forming a spatial-semantic co-representation of region relationships. By introducing a graph-aware attention mechanism, information between region nodes is adaptively propagated and aggregated to characterize the influence of different target regions on the cropping candidate regions. The cropping quality score is predicted based on the aggregated cropping candidate region features, and the region with the best score is selected as the final cropping result.

[0057] By using S2CNet to crop the combined image, we can ensure that, within the constraint of the target size, the area of ​​key semantics in the combined image is maximized in the cropped image, while ensuring the integrity of the semantic information.

[0058] In one embodiment, after obtaining the cropped image, the method of the present invention further includes: applying a preset image coordination model to soften the cropped image to obtain a softened cropped image. The image coordination model can be Diff-Harmonization; since Diff-Harmonization is existing technology, the specific softening process will not be described in detail.

[0059] By introducing Diff-Harmonization, the jagged edges and abruptness caused by simple layer overlay are resolved, achieving a smooth blending and color harmony between the background image and the main image, thereby improving the aesthetics of the composition.

[0060] S6. Input the clipping image and decorative element information of the decorative layer into the layout generation model to obtain the second layout result.

[0061] In this embodiment, decorative element information refers to the information of the decorative layer element image, such as its length, width, shape, etc. The decorative element information is also pre-stored in the database and input into the layout generation model in the form of a Prompt.

[0062] Specifically, under the constraints of the clipping image, a layout tree is constructed for the decoration layer using a layout generation model to obtain the second layout result. The specific processing procedure is similar to S3, so it will not be described in detail here.

[0063] S7. Based on the cropping image, the second layout result is optimized using a preset refined network to obtain the third layout result.

[0064] In this embodiment, the fine-level refinement network adopts the Refiner (Fine-level Refinement) network from SEGA (Stepwise Evolution Paradigm for Content-Aware Layout Generation). The Refiner network can perform fine-level evolution with design prior knowledge based on the initial results generated by the structured layout tree.

[0065] Specifically, based on the second layout result, the elements corresponding to the decorative layer are rendered onto the cropping image to obtain a visual cue image; in one embodiment, the process of obtaining the visual cue image can be represented as:

[0066] In the formula, For visual cue images, For cropping the image, This is the result of the third layout. This is a rendering function used to draw the elements corresponding to the decoration layer onto the clipping image.

[0067] Furthermore, the second layout result is geometrically normalized, that is, the shape of each node (each element) in the second layout result is converted into a rectangular box, and its size and position are mapped to the same coordinate system. In other words, the second layout result (represented by SVG) is converted into a sequence of bounding boxes.

[0068] Furthermore, the visual cue image, the geometrically normalized second layout result, and the corresponding task instructions are input into the Refiner network to optimize and obtain the third layout result and layout quality score.

[0069] In one embodiment, the optimization process can be represented as:

[0070] In the formula, The layout quality score predicted by the Refiner network serves as a self-evaluation of the current layout's aesthetic quality by the Refiner network, guiding the optimization direction for the next round of layout optimization. This is the third layout result obtained after refinement. For visual cue images, For Refiner networks, This is a task instruction, derived based on user-input constraints. The layout of the input Refiner network is the second layout result after geometric normalization.

[0071] In this embodiment, the Refiner network is derived from the initial layout estimation model. Initialization is performed to maintain semantic and spatial consistency, and the alignment, non-overlap, and visual balance of the layout are optimized in a multi-step evolution process, so that the generated result achieves higher aesthetic harmony without destroying the semantic structure.

[0072] By employing a Refiner network to optimize the overall layout, the fine-tuning process of a human designer was simulated, maintaining aesthetic consistency.

[0073] S8. Based on the third layout result, combine the decorative layer with the cropping image to generate the target image.

[0074] The specific combination process is similar to that of S4, and will not be repeated here.

[0075] In summary, the method of this invention can automatically rearrange the composition under multi-element and complex backgrounds, maintaining compositional balance and visual integrity, and ultimately quickly generating finished images of multiple sizes with unified style and visual harmony. By first using PosterO to generate an initial layout result, and then using a Refiner network to refine this initial layout result, its progressive evolutionary mechanism optimizes alignment, spacing, balance, and visual weight distribution in the continuous latent space, thereby significantly improving the professional aesthetics and design rationality of the composition while maintaining semantic consistency.

[0076] Figure 3 This is a structural block diagram of the material expansion, cropping, and beautification system in an embodiment of the present invention.

[0077] In a second aspect, the present invention also provides a material expansion, cropping, and enhancement system. For example... Figure 3 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the material expansion, cropping, and enhancement method described in the first aspect of this invention.

[0078] The system also includes other components well known to those skilled in the art, such as communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.

[0079] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.

[0080] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise explicitly specified. Furthermore, the division of steps in the above method is only for clarity of description; in implementation, it can be combined into one step or some steps can be split into multiple steps, as long as they include the same logical relationship. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0081] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

Claims

1. A method for material augmentation, cropping and beautification, characterized in that, include: Obtain the target size and the corresponding layers of the target image material; the layers include a background image, a main image, and decorative layers; The background image is input into a preset image expansion model for expansion and filling to obtain the target background image; Input the main element information of the target background image and the main image into the preset layout generation model to obtain the first layout result; Based on the first layout result, the main image and the target background image are combined to obtain a combined image; The combined images are cropped according to the target size to obtain the cropped image; Input the decorative element information of the clipping image and decorative layer into the layout generation model to obtain the second layout result; Based on the cropping image, a pre-set refined network is used to optimize the second layout result to obtain the third layout result; Based on the third layout result, the decorative layer is combined with the cropped image to generate the target image.

2. The stock augmentation cropping and beautification method of claim 1, wherein, After obtaining the cropped image, the method further includes: softening the cropped image using a preset image coordination model.

3. The stock augmentation cropping and beautification method of claim 1, wherein, Cropping the combined image based on the target size includes: cropping the combined image using S2CNet based on the target size.

4. The stock augmentation cropping and beautification method of claim 1, wherein, The image expansion model is a ProOut model. The background image is input into the preset image expansion model for expansion filling, including: Control signals are generated by the composition planning module; The control signal is injected into the intermediate feature layer of the latent space diffusion network to guide the generation of local expansion and obtain the local expansion result of the current iteration. The local expansion result of the current iteration is pasted back to the corresponding position to update the global intermediate expansion image; the global intermediate expansion image is generated step by step through multiple progressive expansion iterations with the background image as the initial state.

5. The stock augmentation cropping and beautification method of claim 4, wherein, The expression for calculating the control signal is: In the formula, For control signals, This represents the intermediate feature mapping of the current diffusion network U-Net. Zero convolutional layer, To control the encoder, Features that provide compositional hints For global semantic conditions, , , These are all learnable parameters in the control encoder. For time steps.

6. The material expansion, cutting, and beautification method according to claim 5, characterized in that, The calculation expression for composition cue features is: In the formula, For target masking, For reference image, It is a shallow convolutional layer. This is for splicing operations.

7. The material expansion, cropping, and beautification method according to claim 1, characterized in that, The refined network is the Refiner network in SEGA; The second layout result is optimized using a pre-defined refined network, including: Based on the second layout result, render the elements corresponding to the decorative layer onto the clipping image to generate a visual cue image; The visual cue image, the geometrically normalized second layout result, and the corresponding task instructions are input into the Refiner network to obtain the third layout result and the layout quality score.

8. The material expansion, cutting, and beautification method according to claim 1, characterized in that, Input the main element information of the target background image and the main image into the preset layout generation model to obtain the first layout result, including: Extract design intent information from the target background image; Retrieve similar example layouts from a pre-defined database; The main element information, design intent information, and example layout are combined to form a Prompt; Input the Prompt into the LLM model to generate a structured layout tree containing element spatial attributes and hierarchical relationships, thus obtaining the first layout result.

9. The material expansion, cutting, and beautification method according to claim 1 or 8, characterized in that, The layout generation model is PosterO.

10. A material expansion, cutting, and enhancement system, characterized in that, It includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the material expansion, cropping and beautification method according to any one of claims 1-9 is implemented.