Image layout adjustment method and device, model training method and device and computing equipment
Through feature coding and noise reduction processing of images and materials, adaptive adjustment of image layout is achieved, and the problem that fixed templates cannot adapt to them is solved. The generated images are both beautiful and can effectively convey information, improving work efficiency and design quality.
Patent Information
- Application Number
- CN202510292235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, fixed templates cannot be adaptively adjusted according to different content and context, resulting in the generated posters not being able to achieve ideal visual effects or convey expected information in many cases, increasing labor costs and reducing work efficiency.
通过获取初始图像和目标素材的初始布局信息,进行视觉特征和布局信息特征的编码,基于这些特征执行降噪处理,获得目标素材的目标布局信息,并将目标素材添加至初始图像上,生成目标图像。
Adaptive adjustment of image layout is achieved, and the generated target images are more accurate and conform to the context, both beautiful and effectively convey information, reducing the need for manual adjustment, reducing labor costs, and improving work efficiency and design quality.
Smart Images

Figure CN120198528A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of artificial intelligence, and particularly to an image layout adjustment, model training method, device, and computing device. Background Art
[0002] With the continuous growth of digital content creation and personalized needs, the design of image layout has become increasingly important. Users' requirements for visual communication are increasing day by day. They not only focus on aesthetics but also pay more attention to the effective transmission of information and user experience.
[0003] Currently, by using a fixed material template to automatically add materials to an image and complete the image layout, standardized images can be generated, which can meet the needs of rapid design to a certain extent. However, the fixed template cannot be adaptively adjusted according to different contents and context environments, resulting in the generated posters not achieving the ideal visual effect or conveying the expected information in many cases. The fixity and lack of flexibility of the template make it not suitable for the design requirements in all scenarios. In the face of diverse application scenarios, since the fixed template is difficult to meet the complex and changeable design requirements, manual intervention is often required to adjust or re-design the template. This not only increases the labor cost but also reduces the work efficiency. Therefore, there is an urgent need for an efficient image layout adjustment method. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide an image layout adjustment method. One or more embodiments of this specification also relate to a model training method, an image layout adjustment device, a model training device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, an image layout adjustment method is provided, including:
[0006] Obtain the initial image and the initial layout information of the target material;
[0007] Encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features;
[0008] Perform noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material;
[0009] Based on the target layout information, add the target material to the initial image to obtain a target image.
[0010] According to the second aspect of the embodiments of this specification, a model training method is provided, including:
[0011] Obtain a sample image, multiple sample materials, and the sample material categories of the multiple sample materials;
[0012] Extract a first sample material from the multiple sample materials;
[0013] Generate the sample layout information of the first sample material based on the sample material category and the sample image of the first sample material;
[0014] Based on the sample layout information, add the first sample material to the sample image to obtain an updated sample image;
[0015] Using the updated sample image and the sample material category of the first sample material as input, and the sample layout information of the first sample material as the label output, train a vision-language model. Return to the step of extracting the first sample material from the multiple sample materials, and until training is completed with the multiple sample materials, obtain a trained target layout model.
[0016] According to the third aspect of the embodiments of the present specification, there is provided an image layout adjustment device, including:
[0017] A first acquisition module, configured to acquire an initial image and the initial layout information of a target material;
[0018] A first encoding module, configured to encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features;
[0019] A first noise reduction module, configured to perform noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material;
[0020] A first addition module, configured to add the target material to the initial image based on the target layout information to obtain a target image.
[0021] According to the fourth aspect of the embodiments of the present specification, there is provided a model training device, including:
[0022] A second acquisition module, configured to acquire a sample image, multiple sample materials, and the sample material categories of the multiple sample materials;
[0023] A second extraction module, configured to extract a first sample material from the multiple sample materials;
[0024] A second generation module, configured to generate the sample layout information of the first sample material based on the sample material category and the sample image of the first sample material;
[0025] A second addition module, configured to add the first sample material to the sample image based on the sample layout information to obtain an updated sample image;
[0026] The second training module is configured to take the updated sample image and the sample material category of the first sample material as inputs, and the sample layout information of the first sample material as the output label, train the vision-language model, and return to execute the step of extracting the first sample material from multiple sample materials. Until the case where multiple sample materials are trained, a trained target layout model is obtained.
[0027] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, including:
[0028] A memory and a processor;
[0029] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above image layout adjustment method or model training method are implemented.
[0030] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above image layout adjustment method or model training method are implemented.
[0031] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above image layout adjustment method or model training method are implemented.
[0032] In an embodiment of the present specification, an initial image and the initial layout information of a target material are obtained; the initial image is encoded to obtain visual features, and the initial layout information is encoded to obtain layout information features; based on the visual features, noise reduction processing is performed on the layout information features to obtain the target layout information of the target material; based on the target layout information, the target material is added to the initial image to obtain a target image.
[0033] Regarding the initial layout information as the result of the ideal layout information plus noise, on this basis, noise reduction processing is performed on the layout information features of the initial layout information based on the extracted visual features, realizing the adaptive adjustment of the element layout of the target elements, so as to obtain more accurate and context-compliant target material layout information. Based on the target layout information, the target material is added to the initial image to generate a target image that is both beautiful and can effectively convey information. This not only reduces the need for manual adjustment, lowers the labor cost, but also significantly improves the work efficiency and design quality, ensuring that image works that meet user needs can be efficiently created in various application scenarios, achieving a flexible and efficient image layout solution while satisfying visual aesthetics and information transmission efficiency. Description of the Drawings
[0034] Figure 1 is a flowchart of an image layout adjustment method provided by an embodiment of this specification;
[0035] Figure 2 is one of the schematic diagrams of the process of an image layout adjustment method provided by an embodiment of this specification;
[0036] Figure 3 is another schematic diagram of the process of an image layout adjustment method provided by an embodiment of this specification;
[0037] Figure 4 is a flowchart of a model training method provided by an embodiment of this specification;
[0038] Figure 5 is a schematic diagram of the process of a model training method provided by an embodiment of this specification;
[0039] Figure 6 is a flowchart of the processing procedure of an image layout adjustment method applied to commercial poster generation provided by an embodiment of this specification;
[0040] Figure 7 is a schematic diagram of the structure of an image layout adjustment device provided by an embodiment of this specification;
[0041] Figure 8 is a schematic diagram of the structure of a model training device provided by an embodiment of this specification;
[0042] Figure 9 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0043] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0044] The terms used in one or more embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present invention. The singular forms "a", "the" and "said" used in one or more embodiments of the present invention are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more of the associated listed items.
[0045] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0046] In addition, it should be noted that the data involved in one or more embodiments of the present invention are all information and data that have been authorized by the user or fully authorized by all parties, and the statistics, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0047] First, the noun terms involved in one or more embodiments of this specification are explained.
[0048] Diffusion Model: A probability-based generative model that generates new data by gradually adding noise to the data and learning the inverse process. The diffusion model mimics the diffusion noise reduction process in physics, where information gradually spreads from one point to the entire system. The core of the diffusion model lies in its ability to learn how to reverse a process that gradually destroys the data structure, and by denoising the noisy data, recover clear and ideal data.
[0049] Visual Diffusion Model: A variant of the diffusion model, specifically designed to process and generate image data. Similar to the traditional diffusion model, the visual diffusion model adds noise to the original image through a series of steps and learns how to reverse this process to generate new or reconstruct existing images. This model uses deep learning techniques, especially convolutional neural models, to capture and reproduce complex visual features. The visual diffusion model has been proven to have significant effects in fields such as image synthesis, super-resolution, and restoration.
[0050] Transformer Model: A deep learning architecture that is particularly suitable for processing sequence data, such as natural language and time series analysis. Different from previous methods that relied on recurrent neural models, the Transformer model introduces a self-attention mechanism, enabling the model to process the input sequence in parallel, greatly improving the training efficiency and performance. Its core components include the multi-head attention mechanism and the feed-forward neural model layer, and these designs allow the model to efficiently capture long-range dependencies within the sequence, becoming one of the core technologies in the fields of modern natural language processing and computer vision.
[0051] Cross-Attention Mechanism: A technique used in multimodal learning that allows for information interaction between data of different modalities (such as text and images). By calculating the similarity between queries from one modality and key-value pairs from another modality, the cross-attention mechanism can dynamically focus on relevant parts, thus achieving effective fusion between modalities.
[0052] ResNet model (Residual Network, abbreviated as ResNet): A deep convolutional neural model architecture designed to address the problem of vanishing or exploding gradients during the training of deep networks. By introducing "residual learning units", that is, adding a shortcut path directly connecting the input and output in each layer, the network can easily learn the identity mapping, thus effectively alleviating the performance degradation problem caused by the increase in network depth. This design allows ResNet to build deeper network structures (such as 152 layers) than before, while maintaining or even improving the performance of the model. ResNet has not only achieved excellent results in image classification tasks but also been widely applied in multiple computer vision fields such as object detection and segmentation, becoming one of the important bases for modern deep learning model design.
[0053] Vision Transformer model (VisionTransformer, abbreviated as ViT): A vision model based on the Transformer architecture designed to process and understand image data. Different from traditional convolutional neural models, ViT divides an image into multiple small patches, each of which is regarded as part of a word sequence and processed by the Transformer encoder. This method breaks the limitation of traditional convolutional neural models that rely on local receptive fields, enabling ViT to capture global information and enhancing the ability to understand image content. Although ViT requires a large amount of data and computing resources for training, it demonstrates strong transfer learning ability and performs well in various vision tasks, including image classification, object detection, and image segmentation. The emergence of ViT marks the successful extension of the Transformer architecture from the natural language processing field to the computer vision field.
[0054] Large Language Model (LM): A language processing model with a large number of parameters that can be trained on large-scale text datasets. Such models are usually based on the Transformer architecture and can learn rich language structures and semantic information. Due to their large scale and complexity, large language models perform well in understanding context, generating coherent text, answering questions, etc., and are widely used in many fields such as automatic translation, content creation, and dialogue systems. As the model scale increases, their performance also improves accordingly, but at the same time, it brings higher computational costs and resource requirements.
[0055] Vision-Language Model (VLM): Combines the ability to process visual information and natural language understanding, aiming to achieve effective processing of mixed information containing text and images. Such models usually combine information from the two modalities through a shared latent space, enabling the model to not only understand the content of individual images or texts but also the relationships between them. Application scenarios include image caption generation, visual question answering, and cross-modal retrieval. The development of vision-language models has promoted the progress of multimedia information processing, enabling machines to more comprehensively understand the rich content forms created by humans.
[0056] Prompt: An input provided by the user to guide the model to generate a specific type of output or response. The prompt can be a sentence, a question, or a set of instructions, with the aim of inspiring the model to think or create in a preset direction. In the application of large language models, well-designed prompts can help users effectively obtain the required information or guide the model to complete tasks such as text generation, creative writing, and code writing.
[0057] Currently, users rationally layout various types of materials (such as text, Logo, background, etc.) according to the specific content of the image to generate an image in which all materials coexist harmoniously, are in reasonable positions, and meet the aesthetic requirements. This method usually requires a large amount of design resources to ensure the aesthetics of the result. By producing various templates and combining the materials with the image, the final image is formed. Although the efficiency is improved to a certain extent, the fixed templates cannot adapt to all types of images, so sometimes additional resources still need to be invested for adjustment to achieve the ideal effect.
[0058] However, the limitations and lack of flexibility of fixed templates make them unsuitable for design requirements in all scenarios. When faced with diverse application scenarios, since fixed templates are difficult to meet complex and changing design requirements, manual intervention is often required to adjust or redesign the templates, which not only increases labor costs but also reduces work efficiency. To solve this problem, this solution proposes an image layout adaptive adjustment technology, aiming to model the key factors in the layout adjustment process to achieve intelligent and adaptive position adjustment based on the initial layout and image content, so as to meet the needs of efficient design and improve the effectiveness of information transmission and user experience.
[0059] In this specification, an image layout adjustment method is provided. This specification also relates to a model training method, an image layout adjustment device, a model training device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0060] See Figure 1 , Figure 1 shows a flowchart of an image layout adjustment method provided by an embodiment of this specification, including the following specific steps:
[0061] Step 102: Obtain the initial layout information of the initial image and the target material.
[0062] The embodiments of this specification are applied to applications, websites, or system platforms with image layout adjustment functions. For example, on community content applications, commercial posters of merchants can be designed and produced. Another example is that on community content applications, multimedia content of creators can be designed and generated. Still another example is that on community content applications, users' photos can be designed and beautified.
[0063] The initial image is the original image file that needs to be adjusted for image layout, and is used as the subsequent visual modality input. The initial image completes the adaptive adjustment of the layout information as reference content, and the initial image contains the basic visual information to be combined with the target material, such as posters, photographic photos, illustrations, or design drafts.
[0064] The target material is the image element that needs to be adjusted for image layout. The target material has not been added to the initial image yet. The target material includes but is not limited to text, Logo, background (background image), etc. The target material is used to enhance or modify the information transmission effect of the image and needs to be coordinated with the initial image to ensure that the output target image is both beautiful and effective.
[0065] The layout information of the target material is a parametric description of the geometric and / or style attributes of the target material in the image space. The layout information guides how to reasonably layout the target material on the image without destroying the overall aesthetics. The layout information of the target material usually takes into account visual factors such as visual focus, balance, and reading order.
[0066] The initial layout information is a parametric description of the geometric and / or style attributes of the target material in the image space before the image layout is adjusted. The initial layout information is not yet the ideal layout information and needs to be adaptively adjusted, which is the object of adjustment of the visual diffusion model. The initial layout information can be regarded as the result of superimposing a small amount of noise on the ideal layout information. From this perspective, the process of generating the target layout information from the initial layout information is very similar to the process of diffusion denoising. Therefore, it is considered to use diffusion denoising processing to model this process. The initial layout information can be determined by a template, generated by a target layout model, or manually input through a front-end interactive editing tool, which is not limited here.
[0067] To obtain the initial image, one optional way is to receive the initial image sent by the front end. Another optional way is to query the initial image from a database, where the database includes a historical database and / or an open-source database. Another optional way is to generate the initial image through an image generation model, which can be to generate the initial image based on a reference image or based on an image description text, which is not limited here.
[0068] To obtain the initial layout information of the target material, one optional way is to receive the initial layout information of the target material sent by the front end. Another optional way is to obtain the initial layout information of the target material on the template matching the initial image. Another optional way is to generate the initial layout information of the target material based on the initial image through a target layout model, which is not limited here.
[0069] Exemplarily, in a community content application, a certain merchant hopes to promote a certain product by posting product multimedia content and product posters. First, upload a photographic image containing the product main body as the initial image, and then select the text and brand logo to be added as the target material. By analyzing the composition center of gravity and color distribution of the product image, automatically call the template applicable to the electronic product poster to generate the initial layout information: the initial layout information of the text is (10px, 10px, 30px, 30px), and the initial layout information of the brand logo is (20px, 25px, 30px, 40px).
[0070] Obtaining the initial image and the initial layout information of the target material provides reference content and an adjustment object for subsequent adaptive adjustment.
[0071] Step 104: Encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features.
[0072] The visual features are semantic features representing the image content in the initial image, manifested as an encoded vector of a certain dimension, including but not limited to: object detection box confidence distribution, frequency domain features of color histograms, visual saliency heatmaps, etc.
[0073] The layout information features are semantic features representing the geometric attributes and / or style attributes of the target material in the image space, manifested as an encoded vector of a certain dimension, including but not limited to: coordinate discretization features, relative position matrices, hierarchical topology features, etc.
[0074] Exemplarily, extract the global semantic features of the initial image, and by capturing multi-scale visual context information, output a visual feature vector with a dimension of (H×W×C), where H and W are the feature map sizes, and C is the number of channels. Perform word vector embedding on the initial layout information, and inject geometric space information using positional encoding to generate layout information features with dimensions aligned with the visual features.
[0075] Encoding the initial image to obtain visual features and encoding the initial layout information to obtain layout information features lay a feature foundation for subsequent noise reduction processing.
[0076] Step 106: Perform noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material.
[0077] The target layout information is a parametric description of the geometric attributes and / or style attributes of the target material in the image space after image layout adjustment. Compared with the initial layout information, the target layout information can be considered to be closer to the ideal layout information.
[0078] Exemplarily, perform noise reduction processing on the layout information features based on the visual features using a cross-attention mechanism to obtain the target layout information of the target material: the target layout information of the text is (12px, 13px, 50px, 35px), and the target layout information of the brand Logo is (24px, 30px, 36px, 48px).
[0079] Performing noise reduction processing on the layout information features of the initial layout information based on the extracted visual features realizes the adaptive adjustment of the element layout of the target elements, thereby obtaining more accurate target material layout information that conforms to the context environment, providing more accurate information support for subsequent adding the target material to the initial image.
[0080] Step 108: Add the target material to the initial image based on the target layout information to obtain the target image.
[0081] The target image is an optimized image file with the image layout adjusted. Compared with the initial image, the target image has the following features: 1. Spatial rationality, that is, the positions of the materials conform to the distribution of visual focuses and the reading flow line; 2. Style consistency, such as the font, color, and material matching the aesthetic features of the original image; 3. Information transmission efficiency, that is, the visual weight of key information (such as promotional text) is significantly higher than that of auxiliary information.
[0082] Based on the target layout information, add the target materials to the initial image to obtain the target image. One optional method is: based on the target layout information, render the target materials onto the initial image to obtain the target image. Another optional method is: based on the target layout information, adjust the layout information of the target materials on the template, and then add the target materials to the initial image to obtain the target image. Another optional method is: through an image generation model, based on the target layout information, add the target materials to the initial image to generate the target image. This is not limited here.
[0083] Exemplarily, call the rendering engine based on the target layout information (the target layout information of the text is (12px, 13px, 50px, 35px), and the target layout information of the brand logo is (24px, 30px, 36px, 48px)), render the text and the brand logo onto the photographic image containing the main body of the product to obtain a product poster. Edit the product multimedia content according to the product poster, publish it to the community content application, and submit the product poster to the operator of the community content application to display the product poster on the splash screen page and the home page recommendation of the community content application.
[0084] In the embodiments of this specification, the initial layout information is regarded as the result of the ideal layout information plus noise. On this basis, perform noise reduction processing on the layout information features of the initial layout information based on the extracted visual features, realizing the adaptive adjustment of the element layout of the target elements, so as to obtain more accurate target material layout information that conforms to the context environment. Based on the target layout information, add the target materials to the initial image to generate a target image that is both beautiful and can effectively convey information, which not only reduces the need for manual adjustment, lowers the labor cost, but also significantly improves the work efficiency and design quality, ensuring that image works that meet the user's needs can be efficiently created in various application scenarios, realizing a flexible and efficient image layout solution while meeting visual aesthetics and information transmission efficiency.
[0085] In an optional embodiment of this specification, step 104 includes the following specific steps:
[0086] Through the encoding module of the visual diffusion model, encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features;
[0087] Correspondingly, step 106 includes the following specific steps:
[0088] Through the feature processing module of the visual diffusion model, perform noise reduction processing on the layout information features based on visual features to obtain the target layout information of the target material.
[0089] The visual diffusion model is a visual diffusion model of a generative architecture that integrates image content perception and layout optimization capabilities, including an encoding module and a feature processing module.
[0090] The encoding module is a module in the visual diffusion model that vectorially encodes multi-modal input data, including a visual encoding unit and a text encoding unit. Optionally, the text encoding unit includes a layout encoding unit and / or a category encoding unit.
[0091] The feature processing module is a module in the visual diffusion model that performs noise reduction processing on the encoded features based on the diffusion noise reduction process. The feature processing module realizes the dynamic interaction between visual features and layout information features through the cross-attention mechanism. The feature processing module gradually eliminates the noise interference on the layout information features in the diffusion time step. The feature processing module injects constraint conditions to realize the conditional control of prior knowledge such as categories and hierarchical relationships.
[0092] Through the encoding module of the visual diffusion model, encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features. An optional method is: through the encoding module of the visual diffusion model, perform multi-modal joint encoding on the initial image and the initial layout information to obtain visual features and layout information features. Another optional method is: through the encoding module of the visual diffusion model, perform discrete encoding on the initial image and the initial layout information to obtain discrete visual features and discrete layout information features, which is not limited here.
[0093] Exemplarily, the visual diffusion model includes an encoding module and a feature processing unit. The encoding module includes a visual encoding unit and a text encoding unit. The visual encoding unit uses a pre-trained visual backbone network (such as ResNet-50, ViT-B / 16), and the text encoding unit uses the embedding layer of the Transformer model. The initial image is input into the visual encoding unit to extract the global semantic features of the initial image, and by capturing multi-scale visual context information, a visual feature vector is output. The initial layout information (the initial layout information of the text is (10px, 10px, 30px, 30px), and the initial layout information of the brand logo is (20px, 25px, 30px, 40px)) is input into the text encoding unit to perform word vector embedding on the initial layout information, inject geometric space information using a positional encoding layer (Positional Encoding), and generate layout information features through a fully connected network, with the dimension aligned with the visual features.
[0094] Through the feature processing module of the visual diffusion model, noise reduction processing is performed on the layout information features based on the visual features to obtain the target layout information of the target material. An optional method is: through the feature processing module of the visual diffusion model, noise reduction processing of the cross-attention mechanism is performed on the layout information features based on the visual features to obtain the target layout information of the target material.
[0095] Exemplarily, the feature processing module of the visual diffusion model adopts a feature processing unit composed of the encoders and decoders of multiple Transformer models. The visual features and layout information features are input into each feature processing unit, and noise reduction processing of the cross-attention mechanism is performed on the layout information features based on the visual features to obtain the target layout information of the target material: the target layout information of the text is (12px, 13px, 50px, 35px), and the target layout information of the brand logo is (24px, 30px, 36px, 48px).
[0096] In the embodiments of this specification, by using the layout information features as query features and performing multi-level cross-attention interaction with the visual features, the deep coupling between the image semantic information and the layout adjustment requirements is realized. The visual diffusion model can dynamically capture the visually significant regions related to the target material in the initial image, and perform step-by-step noise reduction optimization on the layout parameters according to the visual focus distribution. During the noise reduction process, the model iteratively corrects the position offset and size error of the initial layout, so that the layout of the target material meets the adaptive requirements, thereby achieving excellent design effects in complex composition scenarios.
[0097] In an optional embodiment of this specification, step 106 includes the following specific steps:
[0098] Using the layout information feature as the query feature, and the visual feature as the key feature and value feature, perform cross-attention calculation to obtain the attention-weighted feature, and perform noise reduction processing on the attention-weighted feature to obtain the target layout information of the target material.
[0099] The query feature is a feature input used to complete the similarity matching with the key feature in the attention mechanism. It comes from the layout information feature and represents the current spatial distribution state of the target material.
[0100] The key feature is another feature input used to complete the similarity matching with the query feature in the attention mechanism. It comes from the visual feature and represents the semantic and spatial attributes of the image content of the initial image.
[0101] The value feature is a feature input used to complete the contribution to the output attention-weighted feature. It comes from the visual feature and represents the semantic and spatial attributes of the image content of the initial image.
[0102] By calculating the similarity between the query feature and the key feature, determine the region in the visual feature that is highly correlated with the layout information feature to complete the weighted processing of the value feature.
[0103] The attention-weighted feature is the feature after being processed by the cross-attention mechanism, which contains the result obtained by weighting the visual feature of the initial image according to the layout requirements of the target material. Performing attention calculation at different scales can effectively capture local details and global composition at different scales.
[0104] The calculation formula for cross-attention calculation is shown in Formula 1:
[0105]
[0106] where, Q is the query feature, K is the key feature, V is the value feature, is the feature dimension scaling factor.
[0107] Using the layout information feature as the query feature, and the visual feature as the key feature and value feature, perform cross-attention calculation to obtain the attention-weighted feature, and perform noise reduction processing on the attention-weighted feature to obtain the target layout information of the target material. An optional way is: through the feature processing module of the visual diffusion model, using the layout information feature as the query feature, and the visual feature as the key feature and value feature, perform cross-attention calculation to obtain the attention-weighted feature, and perform noise reduction processing on the attention-weighted feature to obtain the target layout information of the target material.
[0108] Exemplarily, the layout information feature and the visual feature are input into the Transformer feature processing unit. Using the layout information feature as the query feature Q, and the visual feature as the key feature K and the value feature V, cross-attention calculation is performed using Equation 1 to obtain the attention-weighted feature Attention, and noise reduction processing is performed on the attention-weighted feature:
[0109] Layout information of the text:
[0110] Horizontal displacement Δx = +2px, Δx = +3px → new x coordinate = 12px, 13px;
[0111] Vertical displacement Δy = +20px, Δy = +5px → new y coordinate = 50px, 35px.
[0112] Target layout information of the brand Logo
[0113] Horizontal displacement Δx = +4px, Δx = +5px → new x coordinate = 24px, 30px;
[0114] Vertical displacement Δy = +6px, Δy = +8px → new y coordinate = 36px, 48px.
[0115] Obtain the target layout information of the target material: The target layout information of the text is (12px, 13px, 50px, 35px), and the target layout information of the brand Logo is (24px, 30px, 36px, 48px).
[0116] In the embodiments of this specification, through the cross-attention mechanism, the semantic features of the initial image can be perceived, and the semantic information and spatial layout of the initial image can be understood more deeply, so as to achieve a more accurate adaptive adjustment of the target material, which not only improves the ability to understand the image content, but also can dynamically adjust the layout of each material according to specific design requirements.
[0117] In an alternative embodiment of this specification, before step 106, the following specific steps are further included:
[0118] Obtain the material category of the target material;
[0119] Encode the material category to obtain the category information feature;
[0120] Correspondingly, step 106 includes the following specific steps:
[0121] Under the constraint of the category information feature, perform noise reduction processing on the layout information feature based on the visual feature to obtain the target layout information of the target material.
[0122] The material category of the target material is a semantic type division of the functional attributes and layout constraints of the target material, covering the following dimensions: Content type: text, Logo, background image, decorative elements, etc.; Interaction priority: core information (such as promotional text), auxiliary information (such as disclaimers); Visual weight: high-contrast elements (such as warning signs), low-disturbance elements (such as watermarks); Dynamic attributes: fixed-position elements (such as brand logos), adaptive elements (such as dynamic advertising slogans). Prior knowledge is injected through the material category to implement different differential layout strategies for different material categories.
[0123] The category information feature is a semantic feature that characterizes the functional attributes and layout constraint types of the target material, manifested as a coding vector in a certain dimension, including but not limited to: layout constraint features, semantic matching degree with image content, and style control parameters. For layout constraint features, such as spatial constraints (text cannot be covered, Logo icons allow partial occlusion), size elasticity (such as the substrate can be stretched, icons need to maintain the original ratio), for the semantic matching degree with image content, such as the brand Logo and the scene text are within a certain distance, for style control parameters, font / color suggestions (such as dark backgrounds match light text), effect parameters (such as the displacement range of floating animations).
[0124] One optional way to obtain the material category of the target material is: receive the material category of the target material sent by the front end. Another optional way is: obtain the material category of the target material on the template matching the initial image. Another optional way is: through the target layout model, based on the initial image, generate the material category of the target material, which is not limited here.
[0125] To encode the material category and obtain the category information feature, one optional way is: through the encoding module of the visual diffusion model, encode the material category to obtain the category information feature. Further, through the encoding module of the visual diffusion model, encode the material category to obtain the category information feature. One optional way is: through the encoding module of the visual diffusion model, perform multi-modal joint encoding on the initial image, initial layout information, and material category to obtain visual features, layout information features, and category information features. Another optional way is: through the encoding module of the visual diffusion model, perform discrete encoding on the material category to obtain discrete category information features, which is not limited here.
[0126] Under the constraint of the category information feature, noise reduction processing is performed on the layout information feature based on the visual feature to obtain the target layout information of the target material. An optional method is as follows: through the feature processing module of the visual diffusion model, under the constraint of the category information feature, noise reduction processing is performed on the layout information feature based on the visual feature to obtain the target layout information of the target material. Further, through the feature processing module of the visual diffusion model, under the constraint of the category information feature, noise reduction processing is performed on the layout information feature based on the visual feature to obtain the target layout information of the target material. An optional method is as follows: through the feature processing module of the visual diffusion model, under the constraint of the category information feature, noise reduction processing with a cross-attention mechanism is performed on the layout information feature to obtain the target layout information of the target material.
[0127] Optionally, Figure 2 FIG. 5 shows one of the schematic flowcharts of an image layout adjustment method provided by an embodiment of this specification, as Figure 2 shown:
[0128] The initial image is input into the visual encoding unit to obtain visual features. The initial material category of the target material and the initial layout information of the target material, such as, text 10103030|logo 20253040, are input into the text encoding unit to obtain the category information feature and the layout information feature. After 12 sequentially connected Transformer feature processing units, the target material category of the target material and the target layout information of the target material, such as, text 12135035|logo 24303648, are obtained.
[0129] As Figure 2 shown in the model structure, during the process of the visual diffusion model correcting the initial layout information, since the initial material category is added to the diffusion noise reduction process, the material category of some materials may change due to the misadjustment of the visual diffusion model, resulting in an uncontrollable adjustment process. In the actual application process, since the material category is often fixed and the adjustment scheme does not want the material category to be changed, therefore, a more controllable modeling method needs to be adopted to keep the material category unchanged while introducing the category information feature. To address this problem, Figure 3 the control of the material category is completed by improving the model structure of the visual diffusion model.
[0130] Optionally, Figure 3 FIG. 6 shows another schematic flowchart of an image layout adjustment method provided by an embodiment of this specification, as Figure 3 shown:
[0131] Input the initial image into the visual encoding unit to obtain visual features. Input the initial layout information of the target material, such as 10103030|20253040, into the layout encoding unit to obtain layout information features. Input the material category of the target material, such as texttext texttext|logo logo logo logo, into the category encoding unit to obtain category information features.
[0132] In the visual diffusion model, a control network is used to introduce additional constraint conditions into the model, so that the model can perceive specific control content and guide the generated content to be closer to the constraint conditions. During the diffusion denoising process, the category information features of the material category are input as the supervision signal of the control network to guide each Transformer feature processing unit to complete the perception of the material category without changing the material category during the denoising generation process.
[0133] By Figure 3 The shown method, the visual diffusion model can realize while adaptively adjusting the layout information, perceiving the image content and the material category of the target material, and performing denoising processing on the layout information features based on the visual features under the constraint of keeping the material category unchanged, to obtain the target layout information of the target material.
[0134] Exemplarily, the encoding module includes a visual encoding unit and a text encoding unit, and the text encoding unit includes a category encoding unit. The feature processing module of the visual diffusion model uses a feature processing unit composed of the encoders and decoders of multiple Tranformer models. The category encoding unit uses the embedding layer of the Tranformer model. Input the material category (text, brand Logo) of the target material into the category encoding unit to encode the material category and obtain category information features. Input the visual features, layout information features, and category information features into the feature processing units of each Tranformer model respectively. Under the constraint of the category information features, perform denoising processing on the layout information features through the cross-attention mechanism based on the visual features to obtain the target layout information of the target material: the target layout information of the text is (12px, 13px, 50px, 35px), and the target layout information of the brand Logo is (24px, 30px, 36px, 48px).
[0135] In the embodiments of this specification, the category information features of multiple target materials are introduced, so as to realize the adaptive adjustment of the layout position on the premise of not changing the material category of the target material, reduce the misadjustment of the material category during the diffusion denoising process, and at the same time maintain the perception ability of the layout information of the target materials between different categories.
[0136] In an optional embodiment of this specification, step 104 includes the following specific steps:
[0137] Through the discrete layout encoding unit of the discrete vision diffusion model, the coordinate values in the initial layout information are divided into multiple preset intervals to generate the discrete distribution characteristics of the initial layout information;
[0138] Correspondingly, the material category is encoded to obtain the category information characteristics, including the following specific steps:
[0139] Through the discrete category encoding unit of the discrete vision diffusion model, the material category is divided into multiple preset categories to generate the discrete distribution characteristics of the material category;
[0140] Correspondingly, step 106 includes the following specific steps:
[0141] Under the constraint of the category information characteristics, noise reduction processing is performed on the layout information characteristics based on visual features, and the discrete distribution characteristics of the target layout information of the target material are output. Based on the discrete distribution characteristics, the target layout information of the target material is determined.
[0142] The discrete vision diffusion model is a variant of the vision diffusion model used to process visual inputs with discrete values. Different from the continuous-value vision diffusion model, the discrete vision diffusion model can simultaneously process continuous coordinate information and discrete material category information in an image. By abstracting continuous coordinate values into multiple preset intervals, the model can more flexibly adapt to the task requirements involving mixed data types. The discrete vision diffusion model solves the problem that the continuous-value vision diffusion model is difficult to uniformly handle the modality gap between continuous coordinates (such as Logo positions) and discrete categories (such as text / icon classification), and reduces the solution space dimension of the layout optimization problem through discretization, improving the model convergence speed.
[0143] The discrete layout encoding unit is the unit in the discrete vision diffusion model that performs discrete vector quantization encoding on the layout information. The discrete category encoding unit is the unit in the discrete vision diffusion model that performs discrete vector quantization encoding on the material category.
[0144] The multiple preset intervals are a set of standard layout coordinate ranges predefined in the discrete vision diffusion model. The multiple preset categories are a set of standard material category sets predefined in the discrete vision diffusion model. The multiple preset categories cover various possible material categories from text, Logo to background images, and are further subdivided according to the requirements of the actual application scenario.
[0145] The discrete distribution characteristics of the material category are discrete semantic characteristics representing the functional attributes and layout constraint types of the target material, manifested as an encoding vector of a certain dimension.
[0146] The discrete distribution feature of the initial layout information represents the layout probability distribution of the target material before optimization in the image space. The discrete distribution feature of the target layout information is the layout probability distribution of the target material after optimization in the image space generated through the noise reduction process of the discrete vision diffusion model. Specifically, it includes: 1. Coordinate probability matrix: The probability distribution corresponding to each coordinate dimension (x, y, w, h) in a preset interval; 2. Class retention probability: The probability value that the material class maintains the original label; 3. Joint distribution encoding: The joint probability encoding vector of coordinates and classes, with the dimension aligned with the model hidden layer.
[0147] Under the constraint of the class information feature, perform noise reduction processing on the layout information feature based on the visual feature, and output the discrete distribution feature of the target layout information of the target material. Based on the discrete distribution feature, determine the target layout information of the target material. An optional method is: through the feature processing module of the discrete vision diffusion model, under the constraint of the class information feature, perform noise reduction processing on the layout information feature based on the visual feature, and output the discrete distribution feature of the target layout information of the target material. Further, through the feature processing module of the discrete vision diffusion model, under the constraint of the class information feature, perform noise reduction processing on the layout information feature based on the visual feature, and output the discrete distribution feature of the target layout information of the target material. An optional method is: through the feature processing module of the vision diffusion model, under the constraint of the class information feature, perform noise reduction processing on the layout information feature based on the visual feature using the cross-attention mechanism, and output the discrete distribution feature of the target layout information of the target material.
[0148] Based on the discrete distribution feature, determine the target layout information of the target material. An optional method is: based on the discrete distribution feature, determine the target layout information of the target material by maximum probability selection. Another optional method is: based on the discrete distribution feature, determine the target layout information of the target material by weighted average. Another optional method is: based on the discrete distribution feature, determine the target layout information of the target material by multi-step iterative correction, which is not limited here.
[0149] Exemplarily, through the discrete category encoding unit of the discrete vision diffusion model, the material category is converted into discrete distribution features: First, the input category label is one-hot encoded and mapped to a preset classification system containing 8 categories (such as text, Logo, etc.). Subsequently, a discrete distribution feature with semantic relevance is generated through a learnable category embedding layer. This feature is cross-modally fused with the visual feature to form a conditional constraint in the joint semantic space. Through the feature processing module of the discrete vision diffusion model, under the constraint of the category information feature, denoising processing of the layout information feature is performed on the visual feature by executing the cross-attention mechanism: A three-layer cross-attention module with the layout feature as Q, the visual feature as K, and V is constructed. By dynamically calculating the correlation weight between the visual semantics and the layout position, the noise distribution in the coordinate probability matrix is iteratively corrected. After 3 denoising iterations, the discrete distribution feature of the target layout information of the target material is finally output. Based on these discrete distribution features, a hybrid decoding strategy is adopted to determine the coordinate values: For the x / y coordinate dimension, the candidate intervals with a confidence level exceeding the threshold are selected from the probability matrix for weighted averaging. For the width and height w / h dimensions, the median of the maximum probability interval is directly used. This decoding process also considers the category retention probability (for example, the retention probability of the text category reaches 92%, and the Logo category is 85%), and determines the new position of the text to be (12px, 13px, 50px, 35px), and the new position of the brand Logo to be (24px, 30px, 36px, 48px).
[0150] In the embodiments of this specification, by discretely encoding the initial layout information and combining the discrete distribution features of the material category, the spatial distribution and semantic attributes of the target material can be captured more accurately. By performing denoising processing through the cross-attention mechanism, the layout position can be dynamically optimized while maintaining the material category unchanged, thereby generating a more reasonable and beautiful image layout.
[0151] In an optional embodiment of this specification, under the constraint of the category information feature, denoising processing is performed on the layout information feature based on the visual feature to obtain the target layout information of the target material, including the following specific steps:
[0152] Through the first feature processing unit in the feature processing module of the vision diffusion model, under the constraint of the category information feature, denoising processing is performed on the layout information feature based on the visual feature, and the first layout information feature is output, where the feature processing module includes multiple sequentially connected feature processing units;
[0153] Through the second feature processing unit in the feature processing module, under the constraint of the category information feature, denoising processing is performed on the first layout information feature based on the visual feature, and the second layout information feature is output;
[0154] Until the last layout information feature output by the last feature processing unit in the feature processing module is obtained, and based on the last layout information feature, determine the target layout information of the target material.
[0155] The multiple sequentially connected feature processing units are units for implementing multi-level noise reduction processing in the visual diffusion model. Optionally, the multiple sequentially connected feature processing units are composed of the feature processing units of multiple stacked Transformer models. Each feature processing unit gradually optimizes the layout information feature through the cross-attention mechanism and residual connection. Each unit includes the following sub-units: Cross-modal attention layer: Using the layout information feature as the query feature Q, the visual feature as the key feature K and the value feature V, calculate the attention weight, and dynamically fuse the image semantics and layout space information; Conditional injection layer: Inject the category information feature into the attention-weighted feature to constrain the category consistency of the layout adjustment; Feed-forward network layer: Perform non-linear transformation on the feature through a multi-layer perceptron to enhance the model's ability to model complex layout relationships; Residual connection and layer normalization: Retain the residual path of the original layout information feature to avoid the problem of gradient disappearance.
[0156] Exemplarily, such as Figure 2 and Figure 3 The feature processing units of the multiple sequentially connected Transformer models shown. In the first feature processing unit, receive the initial layout information feature, visual feature and category information feature, capture the potential association between the main region of the image and the initial layout through the cross-modal attention layer, and generate the first layout information feature: The text coordinates are corrected from (10px, 10px) to (11px, 12px). In the second feature processing unit, using the first layout information feature as the input, combined with the higher-level semantic information in the visual feature, such as the edge contour of the commodity main body, further refine the layout offset, such as the text width is extended from 30px to 45px to adapt to the long title. The subsequent units sequentially perform multi-granularity optimization on the layout information feature, gradually eliminating coordinate noise, such as the Logo position iterates from (20px, 25px) to (24px, 30px) to avoid overlapping with background elements. The layout information feature output by the last feature processing unit is mapped to the target parameter space through a linear projection layer and decoded into the target layout information (12px, 13px, 50px, 35px).
[0157] In the embodiments of this specification, the layout information features are gradually optimized through multiple sequentially connected feature processing units, achieving highly refined processing of the image layout adjustment task. Each feature processing unit dynamically fuses visual features and layout information features through a cross-modal attention mechanism, and ensures a high degree of consistency in the material categories during the layout adjustment process through a conditional injection layer. This multi-level noise reduction processing can not only effectively remove the noise interference in the initial layout, but also gradually refine the layout information of the target material, making the finally output target layout information more accurate, beautiful and in line with the context environment.
[0158] In an alternative embodiment of this specification, step 102 includes the following specific steps:
[0159] Obtain the material category of the initial image and the target material;
[0160] Based on the initial image and the material category, generate the initial layout information of the target material through a pre-trained target layout model.
[0161] The target layout model is a deep learning model for automatically generating appropriate layout information according to the input image and material category. The target layout model can be a model based on the CLIP model or the Vision Transformer architecture, or a large language model or vision language model that combines natural language processing capabilities. The target layout model can capture rich visual semantic information and understand the layout rules of different material categories through pre-training on a large-scale dataset.
[0162] A method for generating the initial layout information of the target material based on the initial image and the material category through a pre-trained target layout model: input the initial image and the material category into the pre-trained target layout model to generate the initial layout information of the target material. Another alternative method: through a vision language model, generate the initial layout information of the target material based on the initial image under the guidance of the prompt information including the material category. This is not limited here.
[0163] Exemplarily, input the poster picture without layout into the vision language model. Under the guidance of the prompt information "Design the layout of a commercial poster according to the input poster picture: 1 text", the generated initial layout information of the text is (10px, 10px, 30px, 30px).
[0164] In the embodiments of this specification, through the pre-trained target layout model, the initial layout information is generated based on the material category of the image material, realizing efficient and accurate automated layout design, improving the quality and efficiency of image content creation. It can not only quickly generate a preliminary layout that meets the requirements of visual aesthetics and information transmission, but also significantly reduce the time cost and complexity of manual design.
[0165] In an alternative embodiment of this specification, there are multiple target materials, and there is a hierarchical relationship among the multiple target materials; through a pre-trained target layout model, based on the initial image and the material categories, the initial layout information of the target materials is generated, including the following specific steps:
[0166] Through the pre-trained target layout model, based on the initial image and the material categories of the target materials at the first level, the initial layout information of the target materials at the first level is generated;
[0167] Based on the initial layout information, add the target materials at the first level to the initial image to obtain the first initial image;
[0168] Through the target layout model, based on the first initial image and the material categories of the target materials at the second level, the initial layout information of the target materials at the second level is generated;
[0169] Based on the initial layout information, add the target materials at the second level to the first initial image to obtain the second initial image. Until all the multiple target materials are added, the initial layout information of the multiple target materials is obtained.
[0170] The hierarchical relationship among the multiple target materials affects the order of layout adjustment of the multiple target materials. For example, in the design of a commercial poster, the background may be at the lowest level, followed by the text layer, and important elements such as the brand logo are at the highest level. By setting such a hierarchical relationship, it can be ensured that key information is more prominent visually and will not be blocked or interfered with by other elements. The hierarchical relationship can also help the model understand the interaction and dependency relationships among different materials, so as to generate a more reasonable and coordinated layout.
[0171] An alternative way to generate the initial layout information of the target materials at the first level through the pre-trained target layout model based on the initial image and the material categories of the target materials at the first level is: input the initial image and the material categories of the target materials at the first level into the pre-trained target layout model to generate the initial layout information of the target materials at the first level. Another alternative way is: through a vision-language model, under the guidance of a prompt information including the material categories of the target materials at the first level, based on the initial image, generate the initial layout information of the target materials at the first level, which is not limited here.
[0172] Based on the first initial image and the material categories of the target materials at the second level, the initial layout information of the target materials at the second level is generated through the target layout model. An optional method is to input the first initial image and the material categories of the target materials at the second level into a pre-trained target layout model to generate the initial layout information of the target materials at the second level. Another optional method is to generate the initial layout information of the target materials at the second level based on the first initial image under the guidance of a visual language model with the prompt information including the material categories of the target materials at the second level. This is not limited here.
[0173] Exemplarily, an unlaid-out poster picture is input into the visual language model. Under the guidance of the prompt information "Design the layout of a commercial poster based on the input poster picture: 1 piece of text", the initial layout information of the text is generated. Then, the poster picture with the text material pasted on it is used as the input poster picture for the second generation and input into the visual language model. Under the guidance of the prompt information "Design the layout of a commercial poster based on the input poster picture: 1 brand logo", the initial layout information of the brand logo is generated again, obtaining the initial layout information of the two target materials, i.e., the text and the brand logo, on the product poster: the initial layout information of the text is (10px, 10px, 30px, 30px), and the initial layout information of the brand logo is (20px, 25px, 30px, 40px).
[0174] In the embodiments of this specification, through the iterative generation strategy of the target layout model, by completing the material layout in stages, the initial layout information is highly consistent with the image semantics, providing an efficient and reliable solution for automated design.
[0175] As described in the above embodiments, the target layout model can automatically generate the initial layout information. However, there are problems in the model training process such as high training data acquisition cost and large annotation cost, and the problem of insufficient training data needs to be solved.
[0176] Therefore, an iterative layout generation method is proposed. By generating layout information hierarchically and dynamically constructing sample data, it realizes the process of referring to the designer's image design, classifies the materials of the image, and adds materials to the image layer by layer. For example, a background is added first, and then text is added on this basis. In this way, the volume of the data set can be expanded. The data passing rate is significantly improved, and the sample data increases by a multiple level in terms of quantity. Through this method, the cost problem brought by the layout annotation data is alleviated. At the same time, the increase in the data volume can also be reflected in the number of parameters of the visual diffusion model. By increasing the number of parameters of the training model, the training effect of the model is improved. For the specific implementation method, see Figure 4 , Figure 4 shows a flowchart of a model training method provided by an embodiment of this specification, including the following specific steps:
[0177] Step 402: Obtain a sample image, multiple sample materials, and the sample material categories of the multiple sample materials.
[0178] The embodiments of this specification are applied to a system platform with functions of sample data construction and model training.
[0179] The sample image is an original image file for model training, which contains a blank design or a basic design without added layout elements and serves as a benchmark for layout adjustment.
[0180] The sample materials are a set of image elements to be added to the sample image, such as text, Logo, background image, etc. Each material needs to be labeled with its category and layout rules.
[0181] The sample material categories are the classifications of the functional attributes and layout constraints of the sample materials. For example, "promotion text" (to be highlighted), "brand Logo" (to be fixed in position), etc.
[0182] One optional way to obtain the sample image, multiple sample materials, and the sample material categories of the multiple sample materials is: from an open-source sample library, obtain the sample image, multiple sample materials, and the sample material categories of the multiple sample materials. Another optional way is: through a generation model, generate the sample image, multiple sample materials, and the sample material categories of the multiple sample materials. This is not limited here.
[0183] Exemplarily, from the database of a social content application, obtain 1000 product posters as sample images. The sample materials corresponding to each poster include promotion copy (category: text), brand logo (category: Logo), discount icon (category: decorative element), and generate material category labels through a semi-automatic annotation tool.
[0184] Obtaining the sample image, multiple sample materials, and the sample material categories of the multiple sample materials provides initial data support for subsequent training.
[0185] Step 404: Extract a first sample material from the multiple sample materials.
[0186] The first sample material is the sample material extracted during the iterative training process. It can be extracted according to the hierarchical relationship or randomly. This is not limited.
[0187] Exemplarily, perform stratified sampling based on the priority of material categories (such as text > Logo > decoration).
[0188] Extracting the first sample material from the multiple sample materials provides a data basis of the sample material for subsequent generation of the sample layout information of the first sample material.
[0189] Step 406: Generate the sample layout information of the first sample material based on the sample material category and sample image of the first sample material.
[0190] Since the vision-language model has multi-modal characteristics, its input includes pictures and text, and the output is text. This input-output format well adapts to the generation of layout information. At the same time, because the vision-language model has very good open-source pre-trained models, its pre-trained data includes sample images and has good language features for sample images. For the above reasons, using the vision-language model to train the local generation task has unique advantages, and the training method is the same as that of the vision-language model, directly modeling the sample layout information.
[0191] The sample layout information of the first sample material is a parametric description of the geometric attributes and / or style attributes of the first sample material in the image space. The sample layout information is usually represented in the form of coordinate frames, center points, bounding boxes, etc., and is used to guide the precise placement of the sample material on the image. For example, for materials of the "promotion text" category, the layout information may include the starting position of the text, font size, alignment method, etc.; for materials of the "brand logo" category, the layout information may include the center point coordinates and scaling ratio of the logo.
[0192] Based on the sample material category and sample image of the first sample material, to generate the sample layout information of the first sample material, an optional method is: through the vision-language model, based on the sample material category and sample image of the first sample material, generate the sample layout information of the first sample material. Further, based on the sample material category and sample image of the first sample material, to generate the sample layout information of the first sample material, an optional method is: through the vision-language model, under the guidance of the prompt information including the sample material category of the first sample material, based on the sample image, generate the sample layout information of the first sample material.
[0193] Exemplarily, the first sample material is "text". The vision-language model generates the layout information "Place the text in the center at the top of the image, font size 24px, color red" based on the content of the sample image (such as the product display area) and the category prompt of "text".
[0194] Based on the sample material category and sample image of the first sample material, to generate the sample layout information of the first sample material, by taking the sample image and material category as inputs, the vision-language model can understand the semantic relationship between the image content and the material function, thereby generating accurate sample layout information.
[0195] Step 408: Add the first sample material to the sample image based on the sample layout information to obtain the updated sample image.
[0196] The updated sample image is the image obtained after adding the first sample material to the sample image according to the generated layout information. This image retains the basic design of the original sample image while adding the layout elements of the first sample material, providing a new benchmark for the addition of subsequent materials.
[0197] Based on the sample layout information, add the first sample material to the sample image to obtain the updated sample image. One optional method is: based on the sample layout information, render the first sample material onto the sample image to obtain the updated sample image. Another optional method is: based on the sample layout information, adjust the layout information of the first sample material on the template, add the first sample material to the sample image to obtain the updated sample image. Another optional method is: through an image generation model, based on the sample layout information, add the first sample material to the sample image to generate the updated sample image, which is not limited here.
[0198] Exemplarily, for a sample image of a product poster, the first sample material is "text", and its sample layout information is "place the text in the center at the top of the image, with a font size of 24px and a color of red". Use an image processing tool to render the text onto the sample image according to the sample layout information to generate the updated sample image.
[0199] Based on the sample layout information, add the first sample material to the sample image to obtain the updated sample image, providing a new benchmark for the addition of subsequent materials and model training.
[0200] Step 410: Using the updated sample image and the sample material category of the first sample material as input, and the sample layout information of the first sample material as the label output, train the vision-language model, and return to execute the step of extracting the first sample material from multiple sample materials until, in the case of training with multiple sample materials completed, obtain the trained target layout model.
[0201] Using the updated sample image and the sample material category of the first sample material as input, and the sample layout information of the first sample material as the label output, train the vision-language model. One optional method is: calculate the difference between the layout information generated by the model and the true layout information, calculate the loss value through a loss function, and based on the loss value, update the parameters of the vision-language model through backpropagation.
[0202] Exemplarily, by calculating the difference between the layout information generated by the model and the true layout information, use the cross-entropy loss function for optimization and update the model parameters. After multiple iterations of training, the model can gradually improve the generation accuracy of the layout information, thus better completing the poster layout generation task.
[0203] In the embodiments of this specification, by obtaining sample images, multiple sample materials and their categories, layout information is generated, and the model is iteratively trained to obtain a target layout model. This model can efficiently complete the image layout adjustment task, significantly reduce the data collection and annotation costs, and at the same time improve the effect of layout generation. Through the iterative layout generation method, the volume of the data set can be expanded by making full use of the limited training data, and the cost problem brought by the layout annotation data can be alleviated. In addition, by using a vision-language model as a modeling tool, the accuracy and efficiency of the layout generation task can be further improved by making full use of its multi-modal understanding ability and pre-training advantages.
[0204] In an alternative embodiment of this specification, there is a hierarchical relationship among multiple sample materials; step 404 includes the following specific steps:
[0205] According to the hierarchical relationship, the first sample material is extracted layer by layer from multiple sample materials.
[0206] The hierarchical relationship among multiple sample materials is the order of priority for the layout adjustment of multiple sample materials
[0207] Figure 5 The flowchart of a model training method provided by an embodiment of this specification is shown, as Figure 5 shown:
[0208] In the training stage:
[0209] The sample image of a commodity poster can be decomposed into multiple sample images. According to the material category of the sample materials, the sample image can be decomposed into two samples. The first one is the poster picture of the original sample image, and the position of the background needs to be predicted (the position of the red frame in the figure). The second one is the updated sample image with the background material superimposed, that is, the first sample image, and the position of the text needs to be predicted (the position of the three red frames in the figure). The second sample image is obtained by iterative superposition.
[0210] In the inference stage:
[0211] Input the initial image without layout, first generate the position of the background, then use the first initial image with the background material pasted as the input for the second generation, and generate the position of the text again. Finally, complete the second initial image with the text material pasted as the target image.
[0212] Currently, in the field of commercial poster design, manual design or templates are usually used to layout text, logos, decorative elements and other materials. Designers need to comprehensively consider the main visual features of the product, the brand tone and the visual hierarchy, and manually adjust the position, size and stacking order of each material to meet the visual balance and information communication effect. However, the traditional template method has two significant defects: first, fixed templates are difficult to adapt to the main pictures of products of different sizes and composition styles, resulting in materials blocking key visual areas or unreasonable white space; second, the joint layout of multi-level materials lacks a dynamic coordination mechanism, which is prone to overlapping elements or misaligned spacing.
[0213] However, most existing automated layout solutions are limited to the position prediction of a single material type and fail to effectively model the synergistic relationship between poster content and multiple categories of materials. Specifically, (1) the semantic key areas of product images are ignored when making layout decisions, resulting in the obscuration of important visual elements; (2) the layout adjustment of materials at different levels is time-dependent, and adding elements later may destroy the balance of the existing layout; (3) the discrete coordinate prediction is difficult to capture the continuous correlation characteristics between the layout position and the image content, affecting the aesthetics and rationality of the layout solution.
[0214] The following combination Figure 6 Taking the application of the image layout adjustment method provided in this specification in the generation of commercial posters as an example, the image layout adjustment method is further described. Figure 6 A flowchart of a processing process of an image layout adjustment method for commercial poster generation provided by an embodiment of the present specification is shown, including the following specific steps:
[0215] Step 602: Obtain the material categories of the initial product poster and multiple target materials.
[0216] Step 604: Generate initial layout information of the target material of the first level based on the initial product poster and the material category of the target material of the first level through the pre-trained target layout model; add the target material of the first level to the initial product poster based on the initial layout information to obtain a first initial product poster; generate initial layout information of the target material of the second level based on the first initial product poster and the material category of the target material of the second level through the target layout model; add the target material of the second level to the first initial product poster based on the initial layout information to obtain a second initial product poster; until multiple target materials are added, the initial layout information of multiple target materials is obtained.
[0217] Step 606: Encode the initial product poster through the visual coding unit of the discrete visual diffusion model to obtain visual features.
[0218] Step 608: Through the discrete layout encoding unit of the discrete vision diffusion model, divide the coordinate values in the initial layout information into multiple preset intervals, and generate the discrete distribution feature of the initial layout information as the layout information feature.
[0219] Step 610: Through the discrete category encoding unit of the discrete vision diffusion model, divide the material categories into multiple preset categories, and generate the discrete distribution feature of the material categories as the category information feature.
[0220] Step 612: Through the feature processing module of the discrete vision diffusion model, under the constraint of the category information feature, use the layout information feature as the query feature, and the visual feature as the key feature and value feature to perform cross-attention calculation, obtain the attention-weighted feature, and perform noise reduction processing on the attention-weighted feature to obtain the target layout information of multiple target materials.
[0221] Step 614: Based on the target layout information, add multiple target materials to the initial product poster to obtain the target product poster.
[0222] In the embodiments of this specification, based on the hierarchical initial layout generation of the target layout model, by adding materials with different priorities in stages and iteratively updating the layout information, the defect that traditional templates cannot dynamically coordinate the temporal dependencies of multiple elements is overcome, ensuring the progressive balance of the spatial arrangement of materials at each level; secondly, through the dual-channel feature extraction of the discrete vision diffusion model, the visual semantic features of the product poster are decoupled and encoded with the discrete distribution features of the material categories and the initial layout, enabling the layout decision to accurately perceive the semantic weights of the key regions of the image and avoiding important visual elements from being covered; finally, with the help of the cross-attention mechanism and noise reduction processing, deep interaction between the layout feature and the visual feature is achieved under the category constraint, and the discrete coordinate prediction is transformed into layout optimization based on the continuous association of image content through the feature diffusion and noise reduction process in the continuous space, solving the problem of insufficient aesthetics caused by traditional discretization prediction. Through the content-aware layout feature fusion mechanism, this method realizes the dynamic adaptation of the layout position and the visual semantics of the poster while ensuring the harmonious coexistence of multiple elements, significantly improving the artistic rationality and commercial usability of the automated layout scheme.
[0223] Corresponding to the above method embodiments, this specification also provides embodiments of an image layout adjustment device. Figure 7 The structural schematic diagram of an image layout adjustment device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0224] The first acquisition module 702 is configured to acquire the initial image and the initial layout information of the target material;
[0225] The first encoding module 704 is configured to encode the initial image to obtain visual features, and encode the initial layout information to obtain layout information features;
[0226] The first noise reduction module 706 is configured to perform noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material;
[0227] The first addition module 708 is configured to add the target material to the initial image based on the target layout information to obtain the target image.
[0228] Optionally, the first encoding module 704 is further configured to:
[0229] Encode the initial image through the encoding module of the visual diffusion model to obtain visual features, and encode the initial layout information to obtain layout information features;
[0230] Correspondingly, the first noise reduction module 706 is further configured to:
[0231] Perform noise reduction processing on the layout information features based on the visual features through the feature processing module of the visual diffusion model to obtain the target layout information of the target material.
[0232] Optionally, the first noise reduction module 706 is further configured to:
[0233] Use the layout information features as query features, and the visual features as key features and value features to perform cross-attention calculation to obtain attention-weighted features, and perform noise reduction processing on the attention-weighted features to obtain the target layout information of the target material.
[0234] Optionally, the device further includes:
[0235] The first control module is configured to obtain the material category of the target material; encode the material category to obtain category information features;
[0236] Correspondingly, the first noise reduction module 706 is further configured to:
[0237] Under the constraint of the category information features, perform noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material.
[0238] Optionally, the first encoding module 704 is further configured to:
[0239] Through the discrete layout encoding unit of the discrete visual diffusion model, divide the coordinate values in the initial layout information into multiple preset intervals to generate discrete distribution features of the initial layout information;
[0240] Correspondingly, the first control module is further configured to:
[0241] Through the discrete category encoding unit of the discrete visual diffusion model, the material categories are divided into multiple preset categories to generate the discrete distribution characteristics of the material categories.
[0242] Correspondingly, the first noise reduction module 706 is further configured to:
[0243] Under the constraint of the category information feature, perform noise reduction processing on the layout information feature based on the visual feature, output the discrete distribution characteristics of the target layout information of the target material, and determine the target layout information of the target material based on the discrete distribution characteristics.
[0244] Optionally, the first noise reduction module 706 is further configured to:
[0245] Through the first feature processing unit in the feature processing module of the visual diffusion model, under the constraint of the category information feature, perform noise reduction processing on the layout information feature based on the visual feature, and output the first layout information feature, where the feature processing module includes multiple sequentially connected feature processing units; through the second feature processing unit in the feature processing module, under the constraint of the category information feature, perform noise reduction processing on the first layout information feature based on the visual feature, and output the second layout information feature; until the last layout information feature output by the last feature processing unit in the feature processing module is obtained, and based on the last layout information feature, determine the target layout information of the target material.
[0246] Optionally, the first acquisition module 702 is further configured to:
[0247] Acquire the initial image and the material categories of the target material; through the pre-trained target layout model, generate the initial layout information of the target material based on the initial image and the material categories.
[0248] Optionally, there are multiple target materials, and there is a hierarchical relationship between the multiple target materials; the first acquisition module 702 is further configured to:
[0249] Through the pre-trained target layout model, generate the initial layout information of the target material at the first level based on the initial image and the material categories of the target material at the first level; based on the initial layout information, add the target material at the first level to the initial image to obtain the first initial image; through the target layout model, generate the initial layout information of the target material at the second level based on the first initial image and the material categories of the target material at the second level; based on the initial layout information, add the target material at the second level to the first initial image to obtain the second initial image, until the initial layout information of the multiple target materials is obtained when the addition of all the multiple target materials is completed.
[0250] In the embodiments of this specification, the initial layout information is regarded as the result of adding noise to the ideal layout information. On this basis, noise reduction processing is performed on the layout information features of the initial layout information based on the extracted visual features, realizing the adaptive adjustment of the element layout of the target element, thereby obtaining more accurate target material layout information that conforms to the context environment. Based on the target layout information, the target material is added to the initial image to generate a target image that is both beautiful and can effectively convey information. This not only reduces the need for manual adjustment, lowers the labor cost, but also significantly improves the work efficiency and design quality, ensuring that image works that meet user needs can be efficiently created in various application scenarios, achieving a flexible and efficient image layout solution while satisfying visual aesthetics and information transmission efficiency.
[0251] The above is a schematic solution of an image layout adjustment device in this embodiment. It should be noted that the technical solution of this image layout adjustment device and the technical solution of the above image layout adjustment method belong to the same concept. For the details not described in the technical solution of the image layout adjustment device, reference can be made to the description of the technical solution of the above image layout adjustment method.
[0252] Corresponding to the above method embodiment, this specification also provides an embodiment of a model training device. Figure 8 It shows a schematic structural diagram of a model training device provided by an embodiment of this specification. As Figure 8 shown, the device includes:
[0253] A second acquisition module 802, configured to acquire sample images, a plurality of sample materials, and the sample material categories of the plurality of sample materials;
[0254] A second extraction module 804, configured to extract a first sample material from the plurality of sample materials;
[0255] A second generation module 806, configured to generate sample layout information of the first sample material based on the sample material category of the first sample material and the sample image;
[0256] A second addition module 808, configured to add the first sample material to the sample image based on the sample layout information to obtain an updated sample image;
[0257] A second training module 810, configured to use the updated sample image and the sample material category of the first sample material as inputs and the sample layout information of the first sample material as the label output to train a vision-language model, and return to execute the step of extracting the first sample material from the plurality of sample materials until the target layout model is trained and completed when the plurality of sample materials are trained.
[0258] Optionally, there is a hierarchical relationship among multiple sample materials; the second extraction module 804 is further configured to:
[0259] Extract the first sample material layer by layer from the multiple sample materials according to the hierarchical relationship.
[0260] In the embodiments of this specification, by obtaining sample images, multiple sample materials and their categories, generating layout information, and iteratively training the model, a target layout model is obtained. This model can efficiently complete the image layout adjustment task, significantly reduce the data collection and annotation costs, and at the same time improve the layout generation effect. Through the iterative layout generation method, the volume of the data set can be expanded by making full use of the limited training data, and the cost problem brought by the layout annotation data can be alleviated. In addition, by using the vision-language model as the modeling tool, the accuracy and efficiency of the layout generation task can be further improved by making full use of its multi-modal understanding ability and pre-training advantages.
[0261] The above is a schematic solution of a model training device in this embodiment. It should be noted that the technical solution of this model training device and the technical solution of the above model training method belong to the same concept. For the details not described in the technical solution of the model training device, reference can be made to the description of the technical solution of the above model training method.
[0262] Figure 9 The structural block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0263] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0264] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0265] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or Personal Computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0266] Among them, the processor 920 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above image layout adjustment method or model training method are implemented.
[0267] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above image layout adjustment method and model training method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above image layout adjustment method or model training method.
[0268] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, which when executed by a processor implement the steps of the above image layout adjustment method or model training method.
[0269] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above image layout adjustment method and model training method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above image layout adjustment method or model training method.
[0270] An embodiment of this specification also provides a computer program product including computer programs / instructions, which when executed by a processor implement the steps of the above image layout adjustment method or model training method.
[0271] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above image layout adjustment method and model training method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above image layout adjustment method or model training method.
[0272] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0273] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM for short), random access memories (RAM for short), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0274] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0275] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0276] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all details and do not limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and pass through this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for adjusting image layout, characterized in that: include: Obtaining initial layout information of the initial image and target material; Encoding the initial image to obtain visual features, and encoding the initial layout information to obtain layout information features; Performing noise reduction processing on the layout information feature based on the visual feature to obtain target layout information of the target material; Based on the target layout information, the target material is added to the initial image to obtain the target image.
2. The method according to claim 1, characterized in that: The encoding of the initial image to obtain visual features, and the encoding of the initial layout information to obtain layout information features, includes: By using the encoding module of the visual diffusion model, the initial image is encoded to obtain visual features, and the initial layout information is encoded to obtain layout information features; The performing noise reduction processing on the layout information feature based on the visual feature to obtain the target layout information of the target material includes: The feature processing module of the visual diffusion model performs noise reduction processing on the layout information features based on the visual features to obtain the target layout information of the target material.
3. The method according to claim 1, characterized in that The performing noise reduction processing on the layout information feature based on the visual feature to obtain the target layout information of the target material includes: The layout information feature is used as a query feature, and the visual feature is used as a key feature and a value feature to perform cross-attention calculation to obtain an attention weighted feature, and noise reduction processing is performed on the attention weighted feature to obtain the target layout information of the target material.
4. The method according to any one of claims 1 to 3, characterized in that: Before performing noise reduction processing on the layout information feature based on the visual feature to obtain the target layout information of the target material, the method further includes: Get the material category of the target material; Encoding the material category to obtain category information features; The performing noise reduction processing on the layout information feature based on the visual feature to obtain the target layout information of the target material includes: Under the constraint of the category information feature, noise reduction processing is performed on the layout information feature based on the visual feature to obtain target layout information of the target material.
5. The method according to claim 4, characterized in that The encoding of the initial layout information to obtain layout information features includes: The coordinate values in the initial layout information are divided into a plurality of preset intervals by a discrete layout encoding unit of a discrete visual diffusion model to generate discrete distribution features of the initial layout information; The step of encoding the material category to obtain category information features includes: The material category is divided into a plurality of preset categories by a discrete category encoding unit of the discrete visual diffusion model to generate discrete distribution features of the material category; The step of performing noise reduction processing on the layout information feature based on the visual feature under the constraint of the category information feature to obtain the target layout information of the target material includes: Under the constraint of the category information feature, noise reduction processing is performed on the layout information feature based on the visual feature, discrete distribution features of the target layout information of the target material are output, and the target layout information of the target material is determined based on the discrete distribution features.
6. The method according to claim 4, characterized in that The step of performing noise reduction processing on the layout information feature based on the visual feature under the constraint of the category information feature to obtain the target layout information of the target material includes: By using a first feature processing unit in a feature processing module of a visual diffusion model, under the constraint of the category information feature, performing noise reduction processing on the layout information feature based on the visual feature, and outputting a first layout information feature, wherein the feature processing module includes a plurality of feature processing units connected in sequence; By means of a second feature processing unit in the feature processing module, under the constraint of the category information feature, based on the visual feature, a noise reduction process is performed on the first layout information feature to output a second layout information feature; Until the last layout information feature output by the last feature processing unit in the feature processing module is obtained, the target layout information of the target material is determined based on the last layout information feature.
7. The method according to claim 1, characterized in that The obtaining of the initial layout information of the initial image and the target material includes: Get the material category of the initial image and the target material; Initial layout information of the target material is generated based on the initial image and the material category through a pre-trained target layout model.
8. The method according to claim 7, characterized in that There are multiple target materials, and there is a hierarchical relationship between the multiple target materials; The generating the initial layout information of the target material based on the initial image and the material category by using the pre-trained target layout model includes: Generate initial layout information of the target material of the first level based on the initial image and the material category of the target material of the first level by using a pre-trained target layout model; Based on the initial layout information, adding the target material of the first level to the initial image to obtain a first initial image; Generate initial layout information of the target material of the second level through the target layout model based on the first initial image and the material category of the target material of the second level; Based on the initial layout information, the target material of the second level is added to the first initial image to obtain a second initial image, and when all the target materials are added, the initial layout information of the target materials is obtained.
9. A model training method, characterized in that: include: Acquire a sample image, a plurality of sample materials, and sample material categories of the plurality of sample materials; Extracting a first sample material from the plurality of sample materials; generating sample layout information of the first sample material based on the sample material category of the first sample material and the sample image; adding the first sample material to the sample image according to the sample layout information to obtain an updated sample image; The updated sample image and the sample material category of the first sample material are used as inputs, and the sample layout information of the first sample material is used as a label output to train the visual language model, and the step of extracting the first sample material from the multiple sample materials is returned to be executed until the training with the multiple sample materials is completed, and a trained target layout model is obtained.
10. The method according to claim 9, characterized in that There is a hierarchical relationship between the multiple sample materials; The extracting a first sample material from the plurality of sample materials comprises: According to the hierarchical relationship, the first sample material is extracted layer by layer from the multiple sample materials.
11. An image layout adjustment device, characterized in that: include: A first acquisition module is configured to acquire initial layout information of an initial image and a target material; A first encoding module is configured to encode the initial image to obtain visual features, and to encode the initial layout information to obtain layout information features; A first noise reduction module is configured to perform noise reduction processing on the layout information feature based on the visual feature to obtain target layout information of the target material; The first adding module is configured to add the target material to the initial image based on the target layout information to obtain the target image.
12. A model training device, characterized in that: include: A second acquisition module is configured to acquire a sample image, a plurality of sample materials, and sample material categories of the plurality of sample materials; A second extraction module is configured to extract a first sample material from the plurality of sample materials; A second generating module is configured to generate sample layout information of the first sample material based on the sample material category of the first sample material and the sample image; A second adding module is configured to add the first sample material to the sample image according to the sample layout information to obtain an updated sample image; The second training module is configured to train the visual language model with the updated sample image and the sample material category of the first sample material as input and the sample layout information of the first sample material as label output, and return to execute the step of extracting the first sample material from the multiple sample materials until the training with the multiple sample materials is completed, and a trained target layout model is obtained.
13. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 10 are implemented.
14. A computer-readable storage medium, characterized in that: It stores a computer program / instruction, which implements the steps of the method described in any one of claims 1 to 10 when executed by a processor.
15. A computer program product, characterized in that The invention comprises a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Cited By
Personalized response method for pre-training text to image diffusion model
CN121746851A