Collage generation method and device, computer device and readable storage medium
By adjusting the geometric parameters driven by rendering loss on a set of discrete image patches, the problem of traditional collage creation relying on manual operation and difficulty in understanding abstract semantics is solved, and collage generation with high accuracy and editability is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional collage creation relies on the creator's personal aesthetics and manual operation, which is a high barrier to entry for non-professionals. Furthermore, existing technology struggles to understand abstract semantics through geometric features, resulting in low accuracy of the generated collage images.
By acquiring a set of discrete image blocks and collage prompt text, the image blocks are rendered separately, the rendering loss is determined, and the geometric parameters are adjusted until the preset conditions are met, generating the target rendering result. This enables semantic guidance without the need for a preset shape template, directly responding to high-level abstract text prompts.
It enables the accurate generation of collages that match text descriptions without relying on preset shape templates, improving the accuracy of collage generation, preserving the visual characteristics of the original image blocks, supporting independent layer editing, and reducing the cost of secondary creation.
Smart Images

Figure CN121563767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual design technology, and in particular to a collage generation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology
[0002] In the fields of digital art creation and visual design, image collage, as a form of expression that reassembles discrete visual elements into a unified semantic whole, has wide-ranging application value. Traditional collage creation heavily relies on the creator's personal aesthetic sense and manual operation, making it extremely difficult for non-professionals. With the development of generative artificial intelligence, how to automatically generate collage layouts using computational methods has become a research hotspot.
[0003] In related technologies, it is difficult to understand abstract semantics through geometric features, and it is impossible to generate a reasonable visual structure based solely on text prompts, resulting in low accuracy of the generated collage images. Summary of the Invention
[0004] Therefore, it is necessary to provide a collage generation method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can accurately generate collage images, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a method for generating collages, including:
[0006] Retrieve a set of discrete image patches and collage hint text;
[0007] Each image block in the discrete image block set is rendered separately to obtain the mosaic rendering result;
[0008] Based on the collage rendering result and the collage prompt text, determine the rendering loss;
[0009] The geometric parameters of the image patch are adjusted according to the rendering loss to obtain the adjusted geometric parameters;
[0010] The new pose of the image block is determined based on the adjusted geometric parameters, and the image block in the new pose is rendered to obtain a new mosaic rendering result until the preset rendering conditions are met to obtain the target rendering result.
[0011] Secondly, this application also provides a collage generation device, comprising:
[0012] The data acquisition module is used to acquire a set of discrete image patches and collage prompt text;
[0013] The image patch rendering module is used to render each image patch in the discrete image patch set to obtain the patch rendering result;
[0014] The rendering loss determination module is used to determine the rendering loss based on the tiling rendering result and the tiling prompt text;
[0015] A geometric parameter adjustment module is used to adjust the geometric parameters of the image block according to the rendering loss to obtain the adjusted geometric parameters;
[0016] The rendering result acquisition module is used to determine the new pose of the image block based on the adjusted geometric parameters, and to render the image block in the new pose to obtain a new image block rendering result, until the preset rendering conditions are met to obtain the target rendering result.
[0017] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the collage generation method provided in the first aspect.
[0018] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the collage generation method provided in the first aspect.
[0019] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the collage generation method provided in the first aspect.
[0020] The aforementioned collage generation method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a set of discrete image blocks and collage prompt text. They then render each image block in the discrete image block set to obtain a collage rendering result. Based on the collage rendering result and the collage prompt text, they determine a rendering loss, adjust the geometric parameters of the image blocks according to the rendering loss, obtain adjusted geometric parameters, determine a new pose of the image blocks based on the adjusted geometric parameters, and render the image blocks in the new pose to obtain a new collage rendering result. This process continues until preset rendering conditions are met to obtain the target rendering result. This enables semantic guidance through collage prompt text, allowing direct understanding and response to high-level abstract text prompts. Without any preset shape templates or reference images, it automatically collages discrete image block materials into a collage that conforms to the text description, thereby achieving accurate collage generation and improving the accuracy of collage generation. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is an application environment diagram of a collage generation method in one embodiment;
[0023] Figure 2 This is a flowchart illustrating a collage generation method in one embodiment;
[0024] Figure 3 This is a schematic diagram of the image block rendering process in one embodiment;
[0025] Figure 4 This is a flowchart illustrating a collage generation method in another embodiment;
[0026] Figure 5 This is a structural block diagram of a collage generation device in one embodiment;
[0027] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0030] The collage generation method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 can obtain a set of discrete image blocks and collage prompt text input by the user, and send these to server 104. After obtaining the set of discrete image blocks and collage prompt text sent by terminal 102, server 104 renders each image block in the set of discrete image blocks to obtain a collage rendering result. Based on the collage rendering result and collage prompt text, it determines the rendering loss, adjusts the geometric parameters of the image blocks according to the rendering loss, obtains the adjusted geometric parameters, determines the new pose of the image blocks based on the adjusted geometric parameters, and renders the image blocks in the new pose to obtain a new collage rendering result. This process continues until the preset rendering conditions are met, resulting in the target rendering result, i.e., the final generated collage. Server 104 can return the target rendering result to terminal 102, and terminal 102 displays the target rendering result. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. It should be noted that the collage generation method provided in this application embodiment is applicable not only to application scenarios involving interaction between a server and a terminal, but also to application scenarios involving a single terminal, a single server, terminal-to-terminal interaction, or server-to-server interaction.
[0031] In one exemplary embodiment, such as Figure 2 As shown, a collage generation method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 202 to 208. Wherein:
[0032] Step 202: Obtain the set of discrete image patches and the collage prompt text.
[0033] In this context, a discrete image patch set refers to a collection of multiple discrete image patches. The discrete image patch set essentially serves as the image patch material for generating a collage. Based on this set, collage patterns corresponding to the semantics of the collage prompt text can be generated. There is no specific limit to the number of image patches included in the discrete image patch set. Generally, the more image patches in the discrete image patch set, the more realistic the final collage and the closer it is to the semantics of the collage prompt text. The collage prompt text refers to the prompt text used to generate the collage target. Based on the semantics of the collage prompt text, the image patches in the discrete image patch set can be pieced together to form the corresponding pattern. For example, the collage prompt text could be "assemble a horse," "assemble a cup," or "assemble a tree," etc. The collage prompt text can be used to represent the collage intention.
[0034] In practical applications, the collage hint text can be obtained by converting the collage hint speech from speech to text, or it can be directly obtained as text. The server can obtain the discrete image patch set and collage hint text directly from the human-computer interaction interface, or it can obtain the discrete image patch set and collage hint text from other servers or terminals.
[0035] Step 204: Render each image block in the discrete image block set to obtain the mosaic rendering result.
[0036] In a straightforward manner, each image patch in the discrete image patch set has corresponding geometric parameters such as position, angle, and scaling on the canvas. Each image patch is placed on the canvas according to its corresponding geometric parameters, and then the image patches are rendered sequentially in a preset order to obtain the mosaic rendering result. The mosaic rendering result refers to the combined result after each image patch in the discrete image patch set has been rendered individually. After each image patch is rendered, a corresponding image patch rendering result is obtained. The mosaic rendering result is obtained by combining the image patch rendering results corresponding to each image patch in the discrete image patch set. In the initial state, each image patch in the discrete image patch set can be initialized with a geometric parameter, and each image patch is placed on the canvas based on this initialized geometric parameter. In non-initial states, the geometric parameters of each image patch can be determined by the rendering loss corresponding to the previous rendering.
[0037] In an exemplary embodiment, each image block in the discrete image block set can be rendered sequentially according to a preset order to obtain a mosaic rendering result. The preset order can be the same or different in different rendering rounds. In practical applications, a depth value can be set for each image block in each rendering round. The server sorts the image blocks in the image block set according to this depth value, and then renders each image block sequentially according to this sorted order to obtain the image block rendering result for each image block. The Alpha Blending algorithm is then used to superimpose the image block rendering results corresponding to each image block to obtain the mosaic rendering result for the corresponding round.
[0038] In an exemplary embodiment, before rendering the image blocks in the discrete image block set, each image block in the image block set can be preprocessed to obtain preprocessed image blocks, and then the preprocessed image blocks are rendered to obtain the mosaic rendering result. Preprocessing may include, for example, removing image block backgrounds, denoising, smoothing, etc.
[0039] For example, each image block in the discrete image block set can be rendered separately using a differentiable renderer based on Gaussian sputtering to obtain the image block rendering result corresponding to each image block. The image block rendering results of each image block can then be combined to obtain the mosaic rendering result.
[0040] Step 206: Determine the rendering loss based on the tiling rendering results and tiling hint text.
[0041] The rendering loss characterizes the degree of matching between the collage rendering result and the collage semantics corresponding to the collage hint text. The closer the collage rendering result is to the collage semantics corresponding to the collage hint text, the greater the degree of matching between the collage rendering result and the collage semantics corresponding to the collage hint text, and the smaller the rendering loss. Conversely, the further the collage rendering result is from the collage semantics corresponding to the collage hint text, the smaller the degree of matching between the collage rendering result and the collage semantics corresponding to the collage hint text, and the greater the rendering loss.
[0042] In one example, the discriminative matching degree between the collage rendering result and the collage semantics corresponding to the collage hint text can be determined based on a discriminative model. Alignment prediction between the collage rendering result and the collage semantics corresponding to the collage hint text can be performed based on a scene distribution simulation model to obtain the alignment confidence. The rendering loss can then be determined based on the discriminative matching degree and the alignment confidence. For example, the rendering loss can be obtained by weighted summation of the discriminative matching degree and the alignment confidence.
[0043] In one example, the corresponding structural loss can be determined based on the tiling rendering result. The structural loss characterizes the structural layout of the tiling rendering result on the canvas. The semantic loss is determined based on the discriminant matching score and alignment confidence score. The rendering loss is then determined based on the semantic and structural losses. For example, the rendering loss is obtained by weighted summing the semantic and structural losses.
[0044] Step 208: Adjust the geometric parameters of the image block according to the rendering loss to obtain the adjusted geometric parameters; determine the new pose of the image block based on the adjusted geometric parameters, and render the image block in the new pose to obtain a new mosaic rendering result, until the preset rendering conditions are met to obtain the target rendering result.
[0045] The geometric parameters of an image patch include its position, angle, and scaling. The position of an image patch can be represented by its coordinates on the canvas, specifically by the coordinates of its center or centroid pixels. The angle of the image patch indicates its orientation. It's easy to understand that, while the position of the image patch remains constant, its orientation or angle can include multiple values. The scaling factor refers to the ratio by which the image patch is enlarged or reduced to fit the canvas layout. For example, initially, each image patch is positioned at (0,0), which is the center of the canvas. The angles of each image patch are randomly sampled within the range of [-0.1, 0.1] radians. The scaling factor s init =(1 / N) 0.5 +0.05, where N represents the number of image patches, meaning that the scaling ratio of each image patch in the same discrete image patch set can be the same.
[0046] In practical applications, after obtaining the rendering loss, the server can backpropagate based on the rendering loss to obtain the geometric parameters of the corresponding image patch. The original geometric parameters of the image patch are then adjusted to the geometric parameters determined based on the rendering loss (i.e., the adjusted geometric parameters). Based on the adjusted geometric parameters, the new pose of the image patch can be determined. The image patch in the new pose is then rendered again to obtain a new collage rendering result. This process of rendering the image patch to obtain the collage rendering result and determining the new geometric parameters based on the rendering loss is iterated continuously until the preset rendering conditions are met. Rendering then stops, and the collage rendering result obtained from the last rendering is taken as the target rendering result, which is the generated collage.
[0047] Preset rendering conditions are used to characterize the conditions under which rendering of image patches is stopped, that is, the conditions under which the target rendering result that meets the requirements is obtained. Preset rendering conditions may include, for example, reaching a preset number of rendering attempts, or the rendering loss being less than a loss threshold. The preset number of attempts or the loss threshold can be set according to the actual application scenario.
[0048] In the aforementioned collage generation method, a set of discrete image blocks and collage prompt text are obtained. Each image block in the discrete image block set is rendered to obtain a collage rendering result. Based on the collage rendering result and the collage prompt text, a rendering loss is determined. The geometric parameters of the image blocks are adjusted according to the rendering loss to obtain adjusted geometric parameters. Based on the adjusted geometric parameters, a new pose of the image blocks is determined, and the image blocks in the new pose are rendered to obtain a new collage rendering result. This process continues until preset rendering conditions are met to obtain the target rendering result. This method enables semantic guidance through collage prompt text, allowing direct understanding and response to high-level abstract text prompts. Without any preset shape templates or reference images, it automatically collages discrete image block materials into a collage that conforms to the text description, thereby achieving accurate collage generation and improving the accuracy of collage generation. In addition, during the collage generation process, only the geometric parameters of the image blocks are adjusted, without making substantial changes to the content of the image blocks such as color. This ensures that the generated collage fully retains the visual features of the original image block materials, improving the visual effect of the collage. While obtaining the target rendering result, the geometric parameters corresponding to each image block can be obtained, making each image block highly editable. This allows users to independently edit any layer of the generated collage, facilitating secondary creation and reducing the cost of secondary creation.
[0049] In some embodiments, step 204 renders each image patch in the discrete image patch set to obtain a mosaic rendering result, including:
[0050] The background of each image patch in the discrete image patch set is detected and removed to obtain an image patch including transparency information; the image patch including transparency information is rendered to obtain the mosaic rendering result.
[0051] Transparency information refers to the opacity or visibility of each pixel in an image. This transparency information is typically stored in the image's alpha channel, an additional channel parallel to the color channels (such as red (R), green (G), and blue (B). The alpha channel value represents the transparency level of each pixel. For an 8-bit image, the transparency level is typically between 0 and 255, where 0 represents complete transparency (the pixel is not visible) and 255 represents complete opacity (the pixel is fully visible). Standard RGB images with transparency information are often called RGBA images, where A stands for Alpha. When an image patch is overlaid on a background, the alpha channel determines which parts of the foreground image are visible and which are transparent.
[0052] For example, the rembg library can be used to detect and remove the background of an image patch, resulting in an image patch that includes transparency information. Taking an image patch in RGB format as an example, the foreground and background in the image patch can be detected, and a binary mask can be generated based on the detection results. In the mask, white areas correspond to the foreground, and black areas correspond to the background. rembg can then use the binary mask to set the alpha channel value of the foreground area to 255 and the alpha channel value of the background area to 0. By combining the RGB three channels and the alpha channel of the image patch, an RGBA format image patch is obtained, which includes transparency information.
[0053] In this embodiment, by detecting and removing the background of each image block in the discrete image block set, an image block including transparency information is obtained. The image block including transparency information is then rendered to obtain the collage rendering result. The rendering can be superimposed based on the transparency information, ensuring that the occlusion relationship and edge blending between image blocks are more natural during the rendering process, thus improving the visual experience of the generated collage.
[0054] In some embodiments, step 204 renders each image patch in the discrete image patch set to obtain a mosaic rendering result, including:
[0055] Determine the position of the target pixel in the target image block within the discrete image block set; determine the rendering area based on the position of the target pixel and the Gaussian kernel size; determine the distance between each pixel in the rendering area and the target pixel, and determine the rendering weight corresponding to each pixel based on the distance; render the pixels of the target pixel based on the rendering weight and pixel value corresponding to each pixel in the rendering area to obtain the rendering result of the target pixel.
[0056] The target image block can be any image block in the set of discrete image blocks, and the target pixel can be any pixel in the target image block.
[0057] In practical applications, the rendering area is determined based on the location of the target pixel and the Gaussian kernel size. This can be achieved by selecting an area centered on the target pixel and corresponding to a Gaussian kernel size. The Gaussian kernel size can be, for example, 4x4, 5x5, or 6x6. The Gaussian kernel size can be chosen based on the specific application scenario and is not specifically limited here. For example, with a Gaussian kernel size of 5x5, a 5x5 pixel area centered on the target pixel is selected as the rendering area for that pixel. Different pixels generally correspond to different rendering areas.
[0058] The distance between each pixel within the rendering area and the target pixel can be characterized by the Euclidean distance, Manhattan distance, Chebyshev distance, etc., between the coordinates of the two pixels. The distance can be directly used as the rendering weight for each pixel, or the distance can be normalized to obtain the corresponding rendering weight. For example, the rendering weight can be calculated based on a Gaussian kernel function, such as... Where d represents the distance between each pixel in the rendering area and the target pixel, and σ is a constant. Generally, the greater the distance, the smaller the rendering weight; the closer the distance, the greater the rendering weight. The weighted sum or weighted average of the rendering weights and pixel values corresponding to each pixel in the rendering area can be used as the rendering result of the target pixel. Each pixel in the target image block is rendered according to the rendering method of the target pixel to obtain the image block rendering result corresponding to the target image block. Each image block in the discrete image block set is rendered according to the rendering method of the target image block to obtain the image block rendering result corresponding to each image block. The rendering results of each image block are combined in a preset order to obtain the mosaic rendering result.
[0059] In one example, the rendering process diagram is as follows: Figure 3 As shown, the discrete image patch set ɛ includes multiple image patches e1, e2, ..., e N That is, ɛ={e1,e2,…,e N Let's take target image block e1 as an example. For the target pixel in target image block e1, using the target pixel as the center and a Gaussian kernel size of 5*5 as an example, a 5*5 pixel area is selected as the rendering area. The distance between each pixel in the rendering area and the target pixel is calculated, and the rendering weight of each pixel is determined based on this distance. For example, the rendering weight corresponding to the target pixel is 41, the rendering weight corresponding to the pixels horizontally and vertically adjacent to the target pixel is 26, and the rendering weight corresponding to the pixels diagonally adjacent to the target pixel is 16, etc. The rendering weight and pixel value of each pixel in the rendering area are weighted and averaged to obtain the rendering result of the target pixel. Similarly, the rendering results corresponding to each pixel in target image block e1 can be obtained, resulting in the image block rendering result corresponding to target image block e1. The image block rendering results corresponding to each image block are then superimposed and combined to obtain the rendered image x0 (i.e., the mosaic rendering result). Based on the tile rendering results and tile hint text obtained in the current rendering round, the rendering loss can be determined, the geometric parameters of each image patch can be determined based on the rendering loss, and the new pose of the image patch can be determined based on the geometric parameters of the image patch. The next round of rendering can then be performed on the image patch in the new pose.
[0060] In this embodiment, the rendering area is determined by the position of the target pixel in the target image block in the discrete image block set and the size of the Gaussian kernel. The rendering weight of the corresponding pixel is determined according to the distance between each pixel in the rendering area and the target pixel. The target pixel is rendered according to the rendering weight and pixel value of each pixel in the rendering area to obtain the rendering result of the target pixel. This can accurately render each image block and improve the accuracy of the mosaic rendering result.
[0061] In some embodiments, step 206, determining the rendering loss based on the tiling rendering result and the tiling hint text, includes:
[0062] The discriminant matching degree between the collage rendering result and the collage prompt text is determined based on the text-to-image discriminant model; the alignment confidence is obtained by performing alignment prediction on the collage rendering result and the collage prompt text based on the scene distribution simulation model; and the rendering loss is determined based on the discriminant matching degree and the alignment confidence.
[0063] The text-to-image discriminant model is used to determine the matching degree between the collage rendering result and the collage hint text. It generates a reference collage image corresponding to the collage hint text and determines the matching degree between the reference collage image and the collage rendering result. The scene distribution simulation model is used to predict the alignment between the collage rendering result and the collage hint text. It simulates scene distribution based on the collage hint text and predicts the alignment between the simulated scene distribution and the collage rendering result.
[0064] For example, a text-to-image discriminant model may include a text-to-image model (i.e., a generator) and a discriminator. The generator produces a corresponding image from the input text, and the discriminator determines the matching degree between the input image and the generated image. In practical applications, the text-to-image model can be trained using the discriminator until the discriminator determines that the matching degree between the image generated by the text-to-image model and the input text meets a preset condition, thus obtaining a trained text-to-image model.
[0065] In one example, the collage hint text can be input into a trained text-based graph model to generate a reference collage image. A discriminator then determines the discriminative matching degree between the reference collage image and the rendered collage. A scene distribution model is used to simulate the scene distribution of the collage hint text, resulting in a simulated scene distribution (i.e., the simulated collage image). Alignment prediction is then performed between the simulated scene distribution and the rendered collage to obtain the alignment confidence. The semantic loss is determined based on the difference between the alignment confidence and the discriminative matching degree, and can then be directly used as the rendering loss. For example, the difference between the alignment confidence and the discriminative matching degree can be used as the semantic loss, or the product of the difference between the alignment confidence and the discriminative matching degree and a weight can be used as the semantic loss.
[0066] For example, based on the Variational Score Distillation (VSD) technique, a frozen pre-trained text-based graph model (such as DeepFloyd-IF) can be used. As a semantic discriminator (i.e., text-generated image discrimination model), a lightweight LoRA model is also trained. (i.e., scene distribution simulation model) to simulate the distribution of the current scene. The semantic gradient calculation formula is shown in formula (1) below.
[0067] Formula (1)
[0068] Where x(t) represents the noisy tile rendering result at time t; y represents the tile prompt text; and w(t) represents the weight at time t. It is easy to understand that different noises can be added to the corresponding tile rendering results at different times t, which can further enhance the accuracy of the rendering loss. Compared with traditional SDS (Score Distillation Sampling), VSD can provide more stable and diverse gradient guidance, effectively avoiding oversaturation and mode collapse problems in the generated results.
[0069] In this embodiment, the discriminant matching degree between the collage rendering result and the collage prompt text is determined by a text-based image discrimination model. The alignment prediction between the collage rendering result and the collage prompt text is performed based on a scene distribution simulation model to obtain the alignment confidence. The rendering loss is determined based on the discriminant matching degree and the alignment confidence. This can more accurately determine the rendering loss and precisely guide each image block to move towards the position corresponding to the semantics of the collage prompt text, thereby improving the accuracy of the generated collage.
[0070] In some embodiments, determining the rendering loss based on the discriminant matching degree and alignment confidence degree includes:
[0071] The semantic loss is determined based on the discriminant matching degree and alignment confidence; the cumulative transparency value of each pixel in the tiling rendering result is determined, and the overlap penalty value is determined based on the cumulative transparency value; the geometric constraint value is determined based on the positional relationship between the center of the image patch and the preset canvas area; and the rendering loss is determined based on the semantic loss, overlap penalty value, and geometric constraint value.
[0072] For example, the semantic loss can be determined based on the difference between the alignment confidence and the discriminative matching degree. The tiling rendering result is equivalent to the generated tiling image. Since a pixel in the tiling image may superimpose the transparency information of pixels from multiple image blocks, the transparency information of the corresponding pixels from multiple image blocks is accumulated to obtain the transparency accumulation value of the corresponding pixel in the tiling image. For example, a penalty can be imposed on pixels whose transparency accumulation value is greater than the transparency threshold to reduce unnecessary overlap between image blocks. The transparency threshold can be set according to the application scenario, for example, 1. For example, the formula for calculating the overlap penalty value is shown in formula (2) below.
[0073] Formula (2)
[0074] ReLU() is the activation function. (u,v) represents the coordinates of a pixel in the image. This indicates the allowed overlap threshold (i.e., the maximum overlap threshold). This indicates the overlap strength of the image blocks at that pixel. This represents the summation of all pixels to obtain the overall overlap penalty value.
[0075] It is easy to understand that the tiling rendering result includes the geometric parameter information of each image block. The image block center in the geometric parameters of each image block can be obtained. The image block center can be the position coordinate of the image block. Based on the positional relationship between the image block center and the preset canvas area, it is determined whether the image block center exceeds the preset canvas area. The image block center can be constrained by the geometric constraint value so that the image block center does not exceed the preset canvas area. At the same time, the scaling ratio of the image block is constrained within a reasonable range to prevent some image blocks from being too large to occlude other elements or too small to be visible. The preset canvas area refers to the preset area on the canvas. The preset canvas area can be the entire area of the canvas or a local area. For example, the calculation formula of the geometric constraint value is shown in the following formula (3).
[0076] Formula (3)
[0077] in, This represents the distance between the center of the image patch at the current location and the initial location (e.g., (0,0)). Indicates the boundary distance of the preset canvas area (e.g., (The radius of the inscribed circle of the preset canvas area). Indicates the scaling ratio of the image patch; Indicates the reference scaling ratio; This indicates the maximum scaling ratio.
[0078] In one example, the structural loss can be obtained by weighted summing of the overlap penalty and geometric constraint values, and then the rendering loss can be obtained by weighted summing of the semantic loss and structural loss. Alternatively, the rendering loss can be obtained by directly weighted summing of the semantic loss, overlap penalty, and geometric constraint values.
[0079] In this embodiment, semantic loss is determined based on the discriminant matching degree and alignment confidence degree. The cumulative transparency value of each pixel in the collage rendering result is determined, and the overlap penalty value is determined based on the cumulative transparency value. The geometric constraint value is determined based on the positional relationship between the center of the image block and the preset canvas area. Based on the semantic loss, overlap penalty value and geometric constraint value, the rendering loss is determined. This allows the rendering loss to take into account the semantic consistency between the collage rendering result and the collage prompt text, as well as the rationality of the image block layout structure, thereby improving the accuracy of the target rendering result and the collage quality.
[0080] In some embodiments, step 204 renders each image patch in the discrete image patch set to obtain a mosaic rendering result, including:
[0081] The mosaic rendering result is obtained by rendering each image patch in the discrete image patch set separately using a Gaussian sputtering renderer.
[0082] After the preset rendering conditions are met, the above method also includes:
[0083] Obtain the target geometric parameters of the image block that meets the preset rendering conditions; render the target geometric parameters using a high-definition renderer to obtain the high-definition rendering result.
[0084] Among them, the Gaussian sputtering renderer is a differentiable renderer based on Gaussian sputtering. A high-resolution renderer refers to a renderer capable of producing high-resolution tiled images. For example, a high-resolution renderer is based on bilinear interpolation. A high-resolution renderer can utilize the Grid Sample function to resample and composite the original high-resolution material on any high-resolution canvas according to optimal parameters, ultimately generating a clear, sharp tiled image with a transparent background (i.e., the high-resolution rendering result). Since the Gaussian sputtering renderer may cause blurring in the generated target rendering result, image patch rendering is performed using the Gaussian renderer, and after the rendering conditions are met, the high-resolution renderer generates the high-resolution rendering result.
[0085] In practical applications, a Gaussian sputtering renderer is used to render each image patch in a discrete image patch set individually, resulting in a tiled rendering result. Based on the tiled rendering result and tiled text prompts, a rendering loss is determined. The geometric parameters of the image patches are then adjusted according to the rendering loss, resulting in adjusted geometric parameters. Based on the adjusted geometric parameters, a new pose for the image patches is determined, and the image patches in the new poses are rendered to obtain new tiled rendering results. This process continues until preset rendering conditions are met. The target geometric parameters of each image patch when the preset rendering conditions are met are obtained. Finally, a high-definition renderer is used to render the target geometric parameters, resulting in a high-definition rendering result, i.e., a clear tiled image. A clear tiled image can be, for example, a PNG format image.
[0086] In this embodiment, by using a Gaussian sputtering renderer during the iterative rendering process of image blocks, and after meeting the preset rendering conditions, a high-definition renderer is used to render the target geometric parameters of the image blocks to obtain a high-definition rendering result. This allows the generated collage to retain the clarity of the original image block material and improves the quality of the generated collage.
[0087] In some application scenarios, image collage generation methods include geometric feature-based methods (object space) and semantically driven methods (image space). Geometric feature-based methods typically model the collage problem in object space as a geometric constraint satisfaction or a two-dimensional packing problem. These methods rely on hand-designed geometric descriptors or explicit input shapes. For example, they generate collages by filling the outline of a given target image with multiple small images. Other mosaic generation algorithms are also largely based on similar geometric filling logic. Although these methods perform reasonably well when there is an explicit shape template (mask), they have fundamental drawbacks: first, they cannot understand abstract textual semantics and must rely on users to provide precise outlines; second, for complex and irregular semantic concepts, the hard constraints based on geometry often lead to rigid layouts and a lack of flexibility. Semantically driven methods utilize visual-language models (such as CLIP) to directly control layout generation from text, and this is currently the mainstream research direction. For example, CLIP-CLOP utilizes the image-text matching capability of the CLIP model, maps discrete images to a canvas using differentiable rendering techniques, and optimizes the layout through gradient descent. Specifically, it calculates the CLIP similarity score between the rendered image and the text prompt, and uses this as the optimization target. However, CLIP-CLOP has revealed significant problems in practical applications: First, in order to maximize the CLIP score, the algorithm often performs non-rigid mesh deformation or color mapping modifications on the original material, resulting in material distortion. Second, directly optimizing the CLIP score is prone to getting trapped in local optima. To solve the stability problem of semantic guidance, a diffusion model is introduced as a priori guide. However, diffusion model-based methods tend to focus on generating bounding boxes rather than directly optimizing pixels, and the final generated image is a flattened bitmap where each element is blended with the background and cannot be separated. This means that users cannot obtain structured data including independent layer information, nor can they perform secondary editing on specific elements in the generated result, increasing the cost of secondary editing design. Based on the above problems, this application's embodiments innovatively introduce VSD technology into collage generation, combined with structural regularization constraints, to generate high-quality parametric layouts while maintaining material fidelity.
[0088] For example, a flowchart of the collage generation method is shown below. Figure 4 As shown, by obtaining the discrete image patch set ɛ={e1,e2,…,e N} and collage hint text. Preprocessing of image patches within a discrete image patch set is possible, such as automatically detecting and removing the background from each image patch using a background removal algorithm to generate an RGBA image with an alpha channel. Each image patch e i Assign a set of learnable spatial transformation parameters (i.e., geometric parameters) θi ={t i ,r i ,s i}, where t i =(x i ,y i ) represents the translation coordinates of the image patch center on the canvas, which is initialized to the canvas center (0,0) by default; r i This represents the rotation angle of the image patch. During initialization, it is randomly sampled within the range of [-0.1, 0.1] radians. Introducing a small random perturbation can break the symmetry; s i The uniform scaling ratio of the image patch is represented by the heuristic rule s during initialization. init =(1 / N) 0.5 +0.05 ensures that the image patch size is moderate in the initial state, neither too crowded nor too sparse.
[0089] A Gaussian sputtering renderer (i.e., a microrender) is used to render each pixel of each image patch individually, resulting in a patch rendering result. The image patches are then sorted according to their depth values, and an alpha blending algorithm is used to layer them from back to front, resulting in the corresponding round of mosaic rendering. This not only handles the occlusion relationships between patches but also makes the occlusion order itself an optimizable variable.
[0090] After obtaining the tiling rendering result (i.e., the rendered image x0) through a differentiable renderer, it is then processed through a frozen pre-trained textural image model. It serves as a semantic discriminator and simultaneously trains a lightweight LoRA model. To simulate the current scene distribution, semantic loss is determined based on variational score distillation (VSD) technology. To prevent tiles from stacking haphazardly or flying off the canvas, a structure regularization term can be introduced. Structural regularization terms include overlap penalties. and geometric constraints This approach penalizes image patches whose centers extend beyond the canvas's defined area and constrains the patch scaling to a reasonable range. This prevents oversized patches from occluding other elements or undersized patches from being invisible, ensuring a compact layout. The semantic and structural losses are weighted and summed to obtain the joint loss. (i.e., rendering loss). Based on the backpropagation of the joint loss, the adjusted geometric parameters corresponding to each image patch are determined. The new pose of the image patch is determined based on the adjusted geometric parameters. The above rendering process is repeated for the image patch in the new pose until the number of rendering times reaches the preset number (e.g., 200 or 300 times), or the rendering loss is less than the loss threshold, that is, after the preset rendering conditions are met, the final mosaic image is obtained.
[0091] Alternatively, after satisfying the preset rendering conditions, obtain the optimal parameter θ corresponding to each image patch. final (i.e., target geometric parameters) Using a high-definition renderer based on bilinear interpolation, the original high-definition material (i.e., a set of discrete image patches) is resampled and synthesized on an arbitrary high-resolution canvas according to the optimal parameters, ultimately outputting a clear, sharp collage image (PNG format) with a transparent background. Furthermore, due to the randomness of the VSD loss, by changing the random seed, multiple distinct yet semantically consistent layout schemes can be generated without altering the input material and text prompts, providing a wealth of creative options.
[0092] In the above embodiments, a text-driven semantic guidance mechanism based on variational score distillation (VSD) is used. This mechanism generates semantic gradients by utilizing the noise prediction differences between a frozen pre-trained text-to-image diffusion model and a trainable LoRA model. This achieves direct mapping from text to layout without the need for shape templates, accurately guiding each image block to the position that best matches the collage prompt text description. Gaussian kernel weighted sputtering is used to achieve continuous differentiability of discrete pixel mapping, establishing a complete gradient propagation path from image space loss to geometric transformation parameters. During rendering, only the geometric parameters of the image blocks are adjusted, strictly limiting the optimization variables to affine transformation parameters to ensure material fidelity. The final output is editable parameterized layout data rather than a flat bitmap. A structure regularization loss design based on alpha cumulative penalty and boundary constraints is used. This applies soft constraints through global statistics in the image space, jointly optimizing with semantic loss to achieve semantic alignment and structural rationality of the layout. This generates collage images that semantically match the collage prompt text and ensures that the image blocks in the collage image are clear and orderly distributed.
[0093] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0094] Based on the same inventive concept, this application also provides a collage generation apparatus for implementing the collage generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the collage generation apparatus provided below can be found in the limitations of the collage generation method described above, and will not be repeated here.
[0095] In one exemplary embodiment, such as Figure 5 As shown, a collage generation device 500 is provided, including: a data acquisition module 502, an image block rendering module 504, a rendering loss determination module 506, a geometric parameter adjustment module 508, and a rendering result acquisition module 510, wherein:
[0096] Data acquisition module 502 is used to acquire a set of discrete image patches and collage prompt text;
[0097] The image patch rendering module 504 is used to render each image patch in the discrete image patch set to obtain the mosaic rendering result;
[0098] The rendering loss determination module 506 is used to determine the rendering loss based on the tiling rendering result and the tiling hint text;
[0099] The geometric parameter adjustment module 508 is used to adjust the geometric parameters of the image block according to the rendering loss to obtain the adjusted geometric parameters;
[0100] The rendering result acquisition module 510 is used to determine the new pose of the image block based on the adjusted geometric parameters, and to render the image block in the new pose to obtain a new image block rendering result, until the preset rendering conditions are met to obtain the target rendering result.
[0101] In some embodiments, the image block rendering module 504 is further configured to detect and remove the background of each image block in the discrete image block set to obtain an image block including transparency information; and to render the image block including transparency information to obtain a mosaic rendering result.
[0102] In some embodiments, the image block rendering module 504 is further configured to determine the position of the target pixel of the target image block in the discrete image block set; determine the rendering area based on the position of the target pixel and the Gaussian kernel size; determine the distance between each pixel in the rendering area and the target pixel, and determine the rendering weight corresponding to each pixel based on the distance; and render the pixels of the target pixel based on the rendering weight and pixel value corresponding to each pixel in the rendering area to obtain the rendering result of the target pixel.
[0103] In some embodiments, the rendering loss determination module 506 is further configured to determine the discriminant matching degree between the collage rendering result and the collage prompt text based on the text-to-image discriminant model; perform alignment prediction on the collage rendering result and the collage prompt text based on the scene distribution simulation model to obtain the alignment confidence; and determine the rendering loss based on the discriminant matching degree and the alignment confidence.
[0104] In some embodiments, the rendering loss determination module 506 is further configured to determine semantic loss based on the discrimination matching degree and alignment confidence; determine the transparency accumulation value of each pixel in the tile rendering result, and determine the overlap penalty value based on the transparency accumulation value; determine the geometric constraint value based on the positional relationship between the center of the image patch and the preset canvas area; and determine the rendering loss based on the semantic loss, the overlap penalty value and the geometric constraint value.
[0105] In some embodiments, the image block rendering module 504 is further configured to render each image block in the discrete image block set using a Gaussian sputtering renderer to obtain a mosaic rendering result; after satisfying preset rendering conditions, obtain the target geometric parameters of the image block corresponding to the preset rendering conditions; and render the target geometric parameters using a high-definition renderer to obtain a high-definition rendering result.
[0106] Each module in the aforementioned collage generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0107] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data related to the collage generation method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a collage generation method.
[0108] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0109] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0110] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0111] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0113] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0114] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0115] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for generating collages, characterized in that, The method includes: Retrieve a set of discrete image patches and collage hint text; Each image block in the discrete image block set is rendered separately to obtain the mosaic rendering result; The discriminant matching degree between the collage rendering result and the collage prompt text is determined based on the text-to-image discrimination model; the text-to-image discrimination model is used to generate a reference collage image corresponding to the collage prompt text, and to determine the matching degree between the reference collage image and the collage rendering result; Alignment prediction is performed on the collage rendering result and the collage prompt text based on the scene distribution simulation model to obtain alignment confidence; the scene distribution simulation model is used to simulate scene distribution based on the collage prompt text, and to perform alignment prediction on the simulated scene distribution and the collage rendering result. The semantic loss is determined based on the discriminant matching degree and the alignment confidence degree. Determine the cumulative transparency value of each pixel in the collage rendering result, and determine the overlap penalty value based on the cumulative transparency value; Determine the geometric constraint values based on the positional relationship between the center of the image block and the preset canvas area; The rendering loss is determined based on the semantic loss, the overlap penalty value, and the geometric constraint value. The geometric parameters of the image patch are adjusted according to the rendering loss to obtain the adjusted geometric parameters; The new pose of the image block is determined based on the adjusted geometric parameters, and the image block in the new pose is rendered to obtain a new mosaic rendering result until the preset rendering conditions are met to obtain the target rendering result.
2. The method according to claim 1, characterized in that, The step of rendering each image patch in the discrete image patch set to obtain the mosaic rendering result includes: The background of each image block in the discrete image block set is detected and removed to obtain an image block including transparency information; The image blocks, including transparency information, are rendered to obtain a mosaic rendering result.
3. The method according to claim 1, characterized in that, The step of rendering each image patch in the discrete image patch set to obtain the mosaic rendering result includes: Determine the position of the target pixel in the target image block within the set of discrete image blocks; The rendering area is determined based on the position of the target pixel and the size of the Gaussian kernel; Determine the distance between each pixel within the rendering area and the target pixel, and determine the rendering weight corresponding to each pixel based on the distance; Based on the rendering weight and pixel value corresponding to each pixel in the rendering area, the pixels of the target pixel are rendered to obtain the rendering result of the target pixel.
4. The method according to claim 1, characterized in that, The step of determining the semantic loss based on the discriminant matching degree and the alignment confidence degree includes: The difference between the alignment confidence and the discriminative matching degree is used as the semantic loss.
5. The method according to claim 1, characterized in that, The determination of rendering loss based on the semantic loss, the overlap penalty value, and the geometric constraint value includes: The structural loss is obtained by weighted summation of the overlap penalty value and the geometric constraint value; The semantic loss and the structural loss are weighted and summed to obtain the rendering loss.
6. The method according to claim 1, characterized in that, The step of rendering each image patch in the discrete image patch set to obtain the mosaic rendering result includes: Each image block in the discrete image block set is rendered separately using a Gaussian sputtering renderer to obtain the mosaic rendering result; After the preset rendering conditions are met, the method further includes: Obtain the target geometric parameters of the image block that meets the preset rendering conditions; The target geometric parameters are rendered using a high-definition renderer to obtain a high-definition rendering result.
7. A collage generation device, characterized in that, The device includes: The data acquisition module is used to acquire a set of discrete image patches and collage prompt text; The image patch rendering module is used to render each image patch in the discrete image patch set to obtain the patch rendering result; A rendering loss determination module is used to determine the discriminative matching degree between the collage rendering result and the collage prompt text based on a text-based image discrimination model; the text-based image discrimination model is used to generate a reference collage image corresponding to the collage prompt text and to determine the matching degree between the reference collage image and the collage rendering result; an alignment prediction is performed on the collage rendering result and the collage prompt text based on a scene distribution simulation model to obtain an alignment confidence degree; the scene distribution simulation model is used to simulate a scene distribution based on the collage prompt text and to perform alignment prediction on the simulated scene distribution and the collage rendering result; a semantic loss is determined based on the discriminative matching degree and the alignment confidence degree; the transparency accumulation value of each pixel in the collage rendering result is determined, and an overlap penalty value is determined based on the transparency accumulation value; a geometric constraint value is determined based on the positional relationship between the image patch center and a preset canvas area; and a rendering loss is determined based on the semantic loss, the overlap penalty value, and the geometric constraint value. A geometric parameter adjustment module is used to adjust the geometric parameters of the image block according to the rendering loss to obtain the adjusted geometric parameters; The rendering result acquisition module is used to determine the new pose of the image block based on the adjusted geometric parameters, and to render the image block in the new pose to obtain a new image block rendering result, until the preset rendering conditions are met to obtain the target rendering result.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.