Method and device for expanding image based on object reference image
By processing the original image and the object reference image, an image with consistent features and background adaptation with the target object is generated, which solves the problem of inconsistent content after size adjustment and achieves the coherence and consistency of image content.
Patent Information
- Application Number
- CN202510994995.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies struggle to ensure content consistency and coherence when resizing images, especially in cropping or AI-based image enlargement methods, which often result in distorted character features or mismatched backgrounds.
By acquiring the original image and the object reference image, edge expansion processing is performed to generate the target image and mask. Combined with the feature encoding of the object feature map, an image with the same features as the target object and adapted to the background is generated in the area defined by the mask, ensuring the continuity of the content after size adjustment.
It achieves continuity and consistency between image content and original content after any size adjustment, avoiding the problems of content loss or distortion in traditional methods, and ensuring the continuity of character features and background.
Smart Images

Figure CN120953433A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image technology, and in particular to a method and apparatus for expanding an image based on an object reference map. Background Technology
[0002] Once character posters and promotional materials in fields such as animation and games are produced, their size and aspect ratio are usually fixed, lacking flexibility for adjustment. When there is a need to modify the size or aspect ratio (such as converting a landscape poster to a portrait poster), existing technologies have significant limitations: traditional image processing methods can only achieve this through cropping, which leads to the loss of image content; while some AI image expansion methods can automatically fill in edges, they are difficult to guarantee content consistency, often resulting in distortion of character features (such as hairstyles and clothing details), or the filled background not matching the style of the original image, ultimately failing to ensure the continuity and consistency of the content after size adjustment with the original content. Summary of the Invention
[0003] This application provides a method and apparatus for expanding an image based on an object reference diagram to solve the problem of inconsistency between the content after size adjustment and the original content.
[0004] In a first aspect, this application provides a method for augmenting an image based on an object reference graph, the method comprising:
[0005] Obtain the original image and a complete object reference image of the target object, wherein the original image contains a portion of the target object;
[0006] By performing edge expansion processing on the original image, a target image containing the edge expansion area to be filled and a first mask of the target image are generated;
[0007] By performing background removal processing on the complete object reference image, an object feature image with a clean background is obtained, and a second mask of the object feature image is generated.
[0008] The target image and the object feature map are concatenated into a composite image, and the first mask and the second mask are concatenated into a composite mask;
[0009] The background features and object appearance features of the composite image are extracted, and combined with the feature encoding of the object feature image, an image that is consistent with the features of the target object and adapted to the background is generated in the area to be filled and expanded within the composite mask.
[0010] Optionally, generating a target image containing the area to be filled and a first mask of the target image by performing edge-expanding processing on the original image includes:
[0011] The original image is input into the augmentation node, and augmentation parameters are input into the augmentation node, wherein the augmentation parameters are used to indicate the number of pixels that need to be augmented in the corresponding direction;
[0012] According to the expansion parameters, gray pixels are filled around the original image to generate a target image containing the expansion area to be filled, and a first mask of the target image is generated.
[0013] Specifically, when generating the first mask of the target image, the region corresponding to the original image is set as the reserved region, and the region corresponding to the area to be filled and expanded is set as the region to be generated, so as to define the original content to be retained and the expanded content to be generated through the first mask.
[0014] Optionally, the second mask for generating the object feature map includes:
[0015] Generate a second mask that is the same size as the complete object reference map and is filled entirely with black, wherein the black filling is used to indicate that the object feature map region corresponding to the second mask is a feature reference region that does not need to be modified.
[0016] Optionally, extracting background features and object appearance features from the composite image, and combining them with the feature encoding of the object feature map, guides the generation of an image that matches the target object features and adapts to the background within the area to be filled and expanded by the composite mask, including:
[0017] The composite image and the composite mask are input into the feature extraction node, and the output is the condition information that guides the generation of the redrawn image. The condition information includes the background features in the original image, the object appearance features in the complete object reference image, and the redrawn area located in the area to be filled and expanded as indicated by the composite mask.
[0018] After modifying the object feature map to a standard size, input it into the feature encoding node and output the feature encoding used to enhance the core features of the target object;
[0019] The conditional information for generating the guided redrawing image and the feature-encoded input image generation node are used to obtain a complete image containing the original image, the redrawing image of the area to be filled and expanded, and the stitched object feature map.
[0020] The object feature map is cropped to obtain the final image composed of the original image and the redrawn image.
[0021] Optionally, obtaining the redrawn image includes:
[0022] Generate suitable background content based on the background features;
[0023] Generate content that is consistent with the overall shape of the target object based on the object's appearance features;
[0024] The consistency of detailed features of the target object is enhanced based on the feature encoding.
[0025] Optionally, concatenating the target image and the object feature map into a composite image, and concatenating the first mask and the second mask into a composite mask, includes:
[0026] The target image and the object feature map are concatenated side by side to form a composite image, and the first mask and the second mask are concatenated side by side to form a composite mask. The concatenation method of the first mask and the second mask corresponds to the composite image, and the concatenation order of the target image and the object feature map can be interchanged.
[0027] Optionally, the background features include scene style, color tone and environmental element features, and the object appearance features include the outline, shape and core visual features of the target object.
[0028] Secondly, this application provides an image augmentation device based on an object reference map, the device comprising:
[0029] The acquisition module is used to acquire the original image and the complete object reference image of the target object, wherein the original image contains part of the target object;
[0030] The edge expansion module is used to generate a target image containing the edge expansion area to be filled and a first mask of the target image by performing edge expansion processing on the original image;
[0031] The processing module is used to obtain an object feature map with a clean background by performing background removal processing on the complete object reference map, and to generate a second mask for the object feature map;
[0032] The stitching module is used to stitch the target image and the object feature map into a composite image, and to stitch the first mask and the second mask into a composite mask;
[0033] The generation module is used to extract background features and object appearance features from the composite image, combine them with the feature encoding of the object feature image, and guide the generation of an image that is consistent with the features of the target object and adapted to the background in the area to be filled and expanded within the composite mask.
[0034] Thirdly, this application provides an electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.
[0035] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for executing the object reference graph-based image augmentation method described in any of the preceding claims of this application.
[0036] The technical solutions provided in this application have the following advantages compared with the prior art:
[0037] First, the original image is expanded to generate a target image containing the original image and the area to be filled, thus constructing a framework of the required size. Next, an object feature map with a clean background is introduced to define the overall shape of the object to be redrawn, while its detailed features are extracted through feature encoding as a reference. Then, the target image and the object feature map are stitched together to form a composite image, and a composite mask is generated to precisely define the redrawing area. Finally, based on the background features extracted from the composite image, the overall shape of the object, and the detailed encoding, new content is generated within the redrawing area defined by the mask. This process does not require cropping the original image, preserving all information. Through object feature guidance and mask constraints, it ensures that the expanded content is consistent with the target object's features, and the background adapts to the original image, ultimately enabling flexible modification to any size while maintaining the continuity between character features and the background. This ensures that the content of the resized image is consistent with the original content. Attached Figure Description
[0038] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0041] Figure 1 A schematic diagram of an object reference graph-based augmented image system provided in this application embodiment;
[0042] Figure 2 A flowchart illustrating a method for augmenting an image based on an object reference graph, as provided in this application embodiment;
[0043] Figure 3 An overall flowchart of an object reference graph-based image augmentation provided in this application embodiment;
[0044] Figure 4 A method flowchart provided as an example of an embodiment of this application;
[0045] Figure 5 A schematic diagram of the structure of an object reference map-based image augmentation device provided in this application embodiment;
[0046] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0048] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0049] The application scenarios of this application include, but are not limited to: animation and games, film and television characters, virtual idols, secondary creation of IP characters, or e-commerce clothing model images.
[0050] Optionally, in the embodiments of this application, the above-described image augmentation method based on object reference graphs can be applied to, for example... Figure 1 The hardware environment shown consists of terminal 101 and server 103. Figure 1 As shown, server 103 connects to terminal 101 via a network. Terminal 101 uploads the original image of the target object and a complete object reference image. Server 103 resizes the image and generates a redrawn image. Image processing can also be performed by terminal 101 itself. Terminal 101 includes, but is not limited to, PCs, mobile phones, tablets, etc.
[0051] The following will describe in detail, with reference to specific implementation methods, an image augmentation method based on an object reference graph provided in this application, taking its application to a terminal as an example. Figure 2 As shown, the specific steps are as follows:
[0052] Step 201: Obtain the original image and complete object reference image of the target object, wherein the original image contains part of the target object;
[0053] Step 202: By performing edge expansion processing on the original image, a target image containing the area to be filled and the first mask of the target image are generated;
[0054] Step 203: By removing the background from the complete object reference image, an object feature map with a clean background is obtained, and a second mask for the object feature map is generated;
[0055] Step 204: Concatenate the target image and the object feature map into a composite image, and concatenate the first mask and the second mask into a composite mask;
[0056] Step 205: Extract background features and object appearance features from the composite image, combine them with the feature encoding of the object feature map, and guide the generation of an image that is consistent with the features of the target object and matches the background in the area to be filled and expanded within the scope defined by the composite mask.
[0057] First, some terms involved in the embodiments of this application will be explained, including the following content.
[0058] Original image: refers to the image to be processed that contains part of the target object, such as a character poster in anime or game (showing only part of the character's body and part of the background), which is the base image for image expansion operations.
[0059] Complete object reference image: A reference image containing the complete features of the target object, such as a character design (showing the full-body details of the character, clothing, hairstyle, etc.), used to provide complete information about the object's features.
[0060] Area to be filled and expanded: The area added after expanding the edges of the original image, where content needs to be generated. For example, when changing a landscape poster to a portrait poster, the blank areas added at the top and bottom of the image.
[0061] First mask: A black and white mask with the same size as the target image. The area corresponding to the original image is black (the area to be retained), and the area corresponding to the area to be filled and expanded is white (the area to be generated). It is used to define the original content to be retained and the expanded content to be generated.
[0062] Object feature map: The image obtained after removing the background from the complete object reference image. The background is a pure color (such as white), and only the complete shape of the target object is preserved. It is used to extract the core features of the object.
[0063] Second mask: A completely black mask with the same size as the object feature map, indicating that the object feature map area is a feature reference area that does not need to be modified, thus avoiding interference with the reference features during the generation process.
[0064] Composite image: An image formed by stitching the target image and the object feature map side by side, enabling the model to simultaneously obtain the background information of the original image and the complete object information of the object feature map.
[0065] Composite mask: A mask formed by concatenating the first mask and the second mask, simultaneously defining the region to be generated in the target image and the feature reference region of the object feature map.
[0066] Background features: Information extracted from the original image, such as scene style (e.g., ancient style, science fiction), color tone (e.g., cool tone, warm tone), and environmental elements (e.g., trees, buildings).
[0067] Object appearance features: The outline, shape (such as posture, body shape) and core visual features (such as clothing texture, hairstyle) of the target object extracted from the complete object reference image.
[0068] Feature encoding: After standardizing the object feature map (e.g., adjusting it to 1024×1024 size), the encoded information generated by feature encoding nodes (e.g., FLUX Redux) is used to enhance the core features of the target object (e.g., facial details, accessories).
[0069] comfyUI: is a node-based Stable Diffusion interface.
[0070] Pad Image for Outpainting node: A built-in node in ComfyUI, its function is to add a padding area to the original image, so that the subsequent generative model can synthesize a new image that is consistent with the original content in the expanded area.
[0071] FLUX Redux is a style transfer model developed by Black Forest Labs, and there are related nodes in ComfyUI.
[0072] Inpaint Model Conditioning node: A built-in node in ComfyUI, mainly used to guide the model to perform image inpainting or generation in specific areas (such as masked areas).
[0073] KSampler node: A built-in node in ComfyUI that uses the provided model and positive / negative conditions to generate a new version of a given latent image.
[0074] In step 201, the terminal acquires an original image containing a portion of the target object (e.g., a landscape anime character poster showing only the character's upper body and a tree background on the left), and a complete object reference image of the character (e.g., a character design sketch showing the character's entire body, green cloak, short brown hair, and details of the longsword held in hand). The acquired images are then input into the target model (e.g., comfyUI). The original image provides the basic content for the image expansion, while the complete object reference image provides the general and detailed features of the target object, ensuring that subsequent image expansion has a clear original reference and complete feature support, preventing the content from deviating from the original settings.
[0075] In step 202, the target model inputs the original landscape poster into the expansion node (such as the Pad Image forOutpainting node), sets expansion parameters to determine the expansion direction and number of pixels, and fills the top and bottom of the original image with gray pixels according to the expansion parameters to generate the target image (size 1080×1920, with the original image in the middle and gray areas to be filled at the top and bottom); at the same time, a first mask is generated, in which the original image area is black (preserved), and the top and bottom gray areas are white (to be generated), clarifying the original content to be preserved and the expansion content to be generated. For example, when changing a 1920×1080 landscape game poster to a 1080×1920 portrait poster, the expansion node is set to expand by 420 pixels at the top and bottom, generating a target image containing gray areas to be filled and a corresponding first mask.
[0076] In step 203, the target model performs background removal processing on the complete object reference image (character design image). The original background of the design image is removed and replaced with a solid color background, resulting in an object feature image containing only the complete form of the character (highlighting features such as the green cape and brown short hair). Simultaneously, a second mask (completely black) with the same size as the object feature image is generated, indicating that this area is the feature reference area. This avoids interference with the object features during the generation process and ensures the accuracy and reliability of the extracted object features. For example, the background of a character design image with a complex background is removed, the character form is preserved, and it is replaced with a white background, generating the object feature image and the completely black second mask.
[0077] In step 204, the target model stitches the target image (1080×1920) generated in step 202 with the object feature map (1080×1920) to form a composite image (the left side is the target image and the right side is the object feature map), so that the model can simultaneously obtain the original background and complete object information; at the same time, the first mask and the second mask are stitched together to form a composite mask (the left side is the first mask and the right side is the second mask), ensuring that the mask area corresponds to the composite image and accurately defines the area to be generated and the feature reference area.
[0078] In step 205, the target model inputs the composite image and composite mask into the feature extraction node to extract background features (ancient style scene, cool color tone, trees on the left) and object appearance features (character's green cloak, brown short hair, long sword outline) from the original image; after adjusting the object feature map to 1024×1024 size, it is input into the feature encoding node to generate feature encoding (enhancing facial details and cloak texture); the background features, object appearance features, feature encoding, and composite mask are input into the image generation node (such as the KSampler node) to generate content in the upper and lower fill areas defined by the composite mask: based on the background features, grass and distant mountains with the same style as the trees on the left are generated; based on the object appearance features, the lower half of the character is generated (in accordance with the design). Figure 1 The details of the character's waist accessories (such as the green cape hem and boots) are enhanced based on feature encoding; finally, the feature map of the object on the right is cropped to obtain a vertical poster (original upper body + added lower body and background).
[0079] In this application, the original image is first expanded to generate a target image containing the original image and the expanded area to be filled, thus constructing a framework of the required size. Next, an object feature map with a clean background is introduced to clarify the overall shape of the object to be redrawn, while its detailed features are extracted through feature encoding as a reference. Then, the target image and the object feature map are stitched together to form a composite image, and a composite mask is generated to precisely define the redrawing area. Finally, based on the background features extracted from the composite image, the overall shape of the object, and the detailed encoding, new content is generated within the redrawing area defined by the mask. This process does not require cropping the original image, preserving all information. Through object feature guidance and mask constraints, it ensures that the expanded content is consistent with the target object features, and the background adapts to the original image, ultimately achieving flexible modification to any size while ensuring the continuity between character features and the background. This ensures that the image content after size adjustment is consistent with the original content.
[0080] As an optional implementation, step 202, by performing edge expansion processing on the original image to generate a target image containing the area to be filled and an initial mask for the target image, includes the following:
[0081] Step S11: Input the original image into the augmentation node of the target model, and input the augmentation parameters into the augmentation node, wherein the augmentation parameters are used to indicate the number of pixels that need to be augmented in the corresponding direction;
[0082] Step S12: Fill the perimeter of the original image with gray pixels according to the expansion parameters to generate a target image containing the expansion area to be filled, and generate the first mask of the target image.
[0083] Step S13: When generating the first mask of the target image, the region corresponding to the original image is set as the reserved region, and the region corresponding to the area to be filled and expanded is set as the region to be generated, so as to define the original content to be retained and the expanded content to be generated through the first mask.
[0084] In step S11, the target model inputs the original image into its own model's expansion node (such as the Pad Image forOutpainting node), and inputs the number of pixels to be expanded in the four directions (up, down, left, and right) in the four parameters of the node. These expansion parameters directly determine the expansion range of the original image in each direction.
[0085] Step S12: The expansion node fills gray pixels in the four directions of the original image (up, down, left, right) according to the input expansion parameters, thereby generating a target image containing the expansion area to be filled; at the same time, the expansion node also generates a black and white mask1 (i.e., the first mask) with the same size as the target image.
[0086] Step S13: During the generation of the black and white mask1, the area corresponding to the original image is set to black (i.e., the area to be retained), while the area corresponding to the gray area to be filled and expanded is set to white (i.e., the area to be generated). This setting clearly defines the original image content to be retained and the expanded content to be generated, providing clear area guidance for the subsequent image generation process.
[0087] This application can precisely control the expansion range and direction of the original image. The filling of gray pixels provides a transitional visual cue for the area to be generated, while the clear partitioning of the first mask effectively avoids accidental modification of the original image content during subsequent generation, ensuring the integrity of the original content and the relevance of the expanded content, laying the foundation for generating expanded content consistent with the style and features of the original image.
[0088] As an optional implementation, in step 203, generating the second mask of the object feature map includes: generating a second mask with the same size as the complete object reference map and filled entirely with black through the target model, wherein filling entirely with black is used to indicate that the object feature map region corresponding to the second mask is a feature reference region that does not need to be modified.
[0089] The target model creates an image with the exact same size as the complete object reference image as a base carrier, with all pixels of this carrier filled with black to form a second mask. In the image generation logic of this application, black is given the semantic meaning of a feature reference area that does not need to be modified. That is, the area covered by black in the second mask, when mapped to the subsequent complete mask (mask2), will be identified as an area that needs to strictly reference the original object features (such as the character features in a character design drawing). The target model must use the object features of this area as a benchmark during the generation process and must not modify or replace them independently.
[0090] The black area of the second mask is the same size as and completely corresponds to the character design drawing. Through explicit semantics that require no modification, it provides a clear reference benchmark for character features in the target model. In the subsequent feature extraction process of the Inpaint Model Conditioning node, the object details in this area (such as hairstyle, clothing, posture, etc.) are captured first and used as the basis for generation. This avoids the model generating character features that do not match the design drawing due to the lack of clear reference, effectively solving the core problem of poor character consistency in traditional AI image expansion.
[0091] As an optional implementation, in step 205, the background features and object appearance features in the composite image are extracted, and the feature encoding of the object feature map is combined to guide the generation of an image that is consistent with the features of the target object and adapted to the background in the area to be filled and expanded within the composite mask. This includes the following process.
[0092] Step S21: Input the composite image and composite mask into the feature extraction node in the target model, and output the condition information to guide the generation of the redraw image. The condition information includes the background features in the original image, the appearance features of the object in the complete object reference image, and the redraw area located in the area to be filled and expanded as indicated by the composite mask.
[0093] Step S22: After modifying the object feature map to the standard size, input it into the feature encoding node in the target model, and output the feature encoding used to enhance the core features of the target object;
[0094] Step S23: Input the conditional information and feature encoding for guiding the generation of the redrawn image into the image generation node in the target model to obtain a complete image containing the original image, the redrawn image of the area to be filled and expanded, and the stitched object feature map;
[0095] Step S24: Crop the object feature map to obtain the final image composed of the original image and the redrawn image.
[0096] In step S21, the target model inputs the composite image obtained from the previous processing (i.e., the target image after the original image is padded with gray pixels by augmentation parameters and the object feature map of the blank background) and the composite mask (i.e., the mask mask1 of the target image is concatenated with a black mask of the same size) into the feature extraction node (such as the Inpaint Model Conditioning node). This node uses an image feature parsing algorithm to jointly extract the visual information in the composite image and the regional semantics in the composite mask, and finally outputs the conditional information to guide the generation of the redrawn image, which specifically includes the following:
[0097] Background features of the original image: Extracted from the region corresponding to the original image in the composite image, covering the color tone (such as cool tones, warm tones), scene style (such as ancient Chinese ink painting style, science fiction metal style), environmental element features (such as grass texture, building outlines), etc., to ensure that the newly generated expanded edge content is consistent with the original background style.
[0098] The appearance features of the complete object reference image are extracted from the object feature map spliced from the composite image, including the core visual features of the target object (such as clothing patterns, hair color and eye color), shape (such as body posture and range of motion), and outline (such as shoulder lines, skirt hem curvature, and hairstyle outline), providing an accurate reference for generating consistent objects for the model.
[0099] Redrawing areas for the expansion region to be filled: Based on the semantics of the composite mask (white areas are areas to be generated, black areas are areas to be retained), the spatial range of the new content that the model needs to generate is clearly defined, avoiding erroneous modifications to the core areas of the original image and the complete object reference image.
[0100] In step S22, the target model inputs the modified object feature map (sized to a standard 1024*1024, with the background replaced with white to remove interfering elements) into the feature encoding node (such as a FLUX / Redux related node). This node uses style transfer and feature enhancement algorithms to deeply encode the core features of the character in the object feature map: on the one hand, it filters out residual background interference, focusing on key features such as the object's outline, proportions, and clothing texture; on the other hand, through standardized encoding, it transforms these features into vector forms (feature encoding) that the model can directly parse, which is used to anchor character features in subsequent generation processes, avoiding character appearance distortion or style shift due to the expansion of the image range.
[0101] In step S23, the target model inputs the guided redrawing condition information (including background features, object appearance features, and redrawing area) output in step S21 and the object core feature encoding output in step S22 into an image generation node (such as a KSampler node). This node, based on the diffusion model generation logic, uses the condition information as a framework and the feature encoding as a reference to generate a redrawn image in the area to be filled (white mask area). The redrawn image includes the following content.
[0102] For parts of the expanded area that involve character extension (such as the arms or skirt of the original poster character needing to be extended), the image generation node will generate features that are consistent with the target object based on the specific features in the feature encoding (generating limb or clothing extensions that are consistent with the original character based on the character's posture and clothing features); for parts of the expanded area that only need to add background, the image generation node will refer to the background features of the original poster to generate environmental content that is consistent with the original style (e.g., if the original poster background is a forest, the expanded area will generate consistent trees and lighting).
[0103] The final output image consists of three parts: the original image, the redrawn image of the area to be filled and expanded, and the stitched object feature map.
[0104] In step S24, the target model performs cropping processing on the complete image generated in step S23. Specifically, based on the stitching boundaries in the composite image, it precisely removes the object feature maps used for reference, retaining only the original image and the newly generated redrawn image to form the final image. The cropping process is achieved through pixel coordinate positioning, ensuring that the cut boundaries are neat and do not affect the continuity between the original image and the expanded area.
[0105] In this application, the original image background features provide style constraints for the expanded content. Referencing these background features ensures that the new content and the original background are naturally connected in terms of tone, texture, and style (e.g., if the original poster has a retro oil painting style, the expanded area will not have a cartoon pixel style), avoiding a visual disconnect between the background and the expanded content. The core features of the target object are strengthened through feature encoding nodes. The encoded features and conditional information are jointly input into the image generation node, so that the node always uses the object features of the original object feature map as the benchmark when redrawing the image, avoiding the character deformation caused by the model's autonomous guessing in traditional AI image expansion (such as messed-up hairstyles or sudden changes in clothing style).
[0106] As an optional implementation, step S23, obtaining the redrawn image includes the following:
[0107] Step S231: Generate adapted background content based on background features;
[0108] Step S232: Generate content that matches the overall shape of the target object based on its appearance features;
[0109] Step S233: Enhance the consistency of detailed features of the target object based on feature encoding.
[0110] In step S231, within the processing flow of the image generation node (KSampler node), firstly, based on background features (including the color tone, lighting style, and scene element types of the original image), suitable background content is generated in the area to be filled and expanded. When generating background content, the image generation node uses a feature matching algorithm to continue the visual attributes of these background elements in the expanded area, ensuring that the new content seamlessly connects with the original background in terms of style and element correlation.
[0111] In step S232, the image generation node combines the object's appearance features (including the overall outline, limb posture, proportions, and spatial relationship with the background) to generate content that is consistent with the overall form of the target object. For example, if the target object in the original image is a standing figure in ancient style, its appearance features include a sideways stance, a raised sword in the right hand, a slightly bent left leg, and a 1:3 ratio between its height and the background columns. If the expanded area involves the extension of the figure's limbs (e.g., the left side needs to show the figure's left shoulder and back), the image generation node will generate natural limb connections based on the above posture features. The tilt angle of the left shoulder matches the sideways stance, and the direction of the robe folds on the back is consistent with the body's rotation trend. At the same time, the newly generated limbs still maintain a 1:3 ratio with the background columns, avoiding the problem of disproportion between the figure and the background after the expanded area. If the expansion area does not involve the extension of the character's limbs (only the background is expanded), the image generation node will ensure that the relative positions of the generated background elements (such as the grass and objects next to the character) and the character conform to the original spatial logic (such as the height of the grass and the objects being placed within the range that the character can reach).
[0112] In step S233, the image generation node further enhances the detail coherence of the target object during the generation process through feature encoding (containing detailed feature data of the target object, such as facial contours, clothing patterns, jewelry shapes, skin texture, etc.). The feature encoding includes standardized detail parameters: for example, the facial features of the target object may be encoded as willow-leaf eyebrows, almond eyes, a nose bridge curvature of 30°, and full lips; clothing details may be encoded as a brocade fabric gloss parameter of 0.7, an embroidery pattern of lotus scrolls, and a pattern spacing of 2cm. When the expansion area involves the extended parts of the target object (such as the sleeves, hem, and hair of a person), the image generation node calls these encoded data to ensure that the newly generated details are strictly consistent with the original features. For example, the brocade gloss of the sleeves, the shape and arrangement density of the lotus scroll pattern are exactly the same as the original sleeves; the thickness, curl, and color of the hair are seamlessly connected with the original hair details. Even if the expanded area does not directly involve the target object, the image generation node will verify the influence of background elements on the details of the target object through feature encoding (such as whether the transition of the shadow on the face of a person is consistent with the original lighting logic under changes in light and shadow), so as to avoid distortion of details.
[0113] In this application, by analyzing the multi-dimensional features of the original background, the newly generated background content is made to form a unified whole with the original background in terms of style and element correlation, thus solving the problem of the disjointed style between the background of the expanded border area and the original image in traditional image expansion. This fusion makes the final image visually more natural and meets the requirements of image integrity in scenarios such as posters.
[0114] Based on the constraints of the object's appearance features, posture, proportion, and spatial relationship, it is ensured that the extended parts of the target object (such as limbs and clothing) in the expanded area are consistent with the shape of the original image, avoiding problems such as limb distortion, disproportion, and spatial misalignment that are prone to occur when AI generates images autonomously. This ensures the integrity of the target object as the core of the image.
[0115] By enhancing details through feature encoding, the details of the target object in the expanded area are kept highly consistent with the original image, solving the pain point of detail loss or variation in traditional expanded images (such as the disappearance of the character's unique necklace in the expanded area or the change of clothing pattern).
[0116] As an optional implementation, step 204, which involves stitching the target image and the object feature map into a composite image and stitching the first mask and the second mask into a composite mask, includes: stitching the target image and the object feature map side by side into a composite image and stitching the first mask and the second mask side by side into a composite mask, wherein the stitching method of the first mask and the second mask corresponds to the composite image, and the stitching order of the target image and the object feature map can be interchanged.
[0117] The target model stitches the target image and object feature map side-by-side into a complete image, i.e., a composite image. The specific stitching logic is as follows: the two images are arranged adjacently in the horizontal or vertical direction, with clear stitching boundaries, and each maintains its original size ratio. (The specific direction can be set according to actual needs.)
[0118] Simultaneously, the target model concatenates the first mask (mask1, a black-and-white mask where the original character poster area is black and the extended area is white) and the second mask (a pure black mask with the same size as the object feature map) in the same parallel manner to form a composite mask (mask2): the concatenation direction and boundary position correspond completely with the composite image (e.g., if the target image is on the left and the object feature map is on the right in the composite image, then the first mask is on the left and the second mask is on the right in the composite mask), ensuring that the region division of the composite mask matches the content region of the composite image one-to-one. Alternatively, the correspondence between the first mask and the target image, and the correspondence between the second mask and the object feature map, can be set, thus eliminating the need to set a one-to-one correspondence between the mask position in the composite mask and the image position in the composite image. The areas of the target image and the object feature map in the composite image correspond to black (feature reference areas that need to be preserved) in the composite mask; the gray extended areas in the composite image correspond to white (extended areas that need to be redrawn) in the composite mask.
[0119] Furthermore, the stitching order of the target image and the object feature map can be flexibly interchanged (e.g., the target image on the right and the object feature map on the left, or the target image on top and the object feature map on the bottom), as long as the stitching order of the composite mask is consistent with it, in order to adapt to the feature reference requirements in different scenarios.
[0120] Figure 3 This is an overall schematic diagram of the process of expanding the image based on the object reference diagram provided in the embodiments of this application, including the following contents.
[0121] 1. Prepare materials.
[0122] Get the original character poster (1080×1920) and character design (1080×1920, background changed to white), and determine the parameters to be expanded (how many pixels to add in the top, bottom, left and right).
[0123] 2. Basic expansion (Pad Image for Outpainting node).
[0124] Fill the edges of the original poster with gray to create a new poster with borders to be expanded. Figure 1 At the same time, a mask mask1 is generated (black = retain the original content, white = the edge to be redrawn).
[0125] 3. Process character design sketches.
[0126] The original character design is changed to a fixed size (e.g., 1024×1024) to obtain the adjusted character design.
[0127] 4. Assemble the composite image and mask.
[0128] Figure 1 Placed side-by-side with the original character design sketches → composite image ( Figure 2 ).
[0129] mask1 and pure black mask (the same size as the original character design) are placed side by side to form a composite mask (mask2, black = keep, white = redraw).
[0130] 5. Provide features and conditions.
[0131] The FLUX Redux node reads the adjusted character design and extracts the character's core features (posture, facial features, etc.) → feature encoding.
[0132] The Inpaint ModelConditioning node reads the composite image + mask2, extracts background style, character form, etc., and generates conditions for where to redraw and what content to generate.
[0133] 6. AI generates expanded edges (KSampler nodes).
[0134] Using feature encoding (to ensure consistency of object details) + conditions (to ensure background coordination, similar object appearance, and redrawn areas), new content is generated in the blank area of mask2 to obtain a complete spliced image (including the original poster, redrawn image, and character design).
[0135] 7. Eliminate redundancy.
[0136] Cut out the spliced character design, leaving the original poster plus the redrawn image, to get the final image.
[0137] Figure 4 The example flowchart includes the following content.
[0138] Step 1: Input.
[0139] Original character poster: The red-clad female protagonist stands in a corner of a valley (the background needs to be expanded, retaining the character and existing partial features of the valley), which will serve as the basis for the expanded image.
[0140] Character design sketch: Full-body illustration of the female lead (red dress, twin tails, solid color background with no distractions), used to provide accurate reference for character features.
[0141] At the same time, prepare the expansion parameters (set the size to be expanded in the top, bottom, left and right directions of the poster, and clarify the expansion range).
[0142] Step 2: Expand the image and generate a mask.
[0143] Using the PadImage for Outpainting node, the original character poster is processed as follows based on the set expansion parameters.
[0144] Fill the gray placeholder area around the poster to generate the target character poster with expanded borders (the female protagonist and the original valley corner scene are retained, and the surrounding gray area is the area to be expanded and redrawn).
[0145] Simultaneously generate mask1 (the black area corresponds to the original poster content that needs to be retained, the white area corresponds to the gray placeholder and border area, and the area where the background needs to be redrawn is marked).
[0146] Step 3: Combine images and generate a mask.
[0147] The character design drawings are preprocessed to remove distracting elements except for solid color backgrounds, while retaining the core features of the female lead, such as her red dress, twin ponytails, and posture, to obtain a pure character feature image (providing a perfect character template for AI).
[0148] Composite Image: The target character poster with extended borders (including the original valley corner scene and gray extended border area) is stitched together with the pure character feature image to form a composite image (the left side is the valley scene poster to be improved, and the right side is the accurate character template).
[0149] splicing mask2: Combine mask1 (marking the original poster's retained and expanded areas) and pure black mask (the same size as the pure character feature image; pure black indicates that the character template features need to be retained) in the same splicing method to generate mask2 (black corresponds to the retained content, and white corresponds to the expanded areas to be redrawn).
[0150] Step 4: Complete the map expansion.
[0151] The Inpaint ModelConditioning node (feature extraction) reads the composite image + mask2 and extracts two key pieces of information: the style of the valley scene in the original poster (such as the valley vegetation, rock texture, and lighting tone); and the female protagonist's appearance features such as her red dress and twin ponytails. It integrates and generates redrawing conditions to guide the AI in drawing content consistent with the original valley style in the white expanded area, without altering the female protagonist's features.
[0152] Redux Node (Feature Encoding): Reads the pure character feature map, extracts subtle features such as the folds of the female lead's red dress, the curvature of her twin ponytails, and the proportions of her facial features, and generates feature encoding (sets a character detail template for the AI to ensure that the newly generated content is aligned with the template).
[0153] The redrawing conditions and feature codes are input into the KSampler node. The AI then operates within the white border area of mask2 according to the rules: extending the drawing of content consistent with the original valley style (such as expanding valley vegetation and rocks, and continuing the lighting logic); ensuring that features such as the female lead's red dress and twin ponytails are consistent with pure character features. Figure 1 To ensure no distortion occurs, the final image is a complete composite (the left side is the expanded valley scene poster, and the right side is the character template).
[0154] Step 5: Output.
[0155] Cropping and splicing the pure character feature image on the right side of the complete image, only retaining the expanded valley scene poster (the female protagonist stands in the complete valley scene, with her red dress and twin ponytails accurately depicted, and the background extending from the corner into a coherent and grand valley scene), and outputting the final expanded image result.
[0156] Based on the same technical concept, this application provides an image extension device based on an object reference map, such as... Figure 5 As shown, the device includes:
[0157] The acquisition module 501 is used to acquire the original image and the complete object reference image of the target object, wherein the original image contains part of the target object;
[0158] The edge expansion module 502 is used to generate a target image containing the edge expansion area to be filled and a first mask of the target image by performing edge expansion processing on the original image;
[0159] The processing module 503 is used to obtain an object feature map with a clean background by performing background removal processing on the complete object reference map, and to generate a second mask for the object feature map;
[0160] The stitching module 504 is used to stitch the target image and the object feature map into a composite image, and to stitch the first mask and the second mask into a composite mask;
[0161] The generation module 505 is used to extract background features and object appearance features from the composite image, and combine them with the feature encoding of the object feature map to guide the generation of an image that is consistent with the features of the target object and adapted to the background in the area to be filled and expanded within the scope defined by the composite mask.
[0162] Optionally, the edge-expanding module 502 is used for:
[0163] The original image is input into the augmentation node, and augmentation parameters are input into the augmentation node, where the augmentation parameters indicate the number of pixels that need to be augmented in the corresponding direction;
[0164] Based on the expansion parameters, gray pixels are filled around the original image to generate a target image containing the expansion area to be filled, and a first mask of the target image is generated.
[0165] Specifically, when generating the first mask of the target image, the region corresponding to the original image is set as the reserved region, and the region corresponding to the region to be filled and expanded is set as the region to be generated, so as to define the original content to be retained and the expanded content to be generated through the first mask.
[0166] Optionally, the processing module 503 is used for:
[0167] Generate a second mask that is the same size as the complete object reference image and is filled entirely with black. The black filling is used to indicate that the object feature map region corresponding to the second mask is a feature reference area that does not need to be modified.
[0168] Optionally, the generation module 505 is used for:
[0169] The composite image and composite mask are input into the feature extraction node, and the output is the conditional information that guides the generation of the redrawn image. The conditional information includes the background features in the original image, the appearance features of the object in the complete object reference image, and the redrawn area located in the area to be filled and expanded by the composite mask.
[0170] After modifying the object feature map to a standard size, input it into the feature encoding node, and output the feature encoding used to enhance the core features of the target object;
[0171] The conditional information and feature encoding guiding the generation of the redrawn image are input into the image generation node to obtain a complete image containing the original image, the redrawn image of the area to be filled and expanded, and the stitched object feature map.
[0172] The object feature map is cropped to obtain the final image composed of the original image and the redrawn image.
[0173] Optionally, the generation module 505 is specifically used for:
[0174] Generate suitable background content based on background features;
[0175] Generate content that matches the overall shape of the target object based on its appearance features;
[0176] Enhance the consistency of detailed features of the target object by using feature encoding.
[0177] Optionally, the splicing module 504 is used for:
[0178] The target image and the object feature map are stitched side by side into a composite image, and the first mask and the second mask are stitched side by side into a composite mask. The stitching method of the first mask and the second mask corresponds to that of the composite image, and the stitching order of the target image and the object feature map can be interchanged.
[0179] Optionally, background features include scene style, color tone, and environmental element features, while object appearance features include the outline, shape, and core visual features of the target object.
[0180] like Figure 6 As shown, this application provides an electronic device including a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0181] Memory 603 is used to store computer programs.
[0182] In one embodiment of this application, when the processor 601 executes the program stored in the memory 603, it implements the object reference graph-based image augmentation method provided in any of the foregoing method embodiments.
[0183] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the object reference graph-based image augmentation method provided in any of the foregoing method embodiments.
[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0185] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0186] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0187] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for augmenting an image based on an object reference map, characterized in that, The method includes: Obtain the original image and a complete object reference image of the target object, wherein the original image contains a portion of the target object; By performing edge expansion processing on the original image, a target image containing the edge expansion area to be filled and a first mask of the target image are generated; By performing background removal processing on the complete object reference image, an object feature image with a clean background is obtained, and a second mask of the object feature image is generated. The target image and the object feature map are concatenated into a composite image, and the first mask and the second mask are concatenated into a composite mask; The background features and object appearance features of the composite image are extracted, and combined with the feature encoding of the object feature image, an image that is consistent with the features of the target object and adapted to the background is generated in the area to be filled and expanded within the composite mask.
2. The method according to claim 1, characterized in that, By performing edge-expansion processing on the original image, a target image containing the area to be filled and an initial mask of the target image are generated, including: The original image is input into the augmentation node in the target model, and augmentation parameters are input into the augmentation node, wherein the augmentation parameters are used to indicate the number of pixels that need to be augmented in the corresponding direction; According to the expansion parameters, gray pixels are filled around the original image to generate a target image containing the expansion area to be filled, and a first mask of the target image is generated. Specifically, when generating the first mask of the target image, the region corresponding to the original image is set as the reserved region, and the region corresponding to the area to be filled and expanded is set as the region to be generated, so as to define the original content to be retained and the expanded content to be generated through the first mask.
3. The method according to claim 1, characterized in that, The second mask for generating the object feature map includes: A second mask, entirely filled with black, is generated using the target model and has the same size as the complete object reference map. The entire mask is filled with black to indicate that the object feature map region corresponding to the second mask is a feature reference region that does not need to be modified.
4. The method according to claim 1, characterized in that, Extracting background and object appearance features from the composite image, and combining them with the feature encoding of the object feature map, guides the generation of an image that matches the target object features and adapts to the background within the area to be filled and expanded defined by the composite mask, including: The composite image and the composite mask are input into the feature extraction node in the target model, and the output is conditional information to guide the generation of the redrawn image. The conditional information includes the background features in the original image, the object appearance features in the complete object reference image, and the redrawn area located in the area to be filled and expanded as indicated by the composite mask. After modifying the object feature map to a standard size, it is input into the feature encoding node in the target model, and the feature encoding used to enhance the core features of the target object is output. The conditional information for guiding the redrawing image generation and the feature encoding are input into the image generation node in the target model to obtain a complete image containing the original image, the redrawing image of the area to be filled and expanded, and the stitched object feature map. The object feature map is cropped to obtain the final image composed of the original image and the redrawn image.
5. The method according to claim 4, characterized in that, The resulting redrawn image includes: Generate suitable background content based on the background features; Generate content that is consistent with the overall shape of the target object based on the object's appearance features; The consistency of detailed features of the target object is enhanced based on the feature encoding.
6. The method according to claim 1, characterized in that, Concatenating the target image and the object feature map into a composite image, and concatenating the first mask and the second mask into a composite mask, includes: The target image and the object feature map are concatenated side-by-side to form a composite image using the target model, and the first mask and the second mask are concatenated side-by-side to form a composite mask. The concatenation method of the first mask and the second mask corresponds to the composite image, and the concatenation order of the target image and the object feature map can be interchanged.
7. The method according to claim 4, characterized in that, The background features include scene style, color tone and environmental element features, and the object appearance features include the outline, shape and core visual features of the target object.
8. An image augmentation device based on an object reference map, characterized in that, The device includes: The acquisition module is used to acquire the original image and the complete object reference image of the target object, wherein the original image contains part of the target object; The edge expansion module is used to generate a target image containing the edge expansion area to be filled and a first mask of the target image by performing edge expansion processing on the original image; The processing module is used to obtain an object feature map with a clean background by performing background removal processing on the complete object reference map, and to generate a second mask for the object feature map; The stitching module is used to stitch the target image and the object feature map into a composite image, and to stitch the first mask and the second mask into a composite mask; The generation module is used to extract background features and object appearance features from the composite image, combine them with the feature encoding of the object feature image, and guide the generation of an image that is consistent with the features of the target object and adapted to the background in the area to be filled and expanded within the composite mask.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.