Indoor design appearance migration method and system based on diffusion model target perception
By employing a target perception method based on a diffusion model, utilizing image segmentation and visual language models, and combining a multi-contrast loss function, the problems of detail restoration and feature confusion in interior design appearance transfer are solved, achieving efficient and personalized interior design appearance transfer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-03-19
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for transferring interior design appearances are insufficient in their ability to reproduce details, suffer from serious confusion of appearance features, lack diversity and personalization, and are difficult to meet user needs.
By employing a target perception method based on a diffusion model, this method identifies objects using image segmentation and visual language models. Combined with a multi-contrast loss function, it accurately extracts and transfers detailed features such as texture and tone to generate realistic and natural target scene images.
It achieves accurate matching and independent mapping of object appearance, resulting in more realistic and natural output, enhancing visual effects and design diversity and personalization, and meeting user needs.
Smart Images

Figure CN122492429A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine vision and image generation technology, and in particular relates to a method and system for transferring the appearance of interior design based on diffusion model target perception. Background Technology
[0002] With the rise of AIGC (Artificial Intelligence Generative Technology), deep learning models are widely used for various generative tasks, such as generating images and applying the artistic style of one image to another. These models significantly reduce the workload of interior designers, enabling them to quickly produce interior design concept sketches and improving work efficiency and communication with clients. Designers often need to combine sample images from clients with their interior scenes to generate diverse design concept sketches for clients to explore and choose from. Compared to general image appearance transfer, interior design particularly requires object-aware appearance transfer. That is, the visual appearance features of the reference object should accurately and appropriately appear on the corresponding object in the target scene, rather than being scattered throughout the scene.
[0003] However, existing indoor scene image appearance transfer methods have the following main problems: Insufficient detail restoration capability: They lack accurate preservation of the detailed features (such as texture and tone) of the reference image, resulting in unrealistic generated results; Appearance feature confusion: They cannot correctly identify the corresponding objects in the reference image and the target scene, and accurately match and transfer their appearances, which leads to appearance feature confusion or propagation errors in multi-target scenes, resulting in the mixed appearance features of the generated scene images; Insufficient diversity and personalization: The generated scene designs lack sufficient diversity and customization, making it difficult to meet the personalized needs of users. Summary of the Invention
[0004] To address the aforementioned issues, this invention proposes a method and system for interior design appearance transfer based on diffusion model target perception. By segmenting and analyzing example images and target depth images, this invention accurately identifies objects in the reference image and target scene, precisely matching their appearance features to ensure accurate feature transfer and avoid feature mismatch and confusion. Furthermore, the loss function is a weighted superposition of reconstruction loss, attention alignment loss, generalization loss, and background loss, employing a multi-contrast loss optimization mechanism to effectively extract detailed features such as texture and tone from the reference object and faithfully transfer them to the target object, resulting in more realistic and natural output and significantly improved visual effects.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the present invention provides a method for transferring the appearance of interior design based on diffusion model target perception, comprising: Obtain the example image and the target depth image; Image segmentation and analysis are performed on the example image and the target depth image to obtain the correspondence between the instance objects and their appearances in the example image and the target depth image; Based on the correspondence between instance objects and object appearances, the appearance features perceived by the objects are extracted from the example images. Based on the appearance features and target depth image, as well as the preset diffusion model, a target scene image with the expected appearance features is obtained; wherein, the loss function in the diffusion model is a weighted superposition of the reconstruction loss function, attention alignment loss function, generalization loss function and background loss function.
[0006] Furthermore, object instances, cropping regions, and segmentation masks are extracted from the example image and the target depth image; an object spatial relationship graph is constructed, which includes global constraints, distance constraints, position constraints, alignment constraints, and rotation constraints of the objects; the object spatial relationship graph determines the correspondence between objects in the target image and objects in the example image, and generates appropriate appearance descriptions for objects in the target image that do not have a corresponding object.
[0007] Furthermore, for each cropped object instance image, images of the same category are retrieved from the text-image dataset as regularized images; word tags representing the object's appearance and the cross-attention layer of the pre-trained diffusion model are jointly optimized.
[0008] Furthermore, reconstruction loss is used to reduce the difference between predicted noise and real noise in the object region of the noisy image; attention alignment loss is used to ensure that word tags focus on the object region rather than the entire cropped image; generalization loss is used to generalize the learned appearance features to target objects with different visual attributes; and background loss is used to prevent language drift problem in the pre-trained model.
[0009] Furthermore, the attention alignment loss is achieved by aligning cross-attention activation with an object segmentation mask, the generalization loss is achieved by regularizing the learned labels by regularizing the attention activation of category labels on the image, and the background loss is achieved by comparing the noise difference predicted in the background region with and without appearance labels.
[0010] Furthermore, at each time step of the diffusion process, the latent representation of the appearance and geometric conditions of each object is calculated separately; based on the object segmentation mask, the latent representations of each object are combined into a unified latent representation; the combined latent representation is used to denoise the noisy image at the current time step; the denoising process is iterated until the final image is generated.
[0011] Secondly, the present invention also provides an interior design appearance transfer system based on diffusion model target perception, comprising: The data acquisition module is configured to acquire example images and target depth images; The appearance construction module is configured to perform image segmentation and analysis on the example image and the target depth image to obtain the correspondence between the instance objects and the appearance of the objects in the example image and the target depth image; The appearance inversion module is configured to extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object's appearance. The appearance transfer module is configured to: obtain a target scene image with expected appearance features based on appearance features, target depth image, and a preset diffusion model; wherein, the loss function in the diffusion model is a weighted superposition of reconstruction loss function, attention alignment loss function, generalization loss function, and background loss function.
[0012] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the diffusion model-based target perception method for interior design appearance transfer described in the first aspect.
[0013] Fourthly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the steps of the diffusion model-based target perception method for interior design appearance transfer described in the first aspect.
[0014] Fifthly, the present invention also provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the steps of the diffusion model-based target perception method for interior design appearance transfer described in the first aspect.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention first performs image segmentation and analysis on the example image and the target depth image to obtain the correspondence between instance objects and their appearances in the example image and the target depth image. Then, based on the obtained correspondence between instance objects and their appearances, it extracts the object-perceived appearance features from the example image. Finally, based on the appearance features, the target depth image, and a preset diffusion model, it obtains a target scene image with the expected appearance features. The loss function in the diffusion model is a weighted sum of the reconstruction loss function, the attention alignment loss function, the generalization loss function, and the background loss function. By performing image segmentation and analysis on the example image and the target depth image, it can accurately identify objects in the reference image and the target scene, and precisely match their appearance features, ensuring the accuracy of feature transfer and avoiding feature mismatch and confusion. Simultaneously, the loss function, a weighted sum of the reconstruction loss function, the attention alignment loss function, the generalization loss function, and the background loss function, employs a multi-contrast loss optimization mechanism to effectively extract detailed features such as texture and tone of the reference object and faithfully transfer them to the target object, resulting in a more realistic and natural output and significantly improving the visual effect.
[0016] 2. This invention solves the problem of appearance feature confusion in multi-object scenes. Through a multi-layer attention mechanism and conditional generation technology, the appearance features of each object are independently and accurately mapped to the target object, thereby ensuring the overall consistency and visual integrity of the generated results. The generation process introduces geometric constraints such as depth maps, which supports the generation of multiple schemes according to user needs, realizing the diversification and personalization of designs and meeting users' needs for personalized designs.
[0017] 3. The generation process of this invention adopts a phased strategy, which ensures generation efficiency and allows users to intervene to adjust the generation results to better meet actual needs. The overall technical solution not only improves the generation effect, but also significantly enhances the practicality of appearance transfer, providing an efficient, flexible and practical technical tool for interior design. Attached Figure Description
[0018] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0019] Figure 1 Examples of successful and unsuccessful appearance migration of the present invention are provided, wherein columns c and d represent the failed appearance migration effects generated by other existing methods, and column e represents the appearance migration effect of the present invention. Figure 2 This is a structural diagram of each stage of the present invention; Figure 3This invention provides a process diagram and example for guiding the generation of scene correspondences. Figure 4 This is an example of the correspondence between correct and invalid objects in this invention; Figure 5 This is a schematic diagram illustrating the scene transition effect of the present invention; Figure 6 This is a process diagram illustrating the training appearance representation of the present invention; Detailed Implementation The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0021] Example 1: Existing image appearance transfer methods typically rely on global style transfer, injecting low-level image features encoded from a sample image into the target image generation process. This approach struggles to accurately preserve the appearance features of user-specified objects. Alternatively, they employ a method of first describing the appearance features in the sample image using a visual language model, then generating the target scene image using a text-to-image model. However, this approach fails to accurately identify the objects present in both the sample and target scene images, and how to correctly map their appearance features. Furthermore, due to the ambiguity of textual expression, this method struggles to accurately reproduce the appearance details of the sample image, such as texture, tone, and lighting. Moreover, existing methods often suffer from appearance feature confusion in multi-object scenes. Because objects in a scene may share similar geometric shapes, poses, and textures, the generated images often exhibit incorrect appearance feature propagation, resulting in a lack of realism and visual consistency. Additionally, these methods are inconsistent in generating complex scenes, exhibiting poor diversity in the generated results, and failing to meet users' demands for personalized and richer designs.
[0022] To address at least one of the aforementioned problems, this embodiment provides an interior design appearance transfer method based on diffusion model target perception. It extracts and analyzes the spatial relationships of furniture objects in example images and provided scene images (hereinafter referred to as target images) using image segmentation and visual language models. Design principles are injected, and the correspondence or appearance description between objects in the target interior scene and objects in the example images is inferred. Through a diffusion model optimization method using multiple contrastive loss functions, object-perceived appearance features are extracted from a single example image and represented as potential linguistic markers. Based on the appearance features of the target objects and combined with the depth map of the target scene, a final image with accurate appearance transfer is generated.
[0023] S1. Using a pre-trained image segmentation model and visual language model, process the input example image and target depth image to obtain the correspondence or appearance description of the instance objects and objects in the two images.
[0024] Optionally, One-Former is used as the pre-trained image segmentation model, and ChatGPT-4 is used as the visual language model. One-Former is a general-purpose image segmentation model capable of simultaneously handling semantic segmentation, instance segmentation, and panoptic segmentation tasks, providing accurate segmentation masks for furniture objects in indoor scenes. ChatGPT-4 is a powerful visual language model capable of understanding image content and generating detailed semantic descriptions, supporting multi-turn dialogue interaction. In the process of inferring appearance descriptions, the first step is to generate instance object sets for the two images and the corresponding object spatial relationship graph. Each object in the instance set includes information such as object ID, object category, cropped image of the object, and object segmentation mask; the construction rules for the spatial relationship graph are as follows: Global Constraints: Each object needs a global constraint indicating whether it is located at the edge or center of the room; Distance Constraints: Based on the distance between objects, determine whether they are "close" (50cm < distance < 150cm) or "far" (distance >= 150cm) from another object, and use this as a constraint; Position Constraints: Determine whether an object is located "in front of" or "to the side" (left or right) of another object, and describe this spatial positional relationship; Alignment Constraints: Determine whether an object is "center-aligned" with another object, reflecting visual or geometric symmetry; Rotation Constraints: Describe the orientation of an object, such as whether it is "facing" the center of another object; Constraint Combination Rules: Each object must have one global constraint, and multiple other constraints can be selected according to the scene to ultimately form a comprehensive description of spatial relationships; First, use One-Former to obtain objects, cropping regions, and segmentation masks from the image. During segmentation, only instances of furniture categories are obtained, specifically including common indoor furniture such as sofas, chairs, tables, cabinets, beds, bookshelves, and shelves.
[0025] To ensure the acquisition of meaningful instance objects, only instance objects with a confidence level higher than 75% are selected, and the classification results of One-Former are used as the category labels for the objects. Then, the information of this instance set, along with the complete image and the aforementioned spatial relationship graph rules, are input into the visual language model to infer the spatial relationships between instance objects in each scene in the first round.
[0026] In this embodiment, all cropped object instances, along with scene graphs of the reference and target scenes, are input into the visual language model. The visual language model infers the correspondence between each object in the target image and the object in the reference image. For target scene objects without corresponding objects, a text description consistent with their semantics and spatial characteristics is generated to complete their appearance information, thereby generating an object-perceived appearance for all objects. Finally, the user can interactively input appearance modification requests into the visual language model, which will then infer the corresponding modifications to the appearance of objects in the scene.
[0027] The example image is a user-provided image of an interior scene containing the desired appearance, serving as the source of appearance features. The reference scene is the overall scene represented by the example image, containing multiple objects and their spatial relationships. The target depth image is the geometric representation (depth map) of the target scene, defining the layout, shape, and spatial location of furniture, but excluding appearance (color, texture). The target scene is the final scene to be generated, fusing the appearance of the example image and the geometric structure of the target depth image. Instance objects are individual objects (such as a sofa or a chair) identified from an image through image segmentation. Example scene instances are collections of instance objects segmented from the example image. The scene graph is graph-structured data describing all objects in the scene and their spatial relationships (such as top, bottom, left, right, and proximity). The relationship between the example image and the target depth image: they are paired inputs; the former provides the appearance, and the latter provides geometric constraints; together they constitute the source and target framework for appearance transfer. The relationship between the reference scene and the target scene: the former is the input, and the latter is the output; the appearance of objects in the reference scene is transferred to the corresponding objects in the target scene. The relationship between instance objects and scene graph: Instance objects are the basic nodes that make up the scene graph, and the scene graph describes the spatial relationships between these nodes.
[0028] S2. Use the obtained example scene instances to crop images and masks to train corresponding language tags. Cropped object instances are obtained from example images of a reference scene as input and converted into word tokens representing the object's perceived appearance. For the text-to-image latent diffusion model used in this embodiment, the text first needs to be converted into a combination of word tokens and then into latent vectors via a text encoder. In the UNet module of the diffusion model, image generation is performed through a cross-attention layer to generate images that conform to the text conditions. Given cropped object instance images, this embodiment jointly optimizes the injection of new word tokens {V1, V2, ..., V...}. n The text encoder and the pre-trained latent diffusion model use a cross-attention layer to represent the appearance concept of all objects, where V i This represents the i-th instance object.
[0029] In the text encoder vocabulary of the pre-trained diffusion model, words are treated as tokens, each with a corresponding embedding vector representing its semantic information. Converting these tokens to represent the perceived appearance of objects involves artificially adding placeholders (e.g., {V1, V2, ..., Vn}) based on object instances and initializing a corresponding embedding vector, where each Vi uniquely corresponds to the i-th instance object cropped from the reference scene. "Jointly optimizing the text encoder injecting the new tokens {V1, V2, ..., Vn} and the cross-attention layer of the pre-trained latent diffusion model to represent the appearance concept of all objects" means inserting the previously converted token-embedding vector pairs into the text encoder vocabulary of the pre-trained diffusion model. Through training, the appearance of specific instances in the reference image is "encoded" into the embedding vectors corresponding to these tokens.
[0030] Optionally, during training, the system uses the DDIM diffusion process to add noise to the cropped object instance image, and then uses text containing word tags as text conditional input to denoise and restore the image through the diffusion model. The restored result is compared with the original image and the process result to calculate the multiple contrast loss, obtaining the currently used language tag embedding vector and the gradient of the cross attention layer parameters in the diffusion model U-Net with respect to the image at the current time step, and the relevant parameters are continuously updated iteratively accordingly. The optimized appearance feature language tags (i.e., word tag-embedding vector pairs) and the associated diffusion model attention module parameters exist in the form of model weights. The embedding vectors corresponding to the language tags capture and represent the fine-grained object-aware features of the reference object, such as texture, tone, and lighting, and are stored in the vocabulary of the text encoder. The optimized attention layer parameters are part of the parameters of the latent diffusion model network, ensuring that these appearance features can be accurately and independently mapped to the geometric structure of the target object during the generation stage, thereby achieving accurate appearance transfer.
[0031] Since each reference object has only one image crop, and it is necessary to extract object-aware rather than image-aware appearance features in order to transfer them to the corresponding target objects with different visual attributes, such as viewpoint, shape, pose, etc., this embodiment introduces a series of proposed multi-contrast losses to enhance the visual fidelity and editing generalization of word tags. Next, we will first introduce the mask appearance optimization method used, and then introduce the proposed loss function.
[0032] For each cropped instance image Ii, this embodiment uses the CLIP-retrieve tool to retrieve images of the same category from the large text-image dataset LAION-400M as a regularization dataset, and selects one image R from it. i As a regularized image. For each image, this embodiment uses a fixed text combination "a [V i [C]i The image "[]" is used for I i , and "a [C i The image "[]" is used for R i C i Replace with the object's category label, V i The word tag represents the appearance of the object. To ensure that the optimization model is not affected by the cropped image background, this embodiment uses an object segmentation mask M. Ii Limit the optimization area.
[0033] During training, this embodiment uses the DDIM diffusion process for noise addition and denoising. Specifically, for the noisy latent image z at time step t... t This embodiment uses text prompts p i Predict the noise and calculate the reconstruction loss between the actual added noise ε and the instance mask MII: ; in, , ☐ represents the noise predicted by the pre-trained latent diffusion model, and ⊙ represents element-wise multiplication. Furthermore, this embodiment introduces three contrastive loss functions to ensure that the relationship between appearance and semantics is learned during the optimization process, so as to generalize to other objects with different geometries and prevent other content in the image from affecting appearance extraction: Attention alignment loss function ( ): Encourage word tagging {V i Focusing on the object region rather than the entire cropped image is achieved through aligned cross-attention activation and object segmentation masks. ; in, Indicates I at position u i Attention activation, Γ v It is a word marker [V] i The set of M Ii This is a segmentation mask for the reference object. This loss is applied to the attention activations of all layers in the network. The position u represents the pixel coordinates on the attention map; (That is, attention activation) represents the weight of the word tag y on the image features at a specific spatial location u. The numerator in the formula is the object mask region. Summation within ( The denominator is the entire cropped image area. Summation within ( Here, u is the summation index; the physical meaning of this loss function is to calculate "the proportion of attention weights falling within the object mask to the total weights". The goal is to concentrate attention as much as possible within the mask, so there is no conflict, but rather an inclusion relationship.
[0034] Generalization loss function ( This ensures that the learned appearance features can be generalized to corresponding target objects with different visual attributes. This embodiment utilizes a regularized image R... i Category label [C] i Attention activation regularization learning of word tags [V] i ]: ; in, It is a regularized image R i Cross-attention activation of the word tag y; Γ v It is a word tag for learning [V] i The set of Γ c Includes category labels [C i ]. Γ v ×Γ c Represents the set of learned word tags With category label set The Cartesian product; it represents the combination of all “learned label-class label” pairs used for contrastive learning on regularized images. Refers to a set A specific learned word tag (such as [Vi] representing appearance); Refers to a set The loss is a specific category label (such as "chair"); the purpose of this loss is to ensure that the regions activated by the learned appearance label on the general image are consistent with the regions activated by the general label of that category, thereby achieving feature generality.
[0035] Background loss function ( To prevent language drift in the pre-trained model during optimization, this embodiment adds a background phrase [E] to each text prompt, forming two prompts: one with word tags [V]. i Appearance perception prompts and unlabeled category-aware cues The background loss function encourages the two prompts to generate similar backgrounds: ; The reconstruction losses are: ; in, It is the background mask of the reference object (i.e., the complement of the object mask); , The 'z' represents the noise predicted by the pre-trained latent diffusion model, '⊙' represents element-wise multiplication, and z represents the noise predicted by the latent diffusion model. t For the noisy latent image at time step t, p i For text prompts, ε represents noise; I i To crop the image; [] represents the expectation operator, which indicates that the expectation of the terms within [] is calculated; specifically, during training, it refers to the expectation of all possible latent image variables z and cropped instance images I. i It follows a standard normal distribution. noise And sampling is performed at time step t during the diffusion process and the average loss is calculated; Z represents the image data in the latent space, referring to the distribution of the input images during training; The square of the L2 norm is used to calculate the pixel-level difference between predicted noise and actual noise, measuring the consistency of the predicted noise in the background region between two different text prompts (marked and unmarked).
[0036] The four loss functions are finally weighted and summed to form the final loss function of the optimization model: ; The weights used in this disclosure are: λ1=0.1, λ2=0.3, λ3=0.5. Experiments have verified that this set of weights achieves a good balance between maintaining appearance details and achieving generalization ability.
[0037] Segmentation mask for each object It is the core of object-aware appearance extraction, and it is trained by interacting with the language tag V in the following ways. i The optimization correlation is as follows: In the reconstruction loss, the mask acts as a spatial weight, ensuring that gradient updates and feature learning are only applied to the target object region; in the attention alignment loss, the goal is to make the word tag V... i Corresponding cross-attention map and object mask The formula calculates the proportion of attention falling on the mask area and maximizes this proportion through a loss function, which directly establishes the word tag V. i This demonstrates how to train language tags using masks, based on strong associations with specific object regions in an image; that is, using masks as supervision signals to guide V. i The attention is focused on the target object.
[0038] A background mask is used in the formula for background loss. (The complement of the mask), it is more like a V i and without V i The prompt indicates whether the predicted noise in the background region is consistent, which ensures Vi It only carries the appearance information of the object, without contaminating or carrying background information, preventing the model from encoding background features into the V image in order to fit the entire image. i .
[0039] Therefore, mask It's not simply about inputting data, but about guiding learnable markers V i It accurately captures and represents the key spatial constraints of the object's physical appearance, rather than the content of the entire image; without a mask, accurate "object-aware" appearance transfer cannot be achieved.
[0040] S3. Generate a target scene image with an object-aware appearance using conditional combination methods.
[0041] After extracting the object appearances, the third stage uses the depth map of the target scene and the appearances of all objects in the target scene as conditions to generate the final scene image. The object appearances can be latent word tags learned in step S2, or text descriptions generated by a visual language model. This embodiment uses ControlNet as the base model and uses the depth map and object appearances as conditions for image generation.
[0042] However, using a single text prompt that includes the appearance of all objects as a condition for generation can lead to missing objects or feature leakage, especially for indoor scene objects with similar geometric details. To further bind the combined object-aware appearance conditions to specific objects in the depth map, this embodiment adopts the FineControlNet approach, spatially aligning the appearance conditions with the geometric conditions of each target object.
[0043] Specifically, during the generation process at each time step t, a latent representation is computed for the appearance and geometry of each object. }, then based on the instance mask of the target object { Combine them. The combination of time steps t is expressed as: ; in, This is the latent representation following the cross-attention block in U-Net. This makes each It can be controlled by its own appearance and geometry. Then, the combined... Used to denoise the target scene image at time step t.
[0044] in,{ `z` is an intermediate feature tensor computed separately for the i-th target object at time step t in the diffusion model's denoising process; it resides after a specific layer of the U-Net diffusion model. U-Net is the core of the diffusion model, and its input is the noisy image `z`.t The output is the predicted noise. U-Net consists of a series of downsampling and upsampling blocks, embedded with cross-attention layers. The cross-attention layer is a key mechanism for textual conditions influencing image generation; it calculates the correlation between image features (Key / Value) and text embeddings (Query), thus injecting textual semantics (e.g., "a wood-grain sofa") into the image feature generation process. In the computation... In this process, we first input the specific conditions of the i-th object into the condition control module or a similar condition injection module, which outputs a set of feature modulation signals. These signals are fused with the backbone features of the U-Net in a specific layer (usually inside or after a downsampling or upsampling block). Specifically, this refers to the intermediate feature that has been processed by the cross-attention block at this layer after the fusion has occurred. At this point, the feature simultaneously contains: the image structure information at the current time step, the appearance semantics of the i-th object, and the geometric structure of the i-th object; therefore, It is an image feature that has been "modulated" by the exclusive appearance and geometric conditions of "object i".
[0045] Through the above three steps, the appearance of an object in a reference image can be accurately transferred to the corresponding object in the target scene, while maintaining the harmony and rationality of the generated result. This achieves the appearance transfer of object perception and provides an efficient visual communication tool for interior design.
[0046] Example 2: Based on Example 1, this example provides a method for transferring the appearance of interior design based on diffusion model target perception, including: S1. Obtain the instance indoor scene image representing the style and the depth image of the indoor scene to be generated.
[0047] S2. Input the two images into the image segmentation model to obtain furniture instance objects in each image and assign them numbers.
[0048] S3. Input each segmented instance object (example image), the complete scene image (target image), and the task description into the visual language model to obtain a relative spatial relationship diagram of objects in the example image and the target image.
[0049] S4. Input the segmented object images of the example image and the target image, the relative spatial relationship diagram, and the task description into the visual language model to obtain the corresponding object in the example image for each object in the target image. If there is no suitable corresponding object, obtain the appropriate appearance description text.
[0050] S5. Input the user's required modifications and deduce the new appearance description text (new task description) or corresponding relationship. Repeat step S4 until the user finishes making modifications.
[0051] S6. Obtain the segmented image of the example object from which the appearance features language identifiers are to be extracted, determine the number of steps to optimize the language identifiers and initialize the language identifiers used to represent the appearance, and use a graph-to-graph similarity retrieval tool to obtain recommended images of the same type as the example object as the dataset in a large image dataset.
[0052] S7. The complete text combining the language identifier of the appearance features to be extracted with the auxiliary text, the corresponding example image object cropped region image, and a randomly selected single recommended image of the same type are input into the diffusion model. The current cropped image and the recommended image are converted into latent vectors using the compression model and noise adder of the diffusion model, and noise is added at random time steps. The text is converted into latent vectors through the text encoder, and the true latent vector is predicted using this vector and the noisy latent vector.
[0053] S8. Combine the predicted latent vector with the cropped image and the true latent vector obtained by the compressor from the recommended image, and use the multiple loss function to calculate the gradient of the currently used language tag and the parameters of the diffusion model attention module with respect to the image at the current time step, and update the relevant parameters.
[0054] S9. Repeat the first two steps to iterate until the set number of optimization steps is reached, and save the appearance feature language identifier and the attention module of the diffusion model. S10. Input the target image depth map, the target object mask, the target object appearance text or appearance language identifier output by the visual language model into the diffusion model that can inject image and text conditions, and replace the corresponding module parameters of the diffusion model with the aforementioned appearance feature language identifier and attention module, and determine the number of time steps required for image generation. S11. Using object masks and corresponding text descriptions, at each denoising time step of the diffusion model, the original latent vectors are copied n times (the number of objects), and the text conditions corresponding to the object are injected into each vector. Finally, each vector is combined through a mask to obtain the denoising result of that time step. S12. Iterate continuously according to the time step until the time step reaches 0, and output the final generated image.
[0055] Example 3: This embodiment provides an interior design appearance transfer system based on diffusion model target perception, including: The data acquisition module is configured to acquire example images and target depth images; The appearance construction module is configured to perform image segmentation and analysis on the example image and the target depth image to obtain the correspondence between the instance objects and the appearance of the objects in the example image and the target depth image; The appearance inversion module is configured to extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object's appearance. The appearance transfer module is configured to: obtain a target scene image with expected appearance features based on appearance features, target depth image, and a preset diffusion model; wherein, the loss function in the diffusion model is a weighted superposition of reconstruction loss function, attention alignment loss function, generalization loss function, and background loss function.
[0056] The working method of the system is the same as that of the diffusion model-based target perception method for interior design appearance transfer in Embodiment 1 or Embodiment 2, and will not be repeated here.
[0057] Example 4: This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the diffusion model-based target perception method for interior design appearance transfer described in this embodiment.
[0058] Example 5: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the steps of the indoor design appearance migration method based on diffusion model target perception described in Embodiment 1.
[0059] Example 6: This embodiment provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the indoor design appearance migration method based on diffusion model target perception described in Embodiment 1.
[0060] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.
Claims
1. A method for transferring the appearance of interior design based on diffusion model target perception, characterized in that, include: Obtain the example image and the target depth image; Image segmentation and analysis are performed on the example image and the target depth image to obtain the correspondence between the instance objects and their appearances in the example image and the target depth image; Based on the correspondence between instance objects and object appearances, the appearance features perceived by the objects are extracted from the example images. Based on the appearance features and target depth image, as well as the preset diffusion model, a target scene image with the expected appearance features is obtained; wherein, the loss function in the diffusion model is a weighted superposition of the reconstruction loss function, attention alignment loss function, generalization loss function and background loss function.
2. The method for transferring interior design appearance based on diffusion model target perception as described in claim 1, characterized in that, Extract object instances, cropping regions, and segmentation masks from the example image and the target depth image; construct an object spatial relationship graph, which includes global constraints, distance constraints, position constraints, alignment constraints, and rotation constraints of the objects; determine the correspondence between objects in the target image and objects in the example image from the object spatial relationship graph, and generate appropriate appearance descriptions for objects in the target image that do not have a corresponding object.
3. The method for transferring interior design appearance based on diffusion model target perception as described in claim 2, characterized in that, For each cropped object instance image, retrieve images of the same category from the text-image dataset as regularized images; The word tag representing the appearance of the object is jointly optimized by the cross-attention layer of the pre-trained diffusion model.
4. The method for transferring interior design appearance based on diffusion model target perception as described in claim 2, characterized in that, Reconstruction loss is used to reduce the difference between predicted noise and real noise in the object region of the noisy image; Attention alignment loss is used to ensure that word tags focus on the target area rather than the entire cropped image; generalization loss is used to generalize the learned appearance features to target objects with different visual attributes; background loss is used to prevent language drift in the pre-trained model.
5. The method for transferring interior design appearance based on diffusion model target perception as described in claim 4, characterized in that, The attention alignment loss is achieved by aligning cross-attention activation with an object segmentation mask; the generalization loss is achieved by regularizing the learned labels by regularizing the attention activation of class labels on the image; and the background loss is achieved by comparing the noise difference predicted in the background region with and without appearance labels.
6. The method for transferring interior design appearance based on diffusion model target perception as described in claim 4, characterized in that, At each time step of the diffusion process, the latent representation of the appearance and geometry of each object is calculated separately; based on the object segmentation mask, the latent representations of each object are combined into a unified latent representation; the combined latent representation is used to denoise the noisy image at the current time step; the denoising process is iterated until the final image is generated.
7. An interior design appearance transfer system based on diffusion model target perception, characterized in that, include: The data acquisition module is configured to acquire example images and target depth images; The appearance construction module is configured to perform image segmentation and analysis on the example image and the target depth image to obtain the correspondence between the instance objects and the appearance of the objects in the example image and the target depth image; The appearance inversion module is configured to extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object's appearance. The appearance transfer module is configured to: obtain a target scene image with expected appearance features based on appearance features, target depth image, and a preset diffusion model; wherein, the loss function in the diffusion model is a weighted superposition of reconstruction loss function, attention alignment loss function, generalization loss function, and background loss function.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the interior design appearance transfer method based on diffusion model target perception as described in any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the program, it implements the steps of the interior design appearance transfer method based on diffusion model target perception as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the interior design appearance migration method based on diffusion model target perception as described in any one of claims 1-6.