Interior design appearance migration method and system based on diffusion model target perception

By applying a target perception method based on diffusion model in interior design appearance migration, the problem of insufficient detail restoration capabilities and confusion of appearance characteristics in the prior art is solved, and a more realistic and natural appearance migration effect and personalized design satisfaction are achieved.

CN120107058AInactive Publication Date: 2025-06-06SHANDONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510327767.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing interior design appearance migration methods have problems such as insufficient detail restoration capabilities, confusion of appearance features, diversity and personalization, and it is difficult to accurately identify and match the appearance characteristics of objects in the reference image and the target scene.

Method used

Using a target perception method based on the diffusion model, the object is accurately identified and appearance features are extracted by image segmentation and analysis of the sample images and the target depth images. Using weighted overlays of reconstruction loss, attention alignment loss, generalization loss, and background loss functions, the diffusion model is optimized to extract and migrate detailed features of the reference object's texture, tone, and other aspects.

Benefits of technology

It significantly improves the visual effect, ensures the accuracy of feature migration, avoids feature mismatch and confusion, and the generated scene images are more realistic and natural, meeting users' personalized design needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107058A_ABST
    Figure CN120107058A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of machine vision and image generation, and provides an interior design appearance migration method and system based on diffusion model target perception, and the method comprises the steps: carrying out the image segmentation and analysis of an example image and a target depth image, and obtaining the corresponding relation between an example object and an object appearance in the example image and the target depth image; on the basis of obtaining the corresponding relationship between the instance object and the object appearance, extracting the appearance feature perceived by the object from the instance image; and obtaining a target scene image with expected appearance characteristics according to the appearance characteristics, the target depth image and a preset diffusion model, so that objects in the reference image and the target scene can be accurately identified, the appearance characteristics of the objects can be accurately matched, the problems of mismatching and confusion of the characteristics can be avoided, and meanwhile, a multi-contrast loss optimization mechanism is adopted. The detail features such as texture and hue of the reference object are effectively extracted and are faithfully migrated to the target object, the generated result is more real and natural, and the visual effect is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision and image generation, and in particular relates to a method and system for interior design appearance migration based on diffusion model target perception. Background Art

[0002] With the rise of AIGC (artificial intelligence) technology, deep learning models are widely used in various generation tasks, such as generating images, applying the artistic style of one image to another image, etc. Such models greatly reduce the workload of interior designers, facilitate interior designers to quickly produce interior design concept drawings, and improve work efficiency and communication efficiency with customers. Designers often need to combine customers' example images and customers' interior scenes to generate a variety of design concept drawings in order to show them to customers for further exploration and selection. Compared with general image appearance transfer, interior design particularly requires object-aware appearance transfer. That is, the visual appearance features of the reference object should appear accurately and appropriately on the corresponding object in the target scene, rather than being scattered throughout the scene.

[0003] However, existing indoor scene image appearance transfer methods have the following major problems: insufficient detail restoration capability: lack of accurate preservation of reference image detail features (such as texture and tone), resulting in unrealistic results; appearance feature confusion: inability to correctly identify corresponding objects in reference images and target scenes, and to accurately match and transfer their appearance, which results in appearance features being easily confused or propagating errors in multi-target scenes, causing the appearance features of generated scene images to be mixed with each other; insufficient diversity and personalization: the generated scene designs lack sufficient diversity and customization, making it difficult to meet users' personalized needs. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes a method and system for interior design appearance migration based on diffusion model target perception. The present invention can accurately identify objects in reference images and target scenes by performing image segmentation and analysis on sample images and target depth images, and accurately match their appearance features to ensure the accuracy of feature migration and avoid feature mismatch and confusion problems. At the same time, the loss function is a weighted superposition of reconstruction loss function, attention alignment loss function, generalization loss function and background loss function, and adopts a multi-contrast loss optimization mechanism to effectively extract detail features such as texture and tone of the reference object, and faithfully migrate them to the target object. The generated result is more realistic and natural, and the visual effect is significantly improved.

[0005] In order to achieve the above object, the present invention is implemented through the following technical solutions:

[0006] In a first aspect, the present invention provides an interior design appearance migration method based on diffusion model target perception, comprising:

[0007] Get sample image and target depth image;

[0008] Performing image segmentation and analysis on the sample image and the target depth image to obtain a correspondence between instance objects and object appearances in the sample image and the target depth image;

[0009] Based on the correspondence between the instance object and the object appearance, the object-perceived appearance features are extracted from the example image;

[0010] According to the appearance features and the target depth image, as well as a preset diffusion model, a target scene image with expected appearance features is obtained; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function.

[0011] Furthermore, object instances, cropped regions, and segmentation masks are extracted from the example image and the target depth image; an object space relationship graph is constructed, wherein the object space relationship graph includes global constraints, distance constraints, position constraints, alignment constraints, and rotation constraints of the object; the object space relationship graph determines the correspondence between the object in the target image and the object in the example image, and generates a suitable appearance description for the object without a corresponding object in the target image.

[0012] Furthermore, for each cropped object instance image, images of the same category are retrieved from the text-image dataset as regularization images; the word tags representing the object appearance and the cross-attention layer of the pre-trained diffusion model are jointly optimized.

[0013] Furthermore, the reconstruction loss is used to reduce the difference between the predicted noise and the real noise of the noisy image in the object area; the attention alignment loss is used to ensure that the word tag focuses on the object area rather than the entire cropped image; the generalization loss is used to generalize the learned appearance features to target objects with different visual attributes; the background loss is used to prevent the language drift problem of the pre-trained model.

[0014] Furthermore, the attention alignment loss is implemented by aligning the cross-attention activations with the object segmentation mask, the generalization loss regularizes the learned markers by regularizing the attention activations of the category labels on the image, and the background loss is implemented by comparing the noise difference predicted in the background area by cues with and without appearance markers.

[0015] Furthermore, in each time step of the diffusion process, the latent representations of the appearance and geometric conditions of each object are calculated respectively; based on the object segmentation mask, the latent representations of each object are combined into a unified latent representation; the noisy image of the current time step is denoised using the combined latent representation; and the denoising process is iterated until the final image is generated.

[0016] In a second aspect, the present invention further provides an interior design appearance migration system based on diffusion model target perception, comprising:

[0017] The data acquisition module is configured to: acquire a sample image and a target depth image;

[0018] The appearance construction module is configured to: perform image segmentation and analysis on the example image and the target depth image to obtain a correspondence between instance objects and object appearances in the example image and the target depth image;

[0019] The appearance inversion module is configured to: extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object appearance;

[0020] The appearance transfer module is configured to obtain a target scene image with expected appearance features based on appearance features and a target depth image, as well as a preset diffusion model; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function.

[0021] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for interior design appearance migration based on diffusion model target perception described in the first aspect.

[0022] In a fourth aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein when the processor executes the program, the steps of the interior design appearance migration method based on diffusion model target perception described in the first aspect are implemented.

[0023] In a fifth aspect, the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the interior design appearance migration method based on diffusion model target perception described in the first aspect are implemented.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] 1. The present invention first performs image segmentation and analysis on the example image and the target depth image to obtain the correspondence between the instance object and the object appearance in the example image and the target depth image; then, based on the correspondence between the instance object and the object appearance, the object-perceived appearance features are extracted from the example image; finally, according to the appearance features and the target depth image, and a preset diffusion model, a target scene image with expected appearance features is obtained; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function; by performing image segmentation and analysis on the example image and the target depth image, the objects in the reference image and the target scene can be accurately identified, and their appearance features can be accurately matched, so as to ensure the accuracy of feature migration and avoid feature mismatch and confusion problems; at the same time, the loss function is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function, and a multi-contrast loss optimization mechanism is adopted to effectively extract the texture, hue and other detail features of the reference object, and faithfully migrate them to the target object, so that the generated result is more realistic and natural, and the visual effect is significantly improved.

[0026] 2. The present invention solves the problem of appearance feature confusion in multi-target scenes. Through a multi-layer attention mechanism and conditional generation technology, the appearance features of each object are independently and accurately mapped to the target object, thereby ensuring the overall consistency and visual integrity of the generated results. The generation process introduces geometric constraints such as depth maps to support the generation of multiple solutions based on user needs, thereby achieving design diversification and personalization and meeting users' needs for personalized design.

[0027] 3. The generation process of the present invention adopts a phased strategy to ensure generation efficiency and allow users to intervene to adjust the generation results to make them more in line with actual needs; the overall technical solution not only improves the generation effect, but also significantly enhances the practicality of appearance migration, providing an efficient, flexible and practical technical tool for interior design. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings in the specification that constitute a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments of this embodiment and their descriptions are used to explain this embodiment and do not constitute improper limitations on this embodiment.

[0029] Figure 1 are examples of correct and failed appearance migration of the present invention, wherein columns c and d are failed appearance migration effects generated by other existing methods, and column e is the appearance migration effect of the present invention;

[0030] Figure 2 It is the structural diagram of each stage of the present invention;

[0031] Figure 3 Process diagram and examples of generating scene correspondence for the guidance of the present invention;

[0032] Figure 4 This is an example of the correct and invalid object correspondence relationship of the present invention;

[0033] Figure 5 A schematic diagram of the scene migration effect of the present invention;

[0034] Figure 6 A process diagram for the training appearance representation of the present invention; DETAILED DESCRIPTION

[0035] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0036] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0037] Embodiment 1:

[0038] Existing image appearance transfer methods are usually based on global style transfer. By encoding the low-level image features of the sample image and injecting them into the generation process of the target image, it is difficult to accurately retain the appearance features of the user-specified object; or adopt a method of first using a visual language model to describe the appearance features in the sample image, and then generating the target scene image through a text-generated image model. However, this method cannot accurately identify which objects are in the sample image and the target scene image, and how to correctly correspond to the appearance features of these objects. At the same time, due to the ambiguity of text expression, this method is difficult to accurately restore the appearance details of the sample image, such as texture, tone, and light and shadow. In addition, existing methods often have the problem of appearance feature confusion in multi-target scenes. Since the geometric shapes, postures, and texture features of objects in the scene may be similar, the generated images often have appearance feature propagation errors, resulting in a lack of authenticity and visual consistency in the generated results. At the same time, these methods have unstable generation effects on complex scenes, and the diversity of the generated results is also poor, which makes it difficult to meet users' needs for personalized and rich designs.

[0039] In order to solve at least one of the above problems, the present embodiment provides an interior design appearance migration method based on diffusion model target perception, which extracts and analyzes the spatial relationship between the example image and the provided scene image (hereinafter referred to as the target image) through an image segmentation model and a visual language model, injects design principles and infers the correspondence or appearance description between the objects in the target indoor scene and the objects in the example image; extracts object-perceived appearance features from a single example image through a diffusion model optimization method of a multi-contrast loss function and represents them as potential language markers; generates a final image with precise appearance migration based on the appearance features of the target object and combined with the depth map of the target scene.

[0040] S1. Use the pre-trained image segmentation model and visual language model to process the input sample image and target depth image to obtain the corresponding relationship or appearance description of the instance objects and object appearance in the two images.

[0041] Optionally, use One-Former as the pre-trained image segmentation model and ChatGPT-4 as the visual language model. One-Former is a general image segmentation model that can handle semantic segmentation, instance segmentation, and panoptic segmentation tasks simultaneously, and provide accurate segmentation masks for furniture objects in indoor scenes; ChatGPT-4 is a powerful visual language model that can understand image content and generate detailed semantic descriptions, supporting multi-round dialogue interactions. In the process of reasoning about appearance descriptions, the first step is to generate instance object sets of the two images and the corresponding object spatial relationship graphs. The information included in each object in the instance set includes the object number, object category, cropped image of the object, and segmentation mask of the object; the construction rules of the spatial relationship graph are as follows:

[0042] Global constraint: Each object needs a global constraint to indicate whether it is at the edge or in the middle of the room; Distance constraint: Based on the distance between objects, determine whether it is "close" (50cm<distance<150cm) or "far away" (distance>=150cm) from another object, and use this as a constraint; Position constraint: Determine whether an object is in the "front" or "side" (left or right) of another object, and describe this spatial position relationship; Alignment constraint: Determine whether an object is "center-aligned" with another object, reflecting visual or geometric symmetry; Rotation constraint: Describes the orientation of an object, such as whether it is "facing" the center of another object; Constraint combination rules: Each object must have a global constraint, and multiple other constraints can be selected according to the scene, ultimately forming a comprehensive description of spatial relationships;

[0043] First, One-Former is used to obtain objects, cropped areas, and segmentation masks in the image. During segmentation, only instances of the furniture category are obtained, including common indoor furniture such as sofas, chairs, tables, cabinets, beds, bookshelves, and racks.

[0044] To ensure that meaningful instance objects are obtained, only instance objects with a confidence level higher than 75% are selected, and the classification results of One-Former are used as the category labels of the objects. Then the information of the instance set, the complete image, and the aforementioned spatial relationship graph rules are input into the visual language model to infer the spatial relationship between instance objects in each scene in the first round.

[0045] In this embodiment, all cropped object instances, as well as the scene graphs of the reference scene and the target scene are input into the visual language model, which infers the correspondence between each object in the target image and the object in the reference image. For the target scene objects without corresponding objects, a text description consistent with their semantics and spatial characteristics is generated to complete their appearance information, thereby generating object-perceived appearance for all objects. Finally, the user can interactively input appearance modification requirements to the visual language model, and the visual language model will infer the corresponding modifications to the appearance of the objects in the scene.

[0046] S2. Use the obtained sample scene instance to crop the image and mask to train the corresponding language tag

[0047] The cropped object instances are taken as input from the sample images of the reference scene and converted into word tags representing the perceived appearance of the object. For the text-to-image latent diffusion model used in this embodiment, the text needs to be first converted into a combination of word tags and converted into a latent vector through a text encoder, and then injected into the image generation through a cross-attention layer in the UNet module of the diffusion model to generate an image that meets the text conditions. Given the cropped object instance image, this embodiment jointly optimizes the injection of new word tags {V 1 , V 2 , ..., V n} text encoder and the cross-attention layer of the pre-trained latent diffusion model to represent the appearance concepts of all objects, where V i Represents the i-th instance object.

[0048] Since there is only one image crop per reference object, and object-aware rather than image-aware appearance features need to be extracted in order to transfer them to corresponding target objects with different visual attributes, such as viewpoint, shape, pose, etc., this embodiment introduces a series of multiple contrastive losses to enhance the visual fidelity and editing generalization of word tags. Next, we first introduce the adopted mask appearance optimization and then introduce the proposed loss function.

[0049] For each cropped instance image Ii, this embodiment uses the retrieval tool CLIP-retrieve to retrieve images of the same category from the large text-image dataset LAION-400M as a regularized dataset, and selects an image R i As a regularized image. For each image, this embodiment uses a fixed text collocation "a [V i ][C i ] images for I i , and "a [C i ] for R i , where C i Use the object’s category label to replace V i is a word mark representing the appearance of the object. In order to ensure that the optimization model is not affected by the background of the cropped image, this embodiment uses the object segmentation mask M Ii Limit the optimization area.

[0050] During the training process, this embodiment uses the DDIM diffusion process to perform denoising and denoising operations. Specifically, for the noise potential image z at time step t t , this embodiment uses text prompt p i Predict the noise and calculate the reconstruction loss between the instance mask MIi and the actual added noise ε:

[0051]

[0052] Among them, ∈ θ , t (p)=∈ θ (z t ,t,p),∈ θ , t represents the noise predicted by the pre-trained latent diffusion model and ⊙ represents element-wise multiplication.

[0053] Furthermore, this embodiment introduces three contrast loss functions to ensure that the association between appearance and semantics is learned during the optimization process so as to generalize to other objects with different geometries and prevent other contents in the image from affecting the appearance extraction:

[0054] Attention Alignment Loss Function Encourage word mark {V i Focus on the object region instead of the entire crop by aligning the cross-attention activations with the object segmentation mask:

[0055]

[0056] in, Indicates that at position u I i The attention activation, Γv is a word mark [V i ], M Ii is the segmentation mask of the reference object. The loss is applied to the attention activations of all layers of the network.

[0057] Generalization loss function Ensure that the learned appearance features can be generalized to the corresponding target objects with different visual attributes. i The upper category label [C i ]’s attention activation regularization for learning word tokens [V i ]:

[0058]

[0059] in, is the regularized image R i Cross-attention activation of the previous word token y; Γ v is the learned word token [V i ], Γ c Contains category labels [C i ].

[0060] Background loss function Prevent the language drift problem of the pre-trained model during the optimization process. This embodiment adds a background phrase [E] to each text prompt, forming two prompts: one with a word tag [V i ]’s appearance perception tips and unlabeled category-aware cues The background loss function encourages these two cues to generate similar backgrounds:

[0061]

[0062] in, is the background mask of the reference object (i.e., the complement of the object mask).

[0063] Finally, the four loss functions are weighted and superimposed as the final loss function of the optimization model:

[0064]

[0065] Among them, the weight used in this disclosure is: 1 =0.1,λ 2 =0.3,λ 3 =0.5. Experiments have verified that this set of weights achieves a good balance between maintaining appearance details and achieving generalization ability.

[0066] S3. Generate a target scene image with object-perceived appearance by using a conditional combination method.

[0067] After the object appearance is extracted, the third stage uses the depth map of the target scene and the appearance of all objects in the target scene as conditions to generate the final scene image. The object appearance can be the potential word tag learned in step S2 or the text description generated by the visual language model. This embodiment uses ControlNet as the basic model and uses the depth map and object appearance as conditions for image generation.

[0068] However, using a single text hint to include all object appearances as conditions for generation can lead to object missing or feature leakage problems, especially for indoor scene objects with similar geometric details. In order to further bind the combined object-aware appearance condition to a specific object in the depth map, this embodiment adopts the idea of ​​FineControlNet to spatially align the appearance condition with the geometric condition of each target object.

[0069] Specifically, during the generation process at each time step t, a latent representation is computed for each object-level appearance and geometry condition Then based on the instance mask of the target object Combine them. The combination of time steps t is expressed as:

[0070]

[0071] in, is the potential representation after the cross-attention block in U-Net. This makes each can be controlled by its own appearance and geometry. Then, the combined h t Used to denoise the target scene image at time step t.

[0072] Through the processing of the above three steps, the appearance of the object in the reference image can be successfully and accurately transferred to the corresponding object in the target scene, while maintaining the harmony and rationality of the generated results, realizing the appearance transfer of object perception, and providing an efficient visual communication tool for interior design.

[0073] Embodiment 2:

[0074] Based on Example 1, this embodiment provides an interior design appearance migration method based on diffusion model target perception, including:

[0075] S1. Acquire an example indoor scene image representing a style and an indoor scene depth image to be generated.

[0076] S2. Input the two images into the image segmentation model to obtain the furniture instance objects in each image and number them.

[0077] S3. Input each segmented instance object (example image), the complete scene image (target image), and the task description into the visual language model to obtain a relative spatial relationship diagram of the objects in the example image and the target image.

[0078] S4. Input the segmented object images of the example image and the target image, the relative spatial relationship diagram, and the task description into the visual language model to obtain the corresponding object of each object in the target image in the example image. If there is no suitable corresponding object, obtain the appropriate appearance description text.

[0079] S5. Input the modification requirements required by the user, and infer to obtain a new appearance description text (new task description) or a corresponding relationship, and repeat step S4 until the user finishes the modification.

[0080] S6. Obtain a segmented image of an example object whose appearance feature language identifier is to be extracted, determine the number of steps for optimizing the language identifier and initialize the language identifier used to represent the appearance, and use an image similarity retrieval tool in a large image dataset to obtain recommended images similar to the example object as a dataset.

[0081] S7. The complete text combining the language identifier of the appearance features to be extracted and the auxiliary text, the corresponding sample image object cropped area image, and a randomly selected single similar recommended image are input into the diffusion model, and the compression model and denoiser of the diffusion model are used to convert the current cropped image and the recommended image into latent vectors, and denoise them at random time steps. The text is converted into a latent vector through a text encoder and the vector and the noisy latent vector are used to predict the true latent vector.

[0082] S8. Compare the predicted latent vector with the cropped image and the real latent vector obtained by the compressor of the recommended image, and use multiple loss functions to calculate the gradient of the currently used language tag and the parameter value of the attention module of the diffusion model with respect to the image at the current time step, and update the relevant parameters.

[0083] S9, repeat the first two steps and iterate continuously until the set number of optimization steps is reached, and save the appearance feature language identification of the object and the attention module of the diffusion model;

[0084] S10, inputting the target image depth map, the target object mask output by the visual language model, the target object appearance text or the appearance language identifier into the diffusion model that can be injected with image and text conditions, and replacing the corresponding module parameters of the diffusion model with the aforementioned appearance feature language identifier and the attention module, and determining the required time steps for image generation;

[0085] S11, using the object mask and the corresponding text description of the object, at each denoising time step of the diffusion model, the original latent vector is copied n times (the number of objects) and the text condition corresponding to the object is injected into each vector, and finally each vector is combined by means of a mask to obtain the denoising result of the time step;

[0086] S12. Continue to iterate according to the time step until the time step reaches 0, and output the final generated image.

[0087] Embodiment 3:

[0088] This embodiment provides an interior design appearance migration system based on diffusion model target perception, including:

[0089] The data acquisition module is configured to: acquire a sample image and a target depth image;

[0090] The appearance construction module is configured to: perform image segmentation and analysis on the example image and the target depth image to obtain a correspondence between instance objects and object appearances in the example image and the target depth image;

[0091] The appearance inversion module is configured to: extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object appearance;

[0092] The appearance transfer module is configured to obtain a target scene image with expected appearance features based on appearance features and a target depth image, as well as a preset diffusion model; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function.

[0093] The working method of the system is the same as the interior design appearance migration method based on diffusion model target perception in Example 1 or Example 2, and will not be repeated here.

[0094] Embodiment 4:

[0095] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the interior design appearance migration method based on diffusion model target perception described in the embodiment are implemented.

[0096] Embodiment 5:

[0097] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, the steps of the interior design appearance migration method based on diffusion model target perception described in Example 1 are implemented.

[0098] Embodiment 6:

[0099] This embodiment provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the interior design appearance migration method based on diffusion model target perception described in Embodiment 1 are implemented.

[0100] The above description is only a preferred embodiment of the present embodiment and is not intended to limit the present embodiment. For those skilled in the art, the present embodiment may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present embodiment shall be included in the protection scope of the present embodiment.

Claims

1. A method for interior design appearance transfer based on diffusion model target perception, characterized in that: include: Get sample image and target depth image; Performing image segmentation and analysis on the sample image and the target depth image to obtain a correspondence between instance objects and object appearances in the sample image and the target depth image; Based on the correspondence between the instance object and the object appearance, the object-perceived appearance features are extracted from the example image; According to the appearance features and the target depth image, as well as a preset diffusion model, a target scene image with expected appearance features is obtained; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function.

2. The method for interior design appearance migration based on diffusion model target perception according to claim 1, characterized in that: Extract object instances, cropped regions, and segmentation masks from sample images and target depth images; construct an object spatial relationship graph, which includes global constraints, distance constraints, position constraints, alignment constraints, and rotation constraints for objects; determine the correspondence between objects in the target image and objects in the sample image, and generate appropriate appearance descriptions for objects without corresponding objects in the target image.

3. The method for interior design appearance transfer based on diffusion model target perception according to claim 2, characterized in that: For each cropped object instance image, retrieve images of the same category from the text-image dataset as regularization images; Jointly optimize word tokens representing object appearance and criss-cross attention layers of a pre-trained diffusion model.

4. The method for interior design appearance migration based on diffusion model target perception as claimed in claim 2, characterized in that: The reconstruction loss is used to reduce the difference between the predicted noise and the real noise of the noisy image in the object area; Attention alignment loss is used to ensure that word tags focus on the object area rather than the entire cropped image; generalization loss is used to generalize the learned appearance features to target objects with different visual attributes; background loss is used to prevent the language drift problem of the pre-trained model.

5. The method for interior design appearance transfer based on diffusion model target perception according to claim 4, characterized in that: The attention alignment loss is implemented by aligning the cross-attention activations with the object segmentation mask, the generalization loss regularizes the learned labels by regularizing the attention activations of the category labels on the image, and the background loss is implemented by comparing the noise difference predicted in the background region by cues with and without appearance labels.

6. The method for interior design appearance transfer based on diffusion model target perception according to claim 4, characterized in that: In each time step of the diffusion process, the latent representations of the appearance and geometry of each object are calculated separately; based on the object segmentation mask, the latent representations of each object are combined into a unified latent representation; the noisy image of the current time step is denoised using the combined latent representation; the denoising process is iterated until the final image is generated.

7. The interior design appearance transfer system based on diffusion model target perception is characterized by: include: The data acquisition module is configured to: acquire a sample image and a target depth image; The appearance construction module is configured to: perform image segmentation and analysis on the example image and the target depth image to obtain a correspondence between instance objects and object appearances in the example image and the target depth image; The appearance inversion module is configured to: extract object-perceived appearance features from the example image based on the correspondence between the instance object and the object appearance; The appearance transfer module is configured to obtain a target scene image with expected appearance features based on appearance features and a target depth image, as well as a preset diffusion model; wherein the loss function in the diffusion model is a weighted superposition of a reconstruction loss function, an attention alignment loss function, a generalization loss function, and a background loss function.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the interior design appearance migration method based on diffusion model target perception as described in any one of claims 1 to 6 are implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the program, the steps of the interior design appearance migration method based on diffusion model target perception are implemented as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the interior design appearance migration method based on diffusion model target perception are implemented as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-instance controllable image generation method based on cross attention redistribution

    CN118628611A