Decoration effect diagram generation method and device based on generative artificial intelligence

Through a generative artificial intelligence-based method, using interference item removal, line extraction and depth recognition, combined with a pre-trained model to generate decoration renderings, the problems of low efficiency and resource waste in existing technologies are solved, and efficient and accurate decoration rendering generation is achieved.

CN119478251BActive Publication Date: 2025-09-12BEIJING NAT STANDARD CONSTR TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411852603.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-09-12
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

When generating decoration renderings, existing technologies have low efficiency in manual measurement and modeling, consume a lot of computing resources, and require redesign and rendering when customers require renderings of different styles, which increases time and resource consumption.

Method used

A generative artificial intelligence-based method is used to obtain images of the interior to be renovated and desired style prompts, and then perform interference removal, line extraction, depth recognition, and rendering generation. The pre-trained decoration image generation model is used to quickly output high-quality decoration renderings.

Benefits of technology

It reduces computing resource consumption, improves the accuracy and efficiency of generating decoration renderings, avoids the step of designers manually processing pictures, and ensures the quality and work efficiency of generating renderings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478251B_ABST
    Figure CN119478251B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a method and apparatus for generating decoration renderings based on generative artificial intelligence. A specific implementation of this method includes: obtaining an image of a room to be renovated and a desired style prompt; removing interference from the image of the room to be renovated to obtain an interior structural image; extracting lines from the interior structural image to obtain an interior line image; performing depth recognition on the interior structural image to obtain an interior depth image; generating a rendering of the interior structural image based on the desired style prompt, the interior line image, and the interior depth image based on a pre-trained decoration rendering generation model to obtain an initial interior rendering; and partially redrawing the initial interior rendering to obtain an interior decoration rendering. This implementation improves the efficiency of generating decoration renderings and avoids wasting computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method and device for generating decoration renderings based on generative artificial intelligence. Background Art

[0002] Before deciding on a home renovation style, renovation companies typically create renderings and submit them to clients, allowing them to review the design and choose the most satisfying option. These renderings help clients intuitively understand different design options and make informed decisions. Currently, creating renderings is typically done through manual on-site measurement and modeling. Designers create interior scene models based on the on-site measurements, then create renderings through rendering.

[0003] However, when using the above method to generate decoration renderings, there is often a technical problem: manual measurement and modeling require designers to invest a lot of time and energy, which is inefficient. Moreover, if customers want to view decoration renderings of different styles, they need to redesign and render, which increases the consumption of computing resources.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose a method and device for generating decoration renderings based on generative artificial intelligence to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide a method for generating decoration renderings based on generative artificial intelligence, the method comprising: obtaining an image of an interior to be renovated and a desired style prompt word; removing interference items from the image of the interior to be renovated to obtain an interior structure image; extracting lines from the image of the interior structure to obtain an interior line image; performing depth recognition on the image of the interior structure to obtain an interior depth image; based on a pre-trained decoration drawing generation model, generating a rendering for the interior structure image according to the desired style prompt word, the interior line image and the interior depth image to obtain an initial interior rendering; and partially redrawing the initial interior rendering to obtain an interior decoration rendering.

[0008] In a second aspect, some embodiments of the present disclosure provide a device for generating decoration renderings based on generative artificial intelligence, the device comprising: an acquisition unit, configured to acquire an image of an interior to be renovated and a desired style prompt word; an interference item removal unit, configured to remove interference items from the image of the interior to be renovated to obtain an interior structure image; a line extraction unit, configured to extract lines from the interior structure image to obtain an interior line image; a depth recognition unit, configured to perform depth recognition on the interior structure image to obtain an interior depth image; a rendering generation unit, configured to generate a rendering for the interior structure image based on a pre-trained decoration drawing generation model according to the desired style prompt word, the interior line image and the interior depth image to obtain an initial interior rendering; and a partial redrawing unit, configured to partially redraw the initial interior rendering to obtain an interior decoration rendering.

[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.

[0011] The above-described embodiments of the present disclosure have the following beneficial effects: The generative AI-based decoration rendering generation method of some embodiments of the present disclosure can reduce computing resource consumption. Specifically, the increased computing resource consumption is due to the fact that manual measurement and modeling require designers to invest a significant amount of time and effort, resulting in low efficiency. Furthermore, if a client wishes to view a decoration rendering in a different style, the design and rendering must be redesigned and re-rendered. Based on this, the generative AI-based decoration rendering generation method of some embodiments of the present disclosure first obtains an image of the interior to be renovated and a desired style prompt. Next, interference items are removed from the image to be renovated to obtain an interior structure image. This reduces irrelevant elements in the interior image, helps improve the accuracy of the generated rendering, and eliminates manual image processing by the designer, thereby improving work efficiency. Then, line extraction is performed on the interior structure image to obtain an interior line image. This allows the interior structure lines to be obtained as key information for generating the decoration rendering. Next, depth recognition is performed on the interior structure image to obtain an interior depth image. Automatic depth recognition can enhance the spatial expressiveness of the generated decoration rendering. Next, based on the pre-trained interior design generation model, the interior structure image is generated based on the desired style cues, the interior line image, and the interior depth image, resulting in an initial interior design. This model can quickly generate a high-quality interior design that meets the requirements. Finally, the initial interior design is partially redrawn to produce the final interior design. By making detailed adjustments to the initial interior design, the final design is generated, ensuring the quality of the generated design while significantly improving work efficiency and avoiding wasted computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0013] Figure 1 is a flowchart of some embodiments of a method for generating decoration renderings based on generative artificial intelligence according to the present disclosure;

[0014] Figure 2 It is a structural schematic diagram of some embodiments of the device for generating decoration effect pictures based on generative artificial intelligence according to the present disclosure;

[0015] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure;

[0016] Figure 4 This is a test case diagram of some embodiments of the method for generating decoration renderings based on generative artificial intelligence according to the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 The process 100 of some embodiments of the method for generating decoration effect pictures based on generative artificial intelligence according to the present disclosure is shown. The method for generating decoration effect pictures based on generative artificial intelligence includes the following steps:

[0024] Step 101: Obtain an indoor image to be transformed and desired style prompt words.

[0025] In some embodiments, the execution subject of the decoration rendering generation method based on generative artificial intelligence can obtain an image of the interior to be renovated and a desired style prompt word. The above-mentioned image of the interior to be renovated can be a photo taken of the interior status of the renovation project. The above-mentioned desired style prompt word includes positive prompt words and reverse prompt words. The above-mentioned positive prompt words are decoration style prompt words that meet the user's expectations. The above-mentioned reverse prompt words are decoration style prompt words that do not meet the user's expectations. For example, assuming that the user's desired decoration style is a white and simple style, the positive prompt words can be: "white", "minimalism", "modern", "bright", and the reverse prompt words can be: "messy", "complex", "bright", and the reverse prompt words can be: "messy", "complex", "bright", and "gorgeous". Here, the above-mentioned positive prompt words and reverse prompt words can be converted into English words.

[0026] In practice, the interior images to be transformed should fully cover the entire interior scene. The images are shot at a horizontal and vertical angle, ensuring that all interior details are accurately captured. The high-quality original images significantly improve the quality of the final rendering.

[0027] In practice, we've discovered that using AI models to generate interior design renderings often presents the following technical issues: Interference from interior scenes, such as furniture and appliances, can affect the resulting renderings, resulting in lower accuracy. To avoid these issues, we propose the following steps.

[0028] Step 102: remove interference items from the indoor image to be transformed to obtain an indoor structure image.

[0029] In some embodiments, the execution entity may remove interference items from the indoor image to be transformed to obtain an indoor structural image.

[0030] In some optional implementations of some embodiments, the execution subject removes interference items from the indoor image to be transformed to obtain the indoor structure image, which may include the following steps:

[0031] The first step is to perform semantic segmentation on the image of the room to be remodeled, obtaining a segmented image and a corresponding set of category labels. This can be done using a pre-defined semantic segmentation algorithm. The segmented image contains at least one segmented region, each of which corresponds to a category label.

[0032] As an example, the semantic segmentation algorithm may include but is not limited to at least one of the following: UperNet, DeepLab v3+, etc.

[0033] In the second step, each category label in the above-mentioned category label set that meets the preset conditions is determined as an interference label to obtain an interference label set. Among them, each category label in the above-mentioned category label set that corresponds to a furniture label can be determined as an interference label to obtain an interference label set. Each category label in the above-mentioned category label set that does not correspond to a background label can also be determined as an interference label to obtain an interference label set. The above-mentioned furniture labels can include but are not limited to "sofa", "TV", "air conditioner", etc. The above-mentioned background labels can include "wall", "floor", "ceiling", etc.

[0034] In the third step, interference items are removed from the indoor segmentation result image based on the interference label set to obtain an indoor structure image. First, the segmented regions corresponding to each interference label in the interference label set in the indoor segmentation result image can be identified as regions to be removed, thereby obtaining a set of regions to be removed. Subsequently, interference items can be removed from each region to be removed in the indoor segmentation result image based on a preset image inpainting algorithm to obtain an indoor structure image. Here, the indoor structure image only contains indoor structural elements such as walls, windows, and floors, and does not contain interference elements such as furniture and appliances.

[0035] As an example, the image repair algorithm may include but is not limited to at least one of the following: DeepFill v2 (depth repair 2), HiFill (high resolution repair), etc.

[0036] Optionally, the execution subject performs semantic segmentation on the indoor image to be transformed to obtain an indoor segmentation result image and a corresponding category label set, which may include the following steps:

[0037] The first step is to input the indoor image to be transformed into the multi-scale feature extraction network included in the indoor furniture segmentation model to obtain a multi-scale fusion feature map. The indoor furniture segmentation model includes a trained multi-scale feature extraction network and a semantic segmentation head network (Segmentation Head).

[0038] In the second step, the multi-scale fused feature map is input into the semantic segmentation head network to generate an indoor segmentation probability map and the corresponding initial set of class labels. Each pixel in the indoor segmentation probability map corresponds to a probability value. This probability value represents the probability that the current pixel belongs to a certain class. For example, if a pixel in the indoor segmentation probability map has a value of 0.8 and the pixel is labeled as a sofa, then the probability of the pixel belonging to the sofa class is 0.8.

[0039] The third step is to determine the indoor segmentation result image and the corresponding category label set corresponding to the indoor segmentation probability map and the initial category label set. First, the pixels in the indoor segmentation probability map with probability values ​​greater than a preset threshold can be determined as category pixels of the corresponding category. Secondly, the pixels in the indoor segmentation probability map with probability values ​​less than a preset threshold can also be determined as background pixels of the background category. Afterwards, all category pixels of a category are determined as a segmentation area, and the segmentation area corresponds to the category label of the category. For example, all background pixels corresponding to the background category can be determined as a background segmentation area, and the background segmentation area corresponds to the background category label. Finally, the indoor segmentation result image and the corresponding category label set can be determined based on each segmentation area and the category label corresponding to the segmentation area.

[0040] Optionally, the indoor furniture segmentation model is trained by the following steps:

[0041] The first step is to obtain an initial indoor furniture segmentation model. This initial indoor furniture segmentation model includes an initial multi-scale feature extraction network and an initial semantic segmentation head network. The initial multi-scale feature extraction network may include, but is not limited to, at least one of the following: a Feature Pyramid Network (FPN) or an Atrous Spatial Pyramid Pooling (ASPP) network. The initial semantic segmentation head network may include, but is not limited to, at least one of the following: a deconvolution layer or an upsampling layer.

[0042] The second step is to obtain an indoor training dataset, which includes sample indoor images, sample ground-truth segmentation images, and corresponding sample ground-truth category label sets.

[0043] In the third step, based on the preset rounds, the initial indoor furniture segmentation model is trained as follows:

[0044] In the first sub-step, a sample indoor image corresponding to each indoor training data in the indoor training dataset is input into the initial multi-scale feature extraction network to generate a sample multi-scale feature map, thereby obtaining a set of sample multi-scale feature maps. Each sample indoor image corresponds to one sample multi-scale feature map.

[0045] In a second sub-step, in response to the current round not being the first round, each sample multi-scale feature map in the sample multi-scale feature map set is fused with the sample segmentation probability map from the previous round to generate a fused feature map, thereby obtaining a fused feature map set. The fused feature map can be obtained by element-wise addition of the sample multi-scale feature map and the sample segmentation probability map. Each sample multi-scale feature map corresponds to one fused feature map.

[0046] If the current round is the first round, each sample multi-scale feature map in the above sample multi-scale feature map set is input into the above initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set.

[0047] In the third sub-step, each fused feature map in the fused feature map set is input into the initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set. Each sample segmentation probability map corresponds to an initial class label set, and each sample segmentation probability map contains at least one segmentation region. Each segmentation region corresponds to an initial class label.

[0048] The fourth sub-step is to determine a sample loss value corresponding to each sample segmentation probability map in the sample segmentation probability map set based on the sample true value segmentation image and the sample true value category label set corresponding to each indoor training data in the indoor training dataset. The sample loss value corresponding to each sample segmentation probability map in the sample segmentation probability map set can be determined using a preset loss function based on the sample true value segmentation image and the sample true value category label set corresponding to each indoor training data in the indoor training dataset.

[0049] As an example, the above loss function may include but is not limited to at least one of the following: cross-entropy loss (Cross-Entropy Loss), absolute error loss (L1 Loss), etc.

[0050] In the fifth sub-step, based on the sample loss value, the initial indoor furniture segmentation model is back-propagated to obtain the indoor furniture segmentation model of the current round.

[0051] In a sixth sub-step, in response to the current round being equal to the preset round, the current round indoor furniture segmentation model is used as the indoor furniture segmentation model. If the current round is equal to the preset round, the current round indoor furniture segmentation model is determined as the indoor furniture segmentation model, and training ends. If the current round is less than the preset round, sub-steps 1 to 6 are performed again.

[0052] The above step 102 and its related contents, as an inventive point of an embodiment of the present disclosure, solve the above technical problem that "the accuracy of the generated decoration rendering is low." The factors that lead to the above technical problem are often as follows: interference content such as furniture and home appliances in the indoor scene may affect the generated decoration rendering. If the above factors are solved, the accuracy of the generated decoration rendering can be improved. In order to achieve this effect, first, semantic segmentation can be performed on the indoor image to be transformed. By fully mining and utilizing the potential information in the probability map, the accuracy of semantic segmentation in complex scenes can be improved, so that an accurate indoor segmentation result map can be obtained. Afterwards, the label corresponding to the furniture category in the category label corresponding to the indoor segmentation result map can be determined as an interference label so that the interference content can be removed later. Then, through the image repair algorithm, the interference items are removed and the blank areas are filled to restore the original structure of the indoor environment. As a result, the subsequent generation of decoration renderings will not be affected by interference items such as furniture and home appliances, thereby improving the accuracy.

[0053] Step 103: extract lines from the indoor structure image to obtain an indoor line image.

[0054] In some embodiments, the execution entity may perform line extraction on the indoor structure image to obtain an indoor line image.

[0055] In some optional implementations of some embodiments, the execution subject extracts lines from the indoor structure image to obtain the indoor line image, which may include the following steps:

[0056] In the first step, for each pixel in the indoor structure image, the following filtering steps are performed to generate filtered pixels and obtain a filtered pixel set:

[0057] The first sub-step is to sort the grayscale values ​​corresponding to each pixel within a preset range of the pixel to obtain a neighborhood grayscale value sequence. First, each pixel covered by a square area centered on the pixel can be determined as a neighboring pixel of the pixel to obtain a neighboring pixel set. For example, all pixels within a 3×3 area centered on the pixel can be determined as neighboring pixels of the pixel. Second, the grayscale values ​​corresponding to each neighboring pixel in the neighboring pixel set can be sorted to obtain a neighborhood grayscale value sequence.

[0058] In a second sub-step, the mean of the neighborhood grayscale values ​​in the neighborhood grayscale value sequence that meet preset requirements is determined as the filtered grayscale value. The filtered grayscale value can be determined as the mean of the neighborhood grayscale values ​​in a middle window of the neighborhood grayscale value sequence. The length of the middle window can be half the size of the square area, rounded up, and the width is 1. The center of the middle window corresponds to the center of the neighborhood grayscale value sequence.

[0059] For example, if the preset range is a 3×3 range centered on the pixel, the middle window can be set to a 1×5 window with its center corresponding to the center of the neighborhood grayscale value sequence. The mean of the neighborhood grayscale values ​​covered by the middle window is determined as the filtered grayscale value.

[0060] In practice, by setting an intermediate window, the smaller and larger values ​​in the neighborhood grayscale value sequence are removed, and the final filtered grayscale value is determined only by the intermediate neighborhood grayscale values. In this way, more image details can be retained while ensuring the smoothness of the filtering result.

[0061] The third sub-step is to set the grayscale value of the pixel as the filtered grayscale value to obtain a filtered pixel.

[0062] The second step is to determine the corresponding filtered indoor image based on the above-mentioned filtered pixel set, wherein the corresponding filtered indoor image can be obtained by combining the respective filtered pixels in the above-mentioned filtered pixel set.

[0063] In a third step, a pixel gradient direction of each pixel in the filtered indoor image is determined based on the gradient magnitude between each pixel in the filtered indoor image and its neighboring pixels, thereby obtaining a pixel gradient direction set. The pixel gradient direction set can be obtained by determining the pixel gradient direction of each pixel in the filtered indoor image based on the gradient magnitude between each pixel in the filtered indoor image and its neighboring pixels using a preset edge detection operator. The preset edge detection operator may include a Canny operator.

[0064] Step 4: Apply a non-maximum constraint to each pixel gradient direction in the above pixel gradient direction set to obtain a constrained pixel gradient direction set. If a pixel corresponds to multiple pixel gradient directions, the pixel gradient directions whose gradient magnitudes are not the maximum can be constrained to obtain the constrained pixel gradient directions. The above constraint operation can be performed by setting the pixel gradient directions whose gradient magnitudes are not the maximum to 0.

[0065] In the fifth step, pixels corresponding to constrained pixel gradient directions that meet preset conditions in the constrained pixel gradient direction set are identified as edge points, thereby obtaining an edge point set. First, a gradient histogram of the indoor image can be generated based on the constrained pixel gradient direction set. Second, an adaptive double-threshold setting method and a maximum inter-class variance method can be used to determine edge points in the indoor image gradient histogram, thereby obtaining an edge point set.

[0066] In the sixth step, the edge points in the aforementioned edge point set are connected to generate an interior line image. This interior line image is generated by extracting edges from the interior structure image and representing them as lines. This allows for a precise representation of the interior's architectural features. This approach not only captures the basic outline and layout of the room, but also details the position and shape of walls, doors, windows, and other important structural elements, improving the accuracy of subsequent renovation renderings.

[0067] Step 104: Perform depth recognition on the indoor structure image to obtain an indoor depth image.

[0068] In some embodiments, the execution entity may perform depth recognition on the indoor structure image to obtain an indoor depth image.

[0069] In some optional implementations of some embodiments, the execution subject performs depth recognition on the indoor structure image to obtain the indoor depth image, which may include the following steps:

[0070] In the first step, global feature extraction is performed on the indoor structure image to obtain global indoor structure features. This can be performed using an encoder in a preset depth estimation model. The depth estimation model includes an encoder and a decoder.

[0071] As an example, the depth estimation model includes but is not limited to at least one of the following: Deeper Depth Prediction with Fully Convolutional Residual Networks, Deep Ordinal Regression Network for Monocular Depth Estimation (DORN)

[0072] In the second step, the global indoor structure features are upsampled to obtain an initial indoor depth image. The global indoor structure features can be upsampled by the decoder to obtain the initial indoor depth image. The initial indoor depth image has the same resolution as the indoor structure image.

[0073] The third step is to perform depth smoothing on the initial indoor depth image to obtain an indoor depth image. The depth smoothing can be performed on the initial indoor depth image using a preset depth smoothing algorithm to obtain the indoor depth image.

[0074] As an example, the depth smoothing algorithm may include but is not limited to at least one of the following: Laplacian Smoothing, Depth Map Optimization, etc.

[0075] Step 105 , based on the pre-trained decoration image generation model, and according to the desired style prompt words, the indoor line image, and the indoor depth image, an effect image is generated for the indoor structure image to obtain an initial indoor effect image.

[0076] In some embodiments, the execution entity may generate an effect diagram for the interior structure image based on a pre-trained decoration image generation model according to the desired style prompt words, the interior line image, and the interior depth image to obtain an initial interior effect diagram.

[0077] In some optional implementations of some embodiments, the execution subject generates a rendering of the interior structure image based on a pre-trained decoration image generation model according to the desired style prompt words, the interior line image, and the interior depth image to obtain an initial interior rendering, which may include the following steps:

[0078] In the first step, the indoor line image and the indoor depth image are determined as control signals. The indoor line image and the indoor depth image can be determined as control signals via ControlNet.

[0079] In the second step, based on the control signal, the desired style cues and the interior structure image are input into a pre-trained interior design image generation model to generate an initial interior rendering. First, the desired style cues and the interior structure image can be input into the pre-trained interior design image generation model. Second, ControlNet can be used to control the output of the interior design image generation model based on the control signal to generate the initial interior rendering. To ensure the speed of generating the initial interior rendering, the size of the generated initial interior rendering is controlled within a certain pixel size (for example, 512×512 pixels).

[0080] In some embodiments, the use case test diagram for generating the initial indoor effect diagram is as follows: Figure 4 As shown in the figure, the images generated by the decoration image generation model are highly random and cannot accurately reflect the actual situation of a project. To meet the needs of actual design business, we preprocess the input current photos in various aspects to extract different image information (such as lines and depth). This information is used as the control condition of the decoration image generation model to improve the accuracy of the generated decoration renderings.

[0081] Optionally, the pre-trained decoration image generation model is trained by the following steps:

[0082] The first step is to obtain a preset generative AI model and a sample dataset. The generative AI model can be a stable diffusion model (SD). Each sample data in the sample dataset includes a sample decoration image and a corresponding sample style prompt word set, where each sample decoration image corresponds to at least one sample style prompt word.

[0083] In practice, the sample style prompt words corresponding to the sample decoration images can be stored in a document. The first word in the document can be the trigger word. Sample decoration images with the same style can have the same trigger word. For example, the first word in the sample style prompt word document for all sample decoration images with a white minimalist style can be "white minimalist."

[0084] The second step is to preprocess each sample data in the sample data set to generate training data, thereby obtaining a training data set. First, the sample decoration images in the sample data can be cropped and filtered to generate training decoration images. Second, the training decoration images and their corresponding sample style cue word sets can be determined as training data, thereby obtaining a training data set.

[0085] The third step is to fine-tune the generative AI model to obtain a fine-tuned AI model. The generative AI model can be fine-tuned using a pre-defined fine-tuning technique. The fine-tuning technique can be a Low-Rank Adaptation of Large Language Models (LoRA) technique based on large models.

[0086] In practice, through low-rank adaptation technology, a new data processing layer is inserted into the original large model to avoid modifying the parameters of the original model, thereby achieving lightweight and fast training and improving training efficiency.

[0087] The fourth step is to train the fine-tuned AI model based on the training dataset to obtain a pre-trained model for generating interior decoration images. The fine-tuned AI model can be trained using the sample style cue word set for each training data point in the training dataset as input and the training interior decoration image for each training data point as the desired output to obtain a pre-trained model for generating interior decoration images.

[0088] In practice, the original stable diffusion model usually contains a large number of parameters, resulting in low efficiency and high computational cost when generating decoration renderings. By introducing low-rank adaptation technology and fine-tuning the model in combination with a specific training dataset in the decoration field, the performance of the model can be effectively optimized. The LoRA method adjusts and optimizes some model parameters by introducing a low-rank matrix based on the original model, avoiding the need to retrain the entire large model, thereby significantly reducing the consumption of computing resources and time. For the generation of decoration renderings, the fine-tuned model can more accurately capture the unique visual features and style requirements of the decoration field, improve the accuracy and diversity of the generated effects, while increasing the speed of generating renderings and reducing costs.

[0089] Step 106: Partially redraw the initial interior rendering to obtain an interior decoration rendering.

[0090] In some embodiments, the execution entity may partially redraw the initial interior rendering to obtain an interior decoration rendering. Specifically, the smaller initial interior rendering may be partially redrawn using a preset partial redrawing algorithm and a preset redrawing amplitude to obtain an interior decoration rendering. The redrawing amplitude may be a preset multiple (e.g., 1.7 times) of the original rendering.

[0091] As an example, the local redrawing algorithm may include but is not limited to at least one of the following: LaGAN (Local Attention GAN, local attention generative adversarial network), SRCNN (Super-Resolution Convolutional Neural Network, super-resolution convolutional neural network), etc.

[0092] The above-described embodiments of the present disclosure have the following beneficial effects: The generative AI-based decoration rendering generation method of some embodiments of the present disclosure can reduce computing resource consumption. Specifically, the increased computing resource consumption is due to the fact that manual measurement and modeling require designers to invest a significant amount of time and effort, resulting in low efficiency. Furthermore, if a client wishes to view a decoration rendering in a different style, the design and rendering must be redesigned and re-rendered. Based on this, the generative AI-based decoration rendering generation method of some embodiments of the present disclosure first obtains an image of the interior to be renovated and a desired style prompt. Next, interference items are removed from the image to be renovated to obtain an interior structure image. This reduces irrelevant elements in the interior image, helps improve the accuracy of the generated rendering, and eliminates manual image processing by the designer, thereby improving work efficiency. Then, line extraction is performed on the interior structure image to obtain an interior line image. This allows the interior structure lines to be obtained as key information for generating the decoration rendering. Next, depth recognition is performed on the interior structure image to obtain an interior depth image. Automatic depth recognition can enhance the spatial expressiveness of the generated decoration rendering. Next, based on the pre-trained interior design generation model, the interior structure image is generated based on the desired style cues, the interior line image, and the interior depth image, resulting in an initial interior design. This model can quickly generate a high-quality interior design that meets the requirements. Finally, the initial interior design is partially redrawn to produce the final interior design. By making detailed adjustments to the initial interior design, the final design is generated, ensuring the quality of the generated design while significantly improving work efficiency and avoiding wasted computing resources.

[0093] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for generating decoration effect pictures based on generative artificial intelligence. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0094] like Figure 2As shown, in some embodiments, a device 200 for generating decoration renderings based on generative artificial intelligence includes: an acquisition unit 201, an interference elimination unit 202, a line extraction unit 203, a depth recognition unit 204, a rendering generation unit 205, and a partial redrawing unit 206. The acquisition unit 201 is configured to acquire an image of a room to be renovated and a desired style prompt word; the interference elimination unit 202 is configured to eliminate interference from the image of the room to be renovated to obtain an indoor structure image; the line extraction unit 203 is configured to extract lines from the indoor structure image to obtain an indoor line image; the depth recognition unit 204 is configured to perform depth recognition on the indoor structure image to obtain an indoor depth image; the rendering generation unit 205 is configured to generate a rendering based on the desired style prompt word, the indoor line image, and the indoor depth image based on a pre-trained decoration rendering generation model to obtain an initial indoor rendering; and the partial redrawing unit 206 is configured to partially redraw the initial indoor rendering to obtain an indoor decoration rendering.

[0095] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be repeated here.

[0096] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0097] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory 302 or a program loaded from a storage device 308 into a random access memory 303. Various programs and data required for the operation of the electronic device 300 are also stored in the random access memory 303. The processing device 301, the read-only memory 302, and the random access memory 303 are connected to each other via a bus 304. An input / output interface 305 is also connected to the bus 304.

[0098] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0099] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the read-only memory 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0100] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0101] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0102] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: obtains an image of the interior to be renovated and a desired style prompt; removes interference from the image of the interior to be renovated to obtain an interior structure image; extracts lines from the image of the interior structure to obtain an interior line image; performs depth recognition on the image of the interior structure to obtain an interior depth image; generates a rendering of the interior structure image based on the desired style prompt, the interior line image, and the interior depth image based on a pre-trained decoration image generation model to obtain an initial interior rendering; and partially redraws the initial interior rendering to obtain an interior decoration rendering.

[0103] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0105] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor including an acquisition unit, an interference item removal unit, a line extraction unit, a depth recognition unit, an effect diagram generation unit, and a local redrawing unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the line extraction unit may also be described as a "unit for extracting lines from indoor structural images."

[0106] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0107] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for generating decoration renderings based on generative artificial intelligence, comprising: Obtain an image of the interior to be renovated and desired style prompt words; Eliminating interference items from the indoor image to be transformed to obtain an indoor structural image; Performing line extraction on the indoor structure image to obtain an indoor line image; Performing depth recognition on the indoor structure image to obtain an indoor depth image; Based on the pre-trained decoration image generation model, according to the desired style prompt words, the indoor line image and the indoor depth image, the indoor structure image is subjected to rendering generation to obtain an initial indoor rendering; Partially redrawing the initial interior rendering to obtain an interior decoration rendering; The step of removing interference items from the indoor image to be transformed to obtain an indoor structural image includes: Performing semantic segmentation on the indoor image to be transformed to obtain an indoor segmentation result image and a corresponding category label set, wherein the indoor segmentation result image includes at least one segmentation region, and each segmentation region corresponds to one category label; Determine each category tag in the category tag set that meets a preset condition as an interference tag to obtain an interference tag set, wherein each category tag in the category tag set corresponding to a furniture tag is determined as an interference tag to obtain an interference tag set; According to the interference label set, interference items are removed from the indoor segmentation result image to obtain an indoor structure image, wherein the segmented area corresponding to each interference label in the interference label set in the indoor segmentation result image is determined as a to-be-removed area to obtain a to-be-removed area set; based on a preset image inpainting algorithm, interference items are removed from each to-be-removed area in the indoor segmentation result image in turn to obtain an indoor structure image; the indoor structure image only contains indoor structural elements such as walls, windows, and floors, and does not contain interference elements such as furniture and household appliances; The step of performing semantic segmentation on the indoor image to be transformed to obtain an indoor segmentation result image and a corresponding category label set includes: Inputting the indoor image to be transformed into a multi-scale feature extraction network included in the indoor furniture segmentation model to obtain a multi-scale fusion feature map; Inputting the multi-scale fusion feature map into the semantic segmentation head network to obtain an indoor segmentation probability map and a corresponding initial category label set; Determine an indoor segmentation result image and a corresponding category label set corresponding to the indoor segmentation probability map and the initial category label set; The indoor furniture segmentation model is trained by the following steps: Acquire an initial indoor furniture segmentation model, wherein the initial indoor furniture segmentation model includes an initial multi-scale feature extraction network and an initial semantic segmentation head network; Obtain an indoor training data set, wherein the indoor training data includes sample indoor images, sample true value segmentation images, and corresponding sample true value category label sets; Based on the preset rounds, the following training steps are performed on the initial indoor furniture segmentation model: Inputting a sample indoor image corresponding to each indoor training data in the indoor training dataset into the initial multi-scale feature extraction network to generate a sample multi-scale feature map, thereby obtaining a sample multi-scale feature map set; In response to the current round not being the first round, fusing each sample multi-scale feature map in the sample multi-scale feature map set with the sample segmentation probability map of the previous round to generate a fused feature map, thereby obtaining a fused feature map set, wherein the sample multi-scale feature map and the sample segmentation probability map are element-wise added to obtain the fused feature map; If the current round is the first round, input each sample multi-scale feature map in the sample multi-scale feature map set into the initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set; Inputting each fused feature map in the fused feature map set into the initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set; Determine, according to the sample true value segmentation image and the sample true value category label set corresponding to each indoor training data in the indoor training data set, the sample loss value corresponding to each sample segmentation probability map in the sample segmentation probability map set; Performing backpropagation on the initial indoor furniture segmentation model according to the sample loss value to obtain a current round indoor furniture segmentation model; In response to the current round being equal to the preset round, the current round indoor furniture segmentation model is used as the indoor furniture segmentation model, wherein, if the current round is equal to the preset round, the current round indoor furniture segmentation model is determined as the indoor furniture segmentation model and the training is terminated; if the current round is less than the preset round, the training step is performed again.

2. The method according to claim 1, wherein The pre-trained decoration image generation model is trained by the following steps: Obtaining a preset generative artificial intelligence model and a sample data set, wherein each sample data in the sample data set includes a sample decoration image and a corresponding sample style prompt word set, and the sample decoration image corresponds to at least one sample style prompt word; Preprocessing each sample data in the sample data set to generate training data, thereby obtaining a training data set; Fine-tuning the generative artificial intelligence model to obtain a fine-tuned artificial intelligence model; The fine-tuning artificial intelligence model is trained according to the training data set to obtain a pre-trained decoration picture generation model.

3. The method according to claim 1, wherein The extracting lines from the indoor structure image to obtain an indoor line image includes: For each pixel in the indoor structure image, the following filtering steps are performed to generate a filtered pixel, thereby obtaining a filtered pixel set: Sorting the grayscale values ​​corresponding to each pixel within a preset range of the pixel to obtain a neighborhood grayscale value sequence; Determine the mean value between each neighborhood grayscale value that meets the preset requirements in the neighborhood grayscale value sequence as the filtered grayscale value; Setting the grayscale value of the pixel to the filtered grayscale value to obtain a filtered pixel; determining a corresponding filtered indoor image according to the filtered pixel set; determining a pixel gradient direction of each pixel in the filtered indoor image according to a gradient magnitude between each pixel in the filtered indoor image and its neighboring pixels, to obtain a pixel gradient direction set; performing a non-maximum constraint on each pixel gradient direction in the pixel gradient direction set to obtain a constrained pixel gradient direction set; Determining pixels corresponding to the constrained pixel gradient directions that meet preset conditions in the constrained pixel gradient direction set as edge points to obtain an edge point set; Each edge point in the edge point set is connected to generate an indoor line image.

4. The method according to claim 1, wherein The performing depth recognition on the indoor structure image to obtain the indoor depth image includes: Performing global feature extraction on the indoor structure image to obtain global indoor structure features; Upsampling the global indoor structure feature to obtain an initial indoor depth image, wherein the initial indoor depth image has the same resolution as the indoor structure image; Perform depth smoothing on the initial indoor depth image to obtain an indoor depth image.

5. The method according to claim 1, wherein The expected style prompt words include positive prompt words and negative prompt words, the positive prompt words are prompt words of the decoration style that meets the user's expectations, and the negative prompt words are prompt words of the decoration style that does not meet the user's expectations.

6. The method according to claim 1, wherein The pre-trained decoration image generation model generates a rendering of the indoor structure image according to the desired style prompt word, the indoor line image, and the indoor depth image to obtain an initial indoor rendering, including: determining the indoor line image and the indoor depth image as a control signal; According to the control signal, the desired style prompt word and the indoor structure image are input into a pre-trained decoration picture generation model to generate an initial indoor rendering.

7. A device for generating decoration renderings based on generative artificial intelligence, comprising: an acquisition unit configured to acquire an image of a room to be remodeled and a desired style prompt word; an interference item removal unit configured to remove interference items from the indoor image to be transformed to obtain an indoor structural image; a line extraction unit configured to extract lines from the indoor structure image to obtain an indoor line image; a depth recognition unit configured to perform depth recognition on the indoor structure image to obtain an indoor depth image; An effect image generation unit is configured to generate an effect image for the interior structure image based on a pre-trained decoration image generation model according to the desired style prompt words, the interior line image, and the interior depth image to obtain an initial interior effect image; a partial redrawing unit configured to perform partial redrawing on the initial interior rendering to obtain an interior decoration rendering; The interference item elimination unit is configured to: Performing semantic segmentation on the indoor image to be transformed to obtain an indoor segmentation result image and a corresponding category label set, wherein the indoor segmentation result image includes at least one segmentation region, and each segmentation region corresponds to one category label; Determine each category tag in the category tag set that meets a preset condition as an interference tag to obtain an interference tag set, wherein each category tag in the category tag set corresponding to a furniture tag is determined as an interference tag to obtain an interference tag set; According to the interference label set, interference items are removed from the indoor segmentation result image to obtain an indoor structure image, wherein the segmented area corresponding to each interference label in the interference label set in the indoor segmentation result image is determined as a to-be-removed area to obtain a to-be-removed area set; based on a preset image inpainting algorithm, interference items are removed from each to-be-removed area in the indoor segmentation result image in turn to obtain an indoor structure image; the indoor structure image only contains indoor structural elements such as walls, windows, and floors, and does not contain interference elements such as furniture and household appliances; The step of performing semantic segmentation on the indoor image to be transformed to obtain an indoor segmentation result image and a corresponding category label set includes: Inputting the indoor image to be transformed into a multi-scale feature extraction network included in the indoor furniture segmentation model to obtain a multi-scale fusion feature map; Inputting the multi-scale fusion feature map into the semantic segmentation head network to obtain an indoor segmentation probability map and a corresponding initial category label set; Determine an indoor segmentation result image and a corresponding category label set corresponding to the indoor segmentation probability map and the initial category label set; The indoor furniture segmentation model is trained by the following steps: Acquire an initial indoor furniture segmentation model, wherein the initial indoor furniture segmentation model includes an initial multi-scale feature extraction network and an initial semantic segmentation head network; Obtain an indoor training data set, wherein the indoor training data includes sample indoor images, sample true value segmentation images, and corresponding sample true value category label sets; Based on the preset rounds, the following training steps are performed on the initial indoor furniture segmentation model: Inputting a sample indoor image corresponding to each indoor training data in the indoor training dataset into the initial multi-scale feature extraction network to generate a sample multi-scale feature map, thereby obtaining a sample multi-scale feature map set; In response to the current round not being the first round, fusing each sample multi-scale feature map in the sample multi-scale feature map set with the sample segmentation probability map of the previous round to generate a fused feature map, thereby obtaining a fused feature map set, wherein the sample multi-scale feature map and the sample segmentation probability map are element-wise added to obtain the fused feature map; If the current round is the first round, input each sample multi-scale feature map in the sample multi-scale feature map set into the initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set; Inputting each fused feature map in the fused feature map set into the initial semantic segmentation head network to generate a sample segmentation probability map, thereby obtaining a sample segmentation probability map set; Determine, according to the sample true value segmentation image and the sample true value category label set corresponding to each indoor training data in the indoor training data set, the sample loss value corresponding to each sample segmentation probability map in the sample segmentation probability map set; Performing backpropagation on the initial indoor furniture segmentation model according to the sample loss value to obtain a current round indoor furniture segmentation model; In response to the current round being equal to the preset round, the current round indoor furniture segmentation model is used as the indoor furniture segmentation model, wherein, if the current round is equal to the preset round, the current round indoor furniture segmentation model is determined as the indoor furniture segmentation model and the training is terminated; if the current round is less than the preset round, the training step is performed again.

8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image generation method and device and storage medium

    CN117635760A