An Adaptive Generation Method and System for Indoor Panoramic Images Based on AIGC

Through the combination of MiDiffusion, ControlNet and self-trained LoRA model, the calculation complexity and style adaptation problems in panoramic image generation are solved, and efficient and accurate indoor panoramic image generation is achieved to meet the design needs of domestic users.

CN120070776BActive Publication Date: 2025-07-18COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510543799.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the process of indoor space design, the existing technology has problems such as high computational complexity, long rendering, imperfect image seam processing, insufficient understanding of the three-dimensional spatial structure of AI models, and difficulty in meeting the needs of domestic users in style migration, resulting in low design efficiency and insufficient practicality.

Method used

The MiDiffusion model is used to generate layout data that meets the constraints, and the panoramic image generation is combined with ControlNet and self-trained LoRA model. By extracting line drawings and depth maps as constraint guidance, overlapping sub-image areas are divided and dynamic masks are set. The seams are optimized using horizontal gradient masks to generate a 720° panoramic image that conforms to the domestic design style.

Benefits of technology

It improves the accuracy and style adaptability of panoramic image generation, reduces image distortion and unreasonable layout, realizes smooth transition and visual coherence between images, and improves design efficiency and computing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070776B_ABST
    Figure CN120070776B_ABST
Patent Text Reader

Abstract

The present invention discloses an AIGC-based method and system for adaptively generating indoor panoramic images, belonging to the field of indoor space design, which includes the following steps: S1. Obtain a floor plan attached with a description of the expected design style language and perform preprocessing; S2. Construct a layout: Use a pre-trained MiDiffusion model to generate layout data that meets the constraint conditions; S3. Generate a panoramic image: Based on the layout data, construct a basic 3D model, then supplement key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; S4. Generate a panoramic design plan; S5. Output a 720° panoramic image. By adopting the above AIGC-based method and system for adaptively generating indoor panoramic images, not only can the design efficiency be significantly improved, but also the requirements for rapid iteration and personalized customization can be met, which is conducive to promoting the intelligent and digital development of the indoor space design field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of indoor space design, and particularly to an indoor panoramic view adaptive generation method and system based on AIGC. Background Art

[0002] With the rapid iterative development of virtual reality (VR) and augmented reality (AR) technologies, the indoor space design field is undergoing a transformation from traditional two-dimensional plane displays to immersive three-dimensional visualization solutions. The modern consumer market has put forward higher requirements for indoor space design, not only needing to meet basic functional requirements, but also emphasizing the rationality of space layout, the diversity of design styles, and the refinement of visual presentation effects.

[0003] In the context of the rapid development of information technology and multimedia technology, 720° panoramic images have gradually become the industry standard display solution due to their advantage of all-round three-dimensional display. 720° panoramic images have significant advantages compared with traditional static floor plans or partial perspective display methods. This innovative display method can not only completely present spatial structure information, helping users intuitively understand the overall layout, but also support multi-angle observation. When combined with VR glasses, it can also achieve a truly immersive space experience. It can be seen that panoramic images can accurately display design details and decorative styles, providing strong support for the rapid evaluation and optimization of design schemes.

[0004] However, current methods usually require designers to use CAD drawing, 3D modeling software (such as Maya, C4D, Blender, etc.) and rendering tools in sequence to obtain local and even indoor space panoramic effect drawings. The entire process is cumbersome and time-consuming. It can be seen that the existing methods not only require a large number of professional designers to participate, consuming a huge amount of human resources, but also have a long design cycle, making it difficult to meet the market's needs for creative diversity and rapid design iteration. Especially when carrying out personalized customization, the efficiency of scheme optimization is low, and it is difficult to respond promptly to changes in customer needs.

[0005] To solve the above problems, the prior art further applies artificial intelligence generated content (AIGC) technology to the indoor space design field and has made breakthrough progress. Among them, the technical solutions based on generative adversarial networks (GANs) and diffusion models can significantly improve design efficiency and reduce production costs.

[0006] However, existing AIGC technologies still face many challenges in automatic layout generation. Although the prior art has proposed new solutions through 3D model generation and material parameter optimization, there are still problems such as high computational complexity and long rendering time in practical applications. At the same time, in the face of the need for style transfer, the precise control of layout constraint conditions is also a technical difficulty, making it difficult to achieve rapid response.

[0007] In the field of panoramic image generation, there are still several key problems in the existing technologies that need to be solved urgently. Firstly, the image seam processing technology is imperfect, which is prone to visual breaks. Secondly, the AI models have insufficient understanding of the three-dimensional space structure, especially with obvious defects in dealing with panoramic lens distortion. Finally, it is difficult to control details during the generation process, which directly affects the quality performance of the final image. The above technical difficulties severely restrict the application effect of panoramic image generation technology.

[0008] In the actual use process, there are also problems of insufficient adaptability in the existing AI generation methods. The current models are mainly trained based on internationally common data, and it is difficult to accurately grasp the characteristics and requirements of domestic spatial designs such as the living environment. For example, there are obvious differences in indoor space design at home and abroad in terms of furniture configuration, decoration style, etc. due to differences in living habits, usage scenarios, and aesthetics. This makes it difficult for the design schemes directly generated by existing AI models to fully meet the actual needs of domestic users, affecting the practicality of the technical solutions.

[0009] In summary, the above problems severely restrict the wide application of AI technology in the field of indoor space design. Especially in aspects such as panoramic image generation, spatial structure understanding, and localization adaptation, there is still a large room for improvement in the existing technical solutions. Summary of the Invention

[0010] The purpose of the present invention is to provide an AIGC-based adaptive indoor panoramic map generation method and system to solve the above technical problems.

[0011] To achieve the above purpose, the present invention provides an AIGC-based adaptive indoor panoramic map generation method, including the following steps:

[0012] S1. Obtain a floor plan with a language description of the expected design style and perform preprocessing;

[0013] S2. Construct a layout: Based on the preprocessed floor plan, use the pre-trained MiDiffusion model to generate layout data that meets the constraint conditions;

[0014] S3. Generate a panoramic image: Based on the layout data, construct a basic three-dimensional model, then supplement key structural elements to construct a complete three-dimensional room model, and then perform panoramic image rendering to generate a panoramic white film image;

[0015] S4. Generate a panoramic design plan: Extract the line drawing and depth map of the panoramic white film as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic view. Then divide the overlapping sub-image areas of the preliminary panoramic view, set dynamic masks for optimizing each sub-region, use a horizontal gradient mask to optimize and complete the fusion image output, and obtain a 720° panoramic view;

[0016] S5. Output the 720° panoramic view.

[0017] Preferably, in step S1, the obtained floor plan is from the floor plan in the 3D-FRONT database or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed:

[0018] The first step: Construct CSV data for the room list. The CSV data contains two types of data, room number and training. The training type data is divided into a training set and a test set, and the CSV data is kept in the same format as the room data in the 3D-FRONT database;

[0019] The second step: The user inputs the length and width data of the room, then uses the forward projection method to generate a plane grid matching the room length and width data, and records the projection result of the plane grid from a top-down perspective to generate a room mask image. At the same time, the room type is default set to the user's room, and then the geometric features of the plane grid are entered to generate spatial feature data, and other features of the room are filled with random data of the same dimension to obtain all the spatial data of the custom room;

[0020] The third step: Use the PointNet sampler to collect the object edge data of the custom room spatial data, and then use the pre-trained hybrid diffusion model to verify the usability of the edge feature information, and synchronously modify the configuration file used for layout generation to a custom configuration file.

[0021] Preferably, the MiDiffusion model described in step S2 is a hybrid discrete-continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network. The input layer uses a PointNet feature extractor to convert the spatial distribution of the room into an encoded format; in the denoising network, the VQ-Diffusion architecture is modified to a Transformer decoder module, and at the same time, a multi-head attention mechanism is introduced to introduce constraint conditions during step-by-step denoising;

[0022] And the MiDiffusion model is trained and optimized through the KL divergence loss function and the negative log probability loss function.

[0023] Preferably, step S3 specifically includes the following steps:

[0024] S31. Export the basic 3D model of the room from the generated layout data, where the basic 3D model includes a floor model and a furniture model;

[0025] S32. Supplement key structural elements based on the exported floor model, and the key structural elements include walls and ceilings. Then combine the supplemented wall, ceiling, floor model and furniture model to obtain a complete 3D room model;

[0026] S33. Panoramic image rendering: Place an internal camera in the 3D room model according to the preset perspective, set multiple point lights, and then use the rendering method of the physically based ray tracing rendering technology in the 3D software Blender. Select the equidistant cylindrical projection as the spatial structure projection type of the panoramic image to generate a panoramic white film image with an aspect ratio of 2:1.

[0027] Preferably, step S4 specifically includes the following steps:

[0028] S41. Generate a preliminary panoramic image;

[0029] S411. Extract the constraint image: Input the panoramic white film image, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map to obtain the 3D depth information of the spatial layout;

[0030] S412. Constraint-guided generation: Input the extracted line drawing and depth map into the line constraint model in the ControlNet model and the depth constraint model in the ControlNet model respectively. At the same time, use the self-trained LoRA model to optimize the style and structure of the panoramic white film image to obtain a preliminarily generated panoramic image;

[0031] S413. Image super-resolution processing: Use the RealESRGAN model to quadruple the size of the preliminarily generated panoramic image, retaining the spatial structure and features during generation to obtain a preliminary panoramic image;

[0032] S42. Divide the overlapping sub-image areas of the preliminary panoramic image: Divide the preliminary panoramic image into a left sub-image area, a middle sub-image area, and a right sub-image area arranged in a surrounding manner, a total of three sub-image areas, and there is an overlapping area between each adjacent sub-image area;

[0033] S43. Dynamic mask setting: The mask of the left sub-image area completely covers its entire image area and is set to 0; the mask of the middle sub-image area covers the area except the overlapping area with the left sub-image area; the mask of the right sub-image area is the middle area excluding the overlapping areas of the left and middle sub-image areas and the overlapping areas of the right and left sub-image areas;

[0034] S44. Model-guided optimization: During the generation process, the ControlNet model and the self-trained LoRA model are introduced. The ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used to guide style adjustment. At the same time, image generation is guided by fixing the random seed and the style language description provided by the user.

[0035] S45. Horizontal gradient mask fusion: Set a mask with a transparent gradient in the overlapping area to achieve a completely continuous visual experience and obtain a 720° panoramic image.

[0036] Preferably, in steps S412 and S44, the guiding parameters generated by the self-trained LoRA model are set between 0.35 and 0.65, and the training steps of the self-trained LoRA model are as follows:

[0037] The first step. Collect the panoramic image dataset: Collect six-direction image data of the indoor space. The six-direction image data includes the geographical location, name, type of the room, and the corresponding six-direction images (front, back, left, right, and sky and ground). After classifying and sorting by room type, import the six-direction images into the Blender 3D software, and use the method of converting the cube to equidistant cylindrical projection to convert the six-direction images into panoramic images, and render and export them to obtain the panoramic image dataset.

[0038] The second step. Data annotation and preparation: Use the WD14 tool to annotate the synthesized panoramic images. The annotation data includes furniture, decoration, and spatial structure, and add unified text description labels.

[0039] The third step. Model training: Input the panoramic image dataset and the annotation data into the LoRA model, and train the LoRA model with the cross-entropy loss function. During this process, adjust the pre-trained weights of the LoRA model based on the annotation data, and use low-rank decomposition to complete parameter update.

[0040] Preferably, the formula for adjusting the pre-trained weights of the LoRA model is as follows:

[0041] ;

[0042] ;

[0043] In the formula, represents the original weight matrix of the LoRA model; represents the weight update matrix obtained by low-rank decomposition; and respectively represent matrices with shapes of and ; represents the low-rank dimension; and respectively represent the rows and columns of the weight matrix; represents the output vector of the LoRA model; represents the input vector of the LoRA model;

[0044] The expression of the cross-entropy loss function is as follows:

[0045] ;

[0046] In the formula represents the total loss value of model training; represents the total number of training samples; represents the th sample's loss value; represents the th sample's true label for class ; represents the th sample's predicted probability of being class ; represents the number of classes.

[0047] Preferably, in step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3, and when generating the right sub-image region, the left 1 / 3 overlapping region of the left sub-image region is spliced to the right 1 / 3 overlapping region of the right sub-image region to ensure the spatial structure consistency of the image;

[0048] After the left sub-image region is generated, update the left 1 / 3 overlapping region of the middle sub-image region and the right 1 / 3 overlapping region of the right sub-image region; after the middle sub-image region is generated, update the left 1 / 3 overlapping region of the right sub-image region;

[0049] And the mask edge regions of the three sub-image regions are all feathered.

[0050] An AIGC-based indoor panoramic image adaptive generation system for implementing an AIGC-based indoor panoramic image adaptive generation method, including:

[0051] Data acquisition and preprocessing module: used to acquire the expected design style language description and the floor plan and perform preprocessing;

[0052] Layout construction module, used to generate layout data that meets the constraint conditions based on the preprocessed floor plan by using a pre-trained MiDiffusion model;

[0053] Panoramic image generation module, used to construct a basic 3D model based on the layout data, then supplement key structural elements, construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image;

[0054] The panoramic design solution generation module is used to extract the line drawing and depth map of the panoramic white film as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image. Then, divide the overlapping sub-image regions of the preliminary panoramic image, set dynamic masks for optimizing each sub-region, and use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image;

[0055] The output module is used to output the 720° panoramic image.

[0056] Therefore, the present invention adopts the above-mentioned indoor panoramic image adaptive generation method and system based on AIGC, and the beneficial effects are as follows:

[0057] 1. High generation accuracy: Extract the line drawing and depth map of the panoramic white film, and use ControlNet for constraint guidance, which can accurately control the spatial structure and contour of the image, making the generated panoramic image more in line with the actual situation in terms of geometric structure and spatial layout, reducing distortion and unreasonable layout;

[0058] 2. Strong style adaptability: Use the self-trained LoRA model, train based on the domestic indoor space design style dataset, can specifically optimize the panoramic image style, meet specific style requirements, and make the image style more suitable for the actual application scenario and user preferences;

[0059] 3. Seam processing: Divide the overlapping sub-image regions and use dynamic masks and horizontal gradient masks to optimize the seam fusion, which can effectively eliminate the splicing traces, achieve smooth transition between images, ensure the visual coherence and integrity of the panoramic image, and bring a high-quality visual experience;

[0060] 4. Efficient and flexible training: The LoRA model training adopts low-rank decomposition, reduces the training parameters, improves the training efficiency, and does not change the weights of the pre-trained model. The fine-tuning is flexible, can quickly adapt to different task and style requirements, and saves computing resources and time costs.

[0061] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings

[0062] Figure 1 It is a flowchart of an indoor panoramic image adaptive generation method based on AIGC according to the present invention;

[0063] Figure 2 It is an example flowchart of an indoor panoramic image adaptive generation method based on AIGC according to the present invention;

[0064] Figure 3Layout generation framework diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention;

[0065] Figure 4 Example mask diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is the mask diagram of a certain room in the 3D-FRONT database, and (b) is the user-defined room mask diagram;

[0066] Figure 5 MiDiffusion model training process diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is the change curve diagram of the learning rate, and (b) is the change curve diagram of the loss function;

[0067] Figure 6 Rendering preview example diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention;

[0068] Figure 7 Three-dimensional room model construction diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is an example diagram of the floor model, (b) is an example diagram of the furniture model, and (c) is an example diagram of the complete three-dimensional room model;

[0069] Figure 8 Panoramic white film diagram example of an indoor panoramic view adaptive generation method based on AIGC according to the present invention;

[0070] Figure 9 Panoramic white film diagram feature extraction diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is the extracted line drawing, and (b) is the extracted depth map;

[0071] Figure 10 Preliminary panoramic example diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention;

[0072] Figure 11 Preliminary panoramic diagram division result diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is an example diagram of the left sub-image area, (b) is an example diagram of the middle sub-image area, and (c) is an example diagram of the right sub-image area;

[0073] Figure 12 Surrounding stitching comparison diagram of an indoor panoramic view adaptive generation method based on AIGC according to the present invention, where (a) is the surrounding stitching example diagram before optimization, and (b) is the surrounding stitching example diagram after horizontal gradient mask fusion optimization. Detailed implementation method

[0074] In order to make the purpose, technical solution and advantages of the embodiments disclosed in the present invention more clear and understandable, the following further details the embodiments of the present invention in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention and are not used to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of this application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout.

[0075] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0076] The following further details the embodiments of the present invention in conjunction with the accompanying drawings.

[0077] As Figures 1 - 3 shown, a method for adaptively generating an indoor panoramic view based on AIGC includes the following steps:

[0078] S1. Obtain a floor plan with a description of the expected design style language and perform preprocessing;

[0079] In step S1, the obtained floor plan is from the floor plan in the 3D-FRONT database (3D-FRONT is short for 3DFurnished Rooms with layOuts and semaNTics. This dataset is a large 3D scene dataset jointly open-sourced by Alibaba, Simon Fraser University, and the Institute of Computing, Chinese Academy of Sciences. This dataset contains 6,813 different real house types, which are composed of 51,708 rooms. The room types are rich and can be subdivided into 28 types; among them, 19,775 rooms contain manually verified indoor design information, providing high-quality samples for indoor design-related research and applications), or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed:

[0080] The first step: Construct CSV data for the room list. The CSV data contains two types of data: room number and training. The training type data is divided into a training set and a test set, and the CSV data is consistent with the room data format of the 3D-FRONT database;

[0081] The second step: AsFigure 4 As shown, the length and width data of the room input by the user are used to generate a plane grid that matches the room length and width data using the forward projection method. The projection result of the plane grid is recorded from a top-down perspective to generate a room mask image. At the same time, the room type is default set to the user's room, and then the geometric features of the plane grid are recorded to generate spatial feature data, such as point data, surface data, etc. The other features of the room are filled with random data of the same dimension to obtain custom room spatial data, ensuring that the file format is the same as that of other spatial data in the dataset;

[0082] In the third step, the object edge data of the custom room spatial data is collected using a PointNet sampler, and then the pre-trained hybrid diffusion model is used to verify the usability of the edge feature information. The configuration file used for layout generation is synchronously modified to a custom configuration file (such as mask.png and boxes.npz). These data will be used as conditional inputs for subsequent generation to ensure that the generated layout meets the spatial requirements of the deep learning model.

[0083] S2. Construct a layout: Based on the preprocessed floor plan, use the pre-trained MiDiffusion model to generate layout data that meets the constraint conditions;

[0084] The MiDiffusion model described in step S2 is a hybrid discrete-continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network. The input layer uses a PointNet feature extractor to convert the spatial distribution of the room into an encoded format, providing the basic data representation for the subsequent learning and generation tasks of the model, enabling the model to effectively process and understand spatial information. In the denoising network, the VQ-Diffusion architecture is modified to a Transformer decoder module, and at the same time, a multi-head attention mechanism is introduced to introduce constraint conditions during step-by-step denoising. The introduction of the multi-head attention mechanism enables the model to better consider various spatial relationships and semantic information during denoising, thereby improving the accuracy of denoising and the quality of the generated results.

[0085] Such as Figure 5As shown in the figure, in order to ensure that the model can generate a better spatial layout based on the floor plan, the position limit of MiDiffusion was turned on for retraining. The MiDiffusion model was trained and optimized using the KL divergence loss function and the negative log probability loss function. The KL divergence loss function ensures that the model can accurately infer the next layout structure from the latent space of the current layout and the image data, and conforms to physical constraints and spatial consistency. The negative log probability loss function ensures that the generated layout is consistent with the initial input conditions, thereby optimizing the coherence of the spatial layout and ensuring that the final generated layout conforms to the spatial structure set by the user. During the training process, MiDiffusion will learn the spatial allocation and layout rules between different room types, so as to generate a layout that meets specific requirements. And the generated layout data will be saved as a serialized data file, containing layout data in a specific format, which supports subsequent 3D modeling and panoramic image generation. It is also possible to pull a model that matches the room layout from the 3D-FRONT database based on the feature data in the serialized data file (.pkl file), and then generate a rendering from a top-down perspective as shown in Figure 6 the figure.

[0086] S3. Generate panoramic images: Based on the layout data, construct a basic 3D model, then supplement the key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image;

[0087] Step S3 specifically includes the following steps:

[0088] S31. Export the basic 3D model of the room from the generated layout data, where the basic 3D model includes a floor model and a furniture model;

[0089] S32. Supplement the key structural elements based on the exported floor model, and the key structural elements include walls and ceilings. Then combine the supplemented wall, ceiling, floor model and furniture model to obtain a complete 3D room model; as shown in Figure 7 the figure, the 3D room model not only includes the floor and furniture positions of the room, but also ensures the integrity of the space, providing a basis for the rendering of panoramic images;

[0090] S33. Panoramic image rendering: As shown in Figure 8 the figure, place an internal camera in the 3D room model according to the preset perspective, and set multiple point light sources. These point light sources are used to simulate the light irradiation at different positions, enhancing the three-dimensional sense and layering of the space. Then use the rendering method based on physically based ray tracing rendering technology in the 3D software Blender, and select the equidistant cylindrical projection as the spatial structure projection type of the panoramic image. This projection method helps to accurately express the spatial layout of the room and maintain visual unity, generating a panoramic white film image with an aspect ratio of 2:1.

[0091] S4. Generate a panoramic design plan: Extract the line drawing and depth map of the panoramic white model as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image. Then divide the overlapping sub-image areas of the preliminary panoramic image and set dynamic masks for optimizing each sub-area. Use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image;

[0092] Step S4 specifically includes the following steps:

[0093] S41. Generate a preliminary panoramic image;

[0094] S411. Extract constraint images: Input the panoramic white model, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map to obtain the three-dimensional depth information of the spatial layout;

[0095] S412. Constraint-guided generation: As Figure 9 shown, input the extracted line drawing and depth map into the line constraint model in the ControlNet model and the depth constraint model in the ControlNet model respectively. At the same time, use the self-trained LoRA model to optimize the style and structure of the panoramic white model to obtain a preliminarily generated panoramic image, providing depth guidance for the spatial structure to ensure that during the generation process, the spatial layout structure is restored and any possible spatial distortion or inconsistency is avoided;

[0096] The LoRA model has the following advantages: First, it constructs a dataset for different house types, making the generated spatial layout highly match the requirements of specific house types; Second, the LoRA model adapts to the domestic indoor space design style and can meet the unique aesthetics of the local market in terms of color matching and space arrangement; Finally, when generating panoramic images, the model can optimize the spatial structure, reduce seam problems, and ensure the visual coherence and sense of space of the images.

[0097] S413. Image super-resolution processing: As Figure 10 shown, in order to improve the details and resolution of the image, use the RealESRGAN model to quadruple the size of the preliminarily generated panoramic image, retaining the spatial structure and features during generation to obtain a preliminary panoramic image, which is convenient for subsequent refined processing of each area;

[0098] S42. Divide the overlapping sub-image areas of the preliminary panoramic image: As Figure 11As shown, in order to optimize the seam transition of the image and ensure the coherence of the overall visual effect, the preliminary panoramic image is divided into three sub-image regions, namely the left sub-image region, the middle sub-image region, and the right sub-image region, which are arranged in a circular pattern. There is an overlapping region between each adjacent sub-image region. When generating the right sub-image region, the left 1 / 3 region of the left sub-image region needs to be spliced to the right 1 / 3 region on the left side of the right sub-image region to achieve the end-to-end connection of the panoramic image, ensure the consistency of the spatial structure of the image, and eliminate the seam problem.

[0099] In step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3. When generating the right sub-image region, the left 1 / 3 overlapping region of the left sub-image region is spliced to the right 1 / 3 overlapping region of the right sub-image region to ensure the consistency of the spatial structure of the image.

[0100] After the left sub-image region is generated, the left 1 / 3 overlapping region of the middle sub-image region and the right 1 / 3 overlapping region of the right sub-image region are updated. After the middle sub-image region is generated, the left 1 / 3 overlapping region of the right sub-image region is updated.

[0101] In order to avoid abrupt seams at the mask boundary, the mask edge regions of the three sub-image regions are feathered.

[0102] S43. Dynamic mask setting: The mask of the left sub-image region completely covers its entire image region and is set to 0. The mask of the middle sub-image region covers the region except the overlapping region with the left sub-image region on the left. The mask of the right sub-image region is the middle region excluding the overlapping regions with the left and middle sub-image regions on the left and right.

[0103] S44. Model-guided optimization: During the generation process, the ControlNet model and the self-trained LoRA model are introduced. The ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used for style adjustment guidance. At the same time, the image generation is guided by fixing the random seed and the style language description provided by the user to ensure that the finally generated panoramic image not only meets the spatial structure requirements but also can accurately express the user's design intention, thus achieving higher generation accuracy and style consistency.

[0104] S45. Horizontal gradient mask fusion: As Figure 12As shown, a mask with a gradually changing transparency is set in the overlapping area to achieve a completely continuous visual experience and obtain a 720° panoramic view. When the user looks around, by performing horizontal gradient mask fusion processing in the overlapping area, the user will not perceive the connection between the starting point and the ending point of the image, thus achieving a completely continuous visual experience. As a post-processing method for generating sub-region images, this gradient tiling scheme cooperates with the aforementioned dynamic mask strategy to jointly ensure the visual continuity of the final panoramic image.

[0105] Preferably, in step S412 and step S44, the guiding parameters generated by the self-trained LoRA model are set between 0.35 and 0.65, and the training steps of the self-trained LoRA model are as follows:

[0106] The first step is to collect a panoramic image dataset: collect six-directional map data of the indoor space. The six-directional map data includes the geographical location, name, type of the room, and the corresponding six-directional maps (front, back, left, right, sky, and ground). After classifying and organizing according to the room type, import the six-directional maps into the Blender 3D software, and use the method of converting the cube to an equidistant cylindrical projection to convert the six-directional maps into panoramic images, and render and export them to obtain a panoramic image dataset, providing basic data support for subsequent operations;

[0107] The second step is data annotation and preparation: to ensure the effective loading of the training data of the LoRA model, use the WD14 tool to annotate the synthesized panoramic images. The annotation data includes furniture, decoration, and spatial structure, and add unified text description labels. The above annotations and text descriptions can help the model better understand the elements and their features in the images, providing key semantic information for subsequent training and generation tasks;

[0108] The third step is model training: input the panoramic image dataset and the annotation data into the LoRA model, and train the LoRA model with the cross-entropy loss function. During this process, adjust the pre-trained weights of the LoRA model based on the annotation data, and use low-rank decomposition to complete parameter update. It can be seen that the model training process does not rely on external guiding parameters, but fully explores and optimizes the spatial layout information and design style information contained in the annotation data, and uses this as the main basis for learning.

[0109] And the formula for adjusting the pre-trained weights of the LoRA model is as follows:

[0110] ;

[0111] ;

[0112] In the formula, represents the original weight matrix of the LoRA model; Indicates obtaining the weight update matrix through low-rank decomposition; and respectively represent matrices with shapes of and ; represents the low-rank dimension; and respectively represent the rows and columns of the weight matrix; represents the output vector of the LoRA model; represents the input vector of the LoRA model;

[0113] The expression of the cross-entropy loss function is as follows:

[0114] ;

[0115] In the formula, represents the total loss value of model training; represents the total number of training samples; represents the loss value of the th sample; represents the true label of the th sample for class ; represents the probability that the th sample is predicted as class ; represents the number of classes.

[0116] S5. Output a 720° panoramic view.

[0117] An AIGC-based indoor panoramic view adaptive generation system for performing an AIGC-based indoor panoramic view adaptive generation method, including:

[0118] Data acquisition and preprocessing module: used to acquire the expected design style language description and the floor plan and perform preprocessing;

[0119] Layout construction module, used to generate layout data that meets the constraint conditions based on the preprocessed floor plan by using the pre-trained MiDiffusion model;

[0120] Panoramic image generation module, used to construct a basic 3D model based on the layout data, then supplement key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image;

[0121] The panoramic design scheme generation module is used to extract the line drawing and depth map of the panoramic white film as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image. Then, divide the overlapping sub-image regions of the preliminary panoramic image, set dynamic masks for optimizing each sub-region, and use a horizontal gradient mask to optimize and complete the output of the fused image to obtain a 720° panoramic image;

[0122] The output module is used to output a 720° panoramic image.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An adaptive generation method for indoor panoramic images based on AIGC, characterized in that: It includes the following steps: S1. Obtain a floor plan with a description of the expected design style language and perform preprocessing; S2. Construct a layout: Based on the preprocessed floor plan, use the pre-trained MiDiffusion model to generate layout data that meets the constraint conditions; S3. Generate a panoramic image: Based on the layout data, construct a basic 3D model, then supplement key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; S4. Generate a panoramic design plan: Extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image, then divide the overlapping sub-image areas of the preliminary panoramic image, and set dynamic masks for optimization of each sub-region. Use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image; S5. Output the 720° panoramic image.

2. The adaptive generation method of indoor panoramic images based on AIGC according to claim 1, wherein: In step S1, the obtained floor plan is from the floor plan in the 3D-FRONT database or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed: The first step is to construct CSV data of the room list. The CSV data contains two types of data: room number and training. Among them, the training type data is divided into a training set and a test set, and the CSV data is kept consistent with the room data format in the 3D-FRONT database; The second step is for the user to input the length and width data of the room, then use the orthographic projection method to generate a plane grid that matches the length and width data of the room, and record the projection result of the plane grid from a top-down perspective to generate a room mask image. At the same time, the room type is default set to the user's room, and then the geometric features of the plane grid are recorded to generate spatial feature data, and other features of the room are filled with random data of the same dimension to obtain all the spatial data of the custom room; The third step is to use the PointNet sampler to collect the object edge data of the custom room spatial data, and then use the pre-trained hybrid diffusion model to verify the usability of the edge information, and synchronously modify the configuration file used for layout generation to a custom configuration file.

3. The adaptive generation method of indoor panoramic images based on AIGC according to claim 2, characterized in that: The MiDiffusion model described in step S2 is a hybrid discrete-continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network. Among them, the input layer uses a PointNet feature extractor to convert the spatial distribution of the room into an encoded format; in the denoising network, the VQ-Diffusion architecture is modified to a Transformer decoder module, and at the same time, a multi-head attention mechanism is introduced to introduce constraint conditions during step-by-step denoising; And the MiDiffusion model is trained and optimized through the KL divergence loss function and the negative log probability loss function.

4. The indoor panoramic view adaptive generation method based on AIGC according to claim 3, wherein: Step S3 specifically includes the following steps: S31. Export the basic 3D model of the room from the generated layout data, where the basic 3D model includes a floor model and a furniture model; S32. Supplement key structural elements based on the exported floor model. The key structural elements include walls and ceilings. Then, combine the supplemented wall, ceiling, floor model, and furniture model to obtain a complete three-dimensional room model. S33. Panoramic image rendering: Place an internal camera in the three-dimensional room model according to a preset perspective, set multiple point light sources, and then use the rendering method of the physically based ray tracing rendering technology in the three-dimensional software Blender. Select the equidistant cylindrical projection as the spatial structure projection type of the panoramic view to generate a panoramic white film image with an aspect ratio of 2:

1.

5. The method for adaptively generating an indoor panoramic view based on AIGC according to claim 4, wherein: Step S4 specifically includes the following steps: S41. Generate a preliminary panoramic image. S411. Extract constraint images: Input the panoramic white film image, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map to obtain the three-dimensional depth information of the spatial layout. S412. Constraint-guided generation: Input the extracted line drawing and depth map into the line constraint model in the ControlNet model and the depth constraint model in the ControlNet model respectively. At the same time, use the self-trained LoRA model to optimize the style and structure of the panoramic white film image to obtain a preliminarily generated panoramic image. S413. Image super-resolution processing: Use the RealESRGAN model to quadruple the size of the preliminarily generated panoramic image, retaining the spatial structure and features during generation to obtain a preliminary panoramic image. S42. Divide the overlapping sub-image regions of the preliminary panoramic image: Divide the preliminary panoramic image into a left sub-image region, a middle sub-image region, and a right sub-image region arranged in a surrounding manner, a total of three sub-image regions, and there is an overlapping region between each adjacent sub-image region. S43. Dynamic mask setting: The mask of the left sub-image region completely covers its entire image region and is set to 0; the mask of the middle sub-image region covers the region except the overlapping region with the left sub-image region; the mask of the right sub-image region is the middle region excluding the overlapping regions of the left and middle sub-image regions and the overlapping region of the right and left sub-image regions. S44. Model-guided optimization: During the generation process, introduce the ControlNet model and the self-trained LoRA model. The ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used for style adjustment guidance. At the same time, guide the image generation by fixing the random seed and the style language description provided by the user. S45. Horizontal gradient mask fusion: Set a mask with a gradually changing transparency in the overlapping region to achieve a completely continuous visual experience and obtain a 720° panoramic image.

6. The indoor panoramic view adaptive generation method based on AIGC according to claim 5, characterized in that: In step S412 and step S44, set the guidance parameters generated by the self-trained LoRA model to be between 0.35 and 0.65, and the training steps of the self-trained LoRA model are as follows: Step 1: Collect the panoramic image dataset: Collect six-directional map data of the indoor space. The six-directional map data includes the geographical location, name, type of the room, and the corresponding six-directional map. After classifying and organizing according to the room type, import the six-directional map into the Blender 3D software. Use the method of converting a cube to an equidistant cylindrical projection to convert the six-directional map into a panoramic image, and render and export it to obtain the panoramic image dataset; Step 2: Data annotation and preparation: Use the WD14 tool to annotate the synthesized panoramic image. The annotation data includes furniture, decoration, and spatial structure, and add unified text expression labels; Step 3: Model training: Input the panoramic image dataset and the annotation data into the LoRA model, and train the LoRA model with the cross-entropy loss function. During this process, adjust the pre-trained weights of the LoRA model based on the annotation data, and complete the parameter update using low-rank decomposition.

7. The adaptive generation method of indoor panoramic images based on AIGC according to claim 6, characterized in that: The formula for adjusting the pre-trained weights of the LoRA model is as follows: ; ; Wherein, represents the original weight matrix of the LoRA model; represents the weight update matrix obtained by low-rank factorization; and respectively represent matrices with shapes of and ; represents the low-rank dimension; and respectively represent the rows and columns of the weight matrix; represents the output vector of the LoRA model; represents the input vector of the LoRA model; The expression of the cross-entropy loss function is as follows: ; In the formula, Represents the total loss value of model training; Represents the total number of training samples; Indicates The loss value of samples; Indicates Sample pairs The true label of Indicates samples are predicted to be of class probability; Indicates the number of categories.

8. An indoor panoramic view adaptive generation method based on AIGC according to claim 7, characterized in that: In step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3. When generating the right sub-image region, splice the left 1 / 3 overlapping region of the left sub-image region to the right 1 / 3 overlapping region of the right sub-image region to ensure the spatial structure consistency of the image; After the left sub-image region is generated, update the left 1 / 3 overlapping region of the middle sub-image region and the right 1 / 3 overlapping region of the right sub-image region; after the middle sub-image region is generated, update the left 1 / 3 overlapping region of the right sub-image region; And the mask edge regions of the three sub-image regions are all feathered.

9. An indoor panoramic view adaptive generation system based on AIGC, which is used to execute an indoor panoramic view adaptive generation method based on AIGC described in claim 8, characterized in that: Including: Data acquisition and preprocessing module: Used to acquire the expected design style language description and the floor plan and perform preprocessing; Layout construction module, used to generate layout data that meets the constraint conditions based on the preprocessed floor plan using the pre-trained MiDiffusion model; Panoramic image generation module, used to construct a basic 3D model based on the layout data, then supplement the key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; Panoramic design solution generation module, used to extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image, then divide the overlapping sub-image regions of the preliminary panoramic image, and set dynamic masks to optimize each sub-region. Use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image; Output module, used to output the 720° panoramic image.

Citation Information

Patent Citations

  • Real scene three-dimensional visualization method and system based on semantic point cloud

    CN116912437A

  • Building intelligent design system based on AIGC

    CN117632098A