Indoor panorama adaptive generation method and system based on AIGC

By applying the AIGC-based panoramic image adaptive generation method in indoor space design, a high-quality 720° panoramic image is generated using MiDiffusion, ControlNet and LoRA models, solving the problems of high computational complexity, rendering time and style transfer in the existing technology, and achieving efficient and accurate panoramic image generation and style adaptation.

CN120070776AActive Publication Date: 2025-05-30COMMUNICATION UNIVERSITY OF CHINA

Patent Information

Application Number
CN202510543799.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art has problems such as high computational complexity, long rendering, difficult style transfer and imperfect image seam processing in indoor space design. In addition, the AI ​​model lacks understanding of three-dimensional spatial structure, which makes it difficult to meet the market's demand for creative diversity and rapid iteration of design.

Method used

Adaptive generation method of indoor panoramic images based on AIGC is adopted to generate layout data through the MiDiffusion model, and a 720° panoramic image is generated by combining ControlNet and self-trained LoRA model. Dynamic masks and horizontal gradient masks are used to optimize seam fusion to meet specific style needs.

Benefits of technology

It improves the generation accuracy and style adaptability of panoramic images, reduces the computational complexity and rendering time, realizes the visual coherence and sense of space of the image, and meets the market's demand for creative diversity and rapid iteration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070776A_ABST
    Figure CN120070776A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor panorama self-adaptive generation method and system based on AIGC, and belongs to the field of indoor space design, and the method comprises the following steps: S1, obtaining a plane planning graph with expected design style language description, and carrying out the preprocessing; s2, constructing a layout: generating layout data meeting constraint conditions by adopting a pre-trained MiDiffusion model; s3, generating a panoramic image: constructing a basic three-dimensional model based on the layout data, then supplementing key structural elements, constructing a complete three-dimensional room model, and then performing panoramic image rendering to generate a panoramic white film graph; s4, generating a panoramic design scheme; and S5, outputting a 720-degree panorama. By adopting the AIGC-based indoor panorama adaptive generation method and system, the design efficiency can be remarkably improved, the requirements of rapid iteration and personalized customization can be met, and the intelligent and digital development in the field of indoor space design can be promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of indoor space design, and particularly to a method and system for adaptively generating indoor panoramic images based on AIGC. Background Art

[0002] With the rapid iterative development of virtual reality (VR) and augmented reality (AR) technologies, the indoor space design field is undergoing a transformation from traditional two-dimensional plane displays to immersive three-dimensional visualization solutions. The modern consumer market has put forward higher requirements for indoor space design, not only requiring to meet basic functional needs, but also emphasizing the rationality of space layout, the diversity of design styles, and the refinement of visual presentation effects.

[0003] In the context of the rapid development of information technology and multimedia technology, 720° panoramic images have gradually become the industry standard display solution due to their advantage of all-round three-dimensional display. 720° panoramic images have significant advantages compared with traditional static floor plans or local perspective display methods. This innovative display method can not only completely present spatial structure information, helping users intuitively understand the overall layout, but also support multi-angle observation. Moreover, when combined with VR glasses, it can achieve a truly immersive space experience. It can be seen that panoramic images can accurately display design details and decorative styles, providing strong support for the rapid evaluation and optimization of design schemes.

[0004] However, the current methods usually require designers to use CAD drawing, 3D modeling software (such as Maya, C4D, Blender, etc.) and rendering tools in sequence to obtain local and even indoor space panoramic effect drawings. The entire process is cumbersome and time-consuming. It can be seen that the existing methods not only require a large number of professional designers to participate, consuming a huge amount of human resources, but also have a long design cycle, making it difficult to meet the market's requirements for creative diversity and rapid design iteration. Especially when carrying out personalized customization, the efficiency of scheme optimization is low, and it is difficult to respond in a timely manner to changes in customer needs.

[0005] To solve the above problems, the prior art further applies artificial intelligence generated content (AIGC) technology to the field of indoor space design and has made breakthrough progress. Among them, the technical solutions based on generative adversarial networks (GANs) and diffusion models can significantly improve design efficiency and reduce production costs.

[0006] However, the existing AIGC technology still faces many challenges in automatic layout generation. Although the prior art has proposed new solutions through 3D model generation and material parameter optimization, there are still problems such as high computational complexity and long rendering time in practical applications. At the same time, in the face of the need for style transfer, the precise control of layout constraint conditions is also a technical difficulty, making it difficult to achieve rapid response.

[0007] In the field of panoramic image generation, there are still several key problems in the existing technologies that need to be solved urgently. Firstly, the image seam processing technology is imperfect, and visual breaks are likely to occur. Secondly, the AI models have insufficient understanding of the three-dimensional space structure, especially with obvious defects in dealing with panoramic lens distortion. Finally, it is difficult to control details during the generation process, which directly affects the quality performance of the final image. The above technical difficulties severely restrict the application effect of panoramic image generation technology.

[0008] In the actual use process, there are also problems of insufficient adaptability in the existing AI generation methods. The current models are mainly trained based on international general data and are difficult to accurately grasp the characteristics and requirements of domestic spatial designs such as the living environment. For example, there are obvious differences in indoor space design between domestic and foreign countries in terms of furniture configuration, decoration styles, etc. due to differences in living habits, usage scenarios, and aesthetics. This makes it difficult for the design schemes directly generated by existing AI models to fully meet the actual needs of domestic users and affects the practicality of the technical solutions.

[0009] In summary, the above problems severely restrict the wide application of AI technology in the field of indoor space design. Especially in aspects such as panoramic image generation, spatial structure understanding, and localization adaptation, there is still a large room for improvement in the existing technical solutions. Summary of the Invention

[0010] The purpose of the present invention is to provide an AIGC-based adaptive indoor panoramic map generation method and system to solve the above technical problems.

[0011] To achieve the above purpose, the present invention provides an AIGC-based adaptive indoor panoramic map generation method, including the following steps: S1. Obtain a floor plan with a language description of the expected design style and perform preprocessing; S2. Construct a layout: Based on the preprocessed floor plan, use a pre-trained MiDiffusion model to generate layout data that meets the constraint conditions; S3. Generate a panoramic image: Based on the layout data, construct a basic three-dimensional model, then supplement key structural elements to construct a complete three-dimensional room model, and then perform panoramic image rendering to generate a panoramic white film image; S4. Generate a panoramic design scheme: Extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic map. Then divide the overlapping sub-image areas of the preliminary panoramic map and set dynamic masks for optimization of each sub-region. Use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic map; S5. Output a 720° panoramic map.

[0012] Preferably, in step S1, the obtained floor plan is from the floor plan in the 3D-FRONT database or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed: First step, construct the CSV data of the room list. The CSV data contains two types of data: room number and training. The training type data is divided into a training set and a test set, and the CSV data is kept consistent with the room data format in the 3D-FRONT database; Second step, the user inputs the length and width data of the room, then uses the forward projection method to generate a plane grid matching the room length and width data, and records the projection result of the plane grid from a top-down perspective to generate a room mask image. At the same time, the room type is default set to the user's room, and then the geometric features of the plane grid are recorded to generate spatial feature data, and the other features of the room are filled with random data of the same dimension to obtain all the spatial data of the custom room; Third step, use the PointNet sampler to collect the object edge data of the custom room spatial data, and then use the pre-trained hybrid diffusion model to verify the usability of the edge feature information, and synchronously modify the configuration file used for layout generation to a custom configuration file.

[0013] Preferably, the MiDiffusion model described in step S2 is a hybrid discrete-continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network. The input layer uses a PointNet feature extractor to convert the spatial distribution of the room into an encoded format; in the denoising network, the VQ-Diffusion architecture is modified to a Transformer decoder module, and at the same time, a multi-head attention mechanism is introduced to introduce constraint conditions during step-by-step denoising; And the MiDiffusion model is trained and optimized through the KL divergence loss function and the negative log probability loss function.

[0014] Preferably, step S3 specifically includes the following steps: S31. Export the basic 3D model of the room from the generated layout data, where the basic 3D model includes a floor model and a furniture model; S32. Supplement the key structural elements based on the exported floor model, and the key structural elements include walls and ceilings. Then, combine the supplemented wall, ceiling, floor model, and furniture model to obtain a complete 3D room model; S33. Panoramic image rendering: Place an internal camera in the 3D room model according to a preset perspective, set multiple point light sources, and then use the rendering method of the physically based ray tracing rendering technology in the 3D software Blender. Select the equidistant cylindrical projection as the spatial structure projection type of the panoramic image to generate a panoramic white film image with an aspect ratio of 2:1.

[0015] Preferably, step S4 specifically includes the following steps: S41. Generate a preliminary panoramic image; S411. Extract the constraint image: Input the panoramic white film image, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map to obtain the three-dimensional depth information of the spatial layout; S412. Constraint-guided generation: Input the extracted line drawing and depth map into the line constraint model in the ControlNet model and the depth constraint model in the ControlNet model respectively. At the same time, use the self-trained LoRA model to optimize the style and structure of the panoramic white film image to obtain a preliminarily generated panoramic image; S413. Image super-resolution processing: Use the RealESRGAN model to quadruple the size of the preliminarily generated panoramic image, retaining the spatial structure and features during generation to obtain a preliminary panoramic image; S42. Divide the overlapping sub-image regions of the preliminary panoramic image: Divide the preliminary panoramic image into a left sub-image region, a middle sub-image region, and a right sub-image region arranged in a surrounding manner, a total of three sub-image regions, and there is an overlapping region between each adjacent sub-image region; S43. Dynamic mask setting: The mask of the left sub-image region completely covers its entire image region and is set to 0; the mask of the middle sub-image region covers the region except the overlapping region with the left sub-image region on the left; the mask of the right sub-image region is the middle region excluding the overlapping regions of the left and the middle sub-image regions, and the overlapping regions of the right and the left sub-image regions; S44. Model-guided optimization: During the generation process, introduce the ControlNet model and the self-trained LoRA model. The ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used to perform style adjustment guidance. At the same time, guide the image generation by fixing the random seed and the style language description provided by the user; S45. Horizontal gradient mask fusion: Set a mask with a gradually changing transparency in the overlapping region to achieve a completely continuous visual experience and obtain a 720° panoramic image.

[0016] Preferably, in step S412 and step S44, the guiding parameters generated by the self-trained LoRA model are set to be between 0.35 and 0.65, and the training steps of the self-trained LoRA model are as follows: First step, collect the panoramic image dataset: collect six-direction map data of the indoor space. The six-direction map data includes the geographical location, name, type of the room, and the corresponding six-direction maps (front, back, left, right, sky, and ground). After classifying and sorting by room type, import the six-direction maps into the Blender 3D software, and use the method of converting the cube to an equidistant cylindrical projection to convert the six-direction maps into panoramic images, and render and export them to obtain the panoramic image dataset; Second step, data annotation and preparation: use the WD14 tool to annotate the synthesized panoramic images. The annotation data includes furniture, decoration, and spatial structure, and add unified text expression labels; Third step, model training: input the panoramic image dataset and the annotation data into the LoRA model, and train the LoRA model with the cross-entropy loss function. During this process, adjust the pre-trained weights of the LoRA model based on the annotation data, and complete the parameter update by using low-rank decomposition.

[0017] Preferably, the formula for adjusting the pre-trained weights of the LoRA model is as follows: ; ; In the formula, represents the original weight matrix of the LoRA model; represents the weight update matrix obtained by low-rank decomposition; and respectively represent matrices with shapes of and ; represents the low-rank dimension; and respectively represent the rows and columns of the weight matrix; represents the output vector of the LoRA model; represents the input vector of the LoRA model; The expression of the cross-entropy loss function is as follows: ; In the formula represents the total loss value of model training; represents the total number of training samples; represents the loss value of the th sample; represents the th sample's true label for class ; represents the The probability that a sample is predicted as a class ; represents the number of classes.

[0018] Preferably, in step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3. When generating the right sub-image region, the left 1 / 3 overlapping region of the left sub-image region is spliced to the right 1 / 3 overlapping region of the right sub-image region to ensure the spatial structure consistency of the image; After the left sub-image region is generated, update the left 1 / 3 overlapping region of the middle sub-image region and the right 1 / 3 overlapping region of the right sub-image region; after the middle sub-image region is generated, update the left 1 / 3 overlapping region of the right sub-image region; And the mask edge regions of the three sub-image regions are all feathered.

[0019] An AIGC-based indoor panoramic image adaptive generation system for implementing an AIGC-based indoor panoramic image adaptive generation method, including: Data acquisition and preprocessing module: used to acquire the expected design style language description and the floor plan and perform preprocessing; Layout construction module, used to generate layout data that meets the constraint conditions based on the preprocessed floor plan by using the pre-trained MiDiffusion model; Panoramic image generation module, used to construct a basic 3D model based on the layout data, then supplement the key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; Panoramic design scheme generation module, used to extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image, then divide the overlapping sub-image regions of the preliminary panoramic image, and set dynamic masks to optimize each sub-region, and use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image; Output module, used to output a 720° panoramic image.

[0020] Therefore, the present invention adopts the above-mentioned AIGC-based indoor panoramic image adaptive generation method and system, and the beneficial effects are: 1. High generation accuracy: Extract the line drawing and depth map of the panoramic white film image, and use ControlNet for constraint guidance, which can accurately control the spatial structure and contour of the image, making the generated panoramic image more in line with the actual situation in terms of geometric structure and spatial layout, reducing distortion and unreasonable layout; 2. Strong style adaptability: By using the self-trained LoRA model and training based on a dataset constructed according to domestic indoor space design styles, it can specifically optimize the panorama style to meet specific style requirements, making the image style more in line with the actual application scenarios and user preferences. 3. Seam processing: Divide the overlapping sub-image areas and use dynamic masks and horizontal gradient masks to optimize seam fusion, which can effectively eliminate splicing traces, achieve smooth transitions between images, ensure the visual coherence and integrity of the panorama, and bring a high-quality visual experience. 4. Efficient and flexible training: The LoRA model training uses low-rank decomposition to reduce training parameters, improve training efficiency, and does not change the weights of the pre-trained model. It is flexible in fine-tuning, can quickly adapt to different task and style requirements, and saves computing resources and time costs.

[0021] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0022] Figure 1 It is a flowchart of an AIGC-based indoor panorama adaptive generation method according to the present invention. Figure 2 It is an example flowchart of an AIGC-based indoor panorama adaptive generation method according to the present invention. Figure 3 It is a layout generation framework diagram of an AIGC-based indoor panorama adaptive generation method according to the present invention. Figure 4 It is an example mask diagram of an AIGC-based indoor panorama adaptive generation method according to the present invention, where (a) is a mask diagram of a certain room in the 3D-FRONT database, and (b) is a user-defined room mask diagram. Figure 5 It is a diagram of the training process of the MiDiffusion model of an AIGC-based indoor panorama adaptive generation method according to the present invention, where (a) is a curve diagram of the change in the learning rate, and (b) is a curve diagram of the change in the loss function. Figure 6 It is an example rendering preview diagram of an AIGC-based indoor panorama adaptive generation method according to the present invention. Figure 7 It is a diagram of the construction of a 3D room model of an AIGC-based indoor panorama adaptive generation method according to the present invention, where (a) is an example diagram of the floor model, (b) is an example diagram of the furniture model, and (c) is an example diagram of the complete 3D room model. Figure 8 It is an example diagram of a panoramic white film diagram of an AIGC-based indoor panorama adaptive generation method according to the present invention. Figure 9 This is the feature extraction diagram of the panoramic white film for a method of adaptively generating indoor panoramic images based on AIGC according to the present invention, where (a) is the extracted line drawing and (b) is the extracted depth map; Figure 10 This is the preliminary panoramic example diagram of a method of adaptively generating indoor panoramic images based on AIGC according to the present invention; Figure 11 This is the result diagram of the preliminary panoramic map division of a method of adaptively generating indoor panoramic images based on AIGC according to the present invention, where (a) is the example diagram of the left sub-image area, (b) is the example diagram of the middle sub-image area, and (c) is the example diagram of the right sub-image area; Figure 12 This is the comparison diagram of the surround stitching of a method of adaptively generating indoor panoramic images based on AIGC according to the present invention, where (a) is the example diagram of the surround stitching before optimization and (b) is the example diagram of the surround stitching after optimization by horizontal gradient mask fusion. Detailed implementation manners

[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the following further describes the embodiments of the present invention in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of this application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end.

[0024] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0025] The following describes the embodiments of the present invention in detail with reference to the drawings.

[0026] As Figures 1 - 3 shown, a method of adaptively generating indoor panoramic images based on AIGC includes the following steps: S1. Obtain a floor plan with a description of the expected design style language and perform preprocessing; In step S1, the obtained floor plan is from the 3D-FRONT database (3D-FRONT stands for 3DFurnished Rooms with layOuts and semaNTics, which is a large 3D scene dataset jointly open-sourced by Alibaba, Simon Fraser University, and the Institute of Computing, Chinese Academy of Sciences. This dataset contains 6,813 different real house types, consisting of 51,708 rooms with a rich variety of room types, which can be subdivided into 28 types; among them, 19,775 rooms contain manually verified interior design information, providing high-quality samples for research and applications related to interior design), or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed: First, construct the CSV data of the room list. The CSV data contains two types of data: room number and training. The training type data is divided into a training set and a test set, and the CSV data is kept consistent with the room data format of the 3D-FRONT database; Second, as Figure 4 shown, for the room length and width data input by the user, use the forward projection method to generate a plane grid matching the room length and width data, and record the projection result of the plane grid from a top-down perspective to generate a room mask image. At the same time, default the room type to the user's room, then input the geometric features of the plane grid to generate spatial feature data, such as point data, face data, etc., and use random data of the same dimension to supplement other features of the room to obtain user-defined room spatial data, ensuring that the file format is the same as other spatial data in the dataset; Third, use the PointNet sampler to collect the object edge data of the user-defined room spatial data, and then use the pre-trained hybrid diffusion model to verify the usability of the edge feature information, and synchronously modify the configuration file used for layout generation to a user-defined configuration file (such as mask.png and boxes.npz). These data will be used as conditional inputs for subsequent generation to ensure that the generated layout meets the spatial requirements of the deep learning model.

[0027] S2. Construct the layout: Based on the preprocessed floor plan, use the pre-trained MiDiffusion model to generate layout data that meets the constraint conditions; The MiDiffusion model described in step S2 is a hybrid discrete - continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network. The input layer uses a PointNet feature extractor to convert the spatial distribution of the room into an encoded format, providing the basic data representation for the subsequent learning and generation tasks of the model, enabling the model to effectively process and understand spatial information. In the denoising network, the VQ - Diffusion architecture is modified into a Transformer decoder module, and at the same time, a multi - head attention mechanism is introduced to introduce constraint conditions during step - by - step denoising. The introduction of the multi - head attention mechanism enables the model to better consider various spatial relationships and semantic information during denoising, thereby improving the accuracy of denoising and the quality of the generated results.

[0028] As Figure 5 shown, in order to ensure that the model can generate a better spatial layout according to the floor plan, the position limit of MiDiffusion is turned on for retraining. The MiDiffusion model is trained and optimized through the KL - divergence loss function and the negative log - probability loss function. The KL - divergence loss function ensures that the model can accurately infer the next layout structure from the latent space of the current layout and image data, and conforms to physical constraints and spatial consistency. The negative log - probability loss function ensures that the generated layout is consistent with the initial input conditions, thereby optimizing the coherence of the spatial layout and ensuring that the final generated layout conforms to the spatial structure set by the user. During the training process, MiDiffusion will learn the spatial allocation and layout rules between different room types to generate a layout that meets specific requirements. And the generated layout data will be saved as a serialized data file, containing layout data in a specific format, supporting subsequent 3D modeling and panoramic image generation. It is also possible to pull a model that matches the room layout from the 3D - FRONT database according to the feature data in the serialized data file (.pkl file), and then generate a top - down view rendering as Figure 6 shown.

[0029] S3. Generate panoramic images: Based on the layout data, construct a basic 3D model, then supplement key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white - film image; Step S3 specifically includes the following steps: S31. Export the basic 3D model of the room from the generated layout data, where the basic 3D model includes a floor model and a furniture model; S32. Based on the exported floor model, supplement key structural elements, and the key structural elements include walls and ceilings. Then combine the supplemented wall, ceiling, floor models and furniture models to obtain a complete 3D room model; As Figure 7As shown in the figure, the three-dimensional room model not only includes the floor and furniture positions of the room, but also ensures the integrity of the space, providing a basis for the rendering of panoramic images; S33. Panoramic image rendering: As Figure 8 shown in the figure, place an internal camera in the three-dimensional room model according to the preset perspective, and set multiple point light sources. These point light sources are used to simulate the light irradiation at different positions, enhancing the three-dimensional sense and layering of the space. Then, use the rendering method of the physically based ray tracing rendering technology in the three-dimensional software Blender, and select the equidistant cylindrical projection as the spatial structure projection type of the panoramic image. This projection method helps to accurately express the spatial layout of the room and maintain visual unity, generating a panoramic white film image with an aspect ratio of 2:1.

[0030] S4. Generate a panoramic design plan: Extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image. Then, divide the overlapping sub-image areas of the preliminary panoramic image, and set dynamic masks for optimizing each sub-area. Use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image; Step S4 specifically includes the following steps: S41. Generate a preliminary panoramic image; S411. Extract constraint images: Input the panoramic white film image, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map to obtain the three-dimensional depth information of the spatial layout; S412. Constraint-guided generation: As Figure 9 shown in the figure, input the extracted line drawing and depth map into the line constraint model in the ControlNet model and the depth constraint model in the ControlNet model respectively. At the same time, use the self-trained LoRA model to optimize the style and structure of the panoramic white film image to obtain a preliminarily generated panoramic image, providing depth guidance for the spatial structure, ensuring that during the generation process, the spatial layout structure is restored, and avoiding any possible spatial distortion or inconsistency; The LoRA model has the following advantages: First, it constructs a dataset for different house types, making the generated spatial layout highly match the requirements of specific house types; Second, the LoRA model adapts to the domestic interior space design style and can meet the unique aesthetics of the local market in terms of color matching and space arrangement; Finally, when generating panoramic images, the model can optimize the spatial structure, reduce seam problems, and ensure the visual coherence and sense of space of the image.

[0031] S413. Image super-resolution processing: As Figure 10As shown, to enhance the details and resolution of the image, the RealESRGAN model is used to quadruple the magnification of the initially generated panoramic image, preserving the spatial structure and features during generation to obtain a preliminary panoramic image, facilitating subsequent refined processing in partitions. S42. Divide the overlapping sub-image regions of the preliminary panoramic image: As Figure 11 shown, to optimize the seam transition of the image and ensure the coherence of the overall visual effect, the preliminary panoramic image is divided into three sub-image regions: the left sub-image region, the middle sub-image region, and the right sub-image region, which are arranged in a circular pattern. There is an overlapping region between each adjacent sub-image region. When generating the right sub-image region, the left 1 / 3 region of the left sub-image region needs to be spliced to the right 1 / 3 region of the right sub-image region to achieve the end-to-end connection of the panoramic image, ensuring the consistency of the spatial structure of the image and eliminating the seam problem.

[0032] In step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3. When generating the right sub-image region, the left 1 / 3 overlapping region of the left sub-image region is spliced to the right 1 / 3 overlapping region of the right sub-image region to ensure the consistency of the spatial structure of the image. After the left sub-image region is generated, the left 1 / 3 overlapping region of the middle sub-image region and the right 1 / 3 overlapping region of the right sub-image region are updated. After the middle sub-image region is generated, the left 1 / 3 overlapping region of the right sub-image region is updated. Moreover, to avoid abrupt seams at the mask boundaries, the mask edge regions of the three sub-image regions are all feathered.

[0033] S43. Dynamic mask setting: The mask of the left sub-image region completely covers its entire image region and is set to 0. The mask of the middle sub-image region covers the region except for the overlapping region with the left sub-image region on the left. The mask of the right sub-image region is the middle region excluding the overlapping regions with the left and middle sub-image regions on the left and right. S44. Model-guided optimization: During the generation process, the ControlNet model and the self-trained LoRA model are introduced. The ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used to guide style adjustment. At the same time, the image generation is guided by fixing the random seed and the style language description provided by the user, ensuring that the finally generated panoramic image not only meets the spatial structure requirements but also can accurately express the user's design intention, thus achieving higher generation accuracy and style consistency. S45. Horizontal gradient mask fusion: As Figure 12As shown, a mask with a gradually changing transparency is set in the overlapping area to achieve a completely continuous visual experience, resulting in a 720° panoramic image. When the user looks around, by performing horizontal gradient mask fusion processing in the overlapping area, the user will not perceive the connection between the starting point and the ending point of the image, thus achieving a completely continuous visual experience. As a post-processing method for generating sub-region images, this gradient tiling scheme cooperates with the aforementioned dynamic masking strategy to jointly ensure the visual continuity of the final panoramic image.

[0034] Preferably, in steps S412 and S44, the guiding parameters generated by the self-trained LoRA model are set to be between 0.35 and 0.65, and the training steps of the self-trained LoRA model are as follows: The first step is to collect a panoramic image dataset: collect six-directional map data of the indoor space. The six-directional map data includes the geographical location, name, type of the room, and the corresponding six-directional maps (front, back, left, right, sky, and ground). After classifying and organizing according to the room type, import the six-directional maps into the Blender 3D software, and use the method of converting the cube to equidistant cylindrical projection to convert the six-directional maps into panoramic images, and render and export them to obtain a panoramic image dataset, providing basic data support for subsequent operations; The second step is data annotation and preparation: to ensure the effective loading of the training data of the LoRA model, use the WD14 tool to annotate the synthesized panoramic images. The annotation data includes furniture, decoration, and spatial structure, and add unified text description labels. The above annotations and text descriptions can help the model better understand the elements and their features in the images, providing key semantic information for subsequent training and generation tasks; The third step is model training: input the panoramic image dataset and the annotation data into the LoRA model, and train the LoRA model with the cross-entropy loss function. During this process, adjust the pre-trained weights of the LoRA model based on the annotation data, and complete the parameter update by using low-rank decomposition. It can be seen that the model training process does not rely on external guiding parameters, but fully explores and optimizes the spatial layout information and design style information contained in the annotation data, and uses this as the main basis for learning.

[0035] And the formula for adjusting the pre-trained weights of the LoRA model is as follows: ; ; In the formula, represents the original weight matrix of the LoRA model; represents the weight update matrix obtained by low-rank decomposition; and respectively represent matrices with shapes of and ; represents the low-rank dimension; and represent the rows and columns of the weight matrix, respectively; represents the output vector of the LoRA model; represents the input vector of the LoRA model; The expression of the cross-entropy loss function is as follows: ; In the formula, represents the total loss value of model training; represents the total number of training samples; represents the th sample's loss value; represents the th sample's true label for class ; represents the probability that the th sample is predicted to be class ; represents the number of classes.

[0036] S5. Output a 720° panoramic image.

[0037] An AIGC-based indoor panoramic image adaptive generation system for performing an AIGC-based indoor panoramic image adaptive generation method, including: Data acquisition and preprocessing module: used to acquire the expected design style language description and the floor plan and perform preprocessing; Layout construction module, used to generate layout data that meets the constraint conditions based on the preprocessed floor plan by using a pre-trained MiDiffusion model; Panoramic image generation module, used to construct a basic 3D model based on the layout data, then supplement key structural elements to construct a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; Panoramic design scheme generation module, used to extract the line drawing and depth map of the panoramic white film image as constraint guiding conditions, and based on the extracted constraint guiding conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panoramic image, then divide the overlapping sub-image areas of the preliminary panoramic image, and set dynamic masks for optimization of each sub-region, and use a horizontal gradient mask to optimize and complete the fusion image output to obtain a 720° panoramic image; Output module, used to output a 720° panoramic image.

[0038] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An AIGC-based indoor panoramic image adaptive generation method, characterized by: The following steps are involved: S1. Obtaining a floor plan with a language description of the expected design style and performing preprocessing; S2. Build layout: Based on the preprocessed floor plan, the pre-trained MiDiffusion model is used to generate layout data that meets the constraints; S3, generate panoramic image: build a basic 3D model based on the layout data, add key structural elements, build a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; S4. Generate panoramic design plan: Extract the line draft and depth map of the panoramic white film image as constraint guidance conditions, and based on the extracted constraint guidance conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panorama, then divide the overlapping sub-image areas of the preliminary panorama, set a dynamic mask to optimize each sub-area, and use the horizontal gradient mask optimization to complete the fused image output to obtain a 720° panorama; S5. Output a 720° panoramic image.

2. The method for adaptively generating indoor panoramic images based on AIGC according to claim 1, characterized in that: In step S1, the obtained floor plan is from a floor plan in the 3D-FRONT database or a user-defined floor plan; and if it is a user-defined floor plan, the following preprocessing is performed: The first step is to construct the CSV data of the room list. The CSV data contains two types of data: room number and training data. The training type data is divided into training set and test set. The CSV data is consistent with the room data format of the 3D-FRONT database. Step 2: The user inputs the room length and width data, and then uses the forward projection method to generate a plane grid that matches the room length and width data. The projection result of the plane grid is recorded from a bird's-eye view angle to generate a room mask image. At the same time, the room type is set to the user's room by default. Then, the geometric features of the plane grid are entered to generate spatial feature data, and random data of the same dimension is used to fill in other features of the room to obtain all the spatial data of the custom room. The third step is to use the PointNet sampler to collect the object edge data of the custom room space data, and then use the pre-trained hybrid diffusion model to verify the usability of the edge information, and simultaneously modify the configuration file used for layout generation to the custom configuration file.

3. The method for adaptively generating indoor panoramic images based on AIGC according to claim 2, characterized in that: The MiDiffusion model described in step S2 is a hybrid discrete-continuous diffusion model. The MiDiffusion model includes an input layer and a denoising network, wherein the input layer uses a PointNet feature extractor to convert the spatial distribution of the room into a coding format; in the denoising network, the VQ-Diffusion architecture is modified to a Transformer decoder module, and a multi-head attention mechanism is introduced to introduce constraints during stepwise denoising; The MiDiffusion model is trained and optimized through the KL divergence loss function and the negative logarithmic probability loss function.

4. The method for adaptively generating indoor panoramic images based on AIGC according to claim 3, characterized in that: Step S3 specifically includes the following steps: S31, deriving a basic three-dimensional model of the room from the generated layout data, wherein the basic three-dimensional model includes a floor model and a furniture model; S32, supplementing key structural elements based on the exported floor model, where the key structural elements include walls and ceilings, and then combining the supplemented wall, ceiling, floor models and furniture models to obtain a complete three-dimensional room model; S33. Panoramic image rendering: Place a built-in camera in the 3D room model according to the preset viewing angle, set up multiple light sources, and then use the physics-based ray tracing rendering technology in the 3D software Blender to render. Select equidistant cylindrical projection as the spatial structure projection type of the panorama to generate a panoramic white film image with an aspect ratio of 2:

1.

5. The method for adaptively generating indoor panoramic images based on AIGC according to claim 4, characterized in that: Step S4 specifically includes the following steps: S41, generating a preliminary panoramic image; S411, extract constraint image: input the panoramic white film image, use the real line tool to extract the line drawing, obtain the structural information of the spatial layout, and use the Depth Anything Vit model to extract the depth map, and obtain the three-dimensional depth information of the spatial layout; S412, constraint-guided generation: the extracted line drawing and depth map are respectively input into the middle line constraint model in the ControlNet model and the middle depth constraint model in the ControlNet model, and the self-trained LoRA model is used to optimize the style and structure of the panoramic white film image to obtain a preliminarily generated panoramic image; S413, image super-resolution processing: the RealESRGAN model is used to enlarge the initially generated panoramic image by four times, retaining the spatial structure and features at the time of generation, and obtaining a preliminary panoramic image; S42, dividing the overlapping sub-image areas of the preliminary panoramic image: dividing the preliminary panoramic image into a left sub-image area, a middle sub-image area, and a right sub-image area arranged in a surrounding manner, a total of three sub-image areas, and leaving an overlapping area between each adjacent sub-image area; S43, dynamic mask setting: the mask of the left sub-image area completely covers its entire image area and is set to 0; the mask of the middle sub-image area covers the area except the overlapping area between the left and left sub-image areas; the mask of the right sub-image area is the area excluding the overlapping area between the left and middle sub-image areas and the middle area between the overlapping area between the right and left sub-image areas; S44, model-guided optimization: In the generation process, the ControlNet model and the self-trained LoRA model are introduced, where the ControlNet model is used to provide spatial structure constraints, and the self-trained LoRA model is used to guide style adjustment. At the same time, the image generation is guided by a fixed random seed and a style language description provided by the user; S45. Horizontal gradient mask fusion: Set a mask with gradient transparency in the overlapping area to achieve a completely continuous visual experience and obtain a 720° panoramic image.

6. The method for adaptively generating indoor panoramic images based on AIGC according to claim 5, characterized in that: In step S412 and step S44, the guidance parameters generated by the self-training LoRA model are set between 0.35-0.65, and the training steps of the self-training LoRA model are as follows: The first step is to collect panoramic image data sets: collect the six-way map data of the indoor space. The six-way map data includes the geographical location, name, type and corresponding six-way map of the room. After being classified and sorted by room type, the six-way map is imported into the Blender 3D software. The six-way map is converted into a panoramic image by using the method of converting a cube into an equidistant cylindrical projection. The six-way map is then rendered and exported to obtain a panoramic image data set. Step 2: Data annotation and preparation: Use the WD14 tool to annotate the synthetic panoramic image. The annotated data includes furniture, decoration, and space structure, and add unified text description labels. Step 3: Model training: Input the panoramic image dataset and annotated data into the LoRA model, and train the LoRA model using the cross entropy loss function. In this process, adjust the pre-trained weights of the LoRA model based on the annotated data, and use low-rank decomposition to complete parameter updates.

7. The method for adaptively generating indoor panoramic images based on AIGC according to claim 6, characterized in that: The pre-training weight adjustment formula of the LoRA model is as follows: ; ; In the formula, Represents the original weight matrix of the LoRA model; Indicates that the weight update matrix is ​​obtained by low-rank decomposition; and Respectively represent the shapes and Matrix of represents low-rank dimension; and Represent the rows and columns of the weight matrix respectively; Represents the output vector of the LoRA model; Represents the input vector of the LoRA model; The cross entropy loss function expression is as follows: ; In the formula, Represents the total loss value of model training; Represents the total number of training samples; Indicates The loss value of samples; Indicates Sample pairs The true label of Indicates samples are predicted to be of class probability; Indicates the number of categories.

8. The method for adaptively generating indoor panoramic images based on AIGC according to claim 7, characterized in that: In step S42, the overlapping ratio of adjacent sub-image regions is 1 / 3, and when generating the right sub-image region, the left 1 / 3 overlapping region of the left sub-image region is spliced ​​to the right 1 / 3 overlapping region of the right sub-image region to ensure the consistency of the spatial structure of the image; After the left sub-image area is generated, the left 1 / 3 overlapping area of ​​the middle sub-image area and the right 1 / 3 overlapping area of ​​the right sub-image area are updated; after the middle sub-image area is generated, the left 1 / 3 overlapping area of ​​the right sub-image area is updated; And the mask edge areas of the three sub-image areas are all feathered.

9. An AIGC-based indoor panoramic image adaptive generation system, used to execute the AIGC-based indoor panoramic image adaptive generation method according to claim 8, characterized in that: include: Data acquisition and preprocessing module: used to obtain the expected design style language description and floor plan and perform preprocessing; The layout construction module is used to generate layout data that meets the constraints based on the pre-processed floor plan using the pre-trained MiDiffusion model; The panoramic image generation module is used to build a basic 3D model based on the layout data, add key structural elements, build a complete 3D room model, and then perform panoramic image rendering to generate a panoramic white film image; The panoramic design scheme generation module is used to extract the line draft and depth map of the panoramic white film image as constraint guidance conditions, and based on the extracted constraint guidance conditions, use the ControlNet model and the self-trained LoRA model to generate a preliminary panorama, then divide the overlapping sub-image areas of the preliminary panorama, set a dynamic mask to optimize each sub-area, and use the horizontal gradient mask optimization to complete the fusion image output to obtain a 720° panorama; Output module, used to output 720° panoramic images.

Citation Information

Patent Citations

  • Real scene three-dimensional visualization method and system based on semantic point cloud

    CN116912437A

  • Building intelligent design system based on AIGC

    CN117632098A

  • Panoramic image generation method and device, electronic equipment and storage medium

    CN119515674A

  • Virtual scene construction method and system, electronic equipment and storage medium

    CN119832152A

  • Video communication with changeable background

    WO2024215546A1

Cited By

  • Mixed diffusion model-based workshop equipment layout generation method and equipment, and storage medium

    CN120509104A

  • A workshop equipment layout generation method, device and storage medium based on a mixed diffusion model

    CN120509104B

  • Multi-conditional constraint fusion generation method and system for planning intention graph generation

    CN120745021A

  • Subway room layout intelligent design method based on Stable Diffusion

    CN120930230A

  • Intelligent generation method and system for territorial space planning effect picture

    CN121304848A