A method and system for generating architectural planning images based on a latent diffusion model

By improving the potential diffusion model, combining cross-attention and multi-scale strategies, high-resolution architectural planning images are generated, and high-resolution architectural planning images are solved in the existing technology, and efficient, detailed and standardized image generation is achieved.

CN119557955BActive Publication Date: 2025-07-08GUANGDONG YUANZHOU CULTURE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411625025.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-07-08
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

When existing AI technologies generate high-precision, rich details and conform to building codes, the computing resources are consumed and the generation efficiency is low. The generated images are prone to geometric distortion and blurred details, which cannot meet the design needs of complex scenarios.

Method used

The improved potential diffusion model is adopted, and data preprocessing and model training are carried out by introducing channel crossing and spatial crossing attention mechanisms, combining architectural style embedding vectors, text description embedding vectors and geometric information embedding vectors, the U-Net model structure is optimized, and multi-scale strategy and layer-by-layer generation strategy are adopted to perform data preprocessing and model training to generate high-resolution architectural planning images.

Benefits of technology

It improves computing efficiency and reduces computing costs. The generated images conform to functional layout and geometric structure specifications, has rich details, can meet diverse architectural design needs, and avoid geometric distortion and waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557955B_ABST
    Figure CN119557955B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for generating architectural planning images based on a latent diffusion model. The method includes steps such as data preprocessing, improving the latent diffusion model structure, optimizing input conditions, model training, and image generation. By preprocessing the architectural design image and urban planning image datasets and introducing channel cross and spatial cross attention mechanisms, the relevance between text descriptions and architectural styles is enhanced. At the same time, combining architectural style embedding vectors, architectural description text embedding vectors, and geometric information embedding vectors as additional input conditions ensures that the generated images contain complete information. During the model training process, a multi-scale strategy and depthwise separable convolutions are adopted to improve the model's recognition ability for large scenes and training efficiency. The system includes modules such as data preprocessing, model optimization, model training, and image generation, and can generate high-resolution architectural images that meet design specifications. The method and system can be widely applied in the fields of architectural design and urban planning to improve design efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method and system for generating architectural planning images based on a latent diffusion model. Background Art

[0002] Traditional architectural design methods are currently difficult to adapt to the changing design requirements brought about by the accelerating urbanization process, especially for the demand for architectural planning images. Currently, there have emerged technologies that apply AI technology in the field of image generation, such as GAN (Generative Adversarial Network) and Stable Diffusion. The former defaults to adapting to an image size of 512x512, and the latter relies on the confrontation between the generator and the discriminator. Both perform poorly when dealing with generating diverse images with high precision, rich details, and compliance with architectural specifications. In order to generate high-resolution and high-precision images, due to the inclusion of a lot of structural information and detailed elements, directly operating in the high-bit pixel space will have a large computational burden. Stable Diffusion chooses to compress the high-bit image to a low dimension for operation. Although it reduces the consumption of computing resources to a certain extent, it greatly sacrifices the quality and details of the image and cannot meet the requirements of high-precision images in complex scenarios. In the iterative calculation during the image generation process of GAN, each step requires the mutual training of the generator and the discriminator, resulting in huge costs in terms of time and computing resources and unable to meet the current efficient design generation requirements.

[0003] In addition, since Stable Diffusion defaults to adapting to an image size of 512x512, for image generation requirements beyond this resolution range, an upscaling algorithm needs to be used to process the low-resolution image. During the processing, the model relies on random seeds and noise vectors to freely perform interpolation and inference, and the generated content does not conform to the design scope and standards, nor can it ensure the accurate proportional relationship and structural layout of the building, that is, geometric distortion; and due to limitations such as the model resolution, when generating complex architectural scenes, some details will be blurred, such as material texture, facade decoration, etc. In the model training process of GAN, due to relying on the confrontation between the generator and the discriminator, not only is the computational efficiency low, but also it is prone to mode collapse. Therefore, it cannot process complex image information, and the generated images tend to be single, lacking architectural details and unable to meet the diverse needs in architectural design. Although GAN performs well in terms of fidelity and can generate "pseudo-real" images that are superficially realistic and difficult to distinguish, it also raises credibility issues in practical applications, and there may be a big difference from the actual design requirements in applications; at the same time, this training mechanism makes it difficult for the generated model to accurately synchronize and update parameters at each stage, resulting in a non-convergent situation, and finally the images generated by the model lack the logic and consistency of the architectural structure, that is, image distortion.

[0004] The latent diffusion model is a deep learning generation model based on AI technology that uses a step-by-step denoising method to generate high-quality images. However, it is currently only applied to general image generation and technical style transfer and has not been perfectly combined with the application methods of "architectural design and urban planning". When generating architectural planning images, it cannot meet some special needs in this field and urgently needs to be optimized.

[0005] Therefore, how to overcome the above-mentioned defects has become an important issue that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0006] The present invention relates to the technical field of image processing, and specifically to a method and system for generating architectural planning images based on a latent diffusion model, aiming to generate diverse architectural planning images with high precision, rich details, and in line with architectural specifications. To achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the embodiment of the present invention discloses a method for generating architectural planning images based on a latent diffusion model, including the following steps:

[0008] S1. Data preprocessing: Obtain dataset A containing architectural design images and dataset B containing urban planning images, and preprocess the images in dataset A and dataset B;

[0009] S2. Improve the structure of the latent diffusion model: Improve the network structure of the U-Net model, and first compress the architectural design images and urban planning images in the dataset input to the latent diffusion model into the latent space for reverse diffusion to obtain the latent representations of the preliminary architectural design images and the latent representations of the preliminary urban planning images;

[0010] Introduce channel cross and spatial cross attention mechanisms in each encoder stage of the U-Net model to enhance the correlation between text descriptions and architectural styles;

[0011] S3. Optimize the input conditions: Additionally introduce the architectural style embedding vector C, the architectural description text embedding vector D, and the geometric information embedding vector E in combination as additional input conditions for the latent diffusion model, and ensure that the generated images contain complete information on style, functional layout, and geometric structure by adding "architectural function segmentation" and "geometric constraint" embeddings;

[0012] S4. Model training: Construct a training dataset according to multiple training samples, and use the dataset to train the improved model; among them, each training sample includes the latent representation of the preliminary architectural image, the corresponding text description, and the latent representation of the preliminary urban planning image;

[0013] S5. Image Generation: After the model training is completed, input the latent representation of the target urban planning image, the architectural style embedding vector C, and the architectural description text embedding vector D into the optimized latent diffusion model to obtain the generated architectural planning image that meets the requirements.

[0014] Preferably, the step of "S1. Data Preprocessing: Obtain the dataset A containing architectural design images and the dataset B containing urban planning images, and preprocess the images in the dataset A and the dataset B" includes:

[0015] For large-scale architectural design images / urban planning images, divide the large-scale images into smaller regional blocks for training; among them, an overlapping part must be introduced at the boundary of each block of images to preserve the continuity and context information of the scene;

[0016] Use a multi-scale strategy to learn multi-resolution versions of the images to enhance the recognition ability of the training model for large-scale scenes;

[0017] Among them, for images with lower resolutions, they are used to train the model for learning the global structure, while for higher resolutions, they are used to train the model to learn local details. Finally, the images will be subjected to a unified normalization operation to ensure that image features at different scales can be uniformly input into the training model for training.

[0018] Preferably, for the architectural design images and urban planning images in the dataset input into the latent diffusion model in step S2, the variational autoencoder VAE is used to compress from high-dimensional to low-dimensional.

[0019] Preferably, the architectural style embedding vector C is generated by the following steps:

[0020] Use the CLIP ImageEncoder to encode the architectural design images to obtain the image embedding vector V0. After the embedding vector V0 is processed by the Transformer model, it is passed to different stages in the U-Net network multiple times through the cross-attention mechanism;

[0021] At the same time, adopt a layer-by-layer generation strategy. According to the complexity of the building, adopt a phased generation strategy. First, generate the basic framework of the building, and then generate the external details and internal functional areas;

[0022] Process V0 through a three-layer self-attention mechanism to extract the architectural style and geometric features, and generate the architectural style embedding vector C.

[0023] Preferably, the architectural text embedding vector D is generated by the following steps:

[0024] For each architectural design image in dataset A, generate descriptive text based on architectural style, structural features, functional area layout, and geometric information. Use CLIPTokenizer and CLIPTextModel to encode the text description and generate the architectural text embedding vector D.

[0025] Preferably, the model training process described in step S4 includes the following steps:

[0026] Perform forward diffusion on the architectural images in the training samples, gradually add noise to generate latent images with different noise levels, and use them to train the model to understand how to recover the images from the noise; adopt the Adam optimizer, and use momentum and weight decay techniques during the parameter update process to accelerate the convergence speed and improve the model stability;

[0027] After multiple rounds of training, when the evaluated architectural planning images meet the requirements, the training is completed. The evaluation criteria include the structural accuracy, style consistency, high resolution, and visual quality of the generated architectural planning images.

[0028] Preferably, in the model training step, using depthwise separable convolution includes a depthwise convolution step and a pointwise convolution step. In the depthwise convolution step, a convolution operation is applied separately to each input channel to capture spatial features, and in the pointwise convolution step, a 1x1 convolutional kernel is used to linearly combine the output of the depthwise convolution to integrate information from different channels.

[0029] The second aspect of the embodiments of the present invention discloses an architectural planning image generation system based on a latent diffusion model, which adopts the image generation method described in the first aspect, and includes:

[0030] A data preprocessing module 1, which is used to read and process the architectural design image and urban planning image datasets;

[0031] A model optimization module 2, which is used to improve the structure of the latent diffusion model and optimize the input conditions;

[0032] A model training module 3, which is used to train the optimized latent diffusion model by combining the architectural description text, architectural images, and planning images;

[0033] An image generation module 4, which is used to receive the input of the architectural image and planning image to be generated after the model training is completed, and generate high-resolution architectural images that meet the design specifications.

[0034] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0035] 1. Through the settings of steps S1 - S5 of the inventive method in this case, the computing efficiency is improved, the computing cost and complexity are reduced, the number of model parameters is significantly reduced during the computing process, the image generation speed is accelerated, meeting the high - efficiency requirements for the generation of building planning images in the real environment.

[0036] 2. The optimized latent diffusion model proposed in the present invention is used to generate building planning images, which can generate diverse image results according to actual needs, with rich details and high resolution. By introducing building style embedding vectors and text description embedding vectors, this model can more accurately understand and generate images that meet the requirements. Therefore, it ensures that the generated images comply with the functional layout and geometric structure specifications, avoiding the problem of geometric distortion, and the generated image results are more in line with the actual requirements.

[0037] 3. The training process of the latent diffusion model provided by the present invention can train on a relatively large number of sample data, different from the scarce training samples in the prior art due to the low model training efficiency. Therefore, the performance of the trained model (i.e., the finally generated image results) is more in line with the complexity or diversity requirements for building design in the current real - world scenarios.

[0038] 4. In the training process of the third module in the present invention, the Adam optimizer is adopted, effectively improving the model stability and overcoming the instability problems existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0040] Figure 1 is a flowchart of Embodiment 1 of the present invention;

[0041] Figure 2 is a structural schematic block diagram of Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0043] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects, rather than to describe a specific order. The terms "include" and "have" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0044] The embodiments of the present invention disclose an image generation method based on a latent diffusion model provided by the present invention, which can be applied to multiple fields such as architectural design, urban planning, virtual reality, etc., with wide application and good prospects.

[0045] Embodiment 1

[0046] As Figure 1 shown, a method for generating architectural planning images based on a latent diffusion model includes the following steps:

[0047] S1. Data preprocessing: Obtain dataset A containing architectural design images and dataset B containing urban planning images, and preprocess the images in dataset A and dataset B;

[0048] S2. Improve the structure of the latent diffusion model: Improve the network structure of the U-Net model. For the architectural design images and urban planning images in the dataset input to the latent diffusion model, first compress them into the latent space for reverse diffusion to obtain the latent representations of the preliminary architectural design images and the latent representations of the preliminary urban planning images;

[0049] Introduce channel cross and spatial cross attention mechanisms in each encoder stage of the U-Net model to enhance the correlation between text descriptions and architectural styles;

[0050] S3. Optimize the input conditions: Additionally introduce the architectural style embedding vector C, the architectural description text embedding vector D, and the geometric information embedding vector E in combination as additional input conditions for the latent diffusion model. By adding "architectural function partition segmentation" and "geometric constraint" embeddings, ensure that the generated images contain complete information on style, functional layout, and geometric structure;

[0051] S4. Model training: Construct a training dataset according to multiple training samples, and use the dataset to train the improved model; wherein, each training sample includes the latent representation of the preliminary architectural image, the corresponding text description, and the latent representation of the preliminary urban planning image;

[0052] S5. Image Generation: After the model training is completed, input the latent representation of the target urban planning image, the architectural style embedding vector C, and the architectural description text embedding vector D into the optimized latent diffusion model to obtain the generated architectural planning image that meets the requirements. Specifically, during model training, reference information on urban layout and planning is also obtained from each urban planning image in dataset B for training.

[0053] As described above, through the settings in step S1, preprocessing the images in dataset A (architectural design images) and dataset B (urban planning images) helps improve the training efficiency of the model and the quality of the generated images. Through the settings in S2, improving the network structure of the U-Net model and introducing the reverse diffusion process can compress the architectural design images and urban planning images into the latent space to obtain their latent representations, which helps the model capture the key information and features in the images. At the same time, introducing the channel cross and spatial cross attention mechanisms in each encoder stage of the U-Net model can enhance the model's ability to understand the correlation between the text description and the architectural style, making the generated images more in line with the requirements of the text description and the architectural style. Through the settings in S3, by additionally introducing the architectural style embedding vector C, the architectural description text embedding vector D, and the geometric information embedding vector E as additional input conditions for the latent diffusion model, more information and guidance can be provided for the model; and the optimized input conditions enable the model to design and generate images that meet the conditions and are diverse according to actual needs after training. Through the settings in S4, constructing a dataset containing multiple training samples and training with the improved model can enable the model to learn the correlation and rules between the architectural design images and the urban planning images. At the same time, obtaining reference information on urban layout and planning from each urban planning image in dataset B for training can further improve the model's understanding ability of urban planning and architectural design, making the generated images more in line with the actual requirements. Through the settings in S5, after the model training is completed, inputting the latent representation of the target urban planning image, the architectural style embedding vector C, and the architectural description text embedding vector D into the optimized latent diffusion model can generate the architectural planning image that meets the requirements. This generation method has high flexibility and customizability, and can generate architectural planning images with different styles and functions according to different needs and scenarios. At the same time, since the model has fully learned the correlation and rules between the architectural design images and the urban planning images, the generated images have high accuracy and credibility.

[0054] As a preferred implementation manner, the “S1. Data Preprocessing: Obtain dataset A containing architectural design images and dataset B containing urban planning images, and preprocess the images in dataset A and dataset B” includes:

[0055] For architectural design images / urban planning images of large scenes, the large scene images are segmented into smaller regional blocks for training; among them, overlapping parts must be introduced at the boundaries of each block of images to preserve the continuity and context information of the scene;

[0056] Using a multi-scale strategy, learn multi-resolution versions of the images to enhance the recognition ability of the training model for large scenes;

[0057] Among them, the lower-resolution images are used to train the model to learn the global structure, while the higher resolution is used to train the model to learn local details. Finally, a unified normalization operation will be performed on the images to ensure that image features at different scales can be uniformly input into the training model for training.

[0058] As described above, by segmenting the architectural design images and urban planning images of large scenes into smaller regions for block training, the computational burden of a single training can be significantly reduced. At the same time, overlapping parts are introduced at the boundaries of each block of images to ensure the continuity of the scene and the retention of context information, avoiding the fracture of image information caused by block segmentation. This not only improves the computational efficiency but also ensures the coherence of the image content. Using a multi-scale strategy to learn multi-resolution versions of the images enables the model to capture image features at different scales. The lower-resolution images are used to train the model to learn the global structure, while the higher resolution is used to capture local details, which can reduce the computational cost of operating directly on high-resolution images and improve the model's ability to capture image details. In addition, through the training of the multi-scale strategy, the model can learn the global structure and local details of the images at different scales, enabling the generated images to maintain the consistency of the global structure while also showing rich local details, such as material textures, facade decorations, etc.

[0059] As a preferred implementation manner, for the architectural design images and urban planning images in the dataset input to the latent diffusion model, the variational autoencoder VAE is used to compress from high dimension to low dimension. In this way, through the learned latent space representation during the compression process, VAE can extract key features and information in the images, which is crucial for generating high-quality and compliant architectural planning images.

[0060] As a preferred implementation manner, the architectural style embedding vector C is generated by the following steps:

[0061] Use the CLIP Image Encoder to encode the architectural design image to obtain the image embedding vector V0. After the embedding vector V0 is processed by the Transformer deep learning model, it is passed to different stages in the U-Net network multiple times through the cross-attention mechanism;

[0062] Meanwhile, a layer-by-layer generation strategy is adopted. According to the complexity of the building, a phased generation strategy is used. First, the basic framework of the building is generated, and then the external details and internal functional areas are generated;

[0063] Process V0 through a three-layer self-attention mechanism to extract architectural style and geometric features, and generate an architectural style embedding vector C.

[0064] As described above, by using the CLIP Image Encoder to convert images into embedding vectors and processing these vectors in the Transformer model, the operation in the high-dimensional pixel space is avoided, further reducing the computational burden. The layer-by-layer generation strategy allows the model to generate images in phases, from the basic framework to the external details and internal functional areas. This way of gradually adding details not only improves the generation efficiency but also enables the model to better handle the requirements of complex scenes and high-precision images. Specifically, by processing the embedding vector V0 through a three-layer self-attention mechanism, the method of this case can extract architectural style and geometric features and generate an architectural style embedding vector C, ensuring that the generated images maintain overall style consistency while also showing rich details and accurate geometric structures; moreover, the layer-by-layer generation strategy ensures that there will be no geometric distortion or detail blurring problems during the image generation process because each stage focuses on generating image content at a specific level. Compared with GAN, the method of this case avoids the problems of mode collapse and image singularity because the combination of the Transformer model and the U-Net network provides stronger image generation capabilities and higher diversity.

[0065] As a preferred implementation manner, the building text embedding vector D is generated by the following steps:

[0066] For each architectural design image in dataset A, generate descriptive text based on architectural style, structural features, functional area layout, and geometric information. Use CLIPTokenizer and CLIPTextModel to encode the text description to generate the architectural text embedding vector D. In specific implementation, CLIPTokenizer is a text encoder used to convert the text description into a series of digital tokens; these tokens are the basis for subsequent processing and can capture the lexical and syntactic information in the text. CLIPTextModel is a text model based on the CLIP (Contrastive Language–Image Pre-training) framework; it can further process the tokens generated by CLIPTokenizer, extract the key information in the text description, and generate the corresponding embedding vector. In specific implementation, the architectural text embedding vector D is a high-dimensional numerical representation that contains all the semantic information about architectural style, structural features, functional area layout, and geometric information in the descriptive text. This vector can be used as one of the input conditions for the latent diffusion model to guide the model to generate architectural images that match the descriptive text.

[0067] As described above, the method in this case converts complex image information into a concise text description by generating the architectural text embedding vector D, and then further encodes it into an embedding vector, avoiding the huge computational burden brought by directly operating in the high-dimensional pixel space, thereby reducing the computational cost. Moreover, the architectural text embedding vector D contains key information such as architectural style, structural features, functional area layout, and geometric information, which can more accurately understand and generate images that meet the requirements, ensure that the generated images comply with the functional layout and geometric structure specifications, avoid problems of geometric distortion, and the generated image results are more in line with the actual requirements.

[0068] As a preferred implementation, in the above architectural planning image generation method based on the latent diffusion model, the geometric information embedding vector E is a special numerical representation that is specifically used to capture and represent the geometric features or shape information of the building. During the process of generating architectural planning images, the geometric information embedding vector E can ensure that the generated images are consistent with the input description or design in terms of geometric structure, helping to avoid generating architectural images that are visually reasonable but geometrically incorrect. Moreover, by introducing the geometric information embedding vector E as an additional input condition for the latent diffusion model, the model can learn more information about the building's geometric structure, which helps the model better handle complex geometric shapes and layouts during the generation process. In practical applications, users may hope to generate architectural images with specific geometric features. By adjusting the geometric information embedding vector E, users can customize the required geometric shapes and layouts, thereby generating architectural planning images that meet specific requirements.

[0069] As a preferred implementation, the model training process described in step S4 includes the following steps:

[0070] Perform forward diffusion on the building images in the training samples, gradually add noise, and generate latent images with different noise levels for training the model to understand how to recover images from noise; use the Adam optimizer, and use momentum and weight decay techniques during the parameter update process to accelerate the convergence speed and improve the model stability;

[0071] After the building planning images after multiple rounds of training meet the requirements after evaluation, the training is completed. The evaluation criteria include the structural accuracy, style consistency, high resolution, and visual quality of the generated building planning images.

[0072] As described above, through the process of forward diffusion and noise addition, the model can learn the key information of how to recover images from noise under a relatively low computational burden, which helps to reduce the computational cost of directly operating in the high-bit pixel space. The use of the Adam optimizer further improves the computational efficiency, enabling the model to complete training in a shorter time. Moreover, the momentum and weight decay techniques in the Adam optimizer contribute to enhancing the stability and reliability of the model.

[0073] As a preferred implementation, in the model training step, the use of depthwise separable convolutions includes a depthwise convolution step and a pointwise convolution step. In the depthwise convolution step, a convolution operation is applied separately to each input channel to capture spatial features, and in the pointwise convolution step, a 1x1 convolution kernel is used to linearly combine the outputs of the depthwise convolution to integrate information from different channels. In this way, in the depthwise convolution step, each channel is independently processed, avoiding the repeated calculations between different channels in traditional convolutions; compared with traditional two-dimensional convolutions, depthwise separable convolutions greatly reduce the amount of computation. Moreover, in the pointwise convolution step, the computational amount of the 1x1 convolution kernel is much smaller than that of traditional convolution kernels, further reducing the computational cost. In addition, due to the reduction in the amount of computation, the computational efficiency of the model during training and inference has been significantly improved. The model can process input data faster, thereby shortening the training and inference time.

[0074] Embodiment 2

[0075] As Figure 2 shown, a building planning image generation system based on a latent diffusion model, adopting the image generation method described in Embodiment 1, includes:

[0076] A data preprocessing module 1 for reading and processing building design image and urban planning image datasets;

[0077] A model optimization module 2 for improving the structure of the latent diffusion model and optimizing the input conditions;

[0078] A model training module 3, configured to train the optimized latent diffusion model by combining building description texts, building images, and planning images;

[0079] An image generation module 4, configured to, after the model training is completed, receive inputs of the building image and the planning image to be generated, and generate high-resolution building images that meet the design specifications.

[0080] As described above, by optimizing the structure and input conditions of the latent diffusion model, the speed and efficiency of image generation are significantly improved, enabling the system to generate a large number of high-quality building planning images in a short time to meet the actual application requirements. The system can accurately capture the key information in the building description texts and the planning images, and generate high-resolution images that meet the design specifications.

[0081] In summary, by adopting the image generation method and system of this case, the computing efficiency is improved, the computing cost and computing complexity are reduced, the number of model parameters is significantly reduced during the computing process, the image generation speed is accelerated, and the high-efficiency requirements for building planning image generation in the real environment are met. Moreover, by applying the optimized latent diffusion model proposed by the present invention to generate building planning images, diverse image results can be generated according to actual needs, with rich details and high resolution. The building style embedding vector and the text description embedding vector are introduced, and this model can more accurately understand and generate images that meet the requirements. Therefore, it is ensured that the generated images meet the functional layout and geometric structure specifications, avoiding the problem of geometric distortion, and the generated image results are more in line with the actual requirements.

[0082] The above has introduced in detail a method and system for generating building planning images based on a latent diffusion model disclosed in the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for generating architectural planning images based on a latent diffusion model, characterized in that, Including the following steps: S1. Data preprocessing: Obtain dataset A containing architectural design images and dataset B containing urban planning images, and preprocess the images in dataset A and dataset B; S2. Improve the structure of the latent diffusion model: Improve the network structure of the U-Net model. For the architectural design images and urban planning images in the dataset input to the latent diffusion model, first compress them into the latent space for reverse diffusion to obtain the latent representations of the preliminary architectural design images and the latent representations of the preliminary urban planning images; Introduce channel cross and spatial cross attention mechanisms in each encoder stage of the U-Net model to enhance the correlation between the text description and the architectural style; S3. Optimize the input conditions: Introduce the combination of the architectural style embedding vector C, the architectural description text embedding vector D, and the geometric information embedding vector E as additional input conditions for the latent diffusion model. By adding "architectural function segmentation" and "geometric constraint" embeddings, ensure that the generated images contain complete information on style, functional layout, and geometric structure; S4. Model training: Construct a training dataset according to multiple training samples, and use the dataset to train the improved model; Among them, each training sample includes the latent representation of the preliminary architectural design image, the corresponding text description, and the latent representation of the preliminary urban planning image; S5. Image generation: After the model training is completed, input the latent representation of the target urban planning image, the architectural style embedding vector C, and the architectural description text embedding vector D into the optimized latent diffusion model to obtain the generated architectural planning image that meets the requirements.

2. The method for generating architectural planning images based on a latent diffusion model according to claim 1, wherein, The "S1. Data preprocessing: Obtain dataset A containing architectural design images and dataset B containing urban planning images, and preprocess the images in dataset A and dataset B" includes: For large-scale architectural design images / urban planning images, divide the large-scale images into smaller regional blocks for training; Among them, overlapping parts must be introduced at the boundaries of each block of images to retain the continuity and context information of the scene; Use a multi-scale strategy to learn multi-resolution versions of the images to enhance the recognition ability of the training model for large-scale scenes; Among them, for images with lower resolution, they are used to train the model for global structure learning, while higher resolution is used to train the model to learn local details. Finally, the images will be uniformly normalized to ensure that image features at different scales can be consistently input into the training model for training.

3. The method for generating architectural planning images based on a latent diffusion model according to claim 1, wherein, In step S2, for the architectural design images and urban planning images in the dataset input to the latent diffusion model, the high dimension is compressed to the low dimension through the variational autoencoder VAE.

4. The method for generating architectural planning images based on a latent diffusion model according to claim 1, characterized in that, The architectural style embedding vector C is generated by the following steps: Use the CLIP Image Encoder to encode the architectural design images to obtain the image embedding vector V0. After the embedding vector V0 is processed by the Transformer model, it is passed to different stages in the U-Net network through the cross attention mechanism multiple times; Meanwhile, adopt a layer-by-layer generation strategy. According to the complexity of the building, adopt a phased generation strategy. First, generate the basic framework of the building, and then generate the external details and internal functional areas. Process V0 through a three-layer self-attention mechanism to extract the architectural style and geometric features, and generate the architectural style embedding vector C.

5. The method for generating an architectural planning image based on a latent diffusion model according to claim 1, wherein The architectural description text embedding vector D is generated by the following steps: Based on each architectural design image in dataset A, generate a description text based on the architectural style, structural features, functional area layout, and geometric information. Use CLIPTokenizer and CLIPTextModel to encode the text description to generate the architectural description text embedding vector D.

6. The method for generating an architectural planning image based on a latent diffusion model according to claim 1, wherein, The model training process described in step S4 includes the following steps: Perform forward diffusion on the architectural planning images in the training samples, gradually add noise to generate latent images with different noise levels, which are used to train the model to understand how to recover the image from the noise; adopt the Adam optimizer, and use momentum and weight decay techniques during the parameter update process to accelerate the convergence speed and improve the model stability. The architectural planning images after multiple rounds of training are evaluated and meet the requirements, then the training is completed. The evaluation criteria include the structural accuracy, style consistency, high resolution, and visual quality of the generated architectural planning images.

7. The method for generating an architectural planning image based on a latent diffusion model according to claim 1, characterized in that, In the model training step, the use of depthwise separable convolution includes a depthwise convolution step and a pointwise convolution step. In the depthwise convolution step, a convolution operation is applied separately to each input channel to capture spatial features. In the pointwise convolution step, a 1x1 convolutional kernel is used to linearly combine the output of the depthwise convolution to integrate the information of different channels.

8. An architectural planning image generation system based on a latent diffusion model, characterized in that, Adopt the image generation method described in any one of claims 1-7, which includes: A data preprocessing module (1) for reading and processing the architectural design image and urban planning image datasets. A model optimization module (2) for improving the structure of the latent diffusion model and optimizing the input conditions. A model training module (3) for training the optimized latent diffusion model by combining the architectural description text, architectural design images, and urban planning images. An image generation module (4) for receiving the input of the architectural design image and urban planning image to be generated after the model training is completed, and generating a high-resolution architectural planning image that meets the design specifications.

Citation Information

Patent Citations

  • Method for generating continuous pictures by long text based on diffusion model

    CN117521672A

  • Diffusion image generation method and system based on retrieval and segmentation enhancement

    CN117725247A