Architectural style-oriented architectural design image generation method and device
By combining the U-Net model with a style encoder, architectural design images are generated using architectural sketches and style text prompts, solving the problems of low efficiency and insufficient accuracy in existing technologies, and achieving efficient and accurate generation of architectural design images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-31
AI Technical Summary
Existing architectural design image generation solutions are inefficient and inaccurate, especially when dealing with specific architectural styles, lacking domain adaptability and semantic understanding capabilities, resulting in insufficient rendering quality and detail.
By acquiring architectural sketches and architectural style text prompts input by the user, the U-Net model is used in conjunction with a style encoder and a cross-modal attention mechanism to extract architectural features and fuse style vectors for feature enhancement, ultimately generating architectural design images.
It improves the efficiency and accuracy of architectural design image generation, achieving high-fidelity and rapid architectural design image generation that meets the geometric constraints and semantic requirements of specific styles.
Smart Images

Figure CN121767490A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and apparatus for generating architectural design images oriented towards architectural styles. Background Technology
[0002] In the traditional architectural design process, designers often rely on a variety of tools and methods to express design concepts, such as hand-drawn sketches, physical models, or computer-aided design (CAD) software. While these methods have some effectiveness, they reveal some inherent limitations when facing current industry demands: such as low efficiency and high iteration costs, limited expressiveness and realistic feeling, and restrictions on creative expression and style exploration.
[0003] In recent years, generative artificial intelligence, represented by Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models, has made significant progress in the field of image generation. However, when these general-purpose models are directly applied to the field of architectural design, they still face challenges such as insufficient domain adaptability, inadequate semantic understanding capabilities, and poor rendering quality and detail representation.
[0004] Therefore, how to efficiently and accurately generate architectural images of specific architectural design styles has become a technical problem that urgently needs to be solved. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and apparatus for generating architectural design images oriented towards architectural styles, in order to solve the problems of poor efficiency and accuracy of existing architectural image generation schemes.
[0006] To address the aforementioned problems, in a first aspect, the present invention provides a method for generating architectural design images oriented towards architectural styles, comprising: Obtain the architectural sketches and architectural style suggestions input by the user; The image generation model takes architectural sketches and architectural style text prompts as inputs and outputs architectural design images. The image generation model is trained on a U-Net model with sample sketches carrying architectural style annotations as samples and architectural design images corresponding to the sample sketches as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
[0007] In one possible implementation, the step of taking architectural sketches and architectural style text prompts as input to an image generation model to obtain an architectural design image output by the image generation model includes: Extract architectural features from architectural sketches and convert architectural style text into style vectors based on a style encoder; The architectural features and style vectors in the architectural sketch are fused based on the cross-modal attention mechanism to obtain the first architectural design image; The first feature vector is obtained by extracting features from the first architectural design image, and the first feature vector is enhanced based on the style vector to obtain the second feature vector; Architectural design images are generated based on the second feature vector.
[0008] In one possible implementation, extracting architectural features from the architectural sketch includes: Extract architectural features from architectural sketches using Control-Net or T2I-Adapter in the U-net model.
[0009] In one possible implementation, the method of fusing architectural features and style vectors in the architectural sketch based on a cross-modal attention mechanism to obtain a first architectural design image includes: The cross-modal attention mechanism between the upsampling and downsampling layers based on the U-net model fuses architectural features and style vectors in the architectural sketch to obtain the first architectural design image. There is a residual connection between the cross-modal attention mechanism and the U-net model.
[0010] In one possible implementation, the step of extracting features from the first architectural design image to obtain a first feature vector includes: The first feature vector is obtained by extracting features from the first architectural design image based on multi-layer convolution function and linear rectification function.
[0011] In one possible implementation, the feature enhancement of the first feature vector based on the style vector includes: Based on multi-scale residual dense blocks or Transformer, the style vector and the first feature vector are fused to enhance the features of the first feature vector.
[0012] In one possible implementation, generating the architectural design image based on the second feature vector includes: The upsampling layer based on the U-net model reconstructs the features of the second feature vector to generate an architectural design image.
[0013] On the other hand, the present invention also provides an architectural design image generation device oriented towards architectural styles, comprising: The acquisition module is used to acquire architectural sketches and architectural style text prompts input by the user; The generation module takes architectural sketches and architectural style text prompts as input to the image generation model and outputs architectural design images. The image generation model is trained on the U-Net model using sample sketches with architectural style annotations as samples and the corresponding architectural design images as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
[0014] Secondly, the present invention also provides an image generation device, including a memory and a processor, wherein, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the architectural design image generation method for architectural styles described in any of the above implementations.
[0015] Thirdly, the present invention also provides a computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, can implement the steps in the architectural design image generation method for architectural styles described in any of the above implementations.
[0016] The beneficial effects of this invention are as follows: The architectural design image generation method and apparatus provided by this invention, which is oriented towards architectural style, applies geometric and architectural style constraints to the generated architectural design image by acquiring the architectural sketch and architectural style text prompts input by the user. Then, the architectural sketch and architectural style text prompts are used as input to an image generation model to obtain the architectural design image output by the image generation model, thereby realizing the generation of architectural design images. The image generation model is equipped with a style encoder for vectorizing the architectural style text prompts, which can integrate architectural style during the architectural design image generation process, thereby improving the accuracy of architectural design image generation. Directly generating architectural design images through the image generation model can also improve the efficiency of architectural design image generation. This invention effectively improves the efficiency and accuracy of architectural design image generation. Attached Figure Description
[0017] Figure 1 A schematic flowchart of an embodiment of the architectural design image generation method for architectural styles provided by the present invention; Figure 2 This is a schematic diagram of an embodiment of the architectural style dataset construction process provided by the present invention; Figure 3 A schematic diagram of an embodiment of the business processing flow of the image generation model provided by the present invention; Figure 4 A schematic diagram illustrating an embodiment of the architectural style integration process provided by the present invention; Figure 5 A schematic diagram of an embodiment of the architectural design image generation system provided by the present invention; Figure 6 A schematic diagram of an embodiment of the architectural design image generation device for architectural styles provided by the present invention; Figure 7 This is a schematic diagram of an embodiment of the image generation device provided by the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0019] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0020] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0021] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0022] This invention provides a method and apparatus for generating architectural design images based on architectural styles, which will be described below.
[0023] Figure 1 This is a schematic flowchart of an embodiment of the architectural design image generation method for architectural styles provided by the present invention, as shown below. Figure 1 As shown, the methods for generating architectural design images based on architectural style include: S101. Obtain the architectural sketch and architectural style text prompts input by the user.
[0024] It should be noted that the architectural design image generation method for architectural styles provided by this invention can be applied to architectural design image generation scenarios, especially architectural design image generation scenarios for specific architectural styles.
[0025] Before generating architectural design images, the image generation device (such as a desktop or portable computer) can first obtain the architectural sketch and architectural style text prompts input by the user. The architectural sketch is used to impose geometric constraints on the generated architectural design image, while the architectural style text prompts are used to impose architectural style constraints on the generated architectural design drawing if it is too small.
[0026] S102. The architectural sketch and architectural style text prompts are used as input to the image generation model to obtain the architectural design image output by the image generation model. The image generation model is trained on the U-Net model with sample sketches carrying architectural style annotations as samples and the architectural design images corresponding to the sample sketches as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
[0027] It should be noted that after obtaining the architectural sketch and architectural style text prompts input by the user, these can be used as input to the image generation model to obtain the architectural design image output by the model, thus achieving the generation of the architectural design image. The image generation model is trained on the U-Net model using sample sketches with architectural style annotations as samples and the corresponding architectural design images as labels. Furthermore, the U-Net model includes a style encoder for vectorizing the architectural style text prompts, which allows for the integration of architectural styles during the architectural design image generation process, improving the accuracy of the generated image. Alternatively, directly generating the architectural design image through the image generation model can also improve the efficiency of architectural design image generation.
[0028] In summary, the architectural design image generation method based on architectural style provided by this invention applies geometric and architectural style constraints to the generated architectural design image by acquiring the architectural sketch and architectural style text prompts input by the user. Then, the architectural sketch and architectural style text prompts are used as input to an image generation model to obtain the architectural design image output by the image generation model, thus realizing the generation of architectural design images. The image generation model is equipped with a style encoder for vectorizing the architectural style text prompts, thereby integrating architectural style during the architectural design image generation process and improving the accuracy of architectural design image generation. Directly generating architectural design images through the image generation model also improves the efficiency of architectural design image generation. Therefore, this invention effectively improves both the efficiency and accuracy of architectural design image generation.
[0029] In some embodiments of the present invention, the step of using architectural sketches and architectural style text prompts as input to an image generation model to obtain an architectural design image output by the image generation model includes: Extract architectural features from architectural sketches and convert architectural style text into style vectors based on a style encoder; The architectural features and style vectors in the architectural sketch are fused based on the cross-modal attention mechanism to obtain the first architectural design image; The first feature vector is obtained by extracting features from the first architectural design image, and the first feature vector is enhanced based on the style vector to obtain the second feature vector; Architectural design images are generated based on the second feature vector.
[0030] It should be noted that in the process of using architectural sketches and architectural style text prompts as input to an image generation model to obtain the architectural design image output by the model, the architectural features in the sketches are first extracted, and the architectural style text is converted into style vectors using a style encoder. Then, a cross-modal attention mechanism is used to fuse the architectural features in the sketches with the style vectors to obtain the first architectural design image. Next, feature extraction is performed on the first architectural design image to obtain the first feature vector. Then, the first feature vector is enhanced using the style vector to obtain the second feature vector. Finally, the architectural design image is generated using the second feature vector, completing the generation of the architectural design image.
[0031] In some embodiments of the present invention, the extraction of architectural features from architectural sketches includes: Extract architectural features from architectural sketches using Control-Net or T2I-Adapter in the U-net model.
[0032] It should be noted that when extracting architectural features from architectural sketches, architectural features can be extracted using Control-Net or T2I-Adapter in the U-net model.
[0033] In some embodiments of the present invention, the step of fusing architectural features and style vectors in an architectural sketch based on a cross-modal attention mechanism to obtain a first architectural design image includes: The cross-modal attention mechanism between the upsampling and downsampling layers based on the U-net model fuses architectural features and style vectors in the architectural sketch to obtain the first architectural design image. There is a residual connection between the cross-modal attention mechanism and the U-net model.
[0034] It should be noted that when fusing architectural features and style vectors from architectural sketches using a cross-modal attention mechanism to obtain the first architectural design image, the cross-modal attention mechanism between the upsampling and downsampling layers of the U-net model can be used to fuse the architectural features and style vectors from the architectural sketches to obtain the first architectural design image. Residual connections exist between the cross-modal attention mechanism and the U-net model.
[0035] In some embodiments of the present invention, the step of extracting features from the first architectural design image to obtain a first feature vector includes: The first feature vector is obtained by extracting features from the first architectural design image based on multi-layer convolution function and linear rectification function.
[0036] It should be noted that when extracting features from the first architectural design image to obtain the first feature vector, the first feature vector can be obtained by using multi-layer convolutional functions and linear rectified functions.
[0037] In some embodiments of the present invention, the feature enhancement of the first feature vector based on the style vector includes: Based on multi-scale residual dense blocks or Transformer, the style vector and the first feature vector are fused to enhance the features of the first feature vector.
[0038] It should be noted that when performing feature enhancement on the first feature vector based on the style vector, the style vector and the first feature vector can be fused using multi-scale residual dense blocks or Transformers to enhance the features of the first feature vector.
[0039] In some embodiments of the present invention, generating the architectural design image based on the second feature vector includes: The upsampling layer based on the U-net model reconstructs the features of the second feature vector to generate an architectural design image.
[0040] It should be noted that when generating architectural design images based on the second feature vector, the second feature vector can be reconstructed using the upsampling layer of the U-net model to generate the architectural design images.
[0041] This invention achieves rapid and accurate generation of high-fidelity architectural design drawings from sketches and style cues through data processing, multi-stage fine-tuning strategies, and the integration of style-aware rendering technology. The specific generation process includes the following steps: 1. Construction and preprocessing of architectural style datasets, combined with Figure 2 The process specifically includes: 1) Acquisition of Multi-Source Architectural Image Resources: Obtain architectural facade images, sections, renderings, and other materials covering various architectural styles (e.g., modern, classical, gothic, industrial, parametric, etc.) from multiple sources, including but not limited to internal project databases, public architectural design libraries, and architectural photography websites. For example, web scraping and API interfaces can be used to acquire a vast amount of architectural images from public architectural design libraries, official websites of renowned architectural firms, and architectural photography websites. These images cover a wide range of mainstream and niche architectural styles, from traditional Chinese and European classical to modern minimalism, industrial, neo-Chinese, parametric, and postmodernism. During acquisition, high-resolution, multi-angle images should be prioritized, and their original source information should be recorded.
[0042] 2) Style Tagging and Semantic Annotation: For the acquired image materials, architectural styles are classified and tagged manually or semi-automatically. Simultaneously, natural language processing technology is used to perform deep semantic annotation on the images, extracting information such as key architectural components, materials, lighting conditions, and environmental background to form structured text prompts. First, a semi-automated image classification algorithm is used for preliminary architectural style classification, generating initial style tags. Subsequently, domain experts can manually refine and verify the preliminary classification results. Experts will assign precise architectural style tags to each image based on the building's history, region, materials, and form.
[0043] For each refined architectural image, deep semantic annotation is performed using natural language processing techniques. The specific implementation process is as follows: First, the image is input into an image-to-text encoder to extract its visual features. Then, leveraging its powerful knowledge base and pre-trained contextual understanding capabilities, the model identifies key architectural components (such as facades, roofs, windows, balconies, colonnades, and arches), main building materials (such as concrete, glass curtain walls, red brick, wood, and stone), lighting conditions (such as ample sunlight, cloudy skies, and sunset), environmental background (such as cityscapes, surrounding mountains and forests, and coastal scenery), and building functions (such as residential, commercial, and cultural centers). Based on these identification results, structured and detailed textual prompts are automatically generated, such as: "High-rise modern minimalist residential building, glass curtain wall, partial concrete facade, ample south-facing lighting, surrounded by urban green landscape" or "Gothic cathedral, pointed arches, rose windows, stone carvings, dim interior lighting, rich historical feel." These prompts not only describe the image content but also contain deep architectural semantics.
[0044] 3) Image Quality Screening and Enhancement: Initial screening of image materials removes low-quality, blurry, or unacceptable images. Standardization processing is then applied to acceptable images, including but not limited to size normalization, color correction, and contrast enhancement, to improve the overall quality of the dataset. Automatic Screening: Image quality assessment algorithms automatically identify and remove low-quality images that are blurry, overexposed, underexposed, or contain severe artifacts. Standardization Processing: The retained images undergo uniform size normalization, color correction, contrast enhancement, and sharpening to ensure visual quality and consistency before being input into the model.
[0045] 4) Sketch-Image Pair Generation: For images lacking corresponding sketches, edge detection algorithms are used to extract line drawings or contour information from the images, generating sketches corresponding to the architectural images. Finally, a dataset containing architectural images, style tags, semantic cues, and corresponding sketches is constructed.
[0046] 2. Model pre-training and architecture integration.
[0047] Combination Figure 3 It can be seen that a pre-trained basic generative network with anisotropic parameter optimization is used to complete domain-specific transfer learning training using a domain-specific sample set (i.e., a private dataset of enterprises with professional building annotations).
[0048] 1) Semantic Text Encoding: The system receives natural language descriptions from user input and converts them into serialized text embedding vectors using a text encoder. During this process, the system automatically injects preset "style trigger words" into the input text. These trigger words correspond to specific weight distributions learned by the model during training on domain-specific sample sets, and are used to activate the model's response to specific building codes.
[0049] 2) Construction of Domain Embedding Vector: Unlike general text descriptions, the system introduces an independent domain knowledge embedding vector. This vector is a global feature representation obtained by extracting and clustering features from a domain-specific sample set (including enterprise-specific facade materials, apartment layout standards, and brand color specifications). The domain knowledge embedding vector is not processed by a text encoder; instead, it is directly mapped as an additional key-value pair to the subsequent attention mechanism to "lock" the overall style of the generated image in line with enterprise standards.
[0050] 3) Geometric Constraint Feature Extraction: To achieve feature-guided processing based on geometric constraints, the system receives control base maps (such as CAD line drawings, color blocks, or volumetric sketches) uploaded by the user. A lateral control network encodes the control base map using zero-convolutional layers to extract geometric condition feature maps containing spatial structural information. These feature maps preserve the building's edges, contours, and spatial layout information to prevent geometric distortion during the generation process.
[0051] 4) Domain-Specific Optimized Latent Space Iterative Denoising: Input data is processed within the latent space through a T-step backdiffusion operation using an improved U-Net network architecture. The network's weight parameters have been adjusted for enterprise materials using parameter anisotropy optimization.
[0052] For each time step, the internal data processing flow is as follows: Feature fusion and noise prediction: The latent image vector at the current time step is input into the backbone network. At each convolutional layer of the network, geometric conditional feature maps are superimposed on the backbone features through residual connections, forcing the model to strictly follow the spatial structure of the control base map (such as column alignment and window-to-wall ratio) when generating pixel details.
[0053] Multi-source cross-attention mechanism: By explicitly using domain knowledge embedding vectors as input to the attention mechanism, the model will index texture features learned from domain-specific sample sets (such as specific aluminum plate joint treatment and glass curtain wall reflectivity) with high weight when calculating the correlation between pixels. This makes the generated architectural details conform to specific aesthetic standards and building codes, rather than the random textures of a general model.
[0054] Dynamic activation of domain-specific weights: During computation, the network not only calls upon the general weights of the pre-trained basic generative network, but also activates the incremental weight matrix fine-tuned for enterprise-specific materials. This mechanism ensures that the model retains its ability to generate general environments (sky, plants) while also possessing the ability to generate professional architectural components that conform to enterprise standards.
[0055] 5) Reconstruction and Decoding of Latent Features: After T iterations of denoising, a clean latent feature vector is obtained. This vector contains highly compressed image information that integrates geometric constraints and domain style.
[0056] 3. Large model integration.
[0057] The large model includes a frozen encoder-decoder structure (responsible for latent representation transformation of the image), a U-Net diffusion model (responsible for the denoising process), and a pre-trained text encoder (responsible for converting text prompts into latent semantic vectors).
[0058] Combination Figure 4 Furthermore, a style adapter module can be embedded in the model. This module is located between each downsampling and upsampling layer of U-Net, receives feature maps from U-Net through residual connections, incorporates external style information for processing, and then outputs it back to U-Net.
[0059] Furthermore, a separate style encoder can be designed for the style tags and semantic annotations constructed in step 1. This encoder can encode architectural style tags and more detailed style description text into high-dimensional style vectors. The style encoder can be a small Transformer network or a multilayer perceptron.
[0060] Within the style adapter module, a cross-modal attention mechanism is integrated. This mechanism allows the U-Net's image feature maps to interact with the style vectors output by the style encoder, enabling the diffusion process to better perceive and fuse specific architectural style features.
[0061] 4. Fine-tuning of the diffusion model for architectural styles in multiple stages.
[0062] 1) Structure-guided fine-tuning (Control-Net / T2I-Adapter integration): The main parameters of the large model are frozen, and training is performed only on structure guidance modules such as Control-Net or T2I-Adapter. Using the sketch-image pairs generated in step 1, the sketch is input as additional condition into Control-Net / T2I-Adapter to guide the diffusion model to generate architectural images with accurate structural outlines. The reconstruction error between the generated image and the real image is minimized while ensuring structural consistency.
[0063] How Control-Net works: Control-Net is a trainable neural network that replicates the U-Net encoder and adds a zero-convolutional layer to each replicated layer, allowing it to conditionally control the original U-Net's denoising process with a sketch as additional input. During fine-tuning, the U-Net parameters are frozen, and only the Control-Net parameters are trained.
[0064] Training data: The dataset was trained using the sketch-image configuration described above. The sketches were used as input to Control-Net, the images as input to the original U-Net, and the original images were used as the reconstruction targets.
[0065] Loss functions: L1 / L2 reconstruction loss (minimizing the pixel difference between the generated image and the real image) and perceptual loss (using a pre-trained VGG network to extract features and minimize the feature space difference) are mainly used to ensure that the generated image is close to the real image at both the pixel level and the semantic level, while maintaining structural accuracy. 2) Style-aware LoRA fine-tuning.
[0066] The main parameters of the large model are frozen, and LoRA is used to fine-tune some key Transformer layers in the style adapter module and U-Net. LoRA significantly reduces the number of trainable parameters by introducing a low-rank matrix to modify the weights of the pre-trained model. Style vectors and semantic vectors are generated by combining the constructed style labels and semantic cues through a text encoder and a style encoder. During denoising, these vectors interact with the feature maps of U-Net through a cross-modal attention mechanism. The style consistency and detail representation of the generated images are further optimized while maintaining semantic relevance to the text cues.
[0067] 5. Style-aware rendering and detail enhancement.
[0068] 1) Shallow Feature Extraction and Style Context Analysis: The architectural design drawing generated in step 4 is input into a shallow feature extraction network, which contains multiple convolutional layers and linear rectified functions. This network aims to capture low-level edge, texture, and preliminary structural information of the image, and combines it with the style vector provided by the style encoder for context analysis to identify style-related regions that require focused rendering. The feature maps extracted by the shallow feature extraction network are fused with the style vector output by the style encoder (e.g., through channel concatenation or feature modulation). A small attention module analyzes the fused features to identify regions in the image that are highly relevant to the currently specified architectural style (e.g., spires and rose windows of Gothic architecture; large glass facades of modern architecture). This analysis result generates a "style attention mask," indicating which regions require stronger stylization and detail enhancement in subsequent processing.
[0069] 2) Deep Feature Enhancement and Multi-Scale Fusion: Shallow features and style context information are input into a deep feature enhancement network. This network employs multi-scale residual dense blocks or Transformer blocks to capture rich high-level semantic and style features from different receptive fields, and performs multi-scale fusion to ensure that style details are preserved and enhanced at different granularities. Multi-scale feature fusion is implemented internally within the network. For example, feature maps of different resolutions are fused through a feature pyramid network structure or U-Net-style skip connections. This ensures that the network can capture fine-grained textures while also understanding the overall architectural structure and style layout, avoiding local inconsistencies that may result from detail enhancement.
[0070] 3) Stylized Reconstruction and Detail Optimization: The feature map after deep feature enhancement is input into a stylized reconstruction module. This module uses upsampling convolutional layers to reconstruct the feature map into a high-resolution architectural design drawing. During the reconstruction process, "style consistency loss" and "detail preservation loss" are introduced to ensure that the generated image follows a specific architectural style while preserving the details of the original design intent to the maximum extent, and enhancing the realism of textures, materials, and lighting. After processing by this module, the feature map is reconstructed into a final high-resolution architectural design drawing, with a resolution reaching 4K or even higher, to meet professional rendering and printing requirements.
[0071] Style consistency loss: The Gram matrix of the reconstructed image and the real image is extracted using the VGG network, and the style consistency loss of the style-related regions that need to be rendered is calculated in detail to ensure that the material, texture and color of these key style regions are highly consistent with the target style.
[0072] Detail Preservation Loss: Using structural similarity index loss or VGG feature distance loss, we ensure that while improving resolution and stylization, we preserve the detailed structural and textual information in the original design intent to the greatest extent.
[0073] 4) Interactive Post-processing and Style Calibration: Provides an interactive user interface that allows designers to post-process rendered architectural drawings, such as adjusting lighting, color balance, and material reflectivity. Simultaneously, based on designer feedback, the parameters of the style adapter or rendering network can be iteratively adjusted on a small scale to achieve more precise style calibration.
[0074] Combination Figure 5 Furthermore, the present invention also provides an architectural design image generation system, which specifically includes: 1. Dataset construction module, including: Multi-source data acquisition submodule: Used to acquire architectural image materials from internal databases, public image libraries, the internet, etc.
[0075] Style and Semantic Annotation Submodule: Used for style classification, key component identification and deep semantic annotation of images.
[0076] Sketch-Image Pair Generation Submodule: Used to generate corresponding sketches from images or process existing sketch-image pairs.
[0077] Data preprocessing submodule: used for image quality screening, size normalization, color correction and other operations.
[0078] 2. Core model integration and fine-tuning module, specifically including: Large Model Integration Submodule: Used to load and integrate the base large model, including its frozen encoder-decoder, U-Net diffusion model, and text encoder.
[0079] Style Adapter Embedded Submodule: Used to embed and manage lightweight style adapter modules in U-Net.
[0080] Style Encoder Submodule: Used to encode architectural style tags and descriptions into style vectors.
[0081] Structure-guided fine-tuning submodule: Used for fine-tuning under the guidance of Control-Net or T2I-Adapter.
[0082] Style-aware LoRA fine-tuning submodule: Used for LoRA fine-tuning of style adapters and U-Net key layers.
[0083] 3. Style-aware rendering and enhancement module, specifically including: Shallow feature extraction network: used to extract low-level features and style context of an image.
[0084] Deep Feature Enhancement Networks: Used to capture and fuse multi-scale high-level semantic and stylistic features.
[0085] Stylized Reconstruction Module: Used to reconstruct feature maps into high-resolution architectural drawings and optimize textures, materials, and lighting.
[0086] Interactive post-processing submodule: Provides a user interface for adjusting lighting, color, materials, etc.
[0087] 4. Interface and interaction module, specifically including: Sketch input interface: Supports users to upload hand-drawn sketches, CAD line drawings, or any form of sketch captured from images.
[0088] Text prompt input interface: Supports users to input descriptions such as architectural style, materials, and environment.
[0089] Visual output interface: Displays the generated architectural design drawings, supporting multiple views and rendering qualities.
[0090] Feedback and Iterative Optimization Interface: Collects user feedback and supports iterative model optimization.
[0091] This invention enables the rapid generation of high-fidelity architectural design drawings from sketches, significantly shortening the cycle from design concept to visual renderings. Designers only need to provide simple sketches and style hints to quickly obtain high-quality architectural design drawings, greatly improving design efficiency.
[0092] This invention can generate a variety of design schemes with different styles based on different style prompts, helping designers to quickly explore and evaluate ideas in the early stages and inspire more innovative ideas.
[0093] This invention significantly improves image clarity, detail, and realism, enhances precise control over architectural design intent, and effectively utilizes existing resources while reducing development costs.
[0094] To better implement the architectural design image generation method oriented towards architectural style in the embodiments of the present invention, based on the architectural design image generation method oriented towards architectural style, correspondingly, such as... Figure 6 As shown, this embodiment of the invention also provides an architectural design image generation device oriented towards architectural styles. The architectural design image generation device 600 oriented towards architectural styles includes: The acquisition module 601 is used to acquire the architectural sketch and architectural style text prompts input by the user. The generation module 602 is used to take architectural sketches and architectural style text prompts as input to the image generation model to obtain the architectural design image output by the image generation model. The image generation model is trained on the U-Net model with sample sketches carrying architectural style annotations as samples and the architectural design images corresponding to the sample sketches as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
[0095] The architectural design image generation device 600 for architectural styles provided in the above embodiments can realize the technical solutions described in the above embodiments of the architectural design image generation method for architectural styles. The specific implementation principles of each module or unit can be found in the corresponding content in the above embodiments of the architectural design image generation method for architectural styles, and will not be repeated here.
[0096] like Figure 7 As shown, the present invention also provides an image generation device 700. The image generation device 700 includes a processor 701, a memory 702, and a display 703. Figure 7 Only some components of the image generation device 700 are shown, but it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.
[0097] In some embodiments, processor 701 may be a central processing unit (CPU), microprocessor, or other data processing chip, used to run program code stored in memory 702 or process data, such as the architectural design image generation method for architectural styles in this invention.
[0098] In some embodiments, processor 701 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 701 may be local or remote. In some embodiments, processor 701 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-cloud, etc., or any combination thereof.
[0099] In some embodiments, memory 702 may be an internal storage unit of image generating apparatus 700, such as a hard disk or memory of image generating apparatus 700. In other embodiments, memory 702 may also be an external storage device of image generating apparatus 700, such as a pluggable hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on image generating apparatus 700.
[0100] Furthermore, the memory 702 may include both internal storage units of the image generating device 700 and external storage devices. The memory 702 is used to store application software and various types of data installed on the image generating device 700.
[0101] In some embodiments, display 703 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an organic light-emitting diode (OLED) touchscreen, etc. Display 703 is used to display information from the image generating device 700 and to display a visual user interface. Components 701-703 of the image generating device 700 communicate with each other via a system bus.
[0102] In one embodiment, when processor 701 executes an architectural design image generation program oriented towards architectural style stored in memory 702, the following steps can be performed: Obtain the architectural sketches and architectural style suggestions input by the user; The image generation model takes architectural sketches and architectural style text prompts as inputs and outputs architectural design images. The image generation model is trained on a U-Net model with sample sketches carrying architectural style annotations as samples and architectural design images corresponding to the sample sketches as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
[0103] It should be understood that when the processor 701 executes the architectural design image generation program oriented towards architectural style in the memory 702, in addition to the functions mentioned above, it can also perform other functions, as can be found in the description of the corresponding method embodiments above.
[0104] Furthermore, this embodiment of the invention does not specifically limit the type of the image generating device 700 mentioned. The image generating device 700 can be a portable electronic device such as a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, or laptop computer. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic devices can also be other portable electronic devices, such as laptop computers with touch-sensitive surfaces (e.g., touch panels). It should also be understood that in some other embodiments of the invention, the image generating device 700 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0105] Accordingly, this application also provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the architectural design image generation method for architectural styles provided in the above-described method embodiments.
[0106] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0107] The above provides a detailed description of the architectural design image generation method and apparatus for architectural styles provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for generating architectural design images oriented towards architectural styles, characterized in that, include: Obtain the architectural sketches and architectural style suggestions input by the user; The image generation model takes architectural sketches and architectural style text prompts as inputs and outputs architectural design images. The image generation model is trained on a U-Net model with sample sketches carrying architectural style annotations as samples and architectural design images corresponding to the sample sketches as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
2. The architectural design image generation method oriented towards architectural style according to claim 1, characterized in that, The process of using architectural sketches and architectural style text prompts as input to an image generation model to obtain an architectural design image output by the image generation model includes: Extract architectural features from architectural sketches and convert architectural style text into style vectors based on a style encoder; The architectural features and style vectors in the architectural sketch are fused based on the cross-modal attention mechanism to obtain the first architectural design image; The first feature vector is obtained by extracting features from the first architectural design image, and the first feature vector is enhanced based on the style vector to obtain the second feature vector; Architectural design images are generated based on the second feature vector.
3. The architectural design image generation method oriented towards architectural style according to claim 2, characterized in that, The extracted architectural features from the architectural sketches include: Extract architectural features from architectural sketches using Control-Net or T2I-Adapter in the U-net model.
4. The architectural design image generation method oriented towards architectural style according to claim 2, characterized in that, The cross-modal attention mechanism fuses architectural features and style vectors from architectural sketches to obtain a first architectural design image, including: The cross-modal attention mechanism between the upsampling and downsampling layers based on the U-net model fuses architectural features and style vectors in the architectural sketch to obtain the first architectural design image. There is a residual connection between the cross-modal attention mechanism and the U-net model.
5. The architectural design image generation method oriented towards architectural style according to claim 2, characterized in that, The step of extracting features from the first architectural design image to obtain the first feature vector includes: The first feature vector is obtained by extracting features from the first architectural design image based on multi-layer convolution function and linear rectification function.
6. The architectural design image generation method oriented towards architectural style according to claim 2, characterized in that, The feature enhancement of the first feature vector based on the style vector includes: Based on multi-scale residual dense blocks or Transformer, the style vector and the first feature vector are fused to enhance the features of the first feature vector.
7. The architectural design image generation method oriented towards architectural style according to claim 2, characterized in that, The generation of architectural design images based on the second feature vector includes: The upsampling layer based on the U-net model reconstructs the features of the second feature vector to generate an architectural design image.
8. An architectural design image generation device oriented towards architectural style, characterized in that, include: The acquisition module is used to acquire architectural sketches and architectural style text prompts input by the user; The generation module takes architectural sketches and architectural style text prompts as input to the image generation model and outputs architectural design images. The image generation model is trained on the U-Net model using sample sketches with architectural style annotations as samples and the corresponding architectural design images as labels. The U-Net model contains a style encoder for vectorizing architectural style text prompts.
9. An image generation device, characterized in that, Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the architectural design image generation method for architectural styles as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the architectural design image generation method for architectural styles as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Building design method based on stable diffusion model multi-modal generation technology and application
CN120974589A
Cited By
A sketch-driven artistic image generation method, system and medium
CN122368221A