A building rendering generation method, device, medium and equipment based on multi-stage conditional control

CN122550849APending Publication Date: 2026-08-11GUANGZHOU URBAN PLANNING & DESIGN SURVEY RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

其中,传统建模渲染方法虽然能够较好体现建筑结构和设计意图,但高度依赖人工建模、材质配置、灯光调节和渲染输出,制作周期较长,难以适应建筑方案阶段频繁修改和快速出图的应用需求;基于生成对抗网络的方法虽能够实现图像生成,但存在训练稳定性不足、对数据分布敏感、泛化能力有限等问题;基于扩散模型的方法虽然在图像真实感和风格兼容性方面具有一定优势,但在建筑场景下仍容易出现门窗位置错乱、构件比例失衡、透视关系异常以及局部细节不一致等情况,导致生成结果的可控性较差,从而导致效果图质量较低

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550849A_ABST
    Figure CN122550849A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, medium, and device for generating architectural renderings based on multi-stage conditional control. The method includes: acquiring a training dataset; training a target conditional control network in stages based on the training dataset; inputting an input reference image corresponding to the building to be generated into the target conditional control network; and outputting a target architectural rendering. This invention improves the quality of generated target architectural renderings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, medium and equipment for generating architectural renderings based on multi-stage condition control. Background Technology

[0002] In existing technologies, architectural renderings are typically generated using traditional modeling and rendering methods, generative adversarial network (GAN)-based methods, or diffusion model-based methods. While traditional modeling and rendering methods can effectively represent architectural structures and design intent, they heavily rely on manual modeling, material configuration, lighting adjustments, and rendering output, resulting in long production cycles and making them unsuitable for applications requiring frequent modifications and rapid rendering during the architectural design phase. GAN-based methods, although capable of image generation, suffer from insufficient training stability, sensitivity to data distribution, and limited generalization ability. While diffusion model-based methods offer advantages in image realism and style compatibility, they are prone to issues in architectural scenes, such as misplaced doors and windows, unbalanced component proportions, abnormal perspective relationships, and inconsistent local details, leading to poor controllability of the generated results and consequently, lower-quality renderings. Summary of the Invention

[0003] The purpose of this invention is to propose a method, apparatus, medium, and device for generating architectural renderings based on multi-stage conditional control. This involves constructing a target conditional control network comprising a first, second, third, and fourth conditional control sub-network, and training the target conditional control network in stages using training data corresponding to block diagrams, line drawings, model diagrams, style diagrams, and rendered renderings. This allows the input reference image corresponding to the building to be generated to gradually generate the target architectural rendering under multi-stage conditional constraints, thereby improving the structural controllability and interpretability of the architectural rendering generation process and enhancing the quality of the renderings.

[0004] To achieve the above objectives, a first aspect of the present invention provides a method for generating architectural renderings based on multi-stage condition control, the method comprising: Obtain the training dataset; Based on the training dataset, a target conditional control network is generated through phased training. The input reference image corresponding to the building to be generated is input into the target condition control network, and the target building rendering is output.

[0005] Furthermore, the training dataset includes at least a first training data, a second training data, a third training data, a fourth training data, and a fifth training data corresponding to the same building object.

[0006] Furthermore, the first training data is a block diagram, the second training data is a line drawing, the third training data is a model diagram, the fourth training data is a style diagram, and the fifth training data is a rendered effect diagram.

[0007] Furthermore, the target condition control network includes a first condition control subnetwork, a second condition control subnetwork, a third condition control subnetwork, and a fourth condition control subnetwork.

[0008] Further, the step of generating the target conditional control network in stages based on the training dataset includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are preprocessed to obtain first preprocessed training data, second preprocessed training data, third preprocessed training data, fourth preprocessed training data, and fifth preprocessed training data. The first preprocessed training data is input into the first conditional control subnetwork, and the second preprocessed training data is used as the output target of the first conditional control subnetwork to train the first conditional control subnetwork. The second preprocessed training data is input into the second conditional control subnetwork, and the third preprocessed training data is used as the output target of the second conditional control subnetwork to train the second conditional control subnetwork. The third preprocessed training data and the fourth preprocessed training data are input into the third conditional control subnetwork, and the fifth preprocessed training data is used as the output target of the third conditional control subnetwork to train the third conditional control subnetwork. The fifth preprocessed training data is input into the fourth conditional control subnetwork, and the fifth training data is used as the output target of the fourth conditional control subnetwork to train the fourth conditional control subnetwork. The target conditional control network is obtained based on the trained first conditional control subnetwork, second conditional control subnetwork, third conditional control subnetwork, and fourth conditional control subnetwork.

[0009] Further, the preprocessing of the first training data, the second training data, the third training data, the fourth training data, and the fifth training data includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are respectively subjected to resolution reduction processing to obtain the first preprocessed training data, the second preprocessed training data, the third preprocessed training data, the fourth preprocessed training data, and the fifth preprocessed training data.

[0010] Furthermore, the input reference image corresponding to the building to be generated includes the block diagram of the building to be generated and the style diagram of the building to be generated. The step of inputting the input reference image corresponding to the building to be generated into the target condition control network and outputting the target building rendering includes: Input the block diagram of the building to be generated into the trained first conditional control subnetwork, and output the target line drawing; The target line drawing is input into the trained second conditional control subnetwork, and the target model diagram is output. The target model image and the style image of the building to be generated are input into the trained third conditional control subnetwork, and the target low-resolution rendering effect image is output. The target low-resolution rendered image is input into the trained fourth conditional control subnetwork, which outputs the target architectural rendering.

[0011] To achieve the above objectives, a second aspect of the present invention also provides an architectural rendering generation apparatus based on multi-stage condition control, used to implement the architectural rendering generation method based on multi-stage condition control as described in any of the first aspects, the apparatus comprising: The training set acquisition module is used to acquire the training dataset; The conditional control network generation module is used to train and generate a target conditional control network in stages based on the training dataset. The target building rendering output module is used to input the input reference image corresponding to the building to be generated into the target condition control network and output the target building rendering.

[0012] A third aspect of the present invention also provides a computer-readable storage medium comprising a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the architectural rendering generation method based on multi-stage condition control described in any of the first aspects above.

[0013] A fourth aspect of the present invention also provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the architectural rendering generation method based on multi-stage condition control as described in any of the first aspects above. Attached Figure Description

[0014] Figure 1 This is a flowchart of a preferred embodiment of a method for generating architectural renderings based on multi-stage condition control provided in the first aspect of the present invention; Figure 2This is a flowchart illustrating another preferred embodiment of a method for generating architectural renderings based on multi-stage condition control, provided in the first aspect of the present invention. Figure 3 This is a schematic diagram of a block diagram, line drawing, model diagram, target style diagram and rendering diagram of another preferred embodiment of the architectural rendering generation method based on multi-stage condition control provided in the first aspect of the present invention; Figure 4 This is a structural block diagram of a preferred embodiment of an architectural rendering generation device based on multi-stage condition control provided in the second aspect of the present invention; Figure 5 This is a structural block diagram of a preferred embodiment of a terminal device provided in the fourth aspect of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] It should be noted that the data involved in this invention (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0017] In this embodiment of the invention, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplarily" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0018] In this invention description, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In this invention description, unless otherwise stated, "a plurality of" means two or more. In this invention description, the term "comprising" and its variations are open-ended, meaning "including but not limited to." The term "based on" means "at least partially based on." The term "according to" means "at least partially according to." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments."

[0019] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0020] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0021] The technical solution of the present invention will be further described below with reference to specific embodiments: The first aspect of this invention provides a method for generating architectural renderings based on multi-stage condition control, see [link to relevant documentation]. Figure 1 The diagram shown is a flowchart of a preferred embodiment of a method for generating architectural renderings based on multi-stage condition control provided by the first aspect of the present invention. The method includes steps S1 to S3, as follows: Step S1: Obtain the training dataset; Step S2: Based on the training dataset, generate the target conditional control network in stages; Step S3: Input the reference image corresponding to the building to be generated, input the target condition control network, and output the target building rendering.

[0022] In one example, the training dataset may consist of multiple architectural project samples. Each architectural project sample contains at least the block diagram, line drawing, model diagram, style diagram, and final rendered image of the same architectural object at different expression stages. The images in the training dataset may be stored in common image formats, such as jpg, jpeg, png, webp, or bmp. PNG format is preferred for storing block diagrams, line drawings, and model diagrams with strong structural information, while jpg or png format is preferred for storing style diagrams and rendered images.

[0023] Furthermore, the target conditional control network in step S2 is constructed using a multi-stage cascaded training method. Specifically, the first conditional control sub-network is trained first using the first and second training data, the second conditional control sub-network is trained using the second and third training data, the third conditional control sub-network is trained using the third, fourth, and fifth training data, and finally the fourth conditional control sub-network is trained using the preprocessed results of the fifth training data and the original fifth training data. This ensures that the entire network can sequentially complete the generation process from block diagram to line drawing, from line drawing to model diagram, from model diagram and style diagram to low-resolution rendered image, and from low-resolution rendered image to high-resolution architectural rendering during the inference stage.

[0024] Furthermore, the input reference image corresponding to the building to be generated in step S3 includes at least the block diagram and style diagram of the building to be generated. During inference, the block diagram of the building to be generated is first fed into the trained first conditional control sub-network, which outputs the target line drawing; then the target line drawing is fed into the trained second conditional control sub-network, which outputs the target model diagram; subsequently, the target model diagram and the style diagram of the building to be generated are jointly input into the trained third conditional control sub-network, which outputs the target low-resolution rendering effect diagram; finally, the target low-resolution rendering effect diagram is input into the trained fourth conditional control sub-network, which outputs the target building rendering diagram.

[0025] It should be noted that the phased training refers to breaking down the architectural rendering generation task into multiple sub-tasks with clear intermediate representations and training them separately, thereby reducing structural distortion and style drift problems that occur during single-stage end-to-end generation.

[0026] It should be noted that the first condition control subnetwork, the second condition control subnetwork, the third condition control subnetwork, and the fourth condition control subnetwork in the target condition control network preferably adopt the ControlNet architecture.

[0027] It is understandable that by breaking down the architectural rendering generation task into multiple intermediate stages with professional semantics and introducing corresponding input reference images for conditional control at each stage, this invention can better balance structural rationality, stylistic consistency and image detail quality compared to methods that rely solely on text prompts or weak conditional inputs for generation.

[0028] In this embodiment, by constructing a multi-stage condition-controlled generation process, the generation process of architectural renderings is simultaneously constrained by the relationships between volumes, line drawing structures, model information, and style information, thereby improving the image quality of architectural renderings.

[0029] In another preferred embodiment, the training dataset includes at least a first training data, a second training data, a third training data, a fourth training data, and a fifth training data corresponding to the same building object.

[0030] In one example, the training dataset can be constructed from manually compiled historical architectural design project data, or from publicly available architectural image datasets, a company's own design database, or multi-view images rendered from 3D models. When constructing the training dataset, several architectural objects can be selected first, and the same or similar viewpoint parameters, perspective parameters, and composition range can be uniformly set for each architectural object. Then, block diagrams, line drawings, model diagrams, style diagrams, and rendered effect diagrams corresponding to each architectural object can be generated respectively, to ensure consistency in spatial location, composition range, and design content among multiple types of training data corresponding to the same architectural object.

[0031] Preferably, the number of samples in the training dataset can be between 5,000 and 100,000; the ratio of the training set, validation set, and test set can be 8:1:1 or 7:2:1. In one specific embodiment, 80% of the samples can be used as the training set, 10% as the validation set, and 10% as the test set. The validation set is used to monitor the model convergence and perform hyperparameter tuning, while the test set is used to evaluate the generalization performance of the finally trained target conditional control network.

[0032] It should be noted that the "correspondence of the same building object" means that the first to fifth training data describe different expressions of the same building scheme, the same scene, or the same perspective, rather than different building images that are unrelated to each other.

[0033] It is understandable that the training dataset can come not only from real projects, but also be automatically generated by software programs or manual processes. The key is that there are learnable correspondences between multiple types of training data under the same building object.

[0034] In this embodiment, by constructing multiple types of training data corresponding to the same building object, the conditional control subnetworks at each stage can learn the mapping relationship between adjacent expression stages, thereby providing a reliable data foundation for subsequent multi-stage cascade generation.

[0035] In another preferred embodiment, the first training data is a block diagram, the second training data is a line drawing, the third training data is a model diagram, the fourth training data is a style diagram, and the fifth training data is a rendered effect diagram.

[0036] In one example, the block diagram is an image reflecting the overall volume relationship and basic spatial outline of the building, which may include at least one of the following: the main building outline, the changes in the number of floors, the roof outline, the staggered height relationship, and the relationship between the main and auxiliary blocks; the block diagram is preferably a monochrome or minimal color block image, and may be stored in jpg or png format.

[0037] The line drawing is an image that reflects at least one of the following: building outline, facade joint lines, main component boundaries, door and window opening boundaries, eaves lines, and curtain wall grid lines. It can be in black and white line drawing format, grayscale line drawing format, or color line drawing format with a small number of auxiliary marks, and is preferably stored in PNG format.

[0038] The model diagram is an image that reflects the three-dimensional form of the building, spatial perspective, light and shadow levels, material partition outlines, or basic surface transition relationships. It can be composed of at least one of gray model rendering, simplified model coloring, and surface relationship diagram, and is preferably stored in png or jpg format.

[0039] The style image is a reference image used to characterize the target design style. It may include at least one of the following information: color style, material texture, lighting atmosphere, architectural vocabulary, and compositional style. It is preferably stored in jpg or png format.

[0040] The rendered image is the target output image, which can be a high-quality architectural image that includes complete materials, lighting, environmental atmosphere, sky background, foreground background and detailed components, preferably stored in JPG or PNG format.

[0041] It should be noted that the model diagram is not limited to a single form of expression. It can be a gray model, wireframe shading diagram, simplified material representation diagram, or an intermediate structure image expressed in the form of a normal diagram or depth diagram, output by 3D modeling software.

[0042] It should be noted that the style image can be a reference rendering of a similar building, or at least one of the following: a material concept board, a color scheme image, a photographic reference image, or a design style collage.

[0043] Understandably, block diagrams, line drawings, model diagrams, style diagrams, and rendered diagrams correspond to different stages in the architectural design expression process, providing structural and stylistic constraints in a progressive manner from coarse to fine.

[0044] It is understandable that the boundaries between various types of training data are not absolutely fixed, and those skilled in the art can make equivalent substitutions for the specific representations of line drawings and model diagrams according to specific business scenarios.

[0045] In this embodiment, by specifically defining the different types of training data, the training objectives and conditional constraints at each stage are made clearer, which helps to improve the interpretability and feasibility of the entire target conditional control network.

[0046] In yet another preferred embodiment, the target condition control network includes a first condition control subnetwork, a second condition control subnetwork, a third condition control subnetwork, and a fourth condition control subnetwork.

[0047] In one example, the first, second, and third conditional control subnetworks all employ a ControlNet architecture based on a latent space diffusion model. Each ControlNet architecture may include: a text encoder, a variational autoencoder (VAE), a denoising U-Net backbone network, and a conditional control branch network connected in parallel with the denoising U-Net backbone network. The conditional control branch network receives the input reference image for the current stage and injects corresponding conditional features into each resolution level of the denoising U-Net backbone network to guide the diffusion denoising process.

[0048] Preferably, the denoising U-Net backbone network can adopt a four-level downsampling and four-level upsampling structure, with the basic number of channels set to 320, and the number of channels in subsequent levels set to 640, 1280 and 1280 respectively; each level can include at least two residual blocks, and a cross-attention layer is introduced in the low-to-medium resolution layer to improve the ability to fuse structural and semantic information.

[0049] The fourth conditional control subnetwork can be a conditional control network based on diffusion super-resolution, with a low-resolution rendered image as input and a high-resolution architectural rendering as output. Preferably, the fourth conditional control subnetwork also includes an encoder, a U-Net backbone, and a conditional control branch, wherein the conditional control branch receives the low-resolution rendered image as a conditional image, and the backbone network outputs the denoising result corresponding to the high-resolution rendered image.

[0050] It should be noted that, in the first conditional control sub-network, the constraint condition of ControlNet is preferably a block diagram; in the second conditional control sub-network, the constraint condition of ControlNet is preferably a line drawing diagram; in the third conditional control sub-network, the constraint condition of ControlNet is preferably a model diagram and a style diagram; and in the fourth conditional control sub-network, the constraint condition of ControlNet is preferably a low-resolution rendering effect diagram.

[0051] It should be noted that in the third conditional control sub-network, the model graph can be used as the structural control condition and the style graph as the style control condition. The two are extracted through independent encoding branches to extract structural features and style features, and then fused in the cross-attention layer or feature splicing layer.

[0052] Understandably, the core of the ControlNet architecture lies in maintaining the generative capabilities of the pre-trained diffusion model while introducing conditional control branches to ensure that the output more strictly follows the geometric structure, layout relationships, and style information corresponding to the input reference image.

[0053] In this embodiment, by dividing the target condition control network into four condition control sub-networks with distinct functions, the architectural rendering generation task can be completed step by step in stages, thereby reducing the difficulty of single-stage generation.

[0054] In yet another preferred embodiment, the step of training and generating the target conditional control network in stages based on the training dataset includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are preprocessed to obtain first preprocessed training data, second preprocessed training data, third preprocessed training data, fourth preprocessed training data, and fifth preprocessed training data. The first preprocessed training data is input into the first conditional control subnetwork, and the second preprocessed training data is used as the output target of the first conditional control subnetwork to train the first conditional control subnetwork. The second preprocessed training data is input into the second conditional control subnetwork, and the third preprocessed training data is used as the output target of the second conditional control subnetwork to train the second conditional control subnetwork. The third preprocessed training data and the fourth preprocessed training data are input into the third conditional control subnetwork, and the fifth preprocessed training data is used as the output target of the third conditional control subnetwork to train the third conditional control subnetwork. The fifth preprocessed training data is input into the fourth conditional control subnetwork, and the fifth training data is used as the output target of the fourth conditional control subnetwork to train the fourth conditional control subnetwork. The target conditional control network is obtained based on the trained first conditional control subnetwork, second conditional control subnetwork, third conditional control subnetwork, and fourth conditional control subnetwork.

[0055] In one example, the first, second, third, and fourth conditional control subnetworks are trained independently. The training objective of the first conditional control subnetwork is to learn the mapping relationship from the block diagram to the line drawing. During training, the first preprocessed training data is used as conditional input, and the second preprocessed training data is encoded as a latent variable and noise is added. The ControlNet branch extracts the contour and volume relationship features of the block diagram, and the U-Net backbone predicts the noise component at the corresponding time step to achieve the line drawing generation task. The training objective of the second conditional control subnetwork is to learn the mapping relationship from the line drawing to the model diagram. During training, the second preprocessed training data is used as conditional input, and the third preprocessed training data is used as the output target. The ControlNet branch extracts the contour lines, door and window boundaries, and component relationships from the line drawing and guides the U-Net backbone to generate a model diagram with surface relationships and spatial light and shadow information. The training objective of the third conditional control subnetwork is to learn the mapping relationship from model maps and style maps to low-resolution rendered images. During training, features can be extracted from the third and fourth preprocessed training data respectively. Model map features are used to maintain architectural structure and perspective relationships, while style map features provide constraints on color, material, and atmosphere. These features are then fused to generate a low-resolution rendered image corresponding to the fifth preprocessed training data. The training objective of the fourth conditional control subnetwork is to learn the mapping relationship from low-resolution rendered images to high-resolution rendered images. During training, the fifth preprocessed training data is used as input, and the original fifth training data is used as the output target. A diffusion super-resolution process is used to restore edge details, material textures, shadow transitions, and background elements.

[0056] Preferably, the first, second, and third conditional control subnetworks can be trained for 100 to 300 epochs each, and the fourth conditional control subnetwork can be trained for 80 to 200 epochs; alternatively, in terms of iteration steps, the first three subnetworks can be trained for 100,000 to 300,000 steps each, and the fourth conditional control subnetwork can be trained for 80,000 to 200,000 steps. Preferably, the training optimizer can be AdamW, the learning rate can be set to 1e-5, and the batch size can be set to 16. Preferably, the loss function includes at least the diffuse noise prediction loss, which can be the mean squared error (MSE) loss.

[0057] Preferably, the total loss function of the first conditional control subnetwork can be expressed as: L1_total = α1·Lnoise + β1·Ledge + γ1·Lrec; The total loss function of the second conditional control subnetwork can be expressed as: L2_total = α2·Lnoise + β2·Lstr + γ2·Lrec; The total loss function of the third conditional control subnetwork can be expressed as: L3_total = α3·Lnoise + β3·Lperc + γ3·Lstyle + δ3·Lrec; The loss function of the fourth conditional control subnetwork can be expressed as: L4_total = α4·Lnoise + β4·Lsr + γ4·Lperc; Where Lnoise is the diffusion denoising loss, Ledge is the line drawing edge consistency loss, Lstr is the structure consistency loss, Lstyle is the style consistency loss, Lrec is the pixel reconstruction loss, Lsr is the super-resolution detail recovery loss, Lperc is the perceptual loss, and α1, β1, γ1, α2, β2, γ2, α3, β3, γ3, δ3, α4, β4 and γ4 are all weight coefficients greater than 0.

[0058] In one specific embodiment, α1=1.0, β1=0.5, γ1=0.2; α2=1.0, β2=0.5, γ2=0.2; α3=1.0, β3=0.3, γ3=0.3, δ3=0.2; α4=1.0, β4=0.4, γ4=0.3.

[0059] It should be noted that the four conditional control subnetworks are preferably trained separately so that each subnetwork focuses on a single mapping task, improving convergence stability and task interpretability. It should also be noted that after each stage of training is completed, each subnetwork can be validated separately, and then the validated subnetworks can be cascaded in a predetermined order to form the target conditional control network.

[0060] It is understood that the phased training method does not preclude subsequent joint fine-tuning of the cascaded overall network; in some implementations, the entire target conditional control network can be fine-tuned end-to-end using a joint loss function after all four sub-networks have been trained individually.

[0061] It is understood that the above loss function and training parameters are merely preferred examples, and those skilled in the art can make adaptive adjustments based on the data scale, hardware resources, and the style complexity of the target image.

[0062] In this embodiment, by training the four conditional control subnetworks separately and setting corresponding loss functions and optimization strategies, the generation tasks at each stage can converge stably, thereby improving the training efficiency and generation reliability of the entire target conditional control network.

[0063] In yet another preferred embodiment, the preprocessing of the first training data, the second training data, the third training data, the fourth training data, and the fifth training data includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are respectively subjected to resolution reduction processing to obtain the first preprocessed training data, the second preprocessed training data, the third preprocessed training data, the fourth preprocessed training data, and the fifth preprocessed training data.

[0064] In one example, the first to fifth training data can be uniformly scaled to a preset resolution, such as 256×256 pixels, 384×384 pixels, or 512×512 pixels. In a specific embodiment, it is preferable to reduce the first to fifth training data to 256×256 pixels to reduce the training cost of the first three ControlNet sub-networks.

[0065] Preferably, the resolution reduction processing can be performed using at least one of bilinear interpolation, bicubic interpolation, Lanczos interpolation, or area interpolation. For line drawings and volumetric drawings with relatively clear structural boundaries, nearest neighbor interpolation or bilinear interpolation is preferred to reduce boundary blurring. For style drawings and rendered effect drawings, bicubic interpolation or Lanczos interpolation is preferred to retain more texture features.

[0066] Before downscaling, images can be cropped, normalized, mapped to pixel value ranges, and randomly enhanced. For example, pixel values ​​can be normalized to the range of [0,1]; training samples can be randomly flipped, randomly brightened, and randomly contrasted to improve the model's robustness to style changes and scene noise.

[0067] It should be noted that uniformly down-resolution processing of multiple types of training data helps to maintain consistency in spatial scale between the first and fifth pre-processed training data, thereby reducing the difference in input and output scales during training at each stage.

[0068] It should be noted that the fourth conditional control subnetwork uses the fifth preprocessed training data as input and the fifth training data as output target. Essentially, it performs a super-resolution learning process to recover a high-resolution rendering image from a low-resolution rendering image.

[0069] It is understandable that, in addition to reducing resolution, the preprocessing stage may also include at least one of the following processing operations: image denoising, edge enhancement, color normalization, perspective correction, and background region mask generation.

[0070] In this embodiment, by uniformly down-resolution processing of multiple types of training data before training the first three levels of the network, the training cost is reduced while ensuring structural control capability. Furthermore, the high-resolution details are restored through the fourth conditional control sub-network, thus balancing generation efficiency and output quality.

[0071] In another preferred embodiment, the input reference image corresponding to the building to be generated includes a block diagram of the building to be generated and a style diagram of the building to be generated. The input reference image corresponding to the building to be generated is input into the target condition control network, which outputs a rendering of the target building, including: Input the block diagram of the building to be generated into the trained first conditional control subnetwork, and output the target line drawing; The target line drawing is input into the trained second conditional control subnetwork, and the target model diagram is output. The target model image and the style image of the building to be generated are input into the trained third conditional control subnetwork, and the target low-resolution rendering effect image is output. The target low-resolution rendered image is input into the trained fourth conditional control subnetwork, which outputs the target architectural rendering.

[0072] In one example, during the inference phase, the user can first input the block diagram and style diagram of the building to be generated. The block diagram can be obtained by at least one of the following methods: CAD planar outline conversion, SketchUp block screenshot, BIM model rasterization output, or exporting images from parametric modeling software; the style diagram can be selected by the designer from reference cases, or it can be automatically retrieved and matched by the system from the style library.

[0073] See Figure 2 This is a flowchart illustrating another preferred embodiment of a method for generating architectural renderings based on multi-stage conditional control, provided in the first aspect of the present invention. After receiving the block diagram, the system inputs it into a trained first conditional control subnetwork. In this stage, the ControlNet branch in the first conditional control subnetwork extracts the contour, hierarchy, and height-to-volume relationships from the block diagram and outputs a target line drawing through a constrained diffusion denoising process. The target line drawing preferably includes at least one of the following: the main architectural outline, the main door and window boundary lines, component separation lines, and basic facade composition lines.

[0074] Subsequently, the target line drawing is input into the trained second conditional control subnetwork. The ControlNet branch in the second conditional control subnetwork extracts the boundaries and geometric relationships in the line drawing and outputs the target model image. The target model image is preferably represented as an intermediate image with three-dimensional surface relationships, perspective relationships, simplified shading, and a sense of volume.

[0075] The target model image and the style image of the building to be generated are then input into the trained third conditional control subnetwork. The third conditional control subnetwork can first encode the target model image and the style image separately, and then perform feature fusion in the fusion layer to generate a low-resolution rendering image of the target. The target low-resolution rendering image can preferably be 256×256 or 512×512 resolution.

[0076] Finally, the target low-resolution rendered image is input into the trained fourth conditional control subnetwork to perform high-resolution reconstruction of the low-resolution image, outputting the target architectural rendering. The resolution of the target architectural rendering can preferably be 1024×1024, 1536×1536, or 2048×2048. See also... Figure 3 This is a schematic diagram of a block diagram, line drawing, model diagram, target style diagram and rendering diagram, which is another preferred embodiment of the architectural rendering generation method based on multi-stage condition control provided in the first aspect of the present invention.

[0077] It should be noted that in some implementations, the system can also set manual confirmation nodes after each stage of output. For example, the designer confirms the target line drawing after the first stage output before continuing with the second and third stages of generation.

[0078] It is understood that the reasoning stage of this invention is not only applicable to single generation, but also to the rapid comparison and selection scenario in the scheme stage. That is, under the premise of fixed block diagram, multiple architectural renderings with different style directions can be generated in batches by replacing different style diagrams, adding text prompts or reference images.

[0079] In this embodiment, by sequentially outputting the target line drawing, target model drawing, target low-resolution rendering, and target architectural rendering during the inference stage, the generation process can reflect the transmission relationship between architectural structure, model information, and style information layer by layer, thereby improving the controllability, consistency, and interpretability of the generated results.

[0080] A second aspect of the present invention provides an architectural rendering generation apparatus based on multi-stage condition control, used to implement the architectural rendering generation method based on multi-stage condition control described in any of the first aspects above. See also... Figure 4 The diagram shown is a structural block diagram of a preferred embodiment of an architectural rendering generation device based on multi-stage condition control provided in the second aspect of the present invention. The device includes: Training set acquisition module 11 is used to acquire the training dataset; Conditional control network generation module 12 is used to train and generate a target conditional control network in stages based on the training dataset; The target building rendering output module 13 is used to input the input reference image corresponding to the building to be generated into the target condition control network and output the target building rendering.

[0081] It should be noted that the architectural rendering generation device based on multi-stage condition control provided in the second aspect of the present invention can realize all the processes of the architectural rendering generation method based on multi-stage condition control described in the first aspect. The functions and technical effects of each module and unit in the device are the same as those of the architectural rendering generation method based on multi-stage condition control described in the first aspect, and will not be repeated here.

[0082] A third aspect of the present invention also provides a computer-readable storage medium, the computer-readable storage medium including a stored computer program; wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the architectural rendering generation method based on multi-stage condition control described in any of the first aspects above.

[0083] The fourth aspect of the present invention also provides a terminal device, see [link to documentation]. Figure 5The diagram shown is a structural block diagram of a preferred embodiment of a terminal device provided in the fourth aspect of the present invention. The terminal device includes a processor 10, a memory 20, and a computer program stored in the memory 20 and configured to be executed by the processor 10. When the processor 10 executes the computer program, it implements a method for generating architectural renderings based on multi-stage condition control as described in any of the above embodiments.

[0084] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2, ...), and the one or more modules / units are stored in the memory 20 and executed by the processor 10 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0085] The processor 10 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 10 may be any conventional processor. The processor 10 is the control center of the terminal device, connecting various parts of the terminal device through various interfaces and lines.

[0086] The memory 20 mainly includes a program storage area and a data storage area. The program storage area can store the operating system, applications required for at least one function, etc., while the data storage area can store related data, etc. Furthermore, the memory 20 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard drive, a smart media card (SMC), a secure digital card (SD), and a flash card, or other volatile solid-state storage devices.

[0087] It should be noted that the aforementioned terminal device may include, but is not limited to, processors and memory. Those skilled in the art will understand that the above content is merely an example describing the structure of the terminal device and does not constitute a limitation on the structure of the aforementioned terminal device. The aforementioned terminal device may include more or fewer components than those described above, or combine certain components, or different components.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary hardware platforms, and of course, it can also be implemented entirely by hardware. Based on this understanding, all or part of the technical solution of the present invention that contributes to the background technology can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0089] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for architectural rendering based on multi-stage conditional control, characterized in that, include: Obtain the training dataset; Based on the training dataset, a target conditional control network is generated through phased training. The target condition control network takes the input reference image corresponding to the building to be generated as input and outputs the target building rendering.

2. The method for building rendering generation based on multi-stage conditional control according to claim 1, characterized in that, The training dataset includes at least a first training data, a second training data, a third training data, a fourth training data, and a fifth training data corresponding to the same building object.

3. The method for building rendering generation based on multi-stage conditional control according to claim 2, characterized in that, The first training data is a block diagram, the second training data is a line drawing, the third training data is a model diagram, the fourth training data is a style diagram, and the fifth training data is a rendered effect diagram.

4. The method of claim 3, wherein the method further comprises: The target condition control network includes a first condition control subnetwork, a second condition control subnetwork, a third condition control subnetwork, and a fourth condition control subnetwork.

5. The method for generating architectural renderings based on multi-stage condition control as described in claim 4, characterized in that, The step of generating the target conditional control network through phased training based on the training dataset includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are preprocessed to obtain first preprocessed training data, second preprocessed training data, third preprocessed training data, fourth preprocessed training data, and fifth preprocessed training data. The first preprocessed training data is input into the first conditional control subnetwork, and the second preprocessed training data is used as the output target of the first conditional control subnetwork to train the first conditional control subnetwork. The second preprocessed training data is input into the second conditional control subnetwork, and the third preprocessed training data is used as the output target of the second conditional control subnetwork to train the second conditional control subnetwork. The third preprocessed training data and the fourth preprocessed training data are input into the third conditional control subnetwork, and the fifth preprocessed training data is used as the output target of the third conditional control subnetwork to train the third conditional control subnetwork. The fifth preprocessed training data is input into the fourth conditional control subnetwork, and the fifth training data is used as the output target of the fourth conditional control subnetwork to train the fourth conditional control subnetwork. The target conditional control network is obtained based on the trained first conditional control subnetwork, second conditional control subnetwork, third conditional control subnetwork, and fourth conditional control subnetwork.

6. The method for generating architectural renderings based on multi-stage condition control as described in claim 5, characterized in that, The preprocessing of the first training data, the second training data, the third training data, the fourth training data, and the fifth training data includes: The first training data, the second training data, the third training data, the fourth training data, and the fifth training data are respectively subjected to resolution reduction processing to obtain the first preprocessed training data, the second preprocessed training data, the third preprocessed training data, the fourth preprocessed training data, and the fifth preprocessed training data.

7. The method for generating architectural renderings based on multi-stage condition control as described in claim 6, characterized in that, The input reference image corresponding to the building to be generated includes the block diagram of the building to be generated and the style diagram of the building to be generated. The step of inputting the input reference image corresponding to the building to be generated into the target condition control network and outputting the target building rendering includes: Input the block diagram of the building to be generated into the trained first conditional control subnetwork, and output the target line drawing; The target line drawing is input into the trained second conditional control subnetwork, and the target model diagram is output. The target model image and the style image of the building to be generated are input into the trained third conditional control subnetwork, and the target low-resolution rendering effect image is output. The target low-resolution rendered image is input into the trained fourth conditional control subnetwork, which outputs the target architectural rendering.

8. A building rendering generation device based on multi-stage condition control, used to implement the building rendering generation method based on multi-stage condition control as described in any one of claims 1 to 7, the device comprising: The training set acquisition module is used to acquire the training dataset; The conditional control network generation module is used to train and generate a target conditional control network in stages based on the training dataset. The target building rendering output module is used to input the input reference image corresponding to the building to be generated into the target condition control network and output the target building rendering.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform a method for generating architectural renderings based on multi-stage condition control as described in any one of claims 1 to 7.

10. A terminal device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a method for generating architectural renderings based on multi-stage condition control as described in any one of claims 1 to 7.