Building image generation method and device, and medium

By preprocessing and extracting features from architectural design images, combining the diffusion model with the applied style model, high-quality architectural design images are generated. This solves the style confusion and image blurring problems generated by the diffusion model under multimodal input, and achieves accurate style output and detail presentation.

CN120807715AActive Publication Date: 2025-10-17CHANGAN UNIV

Patent Information

Application Number
CN202511013960.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-17
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing diffusion models have problems with style confusion and image blur when generating architectural renderings, especially under multimodal input conditions, it is difficult to accurately understand user intentions and generate high-quality architectural design images.

Method used

By obtaining the mask image and the original architectural design image, preprocessing is performed to determine the local architectural design image and the spliced ​​architectural design image, and text features and image features are extracted. The pre-trained application style model is used to generate guidance conditions, which are combined with the diffusion model for inference and post-processing to generate the target architectural design image.

Benefits of technology

It achieves the generation of high-quality architectural design images under multimodal input conditions, solves the problems of style confusion and image blur, and ensures that the generated results meet the architectural style and details expected by users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807715A_ABST
    Figure CN120807715A_ABST
Patent Text Reader

Abstract

The invention discloses a building image generation method and device and a medium, and relates to the technical field of building image generation, and the method comprises the steps: obtaining an original building design image and a mask image; preprocessing the original building design image according to the mask image to obtain a local building design image, splicing the building design images, and synthesizing the mask image according to the local building design image to obtain a synthesized mask image; extracting text features and image features of the local building design image; inputting the text features and the image features into a pre-trained application style model, and outputting a second guide condition; inputting the spliced architectural design image, the first guide condition and the second guide condition into a diffusion model to obtain a reasoning architectural design image; and post-processing the reasoning building design image according to the mask image to obtain a target building design image. The method has the advantages of supporting multi-modal input and generating high-quality building images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of building image generation, in particular to a building image generation method, device and medium. BACKGROUND

[0002] Building image generation technology is mainly based on deep learning, computer vision and generative models, which can automatically generate or assist in designing building renderings according to input conditions such as text descriptions and sketches.

[0003] Diffusion models, as an important generative model, can support multi-modal inputs such as text, images and style references, thereby generating building renderings that meet user requirements. For example, when the input text description is a villa, the image is a hand-drawn sketch, and the style is modern, the diffusion model can generate the corresponding building rendering.

[0004] However, the current diffusion model still has problems such as style disorder (such as combining Chinese-style roof with glass curtain wall) and image blurring (such as insufficient image detail levels) when generating building renderings, resulting in poor image quality. SUMMARY

[0005] The present application aims to solve one of the technical problems in the related art to some extent. To this end, the present application provides a building image generation method, device and medium, which has the advantages of supporting multi-modal input and generating high-quality building images.

[0006] In order to achieve the above purpose, as a first aspect of the present application, a building image generation method is provided, wherein the method comprises: Obtaining an original building design image and a mask image, the mask image comprising an occlusion region and a target region, the shape of the target region matching at least one local image block in the original building design image; According to the mask image, the original building design image is preprocessed to obtain a local building design image and a spliced building design image, and the mask image is synthesized according to the local building design image to obtain a synthesized mask image; wherein the synthesized mask image is a first guide condition; Extracting text features and image features of the local building design image; Inputting the text features and the image features into a pre-trained application style model to output a second guide condition, the second guide condition corresponding to a specific building style; Inputting the spliced building design image, the first guide condition and the second guide condition into a diffusion model to obtain an inference building design image; According to the mask image, the inference building design image is post-processed to obtain a target building design image.

[0007] Optionally, the text feature and the image feature are input into a pre-trained application style model to output a second guidance condition, including: The text feature and the image feature are fused to generate a fusion feature; The fusion feature is input into the application style model to output a specific building style.

[0008] Optionally, the specific building style is selected from style information of a building image, texture information of the building image, environment information of the building image, color and material information of the building image, and form and structure information of the building image.

[0009] Optionally, the spliced building design image, the first guidance condition and the second guidance condition are input into a diffusion model to obtain an inferred building design image, including: The spliced building design image is input into a VAE encoder to output an encoding feature of the spliced building design image in a latent space; The encoding feature in the latent space, the first guidance condition and the second guidance condition are input into a U-Net model to output an encoding feature of the spliced building design image in the latent space based on the first guidance condition and the second guidance condition; The encoding feature of the latent space based on the first guidance condition and the second guidance condition is input into a VAE decoder to obtain the inferred building design image.

[0010] Optionally, the encoding feature in the latent space, the first guidance condition and the second guidance condition are input into the U-Net model to output an image feature of the spliced building design image in the latent space based on the first guidance condition and the second guidance condition, including: The encoding feature in the latent space is input into the U-Net model as a basic carrier; The first guidance condition and the second guidance condition are input into the U-Net model as control conditions to guide the U-Net model to infer the basic carrier; An image feature of the spliced building design image in the latent space based on the first guidance condition and the second guidance condition is output.

[0011] Optionally, the original building design image includes a reference building design image and a redrawn building design image, the mask image includes a reference mask and a redrawn mask, the reference mask has the same size as the reference building design image, the redrawn mask has the same size as the redrawn building design image, and the local building design image includes a local reference building design image and a local redrawn building design image. The pre-processing of the original architectural design image according to the mask image comprises: cropping the reference architectural design image according to the target region of the reference mask to output a local reference architectural design image; cropping the redrawn architectural design image according to the target region of the redrawn mask to output a local redrawn architectural design image; scaling the local reference architectural design image according to the size of the local redrawn architectural design image to output an aligned reference architectural design image; splicing the aligned reference architectural design image and the local redrawn architectural design image in a row or column direction to output a spliced architectural design image; splicing the target region of the redrawn mask and the aligned reference architectural design image in a row or column direction to output a synthesized mask image.

[0012] Optionally, the post-processing of the inference architectural design image according to the mask image to obtain a target architectural design image comprises: cropping the inference architectural design image according to the target region of the redrawn mask to output a cropped inference architectural design image; filling the cropped inference architectural design image into a local image block of the redrawn architectural design image matching the target region of the redrawn mask to output a target architectural design image.

[0013] Optionally, the extraction of the text features and the image features of the local architectural design image comprises: in the case that the local reference architectural design image comprises text content, inputting the local reference architectural design image into a Clip model to output text features of the local reference architectural design image; inputting the local reference architectural design image into a Clip visual model to output image features of the local reference architectural design image; in the case that the local reference architectural design image only comprises image content, inputting the local reference architectural design image into a language model to output text features of the cropped reference architectural design image.

[0014] As a second aspect of the present application, an electronic device is provided, comprising: one or more processors; a memory having one or more computer programs stored thereon, when the one or more computer programs are executed by the one or more processors, the one or more processors implement the generation method provided by the first aspect of the present application.

[0015] In addition, as a third aspect of the present application, a computer readable medium is provided, which stores a computer program, wherein the computer program is executed by a processor to implement the generation method of the first aspect of the present application.

[0016] The generation method of the building image provided by the present application first determines the local building design image and the spliced building design image by means of the mask image, and then synthesizes the mask image according to the local building design image to obtain the first guide condition of the diffusion model, that is, the synthesized mask image; secondly, the text features and image features of the local building design image are extracted, and the text features and image features are fused and input into the pre-trained application style model to obtain the second guide condition of the diffusion model; subsequently, the spliced building design image is taken as the main input of the diffusion model, and the first guide condition and the second guide condition are taken as the control condition of the diffusion model, and the diffusion model is used to infer the spliced building design image to obtain the inferred building design image; finally, the inferred building design image is subjected to post-processing operations such as cropping and padding to generate the target building design image. The generation method provided by the present application can lock the target area by the first guide condition, reduce the influence of the non-target area on the generation result in the generation process, and the second guide condition can guide the diffusion model to generate an accurate building design image, and the double guide conditions as the control condition solve the problems of style disorder, image blur and poor image quality of the building image generated by the diffusion model.

[0017] The features and advantages of the present application will be described in detail in the following specific embodiments and drawings. The best mode or means of the present application will be fully described in conjunction with the drawings, but it is not a limitation on the technical solutions of the present application. In addition, the features, elements and components appearing in each of the following text and drawings are multiple, and different symbols or numbers are marked for convenience of representation, but all represent the same or similar structure or function parts. BRIEF DESCRIPTION OF DRAWINGS

[0018] The present application will be further described below in conjunction with the drawings: Figure 1 A flowchart of the generation method of the building image provided by the present application; Figure 2 A flowchart of one embodiment of the generation method step S140 provided by the present application; Figure 3 A flowchart of one embodiment of the generation method step S150 provided by the present application; Figure 4 A flowchart of one embodiment of the generation method step S152 provided by the present application; Figure 5 A flowchart of one embodiment of the generation method step S120 provided by the present application; Figure 6 An embodiment flow chart of the generating method step S160 provided by the present application; Figure 7 An embodiment flow chart of the generating method step S130 provided by the present application; Figure 8 An implementation flow chart of the generating method provided by the present application; Figure 9 Another implementation flow chart of the generating method provided by the present application; Figure 10 A module diagram of an electronic device provided by the present application; Figure 11 A schematic diagram of a computer readable medium provided by the present application.

[0019] Explanation of reference numerals Among them, 101, processor; 102, memory; 103, I / O interface; 104, bus. DETAILED DESCRIPTION

[0020] Embodiments of the present application will be described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. Based on the embodiments in the embodiments, it is intended to explain the present application, and cannot be understood as a limitation of the present application.

[0021] In this specification, "one embodiment" or "an example" or "an example" means that the specific features, structures or characteristics described in connection with the embodiment itself can be included in at least one embodiment of the present disclosure. The appearance of the phrase "in one embodiment" at various places in the specification does not necessarily refer to the same embodiment.

[0022] Diffusion model as a generative model is widely used in the field of building image generation technology. A single input diffusion model can generate building design images that meet the input conditions, but cannot meet the user's demand for diversity and stylization of generated images. While diffusion models supporting multi-modal input (such as text, sketches, style references, etc.) have difficulty in semantic alignment across modalities (such as the ambiguity of sketches and the accuracy of text descriptions, which prevent the model from accurately understanding the relationship between text and sketches, and the multiple types of style coverage, which prevent the model from accurately understanding user intent), feature attention imbalance (such as the diffusion process may focus on local irrelevant features or ignore key structures), and other reasons, resulting in building design images generated with poor image quality, such as style confusion, architectural structure deviation, and image blur.

[0023] In view of this, in order to meet the user's requirement for generating diversified architectural design images and solve the problem of poor quality of architectural design images generated by the diffusion model for multi-modal input, as a first aspect of the present application, a method for generating an architectural image is provided, as shown in Figure 1 The method comprises the following steps: In step S110, an original architectural design image and a mask image are obtained, the mask image comprising an occlusion region and a target region, the shape of the target region matching at least one local image block in the original architectural design image; In step S120, the original architectural design image is preprocessed according to the mask image to obtain a local architectural design image and a spliced architectural design image, and the mask image is synthesized according to the local architectural design image to obtain a synthesized mask image; wherein the synthesized mask image is a first guide condition; In step S130, text features and image features of the local architectural design image are extracted; In step S140, the text features and the image features are input into a pre-trained application style model to output a second guide condition, the second guide condition corresponding to a specific architectural style; In step S150, the spliced architectural design image, the first guide condition and the second guide condition are input into a diffusion model to obtain an inference architectural design image; In step S160, the inference architectural design image is post-processed according to the mask image to obtain a target architectural design image.

[0024] The building image generation method provided by the application first determines a local building design image and a spliced building design image by means of a mask image, then synthesizes the mask image according to the local building design image to obtain a first guide condition of a diffusion model, that is, a synthesized mask image; secondly, text features and image features of the local building design image are extracted, and the text features and the image features are input into a pre-trained application style model after being fused to obtain a second guide condition of the diffusion model; subsequently, the spliced building design image is taken as a main input of the diffusion model, and the first guide condition and the second guide condition are taken as control conditions of the diffusion model, and the diffusion model is used to infer the spliced building design image to obtain an inferred building design image; finally, the inferred building design image is subjected to post-processing operations such as cropping and padding to generate a target building design image. The generation method provided by the application can lock the target area by the first guide condition, reduce the influence of non-target areas on the generation result in the generation process, and solve the problem of feature attention imbalance of the diffusion model in the generation process; secondly, the second guide condition can output accurate style features, and solve the problem of difficult semantic alignment of multi-modal input, and the double guide conditions as the control conditions solve the problems of style disorder, image blur and poor image quality of the building image generated by the diffusion model.

[0025] According to steps S110-S160, the building image generation method provided by the application mainly includes operations such as preprocessing of an input building design image, feature extraction, feature fusion, guiding the diffusion model to infer by taking the first guide condition and the second guide condition as control conditions of the diffusion model, and post-processing of an inferred building design image.

[0026] The second guide condition output by step S140 is described in detail, and as an optional implementation manner of step S140, as shown in Figure 2 The text features and the image features are input into a pre-trained application style model to output the second guide condition, and the method comprises the following steps of: In step S141, the text features and the image features are fused to generate fused features; In step S142, the fused features are input into the application style model to output a specific building style.

[0027] It needs to be particularly pointed out that the application style model is a self-research model trained by a large amount of text information and image information related to buildings and accurate building styles, wherein the specific building style includes but is not limited to style information of the building image (such as Gothic style, modernist style, Huipai architecture, etc.), texture information of the building image (such as rough texture of stone, natural texture of wood, luster and smooth texture of metal, etc.), environment information of the building image (such as the building image in a natural environment by the sea, the building image in a historical block humanistic environment, etc.), color and material information of the building image (such as overall or local color matching of the building image, and building material, etc.), and form and structure information of the building image (such as the building outline gradually tapering from bottom to top, and the internal structure being designed as a central court, etc.).

[0028] The building image generation method provided by the present application mainly relies on a diffusion model to infer the input building image, and as an optional implementation manner of step S150, the entire inference process is illustrated as shown in the figure. Figure 3 As shown in the figure, the building design image, the first guide condition and the second guide condition are input into the diffusion model to obtain an inferred building design image, which comprises the following steps: In step S151, the spliced building design image is input into a VAE (Variational Autoencoder) encoder to output the encoding features of the spliced building design image in the latent space; In step S152, the encoding features of the latent space, the first guide condition and the second guide condition are input into a U-Net model to output the encoding features of the latent space based on the first guide condition and the second guide condition of the spliced building design image; In step S153, the encoding features of the latent space based on the first guide condition and the second guide condition are input into a VAE decoder to obtain the inferred building design image.

[0029] The diffusion model applied in the present application is described in detail below. Before the input image and the control condition are input into the U-Net model, the VAE encoder is used to encode the high-dimensional input spliced building design image into a low-dimensional latent space representation, which can significantly reduce the data storage and computing cost, making the U-Net model more efficient in the inference process; secondly, the VAE encoding can remove the noise and redundant information in the input spliced building design image, and extract more meaningful features, which helps the U-Net model to generate higher quality samples in the generation process; in addition, the VAE can generalize the input spliced building design image to a certain extent in the encoding process, enhancing the generalization ability of the U-Net model. Similarly, the VAE decoding after the U-Net model can improve the quality of the generated building design image, enhance the robustness of the diffusion model, and provide more flexible generation control.

[0030] After the input data is encoded by the VAE encoder, the main operation object and control condition of the U-Net model need to be determined before the input U-Net model, as an embodiment of step S152, as shown in the figure, the encoded features of the latent space, the first guide condition, and the second guide condition are input into the U-Net model, and the output spliced building design image based on the encoded features of the latent space under the first guide condition and the second guide condition comprises: Figure 4 In step S152a, the encoded features of the latent space are input into the U-Net model as a basic carrier; In step S152b, the first guide condition and the second guide condition are input into the U-Net model as control conditions together to guide the U-Net model to reason the basic carrier; In step S152c, the output spliced building design image based on the encoded features of the latent space under the first guide condition and the second guide condition.

[0031] Unlike the prior art, the present application uses two guide conditions as control conditions of the U-Net model, so that the U-Net model only generates corresponding image features for the target region.

[0032] The generation method provided by the present application not only makes corresponding improvements in diffusion reasoning, but also includes many operations in the preprocessing link. The original building design image provided by the present application includes a reference building design image and a redrawn building design image. The mask image includes a reference mask and a redrawn mask. The size of the reference mask is the same as that of the reference building design image, and the size of the redrawn mask is the same as that of the redrawn building design image. The local building design image includes a local reference building design image and a local redrawn building design image. As an optional embodiment of step S120, as shown in the figure, the original building design image is preprocessed according to the mask image to obtain a local building design image and a spliced building design image, and the mask image is synthesized according to the local building design image to obtain a synthesized mask image, comprising: Figure 5 In step S121, the reference building design image is cropped according to the target region of the reference mask, and a local reference building design image is output; In step S122, the redrawn building design image is cropped according to the target region of the redrawn mask, and a local redrawn building design image is output; In step S123, the local reference building design image is scaled based on the size of the local redrawn building design image, and an aligned reference building design image is output; ​​In step S124, the aligned reference architectural design image and the local redrawn architectural design image are spliced in a row or column direction, and a spliced architectural design image is output; In step S125, the target region of the redrawn mask and the aligned reference architectural design image are spliced in a row or column direction, and a synthesized mask image is output.

[0033] It should be noted that the preprocessing operation needs to be explained. The above embodiment gives that the image size of the reference architectural design image is different from the image size of the redrawn architectural design image, and the sizes of the local image blocks cropped according to the target regions of the corresponding masks are also different, so scaling and alignment operations are needed after cropping. In actual application, if local image blocks of the same size are cropped according to the target regions of the corresponding masks, scaling operation is not needed, and the spliced architectural design image can be directly spliced in a row or column direction. At this point, the preprocessing process ensures that only the region of the redrawn architectural design image that needs to be regenerated and the region of the reference architectural design image that carries the architectural style that needs to be generated are retained. Since the first guide condition with the mask effect needs to be used in the subsequent diffusion model inference, a synthesized mask with the same size as the spliced architectural design image needs to be generated. Considering that the spliced architectural design image includes the aligned reference architectural design image and the local redrawn architectural design image, and we only need to focus on the key features in the diffusion model inference, the synthesized mask is formed by splicing the target region of the redrawn mask and the aligned reference architectural design image, but only uses the size information of the aligned reference architectural design image. The aligned reference architectural design image in the synthesized mask is an occlusion region composed of all-black pixel values.

[0034] After the diffusion model completes the inference on the local architectural design image, the generation method provided by the application further includes a post-processing operation. As an optional implementation manner of the post-processing operation, as shown in Figure 6 The post-processing of the inferred architectural design image according to the mask image to obtain a target architectural design image includes: In step S161, the inferred architectural design image is cropped according to the target region of the redrawn mask, and a cropped inferred architectural design image is output. In step S162, the cropped inferred architectural design image is filled into the local image block of the redrawn architectural design image that matches the target region of the redrawn mask, and a target architectural design image is output.

[0035] In actual application, in order to further improve the image quality of the target architectural design image, correction, optimization and other post-processing operations can also be added.

[0036] It should be particularly noted that step S130 also needs to be explained. As an optional implementation manner of step S130, as shown in Figure 7As shown, the text features and image features of the extracted partial building design image are obtained, including: In step S131, in the case that the partial reference building design image includes text content, the partial reference building design image is input into a Clip (Contrastive Language-Image Pre-Training) model, and the text features of the partial reference building design image are output; In step S132, the partial reference building design image is input into a Clip visual model, and the image features of the partial reference building design image are output; In step S133, in the case that the partial reference building design image only includes image content, the partial reference building design image is input into a language model, and the text features of the partial reference building design image are output.

[0037] In steps S131-S133, it is intended to illustrate that in order to improve the quality of the building design image generated by the diffusion model, in the case that the partial reference building design image has no text information or the user does not input additional text information, the text features corresponding to the partial reference building design image are output by the language model to increase the feature information as much as possible to generate an accurate building style.

[0038] The building image generation method provided by the present application is further described below in combination with two embodiments: Embodiment 1: The following will be described in combination with the accompanying Figure 8To make a detailed description of embodiment 1, first, the reference picture, the image that needs to be redrawn, the reference picture mask and the redrawn mask are obtained, the reference picture mask is consistent with the size of the reference picture, and the target area coordinates of the reference picture mask are used to crop the image block corresponding to the reference picture according to the target area coordinates, the redrawn mask is consistent with the size of the image that needs to be redrawn, and the target area coordinates of the redrawn mask are used to crop the corresponding image block that needs to be redrawn according to the target area coordinates; then the sizes of the cropped images are adjusted respectively; the adjusted reference picture and the adjusted image that needs to be redrawn are spliced in the row or column direction, and the VAE encoder of the VAE model is loaded to encode the spliced image, the adjusted reference picture mask and the redrawn mask are spliced in the row or column direction, and the two splicing operations are performed in the same direction; the language model is loaded to describe the adjusted reference picture, the Clip text encoding of the reference picture is output based on the Clip model, and the Clip visual encoding of the reference picture is output based on the Clip visual model; the style model is loaded, and the Clip visual encoding and the Clip text encoding of the reference picture are input into the style model to obtain the conditional guide (i.e., a specific architectural style); then the VAE encoded result, the spliced result of the reference picture mask and the redrawn mask are taken as the inner filling model condition; the conditional guide and the inner filling model condition are jointly input into the U-Net model (the VAE encoding result is input, the spliced result of the reference picture mask and the redrawn mask is taken as the first guide condition, and the conditional guide in the condition is taken as the second guide condition) for diffusion reasoning, and a latent space image is output and decoded by using the VAE model. After image cropping of the decoded image, a building image meeting the user's expectation is finally obtained, and the building image is saved. Figure 8

[0039] Embodiment 2: The following will be described in detail with reference to the accompanying drawings. Figure 9 To make a detailed description of embodiment 2, Figure 9 The images from top to bottom on the left side are, in order: a redrawn architectural design image, a redrawn mask, a reference architectural design image, and a reference mask. The image size of the redrawn architectural design image and the redrawn mask is 2500x2106 pixels, the image size of the reference architectural design image and the reference mask is 1700x1895 pixels, the target area size of the redrawn mask is 1024x1152 pixels, the redrawn architectural design image is cropped according to the target area of the redrawn mask, the cropped redrawn architectural design image is 1024x1152 pixels, and Figure 9 ​It can be seen that the cropped redrawing architectural design image mainly includes a building of glass curtain wall style, and the target region size of the reference mask is 1344x1056 pixels. The reference architectural design image is cropped according to the target region of the reference image, and the cropped reference architectural design image is 1344x1056 pixels. It can also be seen that the cropped reference architectural design image mainly includes a building of green garden and open staircase. Since the size of the cropped reference architectural design image and the cropped redrawing architectural design image is inconsistent, the cropped reference architectural design image is scaled to an aligned reference architectural design image with the same height as the cropped redrawing architectural design image according to the height of the cropped redrawing architectural design image. At this time, the size of the aligned reference architectural design image is 1472x1152 pixels. The aligned reference architectural design image with the same height and the cropped redrawing architectural design image are spliced in the height direction to generate a spliced architectural design image with a size of 2496x1152 pixels. The spliced architectural design image includes not only the region of the redrawing architectural design image to be generated, but also the region of the reference architectural design image expected to be generated by the user. Since the target region size of the redrawing mask is 1024x1152 pixels, and the spliced architectural design image is 2496x1152 pixels, in order to ensure that the diffusion model can focus on key features, a synthetic mask image with the same size as the spliced architectural design image can be created. The synthetic mask image is 2496x1152 pixels. Since the synthetic mask image and the spliced architectural design image are one-to-one corresponding in pixel coordinates, the pixel values of the target region of the redrawing mask are filled into the corresponding synthetic mask image. It should be noted that although the aligned reference architectural design image is used in the process of generating the synthetic mask image, the size of the aligned reference architectural design image is used, and the pixel values of the aligned reference architectural design image are not relied on. Thus, the preprocessing operation is completed. The cropped reference architectural design image is described using the Clip model to output text features, and the image features of the cropped reference architectural design image are extracted using the Clip visual model. Before inputting into the U-Net model, the spliced architectural design image is encoded using the VAE encoder to obtain the low- latitude latent space encoding features. The latent space encoding features are used as the main input of the U-Net model, and the synthetic mask image, the text features and the image features are collectively used as the control condition input into the U-Net model to guide the U-Net model to infer according to the control condition, and output the latent space image features of the spliced architectural design image. The latent space image features are input into the VAE decoder to obtain a 2496x1152 inference architectural design image. At this time, the inference architectural design image includes the cropped reference architectural design image and the image block of the architectural design image expected to be generated by the user.Finally, the inferenced architectural design image is cropped according to the target area of ​​the redrawing mask of size 1024×1152 to obtain the cropped inferenced architectural design image of 1024×1152. The cropped inferenced architectural design image is the architectural style that the user expects to generate. The cropped inferenced architectural design image is then filled into the local image block of the redrawing architectural design image corresponding to the target area of ​​the redrawing mask to generate the target architectural design image of 2048×1728. Figure 9 It can be seen from the generated target architectural design image that the generated target architectural design image contains the style of the green garden and open stairs in the reference architectural design image, and compared with the redrawn architectural design image, there is no style confusion, unclear boundaries and other problems in the non-target area.

[0040] The generation method provided by the present invention has a first guiding condition that can lock the target area, reduce the influence of non-target areas on the generation results during the generation process, and solve the problem of feature attention imbalance in the diffusion model during the generation process; secondly, the second guiding condition can output accurate style features, solving the problem of difficult semantic alignment of multimodal inputs. This dual guiding condition, as a control condition, solves the problems of poor image quality such as style confusion and image blurring in the architectural image generation effect of the diffusion model.

[0041] As a second aspect of the present invention, there is provided an electronic device, such as Figure 10 As shown, including: One or more processors 101; The memory 102 stores one or more computer programs. When the one or more computer programs are executed by the one or more processors 101, the one or more processors 101 implement the scheduling method provided according to the first aspect and the second aspect of the present invention.

[0042] The tool may further include one or more I / O interfaces 103 connected between the processor 101 and the memory 102 and configured to implement information exchange between the processor 101 and the memory 102 .

[0043] Among them, the processor 101 is a device with data processing capabilities, including but not limited to the central processing unit 101 (CPU); the first memory 102 is a device with data storage capabilities, including but not limited to random access memory 102 (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory 102 (ROM), electrically erasable programmable read-only memory 102 (EEPROM), flash memory (FLASH); the I / O interface 103 (read-write interface) is connected between the processor 101 and the memory 102, and can realize information exchange between the processor 101 and the memory 102, including but not limited to the data bus 104 (Bus), etc.

[0044] In some embodiments, the processor 101, the memory 102 and the I / O interface 103 are connected with each other through the bus 104, and further connected with other components of the computing device.

[0045] Further, as a third aspect of the present application, a computer readable medium is provided, which stores a computer program, such as Figure 11 As shown, the computer program is executed by a processor to implement the generation method according to the first aspect of the present application.

[0046] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. Accordingly, the computer program can be stored in a non-volatile computer readable storage medium, and when executed, the computer program can implement the method of any one of the above-mentioned embodiments. In the embodiments provided by the present application, any reference to the memory, storage, database or other medium can include non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM), etc.

[0047] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Those skilled in the art should understand that the present application includes but is not limited to the contents described in the above specific embodiments and the accompanying drawings. Any modification which does not deviate from the functional and structural principles of the present application shall be included in the scope of the claims.

Claims

1. A method for generating a building image, characterized in that: The method comprises: Acquire an original architectural design image and a mask image, wherein the mask image includes an occlusion area and a target area, and a shape of the target area matches at least one local image block in the original architectural design image; Preprocessing the original architectural design image according to the mask image to obtain a partial architectural design image and a spliced ​​architectural design image, and synthesizing the mask image according to the partial architectural design image to obtain a synthesized mask image; wherein the synthesized mask image is a first guiding condition; Extract text features and image features of local architectural design images; Inputting text features and image features into a pre-trained application style model and outputting a second guidance condition, where the second guidance condition corresponds to a specific architectural style; Inputting the spliced ​​architectural design image, the first guiding condition, and the second guiding condition into the diffusion model to obtain an inferred architectural design image; The inferred architectural design image is post-processed according to the mask image to obtain the target architectural design image.

2. The generation method according to claim 1, characterized in that The step of inputting text features and image features into a pre-trained application style model and outputting a second guiding condition includes: Fusing the text feature and the image feature to generate a fused feature; The fused features are input into an application style model, and the output corresponds to a specific architectural style.

3. The generation method according to claim 2, characterized in that The specific architectural style is selected from the style information of the architectural image, the texture information of the architectural image, the environment information of the architectural image, the color and material information of the architectural image, and the form and structure information of the architectural image.

4. The generation method according to claim 1, characterized in that The step of inputting the spliced ​​architectural design image, the first guiding condition, and the second guiding condition into the diffusion model to obtain the inferred architectural design image includes: Input the spliced ​​architectural design image into the VAE encoder and output the encoded features of the spliced ​​architectural design image in the latent space; Inputting the encoding features of the latent space, the first guiding condition, and the second guiding condition into a U-Net model, and outputting the encoding features of the latent space under the first guiding condition and the second guiding condition for the spliced ​​architectural design image; The encoded features of the latent space based on the first guiding condition and the second guiding condition are input into a VAE decoder to obtain an inferred architectural design image.

5. The generation method according to claim 4, characterized in that Inputting the encoding features of the latent space, the first guiding condition, and the second guiding condition into the U-Net model, and outputting the encoding features of the latent space under the first guiding condition and the second guiding condition for the spliced ​​architectural design image, comprises: The encoded features of the latent space are used as the basic carrier and input into the U-Net model; The first guiding condition and the second guiding condition are inputted into a U-Net model as control conditions to guide the U-Net model to infer the basic carrier; The output stitched architectural design image is based on the encoded features of the latent space under the first guided condition and the second guided condition.

6. The generation method according to claim 1, characterized in that The original architectural design image includes a reference architectural design image and a redrawn architectural design image; the mask image includes a reference mask and a redrawn mask; The reference mask has the same size as the reference building design image, and the redraw mask has the same size as the redrawn building image; The local architectural design image includes a local reference architectural design image and a local redrawn architectural design image; The preprocessing of the original architectural design image according to the mask image to obtain a partial architectural design image and a spliced ​​architectural design image, and synthesizing the mask image according to the partial architectural design image to obtain a synthesized mask image includes: Crop the reference building design image according to the target area of ​​the reference mask and output a local reference building design image; Crop the redrawn architectural design image according to the target area of ​​the redraw mask, and output the partially redrawn architectural design image; Based on the size of the partially redrawn architectural design image, the partial reference architectural design image is scaled and the aligned reference architectural design image is output; splicing the aligned reference architectural design image and the partially redrawn architectural design image in a row or column direction, and outputting a spliced ​​architectural design image; The target area of ​​the redrawn mask and the aligned reference architectural design image are spliced ​​in a row or column direction to output a synthesized mask image.

7. The generation method according to any one of claims 1 to 6, characterized in that The post-processing of the inferred architectural design image according to the mask image to obtain the target architectural design image includes: cropping the inferred architectural design image according to the target area of ​​the redrawn mask, and outputting the cropped inferred architectural design image; The cropped inference architectural design image is filled into a local image block of the redrawn architectural design image that matches the target area of ​​the redrawn mask, and a target architectural design image is output.

8. The generation method according to any one of claims 1 to 6, characterized in that: The extraction of text features and image features of a local architectural design image includes: In the case where the partial reference architectural design image includes text content, the partial reference architectural design image is input into the Clip model, and text features of the partial reference architectural design image are output; Input the local reference architectural design image into the Clip visual model and output the image features of the local reference architectural design image; In the case that the partial reference architectural design image only includes image content, the partial reference architectural design image is input into a language model, and text features of the partial reference architectural design image are output.

9. An electronic device, characterized in that: include: one or more processors; A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the generation method according to any one of claims 1 to 8.

10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the generating method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Image generation method and device, computer equipment and storage medium

    CN117078790A

  • Commodity main graph generation method and device based on Stable Diffusion, equipment and medium

    CN118397140A

  • Attitude guide image synthesis method and device, storage medium and product

    CN118505495A

  • Home style image generation method and device based on diffusion model, and storage medium

    CN119251044A

  • Building planning image generation method and system based on potential diffusion model

    CN119557955A

Cited By

  • Local redrawing method and device of building effect picture, electronic equipment and storage medium

    CN121883651A