A method for automatic generation of residential floor plans based on multimodal input
Through the combination of multimodal feature extraction and adaptive generation modules, multiple input modes are integrated, and the limitations of relying on a single input mode in the existing technology are solved, and efficient and personalized automatic generation of residential floor plans is achieved, which improves design efficiency and output quality.
Patent Information
- Application Number
- CN202510420815.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-06
AI Technical Summary
The existing residential design models rely only on one or two input modes, which are difficult to meet the needs of personalized design, and lack a comprehensive multimodal data set and effective vector graph generation method, which limits design efficiency and output quality.
A residential floor plan automatic generation method based on multimodal input is proposed. Through the multimodal feature extraction module and the adaptive generation module based on domain knowledge, the functional connection morphology, spatial geometric morphology and personal preference constraints are integrated, the contribution weights of each modal feature are dynamically calculated, the generation order is determined, and the vector extraction module is regularized.
It significantly improves the personalization and flexibility of the residential design model, improves design efficiency, reduces the number of modifications by designers, ensures the independence and accuracy of vector graphics, simplifies downstream vector extraction tasks, and performs well under multimodal input conditions.
Smart Images

Figure CN119918163B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to residential plan design technology, and in particular to a residential plan automatic generation method based on multi-modal input. Background Art
[0002] Residential design plays a vital role in determining the quality of life, health, and well-being of residents. It is an important issue in society that is receiving increasing attention. It is also a complex and open problem that usually requires multiple inputs due to the need for personalization and flexibility. However, existing studies only accept one or two input types, mainly due to: 1) the lack of a comprehensive large-scale multimodal dataset, especially one containing vector drawings, which is crucial for the construction industry; and 2) the significant differences between different modal representations, making it extremely challenging to integrate them into a unified model.
[0003] Residential design involves multiple stages, including graphic design and performance simulation. These stages usually require designers to make multiple revisions, which is time-consuming and resource-intensive. In addition, design requirements are constantly influenced by social changes, which increases the need for personalized solutions that integrate diverse information. Therefore, there is a growing demand for residential design models that can handle multiple input modes. However, most existing research in the field of residential generation still relies on only one or two input modes. This limits the ability to meet personalized design needs and reduces the applicability of the model in real-world scenarios. In addition, vector drawings, as the main design medium for architects and the only input that truly reflects the design process, are still neglected in current research.
[0004] In addition to the lack of comprehensive multimodal datasets for residential design, an important reason for the limited progress in multimodal housing generation is that the information characteristics of various data modes vary significantly, making it difficult to integrate them into a single model. In addition, generating vector graphics is essential because they play a vital role in the workflow of actual residential designers, such as downstream tasks such as residential performance simulation. Pixel instability during the generation process affects the output quality, and training additional models to solve this problem will incur significant computational costs. Existing methods have difficulty in efficiently and stably generating high-quality vector graphics, limiting their application in actual design workflows.
[0005] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention
[0006] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a method for automatically generating residential floor plans based on multimodal input.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for automatically generating a residential floor plan based on multimodal input comprises the following steps:
[0009] S1. Acquire multimodal input data, wherein the multimodal input data includes at least one modality information of functional connection form, spatial geometry form and personal preference constraint, wherein the functional connection form includes a phased vector map and an N-graph structure, the spatial geometry form includes a residential outline and a room block, and the personal preference constraint includes a text description and a size parameter;
[0010] S2. Through a multimodal feature extraction module, feature extraction is performed on the functional connection form, spatial geometric form and personal preference constraint respectively, wherein the functional connection form adopts a multi-scale fusion de-redundant attention mechanism to extract linear features, the spatial geometric form adopts a convolutional network to extract geometric features, and the personal preference constraint is converted into a structured input through a language model;
[0011] S3, an adaptive generation module based on domain knowledge, dynamically calculates the contribution weight of each modal feature, and determines the generation order of the residential layout and vector map according to the weight, wherein the target image with a higher contribution weight is first generated as the initial output, and then another part is generated based on the initial output, and the two are superimposed and merged into a complete plan view;
[0012] S4. Regularize the generated vector map through the vector extraction module, including noise filtering, line normalization and vector information extraction, to output a residential floor plan in vector format that can be directly used for downstream tasks.
[0013] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method for automatically generating a residential floor plan is implemented.
[0014] The present invention has the following beneficial effects:
[0015] The present invention proposes an innovative multimodal residential floor plan automatic generation method, which significantly improves the personalization and flexibility of the residential design model by integrating multiple input modalities, including functional connectivity morphology, spatial geometry and personal preference constraints. This method overcomes the limitation of existing studies that only rely on one or two input modes, and meets the construction industry's demand for diversified information integration by introducing large-scale multimodal datasets, especially datasets containing vector drawings. In addition, the present invention effectively extracts key features from multimodal inputs by adopting advanced technologies such as multi-scale fusion de-redundant attention mechanism and convolutional network, and dynamically calculates the contribution weights of each modal feature through an adaptive generation module based on domain knowledge, intelligently determines the generation order, and thus generates practical and personalized floor plans. This method not only improves design efficiency and reduces the number of designer modifications, but also ensures the independence and accuracy of the vector map through a two-step generation process, simplifying the downstream vector extraction task. Experimental results show that the method of the present invention surpasses the existing state-of-the-art methods when integrating all modalities as input, and also performs well on each individual modal trajectory, verifying its robustness and adaptability on different input types. Therefore, the present invention lays a foundation for future research and practical application in the field of residential design automation, and has significant development potential and practical application value.
[0016] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a comparison diagram of the multimodal residential floor plan generation (M-RPG) framework and datasets according to an embodiment of the present invention.
[0018] Figure 2 A diagram showing representative samples of a multimodal dataset according to an embodiment of the present invention.
[0019] Figure 3 This is a statistical analysis chart of a data set according to an embodiment of the present invention.
[0020] Figure 4 This is the overall architecture and module details of the multimodal residential floor plan generation (M-RPG) model of an embodiment of the present invention.
[0021] Figure 5 Graph showing comparison results of models under different input conditions according to an embodiment of the present invention.
[0022] Figure 6 This is an example diagram of model generation results combining different modal vector drawings according to an embodiment of the present invention.
[0023] Figure 7 Performance graph of the multimodal residential floor plan generation (M-RPG) model under mixed input conditions for testing the implementation of the present invention.
[0024] Figure 8 The present invention is a flow chart of a method for automatically generating a residential floor plan based on multimodal input. DETAILED DESCRIPTION
[0025] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.
[0026] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0027] The embodiment of the present invention provides a method for automatically generating a residential floor plan based on multimodal input (referred to as multimodal residential floor plan generation, M-RPG), which includes the following steps (see Figure 8 ):
[0028] Step S1, obtaining multimodal input data, wherein the multimodal input data includes at least one modal information of functional connection form, spatial geometry form and personal preference constraints, wherein the functional connection form includes a phased vector graph and an N-graph structure, the spatial geometry form includes a residential outline and a room block, and the personal preference constraints include a text description and a size parameter;
[0029] Step S2, through a multimodal feature extraction module, feature extraction is performed on the functional connection form, spatial geometric form and personal preference constraint respectively, wherein the functional connection form adopts a multi-scale fusion de-redundant attention mechanism to extract linear features, the spatial geometric form adopts a convolutional network to extract geometric features, and the personal preference constraint is converted into a structured input through a language model;
[0030] Step S3, an adaptive generation module based on domain knowledge dynamically calculates the contribution weight of each modal feature, and determines the generation order of the residential layout and the vector map according to the weight, wherein a target image with a higher contribution weight is first generated as an initial output, and then another part is generated based on the initial output, and the two are superimposed and merged into a complete plan view;
[0031] Step S4: Regularize the generated vector map through a vector extraction module, including noise filtering, line normalization and vector information extraction, to output a residential floor plan in vector format that can be directly used for downstream tasks.
[0032] In a preferred embodiment, the N-graph structure in the functional connection form enhances the spatial relationship expression of the traditional graph structure by adding room location and wall connection information, and uses deformable convolution to extract its node and connection features.
[0033] In a preferred embodiment, the multi-scale fusion and redundant attention removal mechanism includes: performing a multi-scale convolution operation on the input vector map to capture linear features, and removing redundant information through an attention mechanism to achieve feature fusion and enhancement.
[0034] In a preferred embodiment, the adaptive generation module aligns multimodal input features through a channel attention mechanism to eliminate potential conflicting information, calculates the contribution weight of each modality to the layout diagram and vector diagram based on domain knowledge, and dynamically selects the generation order.
[0035] In a preferred embodiment, the determination of the generation order includes: if the total contribution weight is higher than a threshold, a residential layout diagram is generated first, otherwise a vector diagram is generated first, and the initial output is input as a condition into another subtask generation module for iterative optimization.
[0036] In a preferred embodiment, the vector extraction module extracts vector elements of walls, doors and windows through color region recognition, and sequentially performs noise pixel filtering, missing pixel filling and line width normalization to generate standardized vector data.
[0037] In a preferred embodiment, the room block input in the spatial geometry is parsed through a natural language description by a language model, converted into a structured geometric representation, and combined with residential contour features to constrain floor plan generation.
[0038] In a preferred embodiment, the overlay fusion process performs pixel-level alignment of the residential layout map and the vector map, retaining the independent geometric information of the vector map and avoiding precision loss due to format conversion in downstream tasks.
[0039] In a particularly preferred embodiment, see Figure 4 , the multimodal feature extraction module includes a functional connection morphology feature extraction submodule, a spatial geometry morphology feature extraction submodule and a personal preference constraint feature extraction submodule. The functional connection morphology feature extraction submodule includes: a phased vector graph processing unit, which uses a multi-scale fusion redundant attention mechanism (MFDA) to capture linear features through multi-scale convolution kernels; and an N-graph structure processing unit, which uses deformable convolution to extract node positions and wall connection information. The spatial geometry morphology feature extraction submodule includes: a two-dimensional convolutional network, which extracts the geometric features of residential outlines and room blocks; and a language parsing unit, which converts natural language descriptions into structured room block parameters through a pre-trained language model. The personal preference constraint feature extraction submodule converts text descriptions and size parameters into image masks or heat map formats. See Figure 4 The AdapGDK module based on domain knowledge includes: a channel attention mechanism (CAM) for aligning multimodal features and eliminating conflicts; a contribution weight calculation submodule for dynamically allocating weights of each modality to the layout diagram and the vector diagram based on preset domain rules, and determining the target type to be generated first according to the sum of weights; a control valve module for selecting the residential layout diagram or vector diagram to be generated first according to the weight determination result, and inputting the initial output as a condition into another subtask module for iterative optimization. Figure 4 ,The vector extraction module adopts a rule-based extraction method and ,performs the following operations in sequence: identifying the color regions of walls, doors and windows, filtering ,noise pixels, filling missing pixels along lines, normalizing the vector line width to a single pixel, and ,finally extracting standardized vector information.
[0040] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method for automatically generating a residential floor plan is implemented.
[0041] This method not only improves the efficiency of residential floor plan design, but also ensures the accuracy and practicality of the design, and has significant application value and development potential.
[0042] The specific embodiments of the present invention are further described below.
[0043] This invention fills the gap in the prior art and creates a new technical task: unifying residential floor plan generation and introducing a residential design model that is compatible with multiple modal inputs. Specifically, a novel basic method, multimodal residential floor plan generation (M-RPG), is proposed, which uses residential multimodal feature extraction and a domain knowledge-based adaptive generation module (AdapGDK) to achieve seamless integration of multimodal inputs and generate practical and personalized floor plans. Extensive experimental results on the Multi-RP dataset and the RPLAN dataset show that M-RPG not only surpasses the existing state-of-the-art methods when integrating all modalities as input, but also performs well on each individual modal trajectory, verifying its robustness and adaptability on different input types.
[0044] In addition to constructing a multimodal residential dataset Multi-RP, which includes phased vector graphs in the design process, N-graph structure data, and other basic modalities of architectural design, this paper proposes a multimodal residential floor plan generation M-RPG framework, which supports one or more flexible input types and generates well-structured residential floor plans, such as Figure 1b). In order to avoid the misalignment problem that may be caused by single-step generation, the model of the present invention adopts a two-step method, dynamically determining whether to generate a layout or a vector map first, and then using the initial output to complete the other part. The final design is the superposition of these two outputs, ensuring that the vector map is not affected by other features, simplifying the downstream vector extraction task. The process is as follows, first pre-training: 1) artificially set the contribution value of each input condition to the residential layout map and the vector map, 2) randomly delete the complete input conditions during model training, for example, only retain 5 of the 11 input conditions, and the remaining 6 conditions are deleted (set to zero), 3) calculate the contribution value of the remaining 5 conditions to the residential layout map and the vector map, and select the one with the largest contribution value as the generated target image. Then the model is fine-tuned as a whole, that is, it is not fixed to generate a residential layout or a vector map, and the model directly generates a complete residential floor plan as the target image, so that the first stage of the model can determine whether to generate a layout or a vector map first. After the above process model, the generation result of the first subtask can be dynamically determined according to the input conditions.
[0045] Figure 1 and Figure 5 The paper also mentions a variety of models, among which HouseGAN, HouseGAN++, HouseDiffusion, etc. are graph-based input models, Graph2Plan is a residential floor plan and room number input model, Tell2Design, Imagen, LayoutGPT, Holodeck, DALL-E-2 are natural language input models, and FloorplanDiffusion is a room block input model. These models all have limitations when dealing with complex residential design requirements.
[0046] The proposed model is the first to comprehensively address the multimodal residential design generation problem. Extensive experiments were performed to validate the performance of M-RPG. First, it was compared with current models using their respective input modalities. Then, M-RPG was evaluated under mixed-modal conditions. The results show that our model achieves state-of-the-art performance, outperforming all existing models with their respective input modalities, and excels in coordination across multiple modalities.
[0047] The proposed M-RPG framework integrates different input modalities and generates accurate residential floor plans in a usable vector format. Comprehensive experiments (further detailed later) were also conducted to demonstrate that M-RPG matches or exceeds all existing models on their respective input modalities and effectively coordinates multiple modalities. This invention lays a foundation for future research and practical applications in the field of residential design automation.
[0048] Residential floor plan generation models can be divided into four main types based on their input format:
[0049] 1) Graph-based input: Models such as HouseGAN, HouseGAN++, and HouseDiffusion use graph-based input and train models based on the RPLAN dataset. However, these graphs usually only represent room types and door connections, but do not capture key spatial relationships such as room locations and wall connections.
[0050] 2) Residential floor plan and number of rooms input: For example, Graph2Plan is trained based on data from the RPLAN dataset, but the limited input structure limits its ability to handle more complex residential floor plans.
[0051] 3) Natural language input: Tell2Design uses RPLAN floor plans with fixed descriptions, but its performance is limited due to its flexible and free description. The same is true for the Imagen model based on T5. In contrast, newer models such as LayoutGPT and Holodeck aim to generate designs from open language input, although their ability to generate real residential floor plans is still limited.
[0052] The various forms of natural language input can reflect the different levels of design needs of users. Figure 5 ) to further illustrate the performance of the model under different natural language inputs:
[0053] Imperative natural language (accurate description): "The north side of this home wouldn't be complete without a balcony. The approximately 16 square feet balcony area can be accessed through the living room or adjacent public room. Bathroom 1 is located on the east side of the home, adjacent to the living room, and is approximately 15 square feet. Bathroom 2, the larger of the two bathrooms, is approximately 30 square feet and is conveniently located between the master bedroom and the second public area, near the west side of the house. Public Room 1 is located in the northeast corner of the home, is approximately 80 square feet, and is conveniently located adjacent to the balcony. Public Room 2 is nearly 100 square feet and is located in the northwest corner, easily accessible through a shared passageway from the adjacent kitchen or living room. The kitchen is located on the north side of the home, between the living room and the second public area, and is approximately 50 square feet. The living room is conveniently located in the southeast corner of the home, is approximately 250 square feet, and is accessible to almost all rooms in the home. The master bedroom is located in the southwest corner of the home, is approximately 120 square feet, and is adjacent to the living room." This precise description provides the model with rich and clear design information, which helps the model generate a home layout that closely meets the needs.
[0054] Open-ended natural language (fuzzy description): "I want a unique house with three bedrooms. One of the larger bedrooms is located on the north side with a balcony above it. The other bedroom is smaller and located on the lower left. The kitchen and bathroom are close to the large bedroom. The overall layout contrast is more reasonable. The house does not need to be too large, an area of about 80 or 90 square meters will be enough." This type of fuzzy natural language input focuses more on expressing the user's general expectations and preferences.
[0055] 4) Room Block Input: FloorplanDiffusion uses individual room elements extracted from the RPLAN dataset as input. Our research shows that this element-based input has the potential to be combined with natural language descriptions to generate comprehensive residential floor plans.
[0056] Most existing models rely on one or two input modalities, limiting their ability to address the open-endedness and complexity of real-world residential designs. In addition, they often have difficulty generating high-quality designs from natural language descriptions, further limiting their practical applicability.
[0057] Multi-RP Dataset
[0058] The Multi-RP dataset is a large-scale multimodal dataset for automatic residential design. The following is a detailed introduction to the collection and annotation, statistical analysis, and validation of the dataset. Figure 2 The image on the left is a residential floor plan collected by the present invention, and the image on the right illustrates the multimodal input format annotated by the present invention. Below, "Residential Layout" and "Residential Vector Drawing" represent two types of model output annotated by the present invention.
[0059] Collect and label
[0060] To construct the dataset, residential floor plans were collected from multiple sources, including the inventor's own collection, RPLAN, and Prothor-10k. A total of 107,388 residential floor plans were collected, of which 96,760 were reasonable. This includes 25,400 floor plans collected by the inventor, 64,462 from RPLAN, and 6,898 from Prothor-10k. The present invention manually filters out designs with major defects in accordance with architectural design guidelines, and then annotates the remaining schemes with multimodal information to meet the input requirements of real residential design tasks. Each annotation has undergone a final manual review to ensure consistency between different modalities. To this end, the present invention recruited 20 architecture graduate students, and each drawing was screened by one student and reviewed by another student. When annotating the collected residential floor plans, in order to fully and accurately meet the input requirements of the real residential design task, a rich variety of modal information is covered, including residential floor plans, residential layouts, residential vector diagrams, vector information, N-graph structures, residential outlines, room color blocks, completion degree descriptions, residential size descriptions, room descriptions, etc. The data collection and annotation process emphasizes the following key aspects:
[0061] 1) Diversity of residential structures: The dataset includes various typical residential floor plans, such as one-, two-, and three-bedroom apartments, reflecting common residential structures.
[0062] 2) Diversity of input methods: The present invention takes into account the diversity of modes (e.g., phased vector diagrams, diagram structures, and room blocks in the design process) and the diversity of representations within each mode. For example, the same floor plan provides five different sets of vector and graphic annotations to simulate the fact that architects may draw vector diagrams in different orders. These annotations also contain rich detailed information to more accurately reflect the various needs of residential design. Specifically, it may include a description of the degree of completion, such as "20% of the plane vector drawing has been completed", "60% of the plane vector drawing has been completed", and "100% of the plane vector drawing has been completed"; a description of the size of the house, such as "the area of the house is 124 square meters"; a description of the rooms, such as "the number of rooms is 3 bedrooms, 1 living room, 1 kitchen, 1 bathroom and 2 balconies, and their areas (square meters) are: bedrooms [23, 14, 10], living room
[53] , kitchen [9], bathroom [5], balcony [8, 2]".
[0063] 3) Design of vector graphics during the project process: Drawing on architectural practice, this invention establishes three principles for vector annotation: 1. Walls are continuous and uninterrupted, 2. Doors and windows are aligned with walls, and 3. Doors and windows do not appear independently.
[0064] 4) Design of N-graph structure data: The N-graph structure annotation of the present invention goes beyond the traditional numerical nodes, such as Figure 4 As shown in d), more detailed and controllable information is provided.
[0065] 5) Exact and bounding box representation: The present invention provides both an exact representation of features and a representation containing only a bounding box to meet both the needs of detailed and coarse consideration of the room geometry.
[0066] 6) Dataset size: The dataset contains more than 20 million images and text descriptions, providing sufficient scale for training and evaluating multimodal models. Figure 1 a) Comparison of the Multi-RP dataset with existing datasets demonstrates the enhanced scope and diversity of our annotations.
[0067] Statistical analysis
[0068] The Multi-RP dataset contains 11 different types of data, including phased vector graphs, N-graph structures, residential outlines, precise representations and bounding box representations of five types of rooms (living room, bedroom, kitchen, bathroom, and balcony), and the completion rates of vector graph descriptions, residential size descriptions, and room descriptions. Its purpose is to enable users to generate floor plans based on one or more randomly selected input conditions. The present invention defines 10 challenge tasks to test the performance of the model under different input combinations, demonstrating the flexibility and complexity of the dataset.
[0069] The dataset covers a wide range of residential floor plans and statistical analysis was performed at two levels: number of rooms and room types. Figure 3 As shown, the most common floor plans are six-room two-bedroom apartments and seven-room three-bedroom apartments, reflecting real-world housing trends. According to practical experience, when using inputs in real-world residential design tasks, vector inputs are the most commonly used mode. Therefore, the present invention will emphasize the performance of the evaluation model in generating outputs from this mode.
[0070] Dataset Validation
[0071] Since rules can be used to verify graph structure, room size and geometry, the present invention focuses the evaluation on phased vector graphs as a data test set.
[0072] The present invention extracted 100 floor plans from the Multi-RP dataset and invited 10 architectural professionals (including graduate students and practitioners) to draw corresponding vector annotations for comparison. Each expert annotated 10 samples, resulting in 100 annotated drawings. During the evaluation process, each professional compared two sets of images - one from the data test set and the other drawn by the professional - and determined which set was closer to the real performance.
[0073] The results summarized in Table 1 show that the difference between the two representation methods is small, highlighting the strong reference value of the data test set. In order to evaluate whether the dataset achieves the goal of supporting multimodal input driven generation of realistic floor plans, the present invention conducted a professional evaluation experiment, which is detailed in Section 5.4.
[0074] Table 1. Human expert evaluation results comparing vector graphics drawn in Multi-RP with real images
[0075]
[0076] M-RPG Framework
[0077] This paper proposes a multimodal residential floor plan generation (M-RPG) framework. The framework supports flexible multimodal input to generate various types of outputs. The overall architecture of the model is as follows: Figure 4 shown. Figure 4 In the figure, a) is the M-RPG framework structure, including the feature extraction module on the left, the AdapGDK module (yellow square), the control valve module (pink square) and the vector extraction module (gray square). b) Multi-scale fusion and redundant attention mechanism (MFDA) module. c) is the image conversion module. d) is the N-graph structure. The original graph structure only contains door connection information, while the N-graph structure adds wall connection and room location information.
[0078] In order to achieve flexible input processing and high-quality output and avoid error accumulation, the present invention uses a single model to complete the entire generation task, avoiding the use of multiple interconnected models. The present invention decomposes the task into two subtasks: generating the residential layout and generating vector graphics. The model determines which subtask to perform first based on the input conditions and adjusts accordingly. To ensure consistency, the results of the first subtask are used as input to the second subtask. The two outputs are then merged to generate the final residential floor plan without further model intervention. The vector graphics are converted into a geometric representation in bitmap format to support downstream applications of residential design.
[0079] Three types of input are defined: 1) Functional connection form: including phase vector graph and N-graph structure representation, respectively denoted as and 2) Spatial geometry: including the house outline and room blocks, represented as and d 3) Personal preference constraints: including the completion rate of vector map, residential area and room description, expressed as , and In addition, the model includes control valve information input, expressed as For example, in actual application scenarios, the M-RPG model may receive such input: fuzzy description "Please give me a 100 square meter house with two bedrooms and two bathrooms", which reflects the way of expressing requirements through natural language description in the personal preference constraint input type. The output is divided into three types: residential layout, residential vector map and complete residential floor plan, which are represented as , and The present invention asserts that the contribution weights of different input patterns vary due to the production of different outputs, expressed as .
[0080] like Figure 4 As shown in a), the framework consists of four main dimensions: multimodal feature extraction, adaptive generation module based on domain knowledge (AdapGDK), and vector extraction module.
[0081] Multimodal feature extraction: The input information of the present invention has multimodal, abstract, sparse and fuzzy characteristics. Each type of input contains different features, so a special extraction module needs to be designed:
[0082] 1) Functional connection form: This representation usually involves lines and nodes. The present invention uses a multi-scale fusion redundant attention mechanism (MFDA) module to extract features from the vector map. Figure 4 As shown in b), this module uses multi-scale and convolution to capture linear features from the vector map. Figure 4 As shown in d), since lines of different colors in the N-graph structure represent different connection information and nodes contain relative position relationships, the present invention uses deformable convolution to effectively extract their features.
[0083] 2) Spatial geometry: Given that residential outlines and room blocks are usually rectangular and well-structured, the present invention applies 2D convolution to extract their features. In addition, a large language model can be used to process free natural language input into room block input.
[0084] 3) Personal preference constraints: To ensure accurate representation, the present invention converts the above text into image-like data for input conditions, such as Figure 4 As shown in c).
[0085] AdapGDK: Considering 1) the potential conflicting information under multimodal input conditions, and 2) The displacement deviation that may occur during the direct generation process. The present invention uses a channel attention mechanism (CAM) for the fusion and reconstruction of input features. This method aligns the input conditions, thereby reducing the impact of contradictory information on the results. Subsequently, the present invention calculates the contribution weight (CW) based on domain knowledge to guide the output generation of the model to ensure easier fitting. The contribution weight is initialized as follows:
[0086]
[0087] in For example, in Figure 4 a), the present invention first generates a residential layout The model then uses the control valve information Generate residential vector map The output representation of the control valve adjustment model is as follows: the zero-value image represents adaptive generation, and the other two situations correspond to: and Finally, the output representations of the two steps are fused to obtain The loss function is defined as follows:
[0088]
[0089] in, According to the contribution weight Select the target image, To join Noise passing The noisy image after the step, Indicates conditional input, To generate the model.
[0090] Vector extraction module: The present invention designs a vector extraction module to convert the generated vector graphics into vector information. However, pixel instability during the generation process affects the output quality. Training an additional model to solve this problem will incur significant computational costs. Therefore, the present invention adopts a rule-based extraction method to ensure the stability of vector information extraction: First, the color areas representing walls, doors, and windows are extracted. Then the noise pixels are filtered out and the missing pixels are filled along the lines. Next, the width of all vector lines is normalized to a single pixel. Finally, the vector information is extracted from the processed image for use in subsequent tasks. It should be noted that this vector extraction module is not the main contribution of the present invention's work, but is intended to support downstream residential design tasks.
[0091] experiment
[0092] Datasets and evaluation metrics
[0093] The proposed Multi-RP dataset is used for experiments. For evaluation, four key metrics are used: FID (Fréchet Inception Distance), GED (Graph Edit Distance), IoU (Intersection over Union), and PSNR (Peak Signal-to-Noise Ratio). FID is used to measure the similarity of overall image features. GED evaluates the similarity of room connections within a residence. IoU evaluates the geometric similarity between rooms in a residence. PSNR measures the quality of the generated image. In the field of residential design, GED and IoU are particularly valuable because they can better reveal the structural correctness and geometric alignment of the generated residence. These two metrics are more interpretable than FID and PSNR when evaluating the functionality and rationality of residential design.
[0094] Implementation details
[0095] The model of the present invention is based on the Stable Diffusion 2.1 framework and fine-tuned using the ControlNet architecture. The present invention uses the following hyperparameters to train the model: learning rate 1e-5, weight decay 1e-5, batch size 16, optimizer AdamW. The code is implemented in Python and PyTorch, and experiments are conducted on a server equipped with an Intel Xeon Platinum 8383C CPU (2.70 GHz) and an NVIDIA RTX 4090 GPU.
[0096] Comparison with public benchmarks
[0097] To ensure a fair comparison, our model is evaluated using the corresponding modalities of existing residential generation models, most of which rely on the RPLAN dataset. Therefore, we only train our model on the RPLAN part of Multi-RP. The experimental results are summarized in Table 2.
[0098] Table 2. Comparison of model results under the same input conditions
[0099]
[0100] The underlined numbers indicate that the inventive model was the second best.
[0101] Quantitative comparison: Table 2 gives the experimental results.
[0102] 1) Graph structure input: Our model outperforms the current best model by an average of 105.4% on four metrics. This demonstrates the superiority of our model and highlights the added value of the N-graph structure, which provides richer information such as room locations and wall connections, leading to better control over residential floor plans.
[0103] 2) Other inputs: For input formats such as residential floor plans containing the number of rooms, imperative natural language, free natural language, and room blocks, the model of the present invention outperforms the current best model by 31.5%, 26.4%, 45.4%, and 27.7%, respectively. The free natural language input and the reference plane are the text description and residential floor plan provided by professionals, respectively. These results confirm that the model of the present invention surpasses all existing models in their respective input formats.
[0104] 3) Random Input: We fine-tune existing models on the dataset and test their performance. The results show that our model outperforms these baseline models by 20.9% on FID and 7.3% on PSNR. These experiments show that our model performs well under a variety of input conditions commonly used by current models.
[0105] Qualitative comparison: Figure 5 The model comparison results under different input conditions are shown.
[0106] Figure 5 A) shows the performance of the model under two natural language inputs: accurate imperative natural language and fuzzy free description. The results show that the model of the present invention can effectively capture the design requirements embedded in the language input and generate a more reasonable residential layout. Figure 5 The right panels b) to f) of the figure show that the model output of the present invention is very similar to the real residential structure, highlighting its competitive advantage in generating reliable residential designs from diverse inputs compared to current residential generation models. Figure 5 c) shows the results using graph structured input, d) using house outline and number of rooms, e) using room blocks, and f) random input.
[0107] Expert evaluation
[0108] In order to verify whether the proposed database meets the design requirements and evaluate the feasibility of the model in the actual design environment, the present invention conducted an expert evaluation. The evaluation focused on two aspects: 1) the combination of vector graphics and other modal inputs, and 2) the performance of the model under contradictory input conditions. All test input conditions were selected in cooperation with domain experts.
[0109] Performing vector graphics and combining other modal inputs: This paper evaluates the performance of the model when combining phased vector graphics input with other modalities. Figure 6 Shown is an example of the model generation results, combining vector plots of different modes.
[0110] 1) Staged vector diagram and single modality input: The results show that the staged vector diagram input works effectively with residential outline and N-graph structure data to generate floor plans that are consistent with the input conditions and meet the design expectations.
[0111] 2) Phased vector diagram and multimodal input: When the phased vector diagram input is combined with multiple modalities, the model successfully integrates three different input data and generates reasonable residential floor plans that meet the specified requirements.
[0112] 3) Phased vector diagram input and completion description: We further tested the model’s ability to integrate the completion of the vector diagram description. The results showed that the model was able to dynamically adjust the level of detail of the generated floor plans based on the provided description. These experiments confirmed that the model can effectively combine the input of the phased vector diagram with other modalities to generate accurate and practical residential floor plans that meet design expectations.
[0113] Handling contradictory input conditions: In real-world design scenarios, contradictory input conditions often occur. To evaluate the robustness of the model in such situations, we tested the model with inconsistent inputs, specifically, providing graph structure data that did not match the number of rooms information. The results are summarized in Table 3, showing that the model's performance was only slightly degraded from the original scenario. Importantly, the model did not experience significant performance degradation or crash due to contradictory inputs. These results show that the model is robust when handling inconsistent information and maintains stable performance without affecting the quality of the generated residential floor plans.
[0114] Table 3. Model results under contradictory conditions
[0115]
[0116] Ablation experiment
[0117] Feature extraction module: In order to evaluate the effectiveness of the feature extraction module, the present invention used the original ControlNet model as a benchmark (marked as Original original model) and tested the model using all input conditions. Subsequently, an ablation experiment was performed, and the results are shown in Table 4. The results show that the MFDA module of the present invention improves the performance of the baseline model by 3.5%, confirming its effectiveness. In addition, the model combining all feature extraction modules improved by 11.7% compared to the original model, achieving the best overall performance, further verifying the impact of the module. Finally, the present invention evaluated the image transformation module, and its performance improved by 1%, proving its effectiveness in the support generation process.
[0118] Table 4. Ablation results of model internal modules and mechanisms
[0119]
[0120] Adaptive feature selection mechanism: We tested the performance of the model in three scenarios: 1) The method proposed by the present invention: the model first generates an output that is easier to fit (layout or vector map), and then generates a second output based on the first output. 2) Inverse order: the model first generates an output that is more difficult to fit, and then generates the second output. 3) One-step generation: the model generates the entire floor plan at one time. The results shown in Table 4 show that the one-step generation performs poorly, 61.8% lower than the step-by-step method of the present invention. This may be due to the complexity of the input features, which will overwhelm the model when generating all outputs at the same time, resulting in color overflow and degraded output quality. In addition, the adaptive feature selection mechanism of the present invention is 15.3% higher than the reverse scenario, confirming its effectiveness in generating reasonable residential floor plans.
[0121] Mixed input conditions: To evaluate the robustness of the model under different input conditions, we tested it in 10 challenging tasks involving 1 to 10 mixed input conditions. For ease of comparison, we standardized the evaluation metrics as follows: Through the above standardization process, the larger the value of all indicators, the better the performance. Figure 7 The results shown show that: 1) More input conditions improve the generation quality. 2) Although the performance fluctuates slightly as the number of input conditions increases, the model remains stable without catastrophic failures or significant performance degradation. These results confirm that the model is able to handle complex mixed input conditions with a high degree of robustness.
[0122] The present invention discloses a method for automatically generating residential floor plans based on multimodal input, which can complete the unified residential floor plan generation task. To this end, the present invention constructs the Multi-RP dataset, which is the first large-scale multimodal residential floor plan dataset that combines multiple input methods to meet actual design needs. The present invention proposes a multimodal residential floor plan generation M-RPG framework, which supports flexible multimodal input and generates floor plans with vector outputs. Experimental results demonstrate the effectiveness of the M-RPG framework of the present invention under various input conditions. The model of the present invention achieves state-of-the-art performance under the respective modes of existing models. It also exhibits robustness and adaptability under one or more input conditions. In addition, it performs well in expert evaluation, further verifying its effectiveness. These findings indicate that the model of the present invention has significant development potential and practical application value in the field of automatic residential design, and also demonstrates the adaptability of CV technology in design tasks.
[0123] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0124] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0125] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.
[0126] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a ferromagnetic random access memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0127] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0128] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0129] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0130] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.
[0131] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0132] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0133] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0134] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0135] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for automatically generating residential floor plans based on multimodal input, characterized in that: The following steps are involved: S1. Acquire multimodal input data, wherein the multimodal input data includes at least one modal information of functional connection form, spatial geometry form and personal preference constraint, wherein the functional connection form includes a phased vector map and an N-graph structure, the spatial geometry form includes a residential outline and a room block, and the personal preference constraint includes a text description and a size parameter; the N-graph structure in the functional connection form enhances the spatial relationship expression of the traditional graph structure by adding room location and wall connection information, and uses deformable convolution to extract its node and connection features; S2. Through a multimodal feature extraction module, feature extraction is performed on the functional connection form, spatial geometric form and personal preference constraint respectively, wherein the functional connection form adopts a multi-scale fusion de-redundant attention mechanism to extract linear features, the spatial geometric form adopts a convolutional network to extract geometric features, and the personal preference constraint is converted into a structured input through a language model; The multi-scale fusion and redundant attention removal mechanism includes: performing a multi-scale convolution operation on the input vector graph to capture linear features, and removing redundant information through the attention mechanism to achieve feature fusion and enhancement; S3, an adaptive generation module based on domain knowledge, dynamically calculates the contribution weight of each modal feature, and determines the generation order of the residential layout diagram and the vector diagram according to the weight, wherein the target image with a higher contribution weight is first generated as the initial output, and then another part is generated based on the initial output, and the two are superimposed and fused into a complete floor plan; the adaptive generation module aligns the multimodal input features through the channel attention mechanism, eliminates potential conflicting information, and calculates the contribution weight of each modality to the layout diagram and the vector diagram based on domain knowledge, and dynamically selects the generation order; S4. Regularize the generated vector map through the vector extraction module, including noise filtering, line normalization and vector information extraction, to output a residential floor plan in vector format that can be directly used for downstream tasks.
2. The method for automatically generating a residential floor plan according to claim 1, characterized in that: Determining the generation order includes: if the total contribution weight is higher than a threshold, first generating a residential layout diagram, otherwise first generating a vector diagram, and inputting the initial output as a condition into another subtask generation module for iterative optimization.
3. The method for automatically generating a residential floor plan according to claim 1, characterized in that: The vector extraction module extracts the vector elements of walls, doors and windows through color region recognition, and sequentially performs noise pixel filtering, missing pixel filling and line width normalization to generate standardized vector data.
4. The method for automatically generating a residential floor plan according to claim 1, characterized in that: The room block input in the spatial geometry is parsed through a language model to convert the natural language description into a structured geometric representation, and combined with the house outline features to constrain the floor plan generation.
5. The method for automatically generating a residential floor plan according to claim 1, characterized in that: The overlay fusion process aligns the residential layout map with the vector map at the pixel level, retaining the independent geometric information of the vector map to avoid the accuracy loss caused by format conversion in downstream tasks.
6. The method for automatically generating a residential floor plan according to any one of claims 1 to 5, characterized in that: The multimodal feature extraction module comprises: Functional connection morphological feature extraction submodule, which includes: The phased vector graph processing unit uses the multi-scale fusion redundant attention mechanism (MFDA) to capture linear features through multi-scale convolution kernels; N-graph structure processing unit, which uses deformable convolution to extract node positions and wall connection information; The spatial geometric morphology feature extraction submodule includes: A 2D convolutional network to extract the geometric features of residential outlines and room blocks; Language parsing unit, which converts natural language description into structured room block parameters through a pre-trained language model; A personal preference constraint feature extraction submodule, which converts text descriptions and size parameters into image masks or heatmap formats; The domain knowledge-based adaptive generation module (AdapGDK) includes: Channel Attention Mechanism (CAM) to align multimodal features and eliminate conflicts; The contribution weight calculation submodule dynamically allocates the weights of each modality to the layout diagram and vector diagram based on the preset domain rules, and determines the target type to be generated first according to the sum of the weights; The control valve module first generates a residential layout diagram or a vector diagram based on the weight determination result, and uses the initial output as a condition to input into another subtask module for iterative optimization; The vector extraction module adopts a rule-based extraction method and performs the following operations in sequence: identifying the color areas of walls, doors, and windows, filtering noisy pixels, filling missing pixels along lines, normalizing the vector line width to a single pixel, and finally extracting standardized vector information.
7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for automatically generating a residential floor plan as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Engineering drawing information extraction method and device, electronic equipment and storage medium
CN118762381A
Large-scale engineering equipment design intention extraction method based on multi-modal data fusion
CN119068302A