House type identification and three-dimensional reconstruction method and system based on grating image
Through key point detection based on raster images and multimodal optical character recognition technology, combined with WebGL framework, convenient and highly accurate house type recognition and three-dimensional reconstruction are achieved, solving the problems of time-consuming, labor-intensive and large errors in traditional methods.
Patent Information
- Application Number
- CN202510340834.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional house type identification and three-dimensional reconstruction methods require manual participation, which is time-consuming and labor-intensive, and is easy to introduce human error, which cannot be reconstructed in advance, and has low accuracy.
The house type recognition and three-dimensional reconstruction method based on raster images is adopted, walls, doors and windows are identified through key point detection networks, and scale digital recognition is combined with multimodal optical character recognition model. A house type recognition and three-dimensional reconstruction system based on WebGL frame is designed to support the rendering and user interaction of the reconstructed wall and doors and windows view.
Convenient and automated house type identification and three-dimensional reconstruction have been realized, accuracy has been improved, human error has been reduced, and floor plan recognition and reconstruction are suitable for Chinese areas.
Smart Images

Figure CN120279394A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for house type recognition and 3D reconstruction, belonging to the technical field of house type image recognition. Background Technique
[0002] The techniques of house type recognition and 3D reconstruction have important research and application values in related fields such as architectural design and interior design. Traditional methods for house type recognition and 3D reconstruction usually require manual participation and complex measurement work, and even use lidar to scan the entire house type to obtain 3D point cloud data. Such reconstruction work is not only time-consuming and laborious, but also in many cases depends on existing buildings, unable to perform pre-reconstruction, and is prone to introducing human errors, resulting in low accuracy. Summary of the Invention
[0003] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0004] A method for house type recognition and 3D reconstruction based on raster images, including: Step 1. Collect and screen high-quality house type maps suitable for the national conditions of China, and construct a raster house type map vector dataset; Step 2. Based on a key point detection network, identify the walls, doors, windows, and scales in the house type map, and generate and screen candidate primitives through axial alignment rules and constraint conditions; Step 3. Use the combination of Yolov8 and Shi-Tomasi corner detection to detect and locate the scale endpoints, and perform scale digital text recognition through a pre-trained multi-modal optical character recognition model OFA-OCR; Step 4. Design and implement a house type recognition and 3D reconstruction system based on the WebGL framework, supporting the rendering of the reconstructed walls and door / window views, user interaction with the reconstructed house type, and the forward and backward functions of the house type operation history record.
[0005] A house type recognition and 3D reconstruction system based on raster images, implemented based on a method for house type recognition and 3D reconstruction based on raster images, including: a house type recognition and reconstruction module, used to receive the house type map file uploaded by the user, recognize and reconstruct the 2D house type map file uploaded by the user, and generate a 3D house type view; a house type wall and door / window interaction module, used to perform addition, deletion, and modification operations on the walls and doors / windows in the 2D and 3D views; a house type design history record module, used to record the user's design operation history, supporting the revocation, redoing, saving, and loading of design schemes.
[0006] Preferably, the housing type identification and reconstruction module includes: a file upload unit for receiving the housing type diagram file uploaded by the user, supporting.png,.jpg,.jpeg formats, with the file size not exceeding 2MB and the resolution not exceeding 1920×1080; an image processing unit for sending the housing type diagram file uploaded by the user to the backend for identification and reconstruction to generate 2D and 3D housing type views; a scale identification and adjustment unit for identifying the scale in the housing type diagram and allowing the user to adjust the position and length of the scale; and a view switching unit for switching between 2D and 3D views.
[0007] Preferably, the image processing unit realizes housing type identification and reconstruction through the following steps: wrapping the housing type diagram file uploaded by the user into a FormData object and sending it to the backend; the backend saves the housing type diagram file on OSS and forwards the image address and task ID to the message queue; the front end periodically queries the task status according to the task ID. After the backend completes housing type identification, it returns the identification result and loads the housing type diagram in the 2D Canvas; generates a 3D reconstruction view according to the identification result and displays it in the upper right corner.
[0008] Preferably, the housing type wall and door / window interaction module includes: a wall editing unit for adding, deleting, and merging walls in the 2D view, and supporting adjusting the length and position of the walls by dragging; a door / window editing unit for adding, deleting, and modifying doors and windows in the 2D view, and the positions of the doors and windows must be limited to the walls; and a 3D view synchronization unit for synchronously updating the operations in the 2D view to the 3D view.
[0009] Preferably, the wall editing unit supports the following operations: when deleting a wall, automatically deleting the doors and windows on the wall and redrawing the adjacent walls; when merging walls, merging the adjacent axially aligned walls into the same wall; adjusting the length and position of the wall by dragging the dot interaction areas at both ends of the wall.
[0010] Preferably, the door / window editing unit supports the following operations: providing multiple types of doors and windows for the user to choose; adding doors and windows on the wall by clicking the mouse, and supporting adjusting the length of the doors and windows by dragging; the geometric and semantic constraints of the doors and windows limit them to be added and modified only on the walls.
[0011] Preferably, the housing type design history record module includes: a historical operation record unit for recording the user's design operation history, supporting undo and redo operations; a design scheme saving unit for serializing the user's design scheme and saving it to the backend database; and a design scheme loading unit for loading the design scheme saved by the user and supporting deleting the saved design scheme.
[0012] The beneficial effects of the present invention are as follows:
[0013] A raster image is a two-dimensional planar image obtained by devices such as cameras or laser scanners, which contains a large amount of information about the structure and layout of a house. Due to its low price, convenient dissemination, and vividness, the house type raster image exists widely in daily life. The recognition and reconstruction technology based on raster images provides a more convenient and automated method for house type recognition and 3D reconstruction, meeting people's expectations. By analyzing features such as lines, corner points, and textures in the raster image, the house type information of the house can be inferred, such as the location, size, and connection relationship of rooms. At the same time, by using the perspective changes of multiple raster images, the 3D reconstruction of the house can be realized, generating a house model with a geometric structure. Description of the Drawings
[0014] Figure 1 is a schematic diagram of the key points of the wall; Figure 2 is a schematic diagram of the key points of the door; Figure 3 is a schematic diagram of the key points of the window; Figure 4 is an example of the house type diagram in the dataset; Figure 5 is an example of the annotation data of the house type primitive; Figure 6 is an example of the annotation process of the end points of the scale line; Figure 7 is the data file of the annotation of the end points of the scale line; Figure 8 is a diagram defining the categories of the key points of the wall; Figure 9 is the structure diagram of the ConvNeXt network block; Figure 10 is the structure diagram of the FPN; Figure 11 is the network structure of the PANet; Figure 12 is the structure diagram of the BiFPN network; Figure 13 is the structure diagram of the BiFPN network based on spatial attention; Figure 14 is the structure diagram of the spatial attention module; Figure 15 is a schematic diagram of the maximum bounding box of the house type area; Figure 16 is a schematic diagram of the detection of the scale end point area; Figure 17 is a schematic diagram of the detection of the corner points of the scale end points; Figure 18 is a schematic diagram of the recognition result of the OFA-OCR scale number; Figure 19 is the architecture diagram of the house type diagram recognition system; Figure 20 is the flow chart of the recognition and reconstruction of the house type diagram; Figure 21 is the sequence diagram of the house type recognition and reconstruction module; Figure 22 is the flow chart of the user's interaction with the wall and doors and windows; Figure 23 is the sequence diagram of the wall and doors and windows interaction module; Figure 24 is the flow chart of saving the historical record; Figure 25 is the flow chart of restoring the historical record; Figure 26 is the flow chart of the comparison algorithm of the new and old scene trees; Figure 27 is the scale setting interface; Figure 28It is the result of household type reconstruction from a 2D perspective; Figure 29 It is the result of household type reconstruction from a 3D perspective; Figure 30 It is the diagram for adjusting the transparency of the household type floor plan background; Figure 31 It is the result of 3D reconstruction of common household types; Figure 32 It is the schematic diagram of the selected target wall; Figure 33 It is the schematic diagram of wall deletion; Figure 34 It is the axial dragging of the wall edge; Figure 35 It is the process of the wall being dragged along the cross-axis; Figure 36 It is the diagram of the end state of the wall dragging; Figure 37 It is the schematic diagram of adding a sliding door; Figure 38 It is the state after dragging the wall on the historical record operation interface; Figure 39 It is the historical record state before restoring the dragging; Figure 40 It is the user login page; Figure 41 It is the persistent historical design record. Specific implementation manners
[0015] Specific implementation manner 1: This implementation manner discloses a method for identifying and 3D reconstructing a house household type based on a raster image, including the following steps:
[0016] Step 1. Collect and screen high-quality household type floor plans suitable for the national conditions of the Chinese region, and construct a raster household type floor plan vector data set;
[0017] The household type floor plan picture data comes from Internet pictures. After manual screening of the pictures, more than 5,000 household type floor plans are comprehensively selected as the household type floor plan data (3,600 household type floor plans are used as the training set, and the remaining more than 1,400 pictures are used as the test set). The geometric information of each floor plan image is annotated by the method of manual annotation, and the lines representing the walls and doors and windows are marked. During training, the data set annotation is read and the key point positions of each type are calculated in real time, and a Ground-truth heat map representation is constructed to train the network, and the annotation information is converted into a connection layer representation, as follows:
[0018] (1) For the door and window key point objects, directly read the two end position points corresponding to the door and window categories in the annotation, and determine the direction information of the two end position points to determine the key point category at the door and window end positions;
[0019] (2) For the wall key points, they are usually calculated from multiple wall primitives. Therefore, the wall intersection points are determined by calculating the local connection of the wall, that is, checking whether there are other walls on the upper, lower, left, and right sides of the wall, and the results are de-duplicated to accurately calculate the category and position of each wall key point; the results are shown in Figures 1 - 3 , where Figure 1 are the wall key points, Figure 2 are the door key points, Figure 3 are the window key points, inFigures 1 - 3 Among them, each key point can not only express position information, but also be attached with semantic information of categories;
[0020] The house type diagram and its annotation data are as Figure 4 and Figure 5 shown. Each row represents a wall or door / window primitive. The first four columns represent the positions of the two endpoints of the primitive with respect to the picture dataset. The fifth column represents the type of the primitive, where "wall" represents a wall, "opening" represents a window, and "door" represents a door. The sixth and seventh columns are reserved categories. The sixth column is the major category of walls and doors / windows. For walls, some interior walls and exterior walls are reserved. For doors and windows, some sub-categories are reserved, such as single-leaf doors or double-leaf doors, etc. The seventh column is the orientation information of the door;
[0021] The dataset of the end point area of the house type diagram's scale line is labeled in the Yolo format. The end points of the scale area of each house type diagram are manually labeled using the makesense annotation tool, and there is only one category, namely the scale line end points. The annotation process is shown in Figure 6 , and the end point area of the house type diagram's scale line is selected using the bounding box method and fine-tuned and corrected. The annotation data file of the scale line end points is shown in Figure 7 ; In Figure 7 , the first column is the target category. In this task, there is only one category, namely the end point area of the scale line. The following four columns are the x coordinate of the center point, the y coordinate of the center point, the width of the detection box, and the height of the detection box with the upper left corner of the picture as the origin. The length and width are normalized with respect to the original image and mapped between 0 and 1.
[0022] Before officially entering the training, the data images also need to be preprocessed. After filling the pictures into squares, they are then adjusted to a resolution of 512×512. At the same time, in order to prevent the model from overfitting, resulting in the problem of weak generalization ability of the trained neural network, a data augmentation scheme is adopted to expand the dataset. This application adopts two data augmentation schemes. One is noise filling, adding random Gaussian noise to the house type images. The other is randomly rotating the house type images, and the rotation angles are 0 degrees, 90 degrees, 180 degrees, and 270 degrees.
[0023] Step 2. Identify the walls, doors, windows, and scales in the house type diagram based on the key point detection network (CPN-Floor), and generate and filter candidate primitives through the axial alignment rule and constraint conditions;
[0024] Step 21. Define the key points of the raster house type diagram elements;
[0025] The main elements that make up the housing unit structure are walls and doors / windows. The main elements that make up the housing unit structure are encoded as a set of connection points with categories. The wall structure (represented by a set of connection points where walls intersect) has a total of 4 types of wall connection types: I-shaped, L-shaped, T-shaped, and cross-shaped. The connection points of the door are defined to have 8 categories, and the connection points of the window also have 8 categories. The 8 categories of connection points for both doors and windows include: the midpoint connection point between two parallel and collinear walls, the eccentric connection point between two parallel and collinear walls, the connection point at the angle between two perpendicular walls, the non-angle connection point between two perpendicular walls, the connection point between the corner and the wall center line, the wall end point connection point, the wall intersection point connection point, and the free connection point.
[0026] In the housing unit floor plan, the main elements of the housing unit structure are composed of walls, doors, and windows. First, the main elements are encoded as a set of connection points with categories. In the housing unit plan recognition task, the recognition of the wall structure is the most common and important task. In this embodiment, the wall structure is defined to be represented by a set of connection points where walls intersect. There are a total of 4 types of wall connection types, I-shaped, L-shaped, T-shaped, and cross-shaped. A plane rectangular coordinate system is constructed with the wall connection point as the origin. Considering rotation, the I-shaped has four directions, corresponding to four categories, namely the direction from 30 degrees to 60 degrees, the direction from 120 degrees to 150 degrees, the direction from 210 degrees to 240 degrees, and the direction from 300 degrees to 330 degrees. The L-shaped and T-shaped also consider rotation, and each category is rotated 90 degrees in the positive direction from the previous category direction, each having 4 categories. The cross-shaped has only 1 type. There are a total of 13 categories of wall connection points, and the wall categories are defined as Figure 8 .
[0027] Doors and windows are represented as a line in the housing unit plan. Taking 45 degrees as a step, there are a total of 8 directions. Each category of door and window has 8 directions, that is, 8 types of key points. Since the common doors are double-leaf sliding doors and single-leaf doors, both of these doors have common widths. For example, the width of a single-leaf door is between 0.8m and 1.2m, and the length of a double-leaf sliding door is usually between 1.8m and 3m. Therefore, in the housing unit plan, the approximate category of the door can be judged by a classification plus length information. Windows only consider the category of ordinary windows in this task. In this embodiment, the position information of doors and windows is mainly considered. The connection points of the door have 8 categories, and the connection points of the window also have 8 categories.
[0028] After obtaining the key points of each category of elements through the key point detection network, the connection points will be encoded as geometric primitives through alignment rules. Walls, doors, and windows are represented as a line. The effective primitives should be able to be connected in the direction represented by the key point category. At the same time, due to some prior knowledge of the wall, the wall primitive must form a closed one-dimensional loop, and the doors and windows must also be located on the wall. Through a simple heuristic post-processing method, a planar vector representation with a high-level structure can be obtained, thereby realizing the vectorization of the main elements of the housing unit.
[0029] Step 22. Use ConvNeXt-B pre-trained on Image-21K as the basic feature extraction network to extract multi-scale features of the house type;
[0030] In this embodiment, ConvNeXt-B pre-trained on Image-21K is used as the basic feature extraction network and fine-tuned during the training of the downstream dataset in this task.
[0031] VGG proposed that the backbone network is divided into several network block structures. Each network block downsamples the feature map by a fixed multiple through pooling. Each network block consists of several different basic layer operations. ConvNeXt proposed the concept of stage on the basis of the block. Each stage is composed of several network blocks. The ratio of the blocks in each stage is 1:1:9:1. Each stage usually consists of 3 blocks. The final number of blocks and ratio are 3:3:27:3. At the same time, in terms of the activation function, ConvNeXt uses the GELU activation function, effectively avoiding the possible problems of the ReLU activation function. The network block structure diagram of ConvNeXt is as Figure 9 : In the design of the network block, the idea of ResNeXt is adopted, and a 7×7 convolutional kernel is used for depth convolution, and more groups are used to expand the width of the feature vector. Depth convolution is a special case of group convolution, and the number of groups is equal to the number of input channels. This convolution method has been widely used in MobileNetV2 and Xception. In depth convolution, each input channel is convolved with a separate convolutional kernel, and only each channel is convolved. The information between channels is independent. This convolution method is often used to extract spatial information without considering the correlation between channels, which will have better performance in tasks that focus on spatial information such as key point detection of house type diagrams. The hidden layer in the block also adopts the inverted bottleneck proposed in MobileNetV2. The hidden bottleneck dimension is four times wider than the input dimension. This design has brought better effects in the ResNet-200 and Swin-B architectures. This depthwise separable convolution has a lower number of parameters and computational cost without losing too much performance. At the same time, Layer Norm is used instead of Batch Norm in depth convolution and pointwise convolution. Regularization techniques can improve model convergence and reduce overfitting. However, Batch Norm also has some more complex problems that may have an adverse impact on the model. Learning from the simpler LayerNorm used in Transformer, good performance can be obtained in the application scenario of this embodiment. In ConvNeXt, better performance is obtained by using Layer Norm than Batch Norm. The overall structure of the ConvNeXt-B network is shown in Table 1:
[0032] Table 1 Structure diagram of ConvNeXt-Base network
[0033]
[0034] Among them, c represents the input channel, s represents the convolution stride, and Conv represents a convolution block, which is composed of a convolutional layer, a normalization layer, and a ReLU layer.
[0035] Step 23. According to the multi-scale features of the house type extracted in Step 22, a bidirectional feature pyramid network (BiFPN) is used for feature fusion of the neural network. After each feature fusion in BiFPN, a spatial attention weight is calculated and multiplied on the fused feature map. The spatial attention calculation formula is as follows:
[0036] M s (F)=sigmoid(conv 7*7 (concat(AvgPool(F),MaxPool(F))))
[0037] Among them, F is the feature map for which spatial attention needs to be calculated, Ms is the spatial attention feature map, which has only one channel and the same feature width and height as F, conv7*7 is a convolutional layer with a convolution kernel size of 7, concat represents the operation of concatenating by channels, AvgPool is average pooling in the channel direction, and MaxPool represents max pooling in the channel direction.
[0038] Feature Pyramid Networks (FPN) is a feature fusion module of the neural network, which mainly solves the deficiency of object detection in dealing with multi-scale change problems. Shallow features have high spatial resolution for localization, which is beneficial for object localization, but the semantic information for recognition is weak. Deep features have richer semantic information, which is beneficial for classification and recognition, but the spatial resolution is low. Therefore, there is usually a structure similar to U-Net to maintain the spatial resolution and semantic information of the feature layer. FPN is a computationally efficient top-down network structure with lateral connections. It further improves the U-shaped structure using deep supervision information. This structure is used to construct feature maps of different sizes with advanced fused semantic features, which can fuse low-resolution maps and high-resolution maps with less computational effort. The FPN structure is as Figure 10 : Among them, the lateral connection uses a convolution with a convolution kernel size of 1 to adjust the number of feature channels, and performs pixel-level addition with the upsampled feature map to obtain the output feature of each layer.
[0039] In this embodiment, four layers of original features with different scales are used for feature fusion. The output features of each stage of the ConvNeXt backbone are {C1, C2, C3, C4}, and the input features for feature fusion in this task are {C1, C2, C3, C4}. To improve the efficiency of information transmission and maintain the integrity of the final information, after the obtained feature maps are fused by the feature fusion module, the features of all pyramid levels are connected to serve as a HyperNet to integrate information at different levels, rather than simply using the final upsampled result at the end of the hourglass module as in HourglassNet. Subsequently, the number of channels is adjusted to the predicted number of channels through two layers of 1×1 convolutions to serve as the heatmap output by the network.
[0040] PANet is an excellent design in semantic segmentation and object detection tasks. It adds an additional path aggregation network on the basis of FPN to enhance feature fusion and expression. PANet is shown in Figure 11 : This design has been proven to be effective, but this design also increases additional computational costs. To improve computational efficiency, BiFPN was proposed. BiFPN made the following optimizations for the cross-scale connections of PANet:
[0041] (1) If a computational feature is not fused with other features, that is, it has only one input edge, its contribution in the feature fusion process is relatively low and can be deleted for optimization;
[0042] (2) If the original input to the output feature is at the same level, an additional connection will be added to fuse more features without significantly increasing too much computational cost;
[0043] (3) Different from PANet which has only one top-down and one bottom-up path, BiFPN can repeat the feature fusion path multiple times and connect more BiFPNs in series to enable more advanced feature fusion.
[0044] Above, the structure of BiFPN can be obtained, as shown in Figure 12 shown, Figure 12 only one case of BiFPN is shown.
[0045] In addition, when performing Feature Pyramid Network (FPN) feature fusion, feature maps with smaller sizes are often upsampled and then pixel-wise added to the feature maps connected horizontally. However, in BiFPN, different input features have different feature resolutions, and the geometric or semantic information they contain also varies in importance. Therefore, the contributions of input feature maps to the output feature are not equal. So, when performing feature fusion, each feature map will learn a weight and then the pixel values are added after weighting. Traditional unbounded weight parameters can lead to unstable training, while weight calculation with softmax normalization will increase the computational cost and result in slower training. Therefore, this embodiment adopts fast normalization fusion. The formula for calculating the fast normalization fusion weight is as follows:
[0046]
[0047] where α i represents the weight of the feature map to be fused, ω i represents the weight coefficient of the current feature map learned by the network. By passing through the ReLU activation function, it is ensured to be a value greater than 0. ε is a small value greater than 0 to avoid numerical instability, usually taking 0.0001. The advantage of this calculation is to normalize the true feature map weights and avoid the computationally expensive softmax operation. To further improve efficiency, after pooling or upsampling feature maps of different sizes and then performing pixel-wise addition with feature weights, a depthwise separable convolution is required. The depthwise separable convolution block consists of a depthwise convolution with a kernel size of 3 and a pointwise convolution with a kernel size of 1. Taking the M3 feature map in Figure 12 as an example, the calculation formula is as follows:
[0048]
[0049] where dsconv represents the depthwise separable convolution operation, ω1 and ω2 are the weights of the third-layer input feature and the fourth-layer input feature respectively, ε is a small value greater than 0, and are calculated by horizontal connection of C3 and C4 respectively, and upsample represents the upsampling operation. Inspired by BiFPN, this embodiment replaces the feature fusion module in CPN with a four-layer BiFPN.
[0050] After passing through BiFPN, the information of feature maps of different sizes has been fully fused, but the spatial information has not been emphasized and effectively utilized. In the key point detection task of this embodiment, the accuracy requirement for the spatial position of the key points of the house type is relatively high. In the feature pyramid, features with higher resolution and larger feature map sizes often contain richer position information. In the larger feature map, spatial information can be effectively transmitted through the direct connection of BiFPN, but the feature map upsampled from the top of the pyramid may lose some spatial information. To enhance the expression of this part of the feature information, after each feature fusion in BiFPN, a spatial attention weight is calculated and multiplied on the fused feature map. The spatial attention calculation formula is as shown in the following formula:
[0051] M s (F) = sigmoid(conv 7*7 (concat(AvgPool(F), MaxPool(F))))
[0052] The output of the final spatial attention module is as shown in the following formula. In the processing of spatial attention in this embodiment, the feature map directly multiplied by the feature weight is not directly output. To make the attention module more stable, the final output is the superposition of the original feature map and the feature map multiplied by the spatial weight;
[0053] F out = F + F * M s (F)
[0054] As described above, this embodiment uses BiFPN with spatial attention to replace the feature pyramid structure in CPN, and the final structure is as Figure 13 shown, where the SAB module structure is as Figure 14 shown.
[0055] Step 24. According to the house type scale features after feature fusion in step S23, the binary cross-entropy loss is used to measure the difference between the predicted value and the true value. The calculation formula of the binary cross-entropy loss function is:
[0056]
[0057] where L BCE is the binary cross-entropy objective function, represents the true value of the nth category at the pixel position (i, j), is the confidence value at the same position;
[0058] However, in the detection task of this embodiment, there are at most 100 key point positions preset in one feature map. If the L2 loss is used for calculation as in the human key point detection task, it will be difficult to calculate the loss in this task, resulting in the model being difficult to accurately learn the positions of each key point. Therefore, in this embodiment, the method of dilation operation is adopted to construct the true positions of key points to replace the Gaussian distribution. In terms of the loss function, it is more appropriate to replace the L2 loss with the binary cross-entropy loss. The heat map finally output by the network is normalized using Sigmoid and mapped between 0 and 1. The L2 loss is calculated by the following formula:
[0059]
[0060] where N is the total number of samples, y i is the true label, is the model prediction value corresponding to the sample.
[0061] The key point detection network calculates the heat map of key points. Each channel contains the position heat map of a type of house key point. The output is a set of detection heat maps containing N types of house key points. The goal is to obtain the probability distribution of key point positions consistent with the true values. The label at each position is a binary classification problem, indicating whether a key point is included. For the confidence of the output activated by Sigmoid, the binary cross-entropy loss calculation formula is as follows:
[0062]
[0063] Step 25. Adopt the non-maximum suppression method to obtain the specific classification points output by the model;
[0064] Non - maximum suppression is a crucial step in obtaining the specific classification points of the model output. It aims to predict the key point positions within a region by the model and select a point closest to the true key point as the key point within a region. In the model prediction, a neighboring region may generate range scores for a target point position. It is necessary to select a point with the highest confidence as the representative key point within a region, and then no key points of the same category will be selected within the neighborhood around this point, so as to generate only one key point within a range neighborhood for the subsequent primitive construction. For the heat map of the model results in this task, each heat map predicts a type of key point. After fixing the maximum number of predicted key points, each time the position with the highest confidence in the heat map is obtained through Argmax as a candidate key point position. The Argmax function is used to find the position of the input value that makes the given function take the maximum value. Then, starting from this point, perform a depth - first traversal of the graph, and neither the selected candidate key points nor the pixels with a confidence greater than the specified threshold will be used as candidate key point positions. Then, loop through the above steps to select multiple key point positions until the confidence of the latest selected key point position is less than the specified threshold.
[0065] The algorithm flow is shown in Table 2 and Table 3, where M is the network - predicted heat map, c is the confidence threshold, m is the maximum number of predicted key points, p is the candidate key point, and n is the point that needs non - maximum suppression.
[0066] Table 2 Algorithm flow steps of key point non - maximum suppression
[0067]
[0068] Step 26. After obtaining the key points of the wall and doors and windows, obtain the candidates of the wall and doors and windows through axial alignment within a certain threshold. Among them, the wall primitive is formed by aligning the key points of two walls, the key points of two doors form a door primitive, and the key points of two windows form a window primitive.
[0069] Step 27. Perform post - processing of vectorization on the wall primitives, door primitives, and window primitives obtained in Step 26. The specific method is: during the process of primitive selection, add the following constraints:
[0070] (1) Mutual exclusion constraint: When two primitives are close in space, especially within 10 pixels, they cannot be selected simultaneously to ensure that the two primitives are not too close.
[0071] (2) Door and window position constraint: Doors and windows must be located on the wall. For each door and window primitive, the wall primitive where this door and window are located must be found, and the door and window primitives must be forced to align with the wall primitive, otherwise this door and window primitive will be removed.
[0072] (3) Connectivity constraint: For horizontal and vertical wall connection points and all door and window connection points, the degree (i.e., the number of connections) of the connection point must match the number of candidate primitives; the specific formula is as follows:
[0073] J wall (j) = ∑P wall (p)
[0074] Where j represents the connection point, p represents the candidate primitive, and P is an indicator variable indicating whether the p-th primitive exists. During the axial connection of the wall, the primitives are formed following the principle of proximity.
[0075] (4) For the connection points of the diagonal wall, similar to doors and windows, it is preferred to connect with the axial diagonal wall connection points to form wall primitives; considering the completeness issue, it is allowed that a diagonal wall connection point and a key point of another type of wall form a horizontal or vertical wall primitive. In the processing of diagonal walls, it is encouraged to form wall primitives with more connection points of other categories. Even if there may be duplicate primitives, they will be removed during the final deduplication process of wall primitives, thus alleviating to a certain extent the problem of some missing wall primitives caused by classification errors.
[0076] The final output is close to the vector representation of the house type diagram in the real situation, but there are still some problems: the connection points are not well aligned because some coordinate errors are allowed when constructing wall and door / window primitives according to the connection points to encourage more candidate primitives. The alignment problem can be well solved by correcting in the horizontal and vertical directions with a certain threshold. In order to achieve a more beautiful and accurate vectorization result, the midpoint of the wall is taken as the midpoint position of the wall by connecting the midpoints of the starting and ending pixels of the wall that is basically close to horizontal and vertical. The horizontal and vertical walls are corrected with the coordinates of this point. Secondly, the doors and windows are not completely located on the wall. After checking the position constraints of the door and window primitives, they are aligned with the nearest wall in the same direction.
[0077] Step 3. Detect and locate the scale endpoints using the method combining Yolov8 and Shi-Tomasi corner detection, and recognize the scale digital text through the pre-trained multi-modal optical character recognition model OFA-OCR;
[0078] The scale of a floor plan refers to the ratio of the actual size of a certain area of the floor plan to the pixel width of the raster image of that area. It is an important indicator for two-dimensional and three-dimensional reconstruction of the floor plan based on the floor plan. Therefore, it is very important to quickly and accurately identify and calculate the scale in the floor plan area. The scale area consists of scale markings and numbers near the markings. The scale of the floor plan is generally distributed around the outer perimeter of the central floor plan area, and the background is relatively monotonous. However, the scale area has various styles, with multiple positions for scale markings and number areas. The scale calculation process is as follows: First, the scale area is segmented from the floor plan. Then, the dimension number area and the scale marking area are segmented separately. The endpoint positions of the scale markings and their corresponding marked numbers are determined respectively, and the scale size is calculated jointly. Different from the traditional calculation of the floor plan scale, in this embodiment, not only the value of the scale needs to be calculated, but also the accurate position of the scale needs to be calculated, so that the scale can be automatically recognized and the scale size can be manually specified and adjusted in the subsequent reconstruction work. Therefore, this embodiment proposes a method to identify and calculate various styles of scales in the floor plan. This method can cover most of the scale styles in the floor plan, has a high accuracy, can obtain an accurately calculated scale without complex post-processing methods, and has strong anti-interference and generalization capabilities. Specifically:
[0079] Step 31. Segment the scale area from the floor plan. Segmenting the scale area from the floor plan is the first step in scale calculation. By observing the floor plan, it is found that the outer perimeter walls and doors and windows in the floor plan area usually form a closed area, while the scale markings are usually not completely closed. Therefore, a simple and efficient method for segmenting the scale area is proposed: First, find the closed space of the floor plan area, and then remove it to obtain the scale area around the perimeter. Taking its four areas of up, down, left, and right, the scale area images around the perimeter can be obtained for subsequent accurate calculation. The segmentation method proposed in this embodiment is a relatively simple, fast, and accurate method for segmenting the scale area, and is applicable to floor plans with only one floor plan area. In this embodiment, Opencv-python is used as the basic library for image processing. The entire method steps are as follows:
[0080] Step 311. Binary threshold segmentation: For a color floor plan, first convert it to a grayscale image, then calculate the average grayscale of the image as the threshold, and perform binary processing on the image;
[0081] Step 312. Contour detection: Find the largest outer contour, that is, the contour of the floor plan area. Specifically, use the contour finding function findContours provided by opencv to find all contours in the floor plan, and screen out the one with the largest perimeter as the contour of the floor plan area;
[0082] Step 313. Find the maximum bounding rectangle: After finding the contour of the house type area, use the boundingRect function in opencv to obtain the maximum bounding rectangle of the contour of the house type area; clear the area within the maximum bounding rectangle from the original image.
[0083] Through the above algorithm, an image containing only the scale area can be obtained, and in this way, it is very easy to obtain partial images of the scales on the four sides according to the four areas of up, down, left, and right, as Figure 15 shown.
[0084] Step 32. Perform scale endpoint area detection and scale corner point positioning
[0085] Use Yolov8 to detect the endpoint area of the scale marking line to regress the bounding box of the scale endpoint area, and then use the Shi-Tomasi corner detection to accurately locate the position of the scale endpoint area;
[0086] The two endpoints of the scale marking line determine the pixel length of a certain section of the scale. The calculation formula of the scale is:
[0087]
[0088] where L p represents the pixel length; L gt represents the real length;
[0089] Due to the diversification of the scale marking line styles, directly performing corner detection on the scale area will cause very large errors. Therefore, in this embodiment, first use the scale endpoint area detector to locate the range of the two corner point areas at both ends of the scale, and then perform detailed corner detection on the detected corner point areas to reduce the interference of non-scale endpoint areas. Finally, optimize the corner detection results to determine the final scale endpoint positions;
[0090] Yolov8 is a very excellent object detector, which is widely used in downstream tasks and has strong generalization ability. In this embodiment, Yolov8 is selected to detect the endpoint area of the scale marking line to regress the bounding box of the scale endpoint area, as Figure 16As shown; after detecting the endpoint area, the specific positions of the scale endpoints are still not accurate enough. If only the center point of the detection box is used, a large error will occur in the calculated scale. Therefore, in the detected scale endpoint area, according to the characteristics of the scale endpoints, corner detection is used to accurately locate the position of the scale endpoint area. In this embodiment, Shi-Tomasi corner detection is adopted. OpenCV provides an interface function goodFeaturesToTrack for implementing Shi-Tomasi, which can be used to implement corner detection. The parameter src among them is the picture of the segmented scale endpoint area, maxCorners is the number of corners returned for each area, which is set to 10 in this task, qualityLevel is the quality level of the detected corners. In this task, a higher corner quality is required, and setting the threshold to 0.8 is a relatively balanced value. The minimum distance minDistance between corners is set to 3, indicating that only one corner can be generated within a range of 3 pixels. After taking the average of the obtained corner coordinates, it is used as the scale endpoint of this area. As Figure 17 shown, the dots on the scale endpoint line in the figure are the average calculation results after corner detection.
[0091] Step 33. Use the OFA-OCR pre-trained model to recognize the numbers in the scale;
[0092] The numbers in the scale are usually located in the middle or the upper and lower areas of the scale line. The size of a section of the scale can be calculated through the two endpoints of the scale line and the scale numbers in the middle. There are various text and digital information in the house type diagram, and some of them are highly interfering. Traditional optical character recognition (OCR) may perform poorly in complex backgrounds. In this embodiment, OFA-OCR is selected for scale ruler number recognition, which can accurately recognize numbers and Chinese characters in both simple and complex backgrounds, and is more suitable for the house type diagram understanding task in the Chinese region. With the development of multi-modal pre-trained models, vision tasks and natural language processing are gradually converging. Multi-modal models can perform cross-modal text and picture understanding and generation. The OCR task involves two modalities, vision tasks and text tasks. Fine-tuning the model with the Transformer encoder-decoder as the core architecture pre-trained on a large-scale multi-modal dataset for downstream tasks has become a new direction for unifying vision and natural language.
[0093] OFA-OCR is an advanced Chinese multi-modal optical character recognition model based on the Transformer encoder-decoder framework. It is fine-tuned based on the multi-modal pre-trained model OFA-Chinese. Thanks to OFA which is pre-trained on visual and language data in the general domain, OFA-OCR achieves very high-quality OCR performance after fine-tuning on downstream datasets in the Chinese OCR task. It obtains better recognition results than other OCR models in broader and more complex images, reaching the top level. Conducting inference on digital recognition using the OFA-OCR pre-trained model, the results are as Figure 18 , and the boxed numbers are the positions of the recognized numbers and the recognized number contents.
[0094] Step 34. Pair the scale markings and scale numbers to accurately calculate the scale;
[0095] After scale endpoint and corner detection and scale number recognition, it is necessary to pair the scale markings and scale numbers to accurately calculate the scale. Rotate the scales in all regions to the horizontal direction. The pairing method is to find the nearest bounding box containing numbers in the upper and lower regions of the midpoint of the line segment connecting two corner points. If found, the pairing is successful and the scale value can be calculated once. If not found, this scale marking is abandoned. The advantage of doing this is to avoid the discontinuity of scale markings caused by scale styles, which may lead to errors in pairing and calculating the scale.
[0096] Conduct statistical mathematical analysis on the previously calculated scales. There may be incorrect values among them. First, calculate the standard deviation and mean of all calculated scales. The distribution of the calculated scale values can be approximately regarded as a normal distribution. Screen out the scale values within the range of plus or minus one standard deviation. These scale values that pass the one-standard-deviation test can be approximately regarded as the scales of this floor plan. Then, perform an average operation on them to obtain the final scale value. The calculation formulas for the standard deviation and mean are shown in the following formulas:
[0097]
[0098] where N is the number of scale samples, and l i is the calculated scale value.
[0099] Step 4. Design and implement a floor plan recognition and 3D reconstruction system based on the WebGL framework, supporting the rendering of the reconstructed wall and door / window views, user interaction with the reconstructed floor plan, and the forward and backward functions of the floor plan operation history.
[0100] Specific implementation method 2. In this implementation method, the focus is on "designing and implementing a house type recognition and three-dimensional reconstruction system based on a WebGL framework, supporting the rendering of reconstructed wall and door and window views, user interaction with the reconstructed house type, and forward and backward functions of the house type operation history record."
[0101] In the development of large front-end graphic applications, data-driven views can control the flow of data in a more fine-grained manner. Views need to be supported by corresponding data models. When users operate on a view, it will trigger changes in its corresponding data model, thereby triggering the re-rendering of the view. There are two mainstream design patterns for implementing this separation of view and data: MVC and MVVM. In the MVC design pattern, the Controller is responsible for receiving user input and requests, and then dispatching the request to the corresponding Model for processing, and then triggering the update of the corresponding View object, thereby updating the view performance. This update usually requires determining the view object corresponding to a certain data pair. MVVM decouples the data exchange between user-defined data objects and real views through virtual view objects, and obtains the minimum update of view objects through the comparison algorithm of virtual view objects. This design pattern is usually applied to situations where it is difficult to determine which view objects need to be updated when updating data objects. These two design patterns can fully realize the modularization of functions, with each module maintained independently and without affecting each other. High cohesion is achieved within the module and low coupling is achieved between modules.
[0102] The main project is developed using Vue and Element Plus, using Typescript. The graphics rendering part will be used as a dependency package of the front-end system display project. After construction, it will be installed and imported by the front-end main project. Referring to the front-end MVC design pattern, the entire graphics rendering module consists of the following packages: VisualModule package, Schema package, Interaction package, History package, and RenderApp package.
[0103] VisualModule Package: It mainly stores the encapsulation of 2D house type scenes and 3D scenes, which are divided into 2D view models and 3D view models, and includes the following core classes: Visual2dModule, which is used to integrate the applications, containers, and resource loading provided by PixiJS, and receive user interactions with the 2D area and event notifications from the scene; Visual3dModule, which is mainly used to integrate functions such as scenes, cameras, lights, materials, textures, and resource loading of Babylon, and receive various events from data objects in the current scene; VisualWall2d and VisualWall3d are respectively the encapsulations of the 2D and 3D view models of the wall model, which are used to calculate wall geometric data and render wall views in 2D and 3D scenes; VisualWallAttachment2d and VisualWallAttachement3d are respectively the encapsulations of the door and window view models, which are used to calculate the geometric properties of doors and windows and render door and window views in 2D and 3D scenes; VisualRoom2d and VisualRoom3d are used for rendering house views.
[0104] Schema Package: A package used to describe the data model structure of user-defined current scenes in the scene. Data objects are registered in this package to complete the unified serialization and deserialization operations of custom data objects such as walls, doors, and windows. It is a collection of all data objects describing the current house type. It contains important classes: SceneSchema, which is used to describe the collection of data objects and the relationships between data objects in the current scene. It is a container for all data models of walls and doors. Different from the Scene in Babylon.js, SceneSchema describes the relationships between data object elements in the current house type defined by the user, rather than the relationships between view objects specifically rendered in the scene; EntityWall and EntityWallAttachment are the data models of walls and doors, which encapsulate the inherent properties such as the length, thickness, starting position, ending position, and height from the ground of walls and doors.
[0105] Interaction Package: It mainly provides interactions and operations on the 2D scene view, receives and processes user interactions with the scene, and changes the corresponding data objects of walls and doors. It includes core classes: InteractionDrawWall, which is used for wall drawing; InteractionEdit, which is used for property editing of walls and doors; InteractionTransport, which is used for dragging walls and doors; InteractionScale, which is used to specify the background scale of the current house type.
[0106] History Package: Used to manage the historical state of data in the current scene, providing the functions of undo and redo for the data in the scene. Based on the data changes in the scene, it caches the historical state of data objects in the current scene and depends on the data definition Schema Package. It includes the core class HistoryManager, which is mainly used for caching data objects in the current scene tree and comparing the old and new scene trees, thereby realizing the forward and backward movement of the historical record of the current scene tree.
[0107] RenderApp Package: Used to integrate the content of the above packages, encapsulating them into an integrated class as the entry of the rendering module for external projects to use. It contains the core class RenderApp, which is unified and exposed for external use for the integration of the above packages.
[0108] The BFF backend service is developed using Nest.js and communicates with the frontend service as the main backend service. Usually, only one backend is required for a frontend project, and all internal communication is solved by the communication between backend servers. It mainly includes the core classes FileService, which is used for uploading floor plan images, and FloorPlanService, which is used for persistent storage of floor plan design.
[0109] The floor plan recognition module is developed and deployed separately, and uses FastAPI to provide services externally. As a separate service of the backend system, it is called by other required modules through network requests to achieve decoupling between modules, making it more convenient to deploy AI-related services. The floor plan recognition and vectorized data are transmitted to the BFF backend in JSON data format, and then returned to the frontend after corresponding business encapsulation. The floor plan recognition module contains the core class PredictService, which is mainly responsible for loading and inferring the scale recognition model and the floor plan recognition model. The overall system architecture is as Figure 19 shown.
[0110] Design of the floor plan recognition and reconstruction module: The main function of the module is to recognize the raster floor plan image uploaded by the user, draw and render the recognition result on the frontend, and synchronously construct and render a 3D floor plan model that can be displayed in real time in the 3D scene based on the reconstructed 2D vectorized floor plan, so as to achieve the effect of reconstructing a 3D floor plan from a 2D raster floor plan image.
[0111] The flow chart of floor plan recognition and reconstruction is as Figure 20As shown, the user first needs to upload the floor plan to be recognized, package the picture as FormData and send it to the backend for scale recognition of the floor plan and reconstruction and vectorization of wall and door / window primitives. Subsequently, the picture is loaded into a 2D canvas element, and the background floor plan picture is loaded to the center of the window using the AssetsLoader provided by PixiJS, and is centered and scaled according to the width and height of the picture, so that the floor plan uploaded by the user can be located exactly in the center of the user's visible area. The server calls the method of Specific Embodiment 1 for vectorization reconstruction and scale calculation of the floor plan walls and doors / windows. When the BFF receives the front-end request, it will upload the user picture to OSS and forward the request and the OSS address of the picture to the processing module of the floor plan recognition service. Subsequently, the recognized scale and floor plan primitive information are encapsulated in JSON data format. After receiving the response, the front end will first pop up the position and number of the scale ruler for the user to confirm, in order to facilitate fine-tuning of the scale area. After confirmation, the graphics rendering module will create data objects and view objects of the walls and doors / windows according to the vectorization results and display them in the canvas element. If the floor plan uploaded by the user originally has no scale information, a default scale ruler will appear in the center of the 2D canvas. After the user manually specifies the scale, the true size of the recognized primitives is calculated.
[0112] The sequence diagram of the floor plan recognition and reconstruction module is as Figure 21 shown. Among them, the main floor plan reconstruction application and the rendering application are located in the Web front end. The main floor plan reconstruction application is responsible for the page UI, and the rendering application is responsible for the graphics rendering of the 2D and 3D areas of the floor plan. The floor plan upload service and the floor plan recognition service are located in the backend. Among them, the floor plan upload service is located in the backend BFF and undertakes general requests from the main application. The floor plan recognition service is an independently deployed AI floor plan recognition and is called through other backend services.
[0113] In this embodiment, asynchronous communication between the BFF service and the AI server is achieved through the message queue. The message producer and the consumer are decoupled by the message queue. Multiple consumers compete to process the information in the same queue, achieving the effect of load balancing and playing the role of peak shaving and valley filling for traffic. The information consumer processes the information through the message queue and notifies the message producer in the form of an asynchronous network request. This method makes full use of asynchronous message notification, increases the concurrency of the system, is a relatively mature solution, and can be well compatible with both traditional deployment and small-scale containerized deployment. Moreover, this method has little invasiveness and low implementation cost. Considering the mainstream message queue middleware on the market and the compatibility with Nest.js and FastAPI, RabbitMQ is considered as the message queue for this system, and the Work Queue working mode is adopted. This working mode is suitable for distributing time-consuming tasks among multiple workers. After receiving a user request, the BFF will return a unique TaskId to the front end. The front end uses the TaskId to poll the BFF to check if the task is completed. Then the BFF sends the TaskId and the picture address to the message queue together. Multiple AI servers will listen to the bound message queue and compete to process the house type recognition message. The prefetch is set to 1. After processing one task, a new message will be retrieved from the queue for inference. After the inference is completed, to ensure the accuracy and reachability of the message, the method of manually acknowledging the message is adopted to ensure that the message has really been consumed by the worker. To increase the concurrency of the AI service, the message will be acknowledged to the message queue asynchronously. Every time the AI service completes the inference of a task, it will send a network request to notify the BFF service of the TaskId and the recognition result and cache them. When the front-end polling arrives, the result will be returned to the front end, and this TaskId and the result will be deleted from the cache.
[0114] After the house type diagram is recognized and reconstructed, the user may make certain changes to the house type, such as changing the walls and doors and windows in the house type diagram. Therefore, interactive design is carried out on the generated house type diagram in the design tool, enabling the user to change the positions and styles of the wall and door and window elements in the house type. By registering mouse click and drag events in the canvas, the wall or door and window primitive object selected by the user is determined according to the coordinate position. Then, in the mouse drag event, the user's mouse coordinate position is obtained in real time, and the relevant attributes of the data object being operated by the user are uniformly processed through the corresponding interactive object, triggering a node change event to change the view object corresponding to the data object, so as to achieve a series of interactive effects of the user on the house type primitives such as walls and doors and windows.
[0115] The user wall and door and window interaction flowchart is as Figure 22As shown below. First, determine the current interaction object and interaction category, then change the relevant attributes of the data object through the corresponding interaction logic, and then publish a data node update event to update the view. This system contains two parts of view objects, 2D and 3D. The change of the wall and door / window data objects triggers a notification so that both parts of the view objects receive the notification that they need to be re-rendered. This is a specific implementation of the publish-subscribe pattern. Both VisualModule2d and VisualModule3d listen for data node change events from SceneSchema. The sequence diagram of the wall and door / window interaction module is shown in Figure 23 , when dealing with the interaction between the wall and the door / window, it completely processes the mouse events of the user in the Canvas area based on the rendering application in the Web front-end. InteractionManager implements the callback of PixiJS for 2D view interaction, and determines the view object being interacted with in the current user mouse event according to the id of the registered user event callback information. Then InteractionManager judges the operation of the user on the housing unit primitive view according to the user interaction behavior, and adopts different data object update logics. Some primitive data objects may be geometrically related to other data objects. When walls are connected to each other, and there are doors and windows on the walls, etc., they also need to be found out as data objects that need to be updated at the same time. After changing the geometric attributes of the relevant data objects through InteractionManager, the data objects that need to be updated are dispatched with node data update events through SceneSchema. By receiving the changed nodes in VisualModel, multiple view objects are updated at one step to improve the rendering performance.
[0116] The design of the household type design history record module. When users design their household types, they often need to redo or undo previous operations due to different inspirations or drawing errors. Serializing all data objects and saving them in the backend database or file, and then fully updating the scene data of the entire canvas through front-end and back-end network communication requests to achieve the re-rendering of the household type reconstruction view is a very time-consuming task. When the scene is complex and the household type is large, it will cause a certain degree of page lag, and network communication is relatively time-consuming. Based on the data structure of nodes in the scene, this system implements an incremental update method without persistence, which can be updated incrementally in the front-end without sending more time-consuming network requests to obtain persistent data. When the user changes the position of a wall or a window, or adds some nodes, first serialize the scene tree nodes, then save the historical nodes to the node pool, and calculate and cache the Hash tree of the serialized scene tree nodes. The Hash value is used to record the unique state of the nodes in the current rendering scene. The Hash value of each node in the Hash tree identifies the unique state of the corresponding node object in the scene tree. When the user clicks the forward and backward buttons of the historical progress, the differences between the two hash trees will be compared, and it will be calculated which data node objects need to be updated, added, or deleted. Then, the nodes are taken out from the historical data object node pool to restore the historical record state. The flowchart for saving the historical record is shown in Figure 24 As shown, after the user performs view interaction on the household type nodes, similar to the first half of the interaction module, it is also necessary to find the view object, data object, and their associated data objects of the wall or window node that the current user is operating on. After the user interaction is completed, the household type element node data object change event is dispatched together through SceneSchema. The HistoryManager in the historical record module receives the node data object change event from SceneSchema, serializes the current scene tree, and caches the serialized scene tree in the current scene.
[0117] The flowchart for restoring the historical record is shown in Figure 25As shown in the figure, when the user clicks the forward or backward button of the history, the historical scene cache queue will retrieve the Hash tree of the target scene, and compare the differences between the two Hash scene trees through a comparison algorithm. There are four types of changes in the data object nodes as follows: added nodes, deleted nodes, nodes with changed self-properties, and nodes with changed parent nodes. And the changed nodes are sorted according to the following rules: the nodes that need to be added in advance depending on the change of the parent node, the nodes with changed parent nodes, the nodes that need to be deleted, the nodes added normally, and the nodes with changed self-properties. The data nodes that need to be updated are added to the current scene, and their view objects are recreated and re-rendered. Since there is a serialization operation of the model object when saving the history, some properties of the model object are values calculated according to the inherent properties of the model during scene loading, and these calculated properties do not need to be serialized. The method adopted in this embodiment is to use the decorator of Typescript, and use the decorator pattern to define the Serializable decorator as a serializer and a deserializer on each property that needs to be persistently serialized, record the meta-information of the decorated property, and manually control the serialization and deserialization processes of data objects such as house type walls and doors and windows. Each time the data objects of the walls and doors are serialized, they will be serialized according to the properties owned by the decorated property metadata. The same goes for deserialization, and property deep cloning is performed according to the decorated property metadata. The serialization and deserialization of data object nodes in the scene and the comparison operation of the scene tree are CPU-intensive tasks. Since the Web front-end language used is single-threaded JavaScript, it is easy to cause the page to freeze and lose user interaction response. Therefore, the operation of saving the history is placed in the micro-task queue of JavaScript and executed asynchronously to avoid blocking the JavaScript main thread, thereby causing the browser page to freeze or lose response. At the same time, in order to simplify the serialization operation of the tree structure object in this embodiment, the data nodes in the scene tree are flattened, and the tree structure of the scene tree is converted into a flattened array, and all nodes in the array maintain indexes of the parent-child node relationship in the original scene tree through the parent and children properties. After serialization, the data nodes in the scene tree are still relatively large. If the scene tree is directly cached and updated in a full-replacement manner during the history operation, all the old data object nodes of the old scene tree will be deleted, and then the new scene tree will be reloaded to render the 2D and 3D views, which is a very time-consuming operation. When the scene tree is large, high-frequency undo and redo operations are likely to cause the browser to freeze. From a macroscopic perspective, it is not scientific to update all nodes of the scene tree in a full amount. During an operation process, most nodes of the scene tree are exactly the same as the old node scene tree. Therefore, a Hash-Diff algorithm is used to compare the differences between two serialized scene trees, so as to achieve incremental update of the scene tree. The algorithm process is as follows:
[0118] (1) Calculate the Hash values of all data object nodes in the serialized scene tree to obtain a scene Hash tree, which is used to represent the unique state of the data nodes contained in the original scene tree;
[0119] (2) Save the calculated scene Hash tree using a queue and add the Hash tree to the queue;
[0120] (3) Construct a historical node cache pool Map with the node hash values of all data object nodes in the scene tree as the index and the data object nodes as the value, so that the data object node can be found through the node hash value. All data object nodes that have appeared in the historical records are stored in this node pool.
[0121] Through the above Hash algorithm, the bloated scene tree can be simplified into a Hash tree with exactly the same structure and indexed by Hash values. All nodes in the Hash tree can be indexed in the node pool through the hash values of the data object nodes.
[0122] When the undo and redo operations are triggered, the pointer indicating the current scene tree will move forward or backward to find the Hash tree in the previous scene, and perform the Hash tree Diff algorithm between the Hash trees in the new and old scenes to compare which nodes in the two scene trees have changed. The Diff algorithm is as Figure 26 , and the time complexity and space complexity of the Diff algorithm are both O(N), where N is the number of data object nodes in the scene tree. After finding the data nodes in the historical records, these nodes contain all the data for creating their view objects, and they are re-added to the current scene tree, and their view objects are re-created and rendered on the screen.
[0123] Table 4 User Information Table
[0124]
[0125] Table 5 Housing Type Design Record Table
[0126]
[0127] Through the above description, the requirement for a quick operation history record with a short frequency has been solved, mainly solving the problem that designers may need to quickly roll back to the previous house type design state due to operation errors when designing the current house type. When the user has designed the house type or pauses work and hopes to save the current house type state, different from the forward and backward operations of the historical record state, the forward and backward operations of the design historical record are completely based on the operations in the front-end browser memory. Once the user refreshes the browser, the design historical record and the house type design will completely disappear. Therefore, a remote persistent storage solution for the house type design is needed to persistently store the house type design that the user hopes to save, and this persistence is based on the historical record state of the current design house type that the user actively hopes to save. In this embodiment, the MySQL database is selected as the persistent data storage for the house type design, saving the user's house type design plan, verifying with the login Token of the current user, and storing the serialized data of the current house type as a string in the database.
[0128] The design of this persistent database for house type design involves two tables, namely the user table and the house type design record table. The detailed designs of the two tables are shown in Tables 4 and 5. The key data of the user is the username and password. The username uniquely identifies a user, and the user ID is the unique ID assigned by the server to each user. The house type history record table uses the username as the primary key to record the design plan data of each user. At the same time, it also needs to record the plan name and the plan design time. The active field is a record soft deletion control field. When the user deletes a record, it does not really delete the row record, but sets this field to 0 to indicate an unavailable state, in order to maintain the persistence and integrity of the design data.
[0129] According to the above description, the house type reconstruction system is implemented. The system development environment is shown in Table 6.
[0130] Table 6 System Development Environment Table
[0131]
[0132]
[0133] After the system is set up, this system mainly has 3 functional modules, namely the house type recognition and reconstruction module, the house type wall and door / window interaction module, and the house type design historical record module.
[0134] House type recognition and reconstruction module: The function of this module is the house type recognition and reconstruction module. The user first needs to upload a house type drawing file, click the select file button under the Background column on the right, which supports picture files with.png,.jpg,.jpeg format suffixes, and the file size cannot exceed 2MB, and the resolution of the house type drawing uploaded by the user cannot exceed 1920×1080.
[0135] The user selects the floor plan to be identified and uploads it, which is packaged as a FormData object and sent to the BFF backend. The BFF backend receives the image upload request, saves the image on OSS, and forwards the image address and the TaskId of the current task to the message queue, and returns the TaskId and image address to the frontend. The frontend needs to use TaskId to query the BFF service every 5 seconds whether the current task is completed. After waiting for the backend AI floor plan recognition to complete and return a response to the network request, the BFF service will return the recognition result to the frontend at the next frontend polling. The floor plan selected by the user will be loaded into the central 2D canvas using AssetsLoader, and the scale recognition ruler will be rendered, corresponding to the scale line position and scale number size. If there is a correctly identified scale result in the current floor plan, the scale ruler area will appear at the identified scale position, such as Figure 27 shown.
[0136] If the current floor plan does not have a recognizable scale, the scale will appear in the center of the floor plan. The user can drag the scale and click and drag the circle buttons on both sides of the scale to adjust the length of the scale. After the user confirms that the position is correct, use the keyboard to click the Enter key and release it, and the reconstructed 2D floor plan will be drawn in the 2DCanvas area. Figure 28 As shown, the 3D reconstruction result of the apartment type will also be displayed in the upper right corner. By clicking the Switch View button in the upper left corner, you can switch the display position of the 2D and 3D views, as shown in the figure below. Figure 29 After the recognition is completed, the right panel will show the transparency and display ratio of the floor plan background. Lower the transparency of the background image, as shown in the figure below. Figure 30 ,After the user completes the recognition and drawing of the floor plan, the ,transparency of the background image can be reduced to reduce the interference of the ,background image, or the scale background image can be deleted to allow the user to focus on the ,floor design.
[0137] For more 3D reconstruction results of apartment types, see Figure 31 , Figure 31 (a) is the reconstruction result of the apartment with inclined walls. Figure 31 The middle (b) is the reconstruction result of a three-bedroom, one-living room, one-bathroom apartment. Figure 31 The middle (c) is the reconstruction result of a small one-bedroom apartment. Figure 31 Middle (d) is the reconstruction result of a two-bedroom, one-living room, one-bathroom apartment. Figure 31 The basic content includes the common housing types in urban areas of my country. Figure 31 As can be seen from the four figures, this embodiment performs 3D reconstruction based on the two-dimensional vectorization results of the apartment type, and the effect is relatively excellent. Complex apartment types with slanted doors and windows can also be accurately restored.
[0138] Household wall and door / window interaction module: After the user reconstructs the household layout, they may need to modify the structure of the layout, change the wall structure of the household, or change the size or type of doors and windows. Buttons for drawing walls and adding some common doors and windows are provided on the right side of the page. The user can click the buttons to activate the corresponding addition functions and draw in the 2D canvas area. The added door and window models will also be displayed in the 3D view. After clicking on a wall in the 2D canvas, it will become a highlighted blue part, and a black box interaction prompt will pop up above it. This interaction prompt can be dragged to move the position to the left. For walls, deletion operations and wall merging operations can be performed. For example, Figure 32 .
[0139] The deletion operation can delete this wall. For example, Figure 33 , and the doors and windows on the wall will be deleted, and the wall views connected to both ends of this wall will be redrawn. Merging walls will axially merge the axially aligned walls adjacent to this wall into the same wall surface. At this time, the user can click on the green dot interactive area at both ends of the selected wall with the mouse and axially drag the edge of the wall to change the length of the wall. For example, Figure 34 , or the wall can be dragged along the cross-axis direction of the wall to change its position. The dragging process is shown in Figure 35 . When the user drags the wall along the cross-axis to the specified position and releases the left mouse button, the position of the wall will change to the position of the mouse cursor area. At the same time, the lengths of the two walls adjacent to this wall will also change to the current wall position. For example, Figure 36 .
[0140] This system provides functions for drawing walls and doors / windows, and restricts that the positions of doors and windows must be limited to the walls. Three types of doors, three types of windows, and opening interaction buttons are provided on the right side, allowing the addition of doors and windows to the household layout. Taking the sliding door as an example, after clicking the button, a 2D picture of the sliding door drawn using PixiJS will be displayed at the mouse pointer. Move the mouse to the position where you want to add this door and click the left mouse button to add the corresponding door to the mouse pointer position. At the same time, this door will enter the highlighted blue selected state. As shown in Figure 37 . By controlling the 2D view, the 3D structure of the household layout will also be synchronously loaded and updated with the model, and the outline of the applied door model will also enter the highlighted blue selected state. Similar to the axial interaction of the wall, dragging its edge can change the length, but it does not provide cross-axis dragging like the wall. Doors and windows have more strict geometric and semantic constraints.
[0141] Household Design History Record Module: In this module, users can restore their historical operations. The default number of forward and backward operations is 20 times. Each time a user changes the attribute value of any primitive node in the current scene, an operation to save the historical record will be triggered. When the user wants to restore the previous operation, in the right-side History Functions option bar, click the Undo button on the right to restore the previous state of adding, deleting, or modifying the nodes in the current scene. For example, click to select the wall at the lower left and drag it to the right. After that, as Figure 38 .
[0142] If you want to restore the previous position state, click the Undo button to restore the state before the drag as Figure 39 . Similarly, if you want to redo the just-performed operation, click the Redo button at this time to redo the operation just performed by the user, that is, restore the state shown in 38. If the user wants to completely abandon the design of the current household type and start over, they can click Clear Floor in the upper left corner of the 2D household type rendering area or Clear in the right-side History Functions option bar to completely clear all the data model objects and view objects of the walls and doors and windows in the current scene household type area, and save a historical record once.
[0143] When the user needs to persistently save the historical record, they can click the Save Plan button. If the current user is not logged in, they need to first enter their account and password in the dialog box to log in to obtain the user's logged-in state for saving the record, as Figure 40 shown. After verification by the backend user table, a Token will be issued to the user through JWT. The frontend will attach the Token information in the request header every time it sends a request to the backend in the future, so that the backend does not need to save the user's logged-in state to reduce the server storage pressure. When the user logs in successfully and then clicks the Save Plan button, a dialog box will pop up asking the user to enter the name of the current plan. After the user confirms, the user's current design plan will be serialized and sent to the backend in the form of a string and stored in the household design record table.
[0144] When the user logs in successfully, the persistent historical record of the household design will appear below the right-side toolbar. The record contains the name of the user's household design plan and the design time, and the records are sorted from new to old by time, as Figure 41Lower right corner. After the user saves the plan, all design plans will be retrieved again to keep the latest data display. When the user selects each record, a dialog box will pop up, indicating that the user can apply this design plan record or delete this design record. When the user needs to apply the design plan record, since all data in the current house type design area will be cleared, when the house type design is not empty, the user will be prompted to save the plan in advance. When the user needs to delete the plan, the corresponding record row in the database will perform a soft delete to facilitate the administrator to query or directly manage the design data and maintain the persistence and integrity of the data.
Claims
1. A method for identifying and three-dimensionally reconstructing a housing unit type based on a raster image, characterized in that, It includes the following steps: Step 1. Collect and screen high-quality floor plan drawings suitable for the national conditions of the Chinese region, and construct a raster floor plan vector dataset; Step 2. Based on the key point detection network, identify the walls, doors, windows, and scales in the floor plan, and generate and screen candidate primitives through the axial alignment rules and constraint conditions; Step 3. Use the combination of Yolov8 and Shi-Tomasi corner detection to detect and locate the scale endpoints, and perform scale digital text recognition through the pre-trained multi-modal optical character recognition model OFA-OCR; Step 4. Design and implement a floor plan recognition and 3D reconstruction system based on the WebGL framework, which supports the rendering of the reconstructed wall and window views, the interaction of the user with the reconstructed floor plan, and the forward and backward functions of the floor plan operation history record.
2. The method for identifying and three-dimensional reconstructing a housing unit type based on a raster image according to claim 1, wherein: In Step 2, the specific method of identifying the walls, doors, windows, and scales in the floor plan based on the key point detection network and generating and screening candidate primitives through the axial alignment rules and constraint conditions is as follows: Step 21. Define the key points of the raster floor plan elements; The main elements that make up the floor plan structure are the walls and doors and windows. Encode the main elements that make up the floor plan structure as a group of connected points with categories. Define that there are 4 connection types of walls in the wall structure: I-shaped, L-shaped, T-shaped, and cross-shaped; define that there are 8 categories of connection points for doors and 8 categories of connection points for windows; Step 22. Use ConvNeXt-B pre-trained on Image-21K as the basic feature extraction network to extract multi-scale features of the floor plan; Step 23. According to the multi-scale features of the floor plan extracted in Step 22, use a bidirectional feature pyramid network for feature fusion of the neural network. After each feature fusion in BiFPN, calculate a spatial attention weight and multiply it on the fused feature map. The spatial attention calculation formula is: M s (F) = sigmoid(conv 7*7 (concat(AvgPool(F), MaxPool(F)))) where F is the feature map for which spatial attention needs to be calculated, Ms is the spatial attention feature map, which has only one channel and the same feature width and height as F, conv7*7 is a convolutional layer with a convolutional kernel size of 7, concat represents the concatenation operation by channel, AvgPool is average pooling in the channel direction, and MaxPool represents max pooling in the channel direction; Step 24. According to the floor plan scale features after feature fusion in Step S23, use binary cross-entropy loss to measure the difference between the predicted value and the true value. The binary cross-entropy loss function calculation formula is: where L BCE is the binary cross-entropy objective function, represents the ground truth value of the n-th class at the pixel position (i, j), is the confidence value at the same position; Step 25. Use the non-maximum suppression method to obtain the specific classification points output by the model; Step 26. After obtaining the key points of the walls and doors and windows, obtain the wall and door and window candidates through axial alignment within a certain threshold. Among them, the wall primitive is formed by aligning the key points of two walls, the key points of two doors form a door primitive, and the key points of two windows form a window primitive.
3. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 2, characterized in that: In the method of identifying the walls, doors, windows, and scales in the floor plan based on the key point detection network in Step 2 and generating and screening candidate primitives through the axial alignment rules and constraint conditions, it also includes: Step 27. Perform post-processing on the wall primitives, door primitives, and window primitives obtained in Step 26. The specific method is as follows: During the process of selecting primitives, add the following constraints: (1) Mutual exclusion constraint: When two primitives are close in space, especially within 10 pixels, they cannot be selected simultaneously to ensure that the two primitives are not too close; (2) Door and window position constraint: Doors and windows must be located on the wall. For each door and window primitive, the wall primitive where this door and window are located must be found, and the door and window primitives must be forced to align with the wall primitive. Otherwise, remove this door and window primitive; (3) Connectivity constraint: For horizontal and vertical wall connection points and all door and window connection points, the degree of the connection point must match the number of candidate primitives; (4) For the connection points of inclined walls, similar to doors and windows, they are preferentially connected to the axial inclined wall connection points to form wall primitives.
4. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 1, characterized in that: In Step 3, the specific process of detecting and positioning the scale endpoints using the combination of Yolov8 and Shi-Tomasi corner detection and performing scale digital text recognition through the pre-trained multi-modal optical character recognition model OFA-OCR includes: Step 31. Segment the scale area from the floor plan; Step 32. Perform scale endpoint area detection and scale corner point positioning; Use Yolov8 to detect the scale marking endpoint area to regress the bounding box of the scale endpoint area, and then use Shi-Tomasi corner detection to accurately locate the position of the scale endpoint area; Step 33. Use the pre-trained model of OFA-OCR to recognize the numbers in the scale; Step 34. Pair the scale markings and scale numbers to accurately calculate the scale.
5. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 4, characterized in that: The specific method for segmenting the scale area from the floor plan in Step 31 is as follows: Step 311. Binary threshold segmentation: For a color floor plan, first convert it to a grayscale image, then calculate the average grayscale of the image as the threshold, and perform binary processing on the image; Step 312. Contour detection: Find the largest outer contour, that is, the contour of the floor plan area. Specifically, use the contour finding function findContours provided by opencv to find all contours in the floor plan and filter out the one with the largest perimeter as the contour of the floor plan area; Step 313. Find the largest bounding rectangle: After finding the contour of the floor plan area, use the boundingRect function of opencv to obtain the largest bounding rectangle of the floor plan area contour.
6. A house floor plan recognition and 3D reconstruction system based on raster images, which is implemented based on the method for house floor plan recognition and 3D reconstruction based on raster images according to any one of claims 1-5, is characterized in that, Including: A floor plan recognition and reconstruction module, which is used to receive the floor plan file uploaded by the user, recognize and reconstruct the 2D floor plan file uploaded by the user, and generate a 3D floor plan view; A floor plan wall and door / window interaction module, which is used to perform addition, deletion, and modification operations on walls, doors, and windows in 2D and 3D views; A floor plan design history record module, which is used to record the user's design operation history and support undo, redo, save, and load design schemes.
7. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 6, characterized in that: The floor plan recognition and reconstruction module includes: A file upload unit for receiving the house type diagram file uploaded by the user, supporting.png,.jpg,.jpeg formats, with the file size not exceeding 2MB and the resolution not exceeding 1920×1080; An image processing unit for sending the house type diagram file uploaded by the user to the backend for recognition and reconstruction to generate a 3D house type view; A scale recognition and adjustment unit for recognizing the scale in the house type diagram and allowing the user to adjust the position and length of the scale; A view switching unit for switching between 2D and 3D views.
8. A method for recognizing and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 7, characterized in that: The image processing unit realizes house type recognition and reconstruction through the following steps: Package the house type diagram file uploaded by the user into a FormData object and send it to the backend; The backend saves the house type diagram file on OSS and forwards the image address and task ID to the message queue; The front end periodically queries the task status according to the task ID. After the backend completes the house type recognition, it returns the recognition result and loads the house type diagram in the 2DCanvas; Generate a 3D reconstruction view according to the recognition result and display it in the upper right corner.
9. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 6, characterized in that: The house type wall and door / window interaction module includes: A wall editing unit for adding, deleting, and merging walls in the 2D view, and supporting adjusting the length and position of the walls by dragging; A door / window editing unit for adding, deleting, and modifying doors and windows in the 2D view, and the positions of the doors and windows must be limited to the walls; A 3D view synchronization unit for synchronously updating the operations in the 2D view to the 3D view.
10. A method for identifying and three-dimensionally reconstructing a house floor plan based on a raster image according to claim 6, characterized in that: The house type design history record module includes: A historical operation record unit for recording the user's design operation history, supporting undo and redo operations; A design scheme saving unit for serializing the user's design scheme and saving it to the backend database; A design scheme loading unit for loading the design scheme saved by the user and supporting deleting the saved design scheme.
Citation Information
Patent Citations
Method for automatically identifying walls in house type graph
CN110197153A
Three-dimensional model reconstruction and image generation method and device, and storage medium
CN114119839A
House type image recognition method based on region segmentation and target detection
CN118447527A
Method and device for identifying and vectorizing house type image, medium and product
CN119229467A
Scene reconstruction in three-dimensions from two-dimensional images
WO2020254448A1
Cited By
Intelligent house type image recognition method and device based on image recognition
CN120726641A
House plan vectorization method and device, storage medium and electronic device
CN121744458A
A method and apparatus for vectorizing floor plans, a storage medium, and an electronic device.
CN121744458B