Method for automatically generating house type image based on artificial intelligence

Through artificial intelligence-based methods, key information is extracted from point cloud data and building floor plans are automatically generated, which solves the efficiency and accuracy problems of manual measurement and manual drawing in the existing technology, and achieves fast and accurate floor plans generation.

CN120014104APending Publication Date: 2025-05-16INST OF FORENSIC SCI OF MIN OF PUBLIC SECURITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510085709.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art relies on manual measurement, manual drawing or use of CAD software when generating building floor plans, resulting in large dimensional errors, time consumption and complex operation, which poses obstacles to non-professional personnel and rapid drawing tasks.

Method used

Using an artificial intelligence-based method, key information is extracted from point cloud data, and floor plans are automatically generated through orthogonal projection, convolutional neural network and sequence prediction network to reduce manual operation and size errors.

Benefits of technology

It significantly improves the speed of floor plan drawing, reduces manual operation time and workload, reduces dimensional errors caused by human errors, and accurately generates multi-layer structure floor plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014104A_ABST
    Figure CN120014104A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically generating a house type image based on artificial intelligence, and relates to the technical field of image processing, and the method comprises the steps: S1, carrying out the orthogonal projection of the three-dimensional data of a house building, and obtaining a color ground projection drawing and a density ground projection drawing through the orthogonal projection; s2, utilizing a pre-trained convolutional neural network to perform feature extraction on the color earth projection drawing and the density earth projection drawing obtained in the step S1 to obtain a color feature vector FC, a density feature vector FD and a fused feature F; s3, the fused features F are sent to a pre-trained sequence prediction network for sequence prediction, and all angular point sequences are obtained; according to the plane graph modeling method, a room is regarded as a polygon, the angular point sequence of the room is predicted, and a direct and efficient mode is provided for reconstructing the indoor plane graph. The method not only simplifies the modeling process, but also improves the flexibility and accuracy of plane graph reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for automatically generating a floor plan based on artificial intelligence. Background Art

[0002] A floor plan is a floor plan showing the interior layout of a house. It is often used in forensic science, architectural design, interior decoration, and real estate sales. The main purpose of a floor plan is to clearly show the rooms and spaces of a house and their relationships. It can also be used to record the spatial environment of a crime scene and the distribution of physical evidence.

[0003] With the increasing demand for efficient, accurate and automated architectural design tools in the fields of forensic science, architectural design, real estate, smart home, etc., traditional manual drawing, manual measurement and manual design methods can no longer meet the requirements of modern architectural design for efficiency and accuracy. Especially in the field of forensic science, floor plans must objectively reflect the environmental status of the crime scene and the positional relationship between physical evidence, as well as the location and relationship of the main on-site traces, which is crucial for the reconstruction of the case scene. Therefore, AI technology, especially computer vision, deep learning, 3D modeling and automated design, is becoming a key driving force for such tasks.

[0004] With the digital transformation of traditional industries, building design, renovation, construction, and operation are gradually changing from traditional manual operations to digital and automated processes. Building information modeling (BIM), 3D modeling, virtual reality (VR) and other technologies have been widely used in building design and construction management. However, many floor plans of existing buildings are still in the form of paper drawings, old CAD files or scanned drawings, lacking a unified standardized format and accuracy.

[0005] In this context, AI-based technology can quickly extract structural information from these non-digital building materials, automatically generate accurate floor plans, and simplify preliminary preparations.

[0006] The existing methods for generating apartment plans usually include the following: 1. Manual drawing; 2. Using CAD (computer-aided design) software; 3. Using online drawing tools; 4. Using 3D modeling software (SketchUp, Blender, Autodesk Revit, etc.); 5. Using laser measuring instruments and scanning technology to obtain spatial data and generate digital apartment plans.

[0007] However, the existing technology mainly relies on manual measurement, manual drawing or the use of software tools such as CAD when generating building floor plans. These technical methods have some unavoidable technical disadvantages. First, when drawing manually, due to differences in measuring tools or personnel operations, dimensional errors are easily caused, and manual measurement takes a lot of time, especially in cases where the building area is large or the design is complex; secondly, when using software tools to generate floor plans, the professional requirements of the operators are very high, and a lot of professional training is required. This may pose an obstacle to non-professionals (such as ordinary users or novices), especially when the drawing task is required to be completed quickly in a short time. At the same time, even for experienced users, it still takes a certain amount of time to draw floor plans. In addition, the existing method of using laser measuring instruments and scanning technology to obtain spatial data to generate floor plans requires high computing resources and time.

[0008] To this end, a method for automatically generating floor plans based on artificial intelligence is provided to solve the above problems. Summary of the invention

[0009] In view of the problems existing in the above-mentioned prior art, the present invention provides a method for automatically generating floor plans based on artificial intelligence, which can automatically extract key information from existing point cloud data and automatically generate floor plans, greatly improving the drawing speed and greatly reducing the time and workload of manual operations. The AI ​​system can model based on data instead of relying on manual measurement or hand-drawing, thereby reducing dimensional errors caused by human error. For multi-story buildings, AI can not only draw single-story floor plans, but also accurately generate multi-story structures and provide detailed information for each floor.

[0010] In order to achieve the above object, the present invention adopts a method for automatically generating a floor plan based on artificial intelligence, comprising:

[0011] S1. Orthogonally project the three-dimensional data of the building, and obtain a color ground projection map and a density ground projection map through orthogonal projection;

[0012] The three-dimensional data is triangular mesh data, and the triangular mesh is composed of multiple triangles;

[0013] A triangle is made up of corner points and faces;

[0014] The corner points are the basic points that make up the triangle, and the face is the triangle defined by the corner point index;

[0015] Convert the triangular mesh data into point cloud data and downsample the point cloud data.

[0016] Among them, to convert the triangular mesh data into point cloud data, it is necessary to extract the coordinate information of all corner points in the three-dimensional mesh data, that is, to obtain the converted point cloud data;

[0017] Downsample the point cloud data using uniform downsampling;

[0018] The uniform downsampling method is to construct multiple spheres of fixed radius in the point cloud data, and then select the single point cloud closest to the center of each sphere as the representative, and discard the other point clouds; this method selectively retains the required single point cloud;

[0019] In each constructed sphere, the algorithm calculates the distance from all points to the center of the sphere and selects the point with the closest distance as the representative point in the sphere. All points in each sphere will be replaced by this closest point to achieve downsampling.

[0020] S11. Method for obtaining color ground projection map: fuse the triangular mesh data and point cloud data, use the three-dimensional coordinate (x, y, z) information contained in the triangular mesh itself, adopt a triangulation scheme, which uses the selected two-dimensional plane (x, y) to calculate the two-dimensional position of the triangle vertex, and thereby obtain the color information in the triangle piece to obtain a color ground projection map.

[0021] S12. A method for obtaining a density projection map: converting triangular mesh data into point cloud data, projecting the point cloud data onto an XY plane, and using the density value of the plane point as a representation of the point cloud density. The specific method is to divide the measurement area into small cells, and use the ratio of the number of points in the cell to the area of ​​the cell as a representation of the density to obtain a density projection map;

[0022] S2, using a pre-trained convolutional neural network, respectively extracting features from the color projection map and the density projection map obtained in step S1 above, to obtain a color feature vector F_C, a density feature vector F_D, and a fused feature F;

[0023] S3, sending the fused feature F to the pre-trained sequence prediction network for sequence prediction to obtain all corner point sequences;

[0024] S4. The training method of the pre-trained convolutional neural network and the pre-trained sequence neural network is divided into two parts:

[0025] Loss function and label file generation during training;

[0026] Among them, the loss function in training uses the Hungarian matching algorithm, which is responsible for matching the predicted detection results with the real annotations to ensure the consistency of coordinates and categories. Similar to traditional target detection tasks, the Hungarian matching algorithm needs to consider the matching of coordinates and categories. The coordinate matching of target detection is transformed from the common bounding box to a polygon. To calculate the matching error, the real annotated polygons must first be converted into a clockwise order. Then, taking each coordinate point as the starting point, the minimum error between the predicted polygon and the real polygon is calculated, and the L1 loss is used. This minimum error is used as a measure of the matching error.

[0027] Among them, in the generation of label files, the true value with certain noise is generated by manual labeling, and the true value required for network training is formed by parsing the label file;

[0028] S5, topologically organize the obtained initial plane graph to obtain an accurate corner point sequence;

[0029] The predicted image is post-processed, with dilation operations to connect close line segments and straight line correction to obtain more accurate floor plan segmentation results;

[0030] If the predicted corner points still have misalignment problems, the shortest distance algorithm is used to further sort these corner points. The Dijkstra algorithm is used. The core idea of ​​this algorithm is the greedy strategy, that is, at each step, the corner point of the currently known shortest path is selected, and then the shortest path estimate of its adjacent corner points is updated;

[0031] S6. The accurate corner point sequence of S5 is assigned closed interval semantic content through semantic recognition and semantic segmentation, and then the corner point format is converted and submitted to the client to generate the final house floor plan.

[0032] As a further optimization of the above scheme, the color feature vector F_C and the density feature vector F_D in S2 use the residual network Resnet as the feature extraction network, by restating the layer as a learning residual function;

[0033] Each residual block attempts to learn the difference between input and output. In each residual block, the input features are directly passed to the output of the block through an identity mapping.

[0034] The convolutional layers inside the residual block learn the residual between the input and output.

[0035] As a further optimization of the above scheme, the extracted color feature vector F_C and density feature vector F_D are fused in S2 using a feature pyramid network.

[0036] As a further optimization of the above solution, the feature pyramid network includes a bottom-up path and a top-down path;

[0037] Among them, the bottom-up path is part of the feedforward backbone. Each level is downsampled with step=2. The network part with the same output size is called one level. The last layer of feature maps of each level is selected as the corresponding layer of the Up-bottompathway, and the reference of element add after 1x1 convolution.

[0038] Top-down path: high-level feature maps are gradually generated through convolution and upsampling operations, and the top-level semantic information is gradually transferred to the low-level feature maps.

[0039] As a further optimization of the above solution, the feature pyramid network also includes a lateral connection, which is used to fuse the features of the color feature vector F_C and the density feature vector F_D;

[0040] The upsampling operation is used to interpolate the feature map of the higher level to obtain a feature map that matches the size of the corresponding low-level feature map, and then the color feature vector F_C and the density feature vector F_D are fused by element-by-element addition to obtain the fused feature F.

[0041] As a further optimization of the above scheme, the sequence prediction network adopts Transformer, and the extracted fusion feature F is input into the Transformer model

[0042] Using the sequence prediction function of Transformers, we directly output an ordered, variable-length sequence of corner points for each building room. Transformer and positional encoding output multiple ordered sequences of corner points, which can be used to restore the floor plan by simply concatenating them in the predicted order.

[0043] As a further optimization of the above solution, S3 also includes:

[0044] S31. Convert the floor plan reconstruction task into the problem of predicting multiple polygons, where each polygon represents an independent room and these polygons are constructed by a series of orderly arranged corner point sequences;

[0045] Each room is represented as a closed polygon, defined by a sequence of corner points;

[0046] The polygon query information is first interacted in the self-attention module, and then different regions in the density map are queried in the multi-scale deformable cross-attention module;

[0047] Predict the validity of each query position as a corner point through a shared CNN network;

[0048] A two-level query mechanism is implemented: one level for querying polygons and another level for corner points to refine the layout of predicted polygons;

[0049] The polygon set is represented as an M×N×2 matrix, where M is the maximum number of polygons and N is the maximum number of corner points per polygon;

[0050] During the training process, a polygon matching module is introduced to measure the difference between the predicted polygon and the real polygon, achieving end-to-end supervision and optimization.

[0051] As a further optimization of the above solution, S3 also includes:

[0052] S32, encode the extracted image features using a deformable attention mechanism that can flexibly focus on key areas in the image;

[0053] Decoding the acquired image features;

[0054] Decoding is implemented through a decoder, which consists of 6 stacked layers, each of which includes three main modules: a self-attention module, a multi-scale deformable cross-attention module, and a feed-forward network;

[0055] Each decoder layer receives enhanced image features from the encoder and polygon query information from the previous layer.

[0056] As a further optimization of the above solution, S3 also includes:

[0057] S33, the output of the last layer of the decoder is an M*N*C matrix. The features of each room are aggregated by averaging the diagonal features to obtain a feature matrix of M*C size after aggregation;

[0058] Finally, the matrix is ​​input into a linear projection layer and the softmax function is used to find the label probability of each room; M is usually larger than the actual number of rooms in the scene, and additional empty class labels are used to represent invalid rooms;

[0059] Each corner point is modeled as a vector consisting of three parts: cn, pn, and Ln;

[0060] pn represents the specific position of the corner point in two-dimensional space;

[0061] cn is a binary flag used to indicate whether the corner point is a valid corner point on the room boundary, where 0 indicates invalid and 1 indicates valid;

[0062] ln represents the language information of the corner point;

[0063] The output of the model is a set of corner point sequences, each of which represents a closed polygon of a room;

[0064] Once the model predicts these ordered sequences of corner points, we simply connect all the corner points marked as valid to obtain a polygonal representation of each room.

[0065] As a further optimization of the above scheme, the Dijkstra algorithm steps include:

[0066] Step 1: Initialization: Set the distance from the source point to itself to 0, and the distance to all other corner points to infinity; create an unvisited corner point set and add all corner points to it;

[0067] Step 2: Select corner points: Select a corner point closest to the source point from the unvisited corner point set and add it to the visited set.

[0068] Step 3: Update distance: Update the distance of all adjacent corner points of the current corner point. If the distance from the current corner point to the adjacent corner point is shorter than the known shortest distance, update the shortest distance of the adjacent corner point and calculate the direction of the corner point to make the direction of the corner point close to the overall direction trend.

[0069] Step 4: Repeat steps 2 and 3 until all corner points are visited or the target corner point is reached;

[0070] Through the above steps and methods, the corner points are sorted, and the connection relationship between the corner points C and the corner points, that is, the edge E and the closed interval R, is also obtained.

[0071] The method of automatically generating a floor plan based on artificial intelligence of the present invention has the following beneficial effects:

[0072] In the method for automatically generating a floor plan based on artificial intelligence of the present invention, the boundaries of the room are implicitly defined by the order of the corner point sequence, thereby eliminating the step of predicting the edges separately;

[0073] FPN fully utilizes the advantages of features at each scale by fusing features at different scales, thus improving the detection capability of the model;

[0074] By simplifying the representation of the room boundary into a sequence of corner points, the complex edge prediction step is avoided, making the modeling process more efficient;

[0075] This method can be adapted to rooms of different shapes and sizes because it allows each room to have a different number of corner points;

[0076] By accurately predicting the location and validity of corner points, an accurate representation of the room boundaries can be obtained.

[0077] With reference to the following description and drawings, a specific embodiment of the present invention is disclosed in detail, indicating the manner in which the principles of the present invention can be adopted. It should be understood that the scope of the embodiments of the present invention is not limited thereby, and within the spirit and scope of the appended claims, the embodiments of the present invention include many changes, modifications and equivalents. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 A diagram of a method for automatically generating a floor plan based on artificial intelligence according to the present invention;

[0079] Figure 2 It is the color projection map of the earth of the present invention;

[0080] Figure 3 It is the density projection map of the present invention;

[0081] Figure 4 is a corner point sequence diagram of the present invention;

[0082] Figure 5 It is the accurate corner point sequence diagram of the present invention;

[0083] Figure 6 It is the final house plan of the present invention. DETAILED DESCRIPTION

[0084] Please refer to the instruction manual Figure 1-6 The present invention provides a technical solution: a method for automatically generating floor plans based on artificial intelligence.

[0085] The present invention can be used as an important part of the on-site investigation, and can help investigators accurately record the layout of the crime scene, the location of doors and windows, the location of the body and other key information, as well as abnormal information at the scene (whether the doors and windows are damaged, the location and size of the footprints at the scene, the collapse and impact marks on the surfaces of walls and furniture, etc.). This information is of great significance for the detection, litigation and trial of the case.

[0086] The present invention uses orthogonal projection to obtain color ground projection maps and density ground projection maps through existing three-dimensional data of house buildings, uses convolutional neural networks to extract features from projection maps, obtains fusion features of building houses, uses sequence prediction networks to perform sequence prediction on the fusion features, topologically organizes the obtained corner point sequence, and performs semantic recognition and semantic segmentation to identify various functional rooms of the house building. Finally, the focus sequence format is submitted to the client through conversion, and the house floor plan is automatically generated, which greatly improves the efficiency of manually drawing the floor plan, greatly reduces the size error caused by human error, and can reflect the environment, status and relationship between objects on the site and accurately display them in proportion to form a site plane proportional map.

[0087] To achieve the above object, the present invention comprises:

[0088] S1. Orthogonally project the three-dimensional data of the building to obtain the following Figure 2 The color projection map shown in Figure 3 The density projection onto the earth shown in .

[0089] The three-dimensional data is triangular mesh data, which is composed of multiple triangles. Furthermore, a triangle is composed of vertices and faces, where vertices are the basic points that make up a triangle, and faces are triangles defined by vertices indices.

[0090] In order to obtain the color ground projection map and the density ground projection map, the triangular mesh data needs to be converted into point cloud data and the point cloud data needs to be downsampled.

[0091] To convert the triangular mesh data into point cloud data, it is necessary to extract the coordinate information of all corner points (Vertices) in the three-dimensional mesh data to obtain the converted point cloud data.

[0092] Furthermore, the point cloud data is downsampled, and the downsampling scheme adopts uniform downsampling. This method constructs multiple spheres of fixed radius in the point cloud data, and then selects the single point cloud closest to the center of each sphere as a representative, and the other point clouds are discarded. This method does not change the position of the point, but only selectively retains the required single point point cloud.

[0093] In each constructed sphere, the algorithm calculates the distance from all points to the center of the sphere and selects the point with the closest distance as the representative point in the sphere. This means that all points in each sphere will be replaced by this closest point, thus achieving downsampling.

[0094] S11. Method for obtaining color ground projection map: fuse the triangular mesh data and point cloud data, use the three-dimensional coordinate (x, y, z) information contained in the triangular mesh itself, and adopt a triangulation scheme, which uses the selected two-dimensional plane (x, y) to calculate the two-dimensional position of the triangle vertex, and thus obtain the color information in the triangle piece. Figure 2 Color projection map shown in .

[0095] S12. Method for obtaining density projection map: convert triangular mesh data into point cloud data, project the point cloud data onto the XY plane, and use the density value of the plane points as the representation of the point cloud density. The specific method is to divide the measurement area into small cells, and use the ratio of the number of points in the cell to the area of ​​the cell as the representation of the density. Figure 3 The density projection onto the earth shown in .

[0096] In particular, in order to avoid the point cloud data of the roof being projected onto the ground and causing interference with the ground projection map, the point cloud data above the shooting point needs to be deleted and truncated.

[0097] S2. Using a pre-trained convolutional neural network (CNN), feature extraction is performed on the color ground projection map and the density ground projection map obtained in the above step S1 to obtain a color feature vector F_C, a density feature vector F_D, and a fused feature F.

[0098] Among them, the color feature vector F_C and the density feature vector F_D use the residual network Resnet as the feature extraction network, by reformulating the layer as a learning residual function (the difference between input and output). Each residual block attempts to learn the difference between input and output. In each residual block, the input feature is directly passed to the output of the block through an identity mapping (or called a jump connection). The convolution layer inside the residual block learns the residual between the input and output. Specifically, assuming the input is x, the output of the residual block is F(x)+x, where F(x) is the residual function learned by the convolution layer inside the residual block. In this way, the network can learn the identity mapping F(x)=0, so that the output is equal to the input x. At the same time, jump links are also introduced to allow signals in the network to bypass one or more layers and pass directly, which helps solve the gradient vanishing problem and allows the network to be trained deeper. Convolutional neural networks (CNNs) can capture local features and spatial hierarchical information, which is crucial for understanding room layout. After being processed by multiple residual blocks, the network outputs a color feature vector F_C and a density feature vector F_D.

[0099] The extracted color feature vector F_C and density feature vector F_D are fused using the Feature Pyramid Network (FPN). The Feature Pyramid Network (FPN) is a deep learning network structure used for computer vision tasks such as object detection and semantic segmentation. It constructs a multi-scale feature pyramid, fuses the semantic information of high-level feature maps with the spatial information of low-level feature maps, and generates a feature representation with rich multi-scale information.

[0100] Feature Pyramid Network (FPN) consists of two main parts: bottom-up pathway and top-down pathway, as well as lateral connections. Bottom-up pathway: This is part of the feedforward backbone. Each level is downsampled with step=2. The network part with the same output size is called a stage. The last layer of feature maps of each level is selected as the corresponding layer of the Up-bottom pathway. After 1x1 convolution, the reference of element add is passed. Top-down pathway: high-level feature maps are gradually generated through convolution and upsampling operations, and the top-level semantic information is gradually transferred to the low-level feature maps.

[0101] At the same time, the features of F_D and F_C are fused using lateral connections. Specifically, the upsampling operation is used to interpolate the feature maps of the higher level to obtain feature maps that match the size of the corresponding low-level feature maps, and then they are fused by element-by-element addition to obtain the fused feature F.

[0102] S3, send the fused feature F to the pre-trained sequence prediction network for sequence prediction, and obtain all corner point sequences, such as Figure 4 As shown in the corner point sequence diagram.

[0103] The sequence prediction network uses Transformer, and the fused feature F extracted in step S2 above is input into the Transformer model. The Transformer model is well-known for its self-attention mechanism, which can process sequence data and capture long-distance dependencies. Using the sequence prediction function of Transformers, the ordered, variable-length sequence of corner points for each building room is directly output. Transformer and position encoding output multiple ordered sequences of corner points. The floor plan is restored by simply connecting in the predicted order.

[0104] In the area of ​​floor plan modeling, this method treats the reconstruction process as a problem of predicting a series of polygons. Each polygon corresponds to a room and is represented by an ordered sequence of corner points. The key advantage of this method is that the boundaries (edges) of the room are implicitly defined by the order of the corner point sequence, thus eliminating the step of predicting the edges separately. The following is a detailed description of the method, presented in different representations:

[0105] S31. The reconstruction task of the floor plan is transformed into the problem of predicting multiple polygons, where each polygon represents an independent room. These polygons are constructed by a series of ordered corner point sequences, and the boundaries of the room are implicit in the order of these corner point sequences, without the need to predict the edges separately.

[0106] Each room is represented as a closed polygon defined by a sequence of corner points.

[0107] These corner point sequences are of variable length, meaning that each room can have a different number of edges.

[0108] The polygon query information is first interacted in the self-attention module, and then different regions in the density map are queried in the multi-scale deformable cross-attention module.

[0109] The validity of each query location as a corner point is predicted by a shared CNN network.

[0110] A two-level query mechanism is implemented: one level for querying polygons and another level for corner points to refine the layout of predicted polygons.

[0111] A polygon set is represented as an M×N×2 matrix, where M is the maximum number of polygons and N is the maximum number of corner points per polygon.

[0112] During the training process, a polygon matching module is introduced to measure the difference between the predicted polygon and the real polygon, achieving end-to-end supervision and optimization.

[0113] S32. The extracted image features are encoded using a deformable attention mechanism that can flexibly focus on key areas in the image.

[0114] The image features acquired in the above step S32 are decoded.

[0115] The decoder consists of 6 stacked layers, each of which includes three main modules: a self-attention module (SA), a multi-scale deformable criss-cross attention module (MS-DCA), and a feed-forward network (FFN).

[0116] Each decoder layer receives enhanced image features from the encoder and polygon query information from the previous layer.

[0117] S33, the output of the last layer of the Transformer decoder is an M*N*C matrix. The features of each room are aggregated by averaging the diagonal features to obtain a feature matrix of size M*C after aggregation. Finally, the matrix is ​​input into a linear projection layer and the label probability of each room is calculated using the softmax function. M is usually larger than the actual number of rooms in the scene, so additional empty class labels are used to represent invalid rooms.

[0118] Each corner point is modeled as a vector consisting of three parts: cn, pn, and Ln.

[0119] pn represents the specific position of the corner point in two-dimensional space.

[0120] cn is a binary flag indicating whether the corner point is a valid corner point on the room boundary (0 means invalid, 1 means valid).

[0121] ln represents the language information of the corner point.

[0122] The output of the model is a set of corner point sequences, each of which represents a closed polygon of a room.

[0123] Once the model predicts these ordered sequences of corner points, we simply connect all corner points marked as valid (cn=1) to obtain a polygonal representation of each room.

[0124] S4. The training method of the convolutional neural network pre-trained in step S2 and the sequence neural network pre-trained in step S3 can be divided into two parts.

[0125] The loss function during training is:

[0126] The Hungarian matching algorithm is mainly used, which is responsible for matching the predicted detection results with the actual annotations to ensure the consistency of coordinates and categories.

[0127] Matching consistency: Similar to traditional object detection tasks, the Hungarian matching algorithm needs to consider the matching of coordinates and categories. The special feature of this method is that the coordinate matching of object detection is transformed from the common bounding box (bbox) to the polygon (poly).

[0128] In order to calculate the matching error, the real annotated polygon (GTpoly) needs to be converted into a clockwise order. Then, taking each coordinate point as the starting point, the minimum error between the predicted polygon and the real polygon is calculated, and the L1 loss is used. This minimum error is used as a measure of the matching error.

[0129] The Hungarian matching algorithm is used in object detection to ensure consistency between predictions and ground truth annotations. When dealing with polygon matching, special consideration needs to be given to the order of coordinate points and the calculation of matching errors. This paper improves the performance of object detection tasks by effectively integrating these matching errors into the loss function to guide the model to learn more accurate polygon predictions.

[0130] Label file generation:

[0131] The true value with a certain amount of noise is generated through manual labeling, and the true value required for network training is formed by parsing the label file.

[0132] S5, topologically organize the initial plane image (focus sequence) obtained in step 3 to obtain an accurate corner point sequence; Figure 5 The exact corner point sequence is shown in the figure.

[0133] The predicted image needs to undergo post-processing steps, such as dilation operations to connect close line segments and straight line correction to obtain more accurate floor plan segmentation results;

[0134] At this time, the predicted corner points still have problems such as misalignment. The shortest distance algorithm can be used to further sort out these corner points. The Dijkstra algorithm is used. The core idea of ​​this algorithm is the greedy strategy, that is, selecting the corner point of the currently known shortest path at each step, and then updating the shortest path estimate of its adjacent corner points.

[0135] Dijkstra Algorithm Steps

[0136] 1) Initialization: Set the distance from the source point to itself to 0, and the distance to all other corner points to infinity. Create a set of unvisited corner points and add all corner points to it.

[0137] 2) Select corner point: Select a corner point closest to the source point from the unvisited corner point set and add it to the visited set.

[0138] 3) Update distance: Update the distance of all adjacent corner points of the current corner point. If the distance from the current corner point to the adjacent corner point is shorter than the known shortest distance, update the shortest distance of the adjacent corner point and calculate the direction of the corner point to make the direction of the corner point closer to the overall direction trend.

[0139] 4) Repeat: Repeat steps 2 and 3 until all corner points are visited or the target corner point is reached.

[0140] Through the above steps and methods, the corner points are sorted, and the connection relationship between the corner points C and the corner points, that is, the edge E and the closed interval R, is also obtained.

[0141] S6, the corner point sequence I4 in step S5 is assigned closed interval semantic content through semantic recognition and semantic segmentation, and then the corner point format is converted and submitted to the client to generate the final house floor plan, such as Figure 6 as shown in .

[0142] The boundaries (edges) of the room are implicitly defined by the order of the corner point sequence, thus eliminating the need to predict the edges separately.

[0143] FPN fully utilizes the advantages of features at each scale by fusing features at different scales, thus improving the detection capability of the model.

[0144] Simplified process: By simplifying the representation of the room boundary into a sequence of corner points, the complex edge prediction step is avoided, making the modeling process more efficient.

[0145] Flexibility: This method can be adapted to rooms of different shapes and sizes because it allows each room to have a different number of corners.

[0146] Accuracy: By accurately predicting the location and validity of corner points, an accurate representation of the room boundaries can be obtained.

[0147] This floor plan modeling method provides a direct and efficient way to reconstruct indoor floor plans by treating rooms as polygons and predicting their corner point sequences. This method not only simplifies the modeling process, but also improves the flexibility and accuracy of floor plan reconstruction.

[0148] Advantages of decoder:

[0149] The decoder computes self-attention for all corner points, allowing interactions not only between corners of a single room, but also between corners of different rooms, thus achieving global reasoning.

[0150] In the multi-scale deformable attention module, polygon queries are directly used as reference points, enabling the network to use explicit spatial priors to aggregate features in multi-scale feature maps around polygon corners.

[0151] The number of rooms and corners is achieved by classifying each query as valid or invalid.

[0152] It can quickly generate a site-scale plan drawing, replacing the traditional on-site hand-drawn sketch and CAD drawing work mode.

Claims

1. A method for automatically generating floor plan based on artificial intelligence, characterized in that: include: S1. Orthogonally project the three-dimensional data of the building, and obtain a color ground projection map and a density ground projection map through orthogonal projection; The three-dimensional data is triangular mesh data, and the triangular mesh is composed of multiple triangles; A triangle is made up of corner points and faces; The corner points are the basic points that make up the triangle, and the face is the triangle defined by the corner point index; Convert the triangular mesh data into point cloud data and downsample the point cloud data. Among them, to convert the triangular mesh data into point cloud data, it is necessary to extract the coordinate information of all corner points in the three-dimensional mesh data, that is, to obtain the converted point cloud data; Downsample the point cloud data using uniform downsampling; The uniform downsampling method is to construct multiple spheres of fixed radius in the point cloud data, and then select the single point cloud closest to the center of each sphere as the representative, and discard the other point clouds; this method selectively retains the required single point cloud; In each constructed sphere, the algorithm calculates the distance from all points to the center of the sphere and selects the point with the closest distance as the representative point in the sphere. All points in each sphere will be replaced by this closest point to achieve downsampling. S11. Method for obtaining color ground projection map: fuse the triangular mesh data and point cloud data, use the three-dimensional coordinate (x, y, z) information contained in the triangular mesh itself, adopt a triangulation scheme, which uses the selected two-dimensional plane (x, y) to calculate the two-dimensional position of the triangle vertex, and thereby obtain the color information in the triangle piece to obtain a color ground projection map. S12. A method for obtaining a density projection map: converting triangular mesh data into point cloud data, projecting the point cloud data onto an XY plane, and using the density value of the plane point as a representation of the point cloud density. The specific method is to divide the measurement area into small cells, and use the ratio of the number of points in the cell to the area of ​​the cell as a representation of the density to obtain a density projection map; S2, using a pre-trained convolutional neural network, respectively extracting features from the color projection map and the density projection map obtained in step S1 above, to obtain a color feature vector F_C, a density feature vector F_D, and a fused feature F; S3, sending the fused feature F to the pre-trained sequence prediction network for sequence prediction to obtain all corner point sequences; S4. The training method of the pre-trained convolutional neural network and the pre-trained sequence neural network is divided into two parts: Loss function and label file generation during training; Among them, the loss function in training uses the Hungarian matching algorithm, which is responsible for matching the predicted detection results with the real annotations to ensure the consistency of coordinates and categories. Similar to traditional target detection tasks, the Hungarian matching algorithm needs to consider the matching of coordinates and categories. The coordinate matching of target detection is transformed from the common bounding box to a polygon. To calculate the matching error, the real annotated polygons must first be converted into a clockwise order. Then, taking each coordinate point as the starting point, the minimum error between the predicted polygon and the real polygon is calculated, and the L1 loss is used. This minimum error is used as a measure of the matching error. Among them, in the generation of label files, the true value with certain noise is generated by manual labeling, and the true value required for network training is formed by parsing the label file; S5, topologically organize the obtained initial plane graph to obtain an accurate corner point sequence; The predicted image is post-processed, with dilation operations to connect close line segments and straight line correction to obtain more accurate floor plan segmentation results; If the predicted corner points still have misalignment problems, the shortest distance algorithm is used to further sort these corner points. The Dijkstra algorithm is used. The core idea of ​​this algorithm is the greedy strategy, that is, at each step, the corner point of the currently known shortest path is selected, and then the shortest path estimate of its adjacent corner points is updated; S6. The accurate corner point sequence of S5 is assigned closed interval semantic content through semantic recognition and semantic segmentation, and then the corner point format is converted and submitted to the client to generate the final house floor plan.

2. A method for automatically generating floor plan based on artificial intelligence according to claim 1, characterized in that: The color feature vector F_C and the density feature vector F_D in S2 use the residual network Resnet as the feature extraction network, by restating the layer as a learning residual function; Each residual block attempts to learn the difference between input and output. In each residual block, the input features are directly passed to the output of the block through an identity mapping. The convolutional layers inside the residual block learn the residual between the input and output.

3. A method for automatically generating floor plan based on artificial intelligence according to claim 2, characterized in that: In S2, the extracted color feature vector F_C and density feature vector F_D are fused using a feature pyramid network.

4. A method for automatically generating floor plan based on artificial intelligence according to claim 3, characterized in that: The feature pyramid network includes a bottom-up path and a top-down path; Among them, the bottom-up path is part of the feedforward backbone. Each level is downsampled with step=2. The network part with the same output size is called one level. The last layer of feature maps of each level is selected as the corresponding layer of the Up-bottom pathway, and the reference of element add after 1x1 convolution. Top-down path: high-level feature maps are gradually generated through convolution and upsampling operations, and the top-level semantic information is gradually transferred to the low-level feature maps.

5. A method for automatically generating floor plan based on artificial intelligence according to claim 4, characterized in that: The feature pyramid network also includes a lateral connection, which is used to fuse the features of the color feature vector F_C and the density feature vector F_D; The upsampling operation is used to interpolate the feature map of the higher level to obtain a feature map that matches the size of the corresponding low-level feature map, and then the color feature vector F_C and the density feature vector F_D are fused by element-by-element addition to obtain the fused feature F.

6. A method for automatically generating floor plan based on artificial intelligence according to claim 1, characterized in that: The sequence prediction network uses Transformer. The extracted fusion feature F is input into the Transformer model. The sequence prediction function of Transformers is used to directly output an ordered, variable-length corner point sequence for each building room. Transformer and position encoding output multiple ordered corner point sequences, and the floor plan is restored by simply connecting them in the predicted order.

7. A method for automatically generating floor plan based on artificial intelligence according to claim 1, characterized in that: The S3 also includes: S31. Convert the floor plan reconstruction task into the problem of predicting multiple polygons, where each polygon represents an independent room and these polygons are constructed by a series of orderly arranged corner point sequences; Each room is represented as a closed polygon, defined by a sequence of corner points; The polygon query information is first interacted in the self-attention module, and then different regions in the density map are queried in the multi-scale deformable cross-attention module; Predict the validity of each query position as a corner point through a shared CNN network; A two-level query mechanism is implemented: one level for querying polygons and another level for corner points to refine the layout of predicted polygons; The polygon set is represented as an M×N×2 matrix, where M is the maximum number of polygons and N is the maximum number of corner points per polygon; During the training process, a polygon matching module is introduced to measure the difference between the predicted polygon and the real polygon, achieving end-to-end supervision and optimization.

8. A method for automatically generating floor plan based on artificial intelligence according to claim 7, characterized in that: The S3 also includes: S32, encode the extracted image features using a deformable attention mechanism that can flexibly focus on key areas in the image; Decoding the acquired image features; Decoding is implemented through a decoder, which consists of 6 stacked layers, each of which includes three main modules: a self-attention module, a multi-scale deformable cross-attention module, and a feed-forward network; Each decoder layer receives enhanced image features from the encoder and polygon query information from the previous layer.

9. A method for automatically generating floor plan based on artificial intelligence according to claim 8, characterized in that: The S3 also includes: S33, the output of the last layer of the decoder is an M*N*C matrix. The features of each room are aggregated by averaging the diagonal features to obtain a feature matrix of M*C size after aggregation; Finally, the matrix is ​​input into a linear projection layer and the softmax function is used to find the label probability of each room; M is usually larger than the actual number of rooms in the scene, and additional empty class labels are used to represent invalid rooms; Each corner point is modeled as a vector consisting of three parts: cn, pn, and Ln; pn represents the specific position of the corner point in two-dimensional space; cn is a binary flag used to indicate whether the corner point is a valid corner point on the room boundary, where 0 indicates invalid and 1 indicates valid; ln represents the language information of the corner point; The output of the model is a set of corner point sequences, each of which represents a closed polygon of a room; Once the model predicts these ordered sequences of corner points, we simply connect all the corner points marked as valid to obtain a polygonal representation of each room.

10. A method for automatically generating floor plan based on artificial intelligence according to claim 1, characterized in that: The Dijkstra algorithm steps include: Step 1: Initialization: Set the distance from the source point to itself to 0, and the distance to all other corner points to infinity; create an unvisited corner point set and add all corner points to it; Step 2: Select corner points: Select a corner point closest to the source point from the unvisited corner point set and add it to the visited set. Step 3: Update distance: Update the distance of all adjacent corner points of the current corner point. If the distance from the current corner point to the adjacent corner point is shorter than the known shortest distance, update the shortest distance of the adjacent corner point and calculate the direction of the corner point to make the direction of the corner point close to the overall direction trend. Step 4: Repeat steps 2 and 3 until all corner points are visited or the target corner point is reached; Through the above steps and methods, the corner points are sorted, and the connection relationship between the corner points C and the corner points, that is, the edge E and the closed interval R, is also obtained.