Mobile LiDAR indoor room instance segmentation method based on scene structure elements
By combining deep learning with spatial grid information, we detect and complete the missing structural areas of indoor scene point clouds, eliminate non-structural points, and generate accurate room instance labels. This solves the problems of low segmentation accuracy and low efficiency in existing technologies and achieves high-quality indoor room segmentation.
Patent Information
- Application Number
- CN202510635140.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-30
AI Technical Summary
Existing room segmentation methods suffer from low segmentation accuracy and low automated segmentation efficiency in indoor scenes. Especially in the absence of prior knowledge, it is difficult to achieve high-quality room segmentation in complex environments.
A deep learning semantic segmentation method is used to segment the indoor scene point cloud. The spatial grid information is combined to detect and complete the structure missing areas. The weighted height difference information is used to eliminate non-structural points. The room instance labels are generated through image algebraic operations and label propagation strategies.
Accurate room segmentation is achieved in complex indoor scenes, improving segmentation accuracy and automation efficiency without relying on additional data input, and enhancing the integrity and closure of the scene structure.
Smart Images

Figure CN120726314A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a mobile LiDAR indoor room instance segmentation method based on scene structure elements. Background Art
[0002] In recent years, with the acceleration of urbanization and the increasing complexity of the built environment, more than 75% of the world's population now lives in cities and spends most of their time indoors. This makes the accurate acquisition and utilization of indoor spatial information a critical issue that needs to be addressed. Indoor scene reconstruction plays a vital role in meeting this demand. It is widely used in fields such as indoor navigation, emergency response, building maintenance, and indoor positioning, and significantly improves efficiency in design evaluation, cost control, and building management. At the same time, the continuous advancement of laser scanning technology provides an effective means for the measurement and perception of building information. High-density 3D point clouds collected by fixed terrestrial laser scanners (TLS) or indoor mobile laser scanners (IMLS) can provide detailed building structure information for indoor scenes.
[0003] Indoor space is typically defined as an enclosed space composed of rigid structural elements such as floors, ceilings, and walls. Many researchers have worked to reconstruct indoor 3D models by extracting these structural elements. However, due to the complexity of indoor scenes and occlusion by furniture, point cloud data often contains a lot of noise and missing areas, limiting accuracy and automation. Another group of researchers has begun to turn to indoor room segmentation, believing that compared to directly generating a complete 3D model, dividing the indoor scene into independent rooms can simplify the complexity of the problem while providing clearer and more structured information for the indoor model. Room segmentation not only simplifies complex indoor scenes, but also establishes spatial relationships between subspaces and generates topological structures.
[0004] Existing room segmentation methods, mostly focused on robotics and architecture, engineering, and construction (ACE), typically use 2D or 3D data acquired from various platforms as input. Some researchers believe that partitioning indoor scene maps into individual rooms or similar semantic units is central to many tasks in robotics. Room segmentation generates detailed and accurate topological maps, saving the computational effort required to obtain navigation trajectories, reducing redundant work, and avoiding interference between robots. In this field, most segmentation methods convert 3D data into 2D occupancy grid maps as input and then combine morphological methods with graph segmentation techniques. However, these methods are prone to over-segmentation in complex environments. In ACE, processing 3D point clouds is more common because they often contain more prior knowledge, such as the trajectory of mobile scanners and semantic information reflecting structural elements. In the absence of this information, satisfactory room segmentation is only possible under strict constraints, such as the lack of links between rooms at certain height levels or relying on a strong Manhattan assumption. Consequently, existing room segmentation methods suffer from low segmentation accuracy and inefficient automated segmentation.
[0005] Therefore, the existing technology needs to be improved. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that, in response to the defects of the existing technology, the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structure elements to solve the problems of low segmentation accuracy and low automated segmentation efficiency in the existing room segmentation methods.
[0007] The technical solutions adopted by the present invention to solve the technical problems are as follows: In a first aspect, the present invention provides a method for indoor room instance segmentation using a mobile LiDAR based on scene structure elements, comprising: The semantic segmentation method of deep learning is used to perform semantic segmentation on the indoor scene point cloud to obtain the main structural element points, secondary structural element points and non-structural element points; Based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, detect the structure missing areas caused by furniture occlusion or scanning angle limitation, and complete the structure of the structure missing areas; Calculating weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and removing non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information; The retained structural points and the original point cloud are converted into a two-dimensional occupancy grid map. The subtraction operation in image algebra is used to remove the connected areas between rooms. Based on eight-connected region analysis and inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
[0008] In one implementation, the semantic segmentation method of deep learning is used to perform semantic segmentation on the indoor scene point cloud to obtain primary structural element points, secondary structural element points, and non-structural element points, including: Based on the semantic alignment operation defined by the structural elements, the original categories of the training dataset are mapped to the semantic definitions that match the structural elements; According to the mapped semantic definition, the input indoor scene point cloud is semantically segmented by the deep learning semantic segmentation method, and the misclassified points are corrected based on the semantic optimization method based on geometric consistency to obtain the main structural element points, the secondary structural element points and the non-structural element points.
[0009] In one implementation, detecting structure-missing regions due to furniture occlusion or scanning view limitation based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, and completing the structure of the structure-missing regions, includes: Detecting the missing structure region based on information from the fused semantic labels, the two-dimensional occupancy grid, and the three-dimensional voxel grid; The structure missing area is structured and supplementary points are generated to fill the structure missing area, thereby restoring the structural closure and integrity of the point cloud scene.
[0010] In one implementation, the detecting the structure missing area based on the fusion of semantic labels, two-dimensional occupancy grid and three-dimensional voxel grid information includes: Dividing the point cloud with the semantic label into the two-dimensional occupancy grid and the three-dimensional voxel grid, and marking the cells where the main structural element points are located; Based on the satisfied conditions of the structural elements in the two-dimensional occupancy grid and the satisfied conditions of the points in the three-dimensional voxel grid, the structure missing area is identified; wherein the structure missing area is a voxel area where the corresponding position in the two-dimensional occupancy grid is a main structural element point and no point exists in the three-dimensional voxel grid.
[0011] In one implementation, calculating weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and removing non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information, includes: The point cloud after the completion structure is divided into corresponding two-dimensional occupancy grids, the maximum height difference and point density of each grid unit are calculated respectively, and the height difference is fused using the point density as a weighting factor to obtain the weighted average height difference of the entire scene; An adaptive height difference threshold is set according to the weighted average height difference, and points with both large height differences and high-density distribution are screened out as the main structural element points, while non-structural points and redundant points with similar geometric features but no structural attributes are eliminated.
[0012] In one implementation, converting the retained structural points and the original point cloud into a two-dimensional occupancy grid map and removing the connected areas between rooms using a subtraction operation in image algebra includes: Converting the retained structural points and the original point cloud into the two-dimensional occupancy grid map, and mapping the primary structural element points and the secondary structural element points into binary grid maps respectively through the two-dimensional grid; The subtraction operator in image algebra is introduced, and the connection areas between rooms are eliminated by subtracting the primary structure graph from the secondary structure graph.
[0013] In one implementation, generating room instance labels based on eight-connected region analysis and inverse distance weighted label propagation strategy and mapping them back to the scene point cloud includes: An eight-neighborhood connected region analysis is performed on the obtained two-dimensional occupancy grid map. By detecting the connection relationship between grid cells and adjacent cells in eight directions, multiple disconnected grid clusters are identified and a unique room label is assigned to each connected cluster. With the help of the coordinate mapping relationship between the point cloud and the grid, the points falling into the marked grid cells directly inherit the corresponding room labels. For points at the boundary or uncovered area, the labels are inferred and completed through the inverse distance weighted method to generate a coherent and complete room instance segmentation result.
[0014] In a second aspect, the present invention provides a mobile LiDAR indoor room instance segmentation system based on scene structure elements, comprising: The semantic segmentation module is used to perform semantic segmentation on the indoor scene point cloud using a deep learning semantic segmentation method to obtain the main structural element points, secondary structural element points, and non-structural element points; A missing area completion module is used to detect structural missing areas caused by furniture occlusion or scanning angle limitation based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, and to complete the structural missing areas; A point cloud removal module is used to calculate weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and to remove non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information; The instance segmentation module is used to convert the retained structural points and the original point cloud into a two-dimensional occupancy grid map. The subtraction operation in image algebra is used to remove the connected areas between rooms. Based on eight-connected region analysis and inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
[0015] In a third aspect, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores a mobile LiDAR indoor room instance segmentation program based on scene structure elements, and when the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by the processor, it is used to implement the operation of the mobile LiDAR indoor room instance segmentation method based on scene structure elements as described in the first aspect.
[0016] In a fourth aspect, the present invention further provides a medium, which is a computer-readable storage medium, storing a mobile LiDAR indoor room instance segmentation program based on scene structure elements. When the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by a processor, it is used to implement the operation of the mobile LiDAR indoor room instance segmentation method based on scene structure elements as described in the first aspect.
[0017] The present invention adopts the above technical solution to achieve the following effects: 1) This invention integrates the semantic and geometric features of point clouds to extract structural elements from indoor scene point cloud data. This allows the invention to be independent of additional data input, such as prior knowledge of scanning trajectories and room heights.
[0018] 2) The structure completion method adopted by this invention integrates semantic information and 2D / 3D grid division, which greatly reduces the impact of missing areas caused by cluttered indoor environments and furniture occlusion on the scene structure integrity and closure.
[0019] 3) The present invention adopts a density-aware height difference filtering algorithm to distinguish structural points from non-structural points with similar geometric features, making the structural points used for room instance segmentation more accurate and complete, and having better visual performance.
[0020] 4) The present invention can effectively extract structural elements in complex indoor scenes and generate accurate room segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0022] Figure 1 It is a flow chart of the mobile LiDAR indoor room instance segmentation method based on scene structure elements in the present invention.
[0023] Figure 2 It is a schematic diagram of the structure missing area caused by occlusion or scanning angle limitation in the indoor scene point cloud of the present invention.
[0024] Figure 3 Schematic diagram of the semantic label transfer process in the room instance segmentation process of the present invention.
[0025] Figure 4 Detailed information diagram of an indoor scene dataset used to evaluate the effect of the present invention (part 1) is shown in FIG.
[0026] Figure 5 This is a schematic diagram (part 2) of detailed information of an indoor scene dataset used to evaluate the effect of the present invention.
[0027] Figure 6 It is a schematic diagram of the under-segmented area and re-segmentation results in the C2 dataset of the present invention.
[0028] Figure 7 This is a detailed schematic diagram of the evaluation indicators of the present invention and the other four methods on 9 indoor scene datasets.
[0029] Figure 8 This is a comparison chart of the room instance segmentation results of the present invention and the segmentation results of the other four methods.
[0030] Figure 9 It is a functional principle diagram of a terminal in one implementation of the present invention.
[0031] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0033] Exemplary Methods Existing room segmentation methods, mostly focused on robotics and architecture, engineering, and construction (ACE), typically use 2D or 3D data acquired from various platforms as input. Some researchers believe that partitioning indoor scene maps into individual rooms or similar semantic units is central to many tasks in robotics. Room segmentation generates detailed and accurate topological maps, saving the computational effort required to obtain navigation trajectories, reducing redundant work, and avoiding interference between robots. In this field, most segmentation methods convert 3D data into 2D occupancy grid maps as input and then combine morphological methods with graph segmentation techniques. However, these methods are prone to over-segmentation in complex environments. In ACE, processing 3D point clouds is more common because they often contain more prior knowledge, such as the trajectory of mobile scanners and semantic information reflecting structural elements. In the absence of this information, satisfactory room segmentation is only possible under strict constraints, such as the lack of links between rooms at certain height levels or relying on a strong Manhattan assumption. Consequently, existing room segmentation methods suffer from low segmentation accuracy and inefficient automated segmentation.
[0034] In response to the above technical problems, an embodiment of the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structural elements, the method comprising: performing semantic segmentation on indoor scene point clouds through a deep learning semantic segmentation method; detecting structure-missing areas based on the semantic segmentation results and spatial grid information of the indoor scene point clouds, and structurally completing the structure-missing areas; calculating weighted height difference information according to the height difference and point density of the completed indoor scene point clouds, and removing non-structural points and redundant clutter in the scene point clouds based on the weighted height difference information; converting the retained structural points and the original point clouds into a two-dimensional occupancy grid map, using the subtraction operation in image algebra to remove the connected areas between rooms, and generating room instance labels based on eight-connected region analysis and inverse distance weighted label propagation strategy, and mapping them back to the scene point clouds; the embodiment of the present invention can effectively extract structural elements in complex indoor scenes and generate accurate room segmentation results.
[0035] like Figure 1 As shown, an embodiment of the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structure elements, comprising the following steps: Step S100: semantically segment the indoor scene point cloud using a deep learning semantic segmentation method to obtain primary structural element points, secondary structural element points, and non-structural element points.
[0036] In this embodiment, the mobile LiDAR indoor room instance segmentation method based on scene structural elements can effectively identify structural elements (e.g., doors, windows, walls, ceilings, and floors) in indoor scenes and utilize their spatial organization relationships to achieve accurate segmentation of complex multi-room environments.
[0037] In this embodiment, in the process of implementing the method, a pre-trained deep learning model is first used to perform semantic segmentation on the scene point cloud, and the point cloud is divided into three categories: primary structural elements, secondary structural elements, and non-structural elements. Among them, the primary structural elements include doors, windows and walls, and the secondary structural elements are ceilings and floors.
[0038] Specifically, in one implementation of this embodiment, step S100 includes the following steps: Step S101, based on the semantic alignment operation of the structural element definition, the original categories of the training data set are mapped to the semantic definitions matching the structural elements; Step S102, according to the mapped semantic definition, the input indoor scene point cloud is semantically segmented by the deep learning semantic segmentation method, and the misclassified points are corrected based on the semantic optimization method of geometric consistency to obtain the main structural element points, the secondary structural element points and the non-structural element points.
[0039] In this embodiment, semantic segmentation is first performed on the original point cloud to give the scene point cloud semantic information in addition to spatial coordinate information. The semantic information of the point cloud provides a key basis for understanding permanent structural elements such as walls, doors and windows in indoor scenes. However, in actual scenes, dense occlusion or local interference can easily lead to misclassification or partial loss of structural elements; traditional correction methods that rely only on geometric similarity (for example, normal vector consistency) are difficult to distinguish between objects with similar geometric shapes but different semantics (for example, flat walls and regular furniture surfaces). Moreover, the structural elements of indoor scenes (for example, walls, doors and windows) and non-structural elements (for example, furniture) have essential functional differences. Therefore, before using a pre-trained semantic segmentation model (for example, PointNet++ model, RandLA-Net model, etc.) to perform semantic segmentation on the indoor scene point cloud, a semantic alignment operation based on the definition of the structural element is designed in this embodiment to map the original categories of the training data set to semantic definitions that strictly match the structural elements, specifically including: i) major structural elements (including walls, doors, windows and columns); ii) secondary structural elements (including ceilings and floors) and iii) non-structural elements (including other categories in the dataset).
[0040] The above process redefines the label semantics so that the AI model trained based on this dataset can focus on the recognition of scene structure.
[0041] Subsequently, in order to address the possible misclassification problem of the deep learning model in occluded or sparse areas after semantic alignment, this embodiment further integrates the semantic features and geometric features of the scene point cloud, and designs a semantic optimization method based on geometric consistency to correct misclassification by analyzing geometric features such as the normal vector of the point.
[0042] Specifically, this embodiment uses normal vector similarity as a reference to identify points with similar geometric features to the main structural elements and correct misclassified points. The following formula 1 represents this calculation process: (1); Where, and Represent the main structural element points and non-structural element points The normal vector of Is the angle between the two normal vectors in radians. In this embodiment, the angle between them is calculated by the vector cross product and compared with the threshold. If it is less than the threshold, it is considered a non-structural element point. These are points that are incorrectly assigned semantic information and identified as main structural element points.
[0043] In this embodiment, the above-mentioned deep learning semantic segmentation method and semantic optimization method can ensure the accuracy of the main structural element points, so that these structures can be further used to accurately divide the room instances of indoor scenes.
[0044] like Figure 1 As shown, an embodiment of the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structure elements, comprising the following steps: Step S200 , based on the semantic segmentation result and spatial grid information of the indoor scene point cloud, detect the structure missing areas caused by furniture occlusion or scanning angle limitation, and complete the structure of the structure missing areas.
[0045] In this embodiment, after semantic segmentation of the indoor scene point cloud, the semantic information, geometric features and spatial grid are combined to identify and complete the structural missing areas caused by furniture occlusion or limited scanning angle, thereby enhancing the integrity and closure of the scene structure.
[0046] Specifically, in one implementation of this embodiment, step S200 includes the following steps: Step S201 : detecting the structure missing area based on the information of the fused semantic label, the two-dimensional occupancy grid and the three-dimensional voxel grid.
[0047] In one implementation of this embodiment, step S201 includes the following steps: Step S201a, dividing the point cloud with the semantic label into the two-dimensional occupancy grid and the three-dimensional voxel grid, and marking the cells where the main structural element points are located; Step S201b: Identify the structure missing region based on the conditions satisfied by the structural elements in the two-dimensional occupancy grid and the conditions satisfied by the points in the three-dimensional voxel grid; wherein the structure missing region is a voxel region where the corresponding position in the two-dimensional occupancy grid is a main structural element point and no point exists in the three-dimensional voxel grid.
[0048] Step S202 : completing the structure of the missing region, generating supplementary points to fill the missing region, and restoring the structural closure and integrity of the point cloud scene.
[0049] In this embodiment, after ensuring the accuracy of the main structural element points, the goal of this embodiment is to use these main structures to accurately divide the room instances of the indoor scene. One of the important influencing factors is the integrity of the structural elements. Due to the occlusion of indoor elements (for example, non-structural objects such as furniture) and the limitations of the scanning angle, structural points (for example, walls, doors and windows) often appear missing in the point cloud data. Figure 2 As shown, Figure 2 is the structure missing area caused by occlusion or scanning angle limitation, where Figure 2 (a) is a real room scene; Figure 2 (b) is a point cloud room scene. In the real scene, the occluded primary and secondary structures still exist, while in the point cloud room scene, the structures of these occluded parts are missing, such as Figure 2 These missing structures are shown by the dashed line in (b). Although these missing structures lack clear physical sampling points in the point cloud, they often correspond to existing structures or components in the real scene. Existing techniques often employ height slicing of wall point clouds or extracting ceiling and floor boundaries to reduce the interference caused by missing structures. However, these methods rely on prior conditions such as a preset scene height and complete ceiling and floor point cloud data, limiting their applicability in complex or data-incomplete environments.
[0050] To overcome these shortcomings, this embodiment provides a method for repairing missing structure in point clouds based on the synergy of semantic information and grid division. This method combines information from a two-dimensional occupancy grid and a three-dimensional voxel grid to model the spatial distribution of structural elements, thereby identifying potential areas of missing structure in point cloud data. Furthermore, the system generates completion points in these identified missing areas to restore the spatial continuity of structural elements in the scene, thereby restoring structural integrity.
[0051] Specifically, the method for repairing missing point cloud structures in this embodiment includes the following implementation steps: First, the semantically labeled point cloud data are mapped to two-dimensional occupancy grids. and 3D voxel grids In the process of mapping, the grid where the main structural elements in the point cloud are located is specifically marked; then, based on the combined information of the two-dimensional grid and the three-dimensional voxel grid, potential structural missing areas are identified. The specific identification method is: through the structural marking information in the two-dimensional grid Information about whether a point exists in a 3D voxel , the judgment rule for defining the structure missing area is shown in formula (2): (2); in, Indicates the type of structure-occupied voxels. A value of 2 indicates a structure point at the voxel location, a value of 1 indicates a region with missing structure, and a value of 0 indicates a region with no structural elements. Through this rule, this embodiment effectively identifies the locations of regions with missing structure. Within these regions, completion points are generated based on the geometric features and semantic information of adjacent structure points, thereby enhancing the spatial closure of the point cloud and restoring structural integrity.
[0052] The above-mentioned point cloud structure missing repair method based on the synergy of semantic information and grid division in this embodiment does not need to rely on prior conditions such as scene height or complete ceiling and floor point clouds, thereby improving adaptability and versatility to complex indoor layouts.
[0053] like Figure 1 As shown, an embodiment of the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structure elements, comprising the following steps: Step S300 , calculating weighted height difference information according to the height difference and point density of the completed indoor scene point cloud, and removing non-structural points and redundant clutter points in the scene point cloud based on the weighted height difference information.
[0054] In this embodiment, after completing the structure of the structure-missing area, non-structural points and redundant points in the scene point cloud are eliminated by weighted height difference information; wherein, the weighted height difference information is calculated by weighted fusion of the local height difference of the point cloud and the point density.
[0055] Specifically, in one implementation of this embodiment, step S300 includes the following steps: Step S301: Divide the point cloud after completing the structure into corresponding two-dimensional occupancy grids, calculate the maximum height difference and point density of each grid unit respectively, and fuse the height differences using the point density as a weighting factor to obtain the weighted average height difference of the entire scene; In step S302, an adaptive height difference threshold is set according to the weighted average height difference, points with large height difference and high density distribution are selected as the main structural element points, and non-structural points and redundant points with similar geometric features but no structural attributes are eliminated.
[0056] In this embodiment, although a relatively complete scene structure can be obtained through semantic-mesh collaborative restoration, when non-structural elements (for example, office partitions, cabinets) are similar to structural elements in local geometric features, the semantic classification and restoration process may still introduce misjudgment points, affecting the accuracy of room instance segmentation.
[0057] To solve this problem, this embodiment proposes an adaptive density-aware height difference filtering algorithm. Its core assumption is that the real structural element area should have both large height differences and high point density distribution, such as Figure 8 Specifically, this embodiment divides the scene point cloud into 2D grids, and calculates the height difference and point density of each grid unit based on the following formula, using the point density as the weight of the height difference to estimate the height distribution characteristics of the structural element: (3); (4); (5); (6); in, Represents a grid cell height difference; and Represents the maximum and minimum height of the unit respectively; Display unit The point density in ; represents the sum of weighted height differences; and Represents the grid in Axis and the number of units on the axis; Indicates the cumulative density value of the scene point cloud; represents the weighted mean height difference; To filter the height difference threshold of the main structural element points, is the empirical setting coefficient.
[0058] Through this step, the algorithm can make densely populated areas have a greater impact on the overall height characteristics, thereby enhancing the stability and accuracy of structure recognition. It can then accurately identify and extract key structural elements, exclude non-structural elements such as office partitions and cabinets, and minimize the number of misjudgments.
[0059] like Figure 1 As shown, an embodiment of the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structure elements, comprising the following steps: In step S400, the retained structural points and the original point cloud are converted into a two-dimensional occupancy grid map. The connected areas between rooms are removed using the subtraction operation in image algebra. Based on the eight-connected region analysis and the inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
[0060] In this example, after removing non-structural and redundant points from the scene point cloud, image algebraic subtraction is used to eliminate connected regions between rooms. Finally, eight-connected region analysis and an inverse distance weighted label propagation strategy are used to generate room instance labels. These labels are then mapped back to the scene point cloud, and the segmentation results are output as clearly semantically annotated room instances.
[0061] Specifically, in one implementation of this embodiment, step S400 includes the following steps: Step S401, converting the retained structural points and the original point cloud into the two-dimensional occupancy grid map, and mapping the main structural element points and the secondary structural element points into binary grid maps respectively through the two-dimensional grid; In step S402, the subtraction operator in image algebra is introduced to remove the connection areas between rooms by subtracting the primary structure graph from the secondary structure graph.
[0062] In this embodiment, the main structural elements obtained by the methods corresponding to the aforementioned steps S100 to S300 form a closed boundary on the two-dimensional plane, preliminarily defining the basic form of the room.
[0063] Furthermore, this embodiment models the room segmentation problem as a label assignment problem. First, the primary and secondary structural elements are mapped into binary raster images using a two-dimensional grid. To ensure spatial independence between rooms, the subtraction operator from image algebra is introduced. Connected regions between rooms are effectively removed by subtracting the primary structural image from the secondary structural image. Specifically, a grid cell retains its structural attributes only if it is marked as a structural cell in the secondary structural image and unoccupied in the primary structural image; otherwise, it is removed. This operation effectively disconnects connected regions between rooms, enhancing the independence of each room.
[0064] Specifically, in one implementation of this embodiment, step S400 further includes the following steps: Step S403: Perform eight-neighborhood connected region analysis on the obtained two-dimensional occupancy grid map. By detecting the connection relationship between the grid unit and the adjacent units in eight directions, multiple disconnected grid clusters are identified, and a unique room label is assigned to each connected cluster. In step S404, with the help of the coordinate mapping relationship between the point cloud and the grid, the points falling into the marked grid cells directly inherit the corresponding room labels. For points at the boundary or uncovered area, the labels are inferred and completed using the inverse distance weighted method to generate a coherent and complete room instance segmentation result.
[0065] In this embodiment, based on the subtraction of the primary structure graph from the secondary structure graph, a label propagation strategy combining eight-neighborhood connectivity analysis and inverse distance weighted interpolation is proposed to achieve complete label mapping from the two-dimensional structure grid to the three-dimensional point cloud layer. First, the eight-neighborhood connected area analysis is performed on the two-dimensional structure grid graph obtained by the previous operation. By detecting the connection relationship between the grid unit and its eight adjacent units, multiple disconnected grid clusters are identified, and a unique room label is assigned to each connected cluster. Subsequently, with the help of the coordinate mapping relationship between the point cloud and the grid, the label is transferred from the grid level to the point cloud level: points falling within the labeled grid unit directly inherit its room label; while points at the boundary or uncovered area have their labels inferred and completed using the inverse distance weighted method to ensure the generation of a coherent and complete room instance segmentation result.
[0066] The specific process of the inverse distance weighted method is as follows: Figure 3 As shown, Figure 3 The calculation method of label probability distribution of unlabeled points is shown in Figure 2. The label points in the dotted box are the neighboring points of the unlabeled points. This method provides a label probability distribution based on spatial proximity for unlabeled points through inverse distance weighting, thereby achieving room instance segmentation. Figure 3 As shown in Figure 2, through this process, it can be calculated that the room instance labels corresponding to the unlabeled point 1 and the unlabeled point 2 are both room 1.
[0067] (7); (8); in, } represents the set of unlabeled area points, For any point in the set, formula (7) calculates the unlabeled point Its neighboring points with labels The Euclidean distance between them, and the reciprocal of the distance is used as the weight , to quantify the influence of neighboring points in label judgment.
[0068] Since neighboring points may have different room instance labels, formula (8) calculates Weighted probability of belonging to a certain label , that is, point The room instance label is The numerator is the probability of all neighbor points with labels The sum of the weights corresponding to the points, and the denominator is the sum of the weights of all neighboring points, so that the sum of the probabilities is 1. The labels of the neighboring points are given by Indicates that it represents the neighbor point Finally, in this embodiment, The label with the largest value Assigned to each unlabeled point , as the room instance label corresponding to the point.
[0069] As an example, in the actual application scenario of this embodiment, qualitative and quantitative tests were conducted on 9 different indoor scene datasets. The detailed information of the indoor scene datasets used to evaluate the effect of this embodiment is as follows: Figure 4~Figure 5 shown.
[0070] The A1 dataset is derived from a portion of the rooms and corridors in area 4 of the S3DIS dataset. Data from B1 to B5 are from the UZHRooms detection dataset. B1, B2, and B3 are virtual datasets obtained by placing a virtual scanner within a synthetic indoor model, while B4 and B5 correspond to the "Cabin" and "Office 3" data scanned by a Faro Focus 3D laser scanner. Data from C1 and C2 are derived from the Indoor Modeling Benchmark datasets provided by the International Society for Photogrammetry and Remote Sensing (ISPRS), specifically the TUB2 Second Floor and Grainger Museum data. Data from D1 is derived from the dataset provided by the CVPRBIM challenge.
[0071] also, Figure 4~Figure 5 The segmentation results of this embodiment on 9 data sets are also shown. Different room instances are marked with different colors for easy intuitive comparison. From these results, it can be seen that the method proposed in this embodiment shows good robustness and effectiveness when dealing with indoor scenes of various types and sizes. In the C2 data set, since the room height in a certain area is significantly lower than that in other areas, this area has a certain degree of under-segmentation in the first segmentation result. In order to solve this problem, this embodiment marks the area separately as C2-Part and segments it again, thereby generating a complete C2 segmentation result. Specifically, Figure 6 As shown, Figure 6The under-segmented regions and re-segmentation results in the C2 dataset are shown in Figure 2.
[0072] Table 1 Parameter description of this embodiment and parameter settings in different data sets
[0073] Table 1 shows the parameter descriptions and settings in the key steps of this example. The method runs automatically and does not require any user intervention except for parameter selection. In the nine scenes, most of the input parameters remain consistent and enable the method to achieve good segmentation results. In the room instance segmentation step, the key parameter is the number of points to fill the missing structure area. and height difference filter coefficient In this embodiment, Set it to 1 to achieve a better balance between scene integrity improvement and algorithm processing efficiency. In this embodiment, the adjustment is made according to the characteristics of the scene. Specifically, for datasets with rooms of different heights such as B4, C1, and C2, the , in order to retain the main structural points of low height as much as possible while filtering out non-main structural points.
[0074] To verify the accuracy of the room instance generation by the segmentation algorithm proposed in this example, a method with four evaluation metrics was designed to comprehensively evaluate the effectiveness of the algorithm from multiple perspectives. These metrics include intersection over union (IoU), precision, recall, and F1 score. The specific formulas are defined as follows: (9); (10); (11); (12); in, Indicates the number of points in the segmented room instance that belong to the corresponding real room. It indicates the number of error points contained in the segmented room instance. Represents the number of points that should be included in the segmented room instance but are actually omitted.
[0075] Formula (9) describes the calculation process of IoU, which measures the degree of spatial overlap between the predicted room instance and the real room instance. The accuracy of the segmentation result is evaluated by the ratio of the correctly predicted point to the union of all predicted points and the real label points.
[0076] Formula (10) and Formula (11) define precision and recall respectively. Precision reflects the proportion of points predicted to belong to a room that actually belong to the room, while recall measures the proportion of points that actually belong to the room that are successfully predicted.
[0077] The score is defined as described in formula (12), which balances the precision and recall, providing a balanced evaluation index for verifying the overall performance of the segmentation algorithm. Five methods, namely, Cloth Simulation Filter and Regular Grid Analysis-based segmentation (CSF-RGA), Mor segmentation (morphological segmentation, Mor), distance segmentation (distance transformation-based segmentation, Dist), and Voronoi segmentation graph-based segmentation, Vor), were used to test the nine indoor scenes mentioned above for comparison. The precision, recall rate, F1 score, and intersection-over-union ratio of different algorithms are shown in Table 2. The details of the evaluation indicators of the method of this embodiment and the other four methods are as follows: Figure 7 shown.
[0078] Table 2 Comparison results of average quantitative indicators on 9 indoor scenes
[0079] As shown in Table 2, the experimental results demonstrate that this embodiment demonstrates excellent performance across all four metrics, achieving the best results in precision, F1 score, and average intersection over union (IoU). The CSF-RGA method achieves a good balance between precision and boundary delineation in most scenarios. In particular, in clearly structured rooms such as A1, B1, and B2, its precision approaches that of the algorithm proposed in this embodiment. However, it still exhibits certain limitations in balancing precision and boundary delineation. The Mor and Vor methods exhibit relatively unstable performance. While the Mor method achieves recall comparable to that of the present embodiment and the CSF-RGA method across all nine scenarios, its precision and IoU fluctuate significantly. This suggests that the Mor method prioritizes room coverage in the room segmentation task, resulting in a high number of under-segmented regions. The Vor method exhibits fluctuations in all metrics across all nine scenarios. While its recall fluctuates less than that of precision, F1 score, and IoU, it still fails to achieve the recall stability of the Mor method. This may be due to the Vor segmentation strategy's reliance on heuristic merging, which results in a high number of small regions in the segmentation results. The Dist method showed relatively balanced overall performance across all scenarios. Although all metrics showed some fluctuation across the nine scenarios, the magnitude of the fluctuation was not as large as that of the Mor and Vor methods, demonstrating its versatility. However, compared to this example and the CSF-RGA method, its segmentation metrics were somewhat limited in some scenarios.
[0080] like Figure 8 As shown, Figure 8 The comparison chart of the room instance segmentation results provided by this embodiment and the results of four methods including CSF-RGA (segmentation based on cloth simulation and regular grid analysis), Mor (morphological segmentation), Dist (distance segmentation) and Vor (Voronoi graph segmentation) is shown in FIG. Figure 8 It can be seen that CSF-RGA has under-segmentation in B5 and C1. This is because CSF-RGA relies on the ceiling point cloud in the scene as the basis for segmentation. Due to the close connection of the ceiling point clouds in scenes B5 and C1, CSF-RGA achieves poor segmentation results in these two scenes. Mor's under-segmentation is mainly reflected in scenes with long corridors such as A1, B1, B3, B5 and C1. In B4, the connection area between different rooms occupies a relatively large area and is therefore mistakenly identified as the same room. Dist shows a certain degree of over-segmentation in some long corridor scenes. In addition, it also shows a certain degree of over-segmentation in irregular scenes such as B2 and C2. Vor's over-segmentation is more obvious in scenes with long corridors. The above results show that compared with other methods, this embodiment can still maintain superior segmentation performance when processing irregular and complex scenes, and exhibits better generalization ability.
[0081] This embodiment achieves the following technical effects through the above technical solution: 1) This embodiment integrates the semantic and geometric features of point clouds to extract structural elements from indoor scene point cloud data. This eliminates the need for additional data input, such as prior knowledge of scanning trajectories and room heights.
[0082] 2) The structure completion method adopted in this embodiment integrates semantic information and 2D / 3D grid division, which greatly reduces the impact of missing areas caused by cluttered indoor environments and furniture occlusion on the scene structure integrity and closure.
[0083] 3) This embodiment uses a density-aware height difference filtering algorithm to distinguish structural points from non-structural points with similar geometric features, making the structural points used for room instance segmentation more accurate and complete, and providing better visual performance.
[0084] 4) This embodiment can effectively extract structural elements in complex indoor scenes and generate accurate room segmentation results.
[0085] Exemplary devices Based on the above embodiments, the present invention further provides a mobile LiDAR indoor room instance segmentation system based on scene structure elements, comprising: The semantic segmentation module is used to perform semantic segmentation on the indoor scene point cloud using a deep learning semantic segmentation method to obtain the main structural element points, secondary structural element points, and non-structural element points; A missing area completion module is used to detect structural missing areas caused by furniture occlusion or scanning angle limitation based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, and to complete the structural missing areas; A point cloud removal module is used to calculate weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and to remove non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information; The instance segmentation module is used to convert the retained structural points and the original point cloud into a two-dimensional occupancy grid map. The subtraction operation in image algebra is used to remove the connected areas between rooms. Based on eight-connected region analysis and inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
[0086] This embodiment achieves the following technical effects through the above technical solution: 1) This embodiment integrates the semantic and geometric features of point clouds to extract structural elements from indoor scene point cloud data. This eliminates the need for additional data input, such as prior knowledge of scanning trajectories and room heights.
[0087] 2) The structure completion method adopted in this embodiment integrates semantic information and 2D / 3D grid division, which greatly reduces the impact of missing areas caused by cluttered indoor environments and furniture occlusion on the scene structure integrity and closure.
[0088] 3) This embodiment uses a density-aware height difference filtering algorithm to distinguish structural points from non-structural points with similar geometric features, making the structural points used for room instance segmentation more accurate and complete, and providing better visual performance.
[0089] 4) This embodiment can effectively extract structural elements in complex indoor scenes and generate accurate room segmentation results.
[0090] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be shown as follows: Figure 9 shown.
[0091] The terminal includes: a processor, memory, interface, display screen and communication module connected via a system bus; wherein the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the storage medium; the interface is used to connect to external devices; the display screen is used to display corresponding information; and the communication module is used to communicate with a cloud server or other devices.
[0092] When the computer program is executed by a processor, it is used to implement the operation of a mobile LiDAR indoor room instance segmentation method based on scene structure elements.
[0093] It will be understood by those skilled in the art that Figure 9 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0094] In one embodiment, a terminal is provided, comprising: a processor and a memory, wherein the memory stores a mobile LiDAR indoor room instance segmentation program based on scene structure elements, and when the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by the processor, it is used to implement the operations of the mobile LiDAR indoor room instance segmentation method based on scene structure elements as described above.
[0095] In one embodiment, a storage medium is provided, wherein the storage medium stores a mobile LiDAR indoor room instance segmentation program based on scene structure elements. When the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by a processor, it is used to implement the operations of the mobile LiDAR indoor room instance segmentation method based on scene structure elements as described above.
[0096] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include both non-volatile and volatile memory.
[0097] In summary, the present invention provides a mobile LiDAR indoor room instance segmentation method based on scene structural elements, including: semantic segmentation of indoor scene point clouds through a deep learning semantic segmentation method; detection of structure-missing areas based on the semantic segmentation results and spatial grid information of the indoor scene point clouds, and structural completion of the structure-missing areas; calculation of weighted height difference information based on the height difference and point density of the completed indoor scene point clouds, and elimination of non-structural points and redundant clutter in the scene point clouds based on the weighted height difference information; conversion of the retained structural points and the original point clouds into a two-dimensional occupancy grid map, use of the subtraction operation in image algebra to remove the connected areas between rooms, and generation of room instance labels based on eight-connected region analysis and inverse distance weighted label propagation strategy, and mapping back to the scene point clouds; the present invention can effectively extract structural elements in complex indoor scenes and generate accurate room segmentation results.
[0098] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A mobile LiDAR indoor room instance segmentation method based on scene structure elements, characterized by: include: The semantic segmentation method of deep learning is used to perform semantic segmentation on the indoor scene point cloud to obtain the main structural element points, secondary structural element points and non-structural element points; Based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, detect the structure missing areas caused by furniture occlusion or scanning angle limitation, and complete the structure of the structure missing areas; Calculating weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and removing non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information; The retained structural points and the original point cloud are converted into a two-dimensional occupancy grid map. The subtraction operation in image algebra is used to remove the connected areas between rooms. Based on eight-connected region analysis and inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
2. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 1, characterized in that: The semantic segmentation method of deep learning is used to perform semantic segmentation on the indoor scene point cloud to obtain the main structural element points, secondary structural element points and non-structural element points, including: Based on the semantic alignment operation defined by the structural elements, the original categories of the training dataset are mapped to the semantic definitions that match the structural elements; According to the mapped semantic definition, the input indoor scene point cloud is semantically segmented by the deep learning semantic segmentation method, and the misclassified points are corrected based on the semantic optimization method based on geometric consistency to obtain the main structural element points, the secondary structural element points and the non-structural element points.
3. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 1, wherein: The method of detecting structure-missing areas caused by furniture occlusion or scanning view limitation based on the semantic segmentation results and spatial grid information of the indoor scene point cloud and completing the structure of the structure-missing areas includes: Detecting the missing structure region based on information from the fused semantic labels, the two-dimensional occupancy grid, and the three-dimensional voxel grid; The structure missing area is structured and supplementary points are generated to fill the structure missing area, thereby restoring the structural closure and integrity of the point cloud scene.
4. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 3, characterized in that: The detecting of the structure missing area based on the fusion of semantic labels, two-dimensional occupancy grid and three-dimensional voxel grid information includes: Dividing the point cloud with the semantic label into the two-dimensional occupancy grid and the three-dimensional voxel grid, and marking the cells where the main structural element points are located; Based on the satisfied conditions of the structural elements in the two-dimensional occupancy grid and the satisfied conditions of the points in the three-dimensional voxel grid, the structure missing area is identified; wherein the structure missing area is a voxel area where the corresponding position in the two-dimensional occupancy grid is a main structural element point and no point exists in the three-dimensional voxel grid.
5. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 1, wherein: The step of calculating weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and removing non-structural points and redundant clutter points in the scene point cloud based on the weighted height difference information, includes: The point cloud after the completion structure is divided into corresponding two-dimensional occupancy grids, the maximum height difference and point density of each grid unit are calculated respectively, and the height difference is fused using the point density as a weighting factor to obtain the weighted average height difference of the entire scene; An adaptive height difference threshold is set according to the weighted average height difference, and points with both large height differences and high-density distribution are screened out as the main structural element points, while non-structural points and redundant points with similar geometric features but no structural attributes are eliminated.
6. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 1, wherein: The method of converting the retained structural points and the original point cloud into a two-dimensional occupancy grid map and removing the connected areas between rooms using a subtraction operation in image algebra includes: Converting the retained structural points and the original point cloud into the two-dimensional occupancy grid map, and mapping the primary structural element points and the secondary structural element points into binary grid maps respectively through the two-dimensional grid; The subtraction operator in image algebra is introduced, and the connection areas between rooms are eliminated by subtracting the primary structure graph from the secondary structure graph.
7. The method for indoor room instance segmentation based on scene structure elements using mobile LiDAR according to claim 1, wherein: The eight-connected region analysis and inverse distance weighted label propagation strategy are used to generate room instance labels and map them back to the scene point cloud, including: An eight-neighborhood connected region analysis is performed on the obtained two-dimensional occupancy grid map. By detecting the connection relationship between grid cells and adjacent cells in eight directions, multiple disconnected grid clusters are identified and a unique room label is assigned to each connected cluster. With the help of the coordinate mapping relationship between the point cloud and the grid, the points falling into the marked grid cells directly inherit the corresponding room labels. For points at the boundary or uncovered area, the labels are inferred and completed through the inverse distance weighted method to generate a coherent and complete room instance segmentation result.
8. A mobile LiDAR indoor room instance segmentation system based on scene structure elements, characterized by: include: The semantic segmentation module is used to perform semantic segmentation on the indoor scene point cloud using a deep learning semantic segmentation method to obtain the main structural element points, secondary structural element points, and non-structural element points; A missing area completion module is used to detect structural missing areas caused by furniture occlusion or scanning angle limitation based on the semantic segmentation results and spatial grid information of the indoor scene point cloud, and to complete the structural missing areas; A point cloud removal module is used to calculate weighted height difference information based on the height difference and point density of the completed indoor scene point cloud, and to remove non-structural points and redundant clutter in the scene point cloud based on the weighted height difference information; The instance segmentation module is used to convert the retained structural points and the original point cloud into a two-dimensional occupancy grid map. The subtraction operation in image algebra is used to remove the connected areas between rooms. Based on eight-connected region analysis and inverse distance weighted label propagation strategy, room instance labels are generated and mapped back to the scene point cloud.
9. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores a mobile LiDAR indoor room instance segmentation program based on scene structure elements, and when the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by the processor, it is used to implement the operation of the mobile LiDAR indoor room instance segmentation method based on scene structure elements according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a mobile LiDAR indoor room instance segmentation program based on scene structure elements. When the mobile LiDAR indoor room instance segmentation program based on scene structure elements is executed by a processor, it is used to implement the operation of the mobile LiDAR indoor room instance segmentation method based on scene structure elements according to any one of claims 1 to 7.