Method for positioning a mobile construction robot at a construction site using semantic segmentation, construction robot system and computer program product
Patent Information
- Application Number
- CN202280019684.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-30
- Filing Date
- 2022-04-13
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2042-04-13
AI Technical Summary
然而,这种解决方案缺乏灵活性和可用性,并且经常需要手动干预
[0005] Therefore, the object of the present invention is to provide a robust and accurate method for positioning mobile construction robots at construction sites.
Smart Images

Figure CN116982081B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for positioning a mobile construction robot at a construction site. Background Technology
[0002] The usefulness of mobile construction robots on construction sites increases with their autonomy. Therefore, it would be of particular interest if mobile construction robots could autonomously, accurately, and robustly infer their location on the construction site.
[0003] Much effort has been put into this to date. One approach involves using map data from the construction site, such as Building Information Modeling (BIM) data combined with tools like total stations. However, this solution lacks flexibility and usability and often requires manual intervention.
[0004] Even if BIM data for a specific construction site is available, positioning based solely on this BIM data and, for example, distance measurements, often fails due to distracting objects on the construction site, such as people, power tools, and material stacks. Summary of the Invention
[0005] Therefore, the object of the present invention is to provide a robust and accurate method for positioning mobile construction robots at construction sites.
[0006] This objective is achieved in several aspects of the invention, wherein a first aspect is a method for locating a mobile construction robot at a construction site, wherein environmental data of the construction site is captured, wherein the environmental data is used to infer the pose of the mobile robot, wherein the method includes a semantic segmentation step, wherein a semantic classifier having an updatable semantic model classifies the environmental data into at least two semantic categories, wherein the semantic model is updated at least once.
[0007] Attitude can include the position and / or orientation of the mobile construction robot, particularly its position and / or orientation relative to the construction site.
[0008] Particularly preferably, the environmental data includes 3D environmental data. The environmental data can be obtained by a distance measurement sensor, such as a LiDAR scanner or a time-of-flight camera.
[0009] Therefore, the idea of this invention is to semantically separate the different parts of the captured environmental data.
[0010] Specifically, environmental data can be categorized based on its usefulness for locating mobile construction site robots. Therefore, noisy or even interfering data can be detected and discarded. This results in robust and accurate positioning based on the purposeful selection of environmental data, particularly in chaotic environments such as construction sites.
[0011] Environmental data can be captured by the mobile construction robot itself. Therefore, all sensor devices can be mounted on the mobile construction robot. As a result, the autonomy of the mobile construction robot can be improved.
[0012] In a preferred variation of this method, environmental data can be used to update the semantic model. Therefore, environmental data can be used to locate the mobile construction robot and to improve the semantic classifier. This, in turn, can further enhance the robustness and / or accuracy of the localization.
[0013] A semantic classifier categorizes environmental data into a set of categories. This set of categories can contain a finite number of semantic categories.
[0014] Preferably, at least one of the semantic categories can correspond to the construction site background, and at least one of the semantic categories can correspond to the construction site foreground.
[0015] Categories (particularly background and foreground categories) may relate to the invariance of objects detected in environmental data (e.g., walls and tables or moving animals); and / or preferably, to the probability that the environmental data is mapped to a set of available map data (e.g., BIM data). "Background" can include environmental data such as walls, ceilings, floors, etc. In particular, it can include elements typically represented in map data or correspondingly in BIM data. "Foreground" can include other elements, particularly inferred objects such as people, furniture, equipment, etc.
[0016] Preferably, the pose can be inferred using a subset of the environmental data classified as construction site background. Therefore, data classified as construction site foreground can be extracted from the environmental data. Unclassified data can also be extracted. Subsequently, localization can be performed based on the remaining environmental data, particularly based solely on environmental data classified as background. Of course, if more refined classifications are available, additional semantic information can be considered. For example, if there are multiple background categories, such as separate categories for walls, ceilings, floors, etc., this additional semantic information can facilitate matching the environmental data with available BIM data.
[0017] The quality (especially classification accuracy) can be improved by refining the semantic segmentation of the environmental data at least once using at least one additional dataset (preferably an image dataset). For this purpose, superpixels can be generated, for example, by using the Simple Linear Iterative Clustering (SLIC) algorithm.
[0018] Preferably, site map data (e.g., BIM data) can be used to infer the orientation. In particular, environmental data, or a selection thereof, can be matched with map data to infer the position and / or orientation of the mobile construction robot.
[0019] The semantic classifier can be trainable. Therefore, to adapt the semantic classification to different domains, the semantic model can be updated multiple times, preferably periodically. Updates can use current or previous semantic segmentation data and / or current or previous contextual data, or a choice thereof.
[0020] To counteract catastrophic forgetting, the semantic model can be updated by mixing current and previously collected semantic segmentation data.
[0021] Traditional semantic segmentation methods allow the introduction of new categories in different semantic segmentation tasks, while this invention focuses on classifying environmental data, particularly based on its usefulness for localization. Therefore, preferably, the set of semantic categories is fixed for all semantic segments. In this sense, this invention addresses the problem of domain adaptation, where the goal is to adjust the semantic classifier to better generalize to new domains (targets) that may have a large semantic gap with the domain (source) on which the semantic classifier was trained.
[0022] Another aspect of the present invention relates to a construction robot system comprising: a mobile construction robot configured to work on a construction site, for example, for drilling, chiseling, grinding, plastering and / or painting; and a control unit configured to position the mobile construction robot using a method according to the present invention.
[0023] The mobile construction robot may include a mobile base for movement on at least one of the floor, ceiling, or wall at the construction site. The mobile base may be unidirectional, multidirectional, or omnidirectional. It may include wheels and / or tracked wheels. Alternatively, it may also be or include a legged mobile base. Additionally or alternatively, the mobile construction robot (particularly the mobile base) may be an aircraft, particularly an unmanned aerial vehicle (UAV), also known as a "construction drone."
[0024] It can be configured for at least one of gripping or moving an object, drilling, chiseling, grinding, plastering, or painting.
[0025] Mobile construction robots may include control units. Alternatively, the control unit may be, or at least partially, external to the mobile construction robot. The control unit may be, at least partially, part of a cloud-based computing system. The control unit may be configured to control multiple mobile construction robots. In particular, the same semantic classifier, or at least its semantic model, may be used for multiple mobile construction robots. Therefore, training the semantic classifier may benefit several of the multiple mobile construction robots.
[0026] In a preferred embodiment, the mobile construction robot includes a 3D sensor (e.g., a LiDAR scanner or stereo camera system) and at least one camera. Thus, the 3D sensor can provide 3D environmental data as environmental data, which is enriched by image data (e.g., RGB image data) provided by the at least one camera.
[0027] Another aspect of the invention relates to a computer program product comprising a storage device readable by a control unit of a construction robot system, the storage device carrying instructions which, when executed by the control unit, cause the control unit to perform the method according to the invention.
[0028] The invention will be further described by way of example with reference to the accompanying drawings, which illustrate preferred variations of the invention. It should be understood that the following description is illustrative and not intended to limit the scope of the invention. Features shown herein are not necessarily to be construed as to scale and are presented in a manner that makes the particular features of the invention clearly visible. In variations of the invention, various features may be implemented individually or in combination in any desired manner. Attached Figure Description
[0029] Figure 1 An overview of the method according to the present invention is shown;
[0030] Figures 2a to 2d This illustrates the different stages of semantic classification; and
[0031] Figure 3 A construction robot system is shown. Detailed Implementation
[0032] In the specification and drawings, the same reference numerals are used for functionally equivalent elements whenever possible.
[0033] Figure 1 An overview of method 10 according to the invention is shown schematically, which will be described in more detail in the following sections.
[0034] The method 10 according to the invention, and the construction robot system that applies method 10 thereto, combine the concepts of continuous learning and self-supervision.
[0035] Environmental data 12 is captured by a LiDAR scanner. An initial attitude estimate and an initial set of pseudo-labels 16 are generated during the positioning step by matching the environmental data 12 with a BIM model 14 representing, for example, a site plan.
[0036] A semantic classifier 18, which includes a continuously learning semantic model, is used as an input filter for localization, thereby semantically segmenting the scene 20 observed by a camera system 22, which includes one or more cameras (in this embodiment, a camera monitoring the left rear, a camera monitoring the center, and a camera monitoring the right rear), into segments classified as construction site foreground or construction site background, for example, superpixels. This fine segmentation is fed forward into localization.
[0037] This set of semantic categories is not an arbitrary category label, but rather a purposeful selection based on observable functional availability, where some parts of the current scene are mapped (construction site background) while others are not (construction site foreground). Therefore, pseudo-labels are created based on localization to train the semantic classifier 18, referencing the BIM model 14, while the segmentation of the construction site foreground and background provides information for localization.
[0038] This creates a feedback loop that improves two parts of the process, particularly the quality of semantic classification and the robustness and accuracy of localization (i.e., pose detection). On average, experiments show a 60% improvement in semantic segmentation and a 10% improvement in median localization error during deployment.
[0039] Therefore, lifelong self-supervised learning for semantic classification in scenario 20 was achieved.
[0040] This invention addresses several issues, among which the following will be discussed in more detail:
[0041] • Match environmental data with BIM model 14 for location purposes.
[0042] • Based on multi-mode calibration and available BIM model 14, pseudo-labels are generated for self-supervised purposes, and
[0043] Integration of continuous learning methods and domain adaptation in semantic segmentation
[0044] position
[0045] The mobile construction robot is positioned by aligning environmental data 12 from a LiDAR scan with a given floor plan in the form of BIM model data 14.
[0046] Given the building model mesh M representing BIM model 14, the point cloud P representing environmental data 12, and the initial alignment
[0047]
[0048] The subsequent posture of the mobile construction robot can be found as
[0049]
[0050] ICP is a point-to-surface ICP based on the work of Yang Chen and Gerard Medioni, published in 1992 in the journal *Image and Visual Computing*, Vol. 10, No. 3, pp. 145-155. ICP filters out points with large distances and other criteria.
[0051] Point-to-surface ICP is preferably run using three nearest neighbors and initialized based on the previously solved pose. Furthermore, multiple filters are applied to the input. These filters can also be applied after semantic filtering.
[0052] ■ A LiDAR scan requires a minimum of 500 points; therefore, scan results will be discarded if the segmentation classifies almost everything as a construction site foreground.
[0053] ■ The maximum density of secondary sampling for LiDAR scanning is 10,000 points / m². 3 .
[0054] ■ After nearest neighbor association, the 20% of points farthest from BIM model 14 were removed.
[0055] ■ In addition, if the estimated surface normal based on the 10 nearest neighbors has an angular deviation greater than 1.5 rad, the association is discarded.
[0056] Furthermore, for highly complex 3D environments or corresponding construction sites, additional filters can be implemented to enable localization without segmentation: for example, only four degrees of freedom can be inferred, specifically x, y, z, and yaw. Normal directions can be based on the 30 nearest neighbors. A point can only be associated with BIM model data 14 if the angle between normals is less than 0.8 rad.
[0057] In order to further divide the point cloud P into foreground points and background points of the construction site, image data from camera system 22 is used as additional information.
[0058] After semantic classification of the image data from camera system 22, static point cloud P is extracted from point cloud P. 静态 Those points p∈P that are reprojected pixels in their image frames and classified as the background of the construction site.
[0059] Therefore, localization (i.e., pose detection) is based on
[0060] Semantic classification
[0061] Figures 2a to 2d The different stages during pseudo-label generation and therefore during semantic classification are shown.
[0062] For each camera in the camera system 22, semantic classification (“pseudo-labels”) is generated by labeling the captured point cloud of the environment data 12 using the current pose estimation and the BIM model data 14 of the construction site.
[0063] As an example, Figure 2a An overlay image of a grid based on BIM model data 14 and a point cloud corresponding to environmental data 12 from a LiDAR scan is shown schematically.
[0064] In this variation of Method 10, the data is categorized into three distinct classes: construction site foreground, construction site background, and unknown.
[0065] The categories are created in two steps. First, for each point of the located environmental data 12, we calculate the distance to the nearest plane of the mesh based on the BIM model data 14 using fast intersection and distance computation according to “3D fast intersection and distance computation” published by Pierre Alliez, Stephane Tayeb, and Camille Wormser in the CGAL User and Reference Manual, version 5.2, published by the CGAL Editorial Board in 2020; https: / / doc.cgal.org / 5.2 / Manual / packages.html#PkgAABBTree.
[0066] If the distance exceeds a given threshold δ, the point is designated as the foreground category of the construction site; otherwise, it is designated as the background category. Based on experience, the distance threshold is set to δ = 0.1m. Figure 2b The result of this preliminary classification is schematically shown as an overlay of the image from camera system 22 and the projection of the classified point cloud. Black pixels along the LiDAR scan lines correspond to unknown categories, light gray pixels along the LiDAR scan lines correspond to the foreground of the construction site, and the remaining pixels along the LiDAR scan lines correspond to the background of the construction site. Therefore, according to Figure 2b The area, including the workers and some equipment, has been classified as Figure 2b The foreground on the right half.
[0067] In the second stage, superpixels created using the Simple Linear Iterative Clustering (SLIC) algorithm, which utilizes k-means clustering, are used to refine the projection. Figure 2c An example of an image segmented using a grid based on superpixels is illustrated schematically. In particular, the grid may correspond to the outline of the superpixels.
[0068] Specifically, the image is over-segmented into a set S of superpixels. Then, a class is assigned to superpixels s∈S based on a majority vote of the included projection classes. The segmentation is further improved by discarding superpixels whose depth variance exceeds a given threshold.
[0069] Superpixels with a depth variance exceeding 0.5m are discarded. The image is smoothed using a Gaussian kernel (σ = 0.2). With SLIC parameter tightness = 10, the image is over-segmented into approximately 400 superpixels.
[0070] Instead of the SLIC algorithm, other superpixel algorithms can be used, such as the SCALP algorithm described in "SCALP: superpixels with contour adherence using linear path" by Remi Giraud, Vinh-Thong Ta, and Nicolas Papadakis in CoRR in 2019, abs / 1903.07149; http: / / arxiv.org / abs / 1903.07149.
[0071] Furthermore, depending on the type of environment, the standard deviation threshold can be increased to 1m or more, especially when a large amount of clutter is expected.
[0072] Figure 2d An example of the resulting semantic classification is illustrated, in which black pixels are classified as unknown by majority vote, dark gray pixels are classified as foreground of the construction site, and light gray pixels are classified as background of the construction site.
[0073] It has been found that the quality of the classification, or the corresponding quality of the pseudo-labels, is sufficient to serve as a training signal for retraining the semantic model of the semantic classifier 18. A particular advantage of the method 10 according to the invention is that the classification can be generated instantaneously without any external supervision.
[0074] Domain Adaptation
[0075] To improve semantic classification, a semantic classifier 18 with a semantic model in the form of a neural network architecture is trained on different data sources.
[0076] To cater to the goals of online learning, the lightweight network architecture of Fast-SCNN, based on “Fast-SCNN: FastSemantic Segmentation Network” presented by Rudra PK Poudel, Stephan Liwicki and Roberto Cipolla at the 2019 British Conference on Machine Vision (BMVC), was used for the semantic classifier.18
[0077] The semantic model uses labeled data 26 (see Figure 1 Pre-training can be performed using datasets such as the NYUDepth v2 dataset (presented by Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus at the 2012 European Conference on Computer Vision (ECCV) as "Indoor Segmentation and Support Inference from RGBD Images"), which contains 1449 images extracted from video sequences of indoor scenes, each with per-pixel semantic annotations. In this dataset, the categories "walls," "ceiling," and "floor" are mapped to the construction site background category, while the remaining categories are mapped to the construction site foreground category.
[0078] Pre-training step 24 (see Figure 1 This is performed as an initial step to allow the semantic model to acquire prior knowledge, which can then be used as an inductive bias to perform the same semantic segmentation task on subsequently captured data.
[0079] Then, the semantic model is fine-tuned using self-supervision by generating pseudo-labels in real-world scenarios where mobile construction robots are deployed. Thus, the semantic model is continuously updated.
[0080] Following the naming conventions often used in continuous learning, each new environment is named a "task," and it is assumed that the "task boundaries" are known. Whenever the mobile construction robot moves to a new environment, the same scheme described above is applied: the semantic model trained on the previous environment is provided as new training data to generate categories or corresponding pseudo-labels from the current environment.
[0081] To achieve this domain adaptation and allow the semantic model to improve semantic segmentation accuracy in the current environment in which the semantic model is deployed without forgetting information learned from previous tasks, a memory-based replay buffer approach is adopted.
[0082] When adapting to a new environment, each training batch is filled with frames collected in the current environment and a small subset of images collected in previous environments, which in particular allows the semantic model to be continuously updated by mixing current and previously collected semantic segmentation data.
[0083] Therefore, in each training step, the semantic model will be jointly optimized using data from both the current and previous environments. However, storing all observations or semantic segmentation data from past environments in memory would require a large amount of memory.
[0084] To address this issue, according to the present invention, the memory buffer for each previous environment contains only a random subset of all images and pseudo-labels or label attributes.
[0085] Then, training batches are populated from the memory buffer of the previous environment and the self-supervised labels of the current environment.
[0086] According to a preferred variant of method 10 of the invention, batches of size 10 can be used. The semantic model can be trained, for example, up to 100 epochs. Preferably, an early stopping based on validation loss with 20 epochs of patience can be used. In different variants, the replay buffer can be filled according to a ratio between 1:1 and 200:1 between the target and NYU data, such as 1:1, 3:1, 4:1, 10:1, 20:1, or 200:1. The replay buffer can also be filled according to a replay fraction between 1% and 30%, for example, 5% or 10%. Preferably, a 10% replay fraction is used.
[0087] The semantic model based on Fast-SCNN preferably has a total of 1 to 2 × 10 6 Trainable parameters, for example, approximately 1.8 × 10⁻⁶. 6 indivual.
[0088] Group normalization can be used in all layers. Generally speaking, group normalization has been found to perform better than alternative normalization methods such as batch normalization for required transfer learning tasks.
[0089] In a variation of method 10, different strategies for memory replay can be used. The first strategy is the one employed in the variation of method 10 described above, where, on each source (e.g., the NYU dataset) to target (e.g., a construction site) migration, the replay buffer can be refilled with a subset of samples randomly selected from the source dataset(s) (here, the NYU dataset). Then, training batches are filled from the replay buffer and the target dataset according to the relative size of the training batch.
[0090] According to the alternative second strategy, the replay buffer can be filled with the entire source dataset, while the training batches are filled with a predefined target-to-source ratio. For example, a construction site to NYU ratio of 4:1 and a batch size of 10 indicate that the batch contains an average of 8 images from the construction site and 2 images from NYU.
[0091] Higher target-to-source ratios or lower replay scores have been shown to achieve higher performance in the target domain, but at the cost of less information retained from the source domain. Meanwhile, the segmentation quality of pseudo-labels on the target dataset follows the opposite trend; generally, lower replay values result in higher segmentation quality. This highlights the trade-off between accuracy in the target and source domains.
[0092] Other variations of method 10 according to the invention can utilize alternative continuous learning methods, such as regularization methods. One example could be a distillation method, wherein the weighted regularization term L... d Added to cross-entropy loss L ce This encourages semantic models to retain knowledge from previous tasks. Distillation can be applied to the output logit produced by the semantic model (output distillation) or to intermediate features extracted from the model architecture before the final classification module (feature distillation).
[0093] Another alternative could be Elastic Weight Consolidation (EWC), in which network parameter biases across tasks are penalized.
[0094] In summary, replay buffers have been shown to be very effective in minimizing the amount of forgetting on the NYU dataset and generally allow for a good trade-off in the segmentation quality of pseudo-labels from the target domain.
[0095] Construction robot system
[0096] Figure 3 A construction robot system 100 is shown on a construction site. The construction robot system 100 includes a mobile construction robot 102 and a control unit 104, which... Figure 3 The diagram is used to represent this.
[0097] In this embodiment, the control unit 104 is disposed inside the mobile construction robot 102. It includes a computing unit 106 and a computer program product 108, the computer program product including a storage device readable by the computing unit 106. The computing unit 106 includes a neural network unit configured as described in reference... Figure 1 The semantic classifier 18 is described. The storage device carries instructions that, when executed by the computing unit 106, cause the computing unit 106 to perform the method according to the present invention and as described above.
[0098] Furthermore, the mobile construction robot 102 includes a robotic arm 110. This includes an end effector on which a power tool 113 is detachably mounted. The power tool 113 may be a drilling machine. It may include a vacuum cleaning unit for automatically removing dust. The robotic arm 110 and / or the power tool 113 may also include a vibration damping unit.
[0099] The control unit 104 is configured to control the robotic arm 110. It may include other modules, such as a communication module, particularly for communication with, for example, external cloud computing systems. Figure 3 Wireless communication of components such as (not shown in the image), display units, etc.
[0100] Furthermore, the control unit 104 is configured to control the mobile base 116 of the mobile construction robot 102. In this embodiment, the mobile base 118 is a wheeled ground vehicle on which other components of the mobile construction robot 102 are mounted.
[0101] The mobile construction robot 102 can be configured to drill holes in floors, walls, and / or ceilings. For this purpose, the robotic arm 110 is preferably capable of movement with at least six degrees of freedom.
[0102] The mobile construction robot 102 includes multiple sensors and positioning support elements. Specifically, it includes a camera system 112 comprising three 2D cameras. It further includes a LiDAR scanner 114. These two elements, 112 and 114, are aligned with each other and connected to a control unit 104. Furthermore, the cameras of the camera system 112 are time-synchronized with the LiDAR scans from the LiDAR scanner 114.
[0103] The mobile construction robot 102 may further include other attitude sensing sensors or components, such as an inertial measurement unit (IMU). Figure 3 (Not shown in the image).
[0104] The mobile construction robot 102 may also include a reflector 110 or another detectable point, which may be used to additionally position the mobile construction robot 102, for example, by using a total station to measure, for example, the relative displacement between the initial posture and the current posture.
Claims
1. A method (10) for locating a mobile construction robot (102) at a construction site (101), wherein, The environmental data (12) of the construction site is captured, and the pose of the mobile robot (102) is inferred using the environmental data (12). The method (10) includes a semantic segmentation step, wherein a semantic classifier (18) with an updatable semantic model classifies the environmental data (12) into at least two semantic categories, at least one of which corresponds to the construction site background and at least one of which corresponds to the construction site foreground, and the pose is inferred using a subset of the environmental data (12) classified as the construction site background, wherein the semantic model is updated at least once by designating the point of the environmental data as the construction site foreground if the distance from the point of the environmental data to the BIM model data exceeds a threshold, otherwise designating the point of the environmental data as the construction site background, and updating the semantic model by using pseudo-labels generated therefrom as training data.
2. The method according to claim 1, characterized in that, Use the environmental data (12) to update the semantic model.
3. The method according to claim 1 or 2, characterized in that, The semantic segmentation of the environment data (12) is refined at least once using at least one additional dataset.
4. The method according to claim 3, characterized in that, The additional dataset is an image dataset.
5. The method according to claim 1 or 2, characterized in that, The orientation was inferred using map data from the construction site.
6. The method according to claim 5, characterized in that, The map data is BIM data (14).
7. The method according to claim 1 or 2, characterized in that, The semantic model has been updated multiple times.
8. The method according to claim 7, characterized in that, The semantic model is updated regularly multiple times.
9. The method according to claim 1 or 2, characterized in that, The semantic model is updated by mixing current and previously collected semantic segmentation data.
10. The method according to claim 1 or 2, characterized in that, The semantic category is fixed for all semantic segments.
11. A construction robot system (100), comprising a mobile construction robot (102) and a control unit (104), characterized in that, The control unit (104) is configured to position the mobile construction robot (102) using the method (10) according to any one of claims 1 to 10.
12. The construction robot system according to claim 11, characterized in that, The mobile construction robot (102) is used for drilling, chiseling, grinding, plastering and / or painting.
13. The construction robot system according to claim 11 or 12, characterized in that, The mobile construction robot (102) includes 3D sensors and at least one camera.
14. The construction robot system according to claim 13, characterized in that, The 3D sensor is a LiDAR scanner (114) or a stereo camera system.
15. A computer program product (108) comprising a storage device readable by a control unit (104) of a construction robot system (100), the storage device carrying instructions which, when executed by the control unit (104), cause the control unit to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Electronic device, robotic system and method for localizing a robotic system
WO2019185170A1