Method for generating labeled point cloud, model training method and semantic segmentation method

By acquiring real point clouds and virtual synthetic point clouds, and combining neural rendering methods to generate target annotation point clouds, the inefficiency and high cost problems caused by the point cloud annotation dependence on manual in the existing technology are solved, and an automatic and efficient annotation process is realized.

CN119992532APending Publication Date: 2025-05-13BOSHI SHANGHAI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311504023.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the labeling process of point cloud data relies on manual labor, resulting in inefficiency, time-consuming and costly.

Method used

By acquiring real point clouds and virtual synthetic point clouds, a target annotated point cloud is generated in combination with neural rendering methods, including points corresponding to physical objects and their semantic labels, as well as abnormal points.

Benefits of technology

It realizes automatic and efficient generation of labeled point clouds, reduces labor costs, and improves the quality and reliability of labeled point clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992532A_ABST
    Figure CN119992532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for generating an annotation point cloud and a computer readable storage medium. The method comprises the steps that real point cloud is acquired, the real point cloud is a point cloud obtained by performing data acquisition on a physical object, and the real point cloud comprises points corresponding to the physical object and abnormal points; a virtual synthesis point cloud is obtained, the virtual synthesis point cloud is a point cloud obtained based on a virtual model of the physical object, and each point in the virtual synthesis point cloud has a semantic tag corresponding to the physical object; and based on the real point cloud and the virtual synthesis point cloud, generating a target annotation point cloud, the target annotation point cloud including points corresponding to the physical object, semantic tags of the points, and points corresponding to the abnormal points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning technology, and in particular, to a method, device, and computer-readable storage medium for generating annotated point clouds. In addition, the present disclosure relates to a method for training a machine learning model and a method for semantically segmenting a point cloud. Background Art

[0002] Driven by various practical applications, 3D vision has become one of the research hotspots in the field of artificial intelligence in recent years. Among different forms of 3D data, point cloud has been widely used in many 3D scene understanding tasks due to its massive data and rich spatial information expression. Semantic segmentation, as a basic technology for point cloud data processing and analysis, has also received more and more attention.

[0003] With the rapid development of deep learning technology, semantic segmentation of point clouds based on deep learning has become one of the mainstream research directions. In order to achieve semantic segmentation of point clouds through deep learning architecture, it is usually necessary to train the deep learning architecture with labeled point cloud data. However, currently point cloud data is usually labeled manually, which is not only inefficient and time-consuming, but also consumes a lot of manpower costs. Summary of the invention

[0004] In view of the need for improvement of the prior art, embodiments of the present disclosure provide methods, devices and computer-readable storage media for generating annotated point clouds. In addition, embodiments of the present disclosure provide methods for training machine learning models and methods for semantically segmenting point clouds.

[0005] On the one hand, an embodiment of the present disclosure provides a method for generating an annotated point cloud, comprising: obtaining a real point cloud, wherein the real point cloud is a point cloud obtained by collecting data for a physical object, and the real point cloud includes points corresponding to the physical object and abnormal points; obtaining a virtual synthetic point cloud, wherein the virtual synthetic point cloud is a point cloud obtained based on a virtual model of the physical object, and each point in the virtual synthetic point cloud has a semantic label corresponding to the physical object; based on the real point cloud and the virtual synthetic point cloud, generating a target annotated point cloud, wherein the target annotated point cloud includes: each point corresponding to the physical object and the semantic label of each point, and points corresponding to the abnormal points.

[0006] On the other hand, an embodiment of the present disclosure provides a device for generating an annotated point cloud, comprising: a first acquisition unit, configured to acquire a real point cloud, wherein the real point cloud is a point cloud obtained by collecting data for a physical object, and the real point cloud includes points corresponding to the physical object and abnormal points; a second acquisition unit, configured to acquire a virtual synthetic point cloud, wherein the virtual synthetic point cloud is a point cloud obtained based on a virtual model of the physical object, and each point in the virtual synthetic point cloud has a semantic label corresponding to the physical object; a generation unit, configured to generate a target annotated point cloud based on the real point cloud and the virtual synthetic point cloud, wherein the target annotated point cloud includes: each point corresponding to the physical object and the semantic label of each point, and points corresponding to the abnormal points.

[0007] On the other hand, an embodiment of the present disclosure provides a device for generating annotated point clouds, comprising: at least one processor; a memory that communicates with the at least one processor and has executable code stored thereon, wherein the executable code, when executed by the at least one processor, causes the at least one processor to execute the above method.

[0008] On the other hand, an embodiment of the present disclosure provides a computer-readable storage medium storing executable codes, which implement the above method when executed.

[0009] On the other hand, an embodiment of the present disclosure provides a method for training a machine learning model, comprising: obtaining a target annotated point cloud, wherein the target annotated point cloud is generated using the above method; and using the target annotated point cloud as training data to train the machine learning model.

[0010] On the other hand, an embodiment of the present disclosure provides a method for semantically segmenting a point cloud, comprising: acquiring a point cloud obtained by collecting data on a physical object; performing semantic segmentation on the point cloud based on a pre-trained machine learning model; wherein the pre-trained machine learning model is obtained by training using the target annotated point cloud generated by the above method as training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the embodiments of the present disclosure will become more apparent through detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same elements in the various drawings.

[0012] Figure 1 is a schematic flow chart of a method for generating annotated point clouds according to some embodiments.

[0013] Figure 2An example of a graphical representation of a real point cloud and a virtual synthetic point cloud corresponding to the same building is shown.

[0014] Figure 3A A schematic diagram showing the representation results of the two-dimensional geometric features of the real point cloud and the virtual synthetic point cloud.

[0015] Figure 3B A schematic diagram showing the representation results of the three-dimensional geometric features of the real point cloud and the virtual synthetic point cloud.

[0016] Figure 4 Another graphical representation of the virtual synthetic point cloud is shown.

[0017] Figure 5 A graphical representation of the merged real point cloud and virtual synthetic point cloud is shown.

[0018] Figure 6 is a schematic diagram of a process for generating annotated point clouds according to some embodiments.

[0019] Figure 7 is a schematic diagram of an apparatus for generating annotated point clouds according to some embodiments.

[0020] Figure 8 is a schematic structural diagram of an apparatus for generating annotated point clouds according to some embodiments.

[0021] Fig. 9 is a schematic flowchart of a method for training a machine learning model according to some embodiments.

[0022] Fig.10 is a schematic flow chart of a method for semantically segmenting a point cloud according to some embodiments. DETAILED DESCRIPTION

[0023] The subject matter described herein will now be discussed with reference to various embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and are not intended to limit the scope of protection, applicability or examples set forth in the claims.

[0024] Semantic segmentation of point clouds is a technology that associates points with semantic labels, so effective and accurate semantic segmentation is the basis for understanding three-dimensional scenes or spaces. For example, in construction scenes, the real point clouds scanned from the construction site can be understood and analyzed (including semantic segmentation), so that construction projects can be automatically monitored and tracked.

[0025] Nowadays, machine learning models (such as deep learning architectures) have been used to achieve semantic segmentation tasks of point clouds. In order to achieve good semantic segmentation results, it is usually necessary to train the deep learning architecture based on annotated point cloud data. At present, it is usually necessary to manually annotate the point cloud data to obtain the training point cloud data of the deep learning architecture. However, this manual method is not only inefficient and time-consuming, but also consumes a lot of manpower costs.

[0026] In view of this, the embodiments of the present disclosure provide a technical solution for generating annotated point clouds. The embodiments of the present disclosure can be applied to various suitable scenes, such as the above-mentioned building scenes. The following will be described in detail in conjunction with specific embodiments.

[0027] Figure 1 is a schematic flow chart of a method for generating annotated point clouds according to some embodiments.

[0028] like Figure 1 As shown, in step 102, a real point cloud may be obtained.

[0029] The real point cloud may be a point cloud obtained by collecting data for a physical object, and may generally include points corresponding to the physical object and abnormal points.

[0030] In step 104 , a virtual synthetic point cloud may be obtained.

[0031] The virtual synthetic point cloud may be a point cloud obtained based on a virtual model of a physical object. Each point in the virtual synthetic point cloud may generally have a semantic label corresponding to the physical object.

[0032] Specifically, a virtual synthetic point cloud can be a point cloud obtained based on a virtual model of a physical object (e.g., a BIM model), and each point of the point cloud can have a semantic label corresponding to the physical object. A real point cloud can be a point cloud obtained by collecting data (e.g., scanning) on ​​a physical object. For example, a device such as a laser radar can be used to scan a physical object to obtain a real point cloud.

[0033] In step 106 , a target annotated point cloud may be generated based on the real point cloud and the virtual synthetic point cloud.

[0034] The target annotated point cloud may include: points corresponding to physical objects and semantic labels of the points, and points corresponding to abnormal points.

[0035] Since the virtual synthetic point cloud is a point cloud obtained based on a virtual model, the virtual synthetic point cloud is usually a clean point cloud, and each point has a clear semantic label or category information. Compared with the virtual synthetic point cloud, the real point cloud is a point cloud obtained by scanning a physical object. Therefore, due to various factors such as errors in point cloud scanning equipment (for example, lidar, etc.) and the complexity of the on-site environment, the real point cloud may usually contain some abnormal points (such as noise points and / or outliers). In an embodiment of the present disclosure, a target annotated point cloud can be generated based on the real point cloud and the virtual synthetic point cloud, so that the target annotated point cloud includes each point corresponding to the physical object and the semantic labels of these points, and also includes abnormal points. It can be seen that through such a scheme, the annotated point cloud can be obtained automatically and efficiently, thereby greatly reducing the labor cost.

[0036] In some embodiments, abnormal points may include noise points and / or outliers.

[0037] In some embodiments, the target annotated point cloud may further include semantic labels corresponding to points of abnormal points. Thus, each point in the target annotated point cloud will have a corresponding semantic label.

[0038] In an embodiment of the present disclosure, the pose of a virtual synthetic point cloud may be aligned with the pose of a real point cloud. The virtual synthetic point cloud may be obtained in various appropriate ways. For example, an initial synthetic point cloud may be generated based on a virtual model of a physical object, such as by deriving the initial synthetic point cloud from a virtual model of a physical object through various applicable techniques. Since the initial synthetic point cloud is obtained from a virtual model, and the real point cloud is obtained by scanning a physical object, the initial synthetic point cloud and the real point cloud may have different poses (e.g., position and orientation). In order to facilitate subsequent processing, the initial synthetic point cloud may be subjected to a pose transformation to obtain a virtual synthetic point cloud. For example, a pose transformation may include a rotation transformation, a translation transformation, and the like, such as a rotation transformation, a translation transformation, and the like in a three-dimensional coordinate system (on the x-axis, y-axis, and z-axis).

[0039] For example, in a construction scene, the physical object may be a building. Typically, a virtual model corresponding to the building, such as a Building Information Modeling (BIM) model, may be constructed during the building design phase. Based on the BIM model, an initial synthetic point cloud of the buildings included in the construction project may be generated, and then the initial synthetic point cloud may be transformed in position to obtain the virtual synthetic point cloud. Accordingly, it can be understood that the BIM model typically contains an accurate and clear building structure, so each point of the virtual synthetic point cloud obtained from the BIM model will have a clear semantic label. The BIM model may provide integrated information including geometric features and semantic information of the construction project.

[0040] In an embodiment of the present disclosure, the target annotated point cloud may be close to the real point cloud, so that the quality of the target annotated point cloud can be ensured. For example, the degree of proximity between the target annotated point cloud and the real point cloud may be within a first predetermined range. For another example, the degree of proximity between the abnormal points in the target annotated point cloud and the abnormal points in the real point cloud may be within a second predetermined range. The first predetermined range and the second predetermined range may be set according to actual application scenarios, business requirements, etc., and are not limited herein. It can be seen that through such constraints, the quality of the target annotated point cloud can be ensured.

[0041] In some embodiments, the closeness between the target annotated point cloud and the real point cloud can be characterized by the closeness between the three-dimensional geometric features of the target annotated point cloud and the three-dimensional geometric features of the real point cloud. In this case, the closeness between the three-dimensional geometric features of the target annotated point cloud and the three-dimensional geometric features of the real point cloud can be within a first predetermined range.

[0042] In some embodiments, the degree of proximity between the abnormal points in the target annotation point cloud and the abnormal points in the real point cloud can be characterized by the degree of proximity between the three-dimensional geometric features of the abnormal points in the target annotation point cloud and the three-dimensional geometric features of the abnormal points in the real point cloud. In this case, the degree of proximity between the three-dimensional geometric features of the abnormal points in the target annotation point cloud and the three-dimensional geometric features of the abnormal points in the real point cloud can be within a second predetermined range.

[0043] In some embodiments, in step 106, a target annotated point cloud may be generated based on the real point cloud and the virtual synthetic point cloud using a neural rendering method. In such an embodiment, by using a neural rendering method to process the real point cloud and the virtual synthetic point cloud to generate the target annotated point cloud, the annotated point cloud can be made closer to the real point cloud, thereby improving the quality of the annotated point cloud.

[0044] In some embodiments, in step 106, a set of abnormal points may be obtained from the real point cloud. The set of abnormal points may be added to the virtual synthetic point cloud to generate an initial annotated point cloud. Then, the initial annotated point cloud may be processed using a neural network implementing a neural rendering method to generate a target annotated point cloud.

[0045] A real point cloud may generally include points corresponding to physical objects. In addition, a real point cloud may also include abnormal points. For example, abnormal points may include noise points, outliers, and the like. Noise points may be erroneous points caused by factors such as errors in the point cloud scanning device and interference from the on-site environment. Outliers may be caused by factors such as the irregular or non-fixed existence of objects other than physical objects, occlusion, and the like. In some cases, a real point cloud may also include points corresponding to other fixed background objects (background objects are relative to the physical objects) other than physical objects. Since abnormal points such as noise points and / or outliers usually have a significant impact on the training effect of subsequent machine learning models, in an embodiment of the present disclosure, abnormal points can be obtained from a real point cloud by methods such as noise point / outlier point filtering or analysis, without extracting points of fixed background objects.

[0046] In addition, since the virtual synthetic point cloud is a clean point cloud, adding abnormal points to the virtual synthetic point cloud can make the target annotated point cloud closer to the real point cloud. In this way, the abnormal points in the target annotated point cloud can be more conducive to the comprehensive training of the deep learning architecture. It can be seen that in this way, not only can the annotated point cloud be obtained automatically and efficiently, greatly reducing the labor cost, but also the high quality of the annotated point cloud can be ensured.

[0047] In some embodiments, a group of abnormal points and the semantic labels of these points can be added to the virtual synthetic point cloud to generate an initial annotated point cloud. In this way, the abnormal points in the initial annotated point cloud can have corresponding semantic labels.

[0048] In some embodiments, a set of three-dimensional geometric features of non-normal points may be determined. Based on the three-dimensional geometric features of the set of non-normal points, the set of non-normal points may be added to the virtual synthetic point cloud to generate an initial annotated point cloud.

[0049] In some implementations, the three-dimensional geometric features of the real point cloud may be determined. The three-dimensional geometric features of the real point cloud may include three-dimensional geometric features of a set of non-normal points.

[0050] In the embodiments of the present disclosure, the three-dimensional geometric features mentioned may include point density features and point distribution features. Point density features may represent the density of points in a point cloud, and point distribution features may represent the distribution of points in a point cloud. Three-dimensional geometric features may be characterized by various applicable methods, such as an octree structure, a K-dimensional tree (KD tree), an R-star tree, etc., which are not limited herein. For example, the node distribution status and depth of an octree structure may effectively characterize point density features and point distribution features.

[0051] In some embodiments, the point density features and point distribution features of the real point cloud can be determined. In some embodiments, an octree structure corresponding to the real point cloud can be determined. The octree structure can be used to represent the point density features and point distribution features of the real point cloud.

[0052] The three-dimensional geometric features of the group of abnormal points may include the three-dimensional geometric features of the group of abnormal points in the real point cloud, such as the point density features and point distribution features of the group of abnormal points in the real point cloud. In this case, a group of abnormal points may be added (e.g., inpainted) to the virtual synthetic point cloud to obtain an initial annotated point cloud, so that the three-dimensional geometric features of the group of abnormal points in the initial annotated point cloud match their three-dimensional geometric features in the real point cloud.

[0053] For example, the three-dimensional geometric features of the group of abnormal points in the initial annotation point cloud are close to the three-dimensional geometric features in the real point cloud, such as being within a predetermined error range. Then, it can be considered that the three-dimensional geometric features of the group of abnormal points in the initial annotation point cloud match the three-dimensional geometric features in the real point cloud. For example, the point density of the group of abnormal points in the initial annotation point cloud can be close to the point density in the real point cloud, such as the error between the two is within a predetermined point density error range. The point distribution of the group of abnormal points in the initial annotation point cloud can be close to the point distribution in the real point cloud, such as the error between the two is within a predetermined point distribution error range. In this case, it can be considered that the three-dimensional geometric features of the group of abnormal points in the initial annotation point cloud match the three-dimensional geometric features in the real point cloud. The various error ranges mentioned here can be set according to actual application scenarios and business needs, and this article does not limit this. Of course, ideally, the three-dimensional geometric features of the group of abnormal points in the initial annotation point cloud can be the same as the three-dimensional geometric features in the real point cloud.

[0054] As can be seen from the above, the initial annotated point cloud is actually a roughly obtained annotated point cloud. Therefore, in order to obtain a high-quality target annotated point cloud, a neural network can be used to process the initial annotated point cloud. For example, the neural network can be trained based on the real point cloud and the initial annotated point cloud until the neural network processes the initial annotated point cloud to obtain the target annotated point cloud. In this case, the degree of proximity between the target annotated point cloud and the real point cloud is within a predetermined range. The predetermined range can be determined based on factors such as actual application requirements.

[0055] For example, the closeness between the target annotated point cloud and the real point cloud can be characterized by the overall closeness between the target annotated point cloud and the real point cloud and / or the closeness between abnormal points in the target annotated point cloud and abnormal points in the real point cloud.

[0056] The training of the neural network can be understood as a process of continuously adjusting the parameters of the neural network using the real point cloud and the results generated by the neural network until the expected result is achieved. In some implementations, the neural network can be used to predict the initial annotated point cloud to obtain a predicted point cloud. For example, the neural network is used to adjust the abnormal points in the initial annotated point cloud, such as adding randomly generated abnormal points to the initial annotated point cloud, reducing the abnormal points previously added to the initial annotated point cloud, etc., to predict the abnormal points. In some cases, if the neural network adds randomly generated abnormal points to the initial annotated point cloud, the neural network can also add semantic labels to these randomly generated points to indicate that these points belong to abnormal points. Then, it is determined whether the degree of proximity between the predicted point cloud and the real point cloud is within a predetermined range; if the degree of proximity is not within the predetermined range, the parameters of the neural network are adjusted. These processes can be performed repeatedly. For example, after adjusting the parameters, the neural network predicts the initial annotated point cloud again to obtain a predicted point cloud, and the degree of proximity between the predicted point cloud and the real point cloud is determined again. If the degree of proximity is not within the predetermined range, the parameters of the neural network are adjusted again. The above process can be terminated until the degree of proximity between the predicted point cloud generated from the neural network and the real point cloud is within a predetermined range. For the convenience of description, the predicted point cloud whose degree of proximity to the real point cloud is within a predetermined range is called the target predicted point cloud. At this time, the target predicted point cloud can be used as the target annotated point cloud. The entire processing process described above about the neural network can also be understood as the domain adaptation process of the neural network.

[0057] The neural network can be implemented using various applicable technologies. For example, the neural network can be any applicable neural rendering network, such as CycleGAN (Cycle Generative Adversarial Network), MSG-Point-GAN (Multi-Scale Gradient Point Generative Adversarial Network), TreeGAN, etc., which is not limited in this article.

[0058] In some embodiments, the predetermined range may include a first loss range and / or a second loss range. The first loss range and / or the second loss range may be set according to factors such as actual application requirements, and this document does not limit this. Accordingly, determining whether the degree of proximity between the predicted point cloud and the real point cloud is within the predetermined range may include two parts, namely, determining whether the loss between the abnormal point in the predicted point cloud and the above-mentioned group of abnormal points is within the first loss range, and / or determining whether the loss between the predicted point cloud and the real point cloud is within the second loss range. The loss between the abnormal point in the predicted point cloud and the above-mentioned group of abnormal points may include the loss between the three-dimensional geometric features of the abnormal point in the predicted point cloud and the three-dimensional geometric features of the above-mentioned group of abnormal points. The loss between the predicted point cloud and the real point cloud may include the loss between the three-dimensional geometric features of the predicted point cloud and the three-dimensional geometric features of the real point cloud. It can be seen that through any one of the above losses or these two losses, the parameters of the neural network can be adjusted more finely, thereby obtaining a high-quality target annotated point cloud. As mentioned above, the abnormal points in the initial annotated point cloud have corresponding semantic labels. In addition, if the neural network adds randomly generated abnormal points in the initial annotated point cloud, the semantic labels of these points can be added together. Therefore, each point corresponding to the physical object and the abnormal point in the predicted point cloud obtained from the neural network can have a corresponding semantic label. In this way, the abnormal point can be identified from the predicted point cloud, so as to determine the loss between the abnormal point in the predicted point cloud and the abnormal point in the real point cloud. The target annotated point cloud finally generated can also include the semantic labels of the abnormal points.

[0059] In order to more clearly understand the embodiments of the present disclosure, the following will be further described in conjunction with specific examples. In the following examples, the building scene will be used as an example for illustration. It should be understood that the following examples do not impose any limitation on the scope of the technical solution of the present disclosure.

[0060] Figure 2 An example of a graphical representation of a real point cloud and a virtual synthetic point cloud corresponding to the same building is shown.

[0061] exist Figure 2 In the example of , the real point cloud 202 may be obtained by scanning a building (eg, using a laser radar, etc.). The virtual synthetic point cloud 204 may be obtained based on a BIM model of the building.

[0062] Figure 3A A schematic diagram showing the representation results of the two-dimensional geometric features of the real point cloud 202 and the virtual synthetic point cloud 204 is shown.

[0063] Figure 3AThe point density features and point distribution features of the two point clouds on the x-axis and y-axis in the world coordinate system are shown after the poses of the two point clouds are aligned. The result 302A corresponds to the real point cloud 202, and the result 304A corresponds to the virtual synthetic point cloud 204. Figure 3A In the figure, the ruler 306 indicates the relationship between different grayscales and the number of points, and the number on the right side of the ruler 306 indicates the number of points. It can be seen that although the real point cloud 202 and the virtual synthetic point cloud 204 both correspond to the same building, the point density and point distribution of the real point cloud 202 and the virtual synthetic point cloud 204 are obviously different on the x-axis and y-axis.

[0064] Figure 3B A schematic diagram showing the representation results of the three-dimensional geometric features of the real point cloud 202 and the virtual synthetic point cloud 204 is shown.

[0065] Figure 3B The geometric feature representation of the two point clouds in the three-dimensional world coordinate system (x-axis, y-axis and z-axis) after the poses of the two point clouds are aligned is shown. The representation result 302B corresponds to the real point cloud 202, and the representation result 304B corresponds to the virtual synthetic point cloud 204. Similarly, although the real point cloud 202 and the virtual synthetic point cloud 204 both correspond to the same building, the real point cloud 202 and the virtual synthetic point cloud 204 have obvious differences in three-dimensional geometric features.

[0066] Figure 3A and Figure 3B The differences shown in the figure may be due to a variety of reasons. For example, the virtual synthetic point cloud is obtained based on the BIM model, which does not contain noise points and / or outliers. The real point cloud is obtained by scanning on-site in the building. Due to factors such as the error of the acquisition equipment, the complexity of the site, the presence of other objects, etc., the real point cloud usually contains noise points and / or outliers. In addition, due to the different ways of generating the two point clouds, the geometric features will also be different.

[0067] In addition, from Figure 3A and Figure 3B It can be seen that the point cloud contains three-dimensional information, so the two-dimensional geometric features may not be able to fully express the characteristics of the point cloud. Therefore, in the embodiments of the present disclosure, the three-dimensional geometric features of the point cloud are used as indicators to add abnormal points in the virtual synthetic point cloud and evaluate the loss between the predicted point cloud and the real point cloud, which can greatly improve the quality and reliability of the target annotated point cloud obtained in the end.

[0068] Figure 4 Another graphical representation 402 of the virtual synthetic point cloud 204 is shown.

[0069] exist Figure 4In , points with different grayscales can correspond to different structures of the building. It can be seen that the virtual synthetic point cloud is a clean point cloud without noise and outliers, and each point can have a clear semantic label.

[0070] In contrast, the real point cloud 202 contains points that do not belong to buildings, such as noise points and / or outliers, in addition to points corresponding to buildings. Figure 5 It can be seen more clearly in. Figure 5 A graphical representation of a merged real point cloud 202 and a virtual synthetic point cloud 204 is shown.

[0071] It can be seen here that in order to obtain the target annotated point cloud from the virtual synthetic point cloud, not only should abnormal points (e.g., noise points and / or outliers) be added to the virtual synthetic point cloud, but also they should be close to or matched with the real point cloud after addition, so that the obtained target annotated point cloud can be effectively used as a training data set for the subsequent deep learning architecture. In addition, noise and / or outliers are very important for generating the target annotated point cloud, because noise and / or outliers may have a significant impact on the training effect of the deep learning model.

[0072] Figure 6 is a schematic diagram of a process for generating annotated point clouds according to some embodiments. Figure 6 The construction scene is still used as an example for description.

[0073] like Figure 6 As shown, in step 612 of phase 610, an initial synthetic point cloud may be generated based on the BIM model of the building, which may be achieved by using various applicable technologies, which are not limited herein.

[0074] In step 614 of stage 610, the initial synthetic point cloud may be subjected to a pose transformation to obtain a virtual synthetic point cloud. The pose of the virtual synthetic point cloud may be aligned with the pose of the real point cloud. For example, the pose transformation may include a rotation transformation and / or a translation transformation, such as a rotation transformation and / or a translation transformation on the x-axis, y-axis, and z-axis in a three-dimensional coordinate system.

[0075] In stage 620, an abnormal point analysis may be performed on the real point cloud based on the virtual synthetic point cloud to obtain a set of abnormal points from the real point cloud. For example, the virtual synthetic point cloud may be compared with the real point cloud to determine which points in the real point cloud are abnormal points. Figure 5 As shown, the virtual synthetic point cloud can be merged with the real point cloud to determine the abnormal points in the real point cloud. The group of abnormal points can include noise points and / or outliers.

[0076] In step 632 of stage 630, the three-dimensional geometric features of the real point cloud may be determined. For example, an octree structure corresponding to the real point cloud may be generated. The octree structure may represent the three-dimensional geometric features of the real point cloud, such as point density features and point distribution features. Specifically, the distribution status and depth status of the nodes of the octree structure may indicate the point density features and point distribution features of the real point cloud. The three-dimensional geometric features of the real point cloud may include the three-dimensional geometric features of the above-mentioned set of abnormal points, such as the point density features and point distribution features of the above-mentioned set of abnormal points.

[0077] In step 634 of stage 630, the group of abnormal points may be added to the virtual synthetic point cloud based on the three-dimensional geometric features of the real point cloud to obtain an initial annotated point cloud. In addition, the semantic labels of the group of abnormal points may be added to the virtual synthetic point cloud. The group of abnormal points in the initial annotated point cloud has the three-dimensional geometric features of the group of abnormal points in the real synthetic point cloud, such as point density features and point distribution features.

[0078] In stage 640, a neural rendering network (eg, CycleGAN) is used to perform neural rendering processing on the initial labeled point cloud. Specifically, in step 642, a neural rendering network is used to predict the initial labeled point cloud to obtain a predicted point cloud.

[0079] In step 644, it may be determined whether the degree of proximity between the predicted point cloud and the real point cloud is within a predetermined range. For example, the loss between the three-dimensional geometric features of the predicted point cloud and the three-dimensional geometric features of the real point cloud and the loss between the three-dimensional geometric features of the abnormal points in the predicted point cloud and the three-dimensional geometric features of the above-mentioned set of abnormal points may be determined. In this case, the predetermined range may include a first loss range and a second loss range. It may be determined whether the loss between the three-dimensional geometric features of the predicted point cloud and the three-dimensional geometric features of the real point cloud is within a first loss range, and it may be determined whether the loss between the three-dimensional geometric features of the abnormal points in the predicted point cloud and the three-dimensional geometric features of the above-mentioned set of abnormal points is within a second loss range.

[0080] In step 646, if the degree of proximity between the predicted point cloud and the real point cloud is not within the predetermined range, the parameters of the neural rendering network are adjusted. After that, the process returns to step 642, and 642 and 644 are repeatedly performed. For example, the initial labeled point cloud is continuously predicted using the neural rendering network with the adjusted parameters, and it is continuously determined whether the degree of proximity between the predicted point cloud and the real point cloud is within the predetermined range.

[0081] If in step 644, it is determined that the degree of proximity between the predicted point cloud and the real point cloud is within a predetermined range, the process may proceed to step 648. In step 648, the predicted point cloud currently obtained in step 642 may be used as the target annotated point cloud. In other words, the degree of proximity between the target annotated point cloud and the real point cloud is within a predetermined range.

[0082] It can be seen that through the embodiments of the present disclosure, a high-quality annotated point cloud can be generated based on a virtual synthetic point cloud obtained from a BIM model of a building. This method can greatly reduce labor costs and time consumption, and can be effectively used for training deep learning networks for building scenes.

[0083] Figure 7 is a schematic diagram of an apparatus for generating annotated point clouds according to some embodiments.

[0084] like Figure 7 As shown, the apparatus 700 may include a first acquisition unit 702 , a second acquisition unit 704 , and a generation unit 708 .

[0085] The first acquisition unit 702 may acquire a real point cloud. The real point cloud may be a point cloud obtained by collecting data for a physical object. The real point cloud may include points corresponding to the physical object and abnormal points. The second acquisition unit 704 may acquire a virtual synthetic point cloud. The virtual synthetic point cloud may be a point cloud obtained based on a virtual model of the physical object. Each point in the virtual synthetic point cloud may have a semantic label corresponding to the physical object.

[0086] The generating unit 706 may generate a target annotated point cloud based on the real point cloud and the virtual synthetic point cloud. The target annotated point cloud may include: points corresponding to the physical object and semantic labels of these points, and points corresponding to abnormal points.

[0087] Each unit of the device 700 can execute the specific process described above with respect to the method embodiment. Therefore, for the sake of brevity of description, the specific operations and functions of each unit of the device 700 are not repeated here.

[0088] Figure 8 is a schematic structural diagram of an apparatus for generating annotated point clouds according to some embodiments.

[0089] like Figure 8 As shown, the apparatus 800 may include a processor 802, a memory 804, an input interface 806, and an output interface 808, and these modules may be coupled together via a bus 810. However, it should be understood that Figure 8 This is only an example and does not limit the scope of the present disclosure. For example, in different application scenarios, the device 800 may include more or fewer modules, which is not limited herein.

[0090] The memory 804 may be used to store various information related to the functions or operations of the device 800 (such as the real point cloud, virtual synthetic point cloud, three-dimensional geometric features, etc. mentioned above), executable instructions or codes, etc. For example, the memory 804 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), registers, hard disks, etc.

[0091] The processor 802 may be used to execute or implement various functions or operations of the device 800, such as the various operations described herein. For example, the processor 802 may execute executable instructions or codes stored in the memory 804, thereby implementing the various processes described above. The processor 802 may include various applicable processors, such as a general-purpose processor (such as a central processing unit (CPU)), a special-purpose processor (such as a digital signal processor, a special-purpose integrated circuit), and the like.

[0092] The input interface 806 can receive data in various forms, such as the above-mentioned real point cloud, initial synthetic point cloud, etc. The output interface 808 can output data in various forms, such as the above-mentioned target annotated point cloud, etc.

[0093] The embodiments of the present disclosure further provide a computer-readable storage medium. The computer-readable storage medium may store executable codes, and the executable codes implement the above-mentioned various processes when executed.

[0094] For example, computer-readable storage media may include, but are not limited to, RAM, ROM, Electrically-Erasable Programmable Read-Only Memory (EEPROM), Static Random Access Memory (SRAM), hard disk, flash memory, and the like.

[0095] Fig. 9 is a schematic flowchart of a method for training a machine learning model according to some embodiments.

[0096] like Fig. 9 As shown, in step 902, a target annotated point cloud may be obtained. The target annotated point cloud may be generated by using the process described in the above method embodiment.

[0097] In step 904, the target annotated point cloud may be used as training data to train the machine learning model.

[0098] In the embodiments of the present disclosure, by training the machine learning model using the annotated point cloud generated by the aforementioned process, the training effect of the machine learning model can be effectively improved, so that the machine learning model can complete the task more efficiently and with higher quality.

[0099] For example, a machine learning model can use a deep learning architecture. A machine learning model can perform various tasks, such as semantic segmentation of point clouds. For example, a machine learning model can perform semantic segmentation on a point cloud collected from a building scene.

[0100] Fig.10 is a schematic flow chart of a method for semantically segmenting a point cloud according to some embodiments.

[0101] like Fig.10 As shown, in step 1002, a point cloud obtained by collecting data on a physical object may be acquired.

[0102] In step 1004, semantic segmentation may be performed on the point cloud based on a pre-trained machine learning model.

[0103] The pre-trained machine learning model is trained using the target annotated point cloud generated by the process described in the previous method embodiment as training data.

[0104] Specific embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0105] Not all steps and units in the above processes and system structure diagrams are necessary, and some steps or units may be omitted according to actual needs. The device structure described in the above embodiments may be a physical structure or a logical structure, that is, some units may be implemented by the same physical entity, some units may be implemented by multiple physical entities, or may be implemented by some components in multiple independent devices.

[0106] The optional implementation modes of the embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present disclosure are not limited to the specific details in the above implementation modes. Within the technical concept of the embodiments of the present disclosure, various modifications can be made to the technical solutions of the embodiments of the present disclosure, and these modifications all belong to the protection scope of the embodiments of the present disclosure.

Claims

1. A method for generating annotated point clouds, comprising: Acquire a real point cloud, wherein the real point cloud is a point cloud obtained by collecting data on a physical object, and the real point cloud includes points corresponding to the physical object and abnormal points; Acquire a virtual synthetic point cloud, wherein the virtual synthetic point cloud is a point cloud obtained based on a virtual model of the physical object, and each point in the virtual synthetic point cloud has a semantic label corresponding to the physical object; Based on the real point cloud and the virtual synthetic point cloud, a target annotated point cloud is generated, wherein the target annotated point cloud includes: points corresponding to the physical object and semantic labels of the points, and points corresponding to the abnormal points.

2. The method according to claim 1, wherein: Generating the target annotation point cloud includes: Based on the real point cloud and the virtual synthetic point cloud, a neural rendering method is used to generate the target annotated point cloud.

3. The method according to claim 2, wherein: Generating the target annotation point cloud includes: Acquire a set of abnormal points from the real point cloud; adding the set of abnormal points to the virtual synthetic point cloud to generate an initial annotated point cloud; The initial annotated point cloud is processed using a neural network to generate the target annotated point cloud, wherein the neural network implements the neural rendering method.

4. The method according to claim 3, wherein: Generating the initial annotated point cloud includes: determining three-dimensional geometric features of the set of abnormal points; Based on the three-dimensional geometric features of the group of abnormal points, the group of abnormal points are added to the virtual synthetic point cloud to generate the initial annotated point cloud.

5. The method according to claim 4, wherein: Determining the three-dimensional geometric features of the set of abnormal points includes: The three-dimensional geometric features of the real point cloud are determined, wherein the three-dimensional geometric features of the real point cloud include the three-dimensional geometric features of the group of abnormal points.

6. The method according to claim 5, wherein: Determining the three-dimensional geometric features of the real point cloud includes: Determine the point density characteristics and point distribution characteristics of the real point cloud.

7. The method according to claim 6, wherein: Determining the point density feature and the point distribution feature of the real point cloud includes: An octree structure corresponding to the real point cloud is determined, wherein the octree structure is used to represent point density features and point distribution features of the real point cloud.

8. The method according to claim 4, wherein: The three-dimensional geometric features of the group of abnormal points include the three-dimensional geometric features of the group of abnormal points in the real point cloud; Adding the set of abnormal points to the virtual synthetic point cloud to generate an initial annotated point cloud includes: The group of abnormal points is added to the virtual synthetic point cloud to generate an initial annotated point cloud, so that the three-dimensional geometric features of the group of abnormal points in the initial annotated point cloud match the three-dimensional geometric features of the group of abnormal points in the real point cloud.

9. The method according to claim 3, wherein: Adding the set of abnormal points to the virtual synthetic point cloud to generate the initial annotated point cloud includes: The group of abnormal points and the semantic labels of the group of abnormal points are added to the virtual synthetic point cloud to generate the initial annotated point cloud.

10. The method according to claim 3, wherein: The initial annotated point cloud is processed using a neural network to generate a target annotated point cloud, including: The neural network is trained based on the real point cloud and the initial annotated point cloud until the neural network processes the initial annotated point cloud to obtain the target annotated point cloud, wherein the degree of proximity between the target annotated point cloud and the real point cloud is within a predetermined range.

11. The method according to claim 10, wherein: The neural network is trained based on the real point cloud and the initial annotated point cloud until the neural network processes the initial annotated point cloud to obtain a target annotated point cloud, including: Repeat the following process: predict the initial annotated point cloud using the neural network to obtain a predicted point cloud; determine whether the degree of proximity between the predicted point cloud and the real point cloud is within the predetermined range; if the degree of proximity is not within the predetermined range, adjust the parameters of the neural network; Until a target predicted point cloud is generated from the neural network, the target predicted point cloud is used as the target annotated point cloud, wherein the degree of proximity between the target predicted point cloud and the real point cloud is within the predetermined range.

12. The method according to claim 11, wherein: The predetermined range includes a first loss range and a second loss range; Determining whether the degree of proximity between the predicted point cloud and the actual point cloud is within the predetermined range includes: determining whether a loss between an abnormal point in the predicted point cloud and the set of abnormal points is within the first loss range; and / or, It is determined whether the loss between the predicted point cloud and the real point cloud is within the second loss range.

13. The method according to claim 12, wherein: The loss between the abnormal point in the predicted point cloud and the group of abnormal points includes the loss between the three-dimensional geometric features of the abnormal point in the predicted point cloud and the three-dimensional geometric features of the group of abnormal points, The loss between the predicted point cloud and the real point cloud includes the loss between the three-dimensional geometric features of the predicted point cloud and the three-dimensional geometric features of the real point cloud.

14. The method according to claim 1, wherein: The degree of proximity between the target annotated point cloud and the real point cloud is within a first predetermined range; and / or, The degree of proximity between the abnormal points in the target annotation point cloud and the abnormal points in the real point cloud is within a second predetermined range.

15. The method according to claim 1, further comprising: generating an initial synthetic point cloud based on the virtual model of the physical object; The initial synthetic point cloud is subjected to a posture transformation to obtain the virtual synthetic point cloud, wherein the posture of the virtual synthetic point cloud is aligned with the posture of the real point cloud.

16. The method according to claim 1, wherein: The set of abnormal points includes noise points and / or outliers; and / or, The target annotated point cloud also includes: semantic labels of the abnormal points.

17. A device for generating annotated point clouds, comprising: A first acquisition unit is configured to acquire a real point cloud, wherein the real point cloud is a point cloud obtained by collecting data for a physical object, and the real point cloud includes points corresponding to the physical object and abnormal points; A second acquisition unit is configured to acquire a virtual synthetic point cloud, wherein the virtual synthetic point cloud is a point cloud obtained based on a virtual model of the physical object, and each point in the virtual synthetic point cloud has a semantic label corresponding to the physical object; A generating unit is configured to generate a target annotated point cloud based on the real point cloud and the virtual synthetic point cloud, wherein the target annotated point cloud includes: points corresponding to the physical object and semantic labels of the points, and points corresponding to the abnormal points.

18. A method for training a machine learning model, comprising: Acquire a target annotated point cloud, wherein the target annotated point cloud is generated by using the method according to any one of claims 1 to 16; The target annotated point cloud is used as training data to train the machine learning model.

19. The method according to claim 18, wherein: The machine learning model is used to perform semantic segmentation on point clouds collected from architectural scenes.

20. A method for semantic segmentation of a point cloud, comprising: Obtaining a point cloud obtained by collecting data on a physical object; Performing semantic segmentation on the point cloud based on a pre-trained machine learning model; The pre-trained machine learning model is obtained by training using a target annotated point cloud generated by a method according to any one of claims 1 to 16 as training data.

21. A device for generating annotated point clouds, comprising: at least one processor; A memory in communication with the at least one processor, having executable codes stored thereon, which, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 16.

22. A computer-readable storage medium storing executable codes, wherein the executable codes implement the method according to any one of claims 1 to 16 when executed.