Large-scale point cloud completion method and device
By using generative adversarial networks to deeply extract multi-scale point cloud features and fuse them with high-resolution image data, the problem of inaccurate and unreliable data in point cloud completion methods is solved, thereby improving the accuracy and adaptability of point cloud data.
Patent Information
- Application Number
- CN202411509063.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-10-28
Smart Images

Figure CN119559093B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a large-scale point cloud completion method and device. BACKGROUND
[0002] In the actual three-dimensional modeling process, the limitation of sensor resolution and the occlusion factor often lead to the partial loss of geometric features and semantic information. Although the point cloud data contains detailed three-dimensional details, due to the existence of outliers and edge blur problems, it significantly increases the difficulty of object recognition.
[0003] In the prior art, the PointNet architecture captures and integrates the local features of individual points and the overall shape features in the point cloud data through its unique shared multi-layer perceptron (Shared-MLP) module. However, when processing point cloud data sets containing multiple complex topological structures, its performance is not satisfactory. The Point Completion Network (PCN) extracts point cloud feature vectors based on the multi-layer perceptron (MLP) mechanism in PointNet, and innovatively combines a voxel-based generator in the decoding stage to gradually refine and perfect the reconstruction of the point cloud. However, due to the dependence on the voxel generator which tends to produce a uniformly distributed grid, PCN shows significant limitations when applied to dense and non-uniformly distributed point clouds.
[0004] Therefore, the point cloud completion method in the prior art still has the technical problems of inaccurate point cloud data and poor reliability. SUMMARY
[0005] The present application provides a large-scale point cloud completion method and device to solve the technical problems of inaccurate point cloud data and poor reliability in the prior art.
[0006] The present application provides a large-scale point cloud completion method, comprising:
[0007] Obtaining incomplete large-scale point cloud data collected in an actual urban scene;
[0008] Segmenting the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data;
[0009] For each small-scale point cloud data, obtaining first point cloud data with a first resolution and second point cloud data with a second resolution through resampling;
[0010] Inputting the first point cloud data and the second point cloud data, and the RGB-D point cloud data corresponding to the small-scale point cloud data, into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network.
[0011] Fusing the RGB-D point cloud data and the complete point cloud data to obtain the completed point cloud data.
[0012] In some embodiments, the generative adversarial network comprises a generator; the generator comprises a multi-resolution point cloud encoder and a point cloud decoder;
[0013] The point cloud encoder comprises an input transformation layer, a shared multi-layer perception layer and a feature transformation layer;
[0014] The first point cloud data, the second point cloud data and the RGB-D point cloud data sequentially pass through the input transformation layer, the shared multi-layer perception layer and the feature transformation layer to generate a feature vector of a specific dimension;
[0015] The point cloud decoder comprises three different feature decoding levels of decoder;
[0016] The feature vector is input into the three different feature decoding levels of decoder, from coarse to fine, to predict different levels of details of the missing points, and obtain the complete point cloud data.
[0017] In some embodiments, the generative adversarial network comprises a discriminator;
[0018] The method further comprises:
[0019] Training the generator of the generative adversarial network using sample data;
[0020] Comparing the complete point cloud data with a true value through the discriminator to evaluate the accuracy of the generator.
[0021] In some embodiments, the discriminator of the generative adversarial network comprises a six-layer fully connected layer with a dimension of (256, 128, 64, 32, 16, 1) and an average pooling layer.
[0022] In some embodiments, the feature vector comprises local geometric information, global geometric information, information of different scales and XYZ coordinate information.
[0023] In some embodiments, the method further comprises:
[0024] Based on each small-scale point cloud data, using a perspective transformer to generate a corresponding RGB image and depth map from a bird's eye view perspective;
[0025] Fusing the RGB image and the depth map to obtain the RGB-D point cloud data corresponding to the small-scale point cloud data.
[0026] The application further provides a large-scale point cloud completion device, comprising:
[0027] An acquisition module is configured to acquire incomplete large-scale point cloud data collected in an actual urban scene.
[0028] A segmentation module is configured to perform segmentation processing on the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data.
[0029] A resampling module is configured to, for each small-scale point cloud data, obtain first point cloud data with a first resolution and second point cloud data with a second resolution through resampling.
[0030] A completion module is configured to input the first point cloud data and the second point cloud data and RGB-D point cloud data corresponding to the small-scale point cloud data into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network.
[0031] A fusion module is configured to fuse the RGB-D point cloud data and the complete point cloud data to obtain completed point cloud data.
[0032] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the large-scale point cloud completion method according to any one of the above when executing the computer program.
[0033] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the large-scale point cloud completion method according to any one of the above.
[0034] The application further provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the large-scale point cloud completion method according to any one of the above.
[0035] The large-scale point cloud completion method and device provided by the application can improve the understanding and adaptability of regular and irregular object structures by deeply extracting multi-scale point cloud features through a generative adversarial network, generating complete point cloud data, and fusing the generated three-dimensional point cloud data with high-resolution two-dimensional image data, thereby improving the accuracy and reliability of the completed point cloud data. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other accompanying drawings can also be obtained by those skilled in the art without creative effort on the basis of these accompanying drawings.
[0037] Figure 1 is a flowchart of the large-scale point cloud completion method provided by the present application.
[0038] Figure 2 is a schematic diagram of the overall framework of the large-scale point cloud completion network provided by the present application.
[0039] Figure 3 is a schematic diagram of the correspondence between the image and the point cloud in different coordinate systems provided by the present application.
[0040] Figure 4 is a schematic diagram of the fused RGB-D point cloud provided by the present application.
[0041] Figure 5 is a schematic diagram of the improved GAN structure provided by the present application.
[0042] Figure 6 is a schematic diagram of the visualized completion effect comparison provided by the present application.
[0043] Figure 7 is a schematic diagram of the large-scale point cloud completion device provided by the present application.
[0044] Figure 8 is a schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0045] Nowadays, high-quality communication environment is increasingly important in social life. Under ideal conditions, personal handheld devices (such as mobile phones) can effectively access network services through ground base station facilities. However, in high-density hotspot areas such as stadiums, limited communication resources are often insufficient to support simultaneous connection of a large number of user terminal devices.
[0046] With the continuous advancement of unmanned aerial vehicle (UAV) technology, it has become a reality to deploy aerial signal relay nodes using the unique hovering capability of UAVs, and this approach is gradually evolving into a promising air-ground integrated communication strategy. Due to the characteristics of UAVs such as mobility, flexibility and rapid deployment, UAVs are increasingly becoming an ideal supplement to enhance communication coverage, especially in complex geological conditions where it is inconvenient or impractical to establish traditional ground base stations. In order to rapidly deploy UAVs in urban environments, it is essential to conduct three-dimensional modeling of urban landscapes for environmental perception. Currently, the mainstream three-dimensional modeling method is to use laser radar (LiDAR) technology to perform laser scanning of urban building structures, terrain and vehicles, thereby obtaining point cloud data containing spatial XYZ information.
[0047] However, in real urban environments, laser scanning can generate point clouds consisting of millions or at least tens of thousands of points. However, existing technologies are generally designed to handle much smaller point clouds, often limited to approximately 2048 points. This difference requires adaptive adjustments to the real urban scene representation and point cloud completion network to address large-scale challenges.
[0048] In addition, in the actual three-dimensional modeling process, the limitations of sensor resolution and occlusion factors often result in partial loss of geometric features and semantic information. Although point cloud data contains detailed three-dimensional details, due to the presence of outliers and edge blurring issues, this significantly increases the difficulty of object recognition. Existing point cloud completion techniques typically rely on voxel-based generators to infer missing data points within the point cloud region, which tends to produce matrix-like sparse output point sets and has proven effective in handling regular-shaped objects. However, when faced with complex irregular structures such as high-voltage towers or bridges, the performance of voxel-based generators often significantly decreases in practical applications, resulting in technical problems such as inaccurate and unreliable completed point cloud data.
[0049] Based on the above technical problems, the present application provides a large-scale point cloud completion method and device, which deeply extracts multi-scale point cloud features through a generative adversarial network, generates complete point cloud data, and fuses the generated three-dimensional point cloud data with high-resolution two-dimensional image data, thereby improving the understanding and adaptability of regular and irregular object structures and improving the accuracy and reliability of the completed point cloud data.
[0050] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0051] Figure 1 is a flowchart of a large-scale point cloud completion method provided by the present application, as shown in Figure 1 , the method comprises the following:
[0052] Step 101, acquiring incomplete large-scale point cloud data collected in an actual urban scene.
[0053] Step 102, performing segmentation processing on the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data.
[0054] Specifically, in order to solve the problem that large-scale point cloud data cannot be processed in a real urban environment, after acquiring incomplete large-scale point cloud data collected in an actual urban scene, the present application performs segmentation processing on the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data.
[0055] Figure 2 is a schematic diagram of the overall framework of the large-scale point cloud completion network provided by the present application, as shown in Figure 2 , the present application proposes a new large-scale point cloud completion network (PIFC-Net) for point cloud and image fusion, which mainly includes a segmentation module, a completion module and a fusion module, aiming to complete the cropped point cloud and combine the point cloud with the image to generate a color point cloud in a large-scale real urban 3D modeling scene.
[0056] In some embodiments, the method further comprises:
[0057] Based on each small-scale point cloud data, an RGB image and a depth map corresponding to the bird's eye view are generated by using a perspective transformer;
[0058] The RGB image and the depth map are fused to obtain RGB-D point cloud data corresponding to the small-scale point cloud data.
[0059] Figure 3 is a schematic diagram of the correspondence relationship between the image and the point cloud in different coordinate systems, as shown in Figure 3 , taking the sun as an example, the world coordinate system can be represented as , the camera coordinate system can be represented as , the picture coordinate system can be represented as , and the depth map coordinate is defined as .
[0060] From the world coordinate system to the camera coordinate system, the transformation formula is as follows:
[0061]
[0062] The transformation formula from the camera coordinate system to the picture coordinate system is as follows:
[0063]
[0064] The transformation formula from the picture coordinate system to the pixel coordinate system is as follows:
[0065]
[0066] Therefore, the transformation formula from the world coordinate system to the pixel coordinate system is as follows:
[0067]
[0068] At this time, the picture coordinates are , the depth map coordinates are , and the expressions of and are as follows:
[0069]
[0070]
[0071] In the segmentation process, the camera uses the cropped colored object point cloud to generate an RGB picture and a depth map in the bird's eye view (BEV) perspective, and generates the corresponding camera extrinsic matrix and intrinsic matrix. The coordinates of the RGB picture can be upgraded to , and the coordinates of the depth map can be upgraded to . Therefore, the coordinate information and color information of the colored point cloud are fused.
[0072] Figure 4 is a schematic diagram of the fused RGB-D point cloud provided by the present application, as shown in Figure 4 , the RGB-D point cloud coordinates can be obtained from the following formula: , corresponding to the coordinates in the world coordinate system:
[0073]
[0074] Because the object is segmented from the real city point cloud, the XYZ coordinate range is very large, and the point cloud completion network can only accept normalized data. The present application normalizes the point cloud data by the following steps:
[0075] Calculate the point cloud centroid, and the formula is as follows:
[0076]
[0077] Calculate the point cloud offset, and the formula is as follows:
[0078]
[0079] The scaling factor is calculated as follows:
[0080]
[0081] The expression of the regularized coordinates is as follows:
[0082]
[0083] The coordinates can be inverse regularized using the following formula:
[0084]
[0085] Step 103, for each small-scale point cloud data, first point cloud data with a first resolution and second point cloud data with a second resolution are obtained by resampling.
[0086] Step 104, the first point cloud data and the second point cloud data, and the RGB-D point cloud data corresponding to the small-scale point cloud data, are input into the trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network.
[0087] Step 105, the RGB-D point cloud data and the complete point cloud data are fused to obtain the completed point cloud data.
[0088] In some embodiments, the generative adversarial network includes a generator; the generator includes a multi-resolution point cloud encoder and a point cloud decoder;
[0089] The point cloud encoder includes an input transformation layer, a shared multi-layer perception layer, and a feature transformation layer;
[0090] The first point cloud data, the second point cloud data, and the RGB-D point cloud data sequentially pass through the input transformation layer, the shared multi-layer perception layer, and the feature transformation layer to generate a feature vector of a specific dimension;
[0091] The point cloud decoder includes three different feature decoding levels of decoder;
[0092] The feature vector is input into the three different feature decoding levels of decoder, from coarse to fine, to predict different levels of details of the missing points and obtain complete point cloud data.
[0093] In some embodiments, the generative adversarial network includes a discriminator;
[0094] The method further includes:
[0095] training the generator of the generative adversarial network using sample data;
[0096] evaluating the accuracy of the generator by comparing the complete point cloud data with the true value through the discriminator.
[0097] In some embodiments, the generative adversarial network includes a discriminator comprising a six-layer fully connected layer with a dimension of (256, 128, 64, 32, 16, 1) and an average pooling layer.
[0098] In some embodiments, the feature vector contains local geometric information, global geometric information, information of different scales, and XYZ coordinate information.
[0099] Specifically, Figure 5 is the structure diagram of the improved generative adversarial network (GAN) provided by the present application, as Figure 5 As shown, the network architecture of the completion module uses GAN, which is composed of a point cloud completion generator and a completion result discriminator (discriminator). The generator contains a multi-resolution point cloud encoder and a decoder. The original point cloud (2048 points of XYZ data from an incomplete RGB-D point cloud) from the segmentation module and two resampled (downsampled) low-resolution point clouds (for example, 512 and 256 points, respectively) are input into the encoder. The encoder has three branches for processing point clouds of different resolutions. These point clouds are first converted into feature vectors through three-level downsampling. Then, they pass through an input transformation layer, a shared multi-layer perceptron (shared-MLP) layer, and a feature transformation layer to generate a feature vector F of a specific dimension. The output feature vector F contains local and global geometric information, information of different scales, and XYZ coordinate information.
[0100] The feature vector F is input into the decoder with three different feature decoding levels, which are then used to predict different levels of detail of the missing points. The decoding strategy from coarse to fine ensures that the node coordinates are decoded step by step from the coarsest level to the finest point. The generated points are combined with the original points and jointly trained in the discriminator to produce better completion results.
[0101] After the point cloud is generated, the result discriminator compares it with the true value to evaluate whether the error range is acceptable. The present application removes batch normalization to reduce the complexity of the discriminator, and expands the fully connected layer from four layers with a dimension of (256, 128, 16, 1) to six layers with a dimension of (256, 128, 64, 32, 16, 1), and changes the Max Pooling to Average Pooling. These improvements make the generated point cloud have better generation results, which are more accurate and reliable.
[0102] The purpose of the fusion module is to integrate the RGB information into the point cloud and arrange the standardized point cloud according to its original macro position. After coordinate normalization, different objects show obvious RGB differences, which helps subsequent target detection and target segmentation tasks.
[0103] Using the offset and the scaling factor , the coordinates of the object in the city after inverse normalization can be obtained as follows:
[0104]
[0105] From the RGB-D point cloud coordinates and the initial point cloud coordinates , the completed point cloud coordinates can be obtained.
[0106] The loss function of the PIFC-Net in the present application is composed of a completion loss and an adversarial loss. The expression of the completion loss is as follows:
[0107] +
[0108]
[0109] The completion loss is used to measure the error between the generated point cloud by the generator and the GT predicted completed point .
[0110] The expression of the adversarial loss is as follows:
[0111]
[0112] Where E represents the expected value, x is sampled from the distribution of the ground-truth point cloud Y gt , z is sampled from the original point cloud p1(z), and G(z) is a generator for converting z into a preliminary completed point cloud.
[0113] The expression of the total loss function is as follows:
[0114]
[0115] Where L represents the total loss function, w represents the weight, L c represents the completion loss, and L a represents the adversarial loss.
[0116] The large-scale point cloud completion method provided by the application extracts multi-scale point cloud features in depth through a generative adversarial network, generates complete point cloud data, and fuses the generated three-dimensional point cloud data with high-resolution two-dimensional image data, thereby improving the understanding and adaptability of regular and irregular object structures and improving the accuracy and reliability of the completed point cloud data.
[0117] The technical effects of the application are verified by experimental data.
[0118] UrbanBIS provides a high-precision color city model in an ideal case, and the application simulates a single scanning operation of a UAV to extract more than 9000 sets of RGB image, depth map and bird's eye view (BEV) point cloud data. The RGB information in the ideal point cloud is removed, FPS (farthest point sampling) down-sampling is performed to one-tenth to one percent of the original point number, and the point cloud is cropped to simulate the roughly colorless and missing area point cloud obtained by a single scanning operation of a UAV LiDAR in an emergency. All experiments are implemented on a NVIDIA GeForce RTX 3090 Ti GPU with 24GB of video memory using PyTorch 1.0.1.
[0119] The PIFC-Net in the application is divided into three sub-networks: a generator (Generator), a discriminator (Discriminator) and a LiDAR-RGB fusion module, which aims to explore the influence of different sub-network combinations on point cloud completion effect. The results are shown in Table 1. Compared with other modules, the point cloud completion network combined with the generator, the discriminator and the LiDAR-RGB fusion significantly improves the completion performance, which is mainly due to the use of color information. For large overhead objects, such as cubes and pyramid-shaped buildings, the performance of this complete network is better than that of the network without RGB fusion; but for complex geometries like cylindrical structures, its effect decreases. Analysis shows that this network performs well in completing objects with sufficient bird's eye view visibility, such as buildings and bridges. LiDAR-RGB fusion extracts RGB and depth data from the bird's eye view, enhances the bird's eye view representation through more detailed features, structural details and rich topological structures. By integrating RGB-D data with standard point cloud through a multi-layer feature extraction encoder, excellent point cloud completion is achieved in the bird's eye view, which is particularly important for UAV-based urban modeling tasks.
[0120] Table 1 Comparison of average performance of different models in the test set
[0121]
[0122] The experiment focuses on five types of buildings with different geometric shapes: cubic buildings, pyramid buildings, strip buildings, cylindrical buildings, and irregular buildings. The results of the PIFC-Net highlight better reconstruction effects for cubic and cylindrical buildings due to their predictable regularity. However, buildings based on pyramid and strip structures pose challenges to reconstruction due to their small top area, resulting in lower completion. Extending the scope of the study to different types of objects, the PIFC-Net performs best on building and bridge datasets, as simpler geometric shapes help improve reconstruction performance. In addition, the execution time is shown in Table 2, and compared with the most advanced (SOTA) Shape-Inversion method and SnowFlakeNet, the proposed method achieves better completion in a shorter time.
[0123] Table 2 Comparison of average performance of different models in the test set
[0124]
[0125] Figure 6 is a comparison diagram of the visualization completion effect provided by the application, as Figure 6 shown, the PIFC-Net proposed by the application has better point cloud completion effect, and the completed point cloud data is more accurate and reliable.
[0126] The point cloud completion device provided by the application is described below, and the point cloud completion device described below can be correspondingly referred to the point cloud completion method described above.
[0127] Figure 7 is a structure diagram of the large-scale point cloud completion device provided by the application, as Figure 7 shown, the application provides a large-scale point cloud completion device, which comprises:
[0128] The acquisition module 701 is used for acquiring incomplete large-scale point cloud data collected in an actual urban scene;
[0129] The segmentation module 702 is used for segmenting the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data;
[0130] The resampling module 703 is used for obtaining, for each small-scale point cloud data, first point cloud data with a first resolution and second point cloud data with a second resolution through resampling;
[0131] The completion module 704 is used for inputting the first point cloud data and the second point cloud data, and the RGB-D point cloud data corresponding to the small-scale point cloud data, into the trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network.
[0132] The fusion module 705 is configured to fuse the RGB-D point cloud data and the complete point cloud data to obtain the completed point cloud data.
[0133] Specifically, the above large-scale point cloud completion device provided by the embodiments of the present application can realize all the method steps realized by the above large-scale point cloud completion method embodiments, and can achieve the same technical effects. Here, the same parts and beneficial effects in the method embodiments will not be described in detail.
[0134] Figure 8 An example of a schematic diagram of the physical structure of an electronic device is shown in FIG. 8. Figure 8 As shown in FIG. 8, the electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 can communicate with each other through the communications bus 840. The processor 810 can invoke the logical instructions in the memory 830 to execute the large-scale point cloud completion method, which includes:
[0135] obtaining incomplete large-scale point cloud data collected in an actual urban scene;
[0136] segmenting the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data;
[0137] For each small-scale point cloud data, first point cloud data with a first resolution and second point cloud data with a second resolution are obtained through resampling;
[0138] inputting the first point cloud data and the second point cloud data, and RGB-D point cloud data corresponding to the small-scale point cloud data, into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network;
[0139] fusing the RGB-D point cloud data and the complete point cloud data to obtain completed point cloud data.
[0140] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as standalone products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or partially contribute to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0141] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the large-scale point cloud completion method provided by the above-mentioned methods, the method comprising:
[0142] obtaining incomplete large-scale point cloud data collected in an actual urban scene;
[0143] segmenting the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data;
[0144] For each small-scale point cloud data, first point cloud data with a first resolution and second point cloud data with a second resolution are obtained by resampling;
[0145] inputting the first point cloud data and the second point cloud data, and RGB-D point cloud data corresponding to the small-scale point cloud data, into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network;
[0146] fusing the RGB-D point cloud data and the complete point cloud data to obtain completed point cloud data.
[0147] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement a large-scale point cloud completion method provided by the above-mentioned methods, the method comprising:
[0148] obtaining incomplete large-scale point cloud data collected in an actual urban scene;
[0149] Segment the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data;
[0150] For each small-scale point cloud data, first point cloud data with a first resolution and second point cloud data with a second resolution are obtained through resampling;
[0151] The first point cloud data and the second point cloud data, and the RGB-D point cloud data corresponding to the small-scale point cloud data, are input into the trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network;
[0152] The RGB-D point cloud data and the complete point cloud data are fused to obtain the completed point cloud data.
[0153] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0154] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course, can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0155] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for large-scale point cloud completion, characterized in that, The method comprises the following steps: acquiring incomplete large-scale point cloud data collected in an actual urban scene; segmenting the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data; for each small-scale point cloud data, obtaining first point cloud data with a first resolution and second point cloud data with a second resolution through resampling; inputting the first point cloud data and the second point cloud data and RGB-D point cloud data corresponding to the small-scale point cloud data into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network; fusing the RGB-D point cloud data and the complete point cloud data to obtain completed point cloud data; the generative adversarial network comprises a generator; the generator comprises a multi-resolution point cloud encoder and a point cloud decoder; the point cloud encoder comprises an input transformation layer, a shared multi-layer perception layer and a feature transformation layer; the first point cloud data, the second point cloud data and the RGB-D point cloud data are sequentially subjected to the input transformation layer, the shared multi-layer perception layer and the feature transformation layer to generate a feature vector with a specific dimension; the point cloud decoder comprises three decoders of different feature decoding levels; the feature vector is input into the three decoders of different feature decoding levels to predict different levels of details of missing points from coarse to fine, so as to obtain complete point cloud data; the generative adversarial network comprises a discriminator; the method further comprises: training the generator of the generative adversarial network by using sample data; comparing the complete point cloud data with a true value by using the discriminator to evaluate the accuracy of the generator; the method further comprises: based on each small-scale point cloud data, generating a corresponding RGB image and depth map from a bird's eye view by using a view transformer; fusing the RGB image and the depth map to obtain RGB-D point cloud data corresponding to the small-scale point cloud data.
2. The method of claim 1, wherein, The discriminator of the generative adversarial network comprises a six-layer fully connected layer with a dimension of (256, 128, 64, 32, 16, 1) and an average pooling layer.
3. The method of claim 1, wherein, The feature vector contains local geometric information, global geometric information, information of different scales and XYZ coordinate information.
4. A large-scale point cloud completion device applying the large-scale point cloud completion method according to any one of claims 1 to 3, characterized in that, The method comprises the following steps: an acquisition module is configured to acquire incomplete large-scale point cloud data collected in an actual urban scene; a segmentation module is configured to segment the incomplete large-scale point cloud data to obtain a plurality of small-scale point cloud data; a resampling module is configured to, for each small-scale point cloud data, obtain first point cloud data with a first resolution and second point cloud data with a second resolution through resampling; a completion module is configured to input the first point cloud data and the second point cloud data and RGB-D point cloud data corresponding to the small-scale point cloud data into a trained generative adversarial network to obtain complete point cloud data generated by the generative adversarial network; a fusion module is configured to fuse the RGB-D point cloud data and the complete point cloud data to obtain completed point cloud data.
5. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the large-scale point cloud completion method as claimed in any one of claims 1 to 3 when executing the computer program.
6. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the large-scale point cloud completion method as claimed in any one of claims 1 to 3 when executed by the processor.
7. A computer program product comprising a computer program, characterized in that, The computer program implements the large-scale point cloud completion method as claimed in any one of claims 1 to 3 when executed by the processor.
Citation Information
Patent Citations
Target and track augmented reality method and system based on generative adversarial network
CN110941996A
Multi-scale greenhouse plant point cloud completion method based on generative adversarial network inverse mapping
CN115439490A