A Large-Scale Point Cloud Semantic Segmentation Method and System Based on Downscaling Projection
By using downscale projection technology in point cloud semantic segmentation, high-quality projection views are generated and semantic information is extracted, the semantic segmentation stability and accuracy problems in large-scale and large-scale scenarios in the existing technology are solved, and the high-precision and robust point cloud semantic segmentation effect is achieved.
Patent Information
- Application Number
- CN202411689933.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The existing point cloud semantic segmentation method has problems of insufficient general applicability and stability in large-scale and large-scale scenarios, and the projection-based method has insufficient semantic segmentation accuracy of low-quality texture two-dimensional image, resulting in some three-dimensional point clouds missing semantic information.
A large-scale point cloud semantic segmentation method based on downscale projection is adopted. By setting downscale projection view points and line of sight of virtual view, high-quality projected views are generated, sufficient semantic information is extracted, and two-dimensional semantic information is passed to three-dimensional, realizing semantic segmentation of large-scale large-scale point clouds.
This method reduces the quantization loss during the projection process, greatly reduces the information loss of three-dimensional points projected to the view, improves the fault tolerance and stability of the two-dimensional image semantic segmentation model, and realizes high accuracy and robust semantic segmentation of large-scale large-scale point clouds.
Smart Images

Figure CN119169299B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of semantic segmentation, and particularly relates to a large-scale point cloud semantic segmentation method and system based on downscaling projection. Background Art
[0002] With the popularization and performance improvement of three-dimensional perception technologies such as lidar (LiDAR) and depth cameras, it has become possible to efficiently and accurately capture and reconstruct point cloud data of complex three-dimensional scenes. However, the geometric information of only point cloud data far from meets the requirements of advanced applications. How to automatically understand and analyze the semantic information in these point cloud data, that is, to identify the categories, positions of various objects in the scene and even the relationships between them, has become a research hotspot. The point cloud semantic segmentation technology has emerged as the times require. It realizes the leap from the original three-dimensional data to advanced semantic understanding by assigning a semantic label (such as vehicle, pedestrian, tree, building, etc.) to each point in the point cloud, providing a solid foundation for subsequent decision-making, environmental perception and interaction.
[0003] The existing point cloud semantic segmentation methods are generally divided into three types. The first is the point-based semantic segmentation method, the second is the voxel-based semantic segmentation method, and the third is the projection-based semantic segmentation method. The first and second methods are affected by the relatively large point cloud density, and these two methods lack generality and stability in different scale scenes. For the projection-based semantic segmentation method, it first converts the three-dimensional data into multiple groups of two-dimensional data, and then uses the currently mature two-dimensional semantic segmentation algorithm to stably obtain semantic information. Therefore, this method is less affected by the point cloud density and the scene size, and is more suitable for point clouds in remote sensing scenes compared with the first two methods. However, due to the low accuracy of semantic segmentation of low-quality texture two-dimensional images, and the projection image cannot include all three-dimensional points, resulting in the lack of semantic information for some three-dimensional point clouds.
[0004] Therefore, it is urgent to solve the above problems. Summary of the Invention
[0005] Object of the Invention: The first object of the present invention is to provide a large-scale point cloud semantic segmentation method based on downscaling projection. This segmentation method can generate projection views containing as many points as possible and with as high quality as possible, extract sufficient semantic information, and can realize the semantic segmentation of large-scale and large-scale point clouds, improving the semantic segmentation accuracy.
[0006] The second object of the present invention is to provide a large-scale point cloud semantic segmentation system based on downscaling projection.
[0007] The third object of the present invention is to provide an electronic device.
[0008] The fourth object of the present invention is to provide a computer storage medium.
[0009] Technical solution: To achieve the above objectives, the present invention discloses a large-scale point cloud semantic segmentation method based on downscaled projection, including the following steps:
[0010] (1) Obtain the point cloud to be processed, set the downscaled projection viewpoints and lines of sight of the virtual views according to the spatial distribution of the large-scale point cloud, and obtain the external parameters of the virtual camera;
[0011] (2) Generate the corresponding virtual view for the internal and external parameters of the virtual camera corresponding to the m-th viewpoint, and at the same time obtain the mapping relationship between the three-dimensional point cloud and the two-dimensional projection points, where the internal parameters of the virtual camera include focal length, pixel size, and image frame size;
[0012] (3) Extract the semantic information in the virtual view based on the existing two-dimensional image semantic segmentation algorithm, and transfer the semantic information from two dimensions to three dimensions based on the obtained mapping relationship between the three-dimensional point cloud and the two-dimensional projection points to complete the semantic segmentation of the point cloud.
[0013] Optionally, step (1) specifically includes the following steps:
[0014] (1.1) Read the three-dimensional model file to obtain the point cloud to be processed , and set up a multi-layer viewpoint network for generating high-quality multi-view projections of the point cloud ;
[0015] (1.2) Obtain the minimum oriented bounding box of the point cloud corresponding to the k-th layer viewpoint network to be processed currently , determine the vertical direction away from the ground , and two directions along the ground and and , a total of three directions , and the initial value of k is 0;
[0016] (1.3) Determine the image resolution of the virtual view of the current layer based on the point cloud density , ,
[0017] ,
[0018] where is a constant,
[0019] According to the geometric relationship between the ground height , the image resolution , the focal length and the pixel size , determine the ground height of the view projection,
[0020] ,
[0021] According to the image resolution obtain the ground coverage range of the image ,
[0022] ,
[0023] wherein is the number of pixels on the side of the square virtual view Figure 1 ;
[0024] (1.4) Based on the ground height , ground coverage range , three directions and the overlap degree , set the viewpoint network of the k-th layer , the directions of all viewpoints in the viewpoint network are the vertical directions towards the ground , the two directions of the view are the two ground directions of the point cloud to be processed currently, and the rotation matrix is ; the viewpoints in each layer of the viewpoint network are evenly distributed, and the viewpoint coordinates of the -th row and -th column are:
[0025] ,
[0026] The value range of the row index is:
[0027] ,
[0028] The value range of the column index is:
[0029] ,
[0030] The reference viewpoint coordinates are:
[0031] ,
[0032] wherein is the ceiling function, are the eight vertices of the minimum oriented bounding box, is the vertex with index 0, represents the size of the minimum oriented bounding box, is the length of the minimum oriented bounding box, is the width of the minimum oriented bounding box; thus, the external parameters of the virtual camera corresponding to the viewpoint in the i-th row and j-th column of the k-th layer viewpoint network are obtained ;
[0033] (1.5) Based on processing the ground coverage range , ground height , three directions Generate the minimum oriented bounding box for all the point clouds within the field of view of this viewpoint, and extract the point clouds within the box. ; If there are point clouds within the box and the point cloud density and the point cloud density of the point clouds targeted by the previous layer of viewpoints and the point cloud density of the entire point cloud meet the threshold condition for the ratio between them,
[0034] ,
[0035] then the point clouds within the box serve as the point clouds to be targeted by the next layer of viewpoint network , repeat the steps (1.2) to (1.5), otherwise terminate the iteration, and finally generate a multi-layer viewpoint network through iteration. ;
[0036] (1.6) Combine the extrinsic parameters of the virtual cameras corresponding to each layer of the viewpoint network in the multi-layer viewpoint network to form a set of virtual camera extrinsic parameters. Among them, the extrinsic parameters of the virtual camera corresponding to the m-th viewpoint in the set of virtual camera extrinsic parameters are .
[0037] Optionally, step (2) specifically includes the following steps:
[0038] (2.1) Based on the intrinsic and extrinsic parameters of the virtual camera corresponding to the m-th viewpoint, transform the coordinate system of the entire point cloud from the world coordinate system to the camera coordinate system , and retain the points with positive Z values in the coordinate system, that is, retain the point clouds in front of the virtual camera and discard other point clouds;
[0039] (2.2) Convert from the camera coordinate system to the pixel coordinate system , and retain the projected points within the image frame;
[0040] (2.3) Based on the Z values, i.e., depths, of the projected points within the image frame corresponding to the camera coordinate system , retain the projected points with the minimum depth within the same pixel block and discard other points;
[0041] (2.4) Assign the depth or grayscale of the selected 3D point cloud to the corresponding 2D projected points to obtain a one-to-one mapping relationship between the 3D point cloud and the 2D projected points , and achieve the initial rendering of the virtual view ;
[0042] (2.5) Based on the initial rendering , to address the problems of holes and missing textures in the projected image, obtain the minimum axis-aligned bounding box containing all the projected points in the image, and search for holes in the projected view within the bounding box;
[0043] (2.6) Extract the neighborhood of each hole, and based on the gray values of the non-hole pixels in the neighborhood, calculate the gray value of the hole using the method of distance Gaussian weighting , thereby filling the holes and enriching the textures, implementing post-processing of the rendering of the virtual view, and generating the virtual view corresponding to the m-th viewpoint ;
[0044] ,
[0045] where are the pixel coordinates, is 's neighborhood, indicates that there is a value at in the initialized virtual view , represents the pixel 's distance on the virtual view.
[0046] Optionally, step (3) specifically includes the following steps:
[0047] (3.1) Process the point cloud with semantic information for training according to steps (1) and (2), obtain the internal and external parameters of the virtual camera, generate the virtual view corresponding to the point cloud with semantic information, assign the semantic encoding of the three-dimensional point cloud corresponding to the two-dimensional projected points in the virtual view to the corresponding two-dimensional projected points, and complete the first step of generating the semantic label map; to address the problem of holes in the semantic label map, obtain the minimum axis-aligned bounding box containing all the projected points in the semantic label map, search for image holes within the bounding box, extract the neighborhood of each hole, and statistically analyze the semantic characteristics within the neighborhood of each hole based on the semantic information of the semantic label map. The semantic characteristic corresponding to the largest number of points with a certain semantic characteristic in the neighborhood is the semantic characteristic of this point, complete the second step of generating the semantic label map, and achieve the generation of the final semantic label map; based on the existing two-dimensional image semantic segmentation algorithm, use the virtual view and the corresponding semantic label map to train and generate a two-dimensional image semantic segmentation model;
[0048] (3.2) Input the virtual view of step (2) into the two-dimensional image semantic segmentation model, extract the semantic label map , and obtain the semantic information in the virtual view;
[0049] (3.3) Establish an initial voting table, initialize it to zero, and based on the mapping relationship between the three-dimensional point cloud in each view and the two-dimensional projected points and the semantic information extracted from each view , count the points that receive votes, and set the semantic feature with the highest number of votes for each point as the semantic feature of that point to complete the first vote;
[0050] (3.4) Establish a new voting table. For each point that does not receive a vote in the initial voting table, extract its three-dimensional spatial neighborhood and count the points with determined semantic features of various types in the neighborhood. The semantic feature corresponding to the largest number of points with a certain semantic feature in the neighborhood is the semantic feature of that point to complete the second vote;
[0051] (3.5) Combine the information contained in the two voting tables, which is the final voting result, to complete the transmission of semantic information based on the semantic labels of the virtual views.
[0052] Based on the same inventive concept, the present invention discloses a large-scale point cloud semantic segmentation system based on downscaling projection, including:
[0053] An external parameter acquisition module for acquiring the point cloud to be processed, setting the downscaling projection viewpoints and lines of sight of the virtual views according to the spatial distribution of the large-scale point cloud, and obtaining the external parameters of the virtual camera;
[0054] A mapping relationship acquisition module for generating the corresponding virtual view for the internal and external parameters of the virtual camera corresponding to the m-th viewpoint, and simultaneously obtaining the mapping relationship between the three-dimensional point cloud and the two-dimensional projection points, where the internal parameters of the virtual camera include the focal length, pixel size, and image frame size;
[0055] A semantic segmentation module for extracting the semantic information in the virtual view based on the existing two-dimensional image semantic segmentation algorithm, and transmitting the semantic information from two dimensions to three dimensions based on the obtained mapping relationship between the three-dimensional point cloud and the two-dimensional projection points to complete the semantic segmentation of the point cloud.
[0056] Optionally, the external parameter acquisition module reads the three-dimensional model file to acquire the point cloud to be processed ; where the three-dimensional model file is obtained by lidar scanning; to generate a high-quality multi-view projection of the point cloud, a multi-layer viewpoint network is set ;
[0057] Acquire the minimum oriented bounding box of the point cloud corresponding to the k-th layer viewpoint network to be processed currently to determine the vertical direction away from the ground and two directions along the ground and and , a total of three directions , and the initial value of k is 0;
[0058] Based on the point cloud density Determine the image resolution of the current virtual view layer ,
[0059] ,
[0060] where is a constant,
[0061] According to the geometric relationship between the ground height , image resolution , focal length and pixel size , determine the ground height of the view projection,
[0062] ,
[0063] According to the image resolution obtain the ground coverage range of the image,
[0064] ,
[0065] where is the number of pixels on the side of the square virtual view Figure 1 ;
[0066] Based on the ground height , ground coverage range , three directions and overlap degree of the k-th layer virtual view, set the view point network of the k-th layer. The directions of all view points in the view point network are the vertical directions towards the ground . The two directions of the view are the two ground directions of the point cloud to be processed currently. The rotation matrix is ; The view points in each layer of the view point network are evenly distributed. The view point coordinates of the th row and the th column are:
[0067] ,
[0068] The value range of the row index is:
[0069] ,
[0070] The value range of the column index is:
[0071] ,
[0072] The reference view point coordinates are:
[0073] ,
[0074] where is the ceiling function, are the eight vertices of the minimum oriented bounding box, is the vertex with index 0, represents the size of the minimum oriented bounding box, is the length of the minimum oriented bounding box, is the width of the minimum oriented bounding box; thus, the extrinsic parameters of the virtual camera corresponding to the view point at the \(i\)-th row and \(j\)-th column of the \(k\)-th layer view point network are obtained. ;
[0075] Based on processing the ground coverage range , ground height , and three directions , the minimum oriented bounding box of all point clouds within the field of view of this view point is generated, and the point clouds within the box are extracted ; if there are point clouds within the box, and the ratio of the point cloud density to the point cloud density of the point cloud targeted by the view point in the previous layer and the point cloud density of the entire point cloud meets the threshold condition,
[0076] ,
[0077] then the point clouds within the box are used as the point clouds to be targeted by the next layer view point network , and continue the iteration, otherwise terminate the iteration, and finally generate a multi-layer view point network ;
[0078] The extrinsic parameters of the virtual cameras corresponding to each layer of the multi-layer view point network are aggregated to form a set of extrinsic parameters of the virtual camera. Among them, the extrinsic parameters of the virtual camera corresponding to the \(m\)-th view point in the set of extrinsic parameters of the virtual camera are .
[0079] Optionally, in the mapping relationship acquisition module, based on the intrinsic and extrinsic parameters of the virtual camera corresponding to the \(m\)-th view point, the coordinate system of the entire point cloud is transformed from the world coordinate system to the camera coordinate system , and the points with positive \(Z\) values in the coordinate system are retained, that is, the point clouds in front of the virtual camera are retained, and other point clouds are discarded;
[0080] It is transformed from the camera coordinate system to the pixel coordinate system , and the projected points within the image frame are retained;
[0081] Based on the projected points within the image frame corresponding to the camera coordinate system The lower Z value, i.e., depth, retains the projection point with the minimum depth within the same pixel block and discards other points;
[0082] Assign the depth or grayscale of the selected 3D point cloud to the corresponding 2D projection point to obtain a one-to-one mapping relationship between the 3D point cloud and the 2D projection point , and realize the initial rendering of the virtual view ;
[0083] Based on the initial rendering , for the problems of holes and missing textures in the projection image, obtain the minimum axis-aligned bounding box where all projection points in the image are located, and search for holes in the projection view within the bounding box;
[0084] Extract the neighborhood of each hole, and based on the grayscale values of non-hole pixels in the neighborhood, calculate the grayscale value of the hole using the method of distance Gaussian weighting , thereby filling the holes and enriching the textures, realizing the post-processing of the virtual view rendering, and generating the virtual view corresponding to the m-th viewpoint ;
[0085] ,
[0086] where is the pixel point coordinate, is the neighborhood of indicating that there is a value at in the initialized virtual view , indicating the distance of the pixel point on the virtual view.
[0087] Optionally, the semantic segmentation module processes the point cloud with semantic information for training, obtains the internal and external parameters of the virtual camera, generates the virtual view corresponding to the point cloud with semantic information, assigns the semantic encoding of the 3D point corresponding to the 2D projection point included in the virtual view to the corresponding 2D projection point, and completes the first step of generating the semantic label map; for the problem of holes in the semantic label map, obtain the minimum axis-aligned bounding box where all projection points in the semantic label map are located, search for image holes within the bounding box, extract the neighborhood of each hole, and statistically analyze the semantic characteristics within the neighborhood of each hole based on the semantic information of the semantic label map. The semantic characteristic corresponding to the largest number of points with a certain semantic characteristic in the neighborhood is the semantic characteristic of this point, and complete the second step of generating the semantic label map to realize the generation of the final semantic label map; based on the existing 2D image semantic segmentation algorithm, use the virtual view and the corresponding semantic label map to train and generate a 2D image semantic segmentation model;
[0088] Input the virtual view into the 2D image semantic segmentation model to extract the semantic label map , obtain semantic information in the virtual view;
[0089] Create an initial voting table, initialized to zero, and based on the 3D point cloud within each view and the 2D projection points between the mapping relationship and the semantic information extracted from each view , count the points that receive votes, and the semantic feature with the highest number of votes for each point is set as the semantic feature of that point, completing the first vote;
[0090] Create a new voting table. For each point that does not receive votes in the initial voting table, extract its 3D spatial neighborhood and count the points with determined semantic features of various types in the neighborhood. The semantic feature corresponding to the largest number of points with a certain semantic feature in the neighborhood is the semantic feature of that point, completing the second vote;
[0091] Merge the information contained in the two voting tables, which is the final voting result, and complete the semantic information transmission based on the semantic labels of the virtual view.
[0092] The present invention discloses an electronic device, including a processor and a storage medium; the storage medium is used to store instructions;
[0093] The processor is used to operate according to the instructions to execute the steps of a large-scale point cloud semantic segmentation method based on downscaling projection as described above.
[0094] The present invention discloses a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to perform the steps of a large-scale point cloud semantic segmentation method based on downscaling projection as described above.
[0095] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The point cloud semantic segmentation method of the present invention provides a strategy to reduce quantization loss in the projection process, greatly reducing the information loss in the process of projecting 3D points onto the view. By generating high-quality projection views and extracting sufficient semantic information, the fault tolerance of the 2D image semantic segmentation model is improved, and the 2D semantic information is stably transmitted to 3D, realizing the semantic segmentation of large-scale and large-scale point clouds, and improving the accuracy and robustness of point cloud semantic segmentation. Brief Description of the Drawings
[0096] Figure 1 is a flow chart of the present invention;
[0097] Figure 2 is a flow chart of a scheme for setting projection viewpoints and line-of-sight directions in the present invention;
[0098] Figure 3Schematic diagram for storing and indexing the vertices of the minimum oriented bounding box in the present invention, where the serial numbers 0, 1, 2, 3, 4, 5, 6, 7 represent the index numbers of the vertices of the minimum oriented bounding box in the array;
[0099] Figure 4 Flowchart for generating a two-dimensional image semantic segmentation model in the present invention;
[0100] Figure 5 System schematic diagram of the present invention. Detailed implementation manners
[0101] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0102] Embodiment 1: This embodiment recommends using a computer with the following configuration: Intel i5 12400F processor, NVIDIA GeForce RTX 4060 graphics processor, main frequency 3.99 GHz, 32 GB of memory, and the operating system is windows11. The implementation of the point cloud viewpoint setting is based on the Open3D three-dimensional data processing framework toolkit, the implementation of the point cloud projection image semantic segmentation network is based on the PyTorch deep learning framework toolkit, and the inference of the projection image two-dimensional semantic segmentation is based on the ONNX and ONNX Runtime model inference framework toolkits.
[0103] As Figure 1 shown, this embodiment discloses a large-scale point cloud semantic segmentation method based on downscaled projection, including the following steps:
[0104] (1) As Figure 2 shown, obtain the point cloud to be processed, set the downscaled projection viewpoints and lines of sight of the virtual view according to the spatial distribution of the large-scale point cloud, and obtain the external parameters of the virtual camera;
[0105] Step (1) specifically includes the following steps:
[0106] (1.1) Read the three-dimensional model file. Taking the WHU3D data as an example, with the help of pywhu3d, read the h5 data of the als type to obtain the point cloud to be processed ; where the three-dimensional model file is obtained by lidar scanning; to generate high-quality multi-view projections of the point cloud, set a multi-layer viewpoint network ;
[0107] (1.2) As Figure 3 shown, obtain the minimum oriented bounding box of the point cloud corresponding to the k-th layer viewpoint network to be processed currently ; as shown, determine the vertical direction away from the ground Figure 1 ; and two directions along the ground and for a total of three directions The initial value of k is 0;
[0108] (1.3) Based on the point cloud density Determine the image resolution of the current virtual view layer ,
[0109] ,
[0110] where is a constant,
[0111] According to the geometric relationship between the ground height , the image resolution , the focal length and the pixel size , determine the ground height of the view projection ,
[0112] ,
[0113] According to the image resolution obtain the ground coverage range of the image ,
[0114] ,
[0115] where is the number of pixels on the side of the square virtual view Figure 1 ;
[0116] (1.4) Based on the ground height of the k-th layer virtual view, the ground coverage range , three directions and the overlap degree , set the viewpoint network of the k-th layer. The directions of all viewpoints in the viewpoint network are the vertical directions towards the ground . The two directions of the view are the two ground directions of the point cloud to be processed currently, and the rotation matrix is ; The viewpoints in each layer of the viewpoint network are evenly distributed. The viewpoint coordinates of the th row and the th column are:
[0117] ,
[0118] The value range of the row index is:
[0119] ,
[0120] The value range of the column index is:
[0121] ,
[0122] The reference viewpoint coordinates are:
[0123] ,
[0124] where is the ceiling function, are the eight vertices of the minimum oriented bounding box, is the vertex with index 0, represents the size of the minimum oriented bounding box, is the length of the minimum oriented bounding box, is the width of the minimum oriented bounding box; thus, the extrinsic parameters of the virtual camera corresponding to the viewpoint at the \(i\)-th row and \(j\)-th column of the \(k\)-th layer viewpoint network are obtained ;
[0125] (1.5) Based on processing the ground coverage range , ground height , and three directions corresponding to each viewpoint, generate the minimum oriented bounding box of all point clouds within the field of view of this viewpoint, and extract the point clouds within the box ; if there are point clouds within the box, and the ratio between the point cloud density of the point clouds within the box and the point cloud density of the point clouds targeted by the previous layer viewpoint and the point cloud density
[0126] of the entire point cloud
[0127] meets the threshold condition, then the point clouds within the box are used as the point clouds to be targeted by the next layer viewpoint network, and repeat the steps of (1.2) to (1.5), otherwise terminate the iteration. Finally, generate a multi-layer viewpoint network . During the whole process, the ground height of the viewpoint network becomes smaller and smaller, the distance from the ground point clouds gets closer and closer, and the scale decreases, i.e., downscaling projection;
[0128] (1.6) Combine the extrinsic parameters of the virtual cameras corresponding to each layer of the multi-layer viewpoint network to form a set of extrinsic parameters of the virtual cameras. Among them, the extrinsic parameters of the virtual camera corresponding to the \(m\)-th viewpoint in the set of extrinsic parameters of the virtual cameras are .
[0129] (2) Generate the corresponding virtual view for the intrinsic and extrinsic parameters of the virtual camera corresponding to the \(m\)-th viewpoint, and at the same time obtain the mapping relationship between the three-dimensional point cloud and the two-dimensional projection points , where the internal parameters of the virtual camera include focal length, pixel size, and image size;
[0130] Step (2) specifically includes the following steps:
[0131] (2.1) Based on the internal and external parameters of the virtual camera corresponding to the m-th viewpoint, transform the coordinate system of the entire point cloud from the world coordinate system to the camera coordinate system , and retain the points with positive Z values in the coordinate system, that is, retain the point cloud in front of the virtual camera and discard other point clouds;
[0132] (2.2) Convert from the camera coordinate system to the pixel coordinate system , and retain the projection points within the image frame;
[0133] (2.3) Based on the Z value, i.e., depth, corresponding to the projection points within the image frame in the camera coordinate system , retain the projection point with the minimum depth within the same pixel and discard other points;
[0134] (2.4) Assign the depth or grayscale of the selected three-dimensional point cloud to the corresponding two-dimensional projection points to obtain a one-to-one mapping relationship between the three-dimensional point cloud and the two-dimensional projection points to achieve the initial rendering of the virtual view ;
[0135] (2.5) Based on the initial rendering , for the problems of holes and missing textures in the projection image, obtain the minimum axis-aligned bounding box where all projection points in the image are located, and search for the holes in the projection view within the bounding box;
[0136] (2.6) Extract the neighborhood of each hole, and based on the grayscale values of the non-hole points in the neighborhood, calculate the grayscale value of the hole using the method of distance Gaussian weighting to fill the holes and enrich the textures, achieve the post-processing of the virtual view rendering, and generate the virtual view corresponding to the m-th viewpoint ;
[0137] ,
[0138] where is the pixel point coordinate, is 's neighborhood, represents having a value at in the initialized virtual view , represents the pixel point 's distance on the virtual view.
[0139] (3) Extract semantic information in the virtual view based on the existing two-dimensional image semantic segmentation algorithm , based on the mapping relationship between the obtained three-dimensional point cloud and the two-dimensional projection points , transfer the semantic information from two dimensions to three dimensions to complete the semantic segmentation of the point cloud.
[0140] Step (3) specifically includes the following steps:
[0141] (3.1) Process the point cloud with semantic information for training according to steps (1) and (2) to obtain the internal and external parameters of the virtual camera, generate the virtual view corresponding to the point cloud with semantic information, assign the semantic encoding of the three-dimensional point cloud corresponding to the two-dimensional projection points contained in the virtual view to the corresponding two-dimensional projection points to complete the first step of generating the semantic label map; for the problem of holes existing in the semantic label map, obtain the smallest axis-aligned bounding box where all projection points in the semantic label map are located, search for image holes within the bounding box, extract the neighborhood of each hole, and statistically analyze the semantic characteristics within the neighborhood of each hole based on the semantic information of the semantic label map. The semantic characteristic corresponding to the largest number of points with a certain semantic characteristic in the neighborhood is the semantic characteristic of this point, completing the second step of generating the semantic label map and realizing the generation of the final semantic label map; based on the existing two-dimensional image semantic segmentation algorithm, use the virtual view and the corresponding semantic label map to train and generate a two-dimensional image semantic segmentation model. The two-dimensional image semantic segmentation model can be in the onnx format, which has higher efficiency and less space occupancy during model inference;
[0142] (3.2) Input the virtual view in step (2) into the two-dimensional image semantic segmentation model to extract the semantic label map and obtain the semantic information in the virtual view;
[0143] (3.3) Establish an initial voting table, initialize it to zero, and based on the mapping relationship between the three-dimensional point cloud and the two-dimensional projection points in each view and the semantic information extracted from each view , count the points that receive votes, and set the semantic characteristic with the highest number of votes for each point as the semantic characteristic of this point to complete the first vote;
[0144] (3.4) Establish a new voting table. For each point that does not receive votes in the initial voting table, extract its three-dimensional spatial neighborhood and count the points with determined semantic characteristics of various types in the neighborhood in the initial voting result. The semantic characteristic corresponding to the largest number of points with a certain semantic characteristic in the neighborhood is the semantic characteristic of this point to complete the second vote;
[0145] (3.5) Combine the information contained in the two voting tables, which is the final voting result, complete the semantic information transmission based on the virtual view semantic tags, and achieve the improvement of the semantic segmentation accuracy of the large-scale scene point cloud.
[0146] Example 2: As Figure 5 shown, this embodiment discloses a large-scale point cloud semantic segmentation system based on downscaling projection, including:
[0147] An external parameter acquisition module, configured to acquire the point cloud to be processed, set the downscaling projection viewpoints and lines of sight of the virtual view according to the spatial distribution of the large-scale point cloud, and obtain the external parameters of the virtual camera;
[0148] Read the 3D model file in the external parameter acquisition module. Taking WHU3D data as an example, read the h5 data of the als type with the help of pywhu3d to obtain the point cloud to be processed ; where the 3D model file is obtained by lidar scanning; to generate a high-quality multi-view projection of the point cloud, set up a multi-layer viewpoint network ;
[0149] As Figure 3 shown, obtain the minimum oriented bounding box of the point cloud corresponding to the k-th layer viewpoint network to be processed currently of , determine the vertical direction away from the ground and the two directions along the ground and , a total of three directions , and the initial value of k is 0;
[0150] Based on the point cloud density determine the image resolution of the current layer of virtual view ,
[0151] ,
[0152] where is a constant,
[0153] According to the geometric relationship between the ground height , the image resolution , the focal length and the pixel size , determine the ground height of the view projection,
[0154] ,
[0155] According to the image resolution obtain the ground coverage range of the image,
[0156] ,
[0157] where is the number of pixels on the side of the square virtual view Figure 1 ;
[0158] Based on the ground height of the k-th layer virtual view , ground coverage , three directions and overlap degree , set the viewpoint network of the k-th layer , the directions of all viewpoints in the viewpoint network are the vertical directions towards the ground , two directions of the view are the two ground directions of the point cloud to be processed currently, and the rotation matrix is ; The viewpoints in each layer of the viewpoint network are evenly distributed. The viewpoint coordinates of the -th row and -th column are:
[0159] ,
[0160] The value range of the row index is:
[0161] ,
[0162] The value range of the column index is:
[0163] ,
[0164] The reference viewpoint coordinates are:
[0165] ,
[0166] where is the ceiling function, are the eight vertices of the minimum oriented bounding box, is the vertex with index 0, represents the size of the minimum oriented bounding box, is the length of the minimum oriented bounding box, is the width of the minimum oriented bounding box; Thus, the external parameters of the virtual camera corresponding to the viewpoint in the i-th row and j-th column of the k-th layer viewpoint network are obtained ;
[0167] Based on processing the ground coverage , ground height , three directions corresponding to each viewpoint, generate the minimum oriented bounding box of all point clouds within the field of view of this viewpoint, and extract the point clouds within the box ; If there are point clouds within the box and the point cloud density The point cloud density of the point cloud targeted by the upper-layer view point and the point cloud density of the entire point cloud The ratio between them meets the threshold condition,
[0168] ,
[0169] then the point cloud within the frame is used as the point cloud to be targeted by the next-layer view point network and continues to iterate, otherwise the iteration terminates. Finally, a multi-layer view point network is iteratively generated During the whole process, the ground height of the view point network becomes smaller and smaller, the distance from the ground point cloud gets closer and closer, and the scale decreases, that is, downscaling projection;
[0170] The external parameters of the virtual cameras corresponding to each layer of the multi-layer view point network are combined to form a set of external parameters of the virtual cameras. Among them, the external parameters of the virtual camera corresponding to the m-th view point are .
[0171] The mapping relationship acquisition module is used to generate the corresponding virtual view for the internal and external parameters of the virtual camera corresponding to the m-th view point , and at the same time obtain the mapping relationship between the three-dimensional point cloud and the two-dimensional projection points , where the internal parameters of the virtual camera include focal length, pixel size, and image frame size;
[0172] In the mapping relationship acquisition module, based on the internal and external parameters of the virtual camera corresponding to the m-th view point, the coordinate system of the entire point cloud is transformed from the world coordinate system to the camera coordinate system , and the points with positive Z values in the coordinate system are retained, that is, the point cloud in front of the virtual camera is retained, and other point clouds are discarded;
[0173] From the camera coordinate system it is converted to the pixel coordinate system , and the projection points within the image frame are retained;
[0174] Based on the Z value, that is, the depth, of the projection points within the image frame corresponding to the camera coordinate system , the projection point with the minimum depth within the same pixel is retained, and other points are discarded;
[0175] The depth or grayscale of the selected three-dimensional point cloud is assigned to the corresponding two-dimensional projection points to obtain the one-to-one mapping relationship between the three-dimensional point cloud and the two-dimensional projection points , realizing the initial rendering of the virtual view ;
[0176] Based on the initial rendering For the problems of holes and missing textures in the projected image, obtain the minimum axis-aligned bounding box containing all the projection points in the image, and search for the holes in the projected view within the bounding box;
[0177] Extract the neighborhood of each hole, and based on the gray values of the non-hole areas in the neighborhood, calculate the gray value of the hole using the method of distance Gaussian weighting Thereby fill the holes and enrich the textures, implement the post-processing of the rendering of the virtual view, and generate the virtual view corresponding to the m-th viewpoint ;
[0178] ,
[0179] where is the pixel point coordinate, is the neighborhood of indicating that there is a value at in the initialized virtual view , indicating the distance of the pixel point on the virtual view.
[0180] The semantic segmentation module is used to extract the semantic information in the virtual view based on the existing two-dimensional image semantic segmentation algorithm , based on the mapping relationship between the obtained three-dimensional point cloud and the two-dimensional projection points , transfer the semantic information from two dimensions to three dimensions to complete the semantic segmentation of the point cloud.
[0181] Process the point cloud with semantic information for training in the semantic segmentation module, obtain the internal and external parameters of the virtual camera, generate the virtual view corresponding to the point cloud with semantic information, assign the semantic encoding of the three-dimensional point cloud corresponding to the two-dimensional projection point contained in the virtual view to the corresponding two-dimensional projection point, and complete the first step of generating the semantic label map; for the problem of holes in the semantic label map, obtain the minimum axis-aligned bounding box containing all the projection points in the semantic label map, search for the image holes within the bounding box, extract the neighborhood of each hole, and based on the semantic information of the semantic label map, statistically analyze the semantic characteristics in the neighborhood of each hole. The semantic characteristic corresponding to the largest number of points with a certain semantic characteristic in the neighborhood is the semantic characteristic of this point, complete the second step of generating the semantic label map, and realize the generation of the final semantic label map; based on the existing two-dimensional image semantic segmentation algorithm, use the virtual view and the corresponding semantic label map to train and generate a two-dimensional image semantic segmentation model;
[0182] Input the virtual view into the two-dimensional image semantic segmentation model, extract the semantic label map , and obtain the semantic information in the virtual view;
[0183] An initial voting table is established and initialized to zero, based on the three-dimensional point cloud within each view and the two-dimensional projected points between the mapping relationship and the semantic information extracted from each view , count the points that receive votes, and the semantic feature with the highest number of votes for each point is set as the semantic feature of that point, completing the first vote;
[0184] A new voting table is established. For each point that does not receive votes in the initial voting table, its three-dimensional spatial neighborhood is extracted and the points with determined semantic features of various types in the initial voting results in the neighborhood are counted. The semantic feature corresponding to the largest number of points with a certain semantic feature in the neighborhood is the semantic feature of that point, completing the second vote;
[0185] Merge the information contained in the two voting tables, which is the final voting result, complete the semantic information transmission based on the semantic labels of virtual views, and achieve the improvement of the semantic segmentation accuracy of large-scale scene point clouds.
[0186] Embodiment 3: Another embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a large-scale point cloud semantic segmentation method based on downscaling projection as described above.
[0187] The electronic device may include: a processor, a memory, a bus, and a communication interface. The processor, the communication interface, and the memory are connected through the bus; a computer program executable on the processor is stored in the memory. When the processor runs the computer program, it executes a large-scale point cloud semantic segmentation method based on downscaling projection provided by any one of the foregoing embodiments of the present invention.
[0188] Among them, the memory may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface (which can be wired or wireless), the communication connection between the device network element and at least one other network element is realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0189] The bus may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory is used to store the program. After the processor receives the execution instruction, it executes the program. A large-scale point cloud semantic segmentation method disclosed in any one of the foregoing embodiments of the present invention can be applied to the processor or implemented by the processor.
[0190] A processor may be an integrated circuit chip with the ability to process signals. In the implementation process, the above
[0191] Each step of the method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The above-mentioned processor can be a general-purpose processor, which may include a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0192] The electronic device provided by the embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0193] Embodiment 4: Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the method of any of the above embodiments. The computer-readable storage medium may be an optical disc, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method provided by any of the foregoing embodiments.
[0194] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, Phase Change Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.
[0195] The computer-readable storage medium provided by the above embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application programs stored therein.
[0196] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields shall be included within the patent protection scope of the present invention.
Claims
1. A large-scale point cloud semantic segmentation method based on downscaling projection, characterized in that: The following steps are involved: (1) Obtain the point cloud to be processed, set the downscaled projection viewpoint and line of sight of the virtual view according to the spatial distribution of the large-scale point cloud, and obtain the external parameters of the virtual camera; The step (1) specifically includes the following steps: (1.1) Read the 3D model file and obtain the point cloud to be processed , to generate high-quality multi-view projections of point clouds, a multi-layer viewpoint network is set up ; (1.2) Get the point cloud corresponding to the k-th layer of viewpoint network to be processed The minimum directed bounding box of , determine the vertical direction away from the ground and along the ground in two directions and , a total of three directions , the initial value of k is 0; (1.3) Based on point cloud density Determine the image resolution of the current virtual view , , in is a constant, According to the ground height , Image resolution ,focal length and pixel size The geometric relationship between them determines the ground height of the view projection , , According to image resolution Get the ground coverage of the image , , in is the number of pixels on one side of the square virtual view; (1.4) Ground height based on the k-th layer virtual view , Ground Coverage , three directions and overlap , set the viewpoint network of the kth layer , the direction of all viewpoints in the viewpoint network is vertical to the ground , the two directions of the view are the two ground directions of the point cloud to be processed, and the rotation matrix is ; The viewpoints in each layer of the viewpoint network are evenly distributed. Line The viewpoint coordinates of the columns are: , The range of row index values is: , The value range of column index is: , The reference viewpoint coordinates are: , in is the ceiling function, are the eight vertices of the minimum directed bounding box, That is, the vertex with index 0. represents the size of the minimum oriented bounding box, is the length of the minimum directed bounding box, is the width of the minimum directional bounding box; thus, the virtual camera extrinsic parameter corresponding to the viewpoint in the i-th row and j-th column of the k-th viewpoint network is obtained. ; (1.5) Based on processing the ground coverage corresponding to each viewpoint , Ground height , three directions , generate the minimum directed bounding box of all point clouds within the field of view of the viewpoint, and extract the point cloud within the box ; If there is a point cloud in the box and the point cloud density Point cloud density of the point cloud corresponding to the previous viewpoint And the point cloud density of the entire point cloud The ratio between meets the threshold condition. , Then the point cloud in the box As the next layer of viewpoint network Repeat steps (1.2) to (1.5) for the point cloud to be targeted, otherwise terminate the iteration, and finally iterate to generate a multi-layer viewpoint network ; (1.6) The virtual camera extrinsic parameters corresponding to each layer of the multi-layer viewpoint network are collected to form a virtual camera extrinsic parameter set, where the virtual camera extrinsic parameter corresponding to the mth viewpoint in the virtual camera extrinsic parameter set is ; (2) Generate the corresponding virtual view based on the intrinsic and extrinsic parameters of the virtual camera corresponding to the m-th viewpoint, and obtain the mapping relationship between the 3D point cloud and the 2D projection point. The intrinsic parameters of the virtual camera include focal length, pixel size, and image size. (3) Based on the existing 2D image semantic segmentation algorithm, the semantic information in the virtual view is extracted. Based on the mapping relationship between the acquired 3D point cloud and the 2D projection points, the semantic information is transferred from 2D to 3D to complete the semantic segmentation of the point cloud.
2. The large-scale point cloud semantic segmentation method based on downscaling projection according to claim 1, characterized in that: The step (2) specifically includes the following steps: (2.1) Based on the intrinsic and extrinsic parameters of the virtual camera corresponding to the m-th viewpoint, the entire point cloud The coordinate system of the world coordinate system Rotate the camera coordinate system ,reserve Points with positive Z values in the coordinate system, that is, the point clouds in front of the virtual camera are retained and other point clouds are discarded; (2.2) From the camera coordinate system Convert to pixel coordinate system , retain the projection points within the image frame; (2.3) Corresponding to the camera coordinate system based on the projection points in the image frame The lower Z value, i.e. depth, retains the projection point with the smallest depth in the same pixel block and discards other points; (2.4) Assign the depth or grayscale of the selected 3D point cloud to the corresponding 2D projection point to obtain a one-to-one mapping relationship between the 3D point cloud and the 2D projection point. , to achieve the initial rendering of the virtual view ; (2.5) Rendering based on initialization ,To address the problem of holes and missing textures in the projected image, the minimum axis-aligned bounding box of all projection points in the image is obtained, and the projection view holes are searched within the bounding box; (2.6) Extract the neighborhood of each hole, and calculate the grayscale value of the hole based on the grayscale value of the non-hole in the neighborhood using the distance Gaussian weighted method , thereby filling the holes and enriching the texture, realizing the rendering post-processing of the virtual view, and generating the virtual view corresponding to the mth viewpoint ; , in is the pixel coordinate, yes Neighborhood of Represents a virtual view that is being initialized middle There is value everywhere, Represents pixel Distance on virtual view.
3. The large-scale point cloud semantic segmentation method based on downscaling projection according to claim 1, characterized in that: The step (3) specifically includes the following steps: (3.1) Process the point cloud with semantic information used for training according to steps (1) and (2), obtain the intrinsic parameters and extrinsic parameters of the virtual camera, generate a virtual view corresponding to the point cloud with semantic information, assign the semantic coding of the three-dimensional point cloud corresponding to the two-dimensional projection point contained in the virtual view to the corresponding two-dimensional projection point, and complete the first step of generating the semantic label map; in order to solve the problem of holes in the semantic label map, obtain the minimum axis-aligned bounding box where all the projection points in the semantic label map are located, search for image holes in the bounding box, extract the neighborhood of each hole, and count the semantic features in the neighborhood of each hole based on the semantic information of the semantic label map. The semantic feature corresponding to the largest number of points of a certain semantic feature in the neighborhood is the semantic feature of the point, and complete the second step of generating the semantic label map, thereby realizing the generation of the final semantic label map; based on the existing two-dimensional image semantic segmentation algorithm, use the virtual view and the corresponding semantic label map to train and generate a two-dimensional image semantic segmentation model; (3.2) The virtual view of step (2) Input the 2D image semantic segmentation model and extract the semantic label map , get the semantic information in the virtual view; (3.3) Create an initial voting table, initialized to zero, based on the 3D point cloud in each view and the two-dimensional projection point The mapping relationship between And the semantic information extracted from each view , the points that received votes are counted, and the semantic attribute with the highest number of votes at each point is set as the semantic attribute of the point, completing the first vote; (3.4) Create a new voting table. For each point that did not receive a vote in the initial voting table, extract its 3D spatial neighborhood and count the points in the neighborhood that have determined the semantic characteristics in the initial voting results. The semantic characteristic corresponding to the point with the largest number of points of a certain semantic characteristic in the neighborhood is the semantic characteristic of the point, completing the second vote. (3.5) The information contained in the two voting tables is merged to obtain the final voting result, completing the semantic information transmission based on the virtual view semantic label.
4. A large-scale point cloud semantic segmentation system based on downscaling projection, characterized in that: include: The external parameter acquisition module is used to obtain the point cloud to be processed, set the downscaled projection viewpoint and line of sight of the virtual view according to the spatial distribution of the large-scale point cloud, and obtain the external parameters of the virtual camera; The external parameter acquisition module reads the three-dimensional model file to obtain the point cloud to be processed ; The 3D model file is obtained by laser radar scanning; in order to generate high-quality multi-view projection of point cloud, a multi-layer viewpoint network is set up ; Get the point cloud corresponding to the k-th layer of viewpoint network currently being processed The minimum directed bounding box of , determine the vertical direction away from the ground and along the ground in two directions and , a total of three directions , the initial value of k is 0; Based on point cloud density Determine the image resolution of the current virtual view , , in is a constant, According to the ground height , Image resolution ,focal length and pixel size The geometric relationship between them determines the ground height of the view projection , , According to image resolution Get the ground coverage of the image , , in is the number of pixels on one side of the square virtual view; The ground height based on the k-th layer virtual view , Ground Coverage , three directions and overlap , set the viewpoint network of the kth layer , the direction of all viewpoints in the viewpoint network is vertical to the ground , the two directions of the view are the two ground directions of the point cloud to be processed, and the rotation matrix is ; The viewpoints in each layer of the viewpoint network are evenly distributed. Line The viewpoint coordinates of the columns are: , The range of row index values is: , The value range of column index is: , The reference viewpoint coordinates are: , in is the ceiling function, are the eight vertices of the minimum directed bounding box, That is, the vertex with index 0. represents the size of the minimum oriented bounding box, is the length of the minimum directed bounding box, is the width of the minimum directional bounding box; thus, the virtual camera extrinsic parameter corresponding to the viewpoint in the i-th row and j-th column of the k-th viewpoint network is obtained. ; Based on processing the ground coverage corresponding to each viewpoint , Ground height , three directions , generate the minimum directed bounding box of all point clouds within the field of view of the viewpoint, and extract the point cloud within the box ; If there is a point cloud in the box and the point cloud density Point cloud density of the point cloud corresponding to the previous viewpoint And the point cloud density of the entire point cloud The ratio between meets the threshold condition. , Then the point cloud in the box As the next layer of viewpoint network To target the point cloud, continue to iterate, otherwise terminate the iteration, and finally iterate to generate a multi-layer viewpoint network ; The virtual camera extrinsic parameters corresponding to each layer of the multi-layer viewpoint network are collected to form a virtual camera extrinsic parameter set, where the virtual camera extrinsic parameter corresponding to the mth viewpoint in the virtual camera extrinsic parameter set is ; A mapping relationship acquisition module is used to generate a corresponding virtual view for the internal parameters and external parameters of the virtual camera corresponding to the m-th viewpoint, and to obtain the mapping relationship between the three-dimensional point cloud and the two-dimensional projection point, wherein the internal parameters of the virtual camera include focal length, pixel size and image size; The semantic segmentation module is used to extract semantic information in the virtual view based on the existing 2D image semantic segmentation algorithm, and transfer the semantic information from 2D to 3D based on the mapping relationship between the acquired 3D point cloud and the 2D projection points to complete the semantic segmentation of the point cloud.
5. The large-scale point cloud semantic segmentation system based on downscaling projection according to claim 4, characterized in that: The mapping relationship acquisition module converts the entire point cloud into The coordinate system of the world coordinate system Rotate the camera coordinate system ,reserve Points with positive Z values in the coordinate system, that is, the point clouds in front of the virtual camera are retained and other point clouds are discarded; From the camera coordinate system Convert to pixel coordinate system , retain the projection points within the image frame; Corresponding to the camera coordinate system based on the projection points in the image frame The lower Z value, i.e. depth, retains the projection point with the smallest depth in the same pixel block and discards other points; Assign the depth or grayscale of the selected 3D point cloud to the corresponding 2D projection point to obtain a one-to-one mapping relationship between the 3D point cloud and the 2D projection point , to achieve the initial rendering of the virtual view ; Initialized rendering ,To address the problem of holes and missing textures in the projected image, the minimum axis-aligned bounding box of all projection points in the image is obtained, and the projection view holes are searched within the bounding box; Extract the neighborhood of each hole, and calculate the grayscale value of the hole based on the grayscale value of the non-hole in the neighborhood using the distance Gaussian weighted method , thereby filling the holes and enriching the texture, realizing the rendering post-processing of the virtual view, and generating the virtual view corresponding to the mth viewpoint ; , in is the pixel coordinate, yes Neighborhood of Represents a virtual view being initialized middle There is value everywhere, Represents pixel Distance on virtual view.
6. The large-scale point cloud semantic segmentation system based on downscaling projection according to claim 4, characterized in that: The semantic segmentation module processes the point cloud with semantic information for training, obtains the internal and external parameters of the virtual camera, generates a virtual view corresponding to the point cloud with semantic information, assigns the semantic coding of the two-dimensional projection point corresponding to the three-dimensional point cloud contained in the virtual view to the corresponding two-dimensional projection point, and completes the first step of generating the semantic label map; To solve the problem of holes in the semantic label map, the minimum axis-aligned bounding box of all projection points in the semantic label map is obtained, image holes are searched in the bounding box, the neighborhood of each hole is extracted, and the semantic features in the neighborhood of each hole are counted based on the semantic information of the semantic label map. The semantic feature corresponding to the largest number of points of a certain semantic feature in the neighborhood is the semantic feature of the point, completing the second step of semantic label map generation and realizing the generation of the final semantic label map; based on the existing two-dimensional image semantic segmentation algorithm, the virtual view and the corresponding semantic label map are used to train and generate a two-dimensional image semantic segmentation model; Virtual View Input the 2D image semantic segmentation model and extract the semantic label map , get the semantic information in the virtual view; Create an initial voting table, initialized to zero, based on the 3D point cloud in each view and the two-dimensional projection point The mapping relationship between And the semantic information extracted from each view , the points that received votes are counted, and the semantic attribute with the highest number of votes at each point is set as the semantic attribute of the point, completing the first vote; A new voting table is created. For each point that did not receive a vote in the initial voting table, its 3D spatial neighborhood is extracted and the points in the neighborhood whose semantic characteristics have been determined in the initial voting results are counted. The semantic characteristic corresponding to the largest number of points of a certain semantic characteristic in the neighborhood is the semantic characteristic of the point, and the second voting is completed. The information contained in the two voting tables is merged to obtain the final voting result, completing the semantic information transmission based on the virtual view semantic label.
7. An electronic device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 3.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Three-dimensional point cloud semantic segmentation method for optimizing boundary
CN115409989A
Inter-class characterization contrast driven graph convolution point cloud semantic annotation method
CN116206306A