Visual semantic information assisted point cloud feature extraction and pose estimation method and device
By using visual semantic information-assisted methods and processing degradation features with Hessian matrix and orthogonal projection matrix, the positioning drift problem of traditional laser SLAM systems in feature degradation scenarios is solved, achieving high-precision and robust pose estimation.
Patent Information
- Application Number
- CN202511090622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Traditional laser SLAM systems suffer from odometry drift in degraded scenarios such as open roads and long straight corridors. Existing degradation detection methods use fixed threshold strategies, which are difficult to adapt to the different degrees of degradation in different scenarios, leading to error divergence.
By using a visual semantic information-assisted method, the degradation factor is calculated using the Hessian matrix, semantic point clouds are constructed and feature classification is performed. Degenerate features are projected onto the degenerate subspace using an orthogonal projection matrix to constrain them, while preserving the constraint ability of non-degenerate features, thus constructing a pose estimation least squares problem.
It improves the robustness and accuracy of feature extraction, solves the problem of localization drift, achieves high-precision and robust localization and mapping, adapts to different degradation levels in different scenarios, and improves the stability of pose estimation.
Smart Images

Figure CN120599047B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of robot autonomous navigation and computer vision, specifically relating to a method and apparatus for point cloud feature extraction and pose estimation assisted by visual semantic information. Background Technology
[0002] LiDAR odometry, as a core component of Simultaneous Localization and Mapping (SLAM) technology, achieves motion estimation by continuously analyzing LiDAR point cloud data. Existing technologies mainly rely on feature matching methods: establishing inter-frame correspondences by extracting geometric features from the point cloud, and then using feature matching to solve for pose transformation.
[0003] SLAM technology has been widely applied in fields such as intelligent transportation and industrial automation, but in degraded scenarios such as open highways and long straight corridors, there is still a technical challenge of odometer drift. Specifically, traditional laser SLAM systems face difficulties in data extraction in degraded scenarios due to feature sparsity and structural similarity; existing degradation detection methods mostly adopt fixed threshold strategies, which are difficult to adapt to the differences in degradation levels in different scenarios, and there are error divergence problems in severely degraded scenarios such as open highways. Summary of the Invention
[0004] To address the above technical problems, this invention provides a method and apparatus for point cloud feature extraction and pose estimation assisted by visual semantic information, the details of which are as follows:
[0005] A method for point cloud feature extraction and pose estimation assisted by visual semantic information includes the following steps:
[0006] Step 1: Acquire laser point cloud and visual image. Input the visual image into the semantic segmentation network and output a pixel-level semantic label map. Project the laser point cloud onto the pixel-level semantic label map to obtain the semantic point cloud. Use the semantic point cloud to initially distinguish edge and planar features, calculate the local curvature of the semantic point cloud, and then filter to obtain planar feature points and edge feature points. Construct the error function of the filtering process and calculate the Hessian matrix.
[0007] Step 2: Calculate the degradation factor based on the Hessian matrix. Determine the degree of degradation of the current feature based on the degradation factor. Degradation judgment is performed on planar features and edge features respectively. If the degradation factor is less than the degradation detection threshold, the current feature is not degraded and is recorded as a non-degraded feature. If the degradation factor is greater than the degradation detection threshold, the current feature is determined to be degraded and is recorded as a degraded feature.
[0008] Step 3: Construct an orthogonal projection matrix to project the degenerate features onto the degenerate subspace for constraint, retaining the complete constraint capability of non-degenerate features on pose estimation. Combine the orthogonal projection matrices of planar features and edge features to construct a least-squares problem for pose estimation. Repeat the above steps until the carrier motion ends and the best pose estimate is obtained.
[0009] A visual semantic information-assisted point cloud feature extraction and pose estimation device, comprising:
[0010] The Hessian matrix calculation module acquires laser point clouds and visual images, inputs the visual images into a semantic segmentation network, outputs a pixel-level semantic label map, projects the laser point clouds onto the pixel-level semantic label map to obtain a semantic point cloud, uses the semantic point cloud to initially distinguish edge and planar features, calculates the local curvature of the semantic point cloud, further filters to obtain planar feature points and edge feature points, constructs an error function for the filtering process, and calculates the Hessian matrix.
[0011] The degradation feature judgment module calculates the degradation factor based on the Hessian matrix and judges the degree of degradation of the current feature according to the degradation factor. It performs degradation judgment on planar features and edge features respectively. If the degradation factor is less than the degradation detection threshold, the current feature is not degraded and is recorded as a non-degraded feature. If the degradation factor is greater than the degradation detection threshold, the current feature is judged to be degraded and is recorded as a degraded feature.
[0012] The optimal pose estimation module constructs an orthogonal projection matrix, projects degenerate features onto the degenerate subspace for constraint, preserves the complete constraint capability of non-degenerate features on pose solution, and constructs a pose estimation least squares problem by combining the orthogonal projection matrices of planar features and edge features. The above steps are repeated until the carrier motion ends to obtain the optimal pose estimation.
[0013] An electronic device includes: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.
[0014] A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.
[0015] The present invention has the following beneficial effects:
[0016] (1) This invention improves the robustness of feature extraction in degraded environments. Traditional laser SLAM suffers from feature degradation in scenarios such as long straight corridors and open squares, where sparse or repetitive geometric features lead to feature matching failure and a sharp decline in pose estimation accuracy. By comprehensively analyzing the distribution characteristics and semantic label information in lidar point clouds, a multi-level point cloud feature classification system is constructed. At the geometric level, stable geometric structures in low-feature environments are identified through regional distribution analysis, and prior knowledge is integrated at the semantic level. The stability of semantic labels compensates for the deficiencies in geometric features, thereby improving the accuracy and efficiency of feature recognition. This invention solves the technical problem of positioning drift in pose estimation in feature degradation scenarios, and achieves high-precision and robust positioning and mapping in typical degraded environments such as open roads and long straight corridors.
[0017] (2) This invention considers that traditional degradation detection methods usually adopt a fixed threshold strategy, which cannot adapt to the changes in the degree of degradation in different scenarios, leading to misjudgment or missed detection. This invention uses condition number analysis based on the Hessian matrix to achieve dynamic degradation assessment, evaluate the quality of feature distribution in real time, improve the utilization efficiency of limited features in the degradation environment and the stability of pose estimation. For the different characteristics of planar and edge features, the condition number of their observation matrix is calculated separately to quantify the degree of feature degradation and avoid the problem of insufficient adaptability caused by a single threshold.
[0018] (3) Based on the degradation detection results, the problem space is decomposed into a reliable constraint subspace and a degradation subspace. Specifically, when a certain type of feature degradation is detected, it is projected onto the degradation subspace for constraint suppression, while preserving the complete constraint capability of another type of feature in the reliable subspace. The method of this invention has strong environmental adaptability and improves the robustness of pose estimation in degradation environments.
[0019] (4) In the data processing stage, the present invention first extracts semantic information from the visual image, then projects the laser point cloud onto the pixel-level semantic label map to obtain the semantic point cloud, and uses semantic prior information to assist in the feature extraction of the laser point cloud, thus making up for the shortcomings of extracting only geometric features in the degraded environment. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method of the present invention;
[0021] Figure 2 This is a structural diagram of the electronic device of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.
[0023] This invention proposes a method for point cloud feature extraction and pose estimation assisted by visual semantic information, such as... Figure 1 The diagram shown is a flowchart of the invention, which includes the following steps:
[0024] Step 1: Acquire laser point cloud and visual image. Input the visual image into the semantic segmentation network and output a pixel-level semantic label map. Project the laser point cloud onto the pixel-level semantic label map to obtain the semantic point cloud. Use the semantic point cloud to initially distinguish edge and planar features, calculate the local curvature of the semantic point cloud, and then filter to obtain planar feature points and edge feature points. Construct the error function of the filtering process and calculate the Hessian matrix.
[0025] Step 2: Calculate the degradation factor based on the Hessian matrix. Determine the degree of degradation of the current feature based on the degradation factor. Degradation judgment is performed on planar features and edge features respectively. If the degradation factor is less than the degradation detection threshold, the current feature is not degraded and is recorded as a non-degraded feature. If the degradation factor is greater than the degradation detection threshold, the current feature is determined to be degraded and is recorded as a degraded feature.
[0026] Step 3: Construct an orthogonal projection matrix to project the degenerate features onto the degenerate subspace for constraint, retaining the complete constraint capability of non-degenerate features on pose estimation. Combine the orthogonal projection matrices of planar features and edge features to construct a least-squares problem for pose estimation. Repeat the above steps until the carrier motion ends and the best pose estimate is obtained.
[0027] Step 1 specifically involves:
[0028] This method utilizes LiDAR to acquire laser point clouds and a camera to acquire visual images. The timestamps of the laser point clouds and visual images are aligned to ensure time synchronization. Pre-calibrated extrinsic and intrinsic parameter matrices of the LiDAR and camera are loaded, establishing a coordinate transformation relationship between them. The visual images are input into a semantic segmentation network, which outputs a pixel-level semantic label map. The laser point clouds are projected onto this map based on the established coordinate transformation relationship, and a semantic label is assigned to each point cloud. The semantic labels are then used to initially distinguish between edge and planar features. For planar features, a point cloud is identified as planar if it meets a preset line threshold or possesses specific semantic category attributes such as ground, road surface, or wall. For edge feature extraction, detection is performed in the remaining point cloud regions after excluding planar features. Simultaneously, targeted extraction is performed using semantic objects with typical edge characteristics, such as building edges and lampposts, thereby achieving efficient and accurate classification of point cloud features.
[0029] By deeply fusing visual semantic information with the geometric features of laser point clouds, the extraction efficiency of sparse environment features is significantly improved. First, using the extrinsic parameters of the LiDAR and camera, the 3D laser point cloud is projected onto a 2D pixel-level semantic label map, assigning semantic attributes to each laser point cloud to construct a semantic point cloud with environmental understanding capabilities. Specifically, the visual image is input into a lightweight DeeplaV3+ semantic segmentation network for semantic segmentation to obtain a pixel-level semantic label map. The laser point cloud is then projected onto the pixel-level semantic label map using the extrinsic parameters of the LiDAR and camera, enabling the laser point cloud to acquire semantic information corresponding to the pixel-level semantic label map. Specifically, the 3D laser point cloud is projected onto the pixel-level semantic label map using the following formula:
[0030] (1)
[0031] In the formula Let be the coordinates of the point on the laser point cloud. For point Coordinate values projected onto the pixel-level semantic tag map This refers to the pixel coordinates of that point in the pixel-level semantic label map. This is the intrinsic parameter matrix for the LiDAR and camera. Given the extrinsic parameter matrices of the LiDAR and camera, the pixel-level semantic label map is fused with the LiDAR point cloud using the method described above. A semantic label is assigned to each LiDAR point cloud to obtain a semantic point cloud, and subsequent steps are all based on this semantic point cloud.
[0032] This invention proposes a point cloud feature classification method based on the fusion of geometric features and semantic information. By comprehensively analyzing the point cloud distribution characteristics and semantic label information in the semantic point cloud, a dual judgment criterion is adopted for planar features. That is, when the semantic point cloud data meets the set beam distribution requirements (e.g., planar features are distributed in areas where the lidar beam is less than 6), or belongs to planar semantic categories such as ground, road surface, and wall, it is judged as a planar feature. For edge feature recognition, scanning and detection are performed in the non-planar feature point cloud space, and semantic targets with significant edge features such as building outlines and street lamp poles are simultaneously associated for collaborative extraction. Thus, a multi-level point cloud feature classification system is constructed. The planar feature judgment formula is as follows:
[0033] (2)
[0034] in, This represents the i-th point in the point cloud. Represents a set of point clouds. This refers to the lidar beam area where the point is located. For the harness threshold, For point semantic tags, This is a set of planar semantic categories. Similarly, the formula for determining edge features is as follows:
[0035] (3)
[0036] in, Filter the beam threshold for edge features. This is the set of edge semantic categories. By using the semantic features mentioned above to assist feature extraction, the problem of feature loss in scenarios such as long straight corridors and open squares, as described in traditional methods, is effectively solved. After preliminary screening of planar and edge features using formulas (2) and (3), further screening of planar and edge features is performed based on the local curvature of the point cloud. Let... For laser point clouds and Calculate the curvature of a point cloud for a continuous set of points on the same point cloud bundle:
[0037] (4)
[0038] in, For laser point clouds and A continuous set of points on the same point cloud bundle, Let j represent the j-th point in the point cloud. To ensure uniform sampling, the scanning area is divided into several parts for each scan. Within each part, select... The sharpest point is used as the edge feature. The flattest point is used as the planar feature. This ultimately yields the set of edge feature points for the entire target frame. and the set of planar feature points Taking the optimization from point to plane as an example, we define the feature points of the plane as... The point on the target plane is The error function of the optimization process can be expressed as:
[0039] (4)
[0040] Let be the error function. The error to be optimized The number of selected flat points, where The pose transformation rotation and translation parameters to be solved are... The problem is solved using the Gauss-Newton method:
[0041] (5)
[0042] In the formula: For small perturbations, and Error functions The Jacobian and Hessian matrices. It is the first Jacobian coefficients of each residual term with respect to the rotated part. It is the first The Jacobian coefficients of each residual term with respect to the translation portion, for example middle This represents the Jacobian coefficient of the first residual term with respect to the rotated part. This represents the Jacobian coefficient of the first residual term with respect to the translation portion. This represents the second derivative of the error function with respect to the rotation (Hessian submatrix). and This represents the mixed derivative of the error function with respect to rotation and translation. The Hessian submatrix represents the second derivative of the error function with respect to the translation.
[0043] Step 2 is as follows:
[0044] The odometry adaptive degradation detection algorithm based on condition number analysis is as follows:
[0045] First, we use eigenvalue decomposition to obtain the eigenvalues and eigenvectors of the Hessian matrix:
[0046] (6)
[0047] in, It is an eigenvalue. It is the corresponding feature vector, and this invention uses... The ratio of the minimum and maximum eigenvalues (i.e., the condition number) of a matrix determines whether planar and marginal features are degenerate, and the degradation factor is defined as:
[0048] (7)
[0049] The condition number was originally used in numerical analysis to measure the stability of a matrix. In this invention, the maximum eigenvalue ( ) represents the dominant observation direction of the local region, while the minimum eigenvalue ( This reveals the most sensitive and weakest observation dimension for features. In degraded environments, the minimum eigenvalue is usually very small, resulting in a low eigenvalue ratio. By calculating the degradation factor, it is possible to effectively determine whether degradation has occurred in a local area of the point cloud. A smaller degradation factor indicates degraded features, while a larger degradation factor indicates that the features are localizable.
[0050] To adapt to changes in different environments, this invention assesses the degradation status of planar features and edge features based on historical data of degradation factors:
[0051] (8)
[0052] In the formula, To determine the degradation detection threshold, historical data on degradation factors are collected using a sliding window, and their mean and standard deviation are calculated. Based on this statistical information, the degradation detection threshold is dynamically adjusted and set as follows:
[0053] (9)
[0054] The mean of the degradation factor This reflects the average level of degradation factors within the historical window, characterizing the general state of the scene's features. Standard deviation of the degradation factor It quantifies the range of fluctuations caused by environmental changes. As an adjustable factor, it allows for a flexible balance between detection sensitivity and robustness. Priority should be given to ensuring rapid detection at the onset of degradation. , A negative value promptly detects initial signs of deterioration in environmental constraints, ensuring that degraded environments retain more usable features; conversely, when environmental conditions improve and return to a non-degraded state, i.e. hour, Switching to a positive value and appropriately increasing the degradation threshold enhances the stability of the system's pose estimation. Dynamically adjusting the strategy avoids misjudgments caused by instantaneous environmental fluctuations, ensuring the system quickly and reliably recovers to its normal operating state.
[0055] Step 3 specifically involves:
[0056] Complementary geometric constraints are provided for planar and edge features in pose estimation. Planar features, represented by the ground, constrain the vertical degrees of freedom through their normal vectors, including: Axial translation (height change) and rotation The rotation of the axis (pitch and roll angles); while edge features, primarily based on the object's outline, restrict the degrees of freedom in the horizontal plane through tangent vectors: covering... Lateral displacement and circumference of the axis The rotation of the axis (yaw angle). This positive complementarity of the constraint dimensions ensures that the system does not undergo synchronous degradation under normal observation conditions. Even if a single type of feature fails, such as the absence of edge features in a corridor or insufficient planar features in an open scene, the other type of feature can still maintain a minimum level of pose observability. To address this important observation, this invention constructs orthogonal projection matrices to suppress the degrees of freedom in the degradation directions corresponding to the degraded planar or edge features:
[0057] (10)
[0058] In the formula, This represents the constructed orthogonal projection matrix. Represents the identity matrix. Represents the smallest eigenvalue ( The eigenvectors corresponding to ) The expression represents the transpose of the eigenvector corresponding to the minimum eigenvalue. The above formula decomposes the problem space of planar and edge features into a reliable constraint subspace and a degenerate subspace. Taking an open highway scene as an example, the lack of vertical constraints in the ground plane point cloud easily leads to degenerate pose estimation; while lane lines and guardrail features not only provide strong constraints for horizontal motion but also partially compensate for the lack of vertical constraints. In planar features, the eigenvector corresponding to the minimum eigenvalue is... Then, it is projected onto the degenerate subspace, and the lack of constraint in this direction is compensated by the non-degenerate edge features. Taking planar feature degradation as an example, when planar feature degradation is detected, it is projected onto the degenerate subspace for constraint suppression. Then, the degenerate feature is projected as:
[0059] (11)
[0060] Indicates degenerative characteristics The available features are projected onto the degenerate subspace, while preserving the complete constraint capability of edge features in the reliable subspace. The pose estimation solutions of the two types of features are merged to establish a unified description:
[0061] (12)
[0062] In the formula, This indicates the construction of an orthogonal projection matrix based on planar features. This represents the orthogonal projection matrix constructed based on edge features. Let represent the final orthogonal projection matrix. The final optimization problem can be expressed as follows:
[0063] (13)
[0064] In the formula, This represents the rotation and translation parameters in the final optimization solution. , Extracted according to the present invention edge features and By using planar features, the optimal solution for pose estimation can be obtained, thereby improving the utilization efficiency of finite features in degraded environments and the stability of pose estimation.
[0065] Another aspect of the present invention provides a visual semantic information-assisted point cloud feature extraction and pose estimation device, comprising:
[0066] The Hessian matrix calculation module acquires laser point clouds and visual images, inputs the visual images into a semantic segmentation network, outputs a pixel-level semantic label map, projects the laser point clouds onto the pixel-level semantic label map to obtain a semantic point cloud, uses the semantic point cloud to initially distinguish edge and planar features, calculates the local curvature of the semantic point cloud, further filters to obtain planar feature points and edge feature points, constructs an error function for the filtering process, and calculates the Hessian matrix.
[0067] The degradation feature judgment module calculates the degradation factor based on the Hessian matrix and judges the degree of degradation of the current feature according to the degradation factor. It performs degradation judgment on planar features and edge features respectively. If the degradation factor is less than the degradation detection threshold, the current feature is not degraded and is recorded as a non-degraded feature. If the degradation factor is greater than the degradation detection threshold, the current feature is judged to be degraded and is recorded as a degraded feature.
[0068] The optimal pose estimation module constructs an orthogonal projection matrix, projects degenerate features onto the degenerate subspace for constraint, preserves the complete constraint capability of non-degenerate features on pose solution, and constructs a pose estimation least squares problem by combining the orthogonal projection matrices of planar features and edge features. The above steps are repeated until the carrier motion ends to obtain the optimal pose estimation.
[0069] In another aspect, the present invention provides an electronic device, Figure 2The diagram illustrates the structure of an electronic device provided in an embodiment of the present invention. For example, the electronic device may include a processor, a memory, and a transmission device. The processor is used to execute the above-described methods. The processor and the memory can be connected via a bus or other means, taking a bus connection as an example. The transmission device can be connected to the processor and the memory via wired or wireless means. The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules.
[0070] In another aspect, the present invention provides a computer-readable storage medium, which may be a computer-readable storage medium included in the apparatus described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into the device. This computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD ROM, or any other form of storage medium known in the art.
Claims
1. A method of visual semantic information assisted point cloud feature extraction and pose estimation, characterized in that, The method comprises the following steps: Step 1, obtaining laser point cloud and visual image, inputting the visual image into a semantic segmentation network to output a pixel-level semantic label map, projecting the laser point cloud to the pixel-level semantic label map to obtain a semantic point cloud; preliminarily distinguishing edge and plane features by using the semantic point cloud, calculating local curvature of the semantic point cloud, and reselecting to obtain plane feature points and edge feature points; constructing an error function of the reselection process and calculating a Hessian matrix; Step 2, calculating a degeneration factor based on the Hessian matrix, judging the degeneration degree of the current feature according to the degeneration factor, and respectively judging the degeneration of the plane feature and the edge feature; if the degeneration factor is less than a degeneration detection threshold, the current feature is not degenerated and is recorded as a non-degenerated feature; if the degeneration factor is greater than the degeneration detection threshold, it is judged that the current feature is degenerated and is recorded as a degenerated feature; Step 3, constructing an orthogonal projection matrix, so as to decompose the problem space of the plane feature and the edge feature into a reliable constraint subspace and a degenerated subspace, projecting the degenerated feature to the degenerated subspace for constraint, retaining the complete constraint capability of the non-degenerated feature for solving the pose, and combining the orthogonal projection matrices of the plane feature and the edge feature to construct a pose estimation least square problem, and repeating the above steps 1-3 until the best pose estimation is obtained. 2.The visual semantic information aided point cloud feature extraction and pose estimation method according to claim 1, characterized in that, In step 1, laser point cloud and visual image are obtained by using a laser radar and a camera respectively, a pre-calibrated laser radar and camera extrinsic parameter matrix is loaded, a coordinate conversion relationship between the laser radar and the camera is established, the laser point cloud is projected to a pixel-level semantic label map according to the coordinate conversion relationship, a semantic label is assigned to each three-dimensional point cloud, and a semantic point cloud is obtained. 3.The visual semantic information aided point cloud feature extraction and pose estimation method according to claim 1, characterized in that, Eigenvalue decomposition is performed on the Hessian matrix to obtain eigenvalues and eigenvectors thereof; the ratio of the minimum eigenvalue to the maximum eigenvalue is taken as a degeneration factor for judging the degeneration degree of the plane feature and the edge feature. 4.The method of claim 1, wherein, The degeneration detection threshold is dynamically adjusted according to the mean value and the standard deviation of the historical data of the degeneration factor. 5.The visual semantic information aided point cloud feature extraction and pose estimation method according to claim 1, characterized in that, In step 1, an error function of the plane feature points and the points on the target plane is defined, appropriate flat points and plane feature points are selected, the error function is constructed, the error function is solved by using a Gauss-Newton method, and a Jacobian matrix and a Hessian matrix are calculated. 6.The visual semantic information aided point cloud feature extraction and pose estimation method according to claim 1, characterized in that, In step 3, when the plane feature is detected to be degenerated, the plane feature is projected to the degenerated subspace for constraint suppression, and the complete constraint capability of the edge feature in the reliable subspace is retained; conversely, when the edge feature is degenerated, the constraint capability of the plane feature is retained. 7.The visual semantic information aided point cloud feature extraction and pose estimation method according to claim 1, characterized in that, A double judgment criterion is adopted for the plane feature, that is, when the semantic point cloud data meets a set bundle distribution requirement or belongs to a ground, road surface or wall surface plane type semantic category, the plane feature is determined; for the identification of the edge feature, scanning detection is performed in a non-plane feature point cloud space, and semantic targets such as building outlines and street lamp poles with obvious edge features are synchronously associated and cooperatively extracted.
8. A device for visual semantic information aided point cloud feature extraction and pose estimation, which performs the method according to claim 1, characterized in that, The method comprises the following steps: The Hessian matrix calculation module obtains the laser point cloud and the visual image, inputs the visual image into a semantic segmentation network, outputs a pixel-level semantic label map, projects the laser point cloud to the pixel-level semantic label map, and obtains a semantic point cloud; the semantic point cloud is used to preliminarily distinguish edge and plane features, calculate local curvatures of the semantic point cloud, and again screen to obtain plane feature points and edge feature points; an error function of the screening process is constructed, and the Hessian matrix is calculated; The degenerate feature judgment module calculates a degenerate factor based on the Hessian matrix, judges the degenerate degree of the current feature according to the degenerate factor, respectively judges the degeneration of the plane feature and the edge feature, and if the degenerate factor is less than a degeneration detection threshold, the current feature is not degenerated and is recorded as a non-degenerated feature, and if the degenerate factor is greater than the degeneration detection threshold, the current feature is judged to be degenerated and is recorded as a degenerated feature; The optimal pose estimation module constructs an orthogonal projection matrix, projects the degenerated feature to a degenerated subspace for constraint, retains the complete constraint capability of the non-degenerated feature for solving the pose, and constructs a pose estimation least square problem by combining the orthogonal projection matrices of the plane feature and the edge feature.
9. An electronic device, comprising: Comprise: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, An executable instruction is stored thereon, which is executed by a processor to enable the processor to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Point cloud map creation and scene identification method based on static semantic information
CN112767485A
Multi-level semantic map construction method and device based on deep learning perception
CN115655262A