Navigation method based on laser radar in large-scale driving scene

The novel point cloud segmentation network addresses the limitations of small-scale networks in large-scale driving scenarios by integrating adaptive fusion and semantic association, achieving improved accuracy and efficiency in navigating complex environments.

CN120313631AActive Publication Date: 2025-07-15XIANGTAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510810007.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle point cloud data in large-scale driving scenarios, and there are problems of disorder, unstructured and high computational complexity, resulting in low data utilization and serious information loss in traditional methods during the encoding process.

Method used

Using a navigation method based on lidar, feature extraction and aggregation is performed through a codec network, combining attention interpolation module and semantic association mining module to realize multi-resolution fusion and cross-compensation of context information, improving the semantic segmentation accuracy of point clouds.

Benefits of technology

It significantly improves the segmentation accuracy and efficiency of point cloud data in large-scale scenarios, improves the safety and reliability of navigation systems, especially in complex driving scenarios, distinguishes the ability to easily confused categories, reduces the misjudgment rate and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120313631A_ABST
    Figure CN120313631A_ABST
Patent Text Reader

Abstract

The invention discloses a laser radar-based navigation method in a large-scale driving scene, and belongs to the technical field of navigation. The method comprises the following steps: acquiring and inputting a laser point cloud signal; carrying out feature extraction and aggregation on the laser point cloud signals through a coding layer of a coding and decoding network to obtain local features, and carrying out multi-resolution fusion on the local features of different scales through a decoding layer of the coding and decoding network by utilizing an adaptive fusion mechanism and combining an attention interpolation module to output advanced features; inputting the advanced features output by the coding and decoding network into a semantic association mining module, and performing embedding processing on the output features through two MLP layers; and performing semantic segmentation on the final output features, and planning an optimal navigation route based on a semantic segmentation result. According to the network provided by the invention, through modular innovation and system optimization, significant breakthroughs are made in the aspects of precision, efficiency and generalization ability of point cloud semantic segmentation in a large-scale scene, and a high-reliability environment perception solution is provided for automatic driving navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of navigation, and particularly relates to a lidar-based navigation method in large-scale driving scenarios. Background Art

[0002] In the field of autonomous driving, lidar technology is often used for positioning and navigation. The lidar scans the target scene by emitting laser pulses to form a laser point cloud. Each laser point cloud contains information such as the three-dimensional coordinates, color, and laser reflection intensity of the point in space. By processing and analyzing these point cloud data, the perception of the target scene environment can be achieved.

[0003] In recent years, with the increasing demand for the understanding of three-dimensional scenes, the scope of the point cloud dataset has gradually expanded from indoor small scenes to outdoor large scenes and even urban-scale scenes. Compared with the common problems of point cloud data in large-scale scenes, such as disorder, unstructuredness, high computational complexity, and low data utilization rate, the common characteristics of point cloud data in small scenes are relatively uniform density, fewer points, a small range of scene areas, and relatively single target types. Therefore, the original lidar point cloud processing network for small scenes is not suitable for direct expansion and application to large-scale scenes.

[0004] Therefore, it is necessary to provide a lidar-based navigation method in large-scale driving scenarios to solve the problems proposed in the above background art. Summary of the Invention

[0005] The present invention discloses a lidar-based navigation method in large-scale driving scenarios, which provides a new point cloud semantic segmentation network, thereby improving the accuracy of lidar navigation in large-scale driving scenarios, and effectively solving at least one technical problem involved in the background art.

[0006] To achieve the above object, the technical solution of the present invention is as follows: A lidar-based navigation method in large-scale driving scenarios includes the following steps: Step S1, obtaining and inputting a lidar point cloud signal; Step S2, extracting and aggregating features of the lidar point cloud signal through the encoding layer of the encoder-decoder network to obtain local features, and then using an adaptive fusion mechanism in combination with an attention interpolation module through the decoding layer of the encoder-decoder network to perform multi-resolution fusion on local features of different scales and output high-level features; Step S3, inputting the high-level features into a semantic association mining module, and performing embedding processing on the output high-level features through two MLP layers of the semantic association mining module to form final output features, wherein the semantic association mining module integrates spatial information and context information through cross compensation; Step S4, perform semantic segmentation on the final output features, and plan the optimal navigation route based on the semantic segmentation results.

[0007] Optionally, in step S2, the specific implementation of the attention interpolation module includes: Step S21, in the interpolation layer, with the set of points to be interpolated at the m-th layer as the center, perform K-nearest neighbor search on the feature map of the small-scale point set to obtain the local representation: ; ; ; In the formula, KNN is the K-nearest neighbor classification algorithm; is the preliminary interpolation result; is the point cloud information of the γ module; is the parameter matrix; γ is the combined module of linear mapping and batch normalization layer; is the number of points at the m-th layer to be interpolated; is the additional known features included in the small-scale point set ; is the result feature of K-value interpolation; Step S22, map through the combined module γ of linear mapping and batch normalization layer to generate the parameter matrix , and perform softmax normalization and weighted summation in the dimension to output the interpolation result: ; In the formula, is the operation in the dimension; softmax (·) is the softmax operation.

[0008] Optionally, in step S3, the specific implementation of the semantic association mining module includes: Step S31, based on the neighborhood center context and the coordinates , where is the number of points of the neighborhood center c, is the number of points; and the absolute point set context and the coordinates , the relative point set context and the relative coordinates can be obtained; Step S32, encode the context information and geometric space information, and generate the mixed feature F′ and geometric encoding P′ through the splicing operation and the 1×1 convolutional layer; Step S33: Aggregate the hybrid features by combining max pooling and attention pooling to generate the final context feature encoding 。

[0009] Optionally, the encoding layer dynamically adjusts the feature extraction range through a receptive field reconstruction mechanism, and the attention interpolation module is used to improve the accuracy of high-level features.

[0010] Optionally, the number of points in the m-th layer to be interpolated is ,wherein δ m is a preset downsampling factor.

[0011] Optionally, in step S32, the concatenation operation of the context encoding and the geometric encoding satisfies: ; ; ; ; wherein copy (·)represents a shallow copy of the encoding; concat(·)represents the concatenation operation; conv 1×1 represents a 1×1 convolution.

[0012] Optionally, in step S33, the final context feature encoding is represented by the following formula: 。

[0013] Optionally, the semantic segmentation result is generated by a Softmax classifier, and the classification categories include roads, obstacles, vehicles, and pedestrians.

[0014] The beneficial effects of the present invention are as follows: 1. Through the semantic association mining module, the present invention realizes the cross compensation and gradient parallel propagation of spatial geometric information and context information, effectively solves the semantic ambiguity problem caused by the sparsity and unstructuredness of point clouds in large-scale scenes, ensures the coordination of semantic association mining, and thus enables the network to learn relatively more correct information. Experimental data shows that in the six-fold cross-validation of the S3DIS dataset, the overall accuracy (OA) of this method reaches 89.4%, and the mean intersection over union (mIoU) reaches 73.6%, which is about 1.4%-2.0% higher than the existing baseline models (such as BAAF-Net); on the SemanticKITTI validation set, for the first time, 60% mIoU is achieved in the point-wise MLP method without data augmentation, and the recognition accuracy of key categories (such as people and vehicles) is significantly better than that of similar methods.

[0015] 2. The attention interpolation module provided by the present invention adopts a reverse scale interpolation mechanism (interpolating from a small scale to a large scale), combines KNN local aggregation and adaptive weighting, and realizes the inverse inference of context information from coarse-grained to fine-grained. This design effectively improves the resolution of sparse point clouds and solves the problem of information loss in the encoding process of traditional methods. Ablation experiments show that after introducing this module, the mIoU of the model is increased by 1.7%, the inference time is only increased by 0.1 second, and the parameter scale is reduced by 57% (from 4.97M to 2.15M).

[0016] 3. The present invention replaces the traditional dense KNN and downsampling strategies through the receptive field reconstruction (RFR) mechanism, while reducing the computational complexity, retains key features. Combining the hybrid pooling (max pooling and attention pooling) of the semantic context encoding (SCE) module, the number of model parameters is reduced to less than 50% of the baseline model, and the inference speed is increased by 30% (from 0.8 second to 0.6 second), which is especially suitable for autonomous driving scenarios with high real-time requirements.

[0017] 4. In complex driving scenarios (such as parking lots and areas with dense traffic signs), the method provided by the present invention significantly improves the ability to distinguish easily confused categories (such as doors, sundries, and pillars) through the deep fusion of geometric and context information. Visualization experiments show that in scenarios such as corridors and meeting rooms, the misjudgment rate of targets such as doors and bookshelves is reduced by more than 40% compared with BAAF-Net, and the recognition accuracy of dynamic targets (such as pedestrians and vehicles) is increased by 15%-20%, ensuring the safety and reliability of the navigation system.

[0018] In summary, the network proposed by the present invention has made significant breakthroughs in terms of the accuracy, efficiency, and generalization ability of point cloud semantic segmentation in large-scale scenarios through modular innovation and system optimization, providing a highly reliable environment perception solution for autonomous driving navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, where: Figure 1 is a flowchart of a lidar-based navigation method in a large-scale driving scenario provided by the present invention; Figure 2 is a diagram of the semantic association mining module provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be noted that all the directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.

[0022] In addition, the descriptions such as "first" and "second" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0023] In the present invention, unless otherwise clearly defined and limited, the terms "connected", "fixed", etc. should be understood in a broad sense. For example, "fixed" may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection; it may be directly connected, or indirectly connected through an intermediate medium, and may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0024] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0025] Please refer to Figure 1 As shown, the embodiments of the present invention provide a lidar-based navigation method in a large-scale driving scenario, including the following steps: Step S1, obtaining and inputting a lidar point cloud signal; Step S2, extracting and aggregating features of the lidar point cloud signal through the encoding layer of the encoding and decoding network to obtain local features, and then using an adaptive fusion mechanism in combination with an attention interpolation module through the decoding layer of the encoding and decoding network to perform multi-resolution fusion on local features of different scales and output high-level features; Step S3: input the high-level features into the semantic association mining module. Figure 2 The two MLP layers (as shown) embed the output high-level features to form the final output features, wherein the semantic association mining module integrates spatial information and context information through cross compensation; Step S4, performing semantic segmentation on the final output features, and planning the optimal navigation route based on the semantic segmentation results.

[0026] In step S2, the specific implementation of the attention interpolation module includes: Step S21: In the interpolation layer, the mth layer of the to-be-interpolated point set As the center, for a small-scale point set Feature map Perform a K nearest neighbor search and obtain a local representation: ; ; In the formula, KNN It is the K nearest neighbor classification algorithm; is the preliminary interpolation result; is the point cloud information of the γ module; is the parameter matrix; γ is the combination module of linear mapping and batch normalization layer; is the number of points in the mth layer to be interpolated; is a small-scale point set Additional known features included; The result feature of K value interpolation; Step S22, through the linear mapping and batch normalization layer combination module γ Mapping to generate parameter matrix , and in Dimensions are normalized by softmax and weighted summation, and interpolation results are output : ; In the formula, For Dimensions are used for operations; softmax (·)for softmax operate.

[0027] In step S3, the specific implementation of the semantic association mining module includes: Step S31, based on the neighborhood center context and coordinates ,in, is the number of points in the neighborhood center c, is the number of points; and the absolute point set context and coordinates , the relative point set context can be obtained and the relative coordinates ; Step S32: Encode the context information and geometric space information, and generate the mixed feature F′ and geometric encoding P′ through a concatenation operation and a 1×1 convolutional layer; Step S33: Aggregate the mixed features by combining max pooling and attention pooling to generate the final context feature encoding .

[0028] The encoding layer dynamically adjusts the feature extraction range through a receptive field reconstruction mechanism, and the attention interpolation module is used to improve the accuracy of high-level features.

[0029] The number of points in the m-th layer to be interpolated is , where δ m is the preset downsampling factor.

[0030] In step S32, the concatenation operation of the context encoding and the geometric encoding satisfies: ; ; ; ; where copy (·) represents a shallow copy of the encoding; concat (·) represents the concatenation operation; conv 1×1 represents a 1×1 convolution.

[0031] In step S33, the final context feature encoding is represented by the following formula: .

[0032] Furthermore, the above-mentioned mixed local aggregation method is used, that is, max pooling and attention pooling are combined: ; where is the encoding after max pooling and attention pooling aggregation; max (·) represents the maximum value calculation process; attention (·) represents the three-dimensional attention mechanism calculation process.

[0033] The semantic segmentation result is generated by a Softmax classifier, and the classification categories include roads, obstacles, vehicles, and pedestrians.

[0034] The following uses specific embodiments to elaborate in detail on the lidar-based navigation method provided by the present invention in large-scale driving scenarios.

[0035] Embodiment 1 Table 1 shows the settings of some hyperparameters when using S3DIS as the experimental dataset. N is the scale of the input point cloud.

[0036] Table 1 Network hyperparameter settings

[0037] Since there are a total of 6 regions in S3DIS, a 6-fold cross-validation experiment is conducted. The 6-fold cross-validation experiment is to traverse all 6 regions, extract one of them as the test set each time, and use the other 5 regions as the training set for training and testing. Finally, the test results of the 6 regions are aggregated together point by point to calculate an average value to obtain the final performance. The data packing strategy is a probability-based training sample selection mechanism, which can maintain the randomness of training sample acquisition.

[0038] Table 2 shows the 6-fold cross-validation experiment results of the proposed model and other similar works in recent years. The experimental results show that the method using the network of the present application is the best among the above methods in terms of overall accuracy and average cross-linking ratio, and 8 out of 13 categories exceed BAAF-Net.

[0039] Table 2 S3DIS six-fold cross-validation experiment results (%)

[0040] Table 3 shows the verification results of S3DIS region 5. As shown in Table 3, the mIoU of the present application increased by about 2.0% compared to the baseline, exceeding recent works.

[0041] Table 3 S3DIS region 5 test experiment results (%)

[0042] Embodiment 2 The SemanticKITTI validation set experiment was conducted to obtain results and compare them with recent works based on pure lidar without data augmentation and without large-scale parameter networks.

[0043] It can be analyzed from the experimental results that the network of the present application has achieved advanced performance, reaching an average intersection over union of 60% for the first time in a pure point-wise MLP-based method without using data augmentation and large-parameter models, and in 8 out of all 19 categories, it exceeds all other methods listed.

[0044] Through visual experiments, the manifestation of this advantage is more obvious. As is well known, people and vehicles are the two most crucial categories in autonomous driving scene prediction. Whether they can be correctly recognized is related to the driving safety of autonomous vehicles. The recognition effect of the proposed model on these two categories is significantly better than that of BAAF-Net. In addition, the geometric structures of parking lots and roads are relatively similar, and the proposed semantic association mining module better captures the connection between the sparse structure of the perceived area and the context information, resulting in a significant improvement in recognition performance.

[0045] Example 3 Table 4 shows the results of the module ablation experiment. Among them, K&D represents the mechanism of dense KNN and layer-by-layer downsampling. BA is the bilateral module. FP refers to the nearest neighbor interpolation decoder used by BAAF-Net. The inference time shows the time required for a batch of data to pass through the model of this application.

[0046] Table 4 Module Ablation Experiment (%)

[0047] By ablating the module to construct different sub-models to evaluate the superiority of the proposed module, as shown in Table 4. Model Ⅰ is the baseline model (BAAF-Net) without module replacement. On Model Ⅱ, the RFR mechanism is used to replace the K&D mechanism, and the mIoU is increased by 1.7%. This is due to the contribution of the RFR mechanism in reducing information loss. Model Ⅲ replaces the BA module with the proposed SCE module. The mIoU increases by 1.2% compared to the baseline, and the number of trainable parameters drops significantly to about half of the original, and the inference time also drops by 1 / 3. This is attributed to the contribution of SCE in mining the association between spatial geometric features and context features. Model Ⅳ replaces FP with the AFP module, and the mIoU steadily increases by 1.6%, thanks to the module's improvement of the model's adaptive decoding ability. However, due to the need for additional KNN operations, the inference time increases.

[0048] For Models Ⅴ to Ⅶ, two of the proposed modules are randomly replaced, and the performance is better than that of the sub-models that only replace one module. Model Ⅷ is the complete model that replaces all three modules. The mIoU and OA are both the best. The number of parameters is slightly increased compared to the model Ⅱ with the least number of parameters, but it is still significantly better than Model Ⅰ. The inference time is slightly increased compared to Model Ⅰ, which is due to the introduction of additional KNN operations. In summary, the more proposed modules are replaced, the better the segmentation performance of the model.

[0049] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.

Claims

1. A lidar-based navigation method in large-scale driving scenarios, characterized in that, It includes the following steps: Step S1, obtaining and inputting a laser point cloud signal; Step S2, extracting and aggregating features of the laser point cloud signal through an encoding layer of an encoding-decoding network to obtain local features, and then using an adaptive fusion mechanism and an attention interpolation module through a decoding layer of the encoding-decoding network to perform multi-resolution fusion on local features of different scales to output high-level features; Step S3, inputting the high-level features into a semantic association mining module, and performing embedding processing on the output high-level features through two MLP layers of the semantic association mining module to form final output features, wherein the semantic association mining module integrates spatial information and context information through cross compensation; Step S4, performing semantic segmentation on the final output features, and planning an optimal navigation route based on the semantic segmentation results.

2. The method according to claim 1, wherein In step S2, the specific implementation of the attention interpolation module includes: Step S21. In the interpolation layer, centered on the set of points to be interpolated in the m-th layer perform K-nearest neighbor search on the feature map of the small-scale point set to obtain a local representation: ​ ; ; In the formula, KNN is the K-nearest neighbor classification algorithm; is the preliminary interpolation result; is the point cloud information of the γ module; is the parameter matrix; γ is the combined module of linear mapping and batch normalization layer; is the number of points in the m-th layer to be interpolated; is the small-scale point set contains additional known features; is the result feature of K-value interpolation; Step S22, through the combined module γ of the linear mapping and batch normalization layer, is mapped to generate the parameter matrix , and softmax normalization and weighted summation are performed in the dimension to output the interpolation result : ; In the formula, is for operation in dimension; softmax (·) is softmax operation.

3. The method according to claim 2, wherein In step S3, the specific implementation of the semantic association mining module includes: Step S31, based on the neighborhood center context and coordinates , where is the number of points of the neighborhood center c, is the number of points; and the absolute point set context and coordinates , the relative point set context and the relative coordinates can be obtained; Step S32, encoding context information and geometric space information, and generating a mixed feature F′ and a geometric encoding P′ through a splicing operation and a 1×1 convolutional layer; Step S33, aggregate the mixed features by combining max pooling and attention pooling to generate the final context feature encoding .

4. The method according to claim 1, characterized in that In step S2, the encoding layer dynamically adjusts the feature extraction range through a receptive field reconstruction mechanism, and the attention interpolation module is used to improve the accuracy of high-level features.

5. The method according to claim 2, wherein The number of points in the m-th layer to be interpolated is , where δ m is the preset downsampling factor.

6. The method according to claim 3, characterized in that In step S32, the splicing operation of context encoding and geometric encoding satisfies: ; ; ; ; In the formula, copy (·) represents a shallow copy of the encoding; concat(·) represents a concatenation operation; conv 1×1 represents a 1×1 convolution.

7. The method according to claim 6, characterized in that In step S33, the final context feature encoding is represented by the following formula: 。 8. The method according to claim 6, characterized in that, The semantic segmentation results are generated by a Softmax classifier, and the classification categories include roads, obstacles, vehicles, and pedestrians.

Citation Information

Patent Citations

  • Vehicle front elevation information extraction method and system based on multi-sensor fusion

    CN116697981A

  • Lidar-assisted multi-image matching for 3-d model and sensor pose refinement

    US20100204964A1