Radar data semantic segmentation method and system for dense disadvantaged road users

CN122821118APending Publication Date: 2026-09-25CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610908131.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,在面向弱势道路使用者的感知任务中仍存在明显缺陷,主要体现在数据集与算法模型两方面

Benefits of technology

通过构建点云分支、距离视图分支和鸟瞰视图分支的并行多分支特征提取架构,以及将距离视图初始特征和鸟瞰视图初始特征分别输入至二维特征提取网络,输出深层次特征的距离视图特征图和鸟瞰视图特征图。本发明的二维特征提取网络采用编码器-解码器结构,编码器用于提取局部细节特征,解码器用于进行全局上下文建模,能够同时兼顾目标的局部几何细节保留和大范围空间依赖关系建模,显著提升了对遮挡、形态多变的弱势道路使用者目标的推理能力。在此基础上,多视角融合模块通过自适应逐点卷积(APWConv)算子对每个激光点及其邻域点进行处理,以自适应学习局部空间几何信息并融合多视角特征,不依赖人工设计的固定核点,能够通过训练自适应地学习每个点邻域的几何结构,并以可学习的方式对特征进行调制,从而更有效地捕捉稀疏、弱特征目标的空间信息,克服了传统固定融合方式适应性差的缺陷。除此之外,本发明通过构建两阶段级联精炼架构,对特征进行递进式优化,有效防止了小目标特征在网络深层中的信息丢失,强化了对行人和骑行者等关键类别的监督。综合以上几个方面,本发明能够有效应对高密度交通场景下的频繁遮挡和点云稀疏问题,显著降低漏检和误检率,大幅提升自动驾驶系统对弱势道路使用者的感知精度与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821118A_ABST
    Figure CN122821118A_ABST
Patent Text Reader

Abstract

The application discloses a radar data semantic segmentation method and system for dense weak road users, and the method comprises the following steps: acquiring original point cloud data collected by a laser radar; constructing a point cloud branch, a distance view and a bird's eye view branch, and generating initial features respectively; inputting the initial features of the distance view and the bird's eye view into a two-dimensional feature extraction network, projecting the two kinds of feature maps into a three-dimensional space, splicing the features with the point cloud branch features, and forming fused three-dimensional features; inputting the fused three-dimensional features into a multi-view fusion module, processing each laser point and its neighborhood points through an adaptive point-by-point convolution operator; constructing a first stage formed by the above steps and a second stage with the same structure, taking the output of the first stage as the input of the second stage, and gradually refining the feature representation through the cascade of the two stages; inputting the refined features output by the second stage into a full connection layer, outputting the semantic category probability distribution of each laser point, and obtaining a semantic segmentation result. The application can effectively improve the perception accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent driving environment perception technology, specifically relating to a radar data semantic segmentation method and system for densely populated vulnerable road users (VRUs). Background Technology

[0002] Environmental perception is a core bottleneck restricting the implementation of autonomous driving technology, while semantic segmentation of LiDAR point clouds is a key technology for achieving accurate identification of targets around vehicles and ensuring driving safety. This technology can assign semantic categories to each point in a 3D LiDAR point cloud, enabling reliable perception of pedestrians, cyclists, vehicles, and the road environment. Compared to visual solutions that are susceptible to lighting conditions and have limited positioning accuracy, LiDAR point clouds can directly provide stable 3D structural information, possessing irreplaceable perception advantages in complex urban road conditions and poor lighting conditions, and are a core support for improving the safety and robustness of autonomous driving.

[0003] Current point cloud semantic segmentation techniques are mostly based on deep learning, using complex network structures to extract local geometric features and global contextual information from point clouds. They often combine multi-view methods with point cloud branching to achieve a more comprehensive scene understanding. However, significant shortcomings remain in perception tasks targeting vulnerable road users, primarily in terms of datasets and algorithmic models. At the dataset level, mainstream open-source point cloud datasets generally lack densely congested scenarios, with sparse pedestrian and cyclist samples and a severely insufficient number of point clouds. Many targets contain only a few laser points, failing to provide complete geometric features and easily confused with interfering targets such as trees and streetlights, resulting in insufficient model generalization ability. At the algorithm level, existing segmentation methods struggle to simultaneously capture local details and model global context, lacking sufficient reasoning ability for occluded or morphologically variable small targets. Multi-view fusion methods often rely on fixed structures and manually set priors, exhibiting poor adaptability to sparse points and targets with weak features, and failing to effectively address information loss caused by occlusion. In high-density traffic scenarios, pedestrians and cyclists frequently obstruct traffic and targets are densely distributed. Existing algorithms generally suffer from missed detections and false detections, making it difficult to meet the actual needs of autonomous driving for highly reliable perception of vulnerable road users. Summary of the Invention

[0004] The purpose of this invention is to address the problems in the prior art by providing a radar data semantic segmentation method and system for dense vulnerable road users. This method can effectively address the problems of frequent occlusion and sparse point clouds in high-density traffic scenarios, significantly reduce the false detection and missed detection rates, and greatly improve the perception accuracy and robustness of autonomous driving systems for vulnerable road users.

[0005] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a semantic segmentation method for radar data targeting densely populated vulnerable road users is provided, including: Obtain the raw point cloud data collected by lidar to obtain the initial point cloud features; Construct point cloud branches, distance view branches, and bird's-eye view branches, and generate initial features for each branch; Initial features of the distance view and the initial features of the bird's-eye view are respectively input into a two-dimensional feature extraction network. The two-dimensional feature extraction network adopts an encoder-decoder structure. The encoder is used to extract local detail features, and the decoder is used to perform global context modeling, and outputs the distance view feature map and the bird's-eye view feature map respectively. The distance view feature map and the bird's-eye view feature map are back-projected into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, which are then stitched together with the point cloud branch features to form a fused three-dimensional feature. The fused 3D features are input into the multi-view fusion module, which processes each laser point and its neighboring points through an adaptive pointwise convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. A first stage consisting of all the above steps is constructed, and a second stage with the same structure as the first stage is constructed. The fused features output by the first stage are used as the input features of the second stage. The feature representation is refined step by step through the cascaded processing of the two stages. The refined features output by the second stage are input into a fully connected layer, and the probability distribution of each laser point belonging to a predefined semantic category is output to obtain the semantic segmentation result of the original point cloud data.

[0006] As a preferred embodiment, in the step of acquiring the raw point cloud data collected by the lidar and obtaining the initial point cloud features, the initial point cloud features are represented as follows: In the formula, N The initial point cloud features represent the total number of laser points in a frame of raw point cloud data, with each laser point having a feature dimension of 4; the initial point cloud features include the three-dimensional spatial coordinates of each laser point. p= ( x , y , z and reflection intensity value remission .

[0007] As a preferred embodiment, the construction of point cloud branches, distance view branches, and bird's-eye view branches, and the generation of initial features corresponding to each branch, includes: The initial features corresponding to the point cloud branches are generated in the following manner: The three-dimensional spatial coordinates of each laser point p= ( x , y ,z ), reflection intensity value remission and grid offset (Δ) x Δ y This forms a feature vector, which is then input into a network composed of multilayer perceptrons to obtain the point cloud branch features. F p,1 ; Initial features corresponding to distance view branches Generate in the following manner: Using the spherical projection formula, ( x , y , z , r , remission The five features are projected onto the distance view, and the expression is as follows:

[0008] In the formula, b rv and h rv These are the width and height of the distance view, respectively; u rv and v rv This represents the coordinates of the projected point in the distance view. This represents the distance of a point in space relative to the laser radar emission point; fov=fov_up+fov_down Indicates the field of view of the lidar; Initial features corresponding to the bird's-eye view branch Generate in the following manner: By filling point attributes in the bird's-eye view grid ( x , y , z , r , remission Δ x Δ y ), where Δ x and Δ y It is the coordinate offset of a point relative to the center of the bird's-eye view grid; the area of ​​the bird's-eye view plane is limited to ( x min , x max , y min , y max The coordinates of each point on the bird's-eye view plane are calculated as follows:

[0009] In the formula, b bev and hbev It is the width and height of the bird's-eye view; u bev and v bev These are the coordinates of the projection point in the bird's-eye view.

[0010] As a preferred embodiment, the encoder of the two-dimensional feature extraction network is a ResNet residual network, and the decoder is a Swin-Transformer. The Swin-Transformer uses a shift window self-attention mechanism for global context modeling. The mathematical expression for the encoder is:

[0011] In the formula, It is the output of the i-th layer encoder, 1≤i≤3; It is the initial feature corresponding to the distance view branch. S rv Or the initial features corresponding to the bird's-eye view branch S bev ; The Swin-Transformer incorporates multi-head self-attention, learning relationships between elements by constructing a matrix of query Q, key K, and value V. Given an input feature map S0, a layer normalization LN is applied to obtain the matrix of query Q, key K, and value V for a windowed multi-head self-attention W-MSA layer and a shifted windowed multi-head self-attention SW-MSA layer. The shifted window strategy restricts attention computation to non-overlapping local windows to improve efficiency. The mathematical expression for one layer of the Swin-Transformer is as follows:

[0012] Four independent attention heads perform attention calculations in parallel, as shown in the following expression:

[0013]

[0014] In the formula, Concat means splicing along the channel; B j Indicates the first j Head bias, 1≤ j ≤4; d j This is the number of channels per end; The output expression of the i-th layer decoder is as follows, where 1 ≤ i ≤ 3:

[0015] In the formula, Flatten flattens the matrix into a vector; Conv is convolution; Up is transpose convolution, used to transform... Upsampling with Same resolution.

[0016] As a preferred approach, the distance view feature map and the bird's-eye view feature map are back-projected into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, which are then stitched together with the point cloud branch features to form fused three-dimensional features. In this step, the distance view feature map obtained from the two-dimensional feature extraction network... and bird's-eye view feature map The distance is obtained using the back projection method. Figure 3 D features F rv and bird's-eye view Figure 3 D features F bev and point cloud branch features F p,1 splicing in the channel dimension yields... C 3D point features of each channel .

[0017] As a preferred embodiment, the multi-view fusion module processes each laser point and its neighboring points using an adaptive pointwise convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. The steps include: Neighborhood search and offset calculation for the current laser point p n ,1≤ n ≤ N It is defined as a center point, and its neighborhood is searched using the K-nearest neighbor algorithm. K Points and their corresponding features ; through each neighbor point p k Using the formula Δ p n,k = p k - p n Calculate the set of position offsets ; Location encoding for each center point p n Neighboring points p k The point position is offset by Δ using a multilayer perceptron. p n,k Encode to a high-dimensional feature space, using a learnable channel attention coefficient. The encoded features are scaled channel by channel; the Sigmoid function is applied to constrain the output to the (0, 1) interval to obtain the position-encoded features. Location coding features m n,k The calculation expression is:

[0018] In the formula, m n,k Indicates the center point p n Local geometric details within the neighborhood are adaptively updated throughout the training process; Feature modulation and output, resulting in position-coded features With the corresponding original neighborhood features The geometric information is injected through addition; then a learnable weight matrix is ​​used. Perform a linear transformation on the obtained features, where, C o It refers to the number of output channels; output characteristics. f n o The following expression is used to calculate:

[0019] The above adaptive pointwise convolution operator is applied to all center points to achieve the fusion of features from multiple perspectives.

[0020] As a preferred approach, the semantic segmentation losses for the first and second stages are calculated separately, and both the loss function for the first and second stages is composed of weighted cross-entropy loss. L wce and Lovasz Softmax loss L ls Combination and composition; Let the first stage prediction be pred 1. The real label is gt The first phase loss L 1. Calculate using the following formula:

[0021] Let the second stage prediction be pred 2, then the second stage loss L 2. Calculate using the following formula:

[0022] The total loss function of the network is: ; The cascaded processing is jointly supervised by two stages of loss functions to progressively refine the feature representation.

[0023] As a preferred approach, the refined features output from the second stage are input into the fully connected layer. The output dimension of the fully connected layer is equal to the total number of semantic categories. The output is converted into the probability of each point belonging to each category through the Softmax function. The category corresponding to the maximum probability is taken as the semantic label of the corresponding point, thereby generating the semantic segmentation result of the entire point cloud frame.

[0024] Secondly, a semantic segmentation system for radar data targeting densely populated vulnerable road users is provided, including: The initial point cloud feature acquisition module is used to acquire the raw point cloud data collected by the lidar and obtain the initial point cloud features; The initial feature generation module for each branch is used to construct the point cloud branch, distance view branch, and bird's-eye view branch, and generate the initial features corresponding to each branch; A two-dimensional feature extraction module is used to input the initial features of the distance view and the initial features of the bird's-eye view into the two-dimensional feature extraction network. The two-dimensional feature extraction network adopts an encoder-decoder structure. The encoder is used to extract local detail features, and the decoder is used to perform global context modeling, and outputs the distance view feature map and the bird's-eye view feature map respectively. The three-dimensional feature fusion module is used to back-project the distance view feature map and the bird's-eye view feature map into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, and then stitch them with the point cloud branch features to form fused three-dimensional features; A multi-view feature fusion module is used to input fused 3D features into the multi-view fusion module. The multi-view fusion module processes each laser point and its neighboring points through an adaptive point-by-point convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. A two-stage cascaded refinement module is used to construct a first stage consisting of all the above steps, and a second stage with the same structure as the first stage. The fused features output by the first stage are used as the input features of the second stage, and the feature representation is refined step by step through the cascaded processing of the two stages. The refined features output by the second stage are input to a fully connected layer, and the probability distribution of each laser point belonging to a predefined semantic category is output to obtain the semantic segmentation result of the original point cloud data.

[0025] Thirdly, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being executed by a processor in an electronic device to implement the radar data semantic segmentation method for densely populated vulnerable road users.

[0026] Compared with the prior art, the present invention has at least the following beneficial effects: By constructing a parallel multi-branch feature extraction architecture consisting of point cloud, distance view, and bird's-eye view branches, and inputting initial features of the distance view and bird's-eye view into a two-dimensional feature extraction network, the network outputs deep-level feature maps of the distance view and bird's-eye view. The two-dimensional feature extraction network of this invention employs an encoder-decoder structure. The encoder extracts local detail features, while the decoder performs global context modeling. This approach simultaneously preserves local geometric details and models large-scale spatial dependencies, significantly improving the inference ability for occluded and morphologically variable vulnerable road user targets. Furthermore, the multi-view fusion module processes each laser point and its neighboring points using an adaptive pointwise convolution (APWConv) operator to adaptively learn local spatial geometric information and fuse multi-view features. Without relying on manually designed fixed kernel points, it can adaptively learn the geometric structure of each point's neighborhood through training and modulate features in a learnable manner, thereby more effectively capturing the spatial information of sparse, weakly featured targets and overcoming the poor adaptability of traditional fixed fusion methods. In addition, this invention constructs a two-stage cascaded refinement architecture to progressively optimize features, effectively preventing information loss of small target features in deeper network layers and strengthening supervision of key categories such as pedestrians and cyclists. In summary, this invention effectively addresses the problems of frequent occlusion and sparse point clouds in high-density traffic scenarios, significantly reducing false negative and false positive rates, and greatly improving the perception accuracy and robustness of autonomous driving systems for vulnerable road users. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the following drawings are only some of the embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.

[0028] Figure 1 Flowchart of a radar data semantic segmentation method for densely populated vulnerable road users according to an embodiment of the present invention; Figure 2 A schematic diagram of the two-dimensional feature extraction network structure according to an embodiment of the present invention; Figure 3 A schematic diagram of the processing flow of the adaptive pointwise convolution operator in an embodiment of the present invention. Detailed Implementation

[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0030] Please see Figure 1 This invention provides a semantic segmentation method for radar data targeting densely populated vulnerable road users (including pedestrians and cyclists). Through two-stage progressive feature refinement, adaptive point convolution operators, and a CNN-Transformer hybrid encoding and decoding structure, it significantly improves the segmentation accuracy and robustness of occluded pedestrians and cyclists.

[0031] Specifically, the radar data semantic segmentation method for densely populated vulnerable road users in this embodiment of the invention mainly includes: Step 1: Point cloud data acquisition; Point cloud data was acquired using a Velodyne HDL-64E LiDAR. A single frame of raw point cloud data is represented as the initial point cloud features. ,in, N This represents the total number of laser points in a frame of original point cloud data. Each laser point has a feature dimension of 4, and the initial point cloud features contain the three-dimensional spatial coordinates of each laser point. p= ( x , y , z and reflection intensity value remission .

[0032] Step 2: Multi-view branch feature generation; Step 2.1: Generation of distance view branch features; Features of Range View (RV) branches It is obtained through the formula for spherical projection, where ( x , y , z , r , remission The expressions for projecting the five features onto the distance view are as follows:

[0033] In the formula, b rv and h rv These are the width and height of the distance view, respectively; u rv and v rv This represents the coordinates of the projected point in the distance view. This represents the distance of a point in space relative to the laser radar emission point; fov=fov_up+fov_down Indicates the field of view of the lidar; Step 2.2: Feature generation of the Bird's-Eye View (BEV) branch; Characteristics of BEV branch This is achieved by filling the BEV's grid with point attributes ( x , y , z , r , remission Δ x Δ y The formula obtained is Δ x and Δ y It is the coordinate offset of a point relative to the center of the bird's-eye view grid; the area of ​​the bird's-eye view plane is limited to ( x min , x max , y min , y max The coordinates of each point on the bird's-eye view plane are calculated as follows:

[0034] In the formula, b bev and h bev It is the width and height of the bird's-eye view; u bev and v bev These are the coordinates of the projection point in the bird's-eye view.

[0035] Step 2.3: Point cloud branch feature generation; The original features of each point, i.e., the three-dimensional spatial coordinates of each laser point. p= ( x , y , z ), reflection intensity value remission And the calculated grid offset (Δ) x Δ y This forms a 7-dimensional feature vector. This vector is then input into a network consisting of two multilayer perceptrons, which outputs high-dimensional point cloud branch features. F p,1 .

[0036] Step 3: Two-dimensional feature extraction; A hybrid CNN-Transformer encoder-decoder architecture is employed in the 2D feature extraction network to generate feature maps for the RV and BEV branches. The encoder uses ResNet, and the decoder uses Swin-Transformer (SwinT). The ResNet encoder extracts local details and preserves detail information, while the SwinT decoder uses shifted window self-attention to aggregate global context in a computationally efficient manner. The combination of these two components enhances inference capabilities for occluded and multi-shaped targets.

[0037] The mathematical expression for the encoder is:

[0038] In the formula, It is the output of the i-th layer encoder, 1≤i≤3; It is the initial feature corresponding to the distance view branch. S rv Or the initial features corresponding to the bird's-eye view branch S bev ; exist Figure 2 The decoder shown employs SwinT with Multihead Self Attention (MSA) to enhance the ability to capture contextual dependencies. Relationships between elements are learned by constructing a query, key, and value matrix (Q, K, V). Given an input feature map... S 0. Apply Layer Normalization (LN) to obtain the Q, K, and V values ​​of the Window-Multihead Self Attention (W-MSA) layer and the Shifted Window-Multihead Self Attention (SW-MSA) layer. The shifted window strategy restricts attention computation to non-overlapping local windows to improve efficiency; the mathematical expression for one layer of the Swin-Transformer is as follows:

[0039] Four independent attention heads perform attention calculations in parallel, as shown in the following expression:

[0040]

[0041] In the formula, Concat means splicing along the channel; B j Indicates the first j Head bias, 1≤ j≤4; d j That is the number of channels per head.

[0042] Furthermore, the output expression of the i-th layer decoder is as follows, 1≤i≤3:

[0043] In the formula, Flatten flattens the matrix into a vector; Conv is convolution; Up is transpose convolution, used to transform... Upsampling with Same resolution.

[0044] Step 4: Multi-view feature fusion; Step 4.1: Feature backprojection and stitching; Distance view feature map obtained from a 2D feature extraction network and bird's-eye view feature map The distance is obtained using the back projection method. Figure 3 D features F rv and bird's-eye view Figure 3 D features F bev and point cloud branch features F p,1 splicing in the channel dimension yields... C 3D point features of each channel .

[0045] Step 4.2: Adaptive pointwise convolutional fusion; In the multi-view fusion module, the Adaptive Pointwise Convolution (APWConv) operator is applied to process point features. F Compared to previously used kernel-based convolutions, APWConv does not rely on predefined kernels, thus requiring less manual parameter tuning. It uses a multilayer perceptron to adaptively learn local detail features within the neighborhood of each laser point and can continuously adjust its representation strategy for neighboring spatial information during training. Figure 3 As shown, this operator enables the network to utilize more local detailed features and improves its ability to capture neighborhood spatial information of small targets.

[0046] Step 4.2.1: Neighborhood search and offset calculation; for the current laser point p n ,1≤ n ≤ N It is defined as a center point, and its neighborhood, including itself, is searched using the K-nearest neighbor algorithm. K Points and their corresponding features ; through each neighbor point p k Using the formula Δ p n,k = p k - p n Calculate the set of position offsets ; .

[0047] Step 4.2.2: Location encoding; For each center point p n Neighboring points p k The point position is offset by Δ using a multilayer perceptron. p n,k Encode to a high-dimensional feature space, using a learnable channel attention coefficient. The encoded features are scaled channel by channel; the Sigmoid function is applied to constrain the output to the (0, 1) interval, resulting in positionally encoded features. This constraint ensures the range of the positionally encoded features, avoids overwriting the original features, and enhances training stability.

[0048] Location coding features m n,k The calculation expression is:

[0049] In the formula, m n,k Indicates the center point p n Local geometric details within the neighborhood are adaptively updated throughout the training process. Unlike manually tuned hyperparameters, m n,k Instead of being predefined, it is adaptively updated throughout the training process, making it more robust.

[0050] Step 4.2.3: Feature Modulation and Output The obtained location encoding features With the corresponding original neighborhood features The geometric information is injected through addition; then a learnable weight matrix is ​​used. Perform a linear transformation on the obtained features, where, C o It is the number of output channels; finally, the output features. f n o The following expression is used to calculate:

[0051] The above adaptive pointwise convolution operator is applied to all center points to achieve the fusion of features from multiple perspectives.

[0052] Step 5: Two-stage cascading and loss calculation; A two-stage cascaded refinement architecture is constructed; the processing flow consisting of steps 1 to 4 is defined as the first stage. A second stage with the exact same structure as the first stage is constructed, and the fused features output from the first stage are used as the input features of the second stage. The semantic segmentation loss of the first and second stages is calculated separately, and the feature representation is gradually refined through the cascaded processing of the two stages. The losses for these two stages are calculated separately so that the second stage can refine the features extracted in the first stage and enhance the overall feature representation capability. Since the number of points for roads and buildings is tens or hundreds of times greater than that for pedestrians and bicycle lanes, a weighted cross-entropy loss is used. L wce and Lovasz Softmax loss L ls To balance the different categories.

[0053] Let the first stage prediction be pred 1. The real label is gt The first phase loss L 1. Calculate using the following formula:

[0054] Let the second stage prediction be pred 2, then the second stage loss L 2. Calculate using the following formula:

[0055] The total loss function of the network is: ; Step 6: Output the semantic segmentation results; The refined feature map output from the second stage F The input is fed into a fully connected layer, whose output dimension is equal to the total number of semantic categories (such as background, pedestrian, cyclist, car, etc.). The output is transformed into the probability of each point belonging to each category through the Softmax function. The category corresponding to the maximum probability is taken as the final semantic label of that point, generating the semantic segmentation result of the entire point cloud frame.

[0056] This invention employs a hybrid CNN-Transformer encoder-decoder structure in the 2D feature extraction network to generate feature maps for the RV and BEV branches. Compared to traditional pure CNN or pure Transformer structures, the hybrid CNN-Transformer structure can preserve the local geometric information of the target while modeling large-scale spatial dependencies, thus improving the segmentation accuracy of occluded targets. In the multi-view fusion module, an adaptive pointwise convolution operator is applied to process point features. The APWConv operator does not require the design of manual kernel points, adaptively learns neighborhood geometric information, and can continuously adjust its representation strategy for neighboring spatial information during training, enabling the network to fully learn spatial detail features and improve its ability to capture the neighboring spatial information of small targets. The two-stage progressive architecture of this invention can strengthen the supervision of small targets and prevent the loss of information of small targets in deep networks. The dedicated dense VRU dataset solves the defects of sparse and few samples in public datasets, improving the generalization ability in real congestion scenarios. The entire semantic segmentation method can effectively improve the segmentation accuracy of occluded pedestrian and cyclist targets under high-density traffic conditions.

[0057] In another embodiment of the present invention, the hyperparameters, optimization strategies, and data augmentation schemes for network training are further defined. During the model training phase, stochastic gradient descent is used as the optimizer, with weight decay set to 0.0001 and momentum set to 0.9. The initial learning rate is set to 0.001, and cosine annealing is used for learning rate adjustment, with a warm-up round. The total number of training rounds is 48. The distance view resolution is set to 64×2048, and the bird's-eye view resolution is set to 600×600. The K value in the K-nearest neighbor algorithm is set to 7. Data augmentation employs a multi-strategy approach, specifically including: random point cloud translation, random rotation, random scaling, random point discarding, and random flipping along the X / Y / Z axes, to improve the model's generalization ability in scenarios with dense occlusion, sparse points, and variable target poses. These parameter settings ensure the stability and convergence speed of model training.

[0058] This invention also proposes a radar data semantic segmentation system for densely populated vulnerable road users, comprising: The initial point cloud feature acquisition module is used to acquire the raw point cloud data collected by the lidar and obtain the initial point cloud features; The initial feature generation module for each branch is used to construct the point cloud branch, distance view branch, and bird's-eye view branch, and generate the initial features corresponding to each branch; A two-dimensional feature extraction module is used to input the initial features of the distance view and the initial features of the bird's-eye view into the two-dimensional feature extraction network. The two-dimensional feature extraction network adopts an encoder-decoder structure. The encoder is used to extract local detail features, and the decoder is used to perform global context modeling, and outputs the distance view feature map and the bird's-eye view feature map respectively. The three-dimensional feature fusion module is used to back-project the distance view feature map and the bird's-eye view feature map into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, and then stitch them with the point cloud branch features to form fused three-dimensional features; A multi-view feature fusion module is used to input fused 3D features into the multi-view fusion module. The multi-view fusion module processes each laser point and its neighboring points through an adaptive point-by-point convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. A two-stage cascaded refinement module is used to construct a first stage consisting of all the above steps, and a second stage with the same structure as the first stage. The fused features output by the first stage are used as the input features of the second stage, and the feature representation is refined step by step through the cascaded processing of the two stages. The refined features output by the second stage are input to a fully connected layer, and the probability distribution of each laser point belonging to a predefined semantic category is output to obtain the semantic segmentation result of the original point cloud data.

[0059] Another embodiment of the present invention also proposes a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the radar data semantic segmentation method for densely populated vulnerable road users.

[0060] For example, the instructions stored in the memory can be divided into one or more modules / units. These modules / units are stored in a computer-readable storage medium and executed by the processor to complete the radar data semantic segmentation method for densely populated vulnerable road users according to the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program on the server.

[0061] The electronic device may be a smartphone, laptop, PDA, or cloud server, among other computing devices. It may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the electronic device may also include more or fewer components, or combinations of certain components, or different components; for example, it may also include input / output devices, network access devices, buses, etc.

[0062] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0063] The memory can be an internal storage unit of the server, such as the server's hard drive or memory. The memory can also be an external storage device of the server, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or FlashCard.

[0064] Furthermore, the memory may include both internal storage units of the server and external storage devices. The memory is used to store the computer-readable instructions and other programs and data required by the server. The memory can also be used to temporarily store data that has been output or will be output.

[0065] It should be noted that the information interaction and execution process between the above-mentioned module units are based on the same concept as the method embodiment. For details on their specific functions and technical effects, please refer to the method embodiment section. They will not be repeated here.

[0066] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0067] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0068] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0069] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A semantic segmentation method for radar data targeting densely populated vulnerable road users, characterized in that, include: Obtain the raw point cloud data collected by lidar to obtain the initial point cloud features; Construct point cloud branches, distance view branches, and bird's-eye view branches, and generate initial features for each branch; Initial features of the distance view and the initial features of the bird's-eye view are respectively input into a two-dimensional feature extraction network. The two-dimensional feature extraction network adopts an encoder-decoder structure. The encoder is used to extract local detail features, and the decoder is used to perform global context modeling, and outputs the distance view feature map and the bird's-eye view feature map respectively. The distance view feature map and the bird's-eye view feature map are back-projected into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, which are then stitched together with the point cloud branch features to form a fused three-dimensional feature. The fused 3D features are input into the multi-view fusion module, which processes each laser point and its neighboring points through an adaptive pointwise convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. A first stage consisting of all the above steps is constructed, and a second stage with the same structure as the first stage is constructed. The fused features output by the first stage are used as the input features of the second stage. The feature representation is gradually refined through the cascaded processing of the two stages. Furthermore, the refined features output from the second stage are input into a fully connected layer, which outputs the probability distribution of each laser point belonging to a predefined semantic category, thus obtaining the semantic segmentation result of the original point cloud data.

2. The semantic segmentation method for radar data targeting densely populated vulnerable road users according to claim 1, characterized in that, In the step of acquiring the raw point cloud data collected by the lidar and obtaining the initial point cloud features, the initial point cloud features are represented as follows: In the formula, N The initial point cloud features represent the total number of laser points in a frame of raw point cloud data, with each laser point having a feature dimension of 4; the initial point cloud features include the three-dimensional spatial coordinates of each laser point. p= ( x , y , z and reflection intensity value remission .

3. The radar data semantic segmentation method for densely populated vulnerable road users according to claim 2, characterized in that, The construction of point cloud branches, distance view branches, and bird's-eye view branches, and the generation of initial features corresponding to each branch, include: The initial features corresponding to the point cloud branches are generated in the following manner: The three-dimensional spatial coordinates of each laser point p= ( x , y , z ), reflection intensity value remission and grid offset (Δ) x Δ y This forms a feature vector, which is then input into a network composed of multilayer perceptrons to obtain the point cloud branch features. F p,1 ; Initial features corresponding to distance view branches Generate in the following manner: Using the spherical projection formula, ( x , y , z , r , remission The five features are projected onto the distance view, and the expression is as follows: In the formula, b rv and h rv These are the width and height of the distance view, respectively; u rv and v rv This represents the coordinates of the projected point in the distance view. This indicates the distance of a point in space relative to the laser radar emission point; fov=fov_up+fov_down Indicates the field of view of the lidar; Initial features corresponding to the bird's-eye view branch Generate in the following manner: By filling point attributes in the bird's-eye view grid ( x , y , z , r , remission Δ x Δ y ), where Δ x and Δ y It is the coordinate offset of a point relative to the center of the bird's-eye view grid; the area of ​​the bird's-eye view plane is limited to ( x min , x max , y min , y max The coordinates of each point on the bird's-eye view plane are calculated as follows: In the formula, b bev and h bev It is the width and height of the bird's-eye view; u bev and v bev These are the coordinates of the projection point in the bird's-eye view.

4. The radar data semantic segmentation method for densely populated vulnerable road users according to claim 3, characterized in that, The encoder of the two-dimensional feature extraction network is a ResNet residual network, and the decoder is a Swin-Transformer. The Swin-Transformer uses a shift window self-attention mechanism for global context modeling. The mathematical expression for the encoder is: In the formula, It is the output of the i-th layer encoder, 1≤i≤3; It is the initial feature corresponding to the distance view branch. S rv Or the initial features corresponding to the bird's-eye view branch S bev ; The Swin-Transformer incorporates multi-head self-attention, learning relationships between elements by constructing a matrix of query Q, key K, and value V. Given an input feature map S0, a layer normalization LN is applied to obtain the matrix of query Q, key K, and value V for a windowed multi-head self-attention W-MSA layer and a shifted windowed multi-head self-attention SW-MSA layer. The shifted window strategy restricts attention computation to non-overlapping local windows to improve efficiency. The mathematical expression for one layer of the Swin-Transformer is as follows: Four independent attention heads perform attention calculations in parallel, as shown in the following expression: In the formula, Concat means splicing along the channel; B j Indicates the first j Head bias, 1≤ j ≤4; d j This is the number of channels per end; The output expression of the i-th layer decoder is as follows, where 1 ≤ i ≤ 3: In the formula, Flatten flattens the matrix into a vector; Conv is convolution; Up is transpose convolution, used to transform... Upsampling with Same resolution.

5. The semantic segmentation method for radar data targeting densely populated vulnerable road users according to claim 4, characterized in that, In the step of back-projecting the distance view feature map and the bird's-eye view feature map into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, and then stitching them with the point cloud branch features to form a fused three-dimensional feature map, the distance view feature map obtained from the two-dimensional feature extraction network is... and bird's-eye view feature map The distance view 3D features are obtained using the back projection method. F rv and bird's-eye view 3D features F bev and point cloud branch features F p,1 splicing in the channel dimension yields... C 3D point features of each channel .

6. The semantic segmentation method for radar data targeting densely populated vulnerable road users according to claim 5, characterized in that, The multi-view fusion module processes each laser point and its neighboring points using an adaptive point-by-point convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. The steps include: Neighborhood search and offset calculation for the current laser point p n ,1≤ n ≤ N It is defined as a center point, and its neighborhood is searched using the K-nearest neighbor algorithm. K Points and their corresponding features ; through each neighbor point p k Using the formula Δ p n,k = p k - p n Calculate the set of position offsets ; Location encoding for each center point p n Neighboring points p k The point position is offset by Δ using a multilayer perceptron. p n,k Encode to a high-dimensional feature space, using a learnable channel attention coefficient. The encoded features are scaled channel by channel; the Sigmoid function is applied to constrain the output to the (0, 1) interval to obtain the position-encoded features. Location coding features m n,k The calculation expression is: In the formula, m n,k Indicates the center point p n Local geometric details within the neighborhood are adaptively updated throughout the training process; Feature modulation and output, resulting in position-coded features With the corresponding original neighborhood features The geometric information is injected through addition; then a learnable weight matrix is ​​used. Perform a linear transformation on the obtained features, where, C o It refers to the number of output channels; output characteristics. f n o The following expression is used to calculate: The above adaptive pointwise convolution operator is applied to all center points to achieve the fusion of features from multiple perspectives.

7. The semantic segmentation method for radar data targeting densely populated vulnerable road users according to claim 1, characterized in that, The semantic segmentation losses for the first and second stages are calculated separately. Both the loss function for the first stage and the loss function for the second stage are derived from the weighted cross-entropy loss. L wce and Lovasz Softmax loss L ls Combination and composition; Let the first stage prediction be pred 1. The real label is gt The first phase loss L 1. Calculate using the following formula: Let the second stage prediction be pred 2, then the second stage loss L 2. Calculate using the following formula: The total loss function of the network is: ; The cascaded processing is jointly supervised by two stages of loss functions to progressively refine the feature representation.

8. The semantic segmentation method for radar data targeting densely populated vulnerable road users according to claim 1, characterized in that, In the step of inputting the refined features output from the second stage into the fully connected layer, the output dimension of the fully connected layer is equal to the total number of semantic categories. The output is converted into the probability of each point belonging to each category through the Softmax function. The category corresponding to the maximum probability is taken as the semantic label of the corresponding point, thereby generating the semantic segmentation result of the entire point cloud frame.

9. A semantic segmentation system for radar data targeting densely populated vulnerable road users, characterized in that, include: The initial point cloud feature acquisition module is used to acquire the raw point cloud data collected by the lidar and obtain the initial point cloud features; The initial feature generation module for each branch is used to construct the point cloud branch, distance view branch, and bird's-eye view branch, and generate the initial features corresponding to each branch; A two-dimensional feature extraction module is used to input the initial features of the distance view and the initial features of the bird's-eye view into the two-dimensional feature extraction network. The two-dimensional feature extraction network adopts an encoder-decoder structure. The encoder is used to extract local detail features, and the decoder is used to perform global context modeling, and outputs the distance view feature map and the bird's-eye view feature map respectively. The three-dimensional feature fusion module is used to back-project the distance view feature map and the bird's-eye view feature map into three-dimensional space to obtain three-dimensional distance view reconstruction features and three-dimensional bird's-eye view reconstruction features, and then stitch them with the point cloud branch features to form fused three-dimensional features; A multi-view feature fusion module is used to input fused 3D features into the multi-view fusion module. The multi-view fusion module processes each laser point and its neighboring points through an adaptive point-by-point convolution operator to adaptively learn local spatial geometric information and fuse multi-view features. A two-stage cascaded refinement module is used to construct a first stage consisting of all the above steps, and a second stage with the same structure as the first stage. The fused features output by the first stage are used as the input features of the second stage, and the feature representation is gradually refined through the cascaded processing of the two stages. Furthermore, the refined features output from the second stage are input into a fully connected layer, which outputs the probability distribution of each laser point belonging to a predefined semantic category, thus obtaining the semantic segmentation result of the original point cloud data.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in an electronic device to implement the radar data semantic segmentation method for densely populated vulnerable road users as described in any one of claims 1 to 8.