Train forward clearance target detection method and system based on point cloud feature aggregation enhancement
Through the detection method based on point cloud feature aggregation enhancement, the problem of insufficient accuracy and real-time performance of traditional detection methods is solved, and more efficient forward clearance target detection of trains is achieved, improving detection accuracy and safety.
Patent Information
- Application Number
- CN202411799289.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-09
AI Technical Summary
Traditional train forward clearance detection methods are insufficient in accuracy and real-time, and are susceptible to environmental factors, making it difficult to effectively detect and identify potential obstacles.
The detection method based on point cloud feature aggregation enhancement is adopted, and the vehicle-mounted point cloud data is obtained, and multi-scale voxel features are extracted using 3D sparse convolution, and feature aggregation and fusion are performed through ensemble sampling and BEV multi-level auxiliary branch feature extraction modules. Finally, the center of mass offset loss of the 3D regression box is introduced for target classification and three-dimensional bounding box regression.
It significantly improves the accuracy and efficiency of forward clearance target detection of trains, can more accurately identify and classify potential obstacles, and provides strong technical support for the safe operation of trains.
Smart Images

Figure CN119992504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of train forward clearance target detection, and in particular to a train forward clearance target detection method and system based on point cloud feature aggregation enhancement. Background Art
[0002] In modern rail transit systems, train safety is the primary task to ensure operational efficiency and passenger safety. With the rapid development of urban rail transit, potential obstacles such as pedestrians, vehicles, and animals faced by trains during operation may cause train collisions, seriously affecting operational safety. Therefore, it is particularly important to carry out the detection and identification of train forward clearance targets.
[0003] Traditional forward clearance detection methods mainly rely on technologies such as manual monitoring and video surveillance, which are not accurate and real-time enough and are easily affected by environmental factors such as lighting and weather. With the gradual maturity of laser radar (LiDAR) and point cloud technology, these problems have been effectively solved. LiDAR can obtain three-dimensional information of the surrounding environment with high precision and generate rich point cloud data. These point cloud data not only contain the position and shape information of objects in the scene, but also provide rich reflection intensity and depth information.
[0004] In order to make full use of point cloud data for target detection, point cloud processing technology based on deep learning has developed rapidly in recent years. Point cloud feature extraction and aggregation technology have become the key to improving detection accuracy. Among them, the voxel feature encoding method converts point cloud data into a regular grid, which can effectively reduce the computational complexity while maintaining spatial structure information. Through the 3D sparse convolutional network, efficient feature extraction and downsampling can be achieved to obtain multi-scale feature representation. However, traditional feature aggregation methods often have the problem of information loss when processing multi-level features, and it is difficult to fully utilize feature information at different scales. Therefore, improving the three-dimensional target detection capability is of great significance for the detection of obstacles in the forward clearance of trains. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for train forward clearance target detection based on point cloud feature aggregation enhancement to solve at least one technical problem existing in the above-mentioned background technology.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for detecting a train forward clearance target based on point cloud feature aggregation enhancement, comprising:
[0008] Obtain vehicle-mounted point cloud data and build a scene point cloud dataset;
[0009] The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0010] Use set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregate the multi-scale voxel features to the voxel center point;
[0011] A fully sparse and highly compressed method is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps, and BEV feature maps of different resolutions are fused through 2D sparse convolution processing and skip connections.
[0012] The centroid shift loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clear object classification and 3D bounding box regression.
[0013] As a further limitation of the first aspect of the present invention, the sparse point cloud is divided into a regular voxel grid, the voxel feature encoding converts the point cloud information into a compact feature representation, and the multi-scale voxel features are extracted by downsampling layer by layer using 3D sparse convolution, including: voxel gridding: inputting N spatial point clouds P = (p1, p2, ..., p n ) area is divided into voxel units with a spatial resolution of L×W×H, where the point cloud features represent (x, y, z, r), i.e., the 3D coordinates and reflection intensity of the point cloud; voxel feature encoding: the point cloud features in the voxel unit are averaged to obtain the initial encoded voxel features, which are represented as a four-dimensional tensor; the voxel features are downsampled layer by layer using 3D submanifold sparse convolution, and the features of non-empty voxels are encoded as multi-scale three-dimensional sparse feature quantities with downsampling ratios of 1x, 2x, 4x, and 8x respectively; where the spatial coordinates of the voxel downsampled at the kth layer are represented as The set of voxel features is expressed as Where N k is the number of non-empty voxels in the kth layer, with the superscript l k (k=1,2,3,4) represents the downsampling layers with 1x, 2x, 4x, and 8x magnifications.
[0014] As a further limitation of the first aspect of the present invention, the multi-level voxel features extracted in step S2 are sampled, grouped and feature extracted using set sampling, and the multi-scale voxel features are aggregated to the voxel center point, including: point query feature grouping: the 8x downsampling layer is used as an aggregation layer after conventional 3D sparse convolution feature encoding, and the center of the voxel unit in this layer is used as the query point Q of the set sampling, and the spatial coordinates of the query point are calculated according to the voxel index and the actual size of the voxel; point query clustering: taking each query point as the center, according to the set k-th layer radius r kFind all non-empty voxels in the layer area, obtain the local grouped non-empty voxel set, and calculate the voxel feature vector set of the k-th layer; point query feature aggregation: use set sampling to extract the features of the voxels of the k-th layer.
[0015] As a further limitation of the first aspect of the present invention, two r of different scales are used in the k layer. k Increase the local variable receptive field and aggregate the voxel features of the initialization voxel layer and all downsampling layers into the query point q i Multi-scale voxel features.
[0016] As a further limitation of the first aspect of the present invention, the fully sparse highly compressed method is used to map the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and the BEV feature maps of different resolutions are fused through 2D sparse convolution processing and jump connection, including: spatial multi-rate feature mapping: using a fully sparse highly compressed operation to map the multi-scale non-empty voxel features to an XY plane grid, and accumulating the features at the same Z position, and finally generating a BEV sparse feature map with 1x, 2x, 4x and 8x downsampling ratios; BEV multi-level feature extraction: extracting the obtained 1 The BEV sparse feature map with x-multiple ratio is used to perform three-level 2D sparse convolution downsampling operations, and BEV sparse convolution features with 2x, 4x and 8x downsampling ratios are extracted layer by layer; Hierarchical feature fusion mechanism: In the downsampling process, a hierarchical feature fusion mechanism is introduced to fuse multi-rate BEV features layer by layer in the order of 1x→2x→4x→8x, extract high-resolution global semantic information, capture more detailed contextual information, and obtain 8x aggregated features with rich contextual information; the obtained 8x aggregated features are converted into 2D dense features and used as the input of the detection head for target classification and positioning tasks.
[0017] As a further limitation of the first aspect of the present invention, the centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss, guiding the detection head to achieve train forward clearance target classification and three-dimensional bounding box regression, including: a 2D feature extraction network is responsible for further processing 2D dense features and generating a final detection output, including two sub-networks: a top-down sub-network performs a 1x1 convolution and two down-samplings on the 2D dense features, and the other sub-network performs upsampling in the opposite direction to restore the spatial resolution of the input, and finally aggregates the features of all layers and the original input features into the detection head to achieve train forward clearance target classification and three-dimensional bounding box regression.
[0018] In a second aspect, the present invention provides a train forward clearance target detection system based on point cloud feature aggregation enhancement, comprising:
[0019] The acquisition module is used to obtain vehicle-mounted point cloud data and build a scene point cloud dataset;
[0020] The downsampling extraction module is used to divide the sparse point cloud into a regular voxel grid. The voxel feature encoding converts the point cloud information into a compact feature representation, and uses 3D sparse convolution to extract multi-scale voxel features by downsampling layer by layer;
[0021] A multi-scale voxel feature aggregation module based on point query, which is used to sample, group and extract features of the extracted multi-level voxel features using set sampling, and aggregate the multi-scale voxel features to the voxel center point;
[0022] BEV multi-level auxiliary branch feature extraction module, which is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps using a fully sparse and highly compressed method, and fuse BEV feature maps of different resolutions through 2D sparse convolution processing and skip connections;
[0023] The detection module is used to introduce the centroid shift loss of the 3D regression box to optimize the multi-task loss and guide the detection head to achieve train forward clearance object classification and 3D bounding box regression.
[0024] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the train forward clearance target detection method based on point cloud feature aggregation enhancement as described in the first aspect is implemented.
[0025] In a fourth aspect, the present invention provides a computer device comprising a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the train forward clearance target detection method based on point cloud feature aggregation enhancement as described in the first aspect.
[0026] In a fifth aspect, the present invention provides an electronic device, comprising: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the train forward clearance target detection method based on point cloud feature aggregation enhancement as described in the first aspect.
[0027] The beneficial effects of the present invention are as follows: the on-board point cloud data is acquired by using the laser radar installed at the front end of the train to construct a scene point cloud data set; the sparse point cloud is divided into a regular voxel grid, and converted into a compact feature representation through voxel feature encoding, and 3D sparse convolution is used to downsample layer by layer to extract multi-scale voxel features; a multi-scale voxel feature aggregation module based on point query uses set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregates the multi-scale features to the voxel center point; a BEV multi-level auxiliary branch feature extraction module uses a completely sparse and highly compressed method to map the multi-scale voxel features into a BEV feature map, and fuses feature maps of different resolutions through 2D sparse convolution processing and jump connection; the centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss to guide the detection head to realize the classification of the train's forward clearance target and the regression of the three-dimensional bounding box; the accuracy and efficiency of the train's forward clearance target detection are significantly improved, providing strong technical support for the safe operation of the train and showing good application prospects.
[0028] Additional advantages of the present invention will be more clearly given in the following description or learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0030] Figure 1 This is a flow chart of the train forward clearance target detection method based on point cloud feature aggregation enhancement described in an embodiment of the present invention.
[0031] Figure 2 Schematic diagram of the experimental platform described in an embodiment of the present invention.
[0032] Figure 3 This is a framework diagram of a multi-scale voxel feature aggregation module based on point query according to an embodiment of the present invention.
[0033] Figure 4 This is a framework diagram of the BEV multi-level auxiliary branch feature extraction module described in an embodiment of the present invention.
[0034] Figure 5 This is a diagram of the detection effect described in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below by the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0036] It should be understood by those skilled in the art that unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.
[0037] It should also be understood that terms, such as those defined in commonly used dictionaries, should be understood to have a meaning consistent with that in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless as defined herein.
[0038] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or groups thereof.
[0039] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. Different embodiments or examples described in this specification and features of different embodiments or examples may be combined and combined by those skilled in the art without contradiction.
[0040] To facilitate understanding of the present invention, the present invention is further explained below with reference to specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0041] Those skilled in the art should understand that the drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily necessary for implementing the present invention.
[0042] The present invention relates to a method and system for detecting a train forward clearance target based on point cloud feature aggregation enhancement. The method first uses a laser radar installed at the front end of the train to obtain on-board point cloud data and construct a scene point cloud data set. Then, the sparse point cloud is divided into a regular voxel grid, and converted into a compact feature representation through voxel feature encoding, and 3D sparse convolution is used to downsample layer by layer to extract multi-scale voxel features. In addition, a multi-scale voxel feature aggregation module based on point query is designed, and the extracted multi-level voxel features are sampled, grouped and feature extracted by set sampling, and the multi-scale features are aggregated to the voxel center point. The present invention also proposes a BEV multi-level auxiliary branch feature extraction module, which uses a completely sparse and highly compressed method to map the multi-scale voxel features into a BEV feature map, and fuses feature maps of different resolutions through 2D sparse convolution processing and jump connection. Finally, the centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss to guide the detection head to achieve the classification of the train forward clearance target and the regression of the three-dimensional bounding box. This method significantly improves the accuracy and efficiency of train forward clearance target detection, provides strong technical support for safe train operation, and shows good application prospects.
[0043] Example 1
[0044] In this embodiment 1, a train forward clearance target detection system based on point cloud feature aggregation enhancement is first provided, including: an acquisition module, used to acquire on-board point cloud data and construct a scene point cloud data set; a downsampling extraction module, used to divide the sparse point cloud into a regular voxel grid, voxel feature encoding converts the point cloud information into a compact feature representation, and uses 3D sparse convolution to downsample layer by layer to extract multi-scale voxel features; a point query-based multi-scale voxel feature aggregation module, used to sample, group and extract features from the extracted multi-level voxel features using set sampling, and aggregate the multi-scale voxel features to the voxel center point; a BEV multi-level auxiliary branch feature extraction module, used to map the extracted hierarchical multi-scale voxel features into a BEV feature map using a completely sparse and highly compressed method, and fuse BEV feature maps of different resolutions through 2D sparse convolution processing and jump connection; a detection module, used to introduce the centroid offset loss of the 3D regression box to optimize the multi-task loss, and guide the detection head to achieve train forward clearance target classification and three-dimensional bounding box regression.
[0045] In this embodiment, the above system is used to implement a train forward clearance target detection method based on point cloud feature aggregation enhancement, including the following steps:
[0046] Step S1: Obtain vehicle-mounted point cloud data through the laser radar installed at the front end of the train to construct a scene point cloud dataset;
[0047] Step S2: The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0048] Step S3: Design a multi-scale voxel feature aggregation module based on point query, use set sampling to sample, group and extract features from the multi-level voxel features extracted in step S2, and aggregate the multi-scale voxel features to the voxel center point;
[0049] Step S4: Design a BEV multi-level auxiliary branch feature extraction module, use a fully sparse and highly compressed method to map the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and fuse BEV feature maps of different resolutions through 2D sparse convolution processing and jump connection;
[0050] Step S5: The centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clearance object classification and 3D bounding box regression.
[0051] The step S2 described in which the sparse point cloud is divided into a regular voxel grid, the voxel feature encoding converts the point cloud information into a compact feature representation, and the 3D sparse convolution is used to downsample layer by layer to extract multi-scale voxel features, including the following steps:
[0052] Step S21: Voxel rasterization: The input N spatial point clouds P = (p1, p2, ..., p n ) area is divided into voxel units with a spatial resolution of L×W×H, where the point cloud features represent (x, y, z, r), i.e., the 3D coordinates and reflection intensity of the point cloud.
[0053] Step S22: Voxel feature encoding: average the point cloud features in the voxel unit to obtain the initial encoded voxel features, which are expressed as a four-dimensional tensor: The corresponding spatial coordinates are expressed as Where N0 is the number of non-empty voxels, and the superscript l0 indicates the initialization voxel layer.
[0054] Step S23: Use 3D submanifold sparse convolution to downsample the voxel features of step S22 layer by layer, and encode the features of non-empty voxels into multi-scale three-dimensional sparse feature quantities with downsampling ratios of 1x, 2x, 4x, and 8x respectively. The spatial coordinates of the voxel sampled at the kth layer are expressed as The set of voxel features is expressed as Where N k is the number of non-empty voxels in the kth layer, with the superscript l k (k=1,2,3,4) represents the downsampling layers with 1x, 2x, 4x, and 8x magnifications.
[0055] The multi-scale voxel feature aggregation module based on point query is designed in step S3, which uses set sampling to sample, group and extract features of the multi-level voxel features extracted in step S2, and aggregates the multi-scale voxel features to the voxel center point, including the following steps:
[0056] Step S31: Point query feature grouping: The 8x downsampling layer in step S23 is encoded with conventional 3D sparse convolutional features as the aggregation layer. The center of the voxel unit in this layer is used as the query point Q of the set sampling (SA). The spatial coordinates of the query point Q = {q1, ..., q m}, where m is The set of voxel feature vectors of this layer is expressed as The superscript q indicates the query layer.
[0057] Step S32: Point query clustering: Take each query point q i (i=1,...,m) as the center, according to the set k-th layer radius r k Find all non-empty voxels in the layer domain and obtain the local grouped non-empty voxel set. The voxel feature vector set of the kth layer is expressed as The calculation formula is as follows:
[0058]
[0059] In the formula, Indicates that in the neighborhood r k Collect the 16 nearest voxels and connect their local relative coordinates With features
[0060] Step S33: Point query feature aggregation: Use Set Abstraction (SA) (through MLP multi-layer perceptron and maximum pooling operation) to extract the features of the voxels in the kth layer:
[0061]
[0062] In the formula, A(.) represents a three-layer MLP network to encode features, and multi-scale features are aggregated into query points through the maximum pooling operation MAX(.) along the channel. In addition, we use two r of different scales in the k layer k Increase the local variable receptive field. Aggregate the voxel features of the initialization voxel layer and all downsampling layers into the query point q i The multi-scale voxel features f i (qv) , the aggregation formula is as follows:
[0063]
[0064] Where i = 1, ..., m, Represents the feature concatenation operation; the generated feature f i (qv) It contains both the multi-scale downsampling features of 3D CNN and the voxel initialization features (i.e. 3D coordinates and reflection intensity).
[0065] The BEV multi-level auxiliary branch feature extraction module described in step S4 uses a completely sparse and highly compressed method to map the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and fuses BEV feature maps of different resolutions through 2D sparse convolution processing and jump connections, including the following steps:
[0066] Step S41: Spatial multi-rate feature mapping: A fully sparse height compression (SHC) operation is used to map the multi-scale non-empty voxel features in step S23 onto an XY plane grid, and the features at the same Z position are accumulated to finally generate BEV sparse feature maps with 1x, 2x, 4x, and 8x downsampling rates.
[0067] Step S42: BEV multi-level feature extraction: 1x BEV sparse feature map obtained in step S41 Perform three-level 2D sparse convolution downsampling operations to extract BEV sparse convolution features with 2x, 4x, and 8x downsampling ratios layer by layer
[0068] Step S43: Hierarchical feature fusion mechanism: In the downsampling process, a hierarchical feature fusion mechanism is introduced to fuse the multi-rate BEV features step by step in the order of 1x→2x→4x→8x Extract high-resolution global semantic information, capture more detailed contextual information, and obtain 8x aggregated features with rich contextual information.
[0069] Step S44: Convert the 8x aggregated features obtained in step S43 into 2D dense features and use them as input to the detection head for target classification and positioning tasks.
[0070] The centroid shift loss of the 3D regression box is introduced as described in step S5 to optimize the multi-task loss, guiding the detection head to achieve train-forward clear object classification and 3D bounding box regression.
[0071] Step S51: The 2D feature extraction network is responsible for further processing the 2D dense features and generating the final detection output, which includes two sub-networks: a top-down sub-network performs a 1x1 convolution and two down-samplings (16x, 32x) on the 2D dense features, and the other sub-network performs up-sampling in the opposite direction to restore the spatial resolution of the input (8x). Finally, the features of all layers and the original input features are aggregated and input into the detection head to realize the forward clearance target classification and 3D bounding box regression of the train.
[0072] Step S52: A multi-task loss function is used during the model convergence process, including classification loss, positioning loss, orientation loss, and centroid offset loss L of the 3D regression box. centroid The calculation is as follows:
[0073]
[0074] In the formula, C box , C gt are the average coordinates of the spatial points in the 3D predicted bounding box and the ground truth box, respectively. γ and τ are the local densities of each point estimated by the kernel density estimation (KDE) method. We introduce the L2 norm ‖.‖2 to calculate the offset between the center of mass of the 3D predicted regression box and the center of mass of the ground truth box.
[0075] The overall loss function is defined as:
[0076]
[0077] Where, L cls is the classification loss, L loc is the positioning loss, L dir is the direction loss, N pos Represents the number of positive anchors, where λ1=1, λ2=2, λ3=0.2, and λ4=1 are weight hyperparameters used to balance the losses of each part.
[0078] Example 2
[0079] In this embodiment, a method and system for detecting train forward clearance targets based on point cloud feature aggregation enhancement are proposed. The method first uses a laser radar installed at the front of the train to obtain on-board point cloud data through a series of steps, thereby constructing a scene point cloud dataset. Then, the sparse point cloud is divided into a regular voxel grid, and the point cloud information is converted into a compact feature representation using voxel feature encoding, and multi-scale voxel features are extracted layer by layer through 3D sparse convolution. In addition, a multi-scale voxel feature aggregation module based on point query is designed, and the extracted multi-level voxel features are sampled, grouped and feature extracted through set sampling, and the multi-scale voxel features are aggregated to the voxel center point. In this embodiment, a BEV multi-level auxiliary branch feature extraction module is also proposed, which uses a completely sparse and highly compressed method to map the hierarchical multi-scale voxel features into a BEV feature map, and fuses BEV feature maps of different resolutions through 2D sparse convolution processing and jump connection. Finally, the centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss, thereby guiding the detection head to achieve the classification of the train forward clearance target and the regression of the three-dimensional bounding box. The detection method based on point cloud feature aggregation enhancement in this embodiment has made a significant contribution to improving the accuracy and efficiency of train forward clearance target detection. At the same time, the method provides strong technical support for the safe operation of trains and provides a practical and effective solution for achieving efficient train forward clearance detection.
[0080] like Figure 1 As shown, the train forward clearance target detection method based on point cloud feature aggregation enhancement described in this embodiment includes the following steps:
[0081] Step S1: Obtain vehicle-mounted point cloud data through the laser radar installed at the front end of the train to construct a scene point cloud dataset;
[0082] Step S2: The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0083] Step S3: Design a multi-scale voxel feature aggregation module based on point query, use set sampling to sample, group and extract features from the multi-level voxel features extracted in step S2, and aggregate the multi-scale voxel features to the voxel center point;
[0084] Step S4: Design a BEV multi-level auxiliary branch feature extraction module, use a fully sparse and highly compressed method to map the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and fuse BEV feature maps of different resolutions through 2D sparse convolution processing and jump connection;
[0085] Step S5: The centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clearance object classification and 3D bounding box regression.
[0086] In this example, the specific technical solution is as follows:
[0087] (1) Multi-scale voxel feature aggregation based on point query
[0088] By designing a multi-level point query mechanism, multi-scale voxel features are effectively aggregated to the center of the voxel. First, the sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to downsample layer by layer to extract multi-scale voxel features. Subsequently, the Set Abstraction (SA) method is used to select voxel features within a specific radius at each scale, and their features are aggregated to the center of the voxel. This process allows the model to capture the feature distribution at different spatial scales, and form a high-quality spatial feature representation through aggregation operations to ensure that the information obtained from multiple levels is not lost. The aggregated features are processed by a multi-layer perceptron (MLP) to achieve efficient feature extraction and enhancement.
[0089] (2) BEV multi-level auxiliary branch feature extraction module
[0090] First, multi-level voxel features are extracted through a 3D sparse convolutional network. Then, these voxel features are converted into corresponding BEV feature maps using sparse height compression technology. 2D sparse convolution processing is used to further extract features and jump connections are introduced to enable low-level and high-level features to be effectively fused in the feature map, thereby improving the expressiveness of features and information fluidity.
[0091] (3) Centroid offset loss of 3D regression box
[0092] The centroid offset loss of the 3D regression box constrains the model's learning process by comparing the difference between the centroid of the predicted 3D bounding box and the centroid of the real box. The multi-task loss that combines classification loss, positioning loss, and orientation loss guides the detection head to perform target classification and 3D bounding box regression more accurately, which is conducive to improving the accuracy of the model's 3D bounding box positioning and enhancing its adaptability to targets of different scales and shapes.
[0093] In this example, the specific implementation is as follows:
[0094] Obtain vehicle-mounted point cloud data and construct scene point cloud datasets: Obtain vehicle-mounted point cloud data through the laser radar installed at the front of the train, such as Figure 2As shown in the figure, in order to better analyze and process these data, a train forward scene point cloud dataset is constructed, which simulates the scene of people and obstacles (such as cartons, foam stones) invading the track, aiming to support subsequent target detection and environmental perception tasks.
[0095] (1) Voxel feature encoding and multi-scale feature extraction of sparse point clouds
[0096] This example divides the sparse point cloud into a regular voxel grid, converts the point cloud information into a compact feature representation, and uses 3D sparse convolution to downsample layer by layer to extract multi-scale voxel features.
[0097] A) Voxel Rasterization
[0098] The input N spatial point clouds P = (p1, p2, ..., p n ) area is divided into voxel units with a spatial resolution of L×W×H, where the point cloud features represent (x, y, z, r), i.e., the 3D coordinates and reflection intensity of the point cloud.
[0099] B) Voxel feature encoding
[0100] The point cloud features within the voxel unit are averaged to obtain the initial encoded voxel features, which are expressed as a four-dimensional tensor: The corresponding spatial coordinates are expressed as Where N0 is the number of non-empty voxels, and the superscript l0 indicates the initialization voxel layer.
[0101] C) Multi-scale feature extraction
[0102] The voxel features of step S22 are downsampled layer by layer using 3D submanifold sparse convolution, and the features of non-empty voxels are encoded into multi-scale three-dimensional sparse feature quantities with downsampling ratios of 1x, 2x, 4x, and 8x respectively. The spatial coordinates of the voxel sampled at the kth layer are expressed as The set of voxel features is expressed as Where N k is the number of non-empty voxels in the kth layer, with the superscript l k (k=1,2,3,4) represents the downsampling layers with 1x, 2x, 4x, and 8x magnifications.
[0103] (2) Design of multi-scale voxel feature aggregation module based on point query
[0104] This example designs a multi-scale voxel feature aggregation module based on point query. It uses set sampling to sample, group and extract features from the multi-level voxel features extracted in step S2, and aggregates the multi-scale voxel features to the voxel center point. The algorithm flow is as follows: Figure 3 shown.
[0105] A) Point query feature grouping
[0106] The 8x downsampling layer in step S23 is encoded with conventional 3D sparse convolutional features as the aggregation layer. The center of the voxel unit in this layer is used as the query point Q of the set sampling (SA). The spatial coordinates of the query point Q = {q1, ..., q m}, where m is The set of voxel feature vectors of this layer is expressed as The superscript q indicates the query layer.
[0107] B) Point Query Clustering
[0108] For each query point q i (i=1,...,m) as the center, according to the set k-th layer radius r k Find all non-empty voxels in the layer domain and obtain the local grouped non-empty voxel set. The voxel feature vector set of the kth layer is expressed as The calculation formula is as follows:
[0109]
[0110] In the formula, Indicates that in the neighborhood r k Collect the 16 nearest voxels and connect their local relative coordinates With features
[0111] C) Point query feature aggregation
[0112] Use Set Abstraction (SA) (through MLP multi-layer perceptron and maximum pooling operation) to extract the features of the voxels in the kth layer:
[0113]
[0114] In the formula, A(.) represents a three-layer MLP network to encode features, and multi-scale features are aggregated into query points through the maximum pooling operation MAX(.) along the channel. In addition, we use two r of different scales in the k layer k Increase the local variable receptive field. Aggregate the voxel features of the initialization voxel layer and all downsampling layers into the query point q i The multi-scale voxel features f i (qv) , the aggregation formula is as follows:
[0115]
[0116] Where i = 1, ..., m, Represents the feature concatenation operation; the generated feature f i (qv) It contains both the multi-scale downsampling features of 3D CNN and the voxel initialization features (i.e. 3D coordinates and reflection intensity).
[0117] (3) Design of BEV multi-level auxiliary branch feature extraction module
[0118] This example designs a BEV multi-level auxiliary branch feature extraction module, which uses a completely sparse and highly compressed method to map the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and fuses BEV feature maps of different resolutions through 2D sparse convolution processing and jump connections. The algorithm flow is as follows: Figure 4 shown.
[0119] A) Spatial multi-rate feature mapping: A fully sparse height compression (SHC) operation is used to map the multi-scale non-empty voxel features in step S23 onto an XY plane grid, and the features at the same Z position are accumulated to finally generate BEV sparse feature maps with 1x, 2x, 4x, and 8x downsampling rates.
[0120] B) BEV multi-level feature extraction: 1x BEV sparse feature map obtained in step S41 Perform three-level 2D sparse convolution downsampling operations to extract BEV sparse convolution features with 2x, 4x, and 8x downsampling ratios layer by layer
[0121] C) Hierarchical feature fusion mechanism: In the downsampling process, a hierarchical feature fusion mechanism is introduced to fuse the multi-rate BEV features step by step in the order of 1x→2x→4x→8x. Extract high-resolution global semantic information, capture more detailed context information, and obtain 8x aggregated features with rich context information. By integrating BEV features of different spatial resolutions, it is helpful to improve the detection accuracy of small targets (such as pedestrians or obstacles) in complex scenes.
[0122] D) Convert the 8x aggregated features obtained in step S43 into 2D dense features and use them as the input of the detection head for target classification and localization tasks.
[0123] (2) 3D regression box centroid offset loss
[0124] This example introduces the centroid offset loss of the 3D regression box to optimize the multi-task loss and guide the detection head to achieve the train forward clearance target classification and 3D bounding box regression. The detection effect is as follows Figure 5 shown.
[0125] A) The 2D feature extraction network is responsible for further processing the 2D dense features and generating the final detection output. It consists of two subnetworks: a top-down subnetwork performs a 1x1 convolution and two downsamplings (16x, 32x) on the 2D dense features, and the other subnetwork performs upsampling in the opposite direction to restore the spatial resolution of the input (8x). Finally, the features of all layers and the original input features are aggregated and input into the detection head to achieve forward clearance object classification and 3D bounding box regression for the train.
[0126] B) A multi-task loss function is used during the model convergence process, including classification loss, positioning loss, orientation loss, and centroid offset loss L of the 3D regression box. centroid The calculation is as follows:
[0127]
[0128] In the formula, C box , C gt are the average coordinates of the spatial points in the 3D predicted bounding box and the ground truth box, respectively. γ and τ are the local density of each point estimated by the kernel density estimation (KDE) method. We introduce the L2 norm ||.||2 to calculate the offset between the center of mass of the 3D predicted regression box and the center of mass of the ground truth box.
[0129] The overall loss function is defined as:
[0130]
[0131] Where, L cls is the classification loss, L loc is the positioning loss, L dir is the direction loss, N pos Represents the number of positive anchors, where λ1=1, λ2=2, λ3=0.2, and λ4=1 are weight hyperparameters used to balance the losses of each part.
[0132] In summary, this example proposes a method and system for train forward clearance target detection based on point cloud feature aggregation enhancement. Preferably, the method first uses a series of steps to obtain on-board point cloud data using a laser radar installed at the front of the train, thereby constructing a scene point cloud dataset. This example uses a sparse point cloud to divide into regular voxel grids, uses voxel feature encoding to convert point cloud information into a compact feature representation, and downsamples layer by layer through 3D sparse convolution to extract multi-scale voxel features, thereby realizing train forward clearance target detection. At the same time, this example uses a designed point query-based multi-scale voxel feature aggregation module to sample, group and extract features from the extracted multi-level voxel features through set sampling, and further aggregates the multi-scale voxel features to the voxel center point. Next, the designed BEV multi-level auxiliary branch feature extraction module uses a completely sparse and highly compressed method to map the hierarchical multi-scale voxel features into a BEV feature map, and uses 2D sparse convolution processing and jump connections to fuse BEV feature maps of different resolutions to further improve the feature expression capability. In particular, this example introduces the centroid offset loss of the 3D regression box to optimize the multi-task loss and guide the detection head to achieve the classification and 3D bounding box regression of the train forward clearance target.
[0133] This example is based on a train forward clearance target detection method and system enhanced by point cloud feature aggregation. Through the collection and processing of lidar data, combined with sparse convolution and feature aggregation methods, accurate detection of forward clearance targets can be achieved, thereby improving the safety and intelligence level of trains and providing important technical support for the construction of intelligent railway systems.
[0134] Example 3
[0135] This embodiment 3 provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the train forward clearance target detection method based on point cloud feature aggregation enhancement as described above is implemented. The method includes:
[0136] Obtain vehicle-mounted point cloud data and build a scene point cloud dataset;
[0137] The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0138] Use set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregate the multi-scale voxel features to the voxel center point;
[0139] A fully sparse and highly compressed method is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps, and BEV feature maps of different resolutions are fused through 2D sparse convolution processing and skip connections.
[0140] The centroid shift loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clear object classification and 3D bounding box regression.
[0141] Example 4
[0142] This embodiment 4 provides a computer device, including a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the above-mentioned train forward clearance target detection method based on point cloud feature aggregation enhancement, the method comprising:
[0143] Obtain vehicle-mounted point cloud data and build a scene point cloud dataset;
[0144] The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0145] Use set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregate the multi-scale voxel features to the voxel center point;
[0146] A fully sparse and highly compressed method is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps, and BEV feature maps of different resolutions are fused through 2D sparse convolution processing and skip connections.
[0147] The centroid shift loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clear object classification and 3D bounding box regression.
[0148] Example 5
[0149] This embodiment 5 provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the above-mentioned method for detecting train forward clearance targets based on point cloud feature aggregation enhancement, the method comprising:
[0150] Obtain vehicle-mounted point cloud data and build a scene point cloud dataset;
[0151] The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer;
[0152] Use set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregate the multi-scale voxel features to the voxel center point;
[0153] A fully sparse and highly compressed method is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps, and BEV feature maps of different resolutions are fused through 2D sparse convolution processing and skip connections.
[0154] The centroid shift loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clear object classification and 3D bounding box regression.
[0155] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0157] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0159] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative work on the basis of the technical solution disclosed in the present invention should be included in the scope of protection of the present invention.
Claims
1. A train forward clearance target detection method based on point cloud feature aggregation enhancement, characterized in that: include: Obtain vehicle-mounted point cloud data and build a scene point cloud dataset; The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and 3D sparse convolution is used to extract multi-scale voxel features by downsampling layer by layer; Use set sampling to sample, group and extract features from the extracted multi-level voxel features, and aggregate the multi-scale voxel features to the voxel center point; A fully sparse and highly compressed method is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps, and BEV feature maps of different resolutions are fused through 2D sparse convolution processing and skip connections. The centroid shift loss of the 3D regression box is introduced to optimize the multi-task loss and guide the detection head to achieve train-forward clear object classification and 3D bounding box regression.
2. The train forward clearance target detection method based on point cloud feature aggregation enhancement according to claim 1 is characterized by: The sparse point cloud is divided into a regular voxel grid, and the voxel feature encoding converts the point cloud information into a compact feature representation, and uses 3D sparse convolution to downsample layer by layer to extract multi-scale voxel features, including: voxel gridding: the input N spatial point clouds P = (p1, p2, ..., p n ) area is divided into voxel units with a spatial resolution of L×W×H, where the point cloud features represent (x, y, z, r), i.e., the 3D coordinates and reflection intensity of the point cloud; voxel feature encoding: the point cloud features in the voxel unit are averaged to obtain the initial encoded voxel features, which are represented as a four-dimensional tensor; the voxel features are downsampled layer by layer using 3D submanifold sparse convolution, and the features of non-empty voxels are encoded as multi-scale three-dimensional sparse feature quantities with downsampling ratios of 1x, 2x, 4x, and 8x respectively; where the spatial coordinates of the voxel downsampled at the kth layer are represented as The set of voxel features is expressed as Where N k is the number of non-empty voxels in the kth layer, with the superscript l k (k=1,2,3,4) represents the downsampling layers with 1x, 2x, 4x, and 8x magnifications.
3. The train forward clearance target detection method based on point cloud feature aggregation enhancement according to claim 2 is characterized by: The method uses set sampling to sample, group and extract features from the multi-level voxel features extracted in step S2, and aggregates the multi-scale voxel features to the voxel center point, including: point query feature grouping: the 8x downsampling layer is used as an aggregation layer after being encoded with conventional 3D sparse convolution features, and the center of the voxel unit in this layer is used as the query point Q of the set sampling, and the spatial coordinates of the query point are calculated according to the voxel index and the actual size of the voxel; point query clustering: taking each query point as the center, according to the set k-th layer radius r k Find all non-empty voxels in the layer area, obtain the local grouped non-empty voxel set, and calculate the voxel feature vector set of the k-th layer; point query feature aggregation: use set sampling to extract the features of the voxels of the k-th layer.
4. The train forward clearance target detection method based on point cloud feature aggregation enhancement according to claim 3 is characterized by: Use two different scales of r in the k layer k Increase the local variable receptive field and aggregate the voxel features of the initialization voxel layer and all downsampling layers into the query point q i Multi-scale voxel features.
5. The train forward clearance target detection method based on point cloud feature aggregation enhancement according to claim 1 is characterized by: The described method of using completely sparse highly compressed maps the hierarchical multi-scale voxel features extracted in step S2 into a BEV feature map, and fuses the BEV feature maps of different resolutions through 2D sparse convolution processing and jump connection, including: spatial multi-rate feature mapping: using completely sparse highly compressed operations to map the multi-scale non-empty voxel features to the XY plane grid, and accumulating the features at the same Z position, and finally generating 1x, 2x, 4x and 8x downsampling magnification BEV sparse feature maps; BEV multi-level feature extraction: extracting the obtained 1x magnification BEV sparse feature maps; Sparse feature mapping, performing three-level 2D sparse convolution downsampling operations, extracting BEV sparse convolution features with 2x, 4x and 8x downsampling ratios layer by layer; Hierarchical feature fusion mechanism: In the downsampling process, a hierarchical feature fusion mechanism is introduced to fuse multi-rate BEV features in the order of 1x→2x→4x→8x step by step, extracting high-resolution global semantic information, capturing more detailed contextual information, and obtaining 8x aggregated features with rich contextual information; the obtained 8x aggregated features are converted into 2D dense features and used as the input of the detection head for target classification and positioning tasks.
6. The train forward clearance target detection method based on point cloud feature aggregation enhancement according to claim 1 is characterized by: The centroid offset loss of the 3D regression box is introduced to optimize the multi-task loss, guiding the detection head to achieve the train forward clearance target classification and 3D bounding box regression, including: the 2D feature extraction network is responsible for further processing the 2D dense features and generating the final detection output, including two sub-networks: a top-down sub-network performs a 1x1 convolution and two down-samplings on the 2D dense features, and the other sub-network performs upsampling in the opposite direction to restore the spatial resolution of the input, and finally aggregates the features of all layers and the original input features into the detection head to achieve the train forward clearance target classification and 3D bounding box regression.
7. A train forward clearance target detection system based on point cloud feature aggregation enhancement, characterized in that: include: The acquisition module is used to obtain vehicle-mounted point cloud data and build a scene point cloud dataset; The downsampling extraction module is used to divide the sparse point cloud into a regular voxel grid. The voxel feature encoding converts the point cloud information into a compact feature representation, and uses 3D sparse convolution to extract multi-scale voxel features by downsampling layer by layer; A multi-scale voxel feature aggregation module based on point query, which is used to sample, group and extract features of the extracted multi-level voxel features using set sampling, and aggregate the multi-scale voxel features to the voxel center point; BEV multi-level auxiliary branch feature extraction module, which is used to map the extracted hierarchical multi-scale voxel features into BEV feature maps using a fully sparse and highly compressed method, and fuse BEV feature maps of different resolutions through 2D sparse convolution processing and skip connections; The detection module is used to introduce the centroid shift loss of the 3D regression box to optimize the multi-task loss and guide the detection head to achieve train forward clearance object classification and 3D bounding box regression.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the train forward clearance target detection method based on point cloud feature aggregation enhancement as described in any one of claims 1-6 is implemented.
9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the train forward clearance target detection method based on point cloud feature aggregation enhancement as described in any one of claims 1-6.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the method for detecting forward clearance targets of a train based on point cloud feature aggregation enhancement as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Visualization method and device for completing multi-task semantic annotation on legal data
CN111651270A
Point cloud target detection method fusing original point cloud and voxel division
CN113378854A
Three-dimensional dynamic target detection method and device based on voxel point cloud fusion
CN113989797A
Lidar point cloud segmentation method, device, apparatus, and storage medium
US20240212374A1
Cited By
Cable joint detection method and system fusing dynamic voxelization and sparse convolution
CN120747072A
Target detection method, electronic equipment, storage medium and vehicle
CN121095537A
A target detection method, electronic equipment, storage medium and vehicle
CN121095537B
Train operation scene-oriented end-to-end target detection method
CN122368940A
A rail transit forward obstacle early warning method for a train scene
CN122724538A