Obstacle detection method and device, electronic equipment and storage medium

By combining feature extraction and fusion technology of point cloud data and image data, the accuracy problem of orbital obstacle detection in complex environments in the prior art is solved, and high-precision and robust obstacle detection are achieved.

CN120198885APending Publication Date: 2025-06-24BEIJING CENTURY DONGFANG COMMUNICATION EQUIPMENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510232243.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision obstacle detection in complex and changeable orbital environments, especially when the light changes greatly, it is prone to false detection or missed detection.

Method used

By obtaining point cloud data and image data of obstacles on the track, using three-dimensional sparse convolution networks and two-dimensional convolution networks to extract the data feature, and through feature alignment and fusion, combining the spatial positioning advantages of point cloud data and the semantic recognition capabilities of image data, accurate detection of obstacles is achieved.

Benefits of technology

It realizes orbital obstacle detection in complex environments, has high accuracy and strong robustness, and can effectively identify the type and location of obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198885A_ABST
    Figure CN120198885A_ABST
Patent Text Reader

Abstract

The invention provides an obstacle detection method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring point cloud data and image data acquired for an obstacle on a track, wherein the difference between acquisition moments of the point cloud data and the image data is within a preset time difference range; extracting features of the point cloud data through the three-dimensional sparse convolutional network to obtain point cloud features, and extracting features of the image data through the two-dimensional convolutional network and the feature pyramid network to obtain image features; aligning and fusing the point cloud features and the image features in sequence to obtain fused features; and determining a first detection result, a second detection result and a third detection result corresponding to the obstacle according to the point cloud feature, the image feature and the fusion feature, and finally determining a target detection result corresponding to the obstacle according to each detection result. By combining the advantages of the point cloud data in spatial positioning and the advantages of the image data in semantic recognition, accurate detection of track obstacles in a complex environment can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of detection technologies, and in particular, to an obstacle detection method, apparatus, electronic device, and storage medium. Background Art

[0002] Rail transit plays an important role in improving transportation efficiency and alleviating traffic congestion. However, its operating environment is complex and there are many potential safety hazards. Obstacles that may appear on the track, such as pedestrians, animals, falling rocks, etc., if not detected in time, may cause the train to collide or derail, seriously threatening the safety of passengers and staff. Therefore, an accurate and efficient obstacle detection system is particularly important.

[0003] In related technologies, most of the track obstacle detection methods rely on a single sensor. If a lidar is used, although the lidar can provide high-precision distance information, it lacks semantic information and is difficult to identify the type of obstacle, and its detection ability for small obstacles and low-reflectivity objects is weak; if a vision sensor is used, although the vision sensor can provide rich semantic information through images, such as the type, color, and texture of objects, etc., it is sensitive to changes in illumination, especially in strong reflection or low-light environments, and is prone to false detection or missed detection, reducing the detection accuracy. Therefore, the existing track obstacle detection methods are difficult to achieve high-precision detection in a complex and changing track environment. Summary of the Invention

[0004] The present application provides an obstacle detection method, apparatus, electronic device, and storage medium, which are used to solve the defect that the existing track obstacle detection methods are difficult to achieve high-precision detection in a complex and changing track environment, and can realize the accurate detection of track obstacles in a complex environment, with high precision and strong robustness.

[0005] The present application provides an obstacle detection method, including the following steps: Obtain point cloud data and image data collected for obstacles on the track, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; Extract features from the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and extract features from the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; Perform alignment processing on the point cloud features and the image features, and perform feature fusion on the aligned point cloud features and image features to obtain fused features; Respectively determine a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fused features; Determine a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes type information and position information of the obstacle.

[0006] According to an obstacle detection method provided by the present application, the aligning the point cloud feature and the image feature, and performing feature fusion on the aligned point cloud feature and image feature to obtain a fusion feature includes: Project the point cloud feature and the image feature into the BEV space respectively to obtain an aligned point cloud BEV feature and an image BEV feature; Perform feature fusion on the point cloud BEV feature and the image BEV feature according to the following formulas (1)-(4) to obtain the fusion feature: F PI =Conv2D(Concat(F PB , F IB )) (1) F F1 =Transformer(Q IB , K PI , V PI ) (2) F F2 =Transformer(Q PB , K PI ’, V PI ’) (3) F BEV =Reshape(Linear(F F2 )) (4) Wherein, Q IB =Linear(F IB ), K PI =V PI =Linear(F PI ), Q PB =Linear(F PB ), K PI ’=V PI ’=Linear(F F1 ), Linear represents a linear transformation, Concat represents feature concatenation, Conv2D represents a two-dimensional convolution operation, Transformer represents feature fusion based on an attention mechanism, Reshape represents mapping F F2 back to the BEV dimension, F PB represents the point cloud BEV feature, F IB represents the image BEV feature, and F BEV represents the fusion feature.

[0007] According to an obstacle detection method provided by the present application, the step of respectively projecting the point cloud feature and the image feature into the BEV space to obtain the aligned point cloud BEV feature and image BEV feature includes: Projecting the point cloud feature into the BEV space according to a preset BEV space resolution; Projecting the image feature into the same dimension as the point cloud feature, and performing feature fusion on the obtained image feature to obtain a first feature; Based on the first feature, the depth information of the image feature, the internal parameter matrix of the acquisition device corresponding to the image data, and the transformation matrix between the coordinate system of the acquisition device corresponding to the image data and the coordinate system of the acquisition device corresponding to the point cloud data, projecting the image feature into the BEV space according to the preset BEV space resolution.

[0008] According to an obstacle detection method provided by the present application, the step of respectively determining a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud feature, the image feature, and the fusion feature includes: Determining the first detection result corresponding to the obstacle through a multi-layer perceptron network according to the point cloud features at multiple different scales; Based on the anchor-free mechanism, determining the second detection result corresponding to the obstacle through two-dimensional convolution operations according to the image features at multiple different scales; Determining the third detection result corresponding to the obstacle through a fusion transformation detection head network according to the fusion feature.

[0009] According to an obstacle detection method provided by the present application, the position information includes three-dimensional space coordinates and two-dimensional plane coordinates; the step of determining a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result includes: Performing a first filtering operation on the type information in the first detection result, the second detection result, and the third detection result that is less than the corresponding confidence threshold through the confidence threshold corresponding to each of the point cloud feature, the image feature, and the fusion feature; Performing a second filtering operation on the three-dimensional space coordinates and two-dimensional plane coordinates in the detection results obtained after the first filtering operation through non-maximum suppression; Correcting the three-dimensional plane coordinates obtained after the second filtering operation according to the two-dimensional plane coordinates obtained after the second filtering operation to obtain a corrected detection result, and the corrected detection result is the target detection result.

[0010] According to an obstacle detection method provided by the present application, a three-dimensional sparse convolutional network includes a plurality of different sparse convolutional layers, and the number of feature channels of each sparse convolutional layer is different; extracting point cloud features from the point cloud data through the three-dimensional sparse convolutional network includes: Performing voxelization processing on the point cloud data to obtain a plurality of voxels; Performing sparse coding on each of the voxels to obtain sparse features corresponding to each of the voxels; Inputting the sparse features into the three-dimensional sparse convolutional network, and extracting features of the point cloud data at multiple different scales through the feature channels of each sparse convolutional layer in the three-dimensional sparse convolutional network to obtain the point cloud features.

[0011] According to an obstacle detection method provided by the present application, a two-dimensional convolutional network includes a plurality of different convolutional layers, and the number of feature channels of each convolutional layer is different. Extracting image features from the image data through the two-dimensional convolutional network and the feature pyramid network includes: Inputting the image data into the two-dimensional convolutional network, and extracting features of the image data at multiple different scales through the feature channels of each convolutional layer in the two-dimensional convolutional network to obtain second features; Inputting the second features into the feature pyramid network, and performing feature fusion on the second features through the top-down feature propagation mechanism of the feature pyramid network to obtain the image features.

[0012] The present application further provides an obstacle detection device, including the following modules: An acquisition module, configured to acquire point cloud data and image data collected for an obstacle on an orbit, and a difference between a collection time of the point cloud data and a collection time of the image data is within a preset time difference range; A feature extraction module, configured to extract point cloud features from the point cloud data through a three-dimensional sparse convolutional network, and extract image features from the image data through a two-dimensional convolutional network and a feature pyramid network; A feature fusion module, configured to perform alignment processing on the point cloud features and the image features, and perform feature fusion on the aligned point cloud features and image features to obtain fusion features; A first determination module, configured to respectively determine a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fusion features; A second determination module, configured to determine a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes type information and location information of the obstacle.

[0013] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the obstacle detection method described in any one of the above is implemented.

[0014] The present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the obstacle detection method described in any one of the above is implemented.

[0015] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the obstacle detection method described in any one of the above is implemented.

[0016] The present application provides an obstacle detection method, device, electronic device, and storage medium. Through the obstacle detection method of the present application, first, point cloud data and image data collected for an obstacle on an orbit are obtained, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; then, the point cloud data is subjected to feature extraction through a three-dimensional sparse convolutional network to obtain point cloud features, and the image data is subjected to feature extraction through a two-dimensional convolutional network and a feature pyramid network to obtain image features; then, the point cloud features and the image features are aligned, and the aligned point cloud features and image features are subjected to feature fusion to obtain fusion features; then, a first detection result, a second detection result, and a third detection result corresponding to the obstacle are determined according to the point cloud features, the image features, and the fusion features respectively; finally, a target detection result corresponding to the obstacle is determined according to the first detection result, the second detection result, and the third detection result, where the target detection result includes type information and location information of the obstacle. Since the point cloud data can provide high-precision distance information and the image data can provide rich semantic information, the obstacle detection method in the present application can achieve accurate detection of track obstacles in a complex environment by simultaneously combining the advantages of the point cloud data in spatial positioning and the advantages of the image data in semantic recognition, and has high precision and strong robustness. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 is a flowchart of an obstacle detection method shown in an embodiment of the present application; Figure 2 is a schematic diagram of the process of feature alignment and fusion shown in an embodiment of the present application; Figure 3 is a schematic diagram of the process of obtaining a target detection result shown in an embodiment of the present application; Figure 4 is a structural block diagram of an obstacle detection device shown in an embodiment of the present application; Figure 5 is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application. Detailed implementation manners

[0019] To make the objectives, technical solutions, and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0020] The execution subject of an obstacle detection method provided by the present application can be an obstacle detection device or an electronic device. The following will take the execution subject being an obstacle detection device as an example to detail the obstacle detection method of the present application.

[0021] Figure 1 is a flowchart of an obstacle detection method shown in an embodiment of the present application. Referring to Figure 1 , the obstacle detection method of the present application may include: Step 101, acquire point cloud data and image data collected for an obstacle on an orbit, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range.

[0022] In this application, point cloud data can be obtained by a three-dimensional scanning device, such as a Light Detection and Ranging (LiDAR). Image data can be obtained by an image acquisition device, such as a camera. This application does not limit the specific types of the device for collecting point cloud data and the device for collecting image data, and can be selected according to actual needs.

[0023] Among them, point cloud data is recorded in the form of points. Each point in the point cloud data contains three-dimensional coordinate (X, Y, Z) information and reflection intensity information r. The intensity refers to the echo intensity collected by the three-dimensional scanning device. Exemplarily, when using a LiDAR to scan in space, every time an object is scanned, the three-dimensional coordinates of a point and the reflection intensity r of this point are recorded. In this way, the entire scene can be described by a data set composed of countless similar points, and the data in this data set is the point cloud data.

[0024] Before performing step 101, it is necessary to synchronously collect point cloud data and image data. During the process of synchronous data collection, it is necessary to ensure the consistency of the three-dimensional scanning device and the image acquisition device in terms of time and space. Taking the three-dimensional scanning device as a LiDAR and the image acquisition device as a camera as an example, the following Step 1 - Step 3 need to be specifically performed before step 101: Step 1: Select appropriate positions on the side of the track to install the LiDAR and the camera, and ensure that the LiDAR and the camera can cover the monitored track area to provide complete perception information.

[0025] Step 2: Perform spatial calibration to determine the conversion relationship between the coordinate system of the LiDAR and the coordinate system of the camera. Specifically, first use the Zhang Zhengyou calibration method to calibrate the internal parameter matrix K of the camera. The internal parameter matrix contains information such as the focal length of the camera, the coordinates of the principal point, and the distortion coefficient, and is an important parameter of the camera imaging model. Then, based on the planar template method, calibrate the external parameter matrix for realizing the coordinate system transformation between the LiDAR and the camera. The external parameter matrix includes the transformation matrix M1 from the coordinate system of the LiDAR to the coordinate system of the camera, and the transformation matrix M2 from the coordinate system of the camera to the coordinate system of the LiDAR. Through M1 or M2, the point cloud data collected by the LiDAR and the image data collected by the camera can be converted into the same coordinate system.

[0026] Among them, the Zhang Zhengyou calibration method is a camera calibration method based on a planar checkerboard. This method extracts the pixel coordinates of the checkerboard corner points by shooting images of a checkerboard of a specific size, and combines the world coordinates (known) of the checkerboard corner points to calculate the internal parameter matrix and the external parameter matrix of the camera.

[0027] The planar template method is a camera calibration method based on a planar template. This method calculates the internal and external parameters of the camera by taking images of the planar template and using the correspondence between the feature points on the template and their projection points on the image. The planar template can be a checkerboard, a dot array, or other patterns with obvious feature points.

[0028] Step 3: Ensure that the lidar and the camera start collecting data at the same time through a hardware trigger system (such as a synchronization signal generator or a trigger controller) to achieve time synchronization and avoid data mismatch problems caused by time asynchrony. During the data collection process, each data point is attached with a timestamp indicating the time of its collection. If multiple groups of point cloud data and image data are collected synchronously, then by comparing the timestamps of the point cloud data and the image data, select the data with the smallest difference in timestamps and within the allowable range of time error (i.e., the preset time difference range) as the synchronized data (P, I), which can further ensure the synchrony of the data and improve the accuracy of subsequent data processing, where P represents the point cloud data and I represents the image data.

[0029] After obtaining the point cloud data and the image data through step 101, it is also necessary to preprocess the point cloud data and the image data. The preprocessing process specifically includes: (1) Denoising the point cloud data: In an orbital environment, the ground point cloud occupies most of the data, and obstacles are usually above the ground. In this application, the Random Sample Consensus (RANSAC) algorithm combined with a height threshold TL can be used to reduce the ground noise point cloud information.

[0030] Among them, the Random Sample Consensus algorithm is an iterative method used to estimate the parameters of a mathematical model from data containing a large amount of noise. In this application, it is used to fit the ground plane model and identify the points that do not belong to this plane (i.e., noise points or obstacle points).

[0031] The height threshold TL is a set threshold used to distinguish ground points from obstacle points. Generally, the height of ground points is relatively low, and the height of obstacle points is relatively high. Therefore, by setting a suitable height threshold, most of the ground points can be removed while retaining the obstacle points.

[0032] (2) Denoising the image data: In this embodiment, algorithms such as Gaussian filtering or median filtering can be used to remove the noise points in the image and improve the image quality.

[0033] Gaussian filtering is a linear smoothing filtering method used to remove Gaussian noise in images. It smooths the image by calculating the weighted average of the pixels within the neighborhood around each pixel point and replacing the original pixel value with this average value.

[0034] Median filtering is a non - linear filtering method commonly used to remove salt - and - pepper noise in images. It eliminates noise by sorting the pixel values within the neighborhood around each pixel point and taking the median value as the new value of that pixel.

[0035] In addition, methods such as adaptive histogram equalization and brightness adjustment can also be used in this application to enhance the contrast and brightness of image data in order to highlight obstacles.

[0036] Adaptive histogram equalization is an image enhancement technique used to improve the contrast of images. It enhances the contrast of each local region of the image through histogram equalization of the local regions, while avoiding the noise amplification caused by excessive enhancement of the global contrast.

[0037] Brightness adjustment refers to adjusting the brightness value of an image to make the image brighter or darker, in order to improve the visual effect of the image and facilitate highlighting obstacles in the case of insufficient or uneven lighting.

[0038] Step 102: Extract features from the point cloud data through a three - dimensional sparse convolutional network to obtain point cloud features, and extract features from the image data through a two - dimensional convolutional network and a feature pyramid network to obtain image features.

[0039] In step 102, the point cloud data and image data used are pre - processed data.

[0040] In this embodiment, multi - scale features of the point cloud image can be extracted through a 3D sparse convolutional network to capture the spatial structure and geometric information of obstacles. At the same time, use a 2D convolutional network to extract multi - scale features in the image data to capture the semantic information, texture, appearance features, etc. of obstacles.

[0041] Step 103: Align the point cloud features and image features, and perform feature fusion on the aligned point cloud features and image features to obtain fused features.

[0042] In this embodiment, the point cloud features and image features are mapped to a unified bird's - eye view space (Bird’s EyeView, BEV) to achieve feature alignment. The point cloud features after feature alignment are point cloud BEV features, and the image features after feature alignment are image BEV features.

[0043] After achieving feature alignment, a Transformer-based feature fusion method is used to fuse the image BEV feature and the point cloud BEV feature through the Cross-Attention mechanism to obtain the BEV fusion feature, which contains both the spatial geometric information of the point cloud data and the semantic information of the image data.

[0044] Step 104: Determine the first detection result, the second detection result, and the third detection result corresponding to the obstacle based on the point cloud feature, the image feature, and the fusion feature respectively.

[0045] For the point cloud feature, the point cloud feature can be processed by a Multilayer Perceptron (MLP) network to output the first detection result. The first detection result includes: the type information of the obstacle and the three-dimensional position information (the three-dimensional center point position, length, width, height, orientation angle, etc. of the obstacle).

[0046] For the image feature, the image feature can be processed by a 2D convolutional neural network to output the second detection result. The second detection result includes: the type information of the obstacle and the two-dimensional position information (the two-dimensional center point position, length, and width of the obstacle).

[0047] For the fusion feature, the third detection result can be obtained through the TransFusion method. The third detection result includes the type information of the obstacle and the three-dimensional position information.

[0048] Step 105: Determine the target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes the type information and the position information of the obstacle.

[0049] By executing Step 105, the first detection result, the second detection result, and the third detection result can be sorted out and analyzed to determine the target detection result of the obstacle.

[0050] To implement the method of the present application, first, point cloud data and image data collected for obstacles on the track are obtained, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; then, the point cloud data is subjected to feature extraction through a three-dimensional sparse convolutional network to obtain point cloud features, and the image data is subjected to feature extraction through a two-dimensional convolutional network and a feature pyramid network to obtain image features; then, the point cloud features and the image features are aligned, and the aligned point cloud features and image features are subjected to feature fusion to obtain fusion features; then, according to the point cloud features, the image features, and the fusion features respectively, a first detection result, a second detection result, and a third detection result corresponding to the obstacle are determined; finally, according to the first detection result, the second detection result, and the third detection result, a target detection result corresponding to the obstacle is determined, where the target detection result includes the type information and the position information of the obstacle. Since the point cloud data can provide high-precision distance information and the image data can provide rich semantic information, the obstacle detection method in the present application can achieve accurate detection of track obstacles in a complex environment by simultaneously combining the advantages of the point cloud data in spatial positioning and the advantages of the image data in semantic recognition, and has high precision and strong robustness.

[0051] Combined with the above embodiments, in one implementation manner, the three-dimensional sparse convolutional network includes a plurality of different sparse convolutional layers, and the number of feature channels of each sparse convolutional layer is different. Correspondingly, in step 102, the process of extracting point cloud features from the point cloud data through the three-dimensional sparse convolutional network may include: Voxelize the point cloud data to obtain a plurality of voxels; Perform sparse coding on each voxel to obtain the sparse feature corresponding to each voxel; Input the sparse features into the three-dimensional sparse convolutional network, and extract the features of the point cloud data at multiple different scales through the feature channels of each sparse convolutional layer in the three-dimensional sparse convolutional network to obtain point cloud features.

[0052] In this embodiment, for the point cloud data P = {pi | pi = (xi, yi, zi, ri)}, a 3D sparse convolutional network is used for feature extraction. Each point cloud is represented as pi = (xi, yi, zi, ri), where (xi, yi, zi) are the three-dimensional spatial coordinates of the point cloud, and ri is the reflection intensity.

[0053] Next, based on the initial grid size (v x , v y , v z), for example, (0.1m, 0.1m, 0.1m), perform voxelization on the point cloud data P to obtain multiple voxels. Usually, the number of points in the point cloud data is large, and it is difficult to directly process the point cloud data. Therefore, the three-dimensional space represented by the point cloud data P can be first divided into multiple boxes of the same size (one box represents a voxel in the 3D space), and then the points falling into each box are grouped into one category. In this way, all the points in the point cloud data can be divided into multiple boxes, and there are a certain number of points in each box. This process is voxelization. The above (v x , v y , v z ) is the size of each box.

[0054] Since the number of points in each voxel may not be the same, for the convenience of processing, this application retains a fixed number of points (for example, 10 points) in each voxel. If the number of points in a voxel is less than this fixed number, some points can be randomly selected to fill the voxel. Similarly, if the number of points in a voxel is more than this fixed number, some points can be randomly removed. In this way, the number of points in each voxel can be guaranteed to be the same.

[0055] Next, perform sparse coding on each voxel to obtain the sparse features corresponding to each voxel, denoted as F V . The sparse features corresponding to each voxel include the central position of the voxel, the average coordinate position of all points in the voxel, etc. Among them, sparse coding is a representation learning method, which aims to find a sparse representation (or coding) to effectively represent the data. Through sparse coding, this application can capture the most valuable information in the point cloud data while removing redundancy and noise.

[0056] Finally, input the sparse features F after sparse coding V into a 3D sparse convolutional network to obtain multi-scale point cloud features.

[0057] In this application, there are three sparse convolutional layers in the 3D sparse convolutional network, so the multi-scale point cloud features can be represented as F P =(F P1 , F P2 , F P3 ), and the feature channel dimensions of the multi-scale features are (64, 128, 256), corresponding to the resolutions of (1 / 2, 1 / 4, 1 / 8).

[0058] In this application, each sparse convolutional layer has a different number of feature channels (or different types of feature information). The first layer has 64 feature channels, the second layer has 128 feature channels, and the third layer has 256 feature channels. Since each layer extracts features from the previous layer, as the network goes deeper, it can capture more and more abstract and complex features. This multi-layer design enables the network to simultaneously extract features at different scales, which helps to more comprehensively understand the data in the subsequent stage and improve the performance and accuracy of the deep learning model. At the same time, the output features of each layer will be more compressed in space than the previous layer, and the resolution will decrease.

[0059] For the feature F of each layer Pi , it can be expressed as: F Pi = ReLU(BatchNorm(SparseConv3D(F P(i-1) , W Pi ))) Among them, W Pi represents the 3D convolution weight of the i-th layer; F P(i-1) represents the output feature of the previous layer; F Pi represents the output feature of the current layer; ReLU is an activation function used for non-linear transformation of the normalized features; BatchNorm represents the batch normalization operation, which is used to normalize the features after the convolution operation, helping the network to converge faster and improve the training efficiency; SparseConv3D represents the 3D sparse convolution operation, which is used to screen out valuable information from the features of the previous layer and combine it with the weight W Pi of the current layer to generate new features.

[0060] In this embodiment, each sparse convolutional layer is composed of multiple 3D sparse convolutional kernels, which are used to extract features from the input data through 3D sparse convolution operations. The 3D sparse convolutional kernel is a window that slides in three-dimensional space and is used to extract features from the input data. Different from the two-dimensional convolutional kernel, the 3D sparse convolutional kernel can capture features in three dimensions simultaneously. The 3D sparse convolution operation is a convolution operation designed for sparse data, which can only process the non-zero value points and the points in their neighborhoods, thus avoiding the scanning of invalid regions in the traditional convolution operation, significantly reducing the computational amount, and improving the computational efficiency.

[0061] In actual implementation, the number of sparse convolutional layers included in the three-dimensional sparse convolutional network can be set according to actual needs, and this application does not make specific restrictions on this.

[0062] In this application, the two-dimensional convolutional network also includes multiple different convolutional layers, and the number of feature channels in each convolutional layer is different. Correspondingly, in step 102, extracting image features from the image data through the two-dimensional convolutional network and the feature pyramid network may include: Input the image data into the two-dimensional convolutional network, and extract the features of the image data at multiple different scales through the feature channels of each convolutional layer in the two-dimensional convolutional network to obtain the second feature; Input the second feature into the feature pyramid network, and perform feature fusion on the second feature through the top-down feature propagation mechanism of the feature pyramid network to obtain the image feature.

[0063] Specifically, for the image data, a 2D convolutional network and a feature pyramid network can be used to extract multi-scale image features.

[0064] In actual implementation, the image I captured by the camera is usually in RGB format, with a size of H I ×W I ×3. In this application, first, the image data is normalized (i.e., the pixel values are scaled to a specific range), and then scaled to the network input size H I ′×W I ′, and input into the 2D convolutional network for multi-layer convolutional operations to obtain the second feature. In the 2D convolutional network, each layer includes two-dimensional convolutional operation Conv2D, batch normalization operation BatchNorm, activation function Leaky ReLU (LReLU), etc., which can gradually extract low-level to high-level semantic features.

[0065] Next, for the second feature, through the top-down feature propagation of the feature pyramid network FPN, the high-level semantic information is transmitted to the low-level features. In this application, the high-level feature map contains rich semantic information but has a low resolution and less detailed information, while the low-level feature map has a high resolution, more detailed information, but less semantic information. Therefore, through the upsampling operation, the resolution of the high-level feature map can be matched with that of the low-level feature map, and the high-resolution features are fused with the adjacent medium-resolution features through the upsampling operation. After each layer of fusion, a convolutional operation is used to further enhance the feature expression ability. There are three convolutional layers in the 2D convolutional network in this application, so the finally output multi-scale image features can be expressed as F I =(F I1 ,F I2 ,F I3 ), and the feature channel dimension is (256, 512, 1024), corresponding to the resolutions of (1 / 8, 1 / 16, 1 / 32).

[0066] In this application, each convolutional layer has a different number of feature channels. The first layer has 256 feature channels, the second layer has 512 feature channels, and the third layer has 1024 feature channels. Since each layer extracts features from the previous layer, as the network goes deeper, it can capture more and more abstract and complex features. This multi-layer design enables the network to extract features at different scales simultaneously, which helps to understand the data more comprehensively in the subsequent process and improve the performance and accuracy of the deep learning model. At the same time, the output features of each layer will be more compressed in space than the previous layer, and the resolution will decrease.

[0067] For the feature F of each layer Ii , it can be expressed as: F Ii = LReLU(BatchNorm(Conv2D(F I(i-1) , W Ii ))) Among them, W Ii is the 2D convolutional weight of the i-th layer, and F I(i-1) represents the output features of the previous layer; F Ii represents the output features of the current layer; ReLU is the activation function used for non-linear transformation of the normalized features; BatchNorm represents the batch normalization operation used for normalizing the features after the convolutional operation; Conv2D represents the 2D convolutional operation, and Conv2D(F I(i-1) , W Ii ) represents applying the 2D convolutional operation to the feature map F I(i-1) of the previous layer, and the convolutional weight is W Ii .

[0068] In actual implementation, the number of convolutional layers included in the two-dimensional convolutional network can be set according to actual needs, and this application does not make specific restrictions on this.

[0069] Combined with the above embodiments, in one implementation, step 103 may include: Project the point cloud features and image features into the BEV space respectively to obtain the aligned point cloud BEV features and image BEV features; According to the point cloud BEV features and image BEV features, perform feature fusion through the following formulas (1)-(4) to obtain the fused features: F PI = Conv2D(Concat(F PB , F IB )) (1) F F1 = Transformer(Q IB , K PI , V PI) (2) F F2 = Transformer(Q PB , K PI ’, V PI ) (3) F BEV = Reshape(Linear(F F2 )) (4) Among them, Q IB = Linear(F IB ), K PI = V PI = Linear(F PI ), Q PB = Linear(F PB ), K PI ’ = V PI ’ = Linear(F F1 ), Linear represents a linear transformation, Concat represents feature concatenation, Conv2D represents a two-dimensional convolution operation, Transformer represents feature fusion based on the attention mechanism, Reshape represents mapping F F2 back to the BEV dimension, F PB represents the point cloud BEV feature, F IB represents the image BEV feature, F BEV represents the fused feature.

[0070] Among them, projecting the point cloud feature and the image feature into the BEV space respectively to obtain the aligned point cloud BEV feature and image BEV feature may include: Projecting the point cloud feature into the BEV space according to the preset BEV space resolution; Projecting the image feature into the same dimension as the point cloud feature, and performing feature fusion on the obtained image feature to obtain the first feature; Based on the first feature, the depth information of the image feature, the internal parameter matrix of the acquisition device corresponding to the image data, and the conversion matrix between the coordinate system of the acquisition device corresponding to the image data and the coordinate system of the acquisition device corresponding to the point cloud data, projecting the image feature into the BEV space according to the preset BEV space resolution.

[0071] Figure 2 is a schematic diagram of a feature alignment and fusion process shown in an embodiment of the present application. The following will be combined with Figure 2 , and step 103 of the present application will be described in detail.

[0072] In the present application, after step 102, multi-scale point cloud features F P = (F P1 , FP2 , F P3 ), and multi-scale image features F I = (F I1 , F I2 , F I3 ). First, project the point cloud features and image features into a unified bird's-eye view space for feature alignment, and then fuse the point cloud BEV features and image BEV features through a Transformer network to obtain the fused features.

[0073] When performing feature alignment, for the point cloud feature F P = (F P1 , F P2 , F P3 ), first transform the dimension of F Pi from C Pi × H Pi × W Pi × L Pi to (C Pi × L Pi ) × H Pi × W Pi , and then map it to the unified BEV space (H BEV , W BEV ) using bilinear interpolation or downsampling. Then, fuse the multi-scale features based on 2D convolution to obtain the point cloud BEV feature F PB . Among them, C pi represents the number of channels, H Pi and W Pi represent the height and width respectively, and L Pi is the dimension related to the intensity of the point cloud. Among them, (H BEV , W BEV ) is the pre-defined image BEV space resolution. H BEV represents the height of the BEV space (the number of pixels in the vertical direction), and W BEV represents the width of the BEV space (the number of pixels in the horizontal direction). At the same time, reduce the feature channel dimension to 256 through 2D convolution.

[0074] Among them, F PB = Conv2D(Concat(Conv2D(F P1 ), Conv2D(F P2 ), Conv2D(F P3 ))) For the multi-scale image feature F I = (F I1 , F I2 , F I3), first, it is mapped to a unified dimension using bilinear interpolation or downsampling, and then multi-scale features are fused based on 2D convolution to obtain multi-scale fused features, and the feature channel dimension is reduced to 512 to obtain the image feature F IF (i.e., the first feature), as follows: F IF =Conv2D(Concat(Conv2D(F I1 ),Conv2D(F I2 ),Conv2D(F I3 ))) Then, the depth information d of the image feature is predicted by the LSS frustum method, and the image feature is mapped to the BEV space. Each pixel point (u, v) in the image is transformed to the coordinate system of the lidar based on the predicted depth information d through the camera intrinsic parameter K and the transformation matrix M2 from the camera coordinate system to the lidar coordinate system: (x,y,z,1) T =M2K -1 (u*d,v*d,d,1) T Then, the (x, y, z) coordinates in the lidar coordinate system are projected to the (u′, v′) position in the BEV space: u′=((x - x min ) / (x max - x min ))*W BEV v′=((y - y min ) / (y max - y min ))*H BEV Finally, according to the coordinate mapping relationship, the image feature F IF is mapped to the BEV space to obtain the image BEV feature F IB : F IB =LSS(F IF ,K,M2) In this embodiment, the LSS frustum method is a method for converting image features from the image view to bird's-eye view features. Its core idea is to use depth information to map image features from the two-dimensional image space to the three-dimensional space and further project them onto the BEV plane, so as to achieve feature alignment.

[0075] When fusing features, for BEV feature fusion, in this application, the interactive attention mechanism based on Transformer is used to fuse the point cloud BEV features and the image BEV features. The feature fusion process specifically includes: First, initialize the fused feature F PI: F PI =Conv2D(Concat(F PB ,F IB )) By transforming the dimension, F IF 、F IB and F PB The dimension changes from (C, H, W) to (H*W, C), where C represents the number of channels, W represents the width, and H represents the height. Then, the image BEV features and the point cloud BEV features are fused successively through the interactive attention mechanism.

[0076] Next, the process of fusing image BEV features is as follows: F F1 =Transformer(Q IB ,K PI ,V PI ) Q IB =Linear(F IB ) K PI =V PI =Linear(F PI ) Among them, Linear represents the linear transformation operation, Q IB represents the query vector, K PI represents the key vector, V PI Represents a value vector. Transformer’s interactive attention mechanism will be based on Q IB and K PI The similarity of V PI Perform weighted summation to achieve image BEV features and initial fusion features F PI fusion of.

[0077] Next, the process of fusing point cloud BEV features is as follows: F F2 =Transformer(Q PB ,K PI ',V PI ') Q PB =Linear(F PB ) K PI =V PI '=Linear(F F1 ) Through this step, the point cloud BEV features can be fused into F F1 .

[0078] Finally, the feature F F2Map back to the original BEV dimension to generate the final fused BEV feature F BEV (i.e., the fused feature): F BEV =Reshape(Linear(F F2 )) Combining the above embodiments, in one implementation, step 104 may include: Determine the first detection result corresponding to the obstacle through a multi-layer perceptron network based on the point cloud features at multiple different scales; Determine the second detection result corresponding to the obstacle through a two-dimensional convolution operation based on the anchor-free mechanism according to the image features at multiple different scales; Determine the third detection result corresponding to the obstacle through a fusion transformation detection head network according to the fused feature.

[0079] Figure 3 FIG. is a schematic diagram of a process for obtaining a target detection result shown in an embodiment of the present application. The following will be combined with Figure 3 , and step 104 of the present application will be described in detail.

[0080] In this embodiment, for the multi-scale point cloud feature F P =(F P1 ,F P2 ,F P3 ), a multi-layer perceptron network is used to predict the detection result for each voxel. According to different target scales of small targets, medium targets, and large targets, voxel features with different resolutions are used for 3D target detection. Each voxel feature is processed through a multi-layer perceptron network, so as to output the type information and 3D bounding box information (such as the position and size of the obstacle in three-dimensional space) of the obstacle corresponding to the voxel feature. This part corresponds to Figure 3 Point cloud feature classification and regression in.

[0081] For the voxel feature F Pi , the type information and 3D bounding box information output based on the MLP are as follows: P cls,i =MLP(F Pi ) P 3Dbox,i =MLP(F Pi ) Among them, P cls,i represents the type information corresponding to the voxel feature F Pi , and P 3Dbox,i represents the 3D bounding box information corresponding to the voxel feature F Pi .

[0082] The first detection result includes the above P cls,i and P3Dbox,i 。

[0083] For the multi-scale image features F I =(F I1 , F I2 , F I3 ), an anchor-free mechanism is adopted, and the 2D detection results of each pixel are directly predicted through 2D convolution. At the same time, a method of predicting according to the object scale (small object, medium object, large object) is adopted, and feature maps with different resolutions are used to detect obstacles of different scales respectively to improve the detection accuracy. This part corresponds to Figure 3 image feature classification and regression in

[0084] In this embodiment, for each feature map F Ii , based on the 2D convolution operation, type information and 2D bounding box information (such as the position and size of the obstacle in the image) are output.

[0085] I cls,i = Conv2D(F Ii ) I 2Dbox,i = Conv2D(F Ii ) where I cls,i represents the type information corresponding to the pixel point F I1 , and I 2Dbox,i represents the 2D bounding box information corresponding to the pixel point F I1 . I cls,i = Conv2D(F Ii ) means that the 2D convolution is used to process the feature map F Ii to obtain the type information corresponding to each pixel point on the feature map I cls,i , such as whether it is a pedestrian or other object.

[0086] I 2Dbox,i = Conv2D(F Ii ), which means that the 2D convolution is used to extract the bounding box information of the obstacle in the two-dimensional image from the feature map F Ii , including the position and size of the obstacle, etc., for determining the specific range of the obstacle in the image.

[0087] The second detection result includes the above I cls,i and I 2Dbox,i .

[0088] For the fused features, the type information and 3D bounding box information of the obstacle are output based on the TransFusion detection head network: E cls , E 3Dbox=TransFusion(F BEV ) Among them, E cls represents the type information of the obstacle, and E 3Dbox represents the 3D bounding box information, including the more precise position and size of the obstacle in the three-dimensional space, etc. This part corresponds to Figure 3 the fusion feature classification and regression in

[0089] The third detection result includes the above E cls and E 3Dbox .

[0090] Combined with the above embodiments, in one implementation manner, the position information includes three-dimensional space coordinates and two-dimensional plane coordinates. Correspondingly, step 105 may include: Performing a first filtering operation on the type information in the first detection result, the second detection result, and the third detection result that is less than the corresponding confidence threshold through the confidence thresholds corresponding to the point cloud feature, the image feature, and the fusion feature respectively; Performing a second filtering operation on the three-dimensional space coordinates and the two-dimensional plane coordinates in the detection result obtained after the first filtering operation through the non-maximum suppression method (Non-Maximum Suppression, NMS); Correcting the three-dimensional plane coordinates obtained after the second filtering operation according to the two-dimensional plane coordinates obtained after the second filtering operation to obtain the corrected detection result, and the corrected detection result is the target detection result.

[0091] In this embodiment, first, through the threshold T P , filtering out the detection results with low confidence generated based on the point cloud feature in the first detection result; through the threshold T I , filtering out the detection results with low confidence generated based on the image feature in the second detection result; through the threshold T E , filtering out the detection results with low confidence generated based on the fusion feature in the third detection result. The above process is the first filtering operation. Among them, the threshold T P , the threshold T I , and the threshold T E can be set according to actual needs.

[0092] Next, use non-maximum suppression NMS to filter out redundant 3D bounding boxes and 2D bounding boxes. This process is the second filtering operation, and the specific process is as follows: B 3D =NMS((E cls ,E 3Dbox ),(P cls,1 ,P 3Dbox,1 ),(P cls,2 ,P3Dbox,2 ),(P cls,3 ,P 3Dbox,3 )) B 2D =NMS((I cls,1 ,I 2Dbox,1 ),(I cls,2 ,I 2Dbox,2 ),(I cls,3 ,I 2Dbox,3 )) Then project the 3D bounding box into the 2D image dimension, and mine the possible 3D bounding boxes, so as to check whether there are some obstacles missed in the 3D detection from the perspective of the 2D image. If the 2D detection result contains obstacles that are not included in the 3D detection result, search for the detection result closest to the 2D bounding box in the 3D bounding box for supplementation (this process is the correction process), and finally output the complete 3D detection result. In this way, the information of 2D and 3D detections can be fully utilized to improve the integrity and accuracy of the detection and avoid missing obstacles.

[0093] Among them, non-maximum suppression is a post-processing technique widely used in the fields of computer vision and image processing. Its core idea is to solve the problem of repeated detections that occur during the object detection process. When the object detection algorithm makes predictions on an image, multiple overlapping or similar bounding boxes may be generated, and these bounding boxes may all correspond to the same object. The purpose of NMS is to select the optimal one from these overlapping bounding boxes, that is, the bounding box with the highest confidence, and suppress other bounding boxes with lower confidence.

[0094] In this application, considering that point cloud data can provide high-precision distance information and image data can provide rich semantic information, an obstacle detection method is designed. This detection method can simultaneously combine the advantages of point cloud data in spatial positioning and the advantages of image data in semantic recognition, and can achieve accurate detection of track obstacles in complex environments, with high accuracy and robustness.

[0095] Next, an obstacle detection device provided by this application will be described. The obstacle detection device described below can be mutually corresponding and referenced to the obstacle detection method described above.

[0096] Figure 4 is a structural block diagram of an obstacle detection device shown in an embodiment of this application. Referring to Figure 4 , the obstacle detection device 400 of this application includes: An acquisition module 401, configured to acquire point cloud data and image data collected for obstacles on the track, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; The feature extraction module 402 is configured to extract features from the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and extract features from the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; The feature fusion module 403 is configured to perform alignment processing on the point cloud features and the image features, and perform feature fusion on the aligned point cloud features and image features to obtain fused features; The first determination module 404 is configured to determine a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fused features respectively; The second determination module 405 is configured to determine a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes the type information and position information of the obstacle.

[0097] According to the obstacle detection device 400 provided by the present application, the feature fusion module 403 includes: A feature alignment sub-module, configured to project the point cloud features and the image features into the BEV space respectively to obtain an aligned point cloud BEV feature and an image BEV feature; A first feature fusion sub-module, configured to perform feature fusion on the point cloud BEV feature and the image BEV feature according to the following formulas (1)-(4) to obtain the fused features: F PI =Conv2D(Concat(F PB ,F IB )) (1) F F1 =Transformer(Q IB ,K PI ,V PI ) (2) F F2 =Transformer(Q PB ,K PI ’,V PI ’) (3) F BEV =Reshape(Linear(F F2 )) (4) where Q IB =Linear(F IB ), K PI =V PI =Linear(F PI ), Q PB =Linear(FPB ),K PI ’ = V PI ’ = Linear(F F1 ), where Linear represents a linear transformation, Concat represents feature concatenation, Conv2D represents a two-dimensional convolution operation, Transformer represents feature fusion based on an attention mechanism, and Reshape represents mapping F F2 back to the BEV dimension. F PB represents the point cloud BEV feature, and F IB represents the image BEV feature, and F BEV represents the fused feature.

[0098] According to the obstacle detection device 400 provided by the present application, the feature alignment sub-module includes: A first projection sub-module for projecting the point cloud feature into the BEV space according to a preset BEV space resolution; A second projection sub-module for projecting the image feature into the same dimension as the point cloud feature and performing feature fusion on the obtained image feature to obtain a first feature; A third projection sub-module for projecting the image feature into the BEV space according to the preset BEV space resolution based on the first feature, the depth information of the image feature, the internal parameter matrix of the acquisition device corresponding to the image data, and the transformation matrix between the coordinate system of the acquisition device corresponding to the image data and the coordinate system of the acquisition device corresponding to the point cloud data.

[0099] According to the obstacle detection device 400 provided by the present application, the first determination module 404 includes: A first determination sub-module for determining a first detection result corresponding to the obstacle through a multi-layer perceptron network according to the point cloud features at multiple different scales; A second determination sub-module for determining a second detection result corresponding to the obstacle through a two-dimensional convolution operation based on an anchor-free mechanism according to the image features at multiple different scales; A third determination sub-module for determining a third detection result corresponding to the obstacle through a fused transformation detection head network according to the fused feature.

[0100] According to the obstacle detection device 400 provided by the present application, the position information includes three-dimensional space coordinates and two-dimensional plane coordinates, and the second determination module 405 includes: The first filtering sub-module is configured to perform a first filtering operation on the type information in the first detection result, the second detection result, and the third detection result that is less than the corresponding confidence threshold through the confidence thresholds corresponding to the point cloud feature, the image feature, and the fusion feature respectively; The second filtering sub-module is configured to perform a second filtering operation on the three-dimensional spatial coordinates and the two-dimensional plane coordinates in the detection result obtained after the first filtering operation through the non-maximum suppression method; The correction sub-module is configured to correct the three-dimensional plane coordinates obtained after the second filtering operation according to the two-dimensional plane coordinates obtained after the second filtering operation to obtain a corrected detection result, and the corrected detection result is the target detection result.

[0101] According to the obstacle detection device 400 provided by the present application, the three-dimensional sparse convolution network includes a plurality of different sparse convolution layers, and the number of feature channels of each sparse convolution layer is different; the feature extraction module 402 includes: The voxel processing sub-module is configured to perform voxelization processing on the point cloud data to obtain a plurality of voxels; The encoding sub-module is configured to perform sparse encoding on each of the voxels to obtain sparse features corresponding to each of the voxels; The first feature extraction sub-module is configured to input the sparse features into the three-dimensional sparse convolution network, and extract the features of the point cloud data at multiple different scales through the feature channels of each sparse convolution layer in the three-dimensional sparse convolution network to obtain the point cloud feature.

[0102] According to the obstacle detection device 400 provided by the present application, the two-dimensional convolution network includes a plurality of different convolution layers, and the number of feature channels of each convolution layer is different. The feature extraction module 402 includes: The second feature extraction sub-module is configured to input the image data into the two-dimensional convolution network, and extract the features of the image data at multiple different scales through the feature channels of each convolution layer in the two-dimensional convolution network to obtain a second feature; The second feature fusion sub-module is configured to input the second feature into the feature pyramid network, and perform feature fusion on the second feature through the top-down feature propagation mechanism of the feature pyramid network to obtain the image feature.

[0103] Figure 5 It is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application. As Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute an obstacle detection method, which includes: obtaining point cloud data and image data collected for obstacles on the track, and the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; Performing feature extraction on the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and performing feature extraction on the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; Performing alignment processing on the point cloud features and the image features, and performing feature fusion on the aligned point cloud features and image features to obtain fused features; Respectively determining a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fused features; Determining a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes the type information and position information of the obstacle.

[0104] In addition, when the logic instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0105] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute an obstacle detection method provided by each of the above methods. The method includes: acquiring point cloud data and image data collected for an obstacle on an orbit, where the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; extracting features from the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and extracting features from the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; performing alignment processing on the point cloud features and the image features, and performing feature fusion on the aligned point cloud features and image features to obtain fused features; respectively determining a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fused features; determining a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes the type information and position information of the obstacle.

[0106] On another aspect, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an obstacle detection method provided by each of the above methods. The method includes: acquiring point cloud data and image data collected for an obstacle on an orbit, where the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; extracting features from the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and extracting features from the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; performing alignment processing on the point cloud features and the image features, and performing feature fusion on the aligned point cloud features and image features to obtain fused features; respectively determining a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud features, the image features, and the fused features; determining a target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result, where the target detection result includes the type information and position information of the obstacle.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. An obstacle detection method, characterized in that: include: Acquire point cloud data and image data collected for obstacles on the track, wherein the difference between the collection time of the point cloud data and the collection time of the image data is within a preset time difference range; Performing feature extraction on the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and performing feature extraction on the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; Aligning the point cloud features and the image features, and fusing the aligned point cloud features and image features to obtain fused features; Determine a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud feature, the image feature, and the fusion feature respectively; A target detection result corresponding to the obstacle is determined according to the first detection result, the second detection result, and the third detection result, wherein the target detection result includes type information and location information of the obstacle.

2. The obstacle detection method according to claim 1, characterized in that: The step of aligning the point cloud features and the image features, and fusing the aligned point cloud features and image features to obtain fused features includes: Projecting the point cloud features and the image features into the BEV space respectively to obtain aligned point cloud BEV features and image BEV features; According to the point cloud BEV feature and the image BEV feature, feature fusion is performed through the following formulas (1)-(4) to obtain the fused feature: F PI =Conv2D(Concat(F PB ,F IB )) (1) F F1 =Transformer(Q IB ,K PI ,V PI ) (2) F F2 =Transformer(Q PB ,K PI ’,V PI ’) (3) F BEV =Reshape(Linear(F F2 )) (4) Among them, Q IB =Linear(F IB ), K PI =V PI =Linear(F PI ), Q PB =Linear(F PB ), K PI =V PI '=Linear(F F1 ), Linear represents linear transformation, Concat represents feature concatenation, Conv2D represents two-dimensional convolution operation, Transformer represents feature fusion based on attention mechanism, and Reshape represents F F2 Mapping back to the BEV dimension, F PB represents the point cloud BEV feature, F IB represents the BEV feature of the image, F BEV Represents fusion features.

3. The obstacle detection method according to claim 2, characterized in that: The step of projecting the point cloud features and the image features to the BEV space to obtain aligned point cloud BEV features and image BEV features comprises: According to a preset BEV space resolution, projecting the point cloud features into the BEV space; Projecting the image feature to the same dimension as the point cloud feature, and performing feature fusion on the obtained image feature to obtain a first feature; Based on the first feature, the depth information of the image feature, the intrinsic parameter matrix of the acquisition device corresponding to the image data, and the conversion matrix between the coordinate system of the acquisition device corresponding to the image data and the coordinate system of the acquisition device corresponding to the point cloud data, the image feature is projected to the BEV space according to the preset BEV spatial resolution.

4. The obstacle detection method according to claim 1, characterized in that: The determining, respectively according to the point cloud feature, the image feature and the fusion feature, a first detection result, a second detection result and a third detection result corresponding to the obstacle comprises: Determine, according to the point cloud features at multiple different scales, a first detection result corresponding to the obstacle through a multi-layer perceptron network; Determining a second detection result corresponding to the obstacle through a two-dimensional convolution operation based on the image features at multiple different scales and an anchor point free mechanism; According to the fused features, a third detection result corresponding to the obstacle is determined through a fusion transformation detection head network.

5. The obstacle detection method according to claim 4, characterized in that: The position information includes three-dimensional space coordinates and two-dimensional plane coordinates; and determining the target detection result corresponding to the obstacle according to the first detection result, the second detection result, and the third detection result includes: Performing a first filtering operation on type information whose confidence level is less than the corresponding confidence level in the first detection result, the second detection result, and the third detection result according to the confidence level corresponding to each of the point cloud feature, the image feature, and the fusion feature; Performing a second filtering operation on the three-dimensional space coordinates and the two-dimensional plane coordinates in the detection result obtained after the first filtering operation by a non-maximum suppression method; According to the two-dimensional plane coordinates obtained after the second filtering operation, the three-dimensional plane coordinates obtained after the second filtering operation are corrected to obtain a corrected detection result, and the corrected detection result is the target detection result.

6. The obstacle detection method according to claim 1, characterized in that: The three-dimensional sparse convolutional network includes a plurality of different sparse convolutional layers, and each of the sparse convolutional layers has a different number of feature channels; The step of extracting features from the point cloud data using a three-dimensional sparse convolutional network to obtain point cloud features includes: voxelize the point cloud data to obtain a plurality of voxels; Performing sparse coding on each of the voxels to obtain sparse features corresponding to each of the voxels; The sparse features are input into the three-dimensional sparse convolutional network, and the features of the point cloud data at multiple scales are extracted through the feature channels of each of the sparse convolutional layers in the three-dimensional sparse convolutional network to obtain the point cloud features.

7. The obstacle detection method according to claim 1, characterized in that: The two-dimensional convolutional network includes a plurality of different convolutional layers, each of which has a different number of feature channels. The image data is subjected to feature extraction through the two-dimensional convolutional network and the feature pyramid network to obtain image features, including: Inputting the image data into the two-dimensional convolutional network, extracting features of the image data at multiple scales through feature channels of each convolutional layer in the two-dimensional convolutional network, and obtaining a second feature; The second feature is input into the feature pyramid network, and the second feature is fused through a top-down feature propagation mechanism of the feature pyramid network to obtain the image feature.

8. An obstacle detection device, characterized in that: include: An acquisition module, used to acquire point cloud data and image data collected for obstacles on the track, wherein the difference between the acquisition time of the point cloud data and the acquisition time of the image data is within a preset time difference range; A feature extraction module, used to extract features from the point cloud data through a three-dimensional sparse convolutional network to obtain point cloud features, and to extract features from the image data through a two-dimensional convolutional network and a feature pyramid network to obtain image features; A feature fusion module is used to align the point cloud features and the image features, and fuse the aligned point cloud features and image features to obtain fused features; A first determination module, used to determine a first detection result, a second detection result, and a third detection result corresponding to the obstacle according to the point cloud feature, the image feature, and the fusion feature, respectively; The second determination module is used to determine the target detection result corresponding to the obstacle according to the first detection result, the second detection result and the third detection result, wherein the target detection result includes the type information and the position information of the obstacle.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, an obstacle detection method as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, an obstacle detection method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Target detection method, electronic equipment, storage medium and vehicle

    CN121095537A