Vision and laser radar fused obstacle detection method and device

Through the obstacle detection method that integrates vision and lidar, combined with the target detection network and lidar point cloud data, the problem of low reliability of obstacle detection in front of the train is solved, and accurate distance detection and safety improvement of obstacles is achieved.

CN119992503APending Publication Date: 2025-05-13WUHU PORT STORAGE & TRANSPORTATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411761531.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the detection method of obstacles in front of the train is not reliable, and it is difficult to ensure the stability and reliability of the forward environment.

Method used

The obstacle detection method that integrates vision and lidar is adopted to process forward images through the target detection network, and combine lidar point cloud data to achieve accurate distance detection of obstacles.

Benefits of technology

It improves the accuracy and reliability of detection of forward obstacles in trains in rail transit scenarios, reduces the dependence on drivers' attention, and reduces the risk of collision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992503A_ABST
    Figure CN119992503A_ABST
Patent Text Reader

Abstract

The invention provides a vision and laser radar fused obstacle detection method and device, and the method comprises the steps: obtaining a visual image; inputting the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; performing normalization operation on the plurality of detection targets to obtain a plurality of normalized detection targets; image classification is carried out on each normalized detection target in the plurality of normalized detection targets, target classifications corresponding to the plurality of detection targets are obtained, and the target classifications comprise obstacle targets; the method comprises the following steps: mapping an obstacle target to a point cloud space based on projection transformation through a laser radar of a train to obtain a view cone space; distance clustering is carried out on the radar point clouds in the view cone space, and a cluster with the maximum density is determined to serve as obstacle point clouds; determining the distance between the obstacle and the train based on the obstacle point cloud; according to the invention, accurate distance detection of forward obstacles in rail transit can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image detection technology, and in particular to an obstacle detection method and device integrating vision and laser radar. Background Art

[0002] As the passenger volume of urban rail transit continues to grow, the departure interval of trains is gradually shortened. Although shorter intervals increase the transportation capacity, they also increase the risk of train collisions. During the operation of the train, the detection of obstacles ahead usually relies on the naked eye observation of the driver.

[0003] However, the driver's attention is easily affected by subjective factors such as fatigue, which makes it difficult to ensure the stability and reliability of observation of the environment ahead.

[0004] It can be seen that the method of detecting obstacles in front of the train in the related art has a technical problem of low reliability. Summary of the invention

[0005] The present invention provides a method and device for obstacle detection by integrating vision and laser radar, so as to solve the defect of low reliability of the detection method of obstacles in front of trains in the prior art and realize accurate distance detection of obstacles in front of rail transit.

[0006] The present invention provides an obstacle detection method integrating vision and laser radar, comprising the following steps: obtaining a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles; inputting the visual image into a preset target detection network to obtain multiple detection targets output by the target detection network; performing normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets; performing image classification on each of the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets respectively, wherein the target classifications include obstacle targets; mapping the obstacle target to a point cloud space based on a projection transformation through the laser radar of the train to obtain a visual cone space; performing distance clustering on the radar point cloud in the visual cone space to determine the cluster with the largest density as the obstacle point cloud; determining the distance between the obstacle and the train based on the obstacle point cloud.

[0007] According to a method for obstacle detection by integrating vision and laser radar provided by the present invention, the visual image is input into a preset target detection network to obtain a plurality of detection targets output by the target detection network, including: normalizing the visual image to obtain a target input image, wherein the size of the target input image is ,in, represents the height of the target input image, is the width of the target input image, 3 represents the number of channels of the target input image; the target input image is feature extracted through the backbone network of the target detection network to obtain a feature map, wherein the backbone network is a fully convolutional neural network, and the size of the feature map is , Represents the downsampling rate of the feature map; through the regression of the heat map, the center offset and the target frame size, the detection target is generated based on the feature map to obtain multiple detection targets.

[0008] According to a vision and laser radar fusion obstacle detection method provided by the present invention, the detection target generation is performed based on the feature map through the regression of the heat map, the center offset and the target frame size to obtain multiple detection targets, including: extracting features from the feature map through a first two-layer neural network to obtain a category feature map, wherein the size of the category feature map is , represents the number of predicted categories; performing maximum pooling on the category feature map of each channel to obtain a heat map, which is used to determine the center point of each detection target in the visual image; performing feature extraction on the feature map through a second two-layer neural network to obtain an offset feature map, wherein the size of the offset feature map is , the offset feature map is used to predict the offset of the center point of each detection target in the x direction and the y direction; the feature map is extracted by a third two-layer neural network to obtain a predicted feature map, wherein the size of the predicted feature map is The predicted feature map is used to predict the width and height of each detection target with reference to the center point of each detection target; detection targets are generated based on the heat map, the offset feature map and the predicted feature map to obtain multiple detection targets.

[0009] According to a vision and lidar fusion obstacle detection method provided by the present invention, the normalization operation is performed on the multiple detection targets respectively to obtain multiple normalized detection targets, including: scale normalization is performed on the multiple detection targets respectively to obtain detection targets of multiple target sizes; pixel normalization is performed on the detection targets of the multiple target sizes to obtain multiple normalized detection targets.

[0010] According to a vision and lidar fusion obstacle detection method provided by the present invention, image classification is performed on each of the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets, including: downsampling each of the multiple normalized detection targets through a first 2D convolution to obtain a feature image of each detection target; feature extraction is performed on the feature image of each detection target through multiple bottleneck structures to obtain deep features of each detection target; image classification is performed on the deep features of each detection target through mean pooling and a second 2D convolution to obtain target classifications corresponding to the multiple inspection targets.

[0011] According to a vision and laser radar fusion obstacle detection method provided by the present invention, the radar point cloud in the cone space is clustered by distance to determine the cluster with the largest density as the obstacle point cloud, including: sorting the radar point cloud in the cone space in descending order according to the distance from the laser radar to obtain a radar point distance sequence; clustering based on the distance between radar points in the radar point distance sequence according to a preset distance threshold to obtain multiple clusters; and determining the cluster containing the most radar points among the multiple clusters as the obstacle point cloud.

[0012] The present invention also provides an obstacle detection device integrating vision and laser radar, comprising the following modules: an acquisition module, used to acquire a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes an obstacle; a detection module, used to input the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; a normalization module, used to perform normalization operations on the plurality of detection targets respectively to obtain a plurality of normalized detection targets; a classification module, used to perform image classification on each of the plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets respectively, wherein the target classifications include obstacle targets; a projection module, used to map the obstacle target to a point cloud space based on a projection transformation through the laser radar of the train to obtain a visual cone space; a clustering module, used to perform distance clustering on the radar point cloud in the visual cone space to determine the cluster with the largest density as the obstacle point cloud; a determination module, used to determine the distance between the obstacle and the train based on the obstacle point cloud.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an obstacle detection method integrating vision and laser radar as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the obstacle detection method of integrating vision and laser radar as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned obstacle detection methods integrating vision and laser radar.

[0016] The obstacle detection method and device integrating vision and laser radar provided by the present invention can automatically identify and detect multiple targets in the image by processing the forward image collected by the train using a preset target detection network; normalizing the multiple detected targets can eliminate the differences in size, shape, etc. between different targets, which helps to improve the stability and accuracy of the subsequent classification algorithm; performing image classification on the normalized detected targets to distinguish obstacle targets, which helps to focus only on obstacles directly related to the safety of train travel in subsequent steps and reduce unnecessary calculations and processing; obstacle targets are removed from the two-dimensional image through the laser radar of the train. The space is mapped to the three-dimensional point cloud space to form a cone space, which makes the positioning of obstacles more accurate and three-dimensional; the radar point cloud in the cone space is clustered by distance, and the cluster with the largest density is determined as the obstacle point cloud, which can effectively extract key information related to the obstacle from a large amount of point cloud data, providing a basis for subsequent distance calculation; the distance between the obstacle and the train is determined based on the obstacle point cloud. Therefore, by combining image processing and lidar, accurate detection and distance determination of obstacles in front of the train in the rail transit scene are achieved, thereby solving the technical problem of low reliability in the detection method of obstacles in front of the train in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a flow chart of the obstacle detection method integrating vision and laser radar provided by the present invention.

[0019] Figure 2 It is a schematic diagram of the structure of the target detection network provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the affine transformation of the size of the detection target provided by the present invention.

[0021] Figure 4 It is a structural schematic diagram of the image classification network provided by the present invention.

[0022] Figure 5 It is a structural schematic diagram of the viewing cone space provided by the present invention.

[0023] Figure 6 It is a schematic diagram of the overall framework of the obstacle detection method of vision and laser radar fusion provided by the present invention.

[0024] Figure 7 It is a structural schematic diagram of the obstacle detection device integrating vision and laser radar provided by the present invention.

[0025] Figure 8 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] With the continuous growth of urban rail transit passenger volume, the departure interval of trains has gradually shortened. Although the shorter interval has increased the transportation capacity, it has also increased the risk of train collisions. During the operation of the train, the detection of obstacles ahead usually relies on the naked eye observation of the driver. However, the driver's attention is easily affected by subjective factors such as fatigue, which makes it difficult to ensure the stability and reliability of the observation of the environment ahead. Therefore, the automatic detection technology of forward obstacles has become the key to ensuring safe operation. It can effectively reduce the dependence on the driver, detect obstacles in advance and issue early warnings, and guide the train to take emergency braking and other countermeasures when necessary, thereby reducing the risk of collision and improving safety.

[0028] At present, many scholars have conducted research on forward obstacle detection in rail transit. For example, a related technology proposes a multi-data fusion train obstacle detection method. This method obtains images and point cloud data in front of the train through a binocular stereo vision camera and a laser radar, performs obstacle detection separately, and then uses spatiotemporal fusion technology to combine the two detection results, and uses a decision model to determine whether the obstacle has an intrusion risk. Another related technology projects the target detection results of the laser point cloud data onto the image plane for verification based on the visual target detection of the image in front of the train, and finally obtains information such as the distance, category and position of the obstacle. Another related technology extracts the track area through image processing, accurately obtains the obstacle detection limit, and uses a laser radar to accurately detect obstacles on the track.

[0029] However, existing research has not fully considered the trade-off between missed detection and false detection of image targets. Meanwhile, the sparsity and noise problems of point cloud data are still the main obstacles affecting detection accuracy.

[0030] Therefore, in view of the shortcomings of existing methods in terms of balance between missed and false detection of obstacles and distance measurement accuracy, the present invention proposes an obstacle detection method based on vision and laser radar technology. This method reduces missed detection through a target detection network, reduces the false detection rate by combining an image classification network, and uses a two-stage method to comprehensively extract targets in the image. In addition, in order to address the problem of inaccurate distance detection caused by the lack of depth information on the image plane, laser radar point cloud data is used to supplement it, ultimately achieving accurate distance detection of forward obstacles in rail transit.

[0031] The present invention focuses on the problem of difficult detection of forward obstacles and inaccurate detection distance in rail transit scenes, and proposes a vision and laser radar fusion obstacle detection method based on vision and laser radar technology; this method uses visual images as input, balances the problem of missed and false detection in image target detection through target detection network, normalization processing and image classification network, and obtains preliminary detection results of obstacles. Subsequently, the radar point cloud in the cone space is adaptively clustered through cone projection, so as to accurately lock the position of the obstacle in three-dimensional space and obtain accurate distance information.

[0032] Optionally, the obstacle detection method integrating vision and laser radar in the embodiment of the present application can be executed by a server, by a terminal device, or jointly by a server and a terminal device, taking the example of the obstacle detection method integrating vision and laser radar in the present embodiment being executed by a train-mounted device.

[0033] Figure 1 FIG. 1 is a flow chart of the obstacle detection method of the present invention by integrating vision and laser radar. Figure 1 As shown, the method includes the following: Step 101, obtaining a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles.

[0034] In rail transit scenarios, trains are usually equipped with forward-facing cameras that can capture real-time visual images of the train ahead. Forward-facing cameras usually have high resolution and a wide-angle field of view to ensure that the details of the track and its surroundings can be clearly captured.

[0035] Rail transit includes subways, light rails, trains, etc. Its operating environment is complex and changeable, including urban streets, viaducts, tunnels, fields, etc. In rail transit scenarios, trains may face various obstacles, such as construction equipment, derailed vehicles, pedestrians, animals, etc.

[0036] When the train is moving, the forward-facing camera captures visual images of possible obstacles ahead. The presentation of these obstacles in the visual image may vary depending on factors such as distance, light, and obstacle type, but they usually appear as shapes, colors, or textures that are different from the background environment.

[0037] Step 102, input the visual image into a preset target detection network to obtain multiple detection targets output by the target detection network.

[0038] Object detection network is an important technology in the field of computer vision. It is used to detect the position and size of the target object in a given image or video and perform related tasks such as classification or recognition.

[0039] The basic framework of the object detection network usually includes three main parts: object localization, object classification, and object box regression.

[0040] Object localization is used to accurately locate the position and size of an object in an image. This usually involves using deep learning models such as convolutional neural networks to extract image features and use these features to determine the bounding box of the target object.

[0041] Object classification is used to match detected objects to pre-defined categories. This is usually performed using fully connected layers or classifiers such as support vector machines (SVMs) or softmax classifiers, which take the extracted features as input and output class probabilities.

[0042] Target box regression is used to correct the position and size of the target box based on the predicted position offset to improve detection accuracy. This usually involves a regression process to optimize the model by minimizing the difference between the predicted bounding box and the true bounding box.

[0043] According to a vision and laser radar fusion obstacle detection method provided by the present invention, a visual image is input into a preset target detection network to obtain multiple detection targets output by the target detection network, including: Normalize the visual image to obtain the target input image, where the size of the target input image is ,in, represents the height of the target input image, is the width of the target input image, and 3 represents the number of channels of the target input image; The target input image is feature extracted through the backbone network of the target detection network to obtain a feature map, where the backbone network is a fully convolutional neural network and the size of the feature map is , Indicates the downsampling rate of the feature map; Through the regression of heat map, center offset and target frame size, detection targets are generated based on the feature map to obtain multiple detection targets.

[0044] In the embodiment of the present invention, target detection is used to obtain the most comprehensive obstacle detection result from the visual image.

[0045] refer to Figure 2 , Figure 2 It is a schematic diagram of the structure of the target detection network provided by the present invention.

[0046] Taking Centernet network as an example, the network structure is as follows Figure 2 As shown, the specific process is: Normalize the input image and transform it to size, reducing the scale imbalance problem in the subsequent network.

[0047] The features of the image are obtained through a fully convolutional neural network. For example, DLA-34 is used as the backbone network to quickly extract the target and obtain , where R is the downsampling rate of the feature map.

[0048] The detection target is generated through three branches: heat map, center offset and target frame size regression, and the detection target is used as the extracted obstacle.

[0049] Here, the heat map branch can generate a probability map related to the target location, which is used to indicate the approximate location of the target in the image; this helps the model locate the target more accurately.

[0050] By predicting the center offset of the target, the model can more accurately determine the specific location of the target, which helps to improve the accuracy of target detection.

[0051] The target box size regression branch can predict the bounding box size of the target, which helps the model to accurately outline the target while detecting the target.

[0052] Through the embodiments of the present invention, the consistency of the pixel values ​​of the input image is ensured through normalization processing; the fully convolutional neural network can effectively extract rich feature information from the input image; the detection target is realized through the three branches of heat map, center offset and target frame size regression, which improves the detection accuracy and efficiency.

[0053] According to a vision and laser radar fusion obstacle detection method provided by the present invention, a detection target is generated based on a feature map through a heat map, a center offset, and a target frame size regression, and multiple detection targets are obtained, including: The feature map is extracted through the first two layers of neural network to obtain the category feature map, where the size of the category feature map is , Indicates the number of predicted categories; Perform maximum pooling on the category feature map of each channel to obtain a heat map, which is used to determine the center point of each detected target in the visual image; The feature map is extracted through the second two-layer neural network to obtain the offset feature map, where the size of the offset feature map is ,The offset feature map is used to predict the offset of the center point of each detected target in the x-direction and y-direction; The feature map is extracted through the third two-layer neural network to obtain the predicted feature map, where the size of the predicted feature map is ,The predicted feature map is used to predict the width and height of each detected object with the center point of each detected object as a reference; Detection targets are generated based on the heat map, the offset feature map, and the predicted feature map to obtain multiple detection targets.

[0054] In the embodiment of the present invention, the three branches are specifically operated as follows: Heatmap branch. The main purpose of the heatmap generation branch is to determine the center position of the target in the image, so as to effectively extract the potential target in the image. In the heatmap branch, the feature map generated by the fully convolutional neural network is first processed by a two-layer neural network to generate a The feature map of , where C represents the number of predicted categories. Then, a 3×3 maximum pooling operation is applied to the feature map of each channel to generate the corresponding heat map.

[0055] Center offset branch. During target detection, the predicted target center usually has a certain offset. To solve this problem, the embodiment of the present invention introduces a center offset branch for calculating and correcting this offset. In the center offset branch, a two-layer neural network is used to process the feature map generated by the fully convolutional neural network to obtain a The feature map is used to predict the offset of each center point in the x and y directions.

[0056] Target box size regression branch. Based on the heat map branch and the center offset branch, it is further necessary to predict the size of the target, that is, the width and height of the target box. The embodiment of the present invention achieves this task by introducing a size box regression branch.

[0057] Specifically, the size box regression branch processes the feature map generated by the full convolutional neural network through a two-layer neural network to generate a A feature map that uses the center point of each object as a reference to predict the width and height of the object.

[0058] Through the embodiment of the present invention, first, highly responsive pixels are found from the heat map as candidate center points. Then, their positions are adjusted according to the offsets of these center points. Finally, the target frame is generated using the corresponding size regression results. Using the three branches of heat map, center offset and target frame size regression can effectively generate detection targets, thereby extracting obstacles.

[0059] Step 103, performing normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets.

[0060] In an embodiment of the present invention, a normalization operation is performed on each inspection target among multiple detection targets so that a subsequent image classification network can better acquire image features. Specifically, the normalization operation of the embodiment of the present invention includes scale normalization and pixel normalization.

[0061] According to a vision and laser radar fusion obstacle detection method provided by the present invention, a plurality of detection targets are respectively normalized to obtain a plurality of normalized detection targets, including: Normalize the scales of multiple detection targets respectively to obtain detection targets of multiple target sizes; Pixel normalization is performed on detection targets of multiple target sizes to obtain multiple normalized detection targets.

[0062] refer to Figure 3 , Figure 3 It is a schematic diagram of the affine transformation of the size of the detection target provided by the present invention.

[0063] An affine transformation is geometrically defined as an affine transformation or affine mapping between two vector spaces, consisting of a non-singular linear transformation (a transformation performed using a linear function) followed by a translation transformation.

[0064] Since the scales of multiple detection targets output by the target detection network vary greatly, the size of targets at close range is usually larger, while the size of targets at far range is smaller. The embodiment of the present invention performs scale normalization through affine transformation, and the process is as follows: Figure 3 shown.

[0065] Specifically, for each extracted detection target, scale normalization is performed to unify its size to the standard scale of H×W×3, which serves as the input of the subsequent image classification network.

[0066] In the embodiment of the present invention, the input detection target (size is ) to complete the scale, and the completed size is ,in , that is, the size of the completion is the larger scale of the original input. Apply an affine transformation to normalize the size of the completion to .

[0067] Based on the scale-normalized image, the embodiment of the present invention further performs pixel normalization processing on the extracted targets (detection targets of multiple target sizes). The standard normalization method is used, that is, ,in, represents the detected target after pixel normalization, The detection target represents the size of the target (the detection target after scale normalization); the pixel value of each target is mapped to the range of (0, 1). Pixel normalization effectively filters out outliers and extreme values ​​in the image, thereby avoiding adverse effects on the subsequent image classification network.

[0068] Through the embodiment of the present invention, by processing scale normalization and pixel normalization, it can be ensured that when the subsequent image classification network receives input, all objects have the same size and pixel value range, which helps the network to better extract image features and improve classification accuracy.

[0069] Step 104 , performing image classification on each of the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets, wherein the target classifications include obstacle targets.

[0070] In an embodiment of the present invention, image classification is performed on each of the multiple normalized detection targets through an image classification network to obtain target classifications corresponding to the multiple inspection targets.

[0071] According to the specific application scenario and data set characteristics, select the appropriate image classification network architecture, such as ResNet, VGG, Inception, etc. in the convolutional neural network (CNN). According to the requirements of the specific task, fine-tune the pre-trained model to adapt to the new data set and classification task.

[0072] The normalized detection target is input into the image classification network, and the features are extracted and classified through the forward propagation process of the network to obtain the probability or score of each detection target belonging to each category output by the image classification network.

[0073] According to a vision and laser radar fusion obstacle detection method provided by the present invention, image classification is performed on each of a plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets, including: Downsampling each of the multiple normalized detection targets through a first 2D convolution to obtain a feature image of each detection target; The feature image of each detection target is extracted through multiple bottleneck structures to obtain the deep features of each detection target; The deep features of each detection target are classified by mean pooling and the second 2D convolution to obtain target classifications corresponding to multiple inspection targets.

[0074] Here, the bottleneck structure in deep learning, especially in convolutional neural networks (CNN), is a special network layer design that aims to reduce the amount of computation and the number of parameters while maintaining or improving the performance of the model.

[0075] The bottleneck structure usually includes: a 1x1 convolution layer to reduce the number of channels of the input feature map (i.e., dimensionality reduction), a 3x3 convolution layer to extract features, and another 1x1 convolution layer to restore (or increase) the number of channels of the feature map to match the input dimension of the residual connection. In this way, the output of the bottleneck structure can be added to the original input to form a residual connection.

[0076] refer to Figure 4 , Figure 4 It is a structural schematic diagram of the image classification network provided by the present invention.

[0077] Each normalized detected target is classified to realize the reconfirmation of each extracted target. The present invention takes the mobilenet-V2 network as an example. The image classification network structure is as follows: Figure 4 As shown, the specific steps are as follows: Each of the multiple normalized detection targets is downsampled through 2D convolution; feature extraction is performed through a series of bottleneck structures; mean pooling and multiple 2D convolutions (before and after mean pooling, respectively) are applied to achieve the final target classification.

[0078] For example, the first 2D convolution layer (i.e., the first 2D convolution) is applied to each normalized detection target. This convolution layer not only extracts image features, but also achieves downsampling (i.e., reduces the size of the image) through sliding and step size settings of the convolution kernel, and outputs feature images corresponding to each detection target. These feature images have lower resolution than the original images but contain richer feature information.

[0079] The feature image is further processed using multiple bottleneck structures (also called residual blocks or bottleneck blocks). The bottleneck structure usually consists of a 1x1 convolution (for reducing the number of channels), a 3x3 convolution (for feature extraction), and another 1x1 convolution (for restoring the number of channels). This structure can effectively reduce the amount of calculation and prevent the gradient disappearance problem; it outputs the deep features of each detected target, which contain more abstract and advanced information in the image.

[0080] The deep features are averaged and pooled to further reduce the size of the feature map and retain important features. Then the second 2D convolution layer (i.e., the second 2D convolution) is applied to further extract and integrate the features after average pooling. Finally, the output of the convolution layer is converted into a probability distribution for each category through a fully connected layer (usually used in conjunction with the Softmax activation function), and the target classification results corresponding to multiple detection targets are obtained.

[0081] Through the embodiments of the present invention, accurate image classification of normalized detection targets can be achieved.

[0082] Step 105 , using the laser radar of the train, the obstacle target is mapped to the point cloud space based on the projection transformation to obtain the cone space.

[0083] refer to Figure 5 , Figure 5 It is a structural schematic diagram of the visual cone space provided by the present invention, which includes a camera, a radar sensor, a target for visual recognition, and a target recognized in a 3D space.

[0084] In the embodiment of the present invention, through the joint calibration of laser radar and visual image, the detected obstacle target is mapped to the point cloud space by using projection transformation to form a cone space such as Figure 5 As shown in Figure 2. In this space, the cluster point with the largest density corresponds to the detected obstacle point cloud.

[0085] Joint calibration is a key step to ensure that the data between different sensors (radar and camera) can be accurately aligned. This process usually includes the following steps: First, the position and attitude (i.e., position and rotation angle) of the radar and camera relative to the same reference point need to be accurately measured.

[0086] Secondly, each sensor is calibrated individually to obtain its internal parameters (such as the focal length and distortion coefficient of the camera, and the scanning angle and resolution of the radar).

[0087] Finally, by taking an image containing known feature points (such as corner points on the calibration plate) and acquiring point cloud data at the same time, the relative transformation matrix (i.e., external parameters) between the radar and the camera is calculated so that the pixels in the image can accurately correspond to the three-dimensional points in the point cloud.

[0088] After the calibration between sensors is completed, the obstacle targets detected in the image can be mapped to the point cloud space using projection transformation.

[0089] For example, in image space, obstacles are detected using image processing algorithms (such as object detection, segmentation, etc.). For each obstacle detected, if the camera provides depth information (such as through stereo vision or depth camera), its depth value can be directly obtained; otherwise, other sensors (such as radar) or algorithms (such as monocular depth estimation) are needed to estimate the depth. Using the depth information and the camera's intrinsic parameters, the two-dimensional obstacle coordinates in the image are converted to coordinates in three-dimensional space. The three-dimensional coordinates are converted to coordinates in the radar point cloud coordinate system to obtain the position of the obstacle in the point cloud space.

[0090] The view cone space is a three-dimensional space region that represents the set of all points that can be observed from the perspective of the camera and radar sensor. In the view cone space, the clustered points with the highest density correspond to the detected obstacle point cloud. This is because each point in the radar point cloud represents a physical point in the environment, and obstacles are usually composed of multiple closely adjacent points. Therefore, through cluster analysis (such as DBSCAN, K-means and other algorithms), high-density areas can be identified in the point cloud space, which correspond to obstacles.

[0091] Through the embodiments of the present invention, through the joint calibration and projection transformation of radar and image, the obstacle target detected in the image can be mapped to the point cloud space to form a cone space. In this space, the point cloud cluster with the highest density can be identified through cluster analysis, so as to determine the position and shape of the obstacle. This method combines the advantages of radar and camera, and improves the accuracy and robustness of obstacle detection in the environment.

[0092] Step 106 , performing distance clustering on the radar point cloud in the viewing cone space, and determining the cluster with the largest density as the obstacle point cloud.

[0093] In an embodiment of the present invention, the radar point cloud in the viewing cone space is adaptively clustered based on the distance between radar points, and the obstacle point cloud is extracted, thereby achieving accurate measurement of the obstacle distance.

[0094] According to a vision and laser radar fusion obstacle detection method provided by the present invention, the radar point cloud in the cone space is clustered by distance, and the cluster with the largest density is determined as the obstacle point cloud, including: The radar point cloud in the cone space is sorted in descending order according to the distance from the laser radar to obtain the radar point distance sequence; According to a preset distance threshold, clustering is performed based on the distances between radar points in the radar point distance sequence to obtain multiple clusters; A cluster containing the most radar points among the multiple clusters is determined as the obstacle point cloud.

[0095] In an embodiment of the present invention, the radar point cloud in the viewing cone space is adaptively clustered to extract the obstacle point cloud, thereby achieving accurate measurement of the obstacle distance. The specific process is as follows: Step 1: sort the radar point cloud within the viewing cone in descending order according to the distance from the vehicle-mounted radar.

[0096] Here, the radar point cloud within the cone of vision is first sorted in descending order according to the distance from the vehicle-mounted radar, so that potential obstacles that are closer can be processed first, which is especially important for emergency braking or obstacle avoidance decisions.

[0097] Step 2: Initialize a cluster and use the first radar point in the cone as the seed point. Calculate the distance between each radar point in the cone and the seed point in turn. If the distance is less than the set threshold, add the point to the current cluster and mark it as a clustered point to avoid repeated calculations. At the same time, update the number of points in the cluster.

[0098] In the embodiment of the present invention, by marking the clustered points, repeated calculation of each radar point in the clustering process is avoided, thereby improving the operation efficiency of the algorithm.

[0099] Step 3: If the distance between a radar point and the seed point is greater than the threshold, a new cluster is created and the radar point is used as the seed point of the new cluster. The distances between other radar points in the cone and the new seed point are calculated according to the steps in step 2.

[0100] In the embodiment of the present invention, clustering is started with the first radar point as a seed point and gradually expanded to other points. This gradual expansion method helps to reduce unnecessary calculation overhead.

[0101] Step 4: Repeat steps 2 to 3 until all radar points are classified into a cluster.

[0102] Step 5: After completing the radar point clustering, select the cluster containing the most radar points as the point cloud of the obstacle, so as to accurately obtain the distance between the obstacle and the train.

[0103] In the embodiment of the present invention, after radar point clustering is completed, a cluster containing the most radar points is selected as the point cloud of the obstacle. The cluster represents the largest obstacle, so that the distance between the obstacle and the train can be accurately measured.

[0104] Through the embodiments of the present invention, by performing distance-based clustering analysis on the radar point cloud within the visual cone, the point cloud belonging to the same obstacle can be accurately aggregated into a cluster, thereby effectively distinguishing different obstacles; even if there is noise or interference in the radar point cloud, the obstacle can be accurately identified, thereby improving the robustness of the system.

[0105] Step 107: Determine the distance between the obstacle and the train based on the obstacle point cloud.

[0106] In the embodiment of the present invention, the position and shape information of the obstacle are extracted from the identified obstacle point cloud.

[0107] After determining the position and shape of the obstacle, the distance between the obstacle and the train can be calculated using geometric relationships. For example, the current position of the train is determined using the positioning system on the train (such as GPS, odometer, etc.). Based on the obstacle point cloud data, the obstacle position (for example, the center of mass position or the point position closest to the train) is calculated. The distance calculation formula in three-dimensional space (such as Euclidean distance) is used to calculate the distance between the obstacle position and the train position.

[0108] Through the embodiment of the present invention, a method for obstacle detection for rail transit scenes is proposed by combining vision and laser radar technology. The method can not only fully identify targets in image data, but also effectively make up for the shortcomings of image target detection in terms of distance information accuracy.

[0109] refer to Figure 6 , Figure 6 It is a schematic diagram of the overall framework of the obstacle detection method of the vision and laser radar fusion provided by the present invention, which includes: performing cone projection on the radar point cloud; performing target detection, normalization processing, and image classification on the visual image; and performing adaptive clustering based on image classification and cone projection to obtain detection results.

[0110] The present invention belongs to the field of rail transit autonomous driving technology, and aims to solve the accuracy and distance estimation problems of forward obstacle detection in rail transit scenarios. By combining visual images and lidar data, the limitations of a single sensor can be overcome and the detection accuracy and range of obstacles can be improved. This method effectively solves the problem of insufficient accuracy of traditional visual detection in obstacle detection, and at the same time, through radar point cloud adaptive clustering technology, it provides more accurate three-dimensional spatial position and distance information, thereby providing technical support for improving the safety of rail transit systems.

[0111] The present invention can effectively solve the problem of insufficient accuracy of a single visual detection method in long-distance obstacle detection, overcome the limitations of a single sensor, and thus significantly improve the accuracy and range of obstacle detection.

[0112] The present invention adopts a two-stage obstacle detection strategy of target detection network and image classification network, and provides a reliable obstacle recognition solution. In forward obstacle detection, there is often a trade-off between misidentification and missed identification. The target detection network is used to extract potential obstacles, reduce missed identification, and the extracted targets are classified through the classification network, thereby improving the stability of the detection system and achieving efficient, long-distance and high-reliability obstacle detection.

[0113] The adaptive clustering algorithm of the present invention can effectively suppress the interference of radar noise points on the detection results. Through the adaptive clustering of the cone space in the perspective projection process, non-target radar points and noise points are effectively filtered, the interference caused by factors such as calibration errors is solved, and the detection accuracy is improved.

[0114] The obstacle detection device integrating vision and laser radar provided by the present invention is described below. The obstacle detection device integrating vision and laser radar described below and the obstacle detection method integrating vision and laser radar described above can be referred to each other.

[0115] refer to Figure 7 , Figure 7 It is a structural schematic diagram of the obstacle detection device integrating vision and laser radar provided by the present invention.

[0116] An acquisition module 701 is used to acquire a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles; A detection module 702 is used to input the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; A normalization module 703 is used to perform normalization operations on multiple detection targets respectively to obtain multiple normalized detection targets; A classification module 704 is used to perform image classification on each of the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets, wherein the target classifications include obstacle targets; The projection module 705 is used to map the obstacle target to the point cloud space based on the projection transformation through the laser radar of the train to obtain the cone space; The clustering module 706 is used to perform distance clustering on the radar point cloud in the viewing cone space and determine the cluster with the largest density as the obstacle point cloud; The determination module 707 is used to determine the distance between the obstacle and the train based on the obstacle point cloud.

[0117] Specifically, the obstacle detection device integrating vision and laser radar provided by the present invention can implement all the method steps implemented in the obstacle detection method embodiment integrating vision and laser radar, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those of the method embodiment will not be described in detail here.

[0118] Figure 8 is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 8 As shown, the electronic device may include: a processor (processor) 810, a communication interface (Communications Interface) 820, a memory (memory) 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logic instructions in the memory 830 to execute the obstacle detection method integrating vision and lidar, the method comprising: acquiring a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles; inputting the visual image into a preset target detection network to obtain multiple detection targets output by the target detection network; performing normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets; performing image classification on each of the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets respectively, wherein the target classifications include obstacle targets; mapping the obstacle target to the point cloud space based on the projection transformation through the lidar of the train to obtain the visual cone space; performing distance clustering on the radar point cloud in the visual cone space to determine the cluster with the largest density as the obstacle point cloud; and determining the distance between the obstacle and the train based on the obstacle point cloud.

[0119] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0120] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the obstacle detection method of vision and laser radar fusion provided by the above methods, the method comprising: acquiring a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes an obstacle; inputting the visual image into a preset target detection network to obtain multiple detection targets output by the target detection network; performing normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets; performing image classification on each normalized detection target in the multiple normalized detection targets to obtain target classifications corresponding to the multiple inspection targets, wherein the target classifications include obstacle targets; mapping the obstacle target to the point cloud space based on the projection transformation through the laser radar of the train to obtain a cone space; performing distance clustering on the radar point cloud in the cone space to determine the cluster with the largest density as the obstacle point cloud; and determining the distance between the obstacle and the train based on the obstacle point cloud.

[0121] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the obstacle detection method of the above-mentioned methods by integrating vision and laser radar, the method comprising: acquiring a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes an obstacle; inputting the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; performing normalization operations on the plurality of detection targets respectively to obtain a plurality of normalized detection targets; performing image classification on each of the plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets respectively, wherein the target classifications include obstacle targets; mapping the obstacle target to a point cloud space based on a projection transformation through a laser radar of the train to obtain a visual cone space; performing distance clustering on the radar point cloud in the visual cone space to determine a cluster with the largest density as an obstacle point cloud; and determining the distance between the obstacle and the train based on the obstacle point cloud.

[0122] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An obstacle detection method integrating vision and laser radar, characterized in that: include: Acquire a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles; Inputting the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; Performing normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets; Performing image classification on each of the plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets, wherein the target classifications include obstacle targets; By using the laser radar of the train, the obstacle target is mapped to the point cloud space based on projection transformation to obtain the visual cone space; Performing distance clustering on the radar point cloud in the cone space, and determining the cluster with the largest density as the obstacle point cloud; Based on the obstacle point cloud, a distance between the obstacle and the train is determined.

2. The obstacle detection method of vision and laser radar fusion according to claim 1 is characterized in that: The step of inputting the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network includes: Normalize the visual image to obtain a target input image, wherein the size of the target input image is ,in, represents the height of the target input image, is the width of the target input image, and 3 represents the number of channels of the target input image; The target input image is feature extracted through the backbone network of the target detection network to obtain a feature map, wherein the backbone network is a fully convolutional neural network and the size of the feature map is , represents the downsampling rate of the feature map; Through the regression of the heat map, the center offset and the target frame size, the detection target is generated based on the feature map to obtain multiple detection targets.

3. The obstacle detection method of vision and laser radar fusion according to claim 2 is characterized in that: The detection target is generated based on the feature map through the regression of the heat map, the center offset and the target frame size, and multiple detection targets are obtained, including: The feature map is subjected to feature extraction by the first two-layer neural network to obtain a category feature map, wherein the size of the category feature map is , Indicates the number of predicted categories; Performing maximum pooling on the category feature map of each channel to obtain a heat map, wherein the heat map is used to determine the center point of each detection target in the visual image; The feature map is subjected to feature extraction by a second two-layer neural network to obtain an offset feature map, wherein the size of the offset feature map is , the offset feature map is used to predict the offset of the center point of each detected target in the x direction and the y direction; The feature map is subjected to feature extraction by a third two-layer neural network to obtain a predicted feature map, wherein the size of the predicted feature map is , the predicted feature map is used to predict the width and height of each detected target with reference to the center point of each detected target; Detection targets are generated based on the heat map, the offset feature map, and the predicted feature map to obtain multiple detection targets.

4. The obstacle detection method of vision and laser radar fusion according to claim 1, characterized in that: The normalizing operation is performed on the multiple detection targets respectively to obtain multiple normalized detection targets, including: Normalizing the scales of the multiple detection targets respectively to obtain detection targets of multiple target sizes; Pixel normalization is performed on the detection targets of the multiple target sizes to obtain multiple normalized detection targets.

5. The obstacle detection method of vision and laser radar fusion according to claim 1, characterized in that: The performing image classification on each of the plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets respectively includes: Downsampling each normalized detection target among the multiple normalized detection targets through a first 2D convolution to obtain a feature image of each detection target; Extracting features from the feature image of each detection target through multiple bottleneck structures to obtain deep features of each detection target; The deep features of each detection target are image classified by mean pooling and a second 2D convolution to obtain target classifications corresponding to multiple inspection targets.

6. The obstacle detection method of vision and laser radar fusion according to claim 1, characterized in that: The performing distance clustering on the radar point cloud in the cone space and determining the cluster with the largest density as the obstacle point cloud includes: Sorting the radar point cloud in the cone space in descending order according to the distance from the laser radar to obtain a radar point distance sequence; According to a preset distance threshold, clustering is performed based on the distances between radar points in the radar point distance sequence to obtain a plurality of clustering clusters; A cluster containing the most radar points among the multiple clusters is determined as an obstacle point cloud.

7. An obstacle detection device integrating vision and laser radar, characterized in that: include: An acquisition module, used to acquire a visual image, wherein the visual image is a forward image collected by a train in a rail transit scene, and the forward image includes obstacles; A detection module, used to input the visual image into a preset target detection network to obtain a plurality of detection targets output by the target detection network; A normalization module, used to perform normalization operations on the multiple detection targets respectively to obtain multiple normalized detection targets; A classification module, used for performing image classification on each of the plurality of normalized detection targets to obtain target classifications corresponding to the plurality of inspection targets, wherein the target classifications include obstacle targets; A projection module, used to map the obstacle target to a point cloud space based on projection transformation through a laser radar of the train to obtain a visual cone space; A clustering module, used for performing distance clustering on the radar point cloud in the vision cone space, and determining the cluster with the largest density as the obstacle point cloud; A determination module is used to determine the distance between the obstacle and the train based on the obstacle point cloud.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the obstacle detection method integrating vision and laser radar as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the obstacle detection method integrating vision and laser radar as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the obstacle detection method integrating vision and laser radar as described in any one of claims 1 to 6 is implemented.