Railway track foreign matter detection method based on millimeter wave radar sensor and camera sensor
By combining the information of millimeter-wave radar and camera sensors in railway track foreign matter detection, multi-frame autonomous fusion, DBSCAN clustering and lightweight YOLOv5s algorithm detection methods are used to solve the problem of insufficient detection accuracy and real-time performance of the existing technology in complex environments, and achieve more efficient and reliable railway track foreign matter detection.
Patent Information
- Application Number
- CN202510120577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-25
AI Technical Summary
The existing railway track foreign object detection methods are susceptible to changes in ambient light and severe weather, and are insufficient in detection accuracy and real-time performance in complex scenarios.
The fusion detection method based on millimeter wave radar sensor and camera sensor is adopted to improve the density and accuracy of radar point cloud through multi-frame autonomous fusion and DBSCAN clustering processing, and visual area of interest detection is combined with the lightweight YOLOv5s algorithm, and the decision-level fusion method of intersection and parallel ratio (IOU) and the extended Kalman filtering algorithm are used for target tracking.
It improves the accuracy and real-time detection of foreign matter on railway tracks, enhances the detection capabilities in harsh environments, and significantly improves railway traffic safety.
Smart Images

Figure CN120028786A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of railway track foreign body detection, and in particular relates to a railway track foreign body detection method based on a millimeter wave radar sensor and a camera sensor. Background Art
[0002] With the rapid development of China's railway infrastructure in recent years, railway traffic safety cannot be ignored. Although the railway adopts a fully closed management, foreign objects such as cattle, sheep, and falling rocks often invade. Foreign object invasion of railways poses a serious threat to the safety of locomotive operation, not only disrupting the transportation order, but also causing casualties and economic losses.
[0003] With the development of deep learning, the current focus on foreign body detection on rails is on image recognition algorithms. Take the one-stage target detection algorithm represented by YOLO as an example: it not only has fast detection speed, but also good real-time performance. However, the use of visual detection algorithms alone is easily affected by the external environment. When encountering large changes in ambient light intensity (such as dark environments) and bad weather such as rain and fog, the visual detection method is greatly affected. Considering this shortcoming, multi-sensor fusion technology has gradually emerged, which can make up for the problem of missing data when a single sensor does not collect enough external data in some cases.
[0004] Millimeter-wave radar is small in size, light in weight, can achieve long-distance detection, and is insensitive to media such as rain, fog, dust, and light. It is widely used in target detection tasks. The combination of millimeter-wave radar and camera can perfectly utilize the advantages of each other to achieve target detection tasks in complex scenes. Summary of the invention
[0005] In view of the defects existing in the prior art, the object of the present invention is to provide a railway track foreign object detection method based on millimeter wave radar sensor and camera sensor to solve the current problem of railway track foreign object identification.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A railway track foreign body detection method based on a millimeter wave radar sensor and a camera sensor comprises the following steps:
[0008] S1. The millimeter wave radar sensor detects obstacles on the front rails, obtains point clouds, and obtains a large frame data set after preprocessing. The specific process of the preprocessing is as follows:
[0009] In view of the sparse point cloud obtained by millimeter-wave radar sensor detection, the acquired point cloud (each point in the point cloud includes the (x, y, z) coordinates, reflection area and speed of the detected obstacle) is first subjected to multi-frame autonomous fusion processing. The data after multi-frame autonomous fusion processing is then processed based on the DBSCAN clustering algorithm of the three-dimensional bounding box to obtain a large frame data set.
[0010] Preferably, in step S1, the specific operation mode of the multi-frame autonomous fusion processing is:
[0011] After reading the acquired point cloud, the point cloud of the current frame and the previous three frames are superimposed to increase the density. That is, the acquired point cloud is superimposed into one large frame every 4 frames. By continuously superimposing 4 frames of point cloud, a large frame set is created. Through multi-frame autonomous fusion processing, a more refined and dense point cloud is obtained while maintaining the original basic form and spatial structure characteristics of the point cloud.
[0012] Preferably, in step S1, the DBSCAN clustering algorithm processing includes the following steps:
[0013] The vertical space with the railway track as the bottom is used as the test area. The millimeter-wave radar scans the test area to obtain a point cloud, which is then processed by autonomous fusion of multiple frames to form a large frame set. After each large frame in the large frame set is processed by the DBSCAN clustering algorithm in turn, the points in the point cloud obtained by the millimeter-wave radar sensor that are real reflection points of obstacles are screened out to form a large frame data set.
[0014] S2: The camera sensor captures the obstacle on the rail ahead, obtains image data, and uses the lightweight YOLOv5s algorithm to process the data to obtain the visual region of interest (ROI). C , the steps are as follows:
[0015] S2.1. Improve the focus layer and CSP (Cross Stage Partial connections) module in the YOLOv5s algorithm to obtain a lightweight YOLOv5s algorithm.
[0016] Preferably, in step S2.1, the process of obtaining the lightweight YOLOv5s algorithm is as follows:
[0017] S2.1.1, the focus layer of the YOLOv5s algorithm converts a feature map of size H×W×C into In order to improve the detection rate and accuracy of the YOLOv5s algorithm, the focus layer is improved to convert a feature map of size H×W×C into size, thus effectively increasing the channel dimension without losing feature map information;
[0018] S2.1.2. Use the MobileNetV3 module to replace the CSP module in the YOLOv5s algorithm. The specific process is as follows: in the YOLOv5s algorithm, use the MobileNetV3 module to replace the CSP module of the original backbone. (1) Perform the depth wise convolution (Depth wiseConvolution) included in the MobileNetV3 module on each feature map channel of the input CSP module, and then perform the point wise convolution (Pointwise Convolution) included in the MobileNetV3 module on all feature map channels of the input CSP module, so as to achieve the standard convolution of the input feature map by using the depth wise convolution included in MobileNetV3 to replace the CSP module; (2) Use the inverted residual structure included in the MobileNetV3 module to replace the residual block connection mechanism in the CSP module. The inverted residual structure expands the number of feature map channels and then performs an expansion convolution operation, and then compresses it back to the original number of channels; (3) Use the SE module included in the MobileNetV3 module to replace the output layer of the CSP module.
[0019] S2.1.3, the improved YOLOv5s algorithm obtained by S2.2.1 and S2.2.2 is optimized using the Detection On Tracks dataset to obtain a lightweight YOLOv5s algorithm; the Detection On Tracks dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and then the training set, validation set, and test set are used in sequence for training, validation, and testing. The number of iterations of the training is 300, until the evaluation index meets the accuracy requirement. The evaluation criteria are:
[0020]
[0021] Among them, the mean average precision M AP It is one of the indicators of target detection and is the mean of the average accuracy of all categories. P is the average accuracy of each class, m is the number of categories, p and r are the accuracy and recall of the algorithm respectively; if the above accuracy evaluation criteria are not met, the Detection On Tracks data set division, training, verification, and testing are repeated, and the parameters are adjusted and updated until the accuracy meets the requirements.
[0022] S2.2, load the lightweight YOLOv5s algorithm obtained in S2.1 into the visual AI board (jetson nano mainboard, B01 version), and process the image data obtained by the camera sensor shooting the obstacle on the front track with the lightweight YOLOv5s algorithm to obtain the visual region of interest ROIC , ROI C The size is S C , S C It is the pixel area size of the obstacle photographed by the camera in the pixel coordinate system.
[0023] S3, synchronize the large frame data set obtained in S1 with the image data in S2, and then project the large frame data set obtained in S1 into the pixel coordinate system to generate the radar region of interest ROI in the pixel coordinate system R Then, the radar region of interest ROI R and visual region of interest ROI C Form a fusion detection frame in the pixel coordinate system;
[0024] Preferably, the specific process of step S3 is as follows:
[0025] S3.1. Time synchronization is performed on the large frame data set obtained by S1 and the image data in S2. That is, the large frame data set obtained by S1 and the image data in S2 are time-aligned and time synchronization is performed using the principle of the minimum common multiple of the period. The specific process is as follows:
[0026] In S1, the point cloud is superimposed into a large frame every 4 frames, that is, the time interval for forming a large frame is T. According to time synchronization, the camera also needs to take an image at an interval of T. Therefore, the lowest common multiple T time is selected as the sampling period of time synchronization. The time interval of millimeter wave radar sampling is obtained according to the following formula (3):
[0027]
[0028] Where T is the duty cycle of the sensor and f is the sampling frequency.
[0029] S3.2, project the large frame data set obtained in S1 to the pixel coordinate system, and generate a radar region of interest ROI in the pixel coordinate system for each point cloud in the large frame data R ;
[0030] Preferably, the specific process of S3.2 is:
[0031] S3.2.1. Convert from millimeter wave radar coordinate system to camera coordinate system:
[0032]
[0033] In formula (4), X C , Y C , Z C Indicates the three-dimensional coordinates of the obstacle's real reflection point in the camera coordinate system; X R , Y R , Z Ris the three-dimensional coordinate of the real reflection point of the obstacle in the millimeter-wave radar coordinate system, R is the rotation matrix, and t is the displacement matrix; is the camera extrinsic matrix.
[0034] S3.2.2. Convert from camera coordinate system to pixel coordinate system:
[0035]
[0036] Among them, u and v are the coordinates of the real reflection point of the obstacle target in the pixel coordinate system. is the camera intrinsic parameter matrix, where f x is the focal length in the horizontal direction, in pixels; f y is the focal length in the vertical direction, in pixels; u o is the pixel coordinate of the image center in the horizontal direction, v o is the pixel coordinate of the center of the image in the vertical direction.
[0037] S3.2.3. The coordinates of each point in the point cloud of the large frame data set obtained by S1 are transformed in S3.1.2.1 and S3.1.2.2 in sequence, so that the point cloud of the large frame data set obtained by S1 in the millimeter wave radar coordinate system is projected to the pixel coordinate system.
[0038] S3.2.4. In the pixel coordinate system, the pixel coordinates of the points in the point cloud of the large frame data set obtained in S1 are used to approximate the size of the three-dimensional bounding box to generate the radar region of interest ROI. R , ROI R Using formula (6), the ROI corresponding to each large frame data in the large frame data set obtained by S1 is R The size is S R , S R The pixel area size of the point cloud in each large frame data, that is, the real reflection point of the obstacle projected into the pixel coordinate system.
[0039] ROI R =(u,v,w,h,t)∈R (6)
[0040] Among them, u and v represent the coordinates of the real reflection point of the obstacle detected by the millimeter wave radar in the pixel coordinate system, w, h, and t are the width, height, and depth of the three-dimensional bounding box in the pixel coordinate system, respectively, and R represents the entire pixel coordinate system.
[0041] S3.3: In the pixel coordinate system, the radar region of interest ROI corresponding to each large frame data is obtained R The visual region of interest ROI corresponding to the image data collected by the camera every 200ms synchronized with it C , radar region of interest ROI Rand visual region of interest ROI C The total pixel area of the union is the fused detection box.
[0042] S4, using the decision-level fusion method of intersection over union (IOU) to calculate the radar region of interest ROI corresponding to each large frame data R The visual region of interest ROI corresponding to the image data acquired by the camera in time synchronization C The overlap degree is obtained to obtain the preliminary fusion result, including the following steps:
[0043] S4.1: Calculate the time-synchronized ROI for each radar R And each visual region of interest ROI C The overlap S IoU , calculated as follows:
[0044]
[0045] Among them, S R ROI for radar region of interest R The pixel area size in the pixel coordinate system, S C ROI is the visual region of interest C The pixel area size in the pixel coordinate system;
[0046] S4.2: Set the detection threshold for evaluation. The evaluation method is as follows:
[0047] S IoU When <0.3, the obstacles detected by the millimeter-wave radar and the camera are considered to be different obstacles and are discarded, that is, false targets;
[0048] 0.3≤S IoU When <0.5, it is considered that the obstacles detected by the millimeter-wave radar and the camera may be the same obstacle, that is, an uncertain target;
[0049] 0.5≤S IoU When ≤1, the obstacles detected by the millimeter-wave radar and the camera are considered to be the same obstacle, that is, the target is determined.
[0050] Therefore, by judging the overlap degree S IoU , the detected point cloud and image data reflecting obstacles are preliminarily fused and processed to obtain preliminary fusion results: false targets, confirmed targets, and uncertain targets.
[0051] S5, using the extended Kalman filter algorithm (EKF) to dynamically track the overlapping boxes corresponding to the uncertain targets in the preliminary fusion results obtained in S4, and output the fusion results, including the following steps:
[0052] The overlapping box state vector corresponding to the uncertain target in S5.1 and S4 is defined as X = (u, v, v u ,v v ) T , where u, v, v u 、v v They represent the horizontal coordinate, vertical coordinate, horizontal axis speed, and vertical axis speed of the overlapping box corresponding to the uncertain target in the pixel coordinate system, respectively. u 、v v It is the speed of the actual reflection point cloud of the obstacle in the overlapping box.
[0053] Using the extended Kalman filter algorithm, the state equation and prediction equation of the overlapping frame are:
[0054]
[0055] Among them, X(k) and X(k-1) represent the state vectors of the overlapping box corresponding to the uncertain target at time k and k-1 respectively; Z(k) represents the observation vector of the overlapping box corresponding to the uncertain target at time k; f(k) represents the state transfer matrix, h(k) represents the observation function; V(k) and W(k) represent Gaussian white noise.
[0056] S5.2. Based on the state of the uncertain target at time k-1, the state equation for predicting the uncertain target at time k is:
[0057]
[0058] in, P(k|k-1) represents the state prediction vector and prediction error vector of the overlapping box corresponding to the uncertain target at time k, F is the Jacobian matrix of f(k) at X(k-1|k-1), P(k-1|k-1) is the prediction error vector of the overlapping box corresponding to the uncertain target at time k-1, and Q is the covariance matrix of the process noise.
[0059] S5.3. The state update equation of the overlapping frame corresponding to the uncertain target is deduced from the state equation of formula 9:
[0060]
[0061] In formula (10), is the state update vector, S(k) represents the innovation covariance, K(k) represents the gain matrix, P(k|k-1) is the prediction error covariance, and H is h(k) in is the Jacobian matrix at , and R is the covariance matrix of the observation noise.
[0062] S5.4. The final state equation of the uncertain target is derived from the state update equation of formula (10):
[0063]
[0064] In formula (11), X(k|k) represents the state estimation vector of the uncertain target at time k, P(k|k) represents the covariance at time k, K(k) represents the gain matrix, is the state update vector, P(k|k-1) is the prediction error covariance, and I represents the identity matrix.
[0065] Formula (8) represents the state equation and prediction equation of the overlapping box corresponding to the uncertain target in the entire EKF tracking process. Formula (9) predicts the state equation at time k based on time k-1. The state update equation (10) of the uncertain target is derived from formula (9), and the final state equation (11) of the uncertain target is derived from formula (10).
[0066] S5.5. After the above steps are iterated and updated for 10 times, if the overlapping frame corresponding to the tracked uncertain target appears 3 times in succession, the overlapping frame is output as the fusion result, indicating that the overlapping frame is the same obstacle identified, that is, the uncertain target is actually a determined target.
[0067] S6, the determined target obtained by S4.2, and the determined target obtained by S4.2 after being processed by S5 together represent the detected obstacle on the railway track.
[0068] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0069] The railway track foreign body detection method based on millimeter wave radar sensor and camera sensor provided by the present invention is to install millimeter wave radar and visual AI board on a 3-meter-high electric pole beside the railway track (the electric pole is 1.2m away from the rail) (the camera and the visual AI board are assembled together at a height of 2m, and data is transmitted between the millimeter wave radar and the visual AI board via the CAN bus). At this time, the camera captures the obstacle on the front rail to form a visual region of interest ROI C The point cloud data obtained by using the millimeter-wave radar to detect obstacles on the front rail is projected onto the visual system to form the millimeter-wave radar region of interest ROI R After that, the two sensors form a fusion detection frame for the detected rail obstacle target, and then calculate the region of interest ROI according to the decision-level fusion method of intersection-over-union (IOU). R ROI CThe overlap degree is calculated to obtain the preliminary fusion result. After obtaining the preliminary fusion result, the uncertain target is dynamically tracked based on the extended Kalman filter algorithm (EKF) and the final fusion result is output. The present invention adopts the information fusion technology combining millimeter wave radar and camera to overcome the shortcomings of traditional track target detection methods, and has achieved significant improvements in detection accuracy and real-time performance, thereby improving the safety of railway track transportation. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 The present invention is a flow chart of a railway track foreign body detection method based on a millimeter wave radar sensor and a camera sensor.
[0071] Figure 2 This is the schematic diagram of the extended Kalman filter algorithm;
[0072] Figure 3 The railway track selected for the experiment of the railway track foreign body detection method of the present invention is the railway track experimental base in the north area of East China Jiaotong University, and a one-to-one restored real railway track scene is used to collect data and conduct tests;
[0073] Figure 4 The density changes of point clouds before and after multi-frame superposition and DBSCAN clustering processing;
[0074] Figure 5 The left picture shows the visual region of interest ROI obtained by the YOLOv5 algorithm C and its confidence. The right picture shows the visual region of interest ROI obtained by the lightweight YOLOv5 algorithm in this invention. C and its confidence level;
[0075] Figure 6 ROI for radar region of interest R and visual region of interest ROI C The fused detection box. DETAILED DESCRIPTION
[0076] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0077] like Figure 1 As shown, a railway track foreign body detection method based on a millimeter wave radar sensor and a camera sensor comprises the following steps:
[0078] S1. The millimeter-wave radar sensor detects obstacles on the front rails, obtains point clouds, and obtains point cloud data sets after preprocessing. The specific process is as follows:
[0079] In view of the sparse point cloud obtained by millimeter-wave radar sensor detection, the obtained point cloud (each point in the point cloud includes the (x, y, z) coordinates, reflection area and speed of the detected obstacle) is first subjected to multi-frame autonomous fusion processing, and the data after multi-frame autonomous fusion processing is then processed based on the DBSCAN clustering algorithm of the three-dimensional bounding box.
[0080] The specific operation mode of the multi-frame autonomous fusion processing is:
[0081] After reading the acquired point cloud, the point cloud of the current frame is first superimposed with the point cloud of the previous three frames to increase the density, that is, the acquired point cloud is superimposed with 4 frames as 1 large frame, and a large frame set is created by continuously superimposing 4 frames of point cloud. Through multi-frame autonomous fusion processing, while maintaining the original basic form and spatial structure characteristics of the point cloud, a more refined and dense point cloud is obtained;
[0082] The DBSCAN clustering algorithm process includes the following steps:
[0083] A vertical space with a straight railway track of 50 meters long and 1.43 meters wide as the bottom is used as the test area ( Figure 3 ), the millimeter wave radar scans the area to be tested to obtain a point cloud, and after multi-frame autonomous fusion processing, a large frame set is formed. For each large frame in the large frame set, the preset neighborhood radius is ε, and the minimum number of points in the ε neighborhood is set to P min ; For a certain large frame, randomly select a point in the point cloud included in the large frame, such as point P, which is located in the point cloud included in the large frame. If point P is determined to be a core point, then use the density accessibility principle to find all points in the point cloud in the large frame that are directly density-reachable and indirectly density-reachable to point P, and classify these points into the same cluster; if point P is a boundary point, classify point P into the cluster of the core points in the neighborhood to which point P belongs; if point P is neither a core point nor a boundary point, it is classified as noise and is not included in the cluster. Therefore, after filtering out abnormal points in the point cloud through the DBSCAN clustering algorithm, the average value of all points retained in the point cloud is used to estimate the center position, and the spatial range of all points retained in the point cloud is determined by searching the outermost points.
[0084] After each of the large frames in the large frame set is processed by the DBSCAN clustering algorithm in turn, points belonging to real reflection points of obstacles in the point cloud detected by the millimeter-wave radar sensor are screened out to form a large frame data set.
[0085] S2: The camera sensor captures the obstacle on the rail ahead, obtains image data, and uses the lightweight YOLOv5s algorithm to process the data to obtain the visual region of interest (ROI). C , the steps are as follows:
[0086] S2.1. Considering the engineering practicality, the focus layer and CSP (Cross Stage Partial connections) module in the YOLOv5s algorithm are improved to lightweight the YOLOv5s algorithm and obtain a lightweight YOLOv5s algorithm to improve the detection rate and accuracy. The process of obtaining the lightweight YOLOv5s algorithm is as follows:
[0087] S2.1.1, the focus layer of the YOLOv5s algorithm converts a feature map of size H×W×C into In order to improve the detection rate and accuracy of the YOLOv5s algorithm, the focus layer is improved to convert a feature map of size H×W×C into size, thus effectively increasing the channel dimension without losing feature map information;
[0088] S2.1.2. Use the MobileNetV3 module to replace the CSP module in the YOLOv5s algorithm. The specific process is as follows: In the YOLOv5s algorithm, use the MobileNetV3 module to replace the original backbone CSP module. (1) Perform the depth wise convolution (Depth wiseConvolution) included in the MobileNetV3 module on each feature map channel of the input CSP module, and then perform the point wise convolution (Pointwise convolution) included in the MobileNetV3 module on all feature map channels of the input CSP module. Convolution) to achieve the goal of using the depth-separable convolution included in MobileNetV3 to replace the CSP module to perform standard convolution on the input feature map; (2) using the inverted residual structure included in the MobileNetV3 module to replace the residual block connection mechanism in the CSP module. The inverted residual structure expands the number of feature map channels and then performs a dilated convolution operation, and then compresses it back to the original number of channels; changing the residual block connection mechanism in the CSP module to perform multiple convolutions on the feature map channels and jump connect them together; (3) using the SE module included in the MobileNetV3 module to replace the output layer of the CSP module. The SE module is an attention mechanism module that can adaptively adjust the channel output weight of the feature map by adjusting the weight of each channel of the feature map. Through the above processing, the number of model parameters is significantly reduced and the computational complexity of the feature map channels in the input and output process is reduced.
[0089] S2.1.3, the YOLOv5s algorithm improved by S2.2.1 and S2.2.2 is optimized using the Detection On Tracks dataset (a public dataset containing 3766 images of human behavior on railway tracks, with a resolution of 1080×1080, annotated in YOLO format (txt), and with detailed class labels) to obtain a lightweight YOLOv5s algorithm. The training conditions are: system environment: windows system, hardware environment: CPU model is Intel(R) Core(TM) i5-12400f, GPU model is NVIDIA GTX 4060 (8GB). The Detection On Tracks dataset is divided into training set, validation set, and test set according to the ratio of 8:1:1 (that is, the training set includes 3012 images, the test set and the validation set each include 374 images), and then the training set, validation set, and test set are used for training, validation, and testing in turn. The number of iterations of the training is 300 times until the evaluation index meets the accuracy requirements. The evaluation criteria are:
[0090]
[0091] Among them, the mean average precision M AP It is one of the indicators of target detection and is the mean of the average accuracy of all categories. P is the average accuracy of each class, m is the number of categories, p and r are the accuracy and recall of the algorithm respectively; if the above accuracy evaluation criteria are not met, the Detection On Tracks data set division, training, verification, and testing are repeated, and the parameters are adjusted and updated until the accuracy meets the requirements.
[0092] S2.2, load the lightweight YOLOv5s algorithm obtained in S2.1 into the visual AI board (jetson nano mainboard, B01 version), and process the image data obtained by the camera sensor shooting the obstacle on the front track with the lightweight YOLOv5s algorithm to obtain the visual region of interest ROI C , ROI C The size is S C , S C It is the pixel area size of the obstacle photographed by the camera in the pixel coordinate system.
[0093] S3, synchronize the large frame data set obtained in S1 with the image data in S2, and then project the large frame data set obtained in S1 into the pixel coordinate system to generate the radar region of interest ROI in the pixel coordinate system R Then, the radar region of interest ROI R and visual region of interest ROI C exist Pixel Coordinate SystemThe fusion detection frame is formed, including the following steps:
[0094] S3.1. Time synchronization is performed on the large frame data set obtained by S1 and the image data in S2. That is, the large frame data set obtained by S1 and the image data in S2 are time-aligned and time synchronization is performed using the principle of the minimum common multiple of the period. The specific process is as follows:
[0095] Due to the different sampling frequencies of different sensors, there is a time difference in the data collection. The sampling frequency of the millimeter wave radar is 20HZ, that is, 20 samples per second, and the sampling frequency of the camera is 30HZ, that is, 30 samples per second. According to formula (3), the sampling time interval of the millimeter wave radar is 50ms, and the sampling time interval of the camera is 33.3ms. In S1, the point cloud is superimposed into a large frame every 4 frames, that is, the time interval for forming a large frame is 200ms. According to time synchronization, the camera also needs to take an image at an interval of 200ms. Therefore, the lowest common multiple of 200ms is selected as the sampling period of time synchronization (specifically, to make the fast camera compatible with the slow radar). The essence of time synchronization is to achieve the consistency of the data obtained by the two sensors in time to prevent large deviations in the fused information.
[0096]
[0097] Where T is the duty cycle of the sensor and f is the sampling frequency.
[0098] S3.2, project the large frame data set obtained in S1 to the pixel coordinate system, and generate a radar region of interest ROI in the pixel coordinate system for each point cloud in the large frame data R , the specific process is:
[0099] S3.2.1. Convert from millimeter wave radar coordinate system to camera coordinate system:
[0100]
[0101] In formula (4), X C , Y C , Z C Indicates the three-dimensional coordinates of the obstacle's real reflection point in the camera coordinate system; X R , Y R , Z R is the three-dimensional coordinate of the real reflection point of the obstacle in the millimeter-wave radar coordinate system, R is the rotation matrix, and t is the displacement matrix; is the camera extrinsic matrix.
[0102] S3.2.2. Convert from camera coordinate system to pixel coordinate system:
[0103]
[0104] Among them, u and v are the coordinates of the real reflection point of the obstacle target in the pixel coordinate system. is the camera intrinsic parameter matrix, where f x is the focal length in the horizontal direction, in pixels; f y is the focal length in the vertical direction, in pixels; u o is the pixel coordinate of the image center in the horizontal direction, v o is the pixel coordinate of the center of the image in the vertical direction.
[0105] S3.2.3. The coordinates of each point in the point cloud of the large frame data set obtained by S1 are transformed in sequence according to S3.1.2.1) and S3.1.2.2) so that the point cloud of the large frame data set obtained by S1 in the millimeter wave radar coordinate system is projected to the pixel coordinate system.
[0106] S3.2.4. In the pixel coordinate system, the pixel coordinates of the points in the point cloud of the large frame data set obtained in S1 are used to approximate the size of the three-dimensional bounding box to generate the radar region of interest ROI. R , ROI R Using formula (6), the ROI corresponding to each large frame data in the large frame data set obtained by S1 is R The size is S R , S R The pixel area size of the point cloud in each large frame data, that is, the real reflection point of the obstacle projected into the pixel coordinate system.
[0107] ROI R =(u,v,w,h,t)∈R (6)
[0108] Among them, u and v represent the coordinates of the real reflection point of the obstacle detected by the millimeter wave radar in the pixel coordinate system, w, h, and t are the width, height, and depth of the three-dimensional bounding box in the pixel coordinate system, respectively, and R represents the entire pixel coordinate system.
[0109] S3.3: In the pixel coordinate system, the radar region of interest ROI corresponding to each large frame data is obtained R The visual region of interest ROI corresponding to the image data collected by the camera every 200ms synchronized with it C , radar region of interest ROI R and visual region of interest ROI C The total pixel area of the union is the fused detection box.
[0110] S4, using the decision-level fusion method of intersection over union (IOU) to calculate the radar region of interest ROI corresponding to each large frame data RThe visual region of interest ROI corresponding to the image data collected by the camera every 200ms synchronized with it C The overlap degree is obtained to obtain the preliminary fusion result, including the following steps:
[0111] S4.1: Calculate the time-synchronized ROI for each radar R And each visual region of interest ROI C The overlap S IoU , calculated as follows:
[0112]
[0113] Among them, S R ROI for radar region of interest R The pixel area size in the pixel coordinate system, S C ROI is the visual region of interest C The pixel area size in the pixel coordinate system;
[0114] S4.2: Set the detection threshold for evaluation. The evaluation method is as follows:
[0115] S IoU When <0.3, the obstacles detected by the millimeter-wave radar and the camera are considered to be different obstacles and are discarded, that is, false targets;
[0116] 0.3≤S IoU When <0.5, it is considered that the obstacles detected by the millimeter-wave radar and the camera may be the same obstacle, that is, an uncertain target;
[0117] 0.5≤S IoU When ≤1, the obstacles detected by the millimeter-wave radar and the camera are considered to be the same obstacle, that is, the target is determined.
[0118] Therefore, by judging the overlap degree S IoU , the detected point cloud and image data reflecting obstacles are preliminarily fused and processed to obtain preliminary fusion results: false targets, confirmed targets, and uncertain targets.
[0119] S5, such as Figure 2 As shown, the extended Kalman filter algorithm (EKF) is used to dynamically track the overlapping boxes corresponding to the uncertain targets in the preliminary fusion results, and the fusion results are output, including the following steps:
[0120] The overlapping box state vector corresponding to the uncertain target in S5.1 and S4 is defined as X = (u, v, v u ,v v ) T , where u, v, v u 、v vThey represent the horizontal coordinate, vertical coordinate, horizontal axis speed, and vertical axis speed of the overlapping box corresponding to the uncertain target in the pixel coordinate system, respectively. u 、v v It is the speed of the actual reflection point cloud of the obstacle in the overlapping box.
[0121] Using the extended Kalman filter algorithm, the state equation and prediction equation of the overlapping frame are:
[0122]
[0123] Among them, X(k) and X(k-1) represent the state vectors of the overlapping box corresponding to the uncertain target at time k and k-1 respectively; Z(k) represents the observation vector of the overlapping box corresponding to the uncertain target at time k; f(k) represents the state transfer matrix, h(k) represents the observation function; V(k) and W(k) represent Gaussian white noise.
[0124] S5.2. Based on the state of the uncertain target at time k-1, the state equation for predicting the uncertain target at time k is:
[0125]
[0126] in, represents the state prediction vector and prediction error vector of the overlapping box corresponding to the uncertain target at time k, F is the Jacobian matrix of f(k) at X(k-1|k-1), P(k-1|k-1) is the prediction error vector of the overlapping box corresponding to the uncertain target at time k-1, and Q is the covariance matrix of the process noise.
[0127] S5.3. The state update equation of the overlapping frame corresponding to the uncertain target is deduced from the state equation of formula 9:
[0128]
[0129] In formula (10), is the state update vector, S(k) represents the innovation covariance, K(k) represents the gain matrix, P(k|k-1) is the prediction error covariance, and H is h(k) in is the Jacobian matrix at , and R is the covariance matrix of the observation noise.
[0130] S5.4. The final state equation of the uncertain target is derived from the state update equation of formula (10):
[0131]
[0132] In formula (11), X(k|k) represents the state estimation vector of the uncertain target at time k, P(k|k) represents the covariance at time k, K(k) represents the gain matrix, is the state update vector, P(k|k-1) is the prediction error covariance, and I represents the identity matrix.
[0133] Formula (8) represents the state equation and prediction equation of the overlapping box corresponding to the uncertain target in the entire EKF tracking process. Formula (9) predicts the state equation at time k based on time k-1. The state update equation (10) of the uncertain target is derived from formula (9), and the final state equation (11) of the uncertain target is derived from formula (10).
[0134] S5.5. After the above steps are iterated and updated for 10 times, if the overlapping frame corresponding to the tracked uncertain target appears 3 times in succession, the overlapping frame is output as the fusion result, indicating that the overlapping frame is the same obstacle identified, that is, the uncertain target is actually a determined target.
[0135] S6, the determined target obtained by S4.2, and the determined target obtained by S4.2 after being processed by S5 together represent the detected obstacle on the railway track.
[0136] In the embodiment, the millimeter-wave radar used is Continental's ARS408-21, which has a working sampling frequency of 20HZ and can realize medium and long-distance detection functions; the camera used is a traffic camera module produced by Yutong Optics, and has a sampling frequency of approximately 30HZ.
[0137] In the actual obstacle detection scenario on railway tracks, a detection system including millimeter-wave radar sensors and camera sensors is first installed on the side of the track: when an obstacle appears on the track, the radar scans the obstacle to obtain a point cloud. After multi-frame superposition and DBSCAN clustering, the number of points in the obstacle's real reflection point cloud increases from the conventional 11 to 44 ( Figure 4 ).
[0138] After the YOLOv5s algorithm was improved and optimized on the training platform, A P The index is ≥88%, M AP The confidence level of the actual detection is 0.82, which is 0.08 higher than the confidence level of the unmodified YOLOv5s (0.74). Figure 5 ). The camera sensor captures the obstacle to obtain image data, and the visual region of interest ROI is obtained after being processed by the lightweight YOLOv5s algorithm. C , ROI C The size is 110px × 280px.
[0139] The information of the two sensors is synchronized in time. After multi-frame superposition and DBSCAN clustering processing, the point cloud is projected to the pixel coordinate system through coordinate transformation, that is, projected onto the picture taken by the camera. The real reflection point cloud of the obstacle forms a region of interest ROI on the picture. R , ROI R The size is 103px × 275px. Figure 6 ROI for radar region of interest R and visual region of interest ROI C The fusion detection frame composed of R ∩S C 24980px,S R ∪S C is 32025px, according to Calculate S IoU It is equal to 0.78, which means the target is determined, indicating that an obstacle on the rail is detected.
[0140] The core innovation of the present invention lies in the combination of its unique sensor fusion strategy and advanced data processing algorithm. First, at the data acquisition and processing level, the present invention adopts multi-frame autonomous fusion technology to deeply integrate the point cloud continuously collected by the millimeter-wave radar, effectively solving the problem of information loss caused by noise or occlusion in a single frame of data. At the same time, the introduction of the DBSCAN clustering algorithm based on the three-dimensional bounding box can accurately extract potential target areas from complex backgrounds, laying a solid foundation for subsequent target recognition and tracking. In terms of sensor fusion, the present invention abandons the traditional simple data superposition method and innovatively proposes a decision-level fusion strategy based on intersection-over-union (IOU). This strategy not only considers the spatial position relationship between the radar and camera detection results, but also evaluates the consistency between the two by calculating the overlap of the region of interest, thereby achieving more accurate target screening and confirmation. This fusion method not only improves the accuracy of detection, but also reduces the occurrence of false alarms and missed reports. In addition, in the target tracking link, the present invention introduces the extended Kalman filter algorithm. The algorithm combines the target's motion model with the observation data to smooth and dynamically predict the fused target trajectory, effectively suppresses noise interference, and improves the stability and continuity of target tracking. The application of this technology enables the present invention to maintain efficient and accurate detection performance in a complex and changeable track environment. In summary, the technical innovation of the present invention covers multiple key links such as data collection, processing, fusion and tracking, bringing a new solution to the field of track foreign body detection.
[0141] The above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any modification or transformation that can be easily thought of by anyone familiar with the technical field within the technical framework disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the claims of the present invention should be used as the standard for determining the protection scope of this application.
Claims
1. A railway track foreign body detection method based on a millimeter wave radar sensor and a camera sensor, comprising the following steps: S1. The millimeter wave radar sensor detects obstacles on the front rails, obtains point clouds, and obtains a large frame data set after preprocessing. The specific process of the preprocessing is as follows: In view of the sparse point cloud obtained by the millimeter-wave radar sensor detection, the acquired point cloud is first subjected to multi-frame autonomous fusion processing. The data after multi-frame autonomous fusion processing is then processed based on the DBSCAN clustering algorithm of the three-dimensional bounding box to obtain a large frame data set; S2: The camera sensor captures the obstacle on the rail ahead, obtains image data, and uses the lightweight YOLOv5s algorithm to process the data to obtain the visual region of interest (ROI). C , the steps are as follows: S2.
1. Improve the focus layer and CSP module in the YOLOv5s algorithm to obtain a lightweight YOLOv5s algorithm. S2.2, load the lightweight YOLOv5s algorithm obtained in S2.1 into the visual AI board (jetson nano mainboard, B01 version), and process the image data obtained by the camera sensor shooting the obstacle on the front track with the lightweight YOLOv5s algorithm to obtain the visual region of interest ROI C , ROI C The size is S C , S C The pixel area of the obstacle photographed by the camera in the pixel coordinate system; S3, synchronize the large frame data set obtained in S1 with the image data in S2, and then project the large frame data set obtained in S1 into the pixel coordinate system to generate the radar region of interest ROI in the pixel coordinate system R Then, the radar region of interest ROI R and visual region of interest ROI C A fusion detection frame is formed in the pixel coordinate system; the specific process of step S3 is as follows: S3.
1. Time synchronization is performed on the large frame data set obtained by S1 and the image data in S2. That is, the large frame data set obtained by S1 and the image data in S2 are time-aligned and time synchronization is performed using the principle of the minimum common multiple of the period. The specific process is as follows: In S1, the point cloud is superimposed into a large frame every 4 frames, that is, the time interval for forming a large frame is T. According to time synchronization, the camera also needs to take an image at an interval of T time. Therefore, the lowest common multiple T time is selected as the sampling period of time synchronization; the time interval of millimeter wave radar sampling is obtained according to the following formula (3): Where T is the duty cycle of the sensor, and f is the sampling frequency; S3.2, project the large frame data set obtained in S1 to the pixel coordinate system, and generate a radar region of interest ROI in the pixel coordinate system for each point cloud in the large frame data R ; S3.3: In the pixel coordinate system, the radar region of interest ROI corresponding to each large frame data is obtained R The visual region of interest ROI corresponding to the image data collected by the camera in time synchronization C , radar region of interest ROI R and visual region of interest ROI C The total pixel area of the union is the fusion detection frame; S4, using the decision-level fusion method of intersection over union (IOU) to calculate the radar region of interest ROI corresponding to each large frame data R The visual region of interest ROI corresponding to the image data collected by the camera in time synchronization C The overlap degree is obtained to obtain the preliminary fusion result, including the following steps: S4.1: Calculate the time-synchronized ROI for each radar R And each visual region of interest ROI C The overlap S IoU , calculated as follows: Among them, S R ROI for radar region of interest R The pixel area size in the pixel coordinate system, S C ROI is the visual region of interest C The pixel area size in the pixel coordinate system; S4.2: Set the detection threshold for evaluation. The evaluation method is as follows: S IoU When <0.3, the obstacles detected by the millimeter-wave radar and the camera are considered to be different obstacles and are discarded, that is, false targets; 0.3≤S IoU When <0.5, it is considered that the obstacles detected by the millimeter-wave radar and the camera may be the same obstacle, that is, an uncertain target; 0.5≤S IoU When ≤1, the obstacle detected by the millimeter-wave radar and the camera is considered to be the same obstacle, that is, the target is determined; Therefore, by judging the overlap degree S IoU , the detected point cloud and image data reflecting obstacles are preliminarily fused and processed to obtain the preliminary fusion results: false target, confirmed target, and uncertain target; S5, using an extended Kalman filter algorithm (EKF) to dynamically track the overlapping frames corresponding to the uncertain targets in the preliminary fusion results obtained in S4, and outputting the fusion results; including the following steps: The overlapping box state vector corresponding to the uncertain target in S5.1 and S4 is defined as X = (u, v, v u ,v v ) T , where u, v, v u 、v v They represent the horizontal coordinate, vertical coordinate, horizontal axis speed, and vertical axis speed of the overlapping box corresponding to the uncertain target in the pixel coordinate system, respectively. u 、v v It is the speed of the real reflection point cloud of the obstacle in the overlapping frame; Using the extended Kalman filter algorithm, the state equation and prediction equation of the overlapping frame are: Among them, X(k) and X(k-1) represent the state vectors of the overlapping box corresponding to the uncertain target at time k and k-1 respectively; Z(k) represents the observation vector of the overlapping box corresponding to the uncertain target at time k; f(k) represents the state transfer matrix, h(k) represents the observation function; V(k) and W(k) represent Gaussian white noise; S5.
2. Based on the state of the uncertain target at time k-1, the state equation for predicting the uncertain target at time k is: in, represents the state prediction vector and prediction error vector of the overlapping box corresponding to the uncertain target at time k, F is the Jacobian matrix of f(k) at X(k-1|k-1), P(k-1|k-1) is the prediction error vector of the overlapping box corresponding to the uncertain target at time k-1, and Q is the covariance matrix of the process noise; S5.
3. The state update equation of the overlapping frame corresponding to the uncertain target is deduced from the state equation of formula 9: In formula (10), is the state update vector, S(k) represents the innovation covariance, K(k) represents the gain matrix, P(k|k-1) is the prediction error covariance, and H is h(k) in The Jacobian matrix at , R is the covariance matrix of the observation noise; S5.
4. The final state equation of the uncertain target is derived from the state update equation of formula (10): In formula (11), X(k|k) represents the state estimation vector of the uncertain target at time k, P(k|k) represents the covariance at time k, K(k) represents the gain matrix, is the state update vector, P(k|k-1) is the prediction error covariance, and I represents the identity matrix; S5.
5. After the above steps are iterated and updated for 10 times, if the overlapping frame corresponding to the tracked uncertain target appears for no less than 3 times in a row, the overlapping frame is output as the fusion result, indicating that the overlapping frame is the same obstacle identified, that is, the uncertain target is actually a determined target; S6, the determined target obtained by S4.2, and the determined target obtained by S4.2 after being processed by S5 together represent the detected obstacle on the railway track.
2. The railway track foreign body detection method according to claim 1, characterized in that: In step S1, the specific operation mode of the multi-frame autonomous fusion processing is: After reading the acquired point cloud, the point cloud of the current frame is first superimposed with the point cloud of the previous three frames to increase the density, that is, the acquired point cloud is superimposed into one large frame every 4 frames. By continuously superimposing 4 frames of point cloud, a large frame set is created.
3. The railway track foreign body detection method according to claim 1, characterized in that: In step S1, the DBSCAN clustering algorithm processing includes the following steps: the vertical space with the railway track as the bottom surface is used as the area to be measured, the millimeter-wave radar scans the area to be measured to obtain a point cloud, and forms a large frame set after multi-frame autonomous fusion processing; after the large frames in the large frame set are processed by the DBSCAN clustering algorithm in turn, the points in the point cloud obtained by the millimeter-wave radar sensor detection that are real reflection points of obstacles are screened out to form a large frame data set.
4. The railway track foreign body detection method according to claim 1, characterized in that: In step S2.1, the process of obtaining the lightweight YOLOv5s algorithm is as follows: S2.1.1, the focus layer of the YOLOv5s algorithm converts a feature map of size H×W×C into size, where H and W represent the height and width of the feature map, respectively, and C represents the number of feature map channels; in order to improve the detection rate and accuracy of the YOLOv5s algorithm, the focus layer is improved to convert a feature map of size H×W×C into size, thus effectively increasing the channel dimension without losing feature map information; S2.1.
2. Use the MobileNetV3 module to replace the CSP module in the YOLOv5s algorithm. The specific process is as follows: In the YOLOv5s algorithm, use the MobileNetV3 module to replace the CSP module in the original backbone. (1) Perform the depthwise separable convolution included in the MobileNetV3 module on each feature map channel of the input CSP module, and then perform the pointwise convolution included in the MobileNetV3 module on all feature map channels of the input CSP module, so as to achieve the standard convolution of the input feature map by using the depthwise separable convolution included in MobileNetV3 to replace the CSP module; (2) Use the inverted residual structure included in the MobileNetV3 module to replace the residual block connection mechanism in the CSP module. The inverted residual structure expands the number of feature map channels and then performs an expansion convolution operation, and then compresses it back to the original number of channels; (3) Use the SE module included in the MobileNetV3 module to replace the output layer of the CSP module; S2.1.3, the improved YOLOv5s algorithm obtained by S2.2.1 and S2.2.2 is optimized using the Detection On Tracks dataset to obtain a lightweight YOLOv5s algorithm; the Detection On Tracks dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and then the training set, validation set, and test set are used in sequence for training, validation, and testing. The number of iterations of the training is 300, until the evaluation index meets the accuracy requirement. The evaluation criteria are: Among them, the mean average precision M AP It is one of the indicators of target detection and is the mean of the average accuracy of all categories. P is the average accuracy of each class, m is the number of categories, p and r are the accuracy and recall of the algorithm respectively; if the above accuracy evaluation criteria are not met, the Detection On Tracks data set division, training, verification, and testing are repeated, and the parameters are adjusted and updated until the accuracy meets the requirements.
5. The railway track foreign body detection method according to claim 1, characterized in that: The specific process of S3.2 is as follows: S3.2.
1. Convert from millimeter wave radar coordinate system to camera coordinate system: In formula (4), X C , Y C , Z C Indicates the three-dimensional coordinates of the obstacle's real reflection point in the camera coordinate system; X R , Y R , Z R is the three-dimensional coordinate of the real reflection point of the obstacle in the millimeter-wave radar coordinate system, R is the rotation matrix, and t is the displacement matrix; is the camera extrinsic matrix; S3.2.
2. Convert from camera coordinate system to pixel coordinate system: Among them, u and v are the coordinates of the real reflection point of the obstacle target in the pixel coordinate system. is the camera intrinsic parameter matrix, where f x is the focal length in the horizontal direction, in pixels; f y is the focal length in the vertical direction, in pixels; u o is the pixel coordinate of the image center in the horizontal direction, v o is the pixel coordinate of the center of the image in the vertical direction; S3.2.3, transform the coordinates of each point in the point cloud of the large frame data set obtained in S1 in accordance with S3.1.2.1 and S3.1.2.2, so that the point cloud of the large frame data set obtained in S1 in the millimeter wave radar coordinate system is projected into the pixel coordinate system; S3.2.
4. In the pixel coordinate system, the pixel coordinates of the points in the point cloud of the large frame data set obtained in S1 are used to approximate the size of the three-dimensional bounding box to generate the radar region of interest ROI. R , ROI R Using formula (6), the ROI corresponding to each large frame data in the large frame data set obtained by S1 is R The size is S R , S R The pixel area size of the point cloud in each large frame data, i.e. the real reflection point of the obstacle, projected into the pixel coordinate system; ROI R =(u,v,w,h,t)∈R (6) Among them, u and v represent the coordinates of the real reflection point of the obstacle detected by the millimeter wave radar in the pixel coordinate system, w, h, and t are the width, height, and depth of the three-dimensional bounding box in the pixel coordinate system, respectively, and R represents the entire pixel coordinate system.
Citation Information
Patent Citations
Fusion sensing method based on millimeter wave radar and monocular camera
CN116246143A
Target detection method based on fusion of millimeter wave radar and monocular camera
CN118015596A
Vehicle multi-target intelligent dynamic fusion tracking method
CN118608563A
Multi-information fusion and heat source assisted medical search and rescue unmanned aerial vehicle positioning method
CN118688874A
Dynamic target detection and tracking method based on camera and laser radar data fusion
CN118711030A
Cited By
Railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion
CN121186807A
Railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion
CN121186807B