Method for detecting invading object on rail based on heterogeneous fusion of millimeter-wave radar sensor and camera sensor
By heterogeneously fusion of the millimeter-wave radar sensor with the data of the camera sensor and processing it using specific algorithms and models, the problem of foreign matter detection on railway tracks is solved, real-time and accurate detection is achieved, and railway safety is improved.
Patent Information
- Application Number
- CN202510124275.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-26
AI Technical Summary
In the prior art, railway track foreign matter detection methods are susceptible to environmental influences and have high misjudgment rates, making it difficult to achieve real-time and accurate detection.
The detection method of heterogeneous fusion of millimeter-wave radar sensors and camera sensors is used to synchronize the timestamps and spatial coordinates to perform data fusion, and data processing and detection is performed using PCA, K-means algorithm and YOLOv8 model.
Real-time and accurate detection of foreign objects on the railway tracks is achieved, the accuracy and robustness of the detection is improved, the false alarm rate is reduced, and the safety of the railway is ensured.
Smart Images

Figure CN119959931A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of railway track safety, and in particular relates to a method for detecting intruders on railway tracks by heterogeneously integrating a millimeter-wave radar sensor and a camera sensor. Background Art
[0002] As a symbol of modernity, railways have played a core role in China's national unification and economic and social development since the late Qing Dynasty. As the main artery of the national economy, railways undertake a large number of tasks in transporting people and goods. Once a safety accident occurs, it will not only cause huge economic losses, but also may cause serious casualties and social impact. Therefore, ensuring railway safety is the primary task of railway work.
[0003] CN118072233A discloses a single-machine track foreign body detection method that integrates computer vision and radar. The scheme mainly emphasizes vision as the main detection method. Only when vision detects foreign body intrusion will it be integrated with radar for detection. The model does not take into account the fact that visual detection is easily affected by the environment and has a certain misjudgment rate.
[0004] In order to more accurately detect the intrusion of foreign objects on the railway and improve the safety factor of the railway, it is necessary to introduce more accurate detection methods, specifically using the fusion of millimeter-wave radar and vision, and detecting both at the same time. No matter which sensor detects the intrusion of foreign objects, fusion detection is required to improve accuracy and reduce the false recognition rate, thereby ensuring railway safety. Summary of the invention
[0005] In view of the defects existing in the prior art, the purpose of the present invention is to provide a method for detecting intruders on rails by heterogeneously integrating millimeter-wave radar sensors and camera sensors, utilizing the respective advantages of millimeter-wave radar sensors and camera sensors for fusion detection, detecting the intrusion of foreign objects on rails, and achieving real-time detection of foreign object intrusion on railways, thereby ensuring real-time safety of railways.
[0006] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is a method for detecting intruders on rails by heterogeneously integrating a millimeter-wave radar sensor and a camera sensor, comprising the following steps:
[0007] S1, sensor synchronization: synchronize the timestamps of the millimeter wave radar sensor and the camera sensor and align the spatial coordinates to ensure that the data obtained by the two sensors correspond to the same target;
[0008] First, the NTP network time protocol is used to synchronize timestamps; then spatial coordinate alignment is performed. The spatial coordinate alignment is to convert the point cloud obtained by the millimeter wave radar sensor from the radar coordinate system to the camera coordinate system, thereby achieving the unification of point cloud and image data in the camera coordinate system for effective data fusion.
[0009] Preferably, in step S1, the specific process of the timestamp synchronization is:
[0010] The clock deviation of the millimeter wave radar sensor is:
[0011]
[0012] The clock bias of the camera sensor is:
[0013]
[0014] The synchronized millimeter wave radar sensor timestamp is:
[0015] T r sync =T r recv +θ r (3)
[0016] The synchronized camera sensor timestamp is:
[0017] T c sync =T c recv +θ c (4)
[0018] Among them, T r req The time when the millimeter wave radar sensor sends the NTP request, T r recv is the time it takes for the millimeter-wave radar sensor to receive a response, T c req The time when the camera sensor sends the NTP request, T c recv is the time when the camera sensor receives the response, T s recv is the time when the time server receives the request, T s send The time at which the server sends a response.
[0019] The specific process of the spatial coordinate alignment is:
[0020]
[0021] in, is the coordinate of the reflection point of the intruder in the camera coordinate system, is the coordinate of the reflection point of the intruder in the radar coordinate system, R is the rotation matrix from the radar coordinate system to the camera coordinate system, and T is the translation vector from the radar coordinate system to the camera coordinate system.
[0022] S2, data collection and preprocessing: The millimeter wave radar sensor and the camera sensor are adjusted to the optimal state, and the intruders on the rails are detected at the same time, and the point cloud and image data are obtained respectively, and then the collected data are preprocessed; the specific process of the preprocessing is as follows:
[0023] The point cloud collected by the millimeter-wave radar sensor includes three-dimensional coordinates, intruder speed, and intruder reflection intensity. Remove the data with intruder reflection intensity less than 5dB, save the processed point cloud, and obtain a point set; crop and scale the image data collected by the camera sensor to a size of 640px×640px, and then normalize the pixel value from [0, 255] to [0, 1] to meet the input requirements of the YOLOv8 model, and then use the labelImg tool to annotate the image data in the YOLO label format. The annotated labels include people, animals, falling rocks, and mudslide blockages, and output the image data set.
[0024] S3, data processing: PCA algorithm and K-means algorithm are used to process the point set obtained in S2, and the features obtained include speed v r 、Position c k and shape to detect intruders; at the same time, the image dataset obtained by S2 is processed using the YOLOv8 model to obtain the two-dimensional coordinates s i =(x i ,y i ), speed v c and shape features to detect intruders.
[0025] Preferably, in step S3, the point set processing method specifically includes the following steps:
[0026] S3.1. The coordinates of the points in the point set are three-dimensional, but the coordinates of the image data are two-dimensional. The PCA algorithm is used to reduce the dimension of the data:
[0027] First, before executing the PCA algorithm, the points are standardized so that each feature of the point (i.e., three-dimensional coordinates, intruder speed) has a mean of zero and a standard deviation of 1. This is because the PCA algorithm relies on the covariance between features, and standardization can eliminate the influence of different scales. The formula used for standardization is:
[0028]
[0029] Where X is the given data matrix, μ is the mean of each feature, and σ is the standard deviation of each feature;
[0030] Then, the core idea of the PCA algorithm is to find the direction with the largest variance in the data. In order to find the direction with the largest variance in the point concentration, it is necessary to calculate the covariance matrix Σ of the data matrix, which is calculated as follows:
[0031]
[0032] Among them, X scaled is the normalized data matrix.
[0033] Next, we use eigenvalue decomposition to find the direction with the largest variance in the point concentration: we perform eigenvalue decomposition on the covariance matrix Σ to obtain the eigenvalues and eigenvectors: Σ = λ i v i , where v i is the i-th eigenvector of the covariance matrix Σ, λ i is the i-th eigenvalue of the covariance matrix Σ.
[0034] The eigenvector represents the direction of data projection, and the eigenvalue represents the variance in that direction. The purpose of the PCA algorithm is to retain the direction with the largest variance, thereby reducing the dimension of the data. Usually, the first k eigenvectors with the largest eigenvalues (i.e., principal components) are selected to reduce the dimensionality of the data. Select the first k eigenvectors with the largest eigenvalues v1, v2, ..., v k , forming a matrix V k , where each column is a feature vector.
[0035] Finally, by projecting the point set into the matrix V k , and get the point set X after dimensionality reduction pca , according to the formula:
[0036] X pca =X scaled V k (8)
[0037] S3.2, the dimension-reduced point set X obtained in S3.1 pca Project it to the pixel coordinate system synchronized with the timestamp for visualization.
[0038] Since the obtained dimensionality reduction point set X pca The midpoint is mainly a two-dimensional coordinate (x ipca ,y ipca ), belongs to two-dimensional data, which can be mapped to the pixel coordinate system through the camera intrinsic parameter matrix L. The intrinsic parameter matrix is:
[0039]
[0040] Among them, f x and f y are the focal lengths of the camera sensor along the x and y axes, respectively, and cx and c y are the coordinates of the principal point of the image (usually the center of the image).
[0041] Based on the internal parameter matrix L, through the formula The reduced point set X can be directly pca Projected onto the pixel coordinate system, we obtain the pixel coordinate set {(u1,v1),(u2,v2),...,(u n ,v n )}, where (u i ,v i ) is the pixel coordinate of the projection of the i-th point in the point set after dimensionality reduction in the pixel coordinate system.
[0042] S3.3, cluster the pixel coordinate set obtained in S3.2 using the K-means clustering algorithm. The K-means clustering algorithm is used to extract key features by dividing the data set into several clusters (a cluster is a cluster or set). The specific process is as follows:
[0043] S3.3.1. Initialize the cluster centers and select m initial cluster centers {c1, c2, ..., c m}, c m =(u m ,v m ), where each c m It is a point in the pixel coordinate system;
[0044] S3.3.2. Calculate each point (u i ,v i ) to the distance of m initial cluster centers, and each point (u i ,v i ) is assigned to the initial cluster center with the smallest distance, thus obtaining m clusters. The distance is the Euclidean distance:
[0045] S3.3.3. Update: Recalculate the center of each of the m clusters. The new cluster center is the mean of all points in each cluster: Among them, S m is the set of points assigned to the mth cluster, |S m | is the number of points in the mth cluster.
[0046] S3.3.4, Iteration: Repeat steps S3.3.2 and S3.3.3. When the cluster center no longer changes or the maximum number of iterations is reached, the K-means algorithm stops and several clusters are obtained.
[0047] After S3.4 and S3.3 obtain several clusters, the location, shape, and speed features of the clusters are extracted to detect intruders. The specific extraction method is as follows:
[0048] First, the location features of the clusters are obtained by calculating the centroid coordinates of the clusters (the centroid coordinates are used as location features). The calculation formula for the centroid coordinates is: And, calculate the width and height of the bounding box of each cluster: W k =u max -u min ,H k =v max -v min ;
[0049] Among them, |S k | is the number of points in the kth cluster, c k is the centroid coordinate of the kth cluster, W k and H k are the width and height of the bounding box of the kth cluster, and the minimum coordinate of the bounding box (u min ,v min ) and the maximum coordinate (u max ,v max )Depend on and Conclude.
[0050] Then, the shape features of the cluster are obtained. The specific process is as follows: Based on the width and height of the bounding box of each cluster, the area A of the cluster is calculated. k : A k =W k ×H k , according to the area A k , calculate the aspect ratio AR k and circularity C k : The shape of the cluster is determined based on the values of the two: for rectangles or ellipses, the value of the aspect ratio is usually large (greater than 1), while for objects close to a circle, the aspect ratio is close to 1; the circularity value is between [0, 1], the closer the value is to 1, the closer the shape of the object is to a circle, and the smaller the value is, the more irregular the shape of the object is; where P k is the boundary perimeter of the kth cluster, which is obtained as follows: the cluster point P = {p1, p2, ..., p n} are all on a two-dimensional plane, and the Graham Scan convex hull algorithm is used to obtain the clustering to obtain the convex hull of the cluster, that is, the convex boundary points of the cluster H = {h1,h2,...,h m}, and then calculate the Euclidean distance between adjacent points in the convex hull: And the Euclidean distance between the mth point and the first point: Finally, the boundary perimeter P can be obtained by the following formula: k :
[0051]
[0052] where h i =(x i ,y i ) is the i-th boundary point.
[0053] S3.5. After the position and shape features of the clusters are extracted in S3.4, the velocity features are extracted according to the displacement and time difference method. The extraction method is as follows: the position features of the intruders, i.e., the centroids c1 and c2, are extracted in the continuous time frames t1 and t2, and the velocity features are obtained by the following formula:
[0054]
[0055] Finally, the feature extraction of the point cloud collected by the millimeter wave radar is completed. The extracted features include the speed v r 、Position c k And shape.
[0056] Preferably, in step S3, the processing method of the image data set specifically includes the following steps:
[0057] S3.5. Obtain CIoU loss function model
[0058] S3.5.1. Divide the image dataset obtained in S2 into a training set, a validation set, and a test set in a ratio of 7:1:2 for optimizing the improved CIoU loss function; the training set is used to train the improved CIoU loss function to obtain the optimized CIoU loss function; the validation set is used to verify the effect of the optimized CIoU loss function offline to facilitate parameter adjustment; the test set is used to test the generalization ability of the CIoU loss function model obtained after the verification is completed, and to determine whether it maintains the same performance on other data.
[0059] S3.5.2. Construction of the improved CIoU loss function. The specific process is as follows:
[0060] The CIoU calculation formula is: Where b = (x b ,y b ,w b ,h b ) is the center coordinate of the ground-truth bounding box (x b ,y b ), width w b and height h b , is the center coordinate of the predicted bounding box width and height is the intersection-over-union ratio of the true bounding box and the predicted bounding box, is the Euclidean distance between the center points: c is the diagonal length of the minimum enclosing box: Among them, x max ,y max and x min ,y min are the diagonal coordinates of the smallest rectangle that contains the true bounding box and the predicted bounding box, and v is the aspect ratio consistency measure: α is an adjustment factor, 0<α<1, which is used to balance the various parts of the loss and is usually set as a hyperparameter.
[0061] Since the initial CIoU is overly dependent on the center point distance, the center point distance weight is too large, especially when the object scales vary greatly, which may lead to inaccurate optimization of the frame of smaller objects and imperfect adjustment of the aspect ratio. Therefore, a stronger scale adaptive mechanism is introduced and a scale adaptive loss function is designed.
[0062] Loss Function Adjust the weights of various indicators (such as center point distance, aspect ratio adjustment) at different scales and improve the aspect ratio loss. Introducing Chamfer distance to replace the original aspect ratio loss can more finely capture the shape differences of the bounding box and improve detection accuracy. Scale-adaptive loss function The calculation formula is:
[0063]
[0064] Among them, A b =w b ×h b is the area of the ground-truth bounding box b, A avg is the average area of all ground-truth bounding boxes, λ is an adjustable hyperparameter that controls the sensitivity to scale differences; P = {(x1,y1),(x2,y2),(x3,y3),(x4,y4)} and They are the four vertex sets of the true bounding box and the predicted bounding box respectively.
[0065] According to equations (15), (16) and (17), the improved CIoU loss function is:
[0066]
[0067] S3.5.3. Optimized and improved CIoU loss function
[0068] The improved CIoU loss function is trained, verified and tested using the training set, validation set and test set respectively to optimize the improved CIoU loss function until the evaluation index meets the accuracy requirements. The evaluation criteria are: the mean average precision (mAP) is improved by more than 0.02, the recall rate (Recall) is improved by more than 0.03, the precision (Precision) is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, repeat steps S3.5.1 and S3.5.2 to adjust and update the parameters until the error meets the requirements or the maximum number of iterations is reached. After the optimization is completed, the CIoU loss function model is obtained, and the optimized parameters in the CIoU loss function model are saved for subsequent intrusion detection.
[0069] S3.6. Use the training set, validation set, and test set described in S3.5.1 to train, validate, and test the YOLOv8 algorithm (including the MobileNetV3 module and the YOLO detection head module) to optimize the YOLOv8 algorithm until the evaluation indicators meet the requirements. The evaluation criteria are: the mean average precision (mAP) is improved by more than 0.02, the recall rate (Recall) is improved by more than 0.03, the precision (Precision) is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, adjust the scale-adaptive loss function. Continue training with λ in until the evaluation criteria are met. The optimized YOLOv8 algorithm has a faster training speed and better training results, so the optimized YOLOv8 algorithm is named the improved YOLOv8 model, and then the optimized parameters obtained in S3.5.3 are loaded into the improved YOLOv8 model; then, the image data normalized in S2 is input into the MobileNetV3 module of the improved YOLOv8 model, and features from low to high levels are extracted from the image data normalized in S2, the features including the two-dimensional coordinates s i =(x i ,y i ), speed v c , and shape features, thereby converting the image data into a multi-level feature map, and saving the feature map for subsequent intruder detection. The speed of the intruder can be obtained based on the two time frames t1 and t2 and the two-dimensional coordinates s1 and s2 of the intruder, that is:
[0070]
[0071] S3.7. The feature map obtained in S3.6 is passed through the YOLO detection head module included in the improved YOLOv8 model, thereby generating multiple prediction boxes in the feature map obtained in S3.6.
[0072] S3.8. Use the non-maximum suppression (NMS) algorithm to process the feature map containing multiple prediction boxes obtained in S3.7, remove redundant boxes from the multiple prediction boxes contained in the feature map, and retain the optimal box, so as to obtain the intruder feature map.
[0073] S4, data fusion: the speed v obtained in S3 r 、Position c k and the two-dimensional coordinates s i =(x i ,y i ), speed v c The shapes and shape features obtained in S3 are fused by weighted averaging after being judged by cosine similarity and then fused by weighted averaging.
[0074] Preferably, in step S4, the speed v obtained in S3 is r 、Position c k and the two-dimensional coordinates s i =(x i ,y i ), speed v c The fusion method of fusion by weighted average is as follows:
[0075] First, different weights are assigned to the features of the point cloud scanned by the millimeter-wave radar sensor and the image data scanned by the camera sensor. Weight Satisfaction and Then, the features corresponding to each intruder i, including radar features {F Radar,1 ,F Radar,2 ,...} and camera features {F Camera,1 ,F Camera,2 ,...}, perform weighted average fusion to obtain the fused features
[0076]
[0077] Among them, F Radar =(c k ,v r ), F Camera =(s i ,v c ), F fusion ={(x,y,z),Speed:_m / s}
[0078] Preferably, in step S4, the shape and shape features obtained in step S3 are judged by cosine similarity and then fused by weighted average, and the specific steps are as follows:
[0079]
[0080] Among them, f radar and f camera They are the shape extracted by S3.4 and the shape features extracted by S3.8. The extracted shapes and shape features are both vector data. ||f radar || and ||f camera || represents the modulus of the vector, also called the norm, and the calculation formula is: Among them, a i are the components of their respective vectors.
[0081] The cosine similarity value range is: [-1,1], 1 means completely similar, -1 means completely dissimilar, and 0 means they are irrelevant. Therefore, when sim(f radar ,f camera )>0.5, weighted average fusion is performed; the fusion formula is:
[0082]
[0083] S5. Detect intruders: The fused features obtained in S4 are input into the YOLOv8 model in S3 again, and intruder detection is performed again to obtain intruder detection results.
[0084] Preferably, the specific process of S5 is:
[0085] First, the fused features obtained by S4 and f fused Input again into the improved YOLOv8 model in S3.6, the improved YOLOv8 model includes a MobileNetV3 module from the input and f fused Extract features from low-level to high-level, and combine the fused features and f fused Converted into a multi-level feature map. Then, the feature map is input into the YOLO detection head module included in the improved YOLOv8 model to generate a prediction box. After being processed by the non-maximum suppression (NMS) algorithm, the redundant boxes in the prediction box are removed, and the optimal box in the prediction box is retained. The optimal box is the detection box, and the detection result is displayed in the detection box, and the confidence is obtained. The detection result is an intruder detection map containing all features, and all the features include the fused position (x, y, z) (z comes from the three-dimensional coordinates obtained by millimeter-wave radar detection), speed Speed: _m / s, and the shape extracted based on the point cloud collected by the millimeter-wave radar sensor, the confidence obtained from the fused information, and other information. Finally, the detection result is output and saved.
[0086] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0087] The present invention provides a method for detecting intruders on rails by heterogeneously fusion of millimeter-wave radar sensors and camera sensors. The millimeter-wave radar sensors and camera sensors are used for real-time detection. After detecting foreign objects, the millimeter-wave radar sensors and camera sensors are associated by Euclidean distance, and then the YOLOv8 model is used for real-time tracking. Finally, the tracking information is fused into a video result by weighted average to make judgments, so as to ensure the safety on the railway in real time. The present invention gives full play to the advantages of millimeter-wave radar and camera for real-time detection, can detect all-weather without being affected by weather, and improves the accuracy and robustness of detection and reduces the false alarm rate by fusing the information of the two sensors. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 The flowchart of the railway foreign object detection method using heterogeneous fusion of millimeter-wave radar and camera;
[0089] Figure 2 This is a diagram of the millimeter wave radar and camera fusion detection framework;
[0090] Figure 3 This is a schematic diagram of the architecture of the improved YOLOv8 model;
[0091] Figure 4 This is the actual processing diagram of the point cloud collected by the millimeter wave radar sensor;
[0092] Figure 5 The multi-level feature map after the improved YOLOv8 model extracts features from the image data at the same time;
[0093] Figure 6 This is the test result obtained in the experiment of the present invention. DETAILED DESCRIPTION
[0094] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0095] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings.
[0096] like Figure 1 , 2As shown, a method for detecting intruders on rails by heterogeneously integrating a millimeter-wave radar sensor and a camera sensor comprises the following steps:
[0097] S1, sensor synchronization: synchronize the timestamps of the millimeter wave radar sensor and the camera sensor and align the spatial coordinates to ensure that the data obtained by the two sensors correspond to the same target;
[0098] First, the NTP network time protocol is used to synchronize timestamps; then spatial coordinate alignment is performed. The spatial coordinate alignment is to convert the point cloud obtained by the millimeter wave radar sensor from the radar coordinate system to the camera coordinate system, thereby achieving the unification of point cloud and image data in the camera coordinate system for effective data fusion.
[0099] The specific process of timestamp synchronization is as follows:
[0100] The clock deviation of the millimeter wave radar sensor is:
[0101]
[0102] The clock bias of the camera sensor is:
[0103]
[0104] The synchronized millimeter wave radar sensor timestamp is:
[0105] T r sync =T r recv +θ r (3)
[0106] The synchronized camera sensor timestamp is:
[0107] T c sync =T c recv +θ c (4)
[0108] Among them, T r req The time when the millimeter wave radar sensor sends the NTP request, T r recv is the time it takes for the millimeter-wave radar sensor to receive a response, T c req The time when the camera sensor sends the NTP request, T c recv is the time when the camera sensor receives the response, T s recv is the time when the time server receives the request, T ssend The time at which the server sends a response.
[0109] The specific process of the spatial coordinate alignment is:
[0110]
[0111] in, is the coordinate of the reflection point of the intruder in the camera coordinate system, is the coordinate of the reflection point of the intruder in the radar coordinate system, R is the rotation matrix from the radar coordinate system to the camera coordinate system, and T is the translation vector from the radar coordinate system to the camera coordinate system.
[0112] S2, data collection and preprocessing: The millimeter wave radar sensor and the camera sensor are adjusted to the optimal state, and the intruders on the rails are detected at the same time, and the point cloud and image data are obtained respectively, and then the collected data are preprocessed; the specific process of the preprocessing is as follows:
[0113] The point cloud collected by the millimeter wave radar sensor includes three-dimensional coordinates, the speed of the intruder, and the reflection intensity of the intruder (each point includes three-dimensional coordinates, the speed of the intruder, and the reflection intensity of the intruder). Since the collected point cloud is very sparse, in order to obtain a denser point cloud, according to the actual detection needs and combined with the optimal detection distance of the millimeter wave radar sensor, the detection distance is set to 50 meters, and when the reflection intensity of the intruder is ≥5dB, the millimeter wave radar sensor can detect the intruder, remove the data with a reflection intensity less than 5dB, save the processed point cloud, and obtain a point set; the image data collected by the camera sensor (image data includes two-dimensional coordinates, image texture, appearance shape features) is cropped and scaled to a size of 640px×640px, and then normalized to normalize the pixel value from [0, 255] to [0, 1] to meet the input requirements of the YOLOv8 model, and then the labelImg tool is used to annotate the image data in the YOLO label format. The annotated labels include people, animals, falling rocks, and mudslide blockages, and the image data set is output.
[0114] S3, data processing: PCA algorithm and K-means algorithm are used to process the point set obtained in S2 to detect intruders; at the same time, the YOLOv8 model is used to process the image data set obtained in S2 to detect intruders.
[0115] The point set processing method specifically includes the following steps:
[0116] S3.1. The coordinates of the points in the point set are three-dimensional, but the coordinates of the image data are two-dimensional. The PCA algorithm is used to reduce the dimension of the data:
[0117] First, before executing the PCA algorithm, the points are standardized so that each feature of the point (i.e., three-dimensional coordinates, intruder speed) has a mean of zero and a standard deviation of 1. This is because the PCA algorithm relies on the covariance between features, and standardization can eliminate the influence of different scales. The formula used for standardization is:
[0118]
[0119] Where X is the given data matrix, μ is the mean of each feature, and σ is the standard deviation of each feature;
[0120] Then, the core idea of the PCA algorithm is to find the direction with the largest variance in the data. In order to find the direction with the largest variance in the point concentration, it is necessary to calculate the covariance matrix Σ of the data matrix, which is calculated as follows:
[0121]
[0122] Among them, X scaled is the normalized data matrix.
[0123] Next, we use eigenvalue decomposition to find the direction with the largest variance in the point concentration: we perform eigenvalue decomposition on the covariance matrix Σ to obtain the eigenvalues and eigenvectors: Σ = λ i v i , where v i is the i-th eigenvector of the covariance matrix Σ, λ i is the i-th eigenvalue of the covariance matrix Σ.
[0124] The eigenvector represents the direction of data projection, and the eigenvalue represents the variance in that direction. The purpose of the PCA algorithm is to retain the direction with the largest variance, thereby reducing the dimension of the data. Usually, the first k eigenvectors with the largest eigenvalues (i.e., principal components) are selected to reduce the dimensionality of the data. Select the first k eigenvectors with the largest eigenvalues v1, v2, ..., v k , forming a matrix V k , where each column is a feature vector.
[0125] Finally, by projecting the point set into the matrix V k , and get the point set X after dimensionality reduction pca , according to the formula:
[0126] X pca =X scaled V k (8)
[0127] S3.2, the dimension-reduced point set X obtained in S3.1 pca Project it to the pixel coordinate system synchronized with the timestamp for visualization.
[0128] Since the obtained dimensionality reduction point set X pca The midpoint is mainly a two-dimensional coordinate (x ipca ,y ipca ), belongs to two-dimensional data, which can be mapped to the pixel coordinate system through the camera intrinsic parameter matrix L. The intrinsic parameter matrix is:
[0129]
[0130] Among them, f x and f y are the focal lengths of the camera sensor along the x and y axes, respectively, and c x and c y are the coordinates of the principal point of the image (usually the center of the image).
[0131] Based on the internal parameter matrix L, through the formula The reduced point set X can be directly pca Projected onto the pixel coordinate system, we obtain the pixel coordinate set {(u1,v1),(u2,v2),...,(u n ,v n )}, where (u i ,v i ) is the pixel coordinate of the projection of the i-th point in the point set after dimensionality reduction in the pixel coordinate system.
[0132] S3.3, cluster the pixel coordinate set obtained in S3.2 using the K-means clustering algorithm. The K-means clustering algorithm is used to extract key features by dividing the data set into several clusters (a cluster is a cluster or set). The specific process is as follows:
[0133] S3.3.1. Initialize the cluster centers and select m initial cluster centers {c1, c2, ..., c m}, c m =(u m ,v m ), where each c m It is a point in the pixel coordinate system;
[0134] S3.3.2. Calculate each point (u i ,v i ) to the distance of m initial cluster centers, and each point (u i ,v i ) is assigned to the initial cluster center with the smallest distance, thus obtaining m clusters. The distance is the Euclidean distance:
[0135] S3.3.3. Update: Recalculate the center of each of the m clusters. The new cluster center is the mean of all points in each cluster: Among them, S m is the set of points assigned to the mth cluster, |S m | is the number of points in the mth cluster.
[0136] S3.3.4, Iteration: Repeat steps S3.3.2 and S3.3.3. When the cluster center no longer changes or the maximum number of iterations is reached, the K-means algorithm stops and several clusters are obtained.
[0137] After S3.4 and S3.3 obtain several clusters, the location, shape, and speed features of the clusters are extracted to detect intruders. The specific extraction method is as follows:
[0138] First, the location features of the clusters are obtained by calculating the centroid coordinates of the clusters (the centroid coordinates are used as location features). The calculation formula for the centroid coordinates is: And, calculate the width and height of the bounding box of each cluster: W k =u max -u min ,H k =v max -v min ;
[0139] Among them, |S k | is the number of points in the kth cluster, c k is the centroid coordinate of the kth cluster, W k and H k are the width and height of the bounding box of the kth cluster, and the minimum coordinate of the bounding box (u min ,v min ) and the maximum coordinate (u max ,v max )Depend on and Conclude.
[0140] Then, the shape features of the cluster are obtained. The specific process is as follows: Based on the width and height of the bounding box of each cluster, the area A of the cluster is calculated. k : A k =W k ×H k , according to the area A k , calculate the aspect ratio AR k and circularity C k : The shape of the cluster is determined based on the values of the two: for rectangles or ellipses, the value of the aspect ratio is usually large (greater than 1), while for objects close to a circle, the aspect ratio is close to 1; the circularity value is between [0, 1], the closer the value is to 1, the closer the shape of the object is to a circle, and the smaller the value is, the more irregular the shape of the object is; where P kis the boundary perimeter of the kth cluster, which is obtained as follows: the cluster point P = {p1, p2, ..., p n} are all on a two-dimensional plane, and the GrahamScan convex hull algorithm is used to obtain the clustering to obtain the convex hull of the cluster, that is, the convex boundary points of the cluster H = {h1,h2,...,h m}, and then calculate the Euclidean distance between adjacent points in the convex hull: And the Euclidean distance between the mth point and the first point: Finally, the boundary perimeter P can be obtained by the following formula: k :
[0141]
[0142] where h i =(x i ,y i ) is the i-th boundary point.
[0143] S3.5. After the position and shape features of the clusters are extracted in S3.4, the velocity features are extracted according to the displacement and time difference method. The extraction method is as follows: the position features of the intruders, i.e., the centroids c1 and c2, are extracted in the continuous time frames t1 and t2, and the velocity features are obtained by the following formula:
[0144]
[0145] Finally, the feature extraction of the point cloud collected by the millimeter wave radar is completed. The extracted features include the speed v r 、Position c k And shape.
[0146] The processing method of the image dataset specifically includes the following steps:
[0147] S3.5. Obtain CIoU loss function model
[0148] S3.5.1. Divide the image dataset obtained in S2 into a training set, a validation set, and a test set in a ratio of 7:1:2 for optimizing the improved CIoU loss function; the training set is used to train the improved CIoU loss function to obtain the optimized CIoU loss function; the validation set is used to verify the effect of the optimized CIoU loss function offline to facilitate parameter adjustment; the test set is used to test the generalization ability of the CIoU loss function model obtained after the verification is completed, and to determine whether it maintains the same performance on other data.
[0149] S3.5.2. Construction of the improved CIoU loss function. The specific process is as follows:
[0150] The CIoU calculation formula is: Where b = (xb ,y b ,w b ,h b ) is the center coordinate of the ground-truth bounding box (x b ,y b ), width w b and height h b , is the center coordinate of the predicted bounding box width and height is the intersection-over-union ratio of the true bounding box and the predicted bounding box, is the Euclidean distance between the center points: c is the diagonal length of the minimum enclosing box: Among them, x max ,y max and x min ,y min are the diagonal coordinates of the smallest rectangle that contains the true bounding box and the predicted bounding box, and v is the aspect ratio consistency measure: α is an adjustment factor, 0<α<1, which is used to balance the various parts of the loss and is usually set as a hyperparameter.
[0151] Since the initial CIoU is overly dependent on the center point distance, the center point distance weight is too large, especially when the object scale difference is large, which may lead to inaccurate optimization of the frame of smaller objects and imperfect adjustment of the aspect ratio. Therefore, a stronger scale adaptation mechanism is introduced and a scale adaptive loss function is designed. Adjust the weights of various indicators (such as center point distance, aspect ratio adjustment) at different scales and improve the aspect ratio loss. Introducing Chamfer distance to replace the original aspect ratio loss can more finely capture the shape differences of the bounding box and improve detection accuracy. Scale-adaptive loss function The calculation formula is:
[0152]
[0153]
[0154] Among them, A b =w b ×h b is the area of the ground-truth bounding box b, A avg is the average area of all ground-truth bounding boxes, λ is an adjustable hyperparameter that controls the sensitivity to scale differences; P = {(x1,y1),(x2,y2),(x3,y3),(x4,y4)} and They are the four vertex sets of the true bounding box and the predicted bounding box respectively.
[0155] According to equations (15), (16) and (17), the improved CIoU loss function is:
[0156]
[0157] S3.5.3. Optimized and improved CIoU loss function
[0158] The improved CIoU loss function is trained, verified and tested using the training set, validation set and test set respectively to optimize the improved CIoU loss function until the evaluation index meets the accuracy requirements. The evaluation criteria are: the mean average precision (mAP) is improved by more than 0.02, the recall rate (Recall) is improved by more than 0.03, the precision (Precision) is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, repeat steps S3.5.1 and S3.5.2 to adjust and update the parameters until the error meets the requirements or the maximum number of iterations is reached. After the optimization is completed, the CIoU loss function model is obtained, and the optimized parameters in the CIoU loss function model are saved for subsequent intrusion detection.
[0159] S3.6. Use the training set, validation set, and test set described in S3.5.1 to train, validate, and test the YOLOv8 algorithm (including the MobileNetV3 module and the YOLO detection head module) to optimize the YOLOv8 algorithm until the evaluation indicators meet the requirements. The evaluation criteria are: the mean average precision (mAP) is improved by more than 0.02, the recall rate (Recall) is improved by more than 0.03, the precision (Precision) is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, adjust the scale-adaptive loss function. Continue training with λ in until the evaluation criteria are met. The optimized YOLOv8 algorithm has a faster training speed and better training results, so the optimized YOLOv8 algorithm is named the improved YOLOv8 model, and then the optimized parameters obtained in S3.5.3 are loaded into the improved YOLOv8 model; then, the image data normalized in S2 is input into the MobileNetV3 module of the improved YOLOv8 model, and features from low to high levels are extracted from the image data normalized in S2, the features including the two-dimensional coordinates s i =(x i ,y i ), speed v c , and shape features, thereby converting the image data into a multi-level feature map, and saving the feature map for subsequent intruder detection. The speed of the intruder can be obtained based on the two time frames t1 and t2 and the two-dimensional coordinates s1 and s2 of the intruder, that is:
[0160]
[0161] S3.7. The feature map obtained in S3.6 is passed through the YOLO detection head module included in the improved YOLOv8 model, thereby generating multiple prediction boxes in the feature map obtained in S3.6.
[0162] S3.8. Use the non-maximum suppression (NMS) algorithm to process the feature map containing multiple prediction boxes obtained in S3.7, remove redundant boxes from the multiple prediction boxes contained in the feature map, and retain the optimal box, so as to obtain the intruder feature map.
[0163] S4: Data fusion: The features obtained in S3.4 include speed v r 、Position c k and shape. The characteristic map of the intruder obtained in S3.6 includes the two-dimensional coordinates s i =(x i ,y i ), speed v c As well as the shape characteristics, the velocity v obtained in S3.4 is r 、Position c k and the two-dimensional coordinates s obtained by S3.6 i =(x i ,y i ), speed v c The fusion is performed by weighted averaging. The fusion method is as follows:
[0164] First, different weights are assigned to the features of the point cloud scanned by the millimeter-wave radar sensor and the image data scanned by the camera sensor. Weight Satisfaction and Then, the features corresponding to each intruder i, including radar features {F Radar,1 ,F Radar,2 ,...} and camera features {F Camera,1 ,F Camera,2 ,...}, perform weighted average fusion to obtain the fused features
[0165]
[0166] Among them, F Radar =(c k ,v r ), F Camera =(s i ,v c ), F fusion={(x,y,z),Speed:_m / s}, where z is derived from the three-dimensional coordinates obtained by millimeter-wave radar detection after processing.
[0167] The shape features obtained by S3.4 and the shape features obtained by S3.6 are judged by cosine similarity and then fused by weighted average:
[0168]
[0169] Among them, f radar and f camera They are the shape extracted by S3.4 and the shape features extracted by S3.8. The extracted shapes and shape features are both vector data. ||f radar || and ||f camera || represents the modulus of the vector, also called the norm, and the calculation formula is: Among them, a i are the components of their respective vectors.
[0170] The cosine similarity value range is: [-1,1], 1 means completely similar, -1 means completely dissimilar, and 0 means they are irrelevant. Therefore, when sim(f radar ,f camera )>0.5, weighted average fusion is performed; the fusion formula is:
[0171]
[0172] S5, detect intruders: input the fused features obtained in S4 into the YOLOv8 model in S3.6 again, perform intruder detection again, and obtain the intruder detection result. The specific process is:
[0173] First, the fused features obtained by S4 and f fused Input again into the improved YOLOv8 model in S3.6, the improved YOLOv8 model includes a MobileNetV3 module from the input and f fused Extract features from low-level to high-level, and combine the fused features and f fusedConverted into a multi-level feature map. Then, the feature map is input into the YOLO detection head module included in the improved YOLOv8 model to generate a prediction box. After being processed by the non-maximum suppression (NMS) algorithm, the redundant boxes in the prediction box are removed, and the optimal box in the prediction box is retained. The optimal box is the detection box, and the detection result is displayed in the detection box, and the confidence is obtained. The detection result is an intruder detection map containing all features, and all the features include the fused position (x, y, z), speed Speed: _m / s, and the shape extracted based on the point cloud collected by the millimeter wave radar sensor, the confidence obtained from the fused information, and other information. Finally, the detection result is output and saved.
[0174] Figure 3 This is a schematic diagram of the improved YOLOv8 model. In addition to the improved CIoU loss function, it also includes the MobileNetV3 module and the YOLO detection head module. The MobileNetV3 module uses deep separable convolution to reduce the amount of calculation and the number of parameters, and combines the reverse residual structure to improve the feature extraction ability and information flow. The introduced SE module (Squeeze-and-Excitation) enhances the network performance through the channel attention mechanism, and the Hard Swish activation function replaces the traditional ReLU activation function to improve the computational efficiency and accuracy. The linear bottleneck structure avoids information loss and enhances the low-dimensional feature expression ability. Through the neural architecture search (NAS), MobileNetV3 combines the EfficientNet idea to automatically design an efficient network structure, and finally classifies the target through a simple classifier layer, further improving the efficiency and accuracy. The YOLO detection head module uses an upsampling layer to improve the resolution of the feature map and enhance the detail information; then, the feature map splices and fuses the feature maps from different scales, combining the low-level and high-level information to enrich the target representation. The feature map is further optimized through the convolution operation to extract more discriminative features. The introduction of attention mechanisms, such as channel attention, enhances the representation ability of important features. The model generates the coordinates and category labels of the predicted boxes through bounding box regression and classification, and finally uses non-maximum suppression (NMS) to remove redundant boxes and optimize the final detection results.
[0175] For millimeter wave radar sensors, the first data collected is 3D point cloud ( Figure 4 (a)), after data standardization and PCA algorithm dimensionality reduction, we can get Figure 4 (b) The 2D point cloud is then projected into the pixel coordinate system to obtain Figure 4 (c) The point cloud is then clustered using the K-means algorithm. Since there is only one man-made intruder in the image, we get Figure 4In the point cloud image (d), there is only one point cloud cluster. That is, the main features can be extracted and the main features can be obtained for fusion.
[0176] Figure 5 The improved YOLOv8 model extracts features from the image at the same time, which is a multi-level feature map, including detection boxes, two-dimensional coordinates, and confidence levels. Figure 6 It is the final detection image extracted by the improved YOLOv8 model after feature fusion, which includes the position (x, y, z), speed Speed: _m / s, shape extracted from the point cloud collected by the millimeter wave radar sensor, and confidence extracted from the image data collected by the camera sensor: Figure 6 As shown, the real-time data of x:2.76, y:9.13, z:3.38, Speed:1.08m / s and confidence level person:0.9, where z is derived from the three-dimensional coordinates obtained by millimeter wave radar detection after processing.
[0177] The present invention detects intruders through the steps of sensor synchronization, data acquisition, data processing, data fusion and detection of intruders. Sensor synchronization mainly synchronizes the timestamps of the millimeter wave radar sensor and the camera sensor and aligns the spatial coordinates to ensure that the two sensors describe the same target; data acquisition is that the millimeter wave radar sensor and the camera sensor collect data at the same time, the millimeter wave radar sensor collects point cloud data, the camera sensor collects image data, and the collected data is made into a data set through preprocessing. Data processing is to process the point cloud data set collected by the millimeter wave radar sensor through the PAC algorithm and the K-means algorithm to obtain the main features. The YOLOv8 model is trained by improving the CIoU loss function to obtain the optimized network weight after the training is completed, and then the image data set collected by the camera sensor is processed according to the network weight to obtain a multi-level feature map. Data fusion is to fuse the features obtained by the two sensors to obtain the fused features. Finally, the obtained fused features are put into the trained improved YOLOv8 model for re-detection to obtain the final output result. The present invention performs intrusion detection by fusing the features of two sensors, combines the advantages of the two sensors, can perform intrusion detection in complex environments, improves redundancy and reliability, and enhances target recognition and positioning accuracy.
[0178] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.
Claims
1. A method for detecting intruders on rails by heterogeneously integrating a millimeter-wave radar sensor and a camera sensor, comprising the following steps: S1. Sensor synchronization: synchronize the timestamps of the millimeter-wave radar sensor and the camera sensor and align the spatial coordinates to ensure that the data obtained by the two sensors correspond to the same target: first use the NTP network time protocol to synchronize the timestamps; Then, spatial coordinate alignment is performed, wherein the point cloud obtained by the millimeter wave radar sensor is converted from the radar coordinate system to the camera coordinate system, thereby realizing the unification of the point cloud and image data in the camera coordinate system; S2, data collection and preprocessing: The millimeter wave radar sensor and the camera sensor are adjusted to the optimal state, and the intruders on the rails are detected at the same time, and the point cloud and image data are obtained respectively, and then the collected data are preprocessed; the specific process of the preprocessing is as follows: The point cloud collected by the millimeter wave radar sensor includes three-dimensional coordinates, the speed of the intruder, and the reflection intensity of the intruder. The data with the reflection intensity of the intruder less than 5dB is removed, and the processed point cloud is saved to obtain a point set. The image data collected by the camera sensor is cropped and scaled to a size of 640px×640px, and then normalized to normalize the pixel value from [0, 255] to [0, 1] to meet the input requirements of the YOLOv8 model. The labelImg tool is then used to annotate the image data in the YOLO label format. The annotated labels include people, animals, falling rocks, and mudslide blockages, and the image data set is output; S3, data processing: PCA algorithm and K-means algorithm are used to process the point set obtained in S2, and the features obtained include speed v r 、Position c k and shape to detect intruders; at the same time, the image dataset obtained by S2 is processed using the YOLOv8 model to obtain the two-dimensional coordinates s i =(x i ,y i ), speed v c and shape features to detect intruders; S4, data fusion: the speed v obtained in S3 r 、Position c k and the two-dimensional coordinates s i =(x i ,y i ), speed v c The shapes and shape features obtained in S3 are fused by weighted average after being judged by cosine similarity and then fused by weighted average; S5. Detect intruders: The fused features obtained in S4 are input into the YOLOv8 model in S3 again, and intruder detection is performed again to obtain intruder detection results.
2. The method for detecting intruders on rails according to claim 1, characterized in that: In step S1, the specific process of timestamp synchronization is: The clock deviation of the millimeter wave radar sensor is: The clock bias of the camera sensor is: The synchronized millimeter wave radar sensor timestamp is: T rsync =T rrecv +θ r (3) The synchronized camera sensor timestamp is: T csync =T crecv +θ c (4) Among them, T rreq The time when the millimeter wave radar sensor sends the NTP request, T rrecv is the time it takes for the millimeter-wave radar sensor to receive a response, T creq The time when the camera sensor sends the NTP request, T crecv is the time when the camera sensor receives the response, T srecv is the time when the time server receives the request, T ssend The time when the time server sends the response; The specific process of the spatial coordinate alignment is: in, is the coordinate of the reflection point of the intruder in the camera coordinate system, is the coordinate of the reflection point of the intruder in the radar coordinate system, R is the rotation matrix from the radar coordinate system to the camera coordinate system, and T is the translation vector from the radar coordinate system to the camera coordinate system.
3. The method for detecting intruders on rails according to claim 1, characterized in that: In step S3, the point set processing method specifically includes the following steps: S3.
1. The coordinates of the points in the point set are three-dimensional, but the coordinates of the image data are two-dimensional. The PCA algorithm is used to reduce the dimension of the data: First, before executing the PCA algorithm, the points are standardized so that each feature of the point has a mean of zero and a standard deviation of 1; the formula used for standardization is: Where X is the given data matrix, μ is the mean of each feature, and σ is the standard deviation of each feature; Then, in order to find the direction with the largest variance in the point set, the covariance matrix Σ of the data matrix is calculated as follows: Among them, X scaled is the standardized data matrix; Next, we use eigenvalue decomposition to find the direction with the largest variance in the point concentration: we perform eigenvalue decomposition on the covariance matrix Σ to obtain the eigenvalues and eigenvectors: Σ = λ i v i , where v i is the i-th eigenvector of the covariance matrix Σ, λ i is the i-th eigenvalue of the covariance matrix Σ; The eigenvector represents the direction of data projection, and the eigenvalue represents the variance in that direction; select the first k eigenvectors with the largest eigenvalues to perform data dimensionality reduction; select the first k eigenvectors with the largest eigenvalues v1, v2, ..., v k , forming a matrix V k , where each column is a feature vector; Finally, by projecting the point set into the matrix V k , and get the point set X after dimensionality reduction pca , according to the formula: X pca =X scaled V k (8) S3.2, the dimension-reduced point set X obtained in S3.1 pca Projecting to the pixel coordinate system synchronized with the timestamp for visualization; Point set X after dimensionality reduction pca The midpoint is the 2D coordinate (x ipca ,y ipca ), belongs to two-dimensional data, and the two-dimensional data is mapped to the pixel coordinate system through the camera intrinsic parameter matrix L. The intrinsic parameter matrix is: Among them, f x and f y are the focal lengths of the camera sensor along the x and y axes, respectively, and c x and c y are the principal point coordinates of the image; Based on the internal parameter matrix L, through the formula Directly reduce the dimension of the point set X pca Projected onto the pixel coordinate system, we obtain the pixel coordinate set {(u1,v1),(u2,v2),...,(u n ,v n )}, where (u i ,v i ) is the pixel coordinate of the projection of the i-th point in the point set after dimensionality reduction in the pixel coordinate system; S3.3, clustering the pixel coordinate set obtained in S3.2 using the K-means clustering algorithm. The K-means clustering algorithm is used to extract key features by dividing the data set into several clusters. The specific process is as follows: S3.3.
1. Initialize the cluster centers and select m initial cluster centers {c1, c2, ..., c m }, c m =(u m ,v m ), where each c m It is a point in the pixel coordinate system; S3.3.
2. Calculate each point (u i ,v i ) to the distance of m initial cluster centers, and each point (u i ,v i ) is assigned to the initial cluster center with the smallest distance, thereby obtaining m clusters; the distance is the Euclidean distance: S3.3.
3. Update: Recalculate the center of each cluster in the m clusters, and the new cluster center is Among them, S m is the set of points assigned to the mth cluster, |S m | is the number of points in the mth cluster; S3.3.4, Iteration: Repeat steps S3.3.2 and S3.3.
3. When the cluster center no longer changes or the maximum number of iterations is reached, the K-means algorithm stops and several clusters are obtained; After S3.4 and S3.3 obtain several clusters, the features of the clusters including position, shape, and speed are extracted. The specific extraction method is as follows: First, the location features of the clusters are obtained by calculating the centroid coordinates of the clusters. The calculation formula for the centroid coordinates is: And, calculate the width and height of the bounding box of each cluster: W k =u max -u min ,H k =v max -v min ; Among them, |S k | is the number of points in the kth cluster, c k is the centroid coordinate of the kth cluster, W k and H k are the width and height of the bounding box of the kth cluster, and the minimum coordinate of the bounding box (u min ,v min ) and the maximum coordinate (u max ,v max )Depend on to conclude; Then, the shape features of the cluster are obtained. The specific process is as follows: Based on the width and height of the bounding box of each cluster, the area A of the cluster is calculated. k : A k =W k ×H k , according to the area A k , calculate the aspect ratio AR k and circularity C k : The shape of the cluster is determined based on the values of the two: for rectangles or ellipses, the value of the aspect ratio is usually larger, while for objects close to a circle, the aspect ratio is close to 1; the circularity value is between [0, 1], the closer the value is to 1, the closer the shape of the object is to a circle, and the smaller the value is, the more irregular the shape of the object is; where P k is the boundary perimeter of the kth cluster, which is obtained as follows: the cluster point P = {p1, p2, ..., p n } are all on a two-dimensional plane, and the Graham Scan convex hull algorithm is used to obtain the clustering to obtain the convex hull of the cluster, that is, the convex boundary points of the cluster H = {h1,h2,...,h m }, and then calculate the Euclidean distance between adjacent points in the convex hull: And the Euclidean distance between the mth point and the first point: Finally, the boundary perimeter P can be obtained by the following formula: k : where h i =(x i ,y i ) is the i-th boundary point; S3.
5. After the position and shape features of the clusters are extracted in S3.4, the velocity features are extracted according to the displacement and time difference method. The extraction method is as follows: the position features of the intruders, i.e., the centroids c1 and c2, are extracted in the continuous time frames t1 and t2, and the velocity features are obtained by the following formula: Finally, the feature extraction of the point cloud collected by the millimeter wave radar is completed. The extracted features include the speed v r 、Position c k And shape.
4. The method for detecting intruders on rails according to claim 1, characterized in that: In step S3, the processing method of the image data set specifically includes the following steps: S3.
5. Obtain CIoU loss function model S3.5.1, divide the image dataset obtained in S2 into training set, validation set and test set in a ratio of 7:1:2 to optimize the improved CIoU loss function; S3.5.
2. Construction of the improved CIoU loss function. The specific process is as follows: Where b = (x b ,y b ,w b ,h b ) is the center coordinate of the ground-truth bounding box (x b ,y b ), width w b and height h b , is the center coordinate of the predicted bounding box width and height is the intersection-over-union ratio of the true bounding box and the predicted bounding box, is the Euclidean distance between the center points: c is the diagonal length of the minimum enclosing box: Among them, x max ,y max and x min ,y min are the diagonal coordinates of the smallest rectangle that contains the true bounding box and the predicted bounding box, and v is the aspect ratio consistency measure: α is an adjustment factor, 0<α<1, which is used to balance the various parts of the loss and is set as a hyperparameter; Introducing a scale-adaptive mechanism, a scale-adaptive loss function Adjust the weights of various indicators at different scales and improve the aspect ratio loss. Introduce Chamfer distance to replace the original aspect ratio loss to capture the shape difference of the bounding box; scale-adaptive loss function The calculation formula is: Among them, A b =w b ×h b is the area of the ground-truth bounding box b, A avg is the average area of all ground-truth bounding boxes, λ is an adjustable hyperparameter that controls the sensitivity to scale differences; P = {(x1,y1),(x2,y2),(x3,y3),(x4,y4)} and These are the four vertex sets of the true bounding box and the predicted bounding box respectively; According to equations (15), (16) and (17), the improved CIoU loss function is: S3.5.
3. Optimized and improved CIoU loss function The improved CIoU loss function is trained, verified and tested using the training set, validation set and test set to optimize the improved CIoU loss function until the evaluation index meets the accuracy requirements. The evaluation criteria are: the average precision mean is improved by more than 0.02, the recall rate is improved by more than 0.03, the accuracy is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, steps S3.5.1 and S3.5.2 are repeated to adjust and update the parameters until the error meets the requirements or the maximum number of iterations is reached. After the optimization is completed, the CIoU loss function model is obtained, and the optimized parameters in the CIoU loss function model are saved for subsequent intrusion detection. S3.
6. Use the training set, validation set and test set described in S3.5.1 to train, validate and test the YOLOv8 algorithm respectively. The YOLOv8 algorithm includes a MobileNetV3 module and a YOLO detection head module to optimize the YOLOv8 algorithm until the evaluation index meets the requirements. The evaluation criteria are: the average precision is improved by more than 0.02, the recall rate is improved by more than 0.03, the precision is improved by more than 0.02, and the positioning error is reduced. If the evaluation criteria are not met, adjust the scale-adaptive loss function. Continue training with λ in until the evaluation criteria are met; name the optimized YOLOv8 algorithm as the improved YOLOv8 model, and then load the optimized parameters obtained in S3.5.3 into the improved YOLOv8 model; then, input the image data normalized in S2 into the MobileNetV3 module of the improved YOLOv8 model, and extract features from low to high levels from the image data normalized in S2, the features including the two-dimensional coordinates s i =(x i ,y i ), speed v c , and shape features, thereby converting the image data into a multi-level feature map, and saving the feature map for subsequent intruder detection; wherein the speed is obtained based on the two time frames t1 and t2 and the two-dimensional coordinates s1 and s2 of the intruder, that is: S3.7, passing the feature map obtained in S3.6 through the YOLO detection head module included in the improved YOLOv8 model, thereby generating multiple prediction boxes in the feature map obtained in S3.6; S3.
8. Use a non-maximum suppression algorithm to process the feature map containing multiple prediction boxes obtained in S3.7, remove redundant boxes from the multiple prediction boxes contained in the feature map, and retain the optimal box, so as to obtain an intruder feature map.
5. The method for detecting intruders on rails according to claim 1, characterized in that: In step S4, the speed v obtained in step S3 is r 、Position c k and the two-dimensional coordinates s i =(x i ,y i ), speed v c The fusion method of fusion by weighted average is as follows: First, different weights are assigned to the features of the point cloud scanned by the millimeter-wave radar sensor and the image data scanned by the camera sensor. Weight Satisfaction and Then, the features corresponding to each intruder i, including radar features {F Radar,1 ,F Radar,2 ,...} and camera features {F Camera,1 ,F Camera,2 ,...}, perform weighted average fusion to obtain the fused features Among them, F Radar =(c k , v r ), F Camera =(s i , v c ), F fusion ={(x, y, z), Speed: _m / s}.
6. The method for detecting intruders on rails according to claim 1, characterized in that: In step S4, the shapes and shape features obtained in step S3 are judged by cosine similarity and then fused by weighted average. The specific steps are as follows: Among them, f radar and f camera are the shape and shape features extracted from S3, and the extracted shape and shape features are both vector data; ||f radar || and ||f camera || represents the modulus of the vector, and the calculation formula is: Among them, a i are the components of their respective vectors; Among them, the cosine similarity value range is: [-1,1], 1 means complete similarity, -1 means complete dissimilarity, and 0 means that the two are irrelevant. Therefore, when sim(f radar ,f camera )>0.5, weighted average fusion is performed; the fusion formula is:
7. The method for detecting intruders on rails according to claim 1, characterized in that: The specific process of S5 is as follows: First, the fused features obtained by S4 and f fused Input again into the improved YOLOv8 model in S3, the improved YOLOv8 model includes the MobileNetV3 module from the input and f fused Extract features from low-level to high-level, and combine the fused features and f fused The method converts the feature map into a multi-level feature map; then, the feature map is input into the YOLO detection head module included in the improved YOLOv8 model to generate a prediction box; then, the non-maximum suppression algorithm is used to remove redundant boxes in the prediction box, and the optimal box in the prediction box is retained, and the optimal box is the detection box. The detection result is displayed in the detection box, and the confidence level is obtained. The detection result is an intruder detection map containing all features; finally, the detection result is output and saved.
Citation Information
Patent Citations
Railway vehicle detection system and method based on millimeter wave radar and camera fusion
CN114814823A
Mining area environment sensing method based on 4D millimeter wave radar
CN115236674A
Single-machine track foreign matter detection method based on computer vision and radar fusion
CN118072233A
Cited By
Track line induction board vehicle-mounted dynamic detection method
CN120673174A
Track traffic abnormal event identification method and device, electronic equipment and storage medium
CN120673351A
Vehicle blind area anti-collision method, device and equipment and storage medium
CN120871175A
Railway intrusion target detection method, device, equipment, medium and product
CN121033382A
Real-time mountain torrent disaster monitoring method and system integrating radar and video
CN121191275A