Urban rail train space foreign matter invasion sensing method based on anchored rail line detection
By setting up cameras and lidars in urban rail trains, establishing a 3D track line extraction network, extracting the 3D track center line and building a boundary space, the problem of insufficient accuracy and robustness of track intrusion detection in complex environments is solved, and high-precision and real-time foreign object intrusion detection is achieved.
Patent Information
- Application Number
- CN202510205791.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art methods for achieving high-precision, real-time sensing of orbit intrusion in complex environments are insufficient, especially in poor performance in unknown categories of foreign object detection.
The urban rail train space foreign object intrusion perception method is adopted based on anchored track line detection. By setting up cameras and lidars, environmental data are collected and track space position information is marked, a 3D track line extraction network is established, and the 3D track center line is extracted. Based on this, the bounded space is constructed, and data is collected in real time by using lidar and cameras to realize foreign object intrusion detection.
It significantly improves the accuracy and robustness of 3D track line detection, can accurately detect any type of obstacle in complex environments, and improves the safety and intelligent management level of rail transit system.
Smart Images

Figure CN120071301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method for perceiving spatial foreign object intrusion of urban rail trains based on the detection of anchored track lines. Background Art
[0002] With the rapid advancement of urbanization, urban rail transit has gradually become an important part of the public transportation system due to its high carrying capacity, high departure frequency, and high punctuality rate, and its safety has received increasing attention. Therefore, it is urgent to achieve real-time and accurate perception of foreign objects invading the limited space of urban rail trains to avoid such accidents.
[0003] Train limited boundary intrusion detection refers to actively identifying abnormal objects within the forward limited boundary that may affect train operation during train operation, providing warning or braking information for the train, thereby ensuring train operation safety. At present, train limited boundary intrusion detection mainly relies on image detection and radar detection of known types of obstacles. There is still a lack of a method that can achieve high-precision and real-time perception of track intrusion in complex environments to improve the safety of rail transit systems and promote the improvement of intelligent management and operation levels of rail transit.
[0004] Currently, the technologies for intrusion detection are divided into two categories: contact type and non-contact type. Contact type intrusion detection technologies detect intrusion behaviors by sensing physical signals such as physical contact, pressure changes, or vibrations. Non-contact type intrusion detection technologies detect obstacle intrusion through a technical route that obtains a large amount of video data through on-vehicle cameras of rail trains and judges the positional relationship between the obstacle and the track. At present, non-contact train limited boundary intrusion detection mainly relies on image detection and radar detection of known types of obstacles. Image-based intrusion detection identifies forward obstacles in the track environment based on common object detection networks through region division. Lidar-based intrusion detection determines the type of obstacle by segmenting and classifying lidar point clouds, demonstrating the possibility of obstacle detection. The above methods face challenges such as inaccurate division of the track limited boundary range and difficulty in detecting foreign object intrusion. Especially when faced with foreign objects of unknown types, the existing detection capabilities are particularly insufficient.
[0005] The Chinese patent application with the publication number CN117237613A discloses a method for detecting foreign object intrusion based on a convolutional neural network. First, radar point cloud data and visual images are obtained and time-aligned to obtain a radar projection image. The constructed improved YOLOv5 network adds a feature extraction branch in the feature extraction part and a splicing module in the original feature extraction branch, and the detection accuracy of small targets is improved. However, the intrusion perception area is not clearly defined, and false detection is likely to occur. In addition, when faced with foreign objects of unknown types, the existing detection capabilities are particularly insufficient. Summary of the Invention
[0006] In view of the above problems, the present invention provides a method for sensing spatial foreign object intrusion of urban rail trains based on the detection of anchored track lines, which solves the technical problems of low accuracy in long-distance track limit sensing and unknown category foreign object detection in the prior art.
[0007] The present invention provides a method for sensing spatial foreign object intrusion of urban rail trains based on the detection of anchored track lines, including the following steps:
[0008] Step S1: Install a camera and a lidar on the urban rail train, and use the camera and the lidar to collect environmental data, and the environmental data has continuous time stamps; label the track spatial position information for the environmental data to form a track line data set;
[0009] Step S2: Establish a 3D track line extraction network, and the 3D track line extraction network includes a feature extraction network, a temporal fusion module, and a detection head; train the 3D track line extraction network based on the track line data set to obtain a trained 3D track line extraction network;
[0010] Step S3: Input the data collected by the camera in real time into the trained 3D track line extraction network to obtain a 3D track line; based on the track width constraint, obtain a 3D track center line from the 3D track line;
[0011] Step S4: Establish a limit space based on a preset plurality of bounding box sizes and the 3D track center line, and the limit space is the space where foreign object intrusion needs to be detected;
[0012] Step S5: Based on the data collected by the lidar in real time and the limit space, obtain limit space intrusion sensing data, and obtain the spatial position of the foreign object based on the limit space intrusion sensing data.
[0013] Preferably, step S1 specifically includes:
[0014] Step S1-1: Install a camera and a lidar in the carriage of the urban rail train. In a variety of urban rail train operation scenarios and a variety of track line forms, the camera and the lidar respectively record image sequences and point cloud sequences within multiple time periods; the image sequences and point cloud sequences have continuous timestamps;
[0015] Step S1-2: Synchronize the time of the image sequence and the point cloud sequence, including: perform matching selection according to the timestamps in the image sequence and the point cloud sequence, and select the image and the point cloud with the same timestamp as the environment data for time synchronization;
[0016] Step S1-3: For the environment data with time synchronization, in the order of timestamps, use the automatic annotation technology to obtain the categories and 3D spatial coordinates of multiple points on the track line in each frame of the image sequence and the point cloud sequence as the ground truth labels;
[0017] Step S1-4: Use each frame of the image, each frame of the point cloud, and the corresponding ground truth labels together as the samples of the dataset to form a track line dataset, and randomly select a preset number of samples from the samples of the track line dataset to form a track line data test set.
[0018] Preferably, the specific steps of Step S2 include:
[0019] Step S2-1: Based on the ResNet-18 network, establish a feature extraction network. The feature extraction network inputs multiple image frames arranged in chronological order to obtain the multi-scale features of the multiple image frames;
[0020] Step S2-2: Establish a temporal fusion module. The temporal fusion module performs temporal network convolution processing, feature dynamic alignment at different time steps, and temporal scale fusion on the multi-scale features of the multiple image frames to finally obtain the fused features;
[0021] Step S2-3: Establish a detection head. The detection head includes a classification head and a regression head. The classification head and the regression head obtain the 3D spatial coordinates and category estimation values of multiple points on the track line based on the fused features;
[0022] Step S2-4: Determine the loss function of the 3D track line extraction network. Based on the track line dataset and the track line data test set, use minimizing the loss function as the optimization goal to train the 3D track line extraction network to obtain the trained 3D track line extraction network.
[0023] Preferably, in Step S2-2, the expression of the temporal network convolution processing is:
[0024] H t = ReLU(W * H t-1 + b)
[0025] where, H t represents the feature map of the t-th frame of the input image I t , W and b are the convolution kernel and the bias respectively, and * represents the one-dimensional convolution operation;
[0026] The expression of the feature dynamic alignment at different time steps is:
[0027] F t→t+1 = FlowNet(I t , I t+1 )
[0028] H′ t = Warp(H t , F t→t+1 )
[0029] where I t , I t+1 are the input images of the t-th and (t + 1)-th frames respectively, H t represents the feature map of I t , F t→t+1 is the optical flow vector field obtained by optical flow estimation, representing the predicted motion vector of each pixel point in I t moving to I t+1 ; Warp(H t , F t→t+1 ) represents the operation of aligning the feature map H t based on the optical flow vector field F t→t+1 ; H′ t represents the aligned H t ;
[0030] The expression of the time-scale fusion is:
[0031]
[0032] where H fuse represents the fused feature, and α i represents the i-th fusion weight coefficient; H′ t-i represents the aligned feature map of the (t - i)-th frame.
[0033] Preferably, step S2-4 specifically includes:
[0034] Step S2-4-1: Establish the loss function of the 3D track line extraction network, and the expression is:
[0035] L total = L cls + L reg + αL match + βL feature_consistency
[0036] where L cls represents the classification loss, L reg represents the regression loss, L match represents the left and right track line matching loss, L feature_consistency represents the temporal consistency loss, and α and β are hyperparameters; The expression of L feature_consistency is:
[0037]
[0038] where H′ t+1 represents the aligned feature map of the (t + 1)-th frame, Denotes the square of the L2 norm;
[0039] Step S2-4-2: Using the samples from the track line dataset as inputs to the 3D track line extraction network, along with the prediction results, the true value labels of the samples, and the loss function, train the 3D track line extraction network with the objective of minimizing the loss function, and perform cross-validation on the prediction results of the 3D track line extraction network using the track line data test set, finally obtaining the trained 3D track line extraction network.
[0040] Preferably, step S3 specifically includes:
[0041] Step S3-1: Input the data collected in real time by the camera into the trained 3D track line extraction network to obtain the 3D spatial coordinates and categories of multiple points on the track line;
[0042] Step S3-2: Based on the 3D spatial coordinates and categories of multiple points on the track line and the track width constraint, pair the points of the left and right track lines to obtain multiple matching point pairs;
[0043] Step S3-3: Calculate the mean of the spatial coordinates of each matching point pair among the multiple matching point pairs to obtain a sequence of central position points as the 3D track center line.
[0044] Preferably, step S3-2 specifically includes:
[0045] Determine the forward direction of the train operation as the y-axis direction. For each point on the left track line, calculate its distance in the y-axis direction from all points on the right track line, and select the point on the right track line that satisfies the track width constraint and has the minimum y-axis distance for pairing. The expression is:
[0046]
[0047] where Find j represents the index j of the searched point on the right track line, argmin j (·) means the index j minimizes the condition, |·| represents the absolute value, and are the y-axis coordinates of the left and right track points, and are the x-axis coordinates of the left and right track points, and are the coordinates of the left and right track points in the z-axis direction, W is the preset track width, and ∈ is the allowable error range.
[0048] Preferably, step S4 specifically includes:
[0049] Step S4-1: For each center point on the 3D track center line, determine the pose of each center point based on the three-dimensional coordinates and tangent vectors of each center point;
[0050] Step S4-2: Determine the bounding box to be transformed corresponding to each center point. Based on the pose of each center point, transform the bounding box to be transformed corresponding to each center point to the position of each center point through rotation and translation to form a sequence of bounding sections, ensuring that each bounding section in the sequence of bounding sections is perpendicular to the track center line;
[0051] The step of determining the bounding box to be transformed corresponding to each center point specifically includes:
[0052] (1) If the center point belongs to the straight track part, determine the SG bounding box as the standard bounding box corresponding to the center point; if the center point belongs to the platform area, determine the KE bounding box as the standard bounding box corresponding to the center point;
[0053] (2) For the standard bounding box corresponding to each center point, as the detection distance increases, shrink it proportionally with the geometric center of the standard bounding box as the shrinking center to finally obtain the bounding box to be transformed corresponding to each center point;
[0054] Step S4-3: Determine the space enclosed by the sequence of bounding sections as the bounding space.
[0055] Preferably, step S5 specifically includes:
[0056] Step S5-1: Use a convex decomposition algorithm to decompose the bounding space into multiple convex polyhedra;
[0057] Step S5-2: Determine whether the point cloud collected in real time by the lidar is inside the multiple convex polyhedra, and remove the point cloud that is not inside the multiple convex polyhedra to obtain the intrusion point cloud inside the bounding space;
[0058] Step S5-3: Cluster the intrusion point cloud, screen the obtained clustering clusters, and use the fitted box of the qualified clustering clusters obtained by screening as the spatial position of the foreign object.
[0059] Preferably, the step of screening the obtained clustering clusters specifically includes:
[0060] Use the length-width-height ratio as the judgment index for the shape characteristics of the clustering clusters to screen the clustering clusters; the expression of the length-width-height ratio is:
[0061]
[0062] where, Lengthto-Width-to-HeightRatio represents the length-width-height ratio, λ 1 、λ 2and λ 3 are the explained variance ratios of the first, second, and third principal components, respectively.
[0063] Compared with the prior art, the present invention has at least the following beneficial effects:
[0064] (1) The present invention constructs a 3D track line extraction network. By designing the structure of the temporal fusion module and introducing the left and right track line matching loss, the accuracy and distance of 3D track line detection in the prior art are significantly improved. The detection accuracy of the track line is improved, and the robustness of the system in a complex environment is enhanced, making the track line extraction more accurate and reliable.
[0065] (2) The total loss function provided by the present invention comprehensively considers the temporal consistency loss, the left and right track line matching loss, the classification loss, and the regression loss, further improving the detection effect of the 3D track line extraction network. By ensuring the smoothness and consistency of the track line extraction result, this method effectively reduces the error in the detection process and ensures the smoothness and consistency of the track line extraction result.
[0066] (3) The present invention provides an algorithm for detecting spatial foreign object intrusion in urban rail trains. Using the long-distance three-dimensional track line extraction technology, the limited space of urban rail trains is accurately established. It can detect the intrusion of any type of obstacle within the limit. Description of the Drawings
[0067] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention.
[0068] Figure 1 It is a schematic diagram of the video data scene of the running environment of the urban rail train provided by the present invention.
[0069] Figure 2 It is a schematic diagram of the anchored 3D track line extraction network provided by the present invention.
[0070] Figure 3 It is a schematic diagram of the limit contour and its scaling provided by the present invention.
[0071] Figure 4 It is a schematic diagram of the construction of the track limit space model provided by the present invention.
[0072] Figure 5 It is a schematic diagram of the foreign object intrusion detection provided by the present invention.
[0073] Figure 6 It is a flowchart of the method for detecting spatial foreign object intrusion in urban rail trains based on anchored track line detection provided by the present invention. Detailed Embodiments
[0074] To better understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other. In addition, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0075] The urban rail train space foreign object intrusion perception method based on anchored track line detection provided by the present invention uses the anchored 3D track line extraction technology to generate the track line extraction result, constructs an accurate urban rail train clearance space based on the track center line, obtains the lidar within the intrusion clearance, and realizes the detection of any type of obstacle through clustering, thereby improving the accuracy and robustness of the detection.
[0076] To illustrate the effectiveness of the method proposed by the present invention, the above technical solution of the present invention will be described in detail through a specific embodiment as follows. Figure 6 As shown, a method for urban rail train space foreign object intrusion perception based on anchored track line detection is disclosed, and the specific implementation steps are as follows:
[0077] Step S1: Install a camera and a lidar on the urban rail train, collect environmental data using the camera and the lidar, and the environmental data has continuous time stamps; label the track space position information for the environmental data to form a track line data set.
[0078] In the present invention, a high-resolution camera and a high-precision lidar are installed inside the carriage of the urban rail train, and the high-resolution camera and the high-precision lidar can cover the entire environment of the forward space of the train. Select a variety of road environments including different urban rail train operation scenarios and different track line forms for dynamic data collection. The urban rail train operation scenarios may include outdoor and in-tunnel scenarios, etc., and the track line forms may include straight tracks, curved tracks, turnouts, and scissors crossings, etc., as Figure 1 shown.
[0079] For each of the above-mentioned various road environments, the camera and the lidar can respectively record image sequences and point cloud sequences within multiple time periods. For the image sequences and point cloud sequences within a period of time, among them, the image sequences have time stamps arranged in chronological order and corresponding images, and the point cloud sequences have time stamps arranged in chronological order and corresponding point clouds. The image sequences and point cloud sequences within the multiple time periods are used as the environmental data.
[0080] The present invention first performs time synchronization processing on the environmental data and then performs track space position information annotation.
[0081] The specific steps for synchronizing the image sequence and the point cloud sequence in the environmental data over a period of time are as follows: perform matching selection according to the timestamps in the image sequence and the point cloud sequence, and select the images and point clouds with the same timestamp as the environmental data for time synchronization. In some embodiments, 10,000 frames of data can be selected, including 5,000 images and point cloud data of the tunnel interior scene and 5,000 images and point cloud data of the outdoor scene.
[0082] For the environmentally synchronized data, in the order of timestamps, obtain the 3D spatial coordinate information of the track line corresponding to each frame of image and lidar point cloud through automatic annotation technology. The 3D spatial coordinate information records the 3D spatial coordinates of the points on the railway track and the categories to which the points on the railway track belong. The categories include the left railway track and the right railway track. The 3D spatial coordinate information of the track line serves as the ground truth label for the location of the track line. Use each frame of image, each frame of point cloud, and the corresponding ground truth label together as samples of the dataset. Randomly select a preset proportion of samples from the track line dataset to form a set, which is used as the track line data test set to test the perception effect of the subsequent model in the actual train operation environment.
[0083] In some embodiments, 10,000 frames of data and the corresponding 3D spatial coordinate information of the track line can be selected as samples, and 5,000 samples in different scenarios and different track line morphology environments are randomly selected from the 10,000 samples to form a set, which is used as the track line data test set.
[0084] In addition, select 867 samples from the track line dataset and complete the obstacle annotation for these samples. The content of the obstacle annotation includes the obstacle category, the central position of the 3D annotation box, the size of the 3D annotation box, and the yaw angle and roll angle of the 3D annotation box. The obstacle annotation is used for the accuracy evaluation of the subsequent foreign object intrusion perception method.
[0085] Step S2: Establish a 3D track line extraction network, which includes a feature extraction network, a temporal fusion module, and a detection head; train the 3D track line extraction network based on the track line dataset to obtain a trained 3D track line extraction network.
[0086] As Figure 2 shown, the 3D track line extraction network includes a feature extraction network, a temporal fusion module, and a detection head. The specific network structure is as follows:
[0087] (1) Feature extraction network
[0088] The feature extraction network of the present invention can adopt a network with an encoder framework, such as the ResNet-18 network. The feature extraction network can achieve multi-scale feature extraction. By modifying the downsampling step size and introducing dilated convolutions, high-resolution feature maps are maintained and the receptive field is enhanced.
[0089] The image frames {I 1 , I 2 , …, I T} arranged in chronological order are input into the feature extraction network to obtain the multi-scale features {H 1 , H 2 , …, H T} of each image frame. Among them, I T represents the input image of the T-th frame, and H T represents the feature map of I T .
[0090] (2) Temporal fusion module
[0091] To improve the network's adaptability to dynamic environments and the stability of track line extraction, the present invention provides a temporal fusion module for integrating the temporal features between consecutive frames. The temporal fusion module processes the feature sequence in the time dimension to perform temporal modeling on the features of consecutive frames to obtain temporal features. The self-attention mechanism is applied to dynamically align the features at different time steps.
[0092] Specifically, for the multi-scale features {H 1 , H 2 , …, H T}, the feature sequence is processed along the time dimension through a temporal network convolution, and the updated feature map H t of the current time step is obtained from the feature maps of each layer to capture long-range temporal dependencies. The expression is:
[0093] H t = ReLU(W * H t-1 + b)
[0094] where H t represents the feature map of the input image I t of the t-th frame, W and b are the convolution kernel and bias respectively, and * represents a one-dimensional convolution operation.
[0095] The present invention uses optical flow estimation to dynamically align the features at different time steps. The feature offset caused by vehicle movement or environmental changes is reduced to ensure the consistency of the features of consecutive frames. The expression is:
[0096] F t→t+1 = FlowNet(I t , I t+1 )
[0097] H' t = Warp(H t , F t→t+1 )
[0098] where F t→t+1 is the optical flow vector field obtained by optical flow estimation, representing the predicted motion vector of each pixel point in I t moving to I t+1 ; Warp(H t , F t→t+1 ) represents the feature map alignment operation on Ht based on the optical flow vector field Ft → t +1 ; H' t represents the aligned H t .
[0099] Fuse the multi-scale features of each aligned image frame in the time scale. The expression is:
[0100]
[0101] where H fuse represents the fused feature, α i represents the i-th fusion weight coefficient, and α i is learnable; H' t-i represents the aligned feature map of the (t - i)-th frame.
[0102] (3) Detection head
[0103] Input the fused feature obtained by the temporal fusion module into the detection head to obtain the 3D spatial coordinate information of the predicted track line. The 3D spatial coordinate information of the track line includes the category of the track line and the 3D coordinates of the points on the track line.
[0104] In some embodiments, the detection head of the present invention is divided into a classification head and a regression head. The classification head is used to determine the category of the extracted fused feature to obtain the category of the points on the track line. The categories of the track line include the left track and the right track. The regression head generates the 3D coordinates of the points on the track line through regression operations.
[0105] After establishing the 3D track line extraction network, train the 3D track line extraction network based on the track line dataset to obtain the trained 3D track line extraction network.
[0106] The present invention establishes a loss function for the 3D track line extraction network. The expression is:
[0107] L total = L cls + L reg + αL match + βL feature_consistency
[0108] Among them, L cls represents the classification loss, and L reg represents the regression loss, and L match represents the left - right track line matching loss, and L feature_consistency represents the temporal consistency loss. α and β are hyperparameters used to balance the weights of different loss terms. The expression of L feature_consistency is as follows:
[0109]
[0110] Among them, represents the square of the L2 norm.
[0111] The present invention inputs an image sequence formed by images in multiple samples in the track line dataset into the 3D track line extraction network to obtain the prediction results of the category of the track line and the 3D coordinates of the points on the track line.
[0112] For the training of the 3D track line extraction network, based on the prediction results, the corresponding ground - truth labels in the track line dataset, and the loss function, error back - propagation algorithms such as the gradient descent method can be used to iteratively adjust the network weights with the goal of minimizing the loss function of the 3D track line extraction network to obtain a trained 3D track line extraction network.
[0113] In some embodiments, a track line data test set can be used to perform cross - validation on the prediction results of the 3D track line extraction network to reduce the prediction error and ensure that the perception effect of the model in the real operating environment is more reliable.
[0114] Step S3: Input the data collected by the camera in real time into the trained 3D track line extraction network to obtain a 3D track line; based on the track width constraint, obtain the 3D track center line from the 3D track line.
[0115] In this step, the trained 3D track line extraction network is used for online processing. The data collected by the camera in real time is input into the trained 3D track line extraction network to obtain a 3D track line. The 3D track line includes the 3D spatial coordinate information of the track line, including the category of the track line and the 3D coordinates of the points on the track line.
[0116] For the category of the track line and the 3D coordinates of the points on the track line, based on the track width constraint, pair the coordinate points of the left and right track lines.
[0117] Determine the forward direction of the train running direction as the y - axis direction. For each left track line point, calculate its distance in the y - axis direction from all right track line points. Select the right track line point that satisfies the track width constraint and has the minimum y - axis distance for pairing. Ensure that each pair of paired points is located at the same track position.
[0118] The specific paired expression is:
[0119]
[0120] Among them, Find j represents the index j of the right track line point obtained by searching, and argmin j (·) means that the index j minimizes the condition, |·| represents the absolute value, and are the y-axis coordinates of the left and right track points, and are the x-axis coordinates of the left and right track points, and are the coordinates of the left and right track points in the z-axis direction. W is the preset track width, and ∈ is the allowable error range.
[0121] Through the above expressions, the matching points on the right track line are selected to minimize the difference in the y direction, and at the same time ensure that the difference between the distances in the x and z directions and the preset track width W is within the allowable error range ∈.
[0122] Pair all the coordinate points of the left and right track lines, calculate the mean value of the positions of all the matching points on the left and right track lines, and obtain a series of coordinate points of the central positions as the 3D track center line.
[0123] Step S4: Establish a bounding space based on a plurality of preset bounding box sizes and the 3D track center line. The bounding space is the space where foreign object intrusion needs to be detected.
[0124] In this step, a bounding space for detecting foreign object intrusion is established based on the 3D track center line. Specifically, according to the regulations of the bounding range of different track areas, a structure gauge (SG) bounding box close to the rail plane is used for the straight track part, and a slightly higher kinematic envelope (KE) bounding box is used for the platform area.
[0125] In order to cope with the deviation of high-precision positioning at different distances, the size of the bounding box is reduced proportionally with the geometric center of the bounding box as the shrinking center as the detection distance increases. As Figure 3 、 Figure 4 shown, the detection range from 0 to 300 meters is divided into several detection sections, and the cross-sectional profile of each section is reduced proportionally according to the distance, effectively reducing the false alarm rate.
[0126] For each center point on the 3D track center line, the attitude of each center point is determined by the three-dimensional coordinates and tangent vectors of each center point, and the bounding box to be transformed corresponding to each center point is determined. Based on the attitude of each center point, the bounding box to be transformed corresponding to each center point is transformed to the position of each center point through rotation and translation to form a sequence of bounding cross-sections, ensuring that each bounding cross-section in the sequence of bounding cross-sections is perpendicular to the track center line. Finally, the real-world non-convex urban rail train bounding space range corresponding to the current image is obtained from all the bounding cross-sections generated by the extraction result of one track center line, as Figure 4 shown.
[0127] The non-convex bounding space is represented as a polyhedron, and its geometric structure of the bounding space is described by its vertices, edges and faces. In some embodiments, the non-convex bounding space is represented as a Polyhedron_3 data structure in the Computational Geometry Algorithms Library (CGAL), and Polyhedron_3 is used to represent a polyhedron in three-dimensional space.
[0128] Step S5: Based on the data collected by the lidar in real time and the bounding space, obtain the intrusion perception data of the bounding space, and obtain the spatial position of the foreign object based on the intrusion perception data of the bounding space.
[0129] After obtaining the non-convex bounding space, apply a convex decomposition algorithm to decompose the non-convex polyhedron into multiple convex polyhedra by processing non-vertical and vertical reflection edges, ensuring the spatial integrity and computational efficiency of the decomposition result. The decomposed convex polyhedra are convenient for subsequent judgment of the relationship between points and the bounding space.
[0130] The data collected by the lidar in real time is time-stamped with the data collected by the camera in real time. For the point cloud data collected by the lidar in real time, an axis-aligned bounding box tree structure is used to organize and optimize the point cloud data, and the point cloud outside the bounding space is quickly removed, and the intrusion point cloud within the bounding space is extracted.
[0131] For the intrusion point cloud within the bounding space, an effective clustering of the intrusion point cloud is performed using the Euclidean distance clustering method accelerated by KD-Tree, ensuring that different types of obstacle samples such as trains, boxes, aerial obstacles and pedestrians are effectively grouped to form independent clustering clusters.
[0132] The principal component analysis method is introduced to evaluate and screen each clustering cluster, and the length-width-height ratio is designed as an index for judging the shape characteristics of the clustering cluster. The specific calculation formula is as follows:[[]]
[0133]
[0134] where, Lengthto-Width-to-HeightRatio represents the length-width-height ratio, λ 1 、λ2 and λ 3 are the explained variance ratios of the first, second, and third principal components, respectively.
[0135] By performing threshold judgment on the aspect ratio, etc., the clustering clusters that do not meet the requirements are screened out, and the fitting box of the qualified clustering clusters is generated as the spatial position of the foreign object. The whole process is as Figure 5 shown.
[0136] It should be understood that the foregoing only illustrates some embodiments, and changes, modifications, additions, and / or variations can be made without departing from the scope and essence of the disclosed embodiments. The embodiments are illustrative rather than restrictive. In addition, the described embodiments relate to the currently considered most practical and preferred embodiments, and it should be understood that the embodiments should not be limited to the disclosed embodiments. On the contrary, it is intended to cover different modifications and equivalent arrangements included in the essence and scope of the embodiments. In addition, the various embodiments described above can be applied in combination with other embodiments. For example, aspects of one embodiment can be combined with aspects of another embodiment to achieve yet another embodiment. Additionally, the independent features or components of any given component can constitute additional embodiments.
[0137] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for sensing foreign body intrusion in urban rail train space based on anchored track line detection, characterized in that: The following steps are involved: Step S1, arranging a camera and a laser radar on the urban rail train, and using the camera and the laser radar to collect environmental data, wherein the environmental data has a continuous time stamp; Marking the environmental data with track spatial position information to form a track line data set; Step S2, establishing a 3D trajectory line extraction network, wherein the 3D trajectory line extraction network includes a feature extraction network, a time series fusion module and a detection head; Training a 3D trajectory line extraction network based on the trajectory line data set to obtain a trained 3D trajectory line extraction network; Step S3, inputting the data collected by the camera in real time into the trained 3D track line extraction network to obtain a 3D track line; based on the track width constraint, obtaining a 3D track center line from the 3D track line; Step S4, establishing a bounded space based on a plurality of preset bounding frame sizes and the 3D track centerline, wherein the bounded space is a space where foreign body intrusion needs to be detected; Step S5: based on the data collected in real time by the laser radar and the bounded space, obtain bounded space intrusion perception data, and obtain the spatial position of the foreign object based on the bounded space intrusion perception data.
2. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 1 is characterized in that: The step S1 specifically includes: Step S1-1, installing a camera and a laser radar in a carriage of an urban rail train, and in various urban rail train operation scenarios and various track line forms, the camera and the laser radar respectively record image sequences and point cloud sequences within multiple time periods; the image sequences and point cloud sequences have continuous timestamps; Step S1-2, time-synchronizing the image sequence and the point cloud sequence, including: matching and selecting according to the timestamps in the image sequence and the point cloud sequence, and selecting the images and point clouds with the same timestamps as the time-synchronized environment data; Step S1-3: For the time-synchronized environmental data, the categories and 3D spatial coordinates of multiple points on the track line in each frame of the image sequence and the point cloud sequence are obtained by automatic annotation technology in the order of timestamps as true value labels; Step S1-4: Take each frame of image, each frame of point cloud and the corresponding true value label as samples of the data set to form a track line data set, and randomly select a preset number of samples from the samples of the track line data set to form a track line data test set.
3. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 2 is characterized in that: The step S2 specifically includes: Step S2-1, establishing a feature extraction network based on the ResNet-18 network, wherein the feature extraction network inputs a plurality of image frames arranged in chronological order to obtain multi-scale features of the plurality of image frames; Step S2-2, establishing a temporal fusion module, wherein the temporal fusion module performs temporal network convolution processing, dynamic alignment of features at different time steps, and time scale fusion on the multi-scale features of the multiple image frames, and finally obtains fused features; Step S2-3, establishing a detection head, wherein the detection head includes a classification head and a regression head, wherein the classification head and the regression head obtain 3D spatial coordinates and category estimation values of multiple points on the track line based on the fusion features; Step S2-4, determining the loss function of the 3D trajectory extraction network, and based on the trajectory data set and the trajectory data test set, training the 3D trajectory extraction network with minimizing the loss function as the optimization goal to obtain a trained 3D trajectory extraction network.
4. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 3 is characterized in that: In step S2-2, the expression of the temporal network convolution processing is: A t =ReLU(W*H t-1 +b) Among them, H t Represents the t-th frame input image I t is the feature map of , W and b are the convolution kernel and bias respectively, * represents a one-dimensional convolution operation; The expression of dynamic alignment of features at different time steps is: F t→t+1 =FlowNet(I t ,I t+1 ) H′ t =Warp(H t ,F t→t+1 ) Among them, I t ,I t+1 are the input images of the tth and t+1th frames respectively, H t Indicates I t The feature map, F t→t+1 is the optical flow vector field obtained by optical flow estimation, indicating I t Each pixel moves to I t+1 Predicted motion vector; Warp (H t ,F t→t+1 ) indicates H t Based on the optical flow vector field F t→t+1 Perform feature map alignment operation; H′ t Indicates the aligned H t ; The expression of the time scale fusion is: Among them, H fuse represents the fusion feature, α i represents the i-th fusion weight coefficient; H′ t-i Represents the aligned feature map of the ti-th frame.
5. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 4 is characterized in that: The step S2-4 specifically includes: Step S2-4-1: Establish the loss function of the 3D trajectory extraction network, expressed as: L total =L cls +L reg +αL match +βL feature_consistency Among them, L cls represents the classification loss, L reg represents the regression loss, L match represents the left and right track line matching loss, L feature_consistency represents the temporal consistency loss, α and β are hyperparameters; L feature_consistency The expression is: Among them, H′ t+1 represents the aligned feature map of the t+1th frame, represents the square of L2 norm; Step S2-4-2, based on the sample of the track line data set, the prediction result, the true value label of the sample and the loss function are input into the 3D track line extraction network, the 3D track line extraction network is trained with minimizing the loss function as the optimization goal, and the prediction result of the 3D track line extraction network is cross-validated using the track line data test set, and finally a trained 3D track line extraction network is obtained.
6. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 5 is characterized in that: The step S3 specifically includes: Step S3-1, inputting the data collected by the camera in real time into the trained 3D track line extraction network to obtain the 3D spatial coordinates and categories of multiple points on the track line; Step S3-2, based on the 3D spatial coordinates and categories of the multiple points on the track line and the track width constraint, pair the points of the left and right track lines to obtain multiple matching point pairs; Step S3-3: Calculate the mean of the spatial coordinates of each of the multiple matching point pairs to obtain a center position point sequence as the 3D track center line.
7. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 6 is characterized in that: The step S3-2 specifically includes: The forward direction of the train running direction is determined as the y-axis direction. For each left track line point, its distance from all right track line points in the y-axis direction is calculated, and the right track line points that meet the track width constraint and have the smallest y-axis distance are selected for pairing. The expression is: Among them, Find j means searching for the index j of the right track line point, argmin j (·) indicates that the index j minimizes the condition, |·| indicates the absolute value, and are the y-axis coordinates of the left and right track points, and are the x-axis coordinates of the left and right track points, and are the coordinates of the left track point and the right track point in the z-axis direction, W is the preset track width, and ∈ is the allowable error range.
8. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 7 is characterized in that: The step S4 specifically includes: Step S4-1, for each center point on the 3D track center line, determine the posture of each center point by the three-dimensional coordinates and tangent vector of each center point; Step S4-2, determining the bounding box to be transformed corresponding to each center point, and based on the posture of each center point, transforming the bounding box to be transformed corresponding to each center point to the position of each center point by rotation and translation to form a bounding section sequence, ensuring that each bounding section in the bounding section sequence is perpendicular to the track centerline; The step of determining the bounding box to be transformed corresponding to each center point specifically includes: (1) If the center point belongs to the straight section, the SG bounding box is determined as the standard bounding box corresponding to the center point; if the center point belongs to the platform area, the KE bounding box is determined as the standard bounding box corresponding to the center point; (2) For the standard bounding box corresponding to each center point, as the detection distance increases, the standard bounding box is reduced proportionally with the geometric center of the standard bounding box as the reduction center, and finally the bounding box to be transformed corresponding to each center point is obtained; Step S4-3: determine the space enclosed by the bounding section sequence as the bounding space.
9. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 8 is characterized in that: The step S5 specifically includes: Step S5-1, using a convex decomposition algorithm to decompose the bounded space into multiple convex polyhedrons; Step S5-2, determining whether the point cloud collected by the laser radar in real time is inside the multiple convex polyhedrons, eliminating the point cloud that is not inside the multiple convex polyhedrons, and obtaining the intrusion point cloud within the bounded space; Step S5-3: clustering the intrusion point cloud, screening the obtained clusters, and using the fitting boxes of the screened qualified clusters as the spatial positions of the foreign matter.
10. The method for sensing foreign body intrusion in urban rail train space based on anchored track line detection according to claim 9 is characterized in that: The step of screening the obtained clusters specifically includes: The aspect ratio is used as a cluster shape feature judgment index to screen the clusters; the expression of the aspect ratio is: Among them, Length-Width-to-Heigh represents the length-to-width-to-height ratio, and λ1, λ2, and λ3 are the explained variance ratios of the first, second, and third principal components, respectively.
Citation Information
Patent Citations
Foreign matter intrusion detection method and device based on convolutional neural network, and storage medium
CN117237613A