Track line detection method based on semantic segmentation
Through the track line detection method based on semantic segmentation, the problem of poor robust performance and inability to adapt to high-speed operation in the prior art is solved, and high-precision and high-real-time track line detection is achieved.
Patent Information
- Application Number
- CN202311548064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art has poor robust performance in track line detection, unable to effectively detect complex backgrounds, shadows and uneven lighting environments, and cannot adapt to high-speed or ultra-high-speed trains.
The track line detection method based on semantic segmentation is adopted, and the track line images are collected through the laser camera, and the preprocessing module is used for preprocessing. The backbone network extracts the track line features, the binary segmentation module performs binary segmentation, the track embedding module divides the track area, and the clustering module is used to cluster pixels belonging to the same track line, and finally the complete track line is fitted and outputted through the least squares post-processing module.
It improves the robust performance, detection accuracy and real-time performance of track line detection, and can accurately identify track lines in complex environments, adapt to high-speed trains, and meet the requirements of high accuracy and high real-time.
Smart Images

Figure CN120020895A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent rail transit, and particularly to a track line detection method based on semantic segmentation. Background Art
[0002] With the continuous opening of high-speed rail transit lines such as railways, maglevs, and pipelines, the safety of rail transit has attracted more and more attention. Unsafe factors such as track surface damage detection and track foreign object intrusion seriously affect the high-speed or ultra-high-speed operation of trains. It will not only cause chaos in the operation order of rail transit, but also may cause huge economic losses. Therefore, it is of great significance to timely detect problems such as track surface damage and foreign object intrusion in the track for the safe operation of trains.
[0003] With the rapid development of artificial intelligence and computer vision technologies, the intelligent inspection of railways has gradually been realized from concept, planning to technical implementation. In the intelligent inspection of rail transit, the detection of track lines is a relatively crucial step. It can not only provide a detection area for the detection function of track surface damage, but also provide a dangerous area for foreign object intrusion into the track. Therefore, the track line detection technology is the basis of each safety detection function, and this technology provides safety guarantees for functions such as track surface damage detection and foreign object intrusion detection in the track. Traditional detection methods have poor detection accuracy and robustness for environments such as complex backgrounds, shadows, and lighting, and cannot adapt to trains running at high speeds or ultra-high speeds.
[0004] Currently, the methods for track line detection mainly include edge detection, template matching, etc. The process of the edge detection method is based on the Hough transform of the gradient image to estimate the main direction of the track, and then the track is detected and located by combining the decision criteria. The template matching method is mainly based on the horizontal and vertical geometric features of the track, selects the key points of the track, constructs a track cross-section model for template matching to determine the track position, and finally locates the track position according to geometric information such as the longitudinal continuity of the track and the track gauge between single tracks.
[0005] However, the recognition method based on edge detection of track lines has poor robustness. Edge detection aims at the geometric features of the edges of objects. When the background is complex, the lighting is uneven, or on rainy days, edge detection will generate many useless edges, resulting in the inability to effectively detect the track line. The track line detection method based on template matching has poor generalization ability. The template matching method requires manual extraction of image features and the establishment of local or global linear models such as straight lines, parabolas, and curves. It has a certain pertinence to different track lines, and using the same model to detect different track lines will result in low detection accuracy. Summary of the Invention
[0006] The present invention provides a track line detection method based on semantic segmentation, which can solve the technical problems in the prior art.
[0007] The present invention provides a track line detection method based on semantic segmentation, wherein the method includes:
[0008] Collecting track line images by using a laser camera, and preprocessing the collected track line images by using a preprocessing module to obtain preprocessed images;
[0009] Extracting track line features from the preprocessed images through a backbone network;
[0010] Performing binary segmentation on the extracted track line features by using a binary segmentation module, dividing the track regions of the extracted track line features by using a track embedding module, and clustering all pixels belonging to the same track line by using the segmentation results and the division results;
[0011] Using a least squares post-processing module to fit the clustered pixels and output a complete visualized track line.
[0012] Preferably, preprocessing the collected track line images to obtain preprocessed images includes:
[0013] Performing camera calibration on the collected track line images;
[0014] Performing inverse perspective transformation on the calibrated track line images to obtain a top view.
[0015] Preferably, performing camera calibration on the collected track line images includes:
[0016] Converting the collected track line images in the order of world coordinate system, camera coordinate system, image coordinate system, and image pixel coordinate system.
[0017] Preferably, performing binary segmentation on the extracted track line features includes:
[0018] Performing binary segmentation on the extracted track line features by using a weighted cross-entropy loss function, a variance loss function, and a distance loss function.
[0019] Preferably, extracting track line features from the preprocessed images through a ResNet34 backbone network, the ResNet34 backbone network includes a convolutional attention module, and the extracted track line features include local features and global features.
[0020] The present invention also provides a track line detection system based on semantic segmentation, wherein the system includes:
[0021] A laser camera for collecting track line images;
[0022] A preprocessing module for preprocessing the collected track line images to obtain preprocessed images;
[0023] A backbone network for extracting track line features based on the preprocessed images;
[0024] A binary segmentation module for binarizing and segmenting the extracted track line features;
[0025] A track embedding module for dividing the extracted track line features into track regions;
[0026] A clustering module for clustering all pixels belonging to the same track line using the segmentation result and the division result;
[0027] A least squares post-processing module for fitting the clustered pixels and outputting a complete and visualized track line.
[0028] Preferably, preprocessing the collected track line images by the preprocessing module to obtain preprocessed images includes:
[0029] Performing camera calibration on the collected track line images;
[0030] Performing inverse perspective transformation on the calibrated track line images to obtain a top view.
[0031] Preferably, performing camera calibration on the collected track line images includes:
[0032] Converting the collected track line images in the order of world coordinate system, camera coordinate system, image coordinate system, and image pixel coordinate system.
[0033] Preferably, binarizing and segmenting the extracted track line features by the binary segmentation module includes:
[0034] Using the binary segmentation module to perform binarizing and segmenting on the extracted track line features by using a weighted cross-entropy loss function, a variance loss function, and a distance loss function.
[0035] Preferably, the backbone network is a ResNet34 backbone network, the ResNet34 backbone network includes a convolutional attention module, and the extracted track line features include local features and global features.
[0036] Through the above technical solutions, the track line category and the background category can be distinguished by semantic segmentation, and each track line category can be distinguished. On the basis of ensuring the accurate fitting of the track line, the timeliness of recognition is ensured. Compared with the traditional edge detection method, its detection accuracy is better and the timeliness is better, meeting the high-precision and high-real-time requirements for actual track line detection. That is, the present invention can improve the robustness, detection accuracy, and real-time performance of track line detection. Description of the Drawings
[0037] The accompanying drawings included are used to provide a further understanding of the embodiments of the present invention, which form a part of the specification, for illustrating the embodiments of the present invention, and for explaining the principles of the present invention together with the written description. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It shows a flowchart of a track line detection method based on semantic segmentation according to an embodiment of the present invention;
[0039] Figure 2 It shows a schematic diagram of the camera calibration coordinate system conversion process according to an embodiment of the present invention;
[0040] Figure 3 It shows a schematic diagram of a track line detection system based on semantic segmentation according to an embodiment of the present invention. Detailed implementation manners
[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. The description of at least one exemplary embodiment below is actually only illustrative and in no way limits the present invention and its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0042] It should be noted that the terms used here are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that for ease of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0044] Figure 1 The flowchart of a track line detection method based on semantic segmentation according to an embodiment of the present invention is shown.
[0045] As Figure 1 shown, an embodiment of the present invention provides a track line detection method based on semantic segmentation, wherein the method includes:
[0046] Collect a track line image using a laser camera, and preprocess the collected track line image using a preprocessing module to obtain a preprocessed image;
[0047] Extract track line features from the preprocessed image through a backbone network;
[0048] Perform binary segmentation on the extracted track line features using a binary segmentation module, perform track area division on the extracted track line features using a track embedding module, and cluster all pixels belonging to the same track line using the segmentation result and the division result;
[0049] Use a least squares post-processing module to fit the clustered pixels and output a complete visual track line.
[0050] Through the above technical solution, the track line category and the background category can be distinguished through semantic segmentation, and each track line category can be distinguished. On the basis of ensuring the accurate fitting of the track line, the timeliness of recognition is ensured. Compared with the traditional edge detection method, its detection accuracy is better and the timeliness is better, meeting the high-precision and high-real-time requirements for actual track line detection. That is, the present invention can improve the robustness, detection accuracy, and real-time performance of track line detection.
[0051] In the present invention, a dual-branch structure of a binary segmentation branch and an orbit embedding branch is adopted. Among them, the purpose of the binary segmentation branch is to extract the pixel points of the orbit line from the background area, and the purpose of the orbit embedding branch is to use the correlation between pixels to divide the orbit area. Then, all the pixels belonging to the same orbit line are clustered by using the segmentation map obtained by the binary segmentation branch and the result of the orbit embedding branch to form a complete orbit line within the field of view.
[0052] According to an embodiment of the present invention, preprocessing the collected orbit line image to obtain a preprocessed image includes:
[0053] Performing camera calibration on the collected orbit line image;
[0054] Performing inverse perspective transformation on the calibrated orbit line image to obtain a top view.
[0055] Among them, the laser camera can be a laser light source linear array camera. Using a laser light source linear array camera to collect orbit line images can effectively suppress the phenomenon of uneven illumination of the image information collected due to the light source. At the same time, the linear array camera has a faster scanning speed, which is relatively more conducive to high-speed detection tasks. The orbit lines in the image appear larger near and smaller far away, converging to a point at the far end. Therefore, inverse perspective transformation can be performed on the collected orbit line image. Inverse perspective transformation can convert the front view to the top view, and the orbit lines become parallel. By inverse perspective transformation, the recognition accuracy can be improved, the interference of irrelevant information can be reduced, and the imbalance between background information and orbit line information data can be suppressed.
[0056] According to an embodiment of the present invention, as Figure 2 shown, performing camera calibration on the collected orbit line image includes:
[0057] Converting the collected orbit line image in the order of world coordinate system, camera coordinate system, image coordinate system, and image pixel coordinate system.
[0058] Thus, the relationship between the camera pixel coordinate system and the spatial position can be established, and the physical information between three-dimensional space objects can be reflected from the limited two-dimensional plane according to the pinhole imaging model.
[0059] Specifically, the four coordinate systems involved in camera calibration are: pixel coordinate system (u, v), the origin coordinate is the pixel point in the upper left corner of the image, and its unit is pixel; image coordinate system O-xy, its origin is the central area of the image to be detected, and its unit is mm; camera coordinate system O c -x c y c z c , its origin is the camera optical center, z c is the camera optical axis, perpendicular to the image plane, and its unit is m; world coordinate system O w -xw y w z w The position is not fixed and can be established with any three-dimensional space point as the origin, and its unit is m. The conversion relationships of the four coordinate systems can be expressed as:
[0060] (1) World coordinate system to camera coordinate system
[0061]
[0062] (2) Camera coordinate system to image coordinate system
[0063]
[0064]
[0065] (3) Image coordinate system to pixel coordinate system
[0066]
[0067]
[0068] Among them, R represents the rotation matrix, f is the camera focal length, c x and c y are the coordinates of the origin of the image coordinate system in pixel coordinates, t represents the translation vector of the camera relative to the world coordinate system, dx represents the physical size of each pixel on the horizontal axis x, and dy represents the physical size of each pixel on the horizontal axis y.
[0069] According to an embodiment of the present invention, the binarization segmentation of the extracted track line features includes:
[0070] Using a weighted cross-entropy loss function, a variance loss function, and a distance loss function to perform binarization segmentation on the extracted track line features.
[0071] That is, in the binary segmentation branch: a weighted cross-entropy loss function can be used to eliminate the influence of class imbalance. A variance loss function can be used to minimize the distance between pixels belonging to the same track line; a distance loss function can be used to maximize the distance between pixels of different track lines. The weighted cross-entropy loss function, the variance loss function, and the distance loss function are respectively:
[0072]
[0073]
[0074]
[0075] The total loss function is: L = L b +L var+L dist .
[0076] Among them, L b represents the weighted cross-entropy loss function, and L var represents the variance loss function, and L dist represents the distance loss function, and w i represents the class balance weight, and y i represents the true class, and p i represents the proportion of class i, N represents the number of lane lines in the binary label, and N i represents the number of elements in the family, and μ i represents the average vector of the cluster centers.
[0077] According to an embodiment of the present invention, the track line features are extracted from the preprocessed image through the ResNet34 backbone network. The ResNet34 backbone network includes a convolutional block attention module (CBAM), and the extracted track line features include local features and global features.
[0078] Using ResNet34 as the backbone network can effectively avoid the degradation of the convolutional layer due to the increase in the number of layers. By adding a convolutional block attention module to the backbone network, the ability to extract target features can be improved. The channel attention in CBAM can calculate the importance of the features in each feature map, and the spatial attention can perform global pooling operations on each channel in the output features of the channel attention to obtain global information. The convolutional features can extract the local details of the track line, and the attention can extract the context feature information of the track line. By fusing the convolutional features and the attention features, both the local feature information of the track line can be retained and the global context feature information can be increased.
[0079] Figure 3 The schematic diagram of a track line detection system based on semantic segmentation according to an embodiment of the present invention is shown.
[0080] As Figure 3 shown, an embodiment of the present invention also provides a track line detection system based on semantic segmentation. Among them, the system includes:
[0081] A laser camera for collecting track line images;
[0082] A preprocessing module for preprocessing the collected track line images to obtain preprocessed images;
[0083] A backbone network for extracting track line features from the preprocessed images;
[0084] A binary segmentation module for binary segmentation of the extracted track line features;
[0085] An orbit embedding module for dividing the orbit region of the extracted orbit line features;
[0086] A clustering module for clustering all pixels belonging to the same orbit line by using the segmentation result and the division result;
[0087] A least squares post-processing module for fitting the clustered pixels and outputting a complete visualized orbit line.
[0088] Through the above technical solutions, the orbit line category and the background category can be distinguished by semantic segmentation, and each orbit line category can be distinguished. On the basis of ensuring the accurate fitting of the orbit line, the timeliness of recognition is ensured. Compared with the traditional edge detection method, its detection accuracy is better and the timeliness is better, meeting the high-precision and high-real-time requirements for actual orbit line detection. That is, the present invention can improve the robustness, detection accuracy and real-time performance of orbit line detection.
[0089] According to an embodiment of the present invention, preprocessing the acquired orbit line image by using a preprocessing module to obtain a preprocessed image includes:
[0090] Performing camera calibration on the acquired orbit line image;
[0091] Performing inverse perspective transformation on the calibrated orbit line image to obtain a top view.
[0092] For example, the preprocessing module includes a camera calibration unit and an inverse perspective transformation unit. The camera calibration unit is used to perform camera calibration on the acquired orbit line image, and the inverse perspective transformation unit is used to perform inverse perspective transformation on the calibrated orbit line image to obtain a top view.
[0093] According to an embodiment of the present invention, performing camera calibration on the acquired orbit line image includes:
[0094] Converting the acquired orbit line image in the order of world coordinate system, camera coordinate system, image coordinate system and image pixel coordinate system.
[0095] According to an embodiment of the present invention, using a binary segmentation module to perform binary segmentation on the extracted orbit line features includes:
[0096] Using the binary segmentation module to perform binary segmentation on the extracted orbit line features by using a weighted cross-entropy loss function, a variance loss function and a distance loss function.
[0097] According to an embodiment of the present invention, the backbone network is a ResNet34 backbone network. The ResNet34 backbone network includes a convolutional attention module, and the extracted orbit line features include local features and global features.
[0098] The following describes the method and system for detecting rail lines based on semantic segmentation according to the present invention in combination with examples.
[0099] Taking a certain rail inspection vehicle as an example, a laser supplementary light device and a laser imaging device are installed at the front center of the inspection vehicle detection platform to collect rail line images during the operation of the inspection vehicle. Among them, the resolution of the laser camera is 2K, and when identifying the rail line, the 2K resolution image needs to be downsampled to 512×288.
[0100] First, data augmentation is performed on the actually collected rail images during the operation of the rail inspection vehicle by methods such as flipping, cropping, spatial transformation, adding noise, scaling, translation, and jittering. Then, the rail line data is divided into a training set, a validation set, and a test set according to 7:2:1.
[0101] Secondly, a network model for rail line recognition is constructed by combining a preprocessing module, a backbone network, a binary segmentation module, a rail embedding module, a clustering module, and a postprocessing module. The backbone network of the network model uses ResNet34, and a CBAM attention mechanism module is added. Deep convolutional layers can extract local details of the rail line, and at the same time, the attention module can extract context feature information of the rail line. By fusing the local detail features and attention features of the rail line, both the local feature information of the rail line can be retained and the global context feature information can be increased. Then, the fused features are input into the binary segmentation branch and the rail line embedding branch respectively. The purpose of the binary segmentation branch is to extract the pixel points of the rail line from the background area, and the purpose of the rail embedding branch is to divide the rail area by using the correlation between pixels. Then, all pixels belonging to the same rail line are clustered by using the segmentation map obtained from the binary segmentation branch and the result of the rail embedding branch. Finally, the least squares method is used to fit all pixels belonging to the same rail line, and the position of the rail line is visually output.
[0102] During model training, the initial learning rate can be set to 0.0001, the minimum learning rate to 1e-6, the epoch to 100, the loss function of each round of training is recorded, evaluated every 5 rounds, the weight value with the optimal mean intersection over union (mIoU) of the evaluation index is saved, the curves of Loss and mIoU versus the number of training rounds are plotted, and the model is evaluated using mIoU, Accuracy, and mean average precision (mAP).
[0103] As can be seen from the above embodiments, the present invention collects the track line images through a laser camera, and converts the images into a top view through camera calibration and inverse perspective transformation. At the same time, a track line segmentation model is constructed, and the detailed features and global information of the track line are obtained through the backbone network. The binary segmentation branch and the track line embedding branch respectively perform binary segmentation and track area division on the features extracted by the backbone network, and cluster all the pixels belonging to the same track line by using the segmentation map obtained by the binary segmentation branch and the result of the track embedding branch. Finally, the least squares method post-processing module fits the pixels of all the track lines belonging to the same track line, and outputs a complete and visual track line. Through the present invention, on the basis of ensuring the accurate fitting of the track line, the timeliness of recognition can be ensured. Compared with the traditional edge detection method, its detection accuracy is better and the timeliness is better, meeting the high-precision and high-real-time requirements for actual track line detection.
[0104] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by orientation words such as "front, rear, upper, lower, left, right", "lateral, vertical, vertical, horizontal" and "top, bottom" is usually based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description. Without contrary description, these orientation words do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the protection scope of the present invention; the orientation words "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0105] For the convenience of description, spatial relative terms such as "above...", "above...", "on the upper surface of...", "above" can be used here to describe the spatial positional relationship between a device or feature shown in the figure and other devices or features. It should be understood that the spatial relative terms are intended to include different orientations in use or operation in addition to the orientation described in the figure for the device. For example, if the device in the figure is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "beneath other devices or structures" afterwards. Thus, the exemplary term "above..." can include both orientations of "above..." and "below...". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the corresponding explanations are made for the spatial relative descriptions used here.
[0106] In addition, it should be noted that the use of words such as "first" and "second" to limit the components is only for the convenience of distinguishing the corresponding components. Without otherwise stating, the above words have no special meanings. Therefore, it cannot be understood as a limitation on the protection scope of the present invention.
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A track line detection method based on semantic segmentation, characterized in that: The method includes: A laser camera is used to collect a track line image, and a preprocessing module is used to preprocess the collected track line image to obtain a preprocessed image; Extract track line features from preprocessed images through the backbone network; The extracted track line features are segmented into two values using the binary segmentation module, and the extracted track line features are divided into track areas using the track embedding module. All pixels belonging to the same track line are clustered using the segmentation and division results. The clustered pixels are fitted using the least squares post-processing module to output a complete visual track line.
2. The method according to claim 1, characterized in that The collected track line image is preprocessed to obtain the preprocessed image including: Perform camera calibration on the collected track line images; Perform inverse perspective transformation on the calibrated track line image to obtain a top view.
3. The method according to claim 2, characterized in that Camera calibration of the acquired track line image includes: The acquired track line image is transformed in the order of world coordinate system, camera coordinate system, image coordinate system and image pixel coordinate system.
4. The method according to claim 1, characterized in that The binary segmentation of the extracted track line features includes: The weighted cross entropy loss function, variance loss function and distance loss function are used to binary segment the extracted track line features.
5. The method according to claim 4, characterized in that The track line features are extracted from the preprocessed image through the ResNet34 backbone network, which includes a convolutional attention module. The extracted track line features include local features and global features.
6. A track line detection system based on semantic segmentation, characterized in that: The system includes: Laser camera, used to collect images of track lines; A preprocessing module, used for preprocessing the collected track line image to obtain a preprocessed image; Backbone network, used to extract track line features based on preprocessed images; Binary segmentation module, used to perform binary segmentation on the extracted track line features; Track embedding module, used to divide the track area based on the extracted track line features; A clustering module, used to cluster all pixels belonging to the same track line using the segmentation results and the division results; The least squares post-processing module is used to fit the clustered pixels and output a complete visual track line.
7. The system according to claim 6, characterized in that The collected track line image is preprocessed using the preprocessing module to obtain the preprocessed image including: Perform camera calibration on the collected track line images; Perform inverse perspective transformation on the calibrated track line image to obtain a top view.
8. The system according to claim 7, characterized in that Camera calibration of the acquired track line image includes: The acquired track line image is transformed in the order of world coordinate system, camera coordinate system, image coordinate system and image pixel coordinate system.
9. The system according to claim 6, characterized in that The binary segmentation module is used to perform binary segmentation on the extracted track line features, including: The binary segmentation module is used to perform binary segmentation on the extracted track line features using weighted cross entropy loss function, variance loss function and distance loss function.
10. The system according to claim 9, characterized in that The backbone network is the ResNet34 backbone network, which includes a convolutional attention module. The extracted track line features include local features and global features.
Citation Information
Cited By
Method and system for detecting forward obstacle of unsupervised train
CN120808302A