Track line detection method and device based on dynamic circle intersection-parallel ratio loss function optimization
Through the orbital line detection method optimized by dynamic circular intersection ratio loss function, the problem of y-axis error neglect and fixed parameters in the existing model is solved, the detection accuracy of curved track lines and the robustness of the model are improved, and the complex environment is adapted to.
Patent Information
- Application Number
- CN202510458092.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-22
AI Technical Summary
The existing track line detection model ignores the y-axis error of the track line and the generalization defects of fixed parameters, resulting in insufficient response to the linear perspective effect in the image, especially in the distant curved track line detection accuracy.
The track line detection method based on dynamic circle intersection and ratio loss function is adopted. By constructing the intersection area and union area ratio between the predicted circle and the real circle of the predicted coordinate point and the real coordinate point, dynamically adjusting the parameter radius and weighting coefficients, constructing the total loss function, and training the track line detection model.
The detection accuracy of curved track lines is improved, the robustness of the model is enhanced in complex environments, the parameter adjustment limitations of traditional model training methods are avoided, and the detection accuracy of distant curved track lines is improved.
Smart Images

Figure CN120356183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rail transit driverless technology, and in particular to a track line detection method and device based on the optimization of a dynamic circular intersection over union loss function. Background Art
[0002] In recent years, due to the frequent occurrence of rail traffic accidents caused by foreign object intrusion, the lives and property safety of passengers have been seriously threatened. Therefore, track line detection, as a prerequisite technology for preventing obstacles from intruding, is crucial for ensuring the accuracy of the train driving path. Accurate track line detection can not only provide a stable decision-making basis for the automatic driving system, but also effectively guarantee the operation of the subsequent obstacle detection system, thereby improving the overall safety of rail transit.
[0003] Although lane line detection technology has made great development in the field of autonomous driving, due to the lack of publicly available rail transit datasets, the technological development of track line detection is relatively slow. Track lines are longer and thinner than lane lines, and most of the rails are exposed to the external environment, making detection more difficult due to being easily affected by changes in the surrounding environment.
[0004] Recent research surveys have shown that when the track line detection model detects mostly straight lines, the detection effect is good, but the effect deteriorates more and more for night, rainy weather, and the curved parts or the presence of obstacles in the distance of the track. For example, in a track detection picture, due to the influence of the linear perspective effect, for the near track at the bottom of the picture, the straight part of the track is detected, and the detection effect of this part is often good; however, for the far track at the top of the picture, the track lines become dense and curved, and are more easily affected by the surrounding environment or obstacles, resulting in a decrease in detection accuracy. How to improve the detection accuracy of the track in the distance and curved conditions is of great significance for track line detection technology and the safe driving of rail transit.
[0005] In existing track line detection models, most models use the SmoothL1 loss function when regressing and predicting track points. The SmoothL1 loss combines the advantages of L1 and L2, uses the parameter β to control the smooth interval, can provide smooth optimization for small errors, and is robust to large errors and outliers. Therefore, it is widely used in regression and track line detection tasks. To further improve the regression accuracy of track lines, CLRNet proposed the Line IOU Loss, that is, the Liou loss. For each point on the track line, this method extends it left and right by a length e to form a line segment. When the predicted line and the true line trajectory perfectly coincide, Liou is 1, and when the two lines are far apart, Liou approaches -1. The researcher regards the track line as a whole rather than a discrete point set for regression, replaces the SmoothL1 loss with the Liou loss and achieves higher accuracy.
[0006] However, both the SmoothL1 loss and the Liou loss fix the position of the y-axis of the track line using preset anchor points, without considering the error in the y-direction, resulting in a significant reduction in the detection effect for distant curved tracks. Moreover, the parameter β in the SmoothL1 loss and the length e in the Liou loss are both determined by researchers through a large number of parameter tuning experiments according to the characteristics of the datasets in their research fields, leading to limited generalization ability of the model and low practicality, making it difficult to apply in actual scenarios. Summary of the Invention
[0007] To this end, the technical problem to be solved by the present invention is to overcome the problem in the prior art that the error of the y-axis of the track line is ignored and there are fixed parameter generalization defects, resulting in insufficient response to the linear perspective effect in the image for track line detection.
[0008] To solve the above technical problem, the present invention provides a track line detection method optimized based on a dynamic circular intersection over union loss function, including:
[0009] Input the annotated track line image into the track line detection model to obtain the predicted coordinate points of the track line;
[0010] Construct a total loss function according to the predicted coordinate points and the true coordinate points of the track line, including:
[0011] Obtain the true coordinate point corresponding to the current predicted coordinate point, obtain the true coordinate point of the other track line on the other side with the same ordinate as this true coordinate point, and calculate the track width corresponding to the current predicted coordinate point;
[0012] Calculate the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point;
[0013] Respectively take the current predicted coordinate point and its corresponding true coordinate point as the centers, and construct a predicted circle and a true circle according to the parameter radius corresponding to the current predicted coordinate point; calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circular intersection over union of the current predicted coordinate point, and construct the circular intersection over union loss of the current predicted coordinate point;
[0014] Calculate the corresponding weighting coefficient according to the track width corresponding to the current predicted coordinate point;
[0015] Construct a total loss function according to the circular intersection over union loss of each predicted coordinate point and its corresponding weighting coefficient;
[0016] Train the track line detection model with the total loss function, and use the trained track line detection model to detect the track line image to be detected.
[0017] Preferably, calculating the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point, the formula is:
[0018]
[0019] Among them, r i represents the parameter radius corresponding to the i-th predicted coordinate point, r max and r min represent the maximum and minimum values of the preset parameter radius respectively, W i represents the track width corresponding to the i-th predicted coordinate point, W min and W max represent the maximum and minimum values of the track width in the current track line image respectively.
[0020] Preferably, calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circle intersection over union ratio of the current predicted coordinate point, including:
[0021]
[0022] A u = 2πr 2 - A j
[0023]
[0024] Among them, (x p , y p ) represents the coordinates of the predicted coordinate point, (x g , y g ) represents the coordinates of the true coordinate point corresponding to the predicted coordinate point, d represents the distance between the predicted coordinate point and its corresponding true coordinate point, r represents the parameter radius corresponding to the predicted coordinate point, A j represents the intersection area between the predicted circle and the true circle, A u represents the union area between the predicted circle and the true circle, and Ciou represents the circle intersection over union ratio of the predicted coordinate point.
[0025] Preferably, after obtaining the circle intersection over union ratio of the current predicted coordinate point, introduce a minimum bounding circle compensation mechanism, and optimize the circle intersection over union ratio of the current predicted coordinate point by calculating the area of the minimum bounding circle of the predicted circle and the true circle. The formula includes:
[0026]
[0027] Among them, d represents the distance between the predicted coordinate point and its corresponding true coordinate point, r represents the parameter radius corresponding to the predicted coordinate point, A c represents the area of the minimum bounding circle of the predicted circle and the true circle, A u represents the union area between the predicted circle and the true circle, Ciou represents the circle intersection over union ratio of the predicted coordinate point, and Riou represents the optimized circle intersection over union ratio of the predicted coordinate point.
[0028] Preferably, the circle intersection over union loss of the current predicted coordinate point is constructed, and the formula is:
[0029]
[0030] where Riou i represents the optimized circle intersection over union of the i-th predicted coordinate point, represents the circle intersection over union loss of the i-th predicted coordinate point.
[0031] Preferably, the weighted coefficient corresponding to the current predicted coordinate point is calculated according to the track width corresponding to the current predicted coordinate point, and the formula is:
[0032]
[0033] where W i represents the track width corresponding to the i-th predicted coordinate point, and ω i represents the weighted coefficient corresponding to the i-th predicted coordinate point.
[0034] Preferably, the total loss function is constructed according to the circle intersection over union loss of each predicted coordinate point and its corresponding weighted coefficient, and the formula is:
[0035]
[0036] where Loss DR represents the total loss function, ω i represents the weighted coefficient corresponding to the i-th predicted coordinate point, represents the circle intersection over union loss of the i-th predicted coordinate point, and N represents the total number of predicted coordinate points.
[0037] Preferably, the track line detection model includes:
[0038] A feature extraction module for extracting feature information from the input track line image and outputting a track line feature map;
[0039] A feature processing module connected to the feature extraction module for extracting the global geometric information of the track line feature map;
[0040] A regression prediction module connected to the feature processing module for outputting the predicted coordinate points of the track line according to the global geometric information.
[0041] Preferably, the backbone network of the feature extraction module adopts ResNet or EfficientNet.
[0042] The present invention also provides a track line detection device optimized based on a dynamic circle intersection over union loss function, including:
[0043] A model construction module for inputting the labeled track line image into the track line detection model to obtain the predicted coordinate points of the track line;
[0044] A loss construction module for constructing a total loss function according to the predicted coordinate points and the true coordinate points of the track line, including:
[0045] A track width acquisition unit for obtaining the true coordinate point corresponding to the current predicted coordinate point, obtaining the true coordinate point of the other track line with the same ordinate as the true coordinate point, and calculating the track width corresponding to the current predicted coordinate point;
[0046] A parameter radius acquisition unit for calculating the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point;
[0047] A circular intersection over union loss calculation unit for constructing a predicted circle and a true circle with the current predicted coordinate point and its corresponding true coordinate point as the centers respectively according to the parameter radius corresponding to the current predicted coordinate point; calculating the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circular intersection over union of the current predicted coordinate point, and constructing the circular intersection over union loss of the current predicted coordinate point;
[0048] A weighted coefficient acquisition unit for calculating the corresponding weighted coefficient according to the track width corresponding to the current predicted coordinate point;
[0049] A total loss calculation unit for constructing a total loss function according to the circular intersection over union loss of each predicted coordinate point and its corresponding weighted coefficient;
[0050] A detection module for training the track line detection model with the total loss function and detecting the to-be-detected track line image with the trained track line detection model.
[0051] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0052] A method for detecting track lines optimized based on a dynamic circle intersection over union loss function constructs a total loss by using predicted coordinate points and true coordinate points to construct a predicted circle and a true circle. By converting the error distance between the predicted coordinate points and the true coordinate points into the ratio of the intersection area and the union area between the predicted circle and the true circle, it makes up for the problem that the traditional regression loss function ignores the error in the y-axis direction and improves the detection accuracy of curved track lines. Moreover, based on the perspective effect during training, the present invention dynamically adjusts the parameter radius and the weighting coefficient for constructing the predicted circle and the true circle according to the track width corresponding to the predicted coordinate points, and can assign different weights to track lines at different longitudinal positions without manual parameter adjustment. By using a larger parameter radius and a lower weighting coefficient at the proximal end of the track and a smaller parameter radius and a higher weighting coefficient at the distal end of the track, the model training can focus on the curved area far away in the track line image, improving the detection accuracy of the model for the curved track line far away. The present invention not only improves the detection accuracy of curved track lines, but also avoids the limitation of manual parameter adjustment required by the traditional model training method, providing a more robust solution for track line detection in complex environments.
[0053] Furthermore, the present invention introduces a minimum bounding circle compensation mechanism to optimize the intersection over union of the predicted circle by calculating the area of the minimum bounding circle of the predicted circle and the true circle, imposing a greater penalty on the predicted coordinate points far from the true coordinate points, which can improve the spatial matching accuracy between the predicted coordinate points and the true coordinate points and further improve the detection accuracy of the curved area at the distal end of the track line. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the drawings, where:
[0055] Figure 1 is a flowchart of a method for detecting track lines optimized based on a dynamic circle intersection over union loss function of the present invention;
[0056] Figure 2 is a structural diagram of a track line detection model;
[0057] Figure 3 is a flowchart of constructing a total loss function according to the predicted coordinate points and the true coordinate points of the track line;
[0058] Figure 4 is a schematic diagram of dynamically calculating the parameter radius;
[0059] Figure 5 is a comparison illustration of the circle intersection over union loss and the existing regression loss function Liou;
[0060] Figure 6 is a schematic diagram of the minimum bounding circle compensation mechanism;
[0061] Figure 7 It is a comparison effect diagram of different hyperparameter selections of the existing regression loss function Liou in the RailSem19 dataset;
[0062] Figure 8 It is a comparison diagram of the TestIOU index obtained by different regression loss functions in the RailSem19 dataset under different training rounds;
[0063] Figure 9 It is a comparison diagram of the TestIOU index obtained by different regression loss functions in the DL-Rail dataset under different training rounds;
[0064] Figure 10 It is a comparison diagram of the feature visualization of different loss functions in the track line detection scenario of the RailSem19 dataset, where Figure 10 The (a) in it is the feature visualization diagram of the loss function SmoothL1, Figure 10 The (b) in it is the feature visualization diagram of the loss function Liou, Figure 10 The (c) in it is the feature visualization diagram of the loss function of the present invention;
[0065] Figure 11 It is a comparison diagram of the feature visualization of different loss functions in the track line detection scenario of the DL-Rail dataset, where Figure 11 The (a) in it is the feature visualization diagram of the loss function SmoothL1, Figure 11 The (b) in it is the feature visualization diagram of the loss function Liou, Figure 11 The (c) in it is the feature visualization diagram of the loss function of the present invention;
[0066] Figure 12 It is a structural diagram of a track line detection device optimized based on a dynamic circular intersection over union loss function according to the present invention. Detailed implementation manners
[0067] The following further describes the present invention with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0068] Embodiment 1
[0069] Referring to Figure 1 as shown, the present invention provides a track line detection method optimized based on a dynamic circular intersection over union loss function, including:
[0070] S1: Input the annotated track line image into the track line detection model to obtain the predicted coordinate points of the track line; the annotation is the true coordinate points of the track line;
[0071] S2: Construct a total loss function based on the predicted coordinate points and the true coordinate points of the track line;
[0072] S3: Train the track line detection model with the total loss function, and use the trained track line detection model to detect the track line image to be detected, so as to obtain the track line coordinate points of the track line image to be detected.
[0073] Specifically, the process of using the labeled track line image to train the track line detection model in S1 includes:
[0074] Collect the track line images of the train's forward driving path in different scenarios, preprocess the track line images, and use coordinate points to perform regression marking on the two track lines in the track line images to obtain a track line image dataset with true coordinate point annotations; divide 80% of this dataset into a training set, 10% into a validation set, and 10% into a test set.
[0075] Specifically, referring to Figure 2 As shown, the structure of the track line detection model adopted in this embodiment includes:
[0076] A feature extraction module, which is used to extract feature information from the input track line image and output a track line feature map;
[0077] A feature processing module, connected to the feature extraction module, which is used to extract the global geometric information of the track line feature map;
[0078] A regression prediction module, connected to the feature processing module, which is used to output the predicted coordinate points of the track line according to the global geometric information.
[0079] Preferably, the backbone network of the feature extraction module adopts ResNet or EfficientNet. The feature extraction module uses ResNet (such as ResNet18, 34, 50) or EfficientNet (such as EfficientNet-b0, b1, b2, b3), and after multiple convolutional and pooling operations, generates a track line feature map with low resolution but high dimension.
[0080] Specifically, referring to Figure 3 As shown, the construction of the total loss function in S2 includes:
[0081] S21: Obtain the true coordinate point corresponding to the current predicted coordinate point, obtain the true coordinate point of the other track line with the same ordinate as this true coordinate point, and calculate the track width corresponding to the current predicted coordinate point.
[0082] The calculation formula of the track width is:
[0083] Wi = x i,right -x i,left
[0084] Wherein, W i represents the track width corresponding to the i-th predicted coordinate point; x i,right and x i,left respectively represent the abscissa of the right track line and the abscissa of the left track line with the same ordinate of the true coordinate point corresponding to the i-th predicted coordinate point.
[0085] S22: Referring to Figure 4 as shown, calculate the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point. The formula is expressed as:
[0086]
[0087] Wherein, r i represents the parameter radius corresponding to the i-th predicted coordinate point, r max and r min respectively represent the maximum value and the minimum value of the preset parameter radius, W i represents the track width corresponding to the i-th predicted coordinate point, W min and W max respectively represent the maximum value and the minimum value of the track width in the current track line image.
[0088] The formula for the maximum value and the minimum value of the track width in the current track line image is expressed as:
[0089] W max = max(W i )
[0090] W min = min(W i )
[0091] Based on the perspective effect of the track line, the present invention dynamically adjusts the parameter radius corresponding to each predicted coordinate point according to the change of the track line width in the longitudinal axis direction of the track line image.
[0092] S23: Respectively take the current predicted coordinate point and its corresponding true coordinate point as the centers, and construct a predicted circle and a true circle according to the parameter radius corresponding to the current predicted coordinate point; calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circle intersection and union ratio of the current predicted coordinate point, and construct the circle intersection and union ratio loss of the current predicted coordinate point.
[0093] Specifically, calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circle intersection and union ratio of the current predicted coordinate point. The formula includes:
[0094]
[0095] A u = 2πr 2 -A j
[0096]
[0097] where (x p , y p ) represents the coordinates of the predicted coordinate point, (x g , y g ) represents the coordinates of the true coordinate point corresponding to the predicted coordinate point, d represents the distance between the predicted coordinate point and its corresponding true coordinate point, r represents the parameter radius corresponding to the predicted coordinate point, A j represents the intersection area between the predicted circle and the true circle, A u represents the union area between the predicted circle and the true circle, and Ciou represents the circle intersection over union of the predicted coordinate point.
[0098] Construct the circle intersection over union loss of the current predicted coordinate point with the circle intersection over union Ciou of the predicted coordinate point. The formula is:
[0099]
[0100] where Ciou i represents the circle intersection over union of the i-th predicted coordinate point, represents the circle intersection over union loss of the i-th predicted coordinate point.
[0101] Referring to Figure 5 as shown, where P represents the predicted coordinate point and G represents the true coordinate point. Existing loss functions based on SmoothL1 or Liou are difficult to balance the errors in the x / y axis directions, resulting in poor smoothness of the detection results. The present invention converts the error distance between the predicted coordinate point and the true coordinate point into the ratio of the intersection area and the union area between the predicted circle and the true circle, compensates for the problem that the traditional regression loss function ignores the error in the y-axis direction, and improves the detection accuracy of the curved track line.
[0102] Preferably, referring to Figure 6 as shown, after obtaining the circle intersection over union of the current predicted coordinate point, introduce the minimum bounding circle compensation mechanism. By calculating the area of the minimum bounding circle of the predicted circle and the true circle, construct an improved intersection over union metric to optimize the circle intersection over union of the current predicted coordinate point. The formula includes:
[0103]
[0104] Construct the circle intersection over union loss of the current predicted coordinate point with the optimized circle intersection over union Riou of the predicted coordinate point. The formula is:
[0105]
[0106] Among them, Riou i represents the optimized intersection over union (IoU) of the i-th predicted coordinate point's circle, and represents the loss of the optimized IoU of the target circle for the i-th predicted coordinate point.
[0107] In this embodiment, a minimum bounding circle compensation mechanism is introduced. By calculating the area of the minimum bounding circle of the predicted circle and the true circle to optimize the IoU of the predicted coordinate points, a greater penalty is imposed on the predicted coordinate points far from the true coordinate points, which can improve the spatial matching accuracy between the predicted coordinate points and the true coordinate points, and further improve the detection accuracy of the curved area at the far end of the track line.
[0108] S24: Calculate the corresponding weighting coefficient according to the track width corresponding to the current predicted coordinate point. The formula is expressed as:
[0109]
[0110] Among them, W i represents the track width corresponding to the i-th predicted coordinate point, and ω i represents the weighting coefficient corresponding to the i-th predicted coordinate point.
[0111] S25: Construct a total loss function according to the IoU loss of each predicted coordinate point and its corresponding weighting coefficient. The formula is expressed as:
[0112]
[0113] Among them, Loss DR represents the total loss function, ω i represents the weighting coefficient corresponding to the i-th predicted coordinate point, represents the IoU loss of the i-th predicted coordinate point, and N represents the total number of predicted coordinate points.
[0114] Existing loss functions, such as the regression loss function Liou, require manual adjustment of parameters (such as e), and need to be repeatedly tuned in different scenarios, with low practicability. Figure 7 It is a comparison effect diagram of the existing regression loss function Liou for different hyperparameter selections in the RailSem19 dataset. Based on the perspective effect during training, the present invention dynamically adjusts and constructs the parameter radii of the predicted circle and the true circle and the weighting coefficient according to the track width corresponding to the predicted coordinate point, and can assign different weights to the track lines at different longitudinal positions without manual repeated tuning. By using a larger parameter radius and a lower weighting coefficient at the near end of the track, and a smaller parameter radius and a higher weighting coefficient at the far end of the track, the model can focus on the curved area far from the track line image during training, improving the detection accuracy of the model for the far curved track line.
[0115] Example 2
[0116] To prove the effectiveness of the track line detection method based on the optimization of the dynamic IoU loss function described in Example 1, this example uses the RailSem19 and DL-Rail public datasets for experiments.
[0117] RailSem19 dataset: This dataset contains 8,500 images with extensive annotations. This dataset covers a wide range of conditions, including images from 530 different video sequences, covering more than 350 hours of train and tram traffic in 38 countries. These cameras are installed at different positions in the front of the train, and it captures all four seasons as well as diverse weather and lighting conditions, and includes images taken by cameras set from different perspectives. RailSem19 is the most popular dataset in the field of computer vision for railway intelligent systems and is particularly suitable for training powerful and general models for track line self-path detection.
[0118] DL-Rail dataset: This dataset is a challenging urban rail transit detection dataset. To collect data, staff installed cameras at the front end of the trams on Line 202 in Dalian, Liaoning Province, and recorded videos of the front scenes during actual train operation with a video resolution of 1920×1080. This dataset carefully selected 50 videos, covering challenging scenes such as insufficient lighting and bad weather. To avoid scene repetition or similarity between adjacent frames, each video was sampled at an interval of 2 seconds, and finally 7,000 images were collected to form the dataset. The sample track lines in this dataset are greatly affected by urban environmental changes, which helps to evaluate the model's ability to handle various types of challenging scenes.
[0119] The detailed structural parameter settings of the feature extraction module, feature processing module, and regression prediction module in the track line detection model are shown in Table 1.
[0120] Table 1. Explanation of the structural parameters of the track line detection model
[0121]
[0122] According to the network structure shown in Table 1, the model processing starts from the input of a 512×512×3 RGB track line image. First, through the feature extraction module, a 7×7 convolutional layer (stride = 2) is used to halve the spatial dimension to 256×256×64 and increase the number of channels. Subsequently, it is further downsampled to 128×128×64 through a 3×3 max pooling layer. Then, progressive feature extraction is carried out through four stages of residual block groups: Stage1 uses 2 residual blocks with 64 channels to maintain the resolution, Stage2 downsamples the feature map to 64×64×128 through a stride convolution and expands the channels, Stage3 continues to downsample to 32×32×256, and finally Stage4 outputs high-dimensional features of 16×16×512; in the feature processing module, after compressing 512 channels to 8 channels through a 1×1 convolution, adaptive average pooling is used to compress the features into a 1×1×8 global vector; finally, in the regression prediction module, regression prediction from 8 dimensions to 2048 dimensions and then to 129 dimensions is achieved through two-level fully connected layers.
[0123] The main parameter settings in the training and fine-tuning stages of the track line detection model are shown in Table 2.
[0124] Table 2. Main parameter settings of the model
[0125]
[0126] To verify the superiority of the method of the present invention, 2 methods are selected for comparison in this embodiment, including a traditional method and a latest method, and comparisons are made on 4 evaluation indicators.
[0127] The two comparison methods are as follows:
[0128] (1) SmoothL1: In the SmoothL1 loss function, the parameter β controls the size of the smooth interval. If β is set unreasonably, the following problems may occur: When β is too large, the SmoothL1 loss function behaves like the squared error (L2 loss) in a large range. This will lead to problems such as the model being sensitive to outliers, gradient explosion, and slow training speed. When β is too small, the SmoothL1 loss function will be more like the absolute error (L1 loss) and lose smoothness. In this case, the model is not easy to converge, the training process becomes noisier, and small errors cannot be effectively punished.
[0129] (2) Liou: In the Liou loss function, the function converts the distance between points into the ratio between line segments to fit the positions of the predicted points and the true points. Expand the positions of the predicted points and the true points by e pixels to the left and right along the x-axis. When the predicted points and the true points are very close, the ratio of their line segments is close to 1. When the predicted points and the true points are very far apart, the ratio of their line segments is close to -1.
[0130] The four evaluation metrics are as follows:
[0131] (1) MAE: MAE represents the average of the absolute values of the errors between the predicted values and the true values. It is simple to calculate and is suitable for scenarios that require a more robust metric.
[0132] (2) RMSE: RMSE represents the square root of the error between the predicted values and the true values. It is very sensitive to outliers, and the smaller the value, the better the model.
[0133] (3) R2: The coefficient of determination R 2 measures the degree of explanation of the variance of the target variable by the model. The closer the value is to 1, the stronger the explanatory power of the model. It is used to measure the overall performance of the regression model, rather than the error of a single prediction point.
[0134] (4) Test IOU: Test IOU is a special metric for measuring the detection accuracy of track lines. By calculating the ratio of the area of the predicted track line region to the area of the true track line region, the overall fitting effect of track line detection can be judged.
[0135] Tables 3 and 4 show the comparison of the effects of the four evaluation metrics using different methods on two datasets.
[0136] Table 3. Comparison of the effects of the present invention and the comparative methods in the RailSem19 dataset
[0137]
[0138] Table 4. Comparison of the effects of the present invention and the comparative methods in the DL-Rail dataset
[0139]
[0140]
[0141] According to Table 3, the loss function method proposed by the present invention has better performance than the currently most commonly used regression loss function method under different backbone network frameworks. Among the detailed evaluation metrics, the method of the present invention achieves the best results in the four metrics of MAE, RMSE, R 2 and Test IOU. In the conventional regression evaluation metrics of the RailSem19 dataset, when using the loss function proposed by the present invention, in the MAE evaluation metric, compared with the SmoothL1 and Liou loss function methods, the present invention can reduce by an average of 1-3 points. In the RMSE evaluation metric, it can be reduced by an average of 5-8 points, and in R 2The evaluation metrics can be improved by 3% - 6% on average. Secondly, in terms of the evaluation metric of test IOU for the detection accuracy of specific track lines, the method of the present invention, compared with the SmoothL1 and Liou loss function methods, has an average accuracy improvement of 2% - 4%. Similarly, in the DL-Rail dataset, according to Table 4, the method of the present invention has an average reduction of 3 - 4 points in the MAE evaluation metric. It can be reduced by 6 - 7 points on average in the RMSE evaluation metric, and in R 2 The evaluation metrics can be improved by 1% - 1.5% on average. In the Test IOU evaluation metric, the method of the present invention has an average accuracy improvement of 3%.
[0142] This embodiment also evaluates the influence of the selection of the parameter radius on the accuracy and compares it with the method for selecting the dynamic parameter radius of the present invention. The results are shown in Tables 5 and 6.
[0143] Table 5. Comparison results of hyperparameter selection for the RailSem19 dataset
[0144]
[0145] Table 6. Comparison results of hyperparameter selection for the DL-Rail dataset
[0146]
[0147]
[0148] As can be seen from Tables 5 and 6, the method of dynamically selecting the parameter radius of the present invention effectively solves the challenge of hyperparameter selection. When the parameter radius r is selected in the range of 0.1 to 0.9, the experiments show that for the RailSem19 dataset, r = 0.3 provides good results, while for the DL-Rail dataset, r = 0.8 is optimal and good results are achieved in all evaluation metrics. However, this method takes a long time and cannot quickly determine the optimal hyperparameters. By using the method of dynamically selecting the parameter radius of the present invention, the approximate range of r can be quickly determined, and good or optimal results can be achieved in all evaluation metrics.
[0149] This embodiment also conducts a comparative study on the detection accuracy of track lines using various loss function schemes under different training epochs. Based on the ResNet18 backbone network, this embodiment compares the detection accuracies of track lines under 20, 50, 100, 200, and 400 training epochs. Figure 8 It is a comparison chart of the Test IOU metrics obtained by different regression loss functions under different training rounds in the RailSem19 dataset, Figure 9This is a comparison chart of the Test IOU indicators obtained by different regression loss functions in different training rounds in the DL-Rail dataset. Figure 8 and Figure 9 It can be seen that the method of the present invention always maintains the highest detection accuracy in all training cycles.
[0150] To evaluate the effectiveness of the loss function proposed in the present invention, this embodiment was tested in representative complex railway scenes of the RailSem19 and DL-Rail datasets. Figure 10 This is a feature visualization comparison of different loss functions in the RailSem19 dataset track line detection scenario. Figure 10 (a) Feature visualization of the behavior loss function SmoothL1, Figure 10 (b) Feature visualization of the behavior loss function Liou, Figure 10 Line (c) in FIG. 1 is a feature visualization diagram of the loss function of the present invention. Figure 11 This is a feature visualization comparison of different loss functions in the DL-Rail dataset track line detection scenario. Figure 11 (a) Feature visualization of the behavior loss function SmoothL1, Figure 11 (b) Feature visualization of the behavior loss function Liou, Figure 11 (c) in FIG. 1 is a characteristic visualization diagram of the loss function of the present invention. Figure 10 and Figure 11 It can be seen that compared with the previous loss function method, which performs poorly or has defects in long-distance, curved and occluded areas, the method of the present invention successfully achieves track line detection in these areas. Compared with the existing methods, the method of the present invention effectively alleviates the roughness of the far-end area of the track curve. By giving a greater weight to the far-end area of the track, the detection results of the curved area are smoother and more accurate.
[0151] Embodiment 3
[0152] Reference Figure 12 As shown, based on the track line detection method based on the optimization of the dynamic circular intersection and ratio loss function described in the first embodiment, this embodiment proposes a track line detection device based on the optimization of the dynamic circular intersection and ratio loss function, including:
[0153] A model building module is used to input the annotated track line image into the track line detection model to obtain the predicted coordinate points of the track line;
[0154] The loss construction module is used to construct the total loss function based on the predicted coordinate points and the true coordinate points of the track line, including:
[0155] An orbital width acquisition unit, configured to acquire the true coordinate point corresponding to the current predicted coordinate point, acquire the true coordinate point of the other orbital line with the same ordinate as that of the true coordinate point, and calculate the orbital width corresponding to the current predicted coordinate point;
[0156] A parameter radius acquisition unit, configured to calculate the corresponding parameter radius according to the orbital width corresponding to the current predicted coordinate point;
[0157] A circular intersection over union loss calculation unit, configured to respectively use the current predicted coordinate point and its corresponding true coordinate point as the centers of circles, construct a predicted circle and a true circle according to the parameter radius corresponding to the current predicted coordinate point; calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circular intersection over union of the current predicted coordinate point, and construct the circular intersection over union loss of the current predicted coordinate point;
[0158] A weighted coefficient acquisition unit, configured to calculate the corresponding weighted coefficient according to the orbital width corresponding to the current predicted coordinate point;
[0159] A total loss calculation unit, configured to construct a total loss function according to the circular intersection over union loss of each predicted coordinate point and its corresponding weighted coefficient;
[0160] A detection module, configured to train an orbital line detection model with the total loss function, and detect the orbital line image to be detected with the trained orbital line detection model.
[0161] In summary, for an orbital line detection method based on the optimization of a dynamic circular intersection over union loss function described in the present invention, when constructing the total loss, a predicted circle and a true circle are constructed with the predicted coordinate point and the true coordinate point. By converting the error distance between the predicted coordinate point and the true coordinate point into the ratio of the intersection area and the union area between the predicted circle and the true circle, the problem that the traditional regression loss function ignores the error in the y-axis direction is compensated, and the detection accuracy of curved orbital lines is improved. Moreover, based on the perspective effect during training, the present invention dynamically adjusts the parameter radius and the weighted coefficient for constructing the predicted circle and the true circle according to the orbital width corresponding to the predicted coordinate point, and can assign different weights to orbital lines at different longitudinal positions without manual repeated parameter adjustment. By using a larger parameter radius and a lower weighted coefficient at the proximal end of the orbit and a smaller parameter radius and a higher weighted coefficient at the distal end of the orbit, the model can focus on the curved areas in the distance of the orbital line image during training, and improve the detection accuracy of the model for curved orbital lines in the distance. The present invention not only improves the detection accuracy of curved orbital lines, but also avoids the limitation of the traditional model training method that requires manual parameter adjustment, and significantly improves the adaptability to different working conditions (such as rainy days, nights) and orbital forms (such as straight lines, curves), providing a more robust solution for orbital line detection in complex environments.
[0162] Furthermore, the present invention introduces a minimum enclosing circle compensation mechanism. By calculating the area of the minimum enclosing circle of the predicted circle and the true circle, the intersection-over-union ratio of the predicted coordinate points is optimized, and a greater penalty is imposed on the predicted coordinate points far from the true coordinate points, which can improve the spatial matching accuracy between the predicted coordinate points and the true coordinate points and further improve the detection accuracy of the curved area at the far end of the track line.
[0163] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. An orbit line detection method optimized based on a dynamic IoU loss function, characterized in that Including: Input the labeled track line image into the track line detection model to obtain the predicted coordinate points of the track line; Construct a total loss function based on the predicted coordinate points and the true coordinate points of the track line, including: Obtain the true coordinate point corresponding to the current predicted coordinate point, use this true coordinate point to obtain the true coordinate point of the other track line with the same ordinate as it, and calculate the track width corresponding to the current predicted coordinate point; Calculate the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point; Respectively use the current predicted coordinate point and its corresponding true coordinate point as the centers, and construct a predicted circle and a true circle according to the parameter radius corresponding to the current predicted coordinate point; calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circle intersection over union ratio of the current predicted coordinate point, and construct the circle intersection over union loss of the current predicted coordinate point; Calculate the corresponding weighting coefficient according to the track width corresponding to the current predicted coordinate point; Construct a total loss function according to the circle intersection over union loss of each predicted coordinate point and its corresponding weighting coefficient; Train the track line detection model with the total loss function, and use the trained track line detection model to detect the track line image to be detected.
2. The orbital line detection method optimized based on the dynamic IoU loss function according to claim 1, wherein Calculate the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point, and the formula is: Among them, r i represents the parameter radius corresponding to the i-th predicted coordinate point, r max and r min respectively represent the maximum and minimum values of the preset parameter radius, W i represents the track width corresponding to the i-th predicted coordinate point, W min and W max respectively represent the maximum and minimum values of the track width in the current track line image.
3. The orbital line detection method optimized based on the dynamic IoU loss function according to claim 1, wherein Calculate the ratio of the intersection area and the union area between the predicted circle and the true circle to obtain the circle intersection over union ratio of the current predicted coordinate point, including: A u = 2πr 2 -A j Among them, (x p , y p ) represents the coordinates of the predicted coordinate point, (x g , y g ) represents the coordinates of the true coordinate point corresponding to the predicted coordinate point, d represents the distance between the predicted coordinate point and its corresponding true coordinate point, r represents the parameter radius corresponding to the predicted coordinate point, A j represents the intersection area between the predicted circle and the true circle, A u represents the union area between the predicted circle and the true circle, and Ciou represents the circle intersection over union of the predicted coordinate point.
4. The orbital line detection method optimized based on the dynamic IoU loss function according to claim 3, characterized in that After obtaining the circle intersection over union ratio of the current predicted coordinate point, introduce the minimum bounding circle compensation mechanism, and optimize the circle intersection over union ratio of the current predicted coordinate point by calculating the area of the minimum bounding circle of the predicted circle and the true circle. The formula includes: Among them, d represents the distance between the predicted coordinate point and its corresponding true coordinate point, r represents the parameter radius corresponding to the predicted coordinate point, and A c represents the area of the minimum enclosing circle of the predicted circle and the true circle, and A u represents the union area between the predicted circle and the true circle. Ciou represents the circle intersection over union of the predicted coordinate point, and Riou represents the optimized circle intersection over union of the predicted coordinate point.
5. The orbital line detection method optimized based on the dynamic IoU loss function according to claim 4, characterized in that Construct the circle intersection over union loss of the current predicted coordinate point, and the formula is: Among them, Riou i represents the optimized intersection over union of the i-th predicted coordinate point, and represents the intersection over union loss of the i-th predicted coordinate point.
6. The method for detecting an orbital line optimized based on a dynamic IoU loss function according to claim 1, wherein Calculate the corresponding weighting coefficient according to the track width corresponding to the current predicted coordinate point, and the formula is: Among them, W i represents the track width corresponding to the i-th predicted coordinate point, ω i represents the weighting coefficient corresponding to the i-th predicted coordinate point.
7. The method for detecting an orbital line optimized based on a dynamic IoU loss function according to claim 1, characterized in that, Construct a total loss function according to the circle intersection over union loss of each predicted coordinate point and its corresponding weighting coefficient, and the formula is: Among them, Loss DR represents the total loss function, ω i represents the weighting coefficient corresponding to the i-th predicted coordinate point, represents the circle intersection over union loss of the i-th predicted coordinate point, and N represents the total number of predicted coordinate points.
8. A method for detecting track lines optimized based on a dynamic IoU loss function, according to claim 1, wherein The track line detection model includes: A feature extraction module, which is used to extract feature information from the input track line image and output a track line feature map; A feature processing module, connected to the feature extraction module, which is used to extract the global geometric information of the track line feature map; A regression prediction module, connected to the feature processing module, which is used to output the predicted coordinate points of the track line according to the global geometric information.
9. A method for detecting an orbital line optimized based on a dynamic IoU loss function, as claimed in claim 8, wherein The backbone network of the feature extraction module adopts ResNet or EfficientNet.
10. An orbit line detection device optimized based on a dynamic IoU loss function, characterized in that, Including: A model construction module, which is used to input the labeled track line image into the track line detection model to obtain the predicted coordinate points of the track line; A loss construction module, which is used to construct a total loss function according to the predicted coordinate points and the true coordinate points of the track line, including: A track width acquisition unit, which is used to obtain the true coordinate point corresponding to the current predicted coordinate point, use this true coordinate point to obtain the true coordinate point of the other track line with the same ordinate as it, and calculate the track width corresponding to the current predicted coordinate point; A parameter radius acquisition unit, which is used to calculate the corresponding parameter radius according to the track width corresponding to the current predicted coordinate point; A circular intersection over union loss calculation unit, which is used to construct a predicted circle and a ground truth circle respectively with the current predicted coordinate point and its corresponding ground truth coordinate point as the centers, and construct a predicted circle according to the parameter radius corresponding to the current predicted coordinate point; calculate the ratio of the intersection area and the union area between the predicted circle and the ground truth circle to obtain the circular intersection over union of the current predicted coordinate point, and construct the circular intersection over union loss of the current predicted coordinate point; A weighted coefficient acquisition unit, which is used to calculate the corresponding weighted coefficient according to the track width corresponding to the current predicted coordinate point; A total loss calculation unit, which is used to construct a total loss function according to the circular intersection over union loss of each predicted coordinate point and its corresponding weighted coefficient; A detection module, which is used to train a track line detection model with the total loss function, and detect the track line image to be detected with the trained track line detection model.