A Deep Learning-Based Method and System for Rail Damage Detection
By improving the bounding box regression loss function and multi-scale feature fusion strategy, the accuracy and robustness of deep learning models in rail damage detection are enhanced, solving the problems of sample imbalance and small target detection, and achieving efficient and accurate rail damage identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-26
AI Technical Summary
Existing track flaw detection methods have shortcomings in identifying a few types of damage, modeling global features, and conducting efficient offline flaw detection, especially in detecting imbalanced samples and small targets, where it is difficult to achieve high accuracy and efficiency.
A deep learning-based rail damage detection method is adopted. By improving the bounding box regression loss function, multi-scale feature fusion and data augmentation strategies, combined with feature extraction module, feature fusion module and multi-branch prediction module, the detection accuracy and robustness of rail damage are improved.
It effectively solves the problems of imbalance between positive and negative samples and uneven distribution of easy and difficult samples during training, improves the detection accuracy and robustness of small targets and hidden damage, and enhances the accuracy and practical value of rail damage identification.
Smart Images

Figure CN121661050B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rail transit monitoring and damage identification technology, specifically to a rail damage detection method and system based on deep learning. Background Technology
[0002] With the rapid growth in demand for rail transportation, rails, as a core component of rail transit, are crucial to the overall system's safety and reliability. Under long-term high loads and complex environments, rails are prone to various damages, including railhead damage, hole cracks, bolt hole defects, joint defects, and railbed damage. Failure to detect and repair these damages in a timely manner can lead to train instability and even serious traffic accidents. Therefore, researching and developing efficient and accurate rail flaw detection methods is of great significance for ensuring the safe operation of rail transit.
[0003] Machine vision-based detection methods utilize camera equipment combined with image processing algorithms for track damage detection. These methods have certain advantages in handling surface damage, but are limited by their relatively simple feature extraction methods, making them difficult to handle the complexity of track structures and the diversity of damage. Their accuracy significantly decreases, especially under conditions of changing lighting and contamination interference. Deep learning methods, such as Convolutional Neural Networks (CNNs) and their improved models, have achieved significant results in object detection and image segmentation, and have therefore been introduced into track flaw detection tasks. Typical methods include rail damage detection based on frameworks such as YOLO, Faster R-CNN, and Ultralytics. These methods achieve automatic damage identification through end-to-end training, offering higher detection accuracy compared to traditional methods. However, in practical applications, convolutional structures focus more on local region features and are insufficient in modeling the global dependencies between different parts of the track, making them prone to false detections in similar backgrounds or complex environments. Meanwhile, while some deep learning models perform excellently under laboratory conditions, their high computational complexity and slow inference speed make them difficult to scale up in real-world scenarios.
[0004] At the same time, in the ultrasonic scanning images of rails, there are also cases where the targets are small and there are a lot of data. Most of the rail damage and rail structure are very regular and similar, and the targets are relatively small. In the acquired B-mode data image set, the number of some damages is too small, which leads to sample imbalance.
[0005] In summary, while existing rail flaw detection methods each have their advantages, they still have shortcomings in areas such as identification of a few types of damage, global feature modeling, and efficient offline flaw detection. Therefore, a novel method is needed that can simultaneously ensure detection accuracy and capture global information to improve the overall performance and practical value of rail damage identification. Summary of the Invention
[0006] The present invention aims to provide a rail damage detection method and system based on deep learning, so as to overcome the problems of sample imbalance and difficulty in detecting small targets in existing models, and improve the accuracy and robustness of rail damage identification.
[0007] In a first aspect, the present invention provides a rail damage detection method based on deep learning, the method comprising the following steps:
[0008] Step S1: Obtain an ultrasonic two-dimensional image of the rail to be inspected;
[0009] Step S2: Input the ultrasonic two-dimensional image into a pre-trained neural network model and output the detection result containing the damage type and location coordinates;
[0010] The training of the neural network model includes: acquiring a training sample set containing labeled information, the labeled information including ground truth boxes representing the true location and size of the damage; inputting the training sample set into the neural network model for forward propagation, and having the neural network model output predicted boxes representing the predicted location and size of the damage; calculating the error between the predicted boxes and the ground truth boxes using a bounding box regression loss function, and updating the model parameters through backpropagation;
[0011] Furthermore, the bounding box regression loss function includes an overlap loss term, a center distance loss term, and a size penalty term, and is adjusted by a dynamic weighting factor;
[0012] The overlap loss term is configured to be negatively correlated with the ratio of the intersection area to the union area of the predicted box and the ground truth box, so as to characterize the degree of geometric overlap between the predicted box and the ground truth box.
[0013] The center distance loss term is configured to represent the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, and is then normalized.
[0014] The size penalty term is configured to calculate the distance between the predicted bounding box and the ground truth bounding box in the width and height dimensions;
[0015] The dynamic weighting factor is configured to change negatively with the overlap value between the predicted box and the ground truth box, so as to increase the weight of low overlap samples in error calculation.
[0016] Preferably, in the above-mentioned deep learning-based rail damage detection method, the neural network model includes a feature extraction module, a feature fusion module, and a multi-branch prediction module connected in sequence; wherein,
[0017] The feature extraction module is configured to perform multi-level convolution and downsampling operations on the ultrasonic two-dimensional image to generate feature maps at least three different resolution levels.
[0018] The feature fusion module is configured to perform cross-scale feature aggregation on feature maps at different resolution levels. The cross-scale feature aggregation process includes: upsampling the low-resolution feature map and fusing it with the high-resolution feature map through a top-down path; and downsampling the high-resolution feature map and fusing it with the low-resolution feature map through a bottom-down path.
[0019] The multi-branch prediction module is configured to receive the feature map processed by the feature fusion module, and output the class probability and bounding box prediction location coordinates through parallel classification and regression branches, respectively.
[0020] Preferably, in the above-mentioned deep learning-based rail damage detection method, the bounding box regression loss function is:
[0021] ,
[0022] in, The bounding box regression loss function;
[0023] It is the ratio of the intersection area to the union area of the predicted bounding box and the ground truth bounding box;
[0024] γ is the dynamic weighting factor, configured to change negatively with the overlap value between the predicted box and the ground truth box;
[0025] For improvement The loss,
[0026] ,
[0027] in, For overlap loss term, = ;
[0028] For the center distance loss term, = , This represents the square of the Euclidean distance. and These represent the center points of the predicted bounding box and the ground truth bounding box, respectively. The diagonal distance represents the minimum closure region; the minimum closure region is the smallest rectangular region that simultaneously contains the predicted bounding box and the ground truth bounding box.
[0029] For size penalty items, = , This represents the square of the Euclidean distance. and These represent the widths of the predicted bounding box and the ground truth bounding box, respectively. and These represent the heights of the predicted bounding box and the ground truth bounding box, respectively. and These represent the width and height of the smallest closure region, respectively.
[0030] Preferably, in the above-mentioned deep learning-based rail damage detection method, the rail damage types corresponding to the ultrasonic two-dimensional images obtained in step S1 include: bolt hole cracks, rail head core damage, conductor hole cracks, rail joint abnormalities, and rail bottom transverse cracks.
[0031] Preferably, in the above-mentioned deep learning-based rail damage detection method, the training of the neural network model further includes data augmentation operations to increase the background complexity of the training samples and the distribution density of small target damage.
[0032] The data augmentation operation involves transforming the image and its corresponding ground truth bounding box using mosaic enhancement, hybrid enhancement, or copy-paste operations before inputting the training sample set into the model.
[0033] The mosaic enhancement operation involves randomly selecting multiple two-dimensional ultrasound images, cropping, scaling, and stitching them together into a single composite image to increase background complexity.
[0034] The hybrid enhancement operation involves pixel-level weighted superposition of two ultrasonic two-dimensional images and their corresponding labels to generate new training samples.
[0035] The copy-paste operation involves cropping the image region containing the damaged target and pasting it into the background image region that does not contain the damage, thereby increasing the number of positive samples containing the damage.
[0036] Preferably, in the above-mentioned deep learning-based rail damage detection method, the training of the neural network model further includes hyperparameter adjustment and model evaluation;
[0037] The hyperparameter tuning and model evaluation involve dynamically setting the batch size based on the GPU memory capacity of the training device and updating the parameters using stochastic gradient descent or the Adam optimizer; and
[0038] During training, the model is evaluated using an independent validation set, the mean accuracy metric is monitored, and the model weights are saved based on the best performance metric on the validation set.
[0039] Preferably, in the above-mentioned deep learning-based rail damage detection method, step S2 includes:
[0040] Determine whether the length of the ultrasonic two-dimensional image of the rail to be inspected exceeds a predetermined threshold;
[0041] If so, perform the following slide-down window operation:
[0042] According to the preset cutting size and step size, the ultrasonic two-dimensional image of the rail that exceeds the set threshold distance is divided into multiple local image blocks; the step size is less than or equal to the cutting size, so as to form an overlapping area between adjacent local image blocks to prevent damage features located at the edge from being truncated.
[0043] The local image patches are sequentially input into the neural network model for inference;
[0044] The detection results of all local image blocks are mapped back to the original image coordinate system based on the step size to obtain complete rail damage detection results;
[0045] If not, the ultrasonic two-dimensional image of the rail to be tested is directly input into the pre-trained neural network model, and the detection result containing the damage type and location coordinates is output.
[0046] Preferably, in the above-described deep learning-based rail damage detection method, the method further includes the following after step S2:
[0047] A confidence threshold filter is applied to all predicted bounding boxes output by the neural network model to remove predicted boxes with a confidence level lower than a preset value.
[0048] The filtered predicted bounding boxes are subjected to non-maximum suppression, and the cross-union ratio between overlapping predicted boxes is calculated. The predicted box with the highest score is retained and the remaining overlapping boxes are suppressed to eliminate redundant detection of the same damaged target.
[0049] Preferably, in the above-mentioned deep learning-based rail damage detection method, the feature extraction module, the feature fusion module, and the multi-branch prediction module are constructed based on the YOLO v8 network architecture.
[0050] Secondly, the present invention also provides a rail damage detection system based on deep learning, the system comprising:
[0051] The training module and the inference module, wherein the inference module includes:
[0052] The image acquisition unit is used to acquire ultrasonic two-dimensional images of the rail to be inspected;
[0053] A detection unit is used to input the ultrasonic two-dimensional image into a pre-trained neural network model and output a detection result including the damage type and location coordinates; and
[0054] The training module is used to train the neural network model;
[0055] The training of the neural network model includes: acquiring a training sample set containing labeled information, the labeled information including ground truth boxes representing the true location and size of the damage; inputting the training sample set into the neural network model for forward propagation, and having the neural network model output predicted boxes representing the predicted location and size of the damage; calculating the error between the predicted boxes and the ground truth boxes using a bounding box regression loss function, and updating the model parameters through backpropagation;
[0056] Furthermore, the bounding box regression loss function includes an overlap loss term, a center distance loss term, and a size penalty term, and is adjusted by a dynamic weighting factor;
[0057] The overlap loss term is configured to be negatively correlated with the ratio of the intersection area to the union area of the predicted box and the ground truth box, so as to characterize the degree of geometric overlap between the predicted box and the ground truth box.
[0058] The center distance loss term is configured to represent the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, and is then normalized.
[0059] The size penalty term is configured to calculate the distance between the predicted bounding box and the ground truth bounding box in the width and height dimensions;
[0060] The dynamic weighting factor is configured to change negatively with the overlap value between the predicted box and the ground truth box, so as to increase the weight of low overlap samples in error calculation.
[0061] Compared with the prior art, the present invention has the following technical effects: by introducing an improved bounding box regression loss function, it effectively solves the problems of imbalance between positive and negative samples and uneven distribution of easy and difficult samples during training; through multi-scale feature fusion and specific data augmentation strategies, it improves the detection accuracy and robustness of small targets and hidden damage such as screw hole cracks and railhead core damage. Attached Figure Description
[0062] In the accompanying drawings of the embodiments of the present invention:
[0063] Figure 1 This is a flowchart illustrating the steps of an embodiment of the rail damage detection method based on deep learning according to the present invention.
[0064] Figure 2 This is a structural block diagram of an embodiment of the rail damage detection system based on deep learning according to the present invention. Detailed Implementation
[0065] To enable those skilled in the art to better understand the technical solutions of the present invention, the clock loop forming processing method and apparatus provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0066] The invention will be described more fully below with reference to the accompanying drawings; however, the embodiments shown may be embodied in different forms, and the invention should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that the invention will be thorough and complete, and that those skilled in the art will fully understand the scope of the invention.
[0067] The accompanying drawings of the embodiments of the present invention are provided to further illustrate the embodiments of the present invention and form part of the specification. They are used together with the detailed embodiments to explain the present invention and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the description of the detailed embodiments with reference to the accompanying drawings.
[0068] The present invention is not limited to the embodiments shown in the accompanying drawings, which illustrate illustrative properties but are not intended to be limiting.
[0069] Reference Figure 1 , Figure 1 This diagram illustrates a step-by-step flowchart of an embodiment of the rail damage detection method based on deep learning according to the present invention. The method includes the following steps:
[0070] Step S0, pre-training the neural network model, includes: obtaining a training sample set containing labeled information, the labeled information including ground truth boxes representing the true location and size of the damage; inputting the training sample set into the neural network model for forward propagation, the neural network model outputting predicted boxes representing the predicted location and size of the damage; calculating the error between the predicted boxes and the ground truth boxes using the bounding box regression loss function, and updating the model parameters through backpropagation; until a preset condition is met, the training of the neural network is completed;
[0071] Step S1: Obtain an ultrasonic two-dimensional image of the rail to be inspected;
[0072] Step S2: Input the ultrasonic two-dimensional image into a pre-trained neural network model and output the detection result containing the damage type and location coordinates;
[0073] Furthermore, the bounding box regression loss function includes an overlap loss term, a center distance loss term, and a size penalty term, and is adjusted by a dynamic weighting factor;
[0074] The overlap loss term is configured to be negatively correlated with the ratio of the intersection area to the union area of the predicted box and the ground truth box, so as to characterize the degree of geometric overlap between the predicted box and the ground truth box.
[0075] The center distance loss term is configured to represent the Euclidean distance between the center point of the predicted box and the center point of the ground box, and is normalized using the diagonal distance of the minimum closure region.
[0076] The size penalty term is configured to calculate the distance between the predicted bounding box and the ground truth bounding box in the width and height dimensions;
[0077] The dynamic weighting factor is configured to change negatively with the overlap value between the predicted box and the ground truth box, so as to increase the weight of low overlap samples in error calculation.
[0078] This invention covers two main stages: data acquisition and model inference. In the data acquisition stage, an ultrasonic two-dimensional image of the rail to be inspected is acquired as input data. In the processing stage, a pre-trained neural network model is used to analyze the input two-dimensional image and directly outputs detection results containing specific damage categories and precise location coordinates.
[0079] The training process of the neural network model involves constructing a training sample set, utilizing labeled information, forward propagation prediction, and backpropagation parameter updates based on the loss function. Specifically, this embodiment employs a composite bounding box regression loss function, which consists of an overlap loss term, a center distance loss term, and a size penalty term. Furthermore, this loss function introduces dynamic weight factors to adjust each term. Through this multi-dimensional error calculation mechanism, pre-training is completed, thereby enabling the location and identification of rail damage targets.
[0080] To enable those skilled in the art to more clearly understand this embodiment, the key technical terms involved are explained below.
[0081] Rail damage detection refers to the process of automatically analyzing rail inspection data using deep learning algorithms to identify the types of damage (such as cracks, core damage, etc.) on the rail and determine their specific physical locations on the rail.
[0082] Ultrasonic 2D image: refers to the B-scan image generated after scanning the rail with ultrasonic flaw detection equipment. This image can intuitively display the cross-sectional information of the rail's internal structure and potential defects.
[0083] Training sample set: refers to the data set used to train the neural network model, which contains a large number of preprocessed two-dimensional ultrasound images, and each image is associated with corresponding annotation information.
[0084] Annotation information: This refers to manually or automatically labeled data of damaged targets in training images. Specifically, it is represented by ground truth boxes, which are rectangular bounding boxes that accurately define the location of the damage and cover its size.
[0085] Overlap loss term: This refers to the loss component calculated based on the Intersection over Union (IoU), which measures the degree of overlap between the predicted bounding box and the ground truth bounding box in geometric space. The smaller the overlap area, the larger the value of this loss term.
[0086] Center distance loss term: This refers to the loss component used to quantify the positioning deviation between the center point of the predicted bounding box and the center point of the true bounding box. This component calculates the Euclidean distance between the two points and is usually normalized using the diagonal length of the minimum closure region to eliminate the influence of the target size on the distance metric.
[0087] Minimum closure region: refers to the smallest rectangular region that can simultaneously contain both the predicted bounding box and the ground truth bounding box. In loss function calculation, the diagonal distance of this region is often used as the normalization denominator to eliminate the influence of image scale on distance metrics.
[0088] Size penalty: This is a loss component calculated based on the difference in width and height between the predicted bounding box and the actual bounding box. It aims to constrain the shape of the predicted bounding box to make it closer to the size of the actual damage.
[0089] Dynamic weighting factor: This refers to a coefficient that adaptively adjusts the loss weight based on the prediction quality. This factor is negatively correlated with the overlap value. That is, when the overlap between the predicted box and the ground truth box is low, the value of this factor increases, thereby giving the sample a higher weight in the loss calculation.
[0090] The working principle of this invention is as follows:
[0091] The first stage is the neural network model training stage.
[0092] During the training phase of the neural network model, closed-loop iterations of forward and backward propagation are performed. After the training data is input into the model, the model outputs predicted bounding boxes representing the predicted location and size of the damage. At this point, the bounding box regression loss function is called to calculate the difference between the predicted bounding boxes and the ground truth bounding boxes in the annotation information.
[0093] The calculation process is not a single-dimensional comparison, but a comprehensive consideration of three dimensions: first, the geometric overlap between the two is evaluated through the overlap loss term; second, the accuracy of the center positioning is evaluated through the center distance loss term, and normalization is used to ensure the consistency of targets at different scales; and third, the fit of the predicted box in the width and height dimensions is evaluated through the size penalty term.
[0094] Building upon this, this embodiment introduces a dynamic weighting factor to adjust the total loss. This factor is configured to be negatively correlated with the overlap between the predicted and ground truth bounding boxes. This means that during training, if the model's prediction of a sample has a low overlap with the ground truth (i.e., it belongs to a sample that is difficult to detect or accurately locate), the dynamic weighting factor will automatically increase, thereby significantly increasing the proportion of that sample in the total loss.
[0095] In response to the calculated weighted total loss, the system executes the backpropagation algorithm to update the parameters of the neural network model based on the error gradient.
[0096] Through this mechanism, the model adaptively focuses more attention on difficult samples with low overlap and poor regressibility during training iterations, while also ensuring the accuracy of center point localization and size fitting. Ultimately, this training method based on multiple constraints and dynamic weight adjustment enables the trained model to more accurately regress the bounding boxes of damaged targets when handling rail damage detection tasks, effectively improving the detection accuracy and robustness against complex backgrounds or targets with minor damage.
[0097] The second stage is the neural network model inference stage.
[0098] During the inspection phase, the system responds to the received ultrasonic two-dimensional image of the rail to be inspected, standardizes it, and inputs it into a pre-deployed neural network model. The model then performs multi-layered feature extraction and calculation to ultimately output the inspection result. This result directly indicates whether damage exists in the image, the type of damage, and the specific location coordinates of the damage in the image coordinate system.
[0099] Those skilled in the art should understand that the specific mathematical expressions of the loss function terms in the embodiments of the present invention can be adjusted according to actual application scenarios. For example, the normalization method of the center distance or the specific calculation formula of the size penalty can be adaptively modified according to computing resources or accuracy requirements. Simultaneously, the specific mapping function of the dynamic weighting factor can also adopt various forms such as power functions and exponential functions, as long as it satisfies the logical trend of negative correlation with overlap, it should be included within the protection scope of the embodiments of the present invention. This embodiment does not limit this.
[0100] In one embodiment of the present invention, the neural network model includes a feature extraction module, a feature fusion module, and a multi-branch prediction module connected in sequence; wherein,
[0101] The feature extraction module is configured to perform multi-level convolution and downsampling operations on the ultrasonic two-dimensional image to generate feature maps at least three different resolution levels.
[0102] The feature fusion module is configured to perform cross-scale feature aggregation on feature maps at different resolution levels. The cross-scale feature aggregation process includes: upsampling the low-resolution feature map and fusing it with the high-resolution feature map through a top-down path; and downsampling the high-resolution feature map and fusing it with the low-resolution feature map through a bottom-down path.
[0103] The multi-branch prediction module is configured to receive the feature map processed by the feature fusion module, and output the class probability and bounding box prediction coordinates through parallel classification and regression branches, respectively.
[0104] This invention provides a specific structure for a neural network model. This architecture enables the model to fully utilize the deep semantic information and shallow detail information in ultrasonic two-dimensional images, thereby achieving accurate identification and location of rail damage targets of different sizes.
[0105] The neural network model consists of a large number of interconnected nodes (or neurons). In this embodiment of the invention, it is a deep learning model configured to receive image data and output detection results.
[0106] The feature extraction module, which is the front end of the neural network model, automatically learns and extracts useful information representations, i.e., features, from the input two-dimensional ultrasound image. It transforms the raw pixel data into more abstract and higher-level semantic features layer by layer through multi-level convolution and downsampling operations.
[0107] Convolution is the core operation for extracting local image features. Downsampling, such as pooling, reduces the resolution of feature maps, helping to expand the receptive field of subsequent convolutional layers, thereby capturing a wider range of contextual information and reducing computational cost. Multi-level convolution and downsampling operations produce a series of feature maps with varying resolutions from high to low and semantic information from shallow to deep.
[0108] Feature maps are the outputs of each convolutional operation in a convolutional neural network. They can be viewed as a two-dimensional array, where each element represents the response intensity of a specific region of the input image to a certain feature. Feature maps at different levels correspond to features with different levels of abstraction.
[0109] The feature fusion module is the middle part of the neural network model in this embodiment, used to integrate feature maps from different levels generated by the feature extraction module. Its goal is to generate richer and more expressive feature representations for use in subsequent prediction tasks.
[0110] The cross-scale feature aggregation in this embodiment is a technique for fusing feature map information at different resolutions (scales). In this embodiment, it includes two paths:
[0111] (1) Top-down path: The low-resolution, high-semantic feature map is enlarged by upsampling operations (such as deconvolution or interpolation) and then combined with the high-resolution, low-semantic feature map; this process aims to pass global context information to the detailed shallow features.
[0112] (2) Bottom-up path: The high-resolution, detailed feature map is reduced in size by downsampling operations (e.g., convolution stride greater than 1 or pooling), and then combined with the low-resolution, semantically rich feature map; this process aims to supplement the deep features with more accurate location information.
[0113] The multi-branch prediction module, as the backend of the neural network model, receives the enhanced feature map from the feature fusion module and inputs it into multiple parallel processing branches to simultaneously execute different prediction tasks. In this embodiment, it includes a classification branch and a regression branch. The classification branch determines what type of damaged target each region of the feature map contains and outputs the probability of the corresponding category. The regression branch predicts the precise location and size of the damaged target in the image and outputs the coordinates, width, and height of the bounding box.
[0114] The workflow of the neural network model in this embodiment of the invention is as follows:
[0115] First, a 2D ultrasonic image of the rail to be inspected is input into the feature extraction module. This module generates a feature pyramid through successive convolutional and downsampling layers, containing feature maps of at least three different resolutions. High-resolution feature maps preserve rich spatial detail information, which helps in locating small targets; low-resolution feature maps have stronger semantic information, which helps in identifying large targets and understanding the global context.
[0116] Subsequently, these multi-scale feature maps are fed into the feature fusion module. This module performs bidirectional cross-scale aggregation: on the one hand, through a top-down path, it upsamples and fuses deep semantic information into shallow feature maps, enhancing their ability to identify targets; on the other hand, through a bottom-up path, it downsamples and fuses shallow detail information into deep feature maps, providing more accurate spatial localization cues. Through this bidirectional information flow, the module ultimately outputs a set of fully fused, semantically and detail-rich enhanced feature maps.
[0117] Finally, the enhanced feature map is passed to the multi-branch prediction module. This module processes the input features in parallel: the classification branch analyzes each location on the feature map to determine the damage category it represents; simultaneously, the regression branch performs precise regression on the identified target locations, outputting their bounding box coordinates. The two branches work together to ultimately generate a complete detection result containing damage category and location information.
[0118] Based on the above working principle, the embodiments of the present invention can bring the following technical effects:
[0119] First, improve multi-scale target detection capability: generate feature maps of different resolutions through the feature extraction module, and perform cross-scale aggregation using the feature fusion module, so that the model can simultaneously and effectively detect large-size damage and small-size damage in the image.
[0120] Second, improve detection accuracy: The bidirectional feature fusion path from top to bottom ensures that the model can accurately classify using high-level semantic information and accurately locate using low-level spatial details when making predictions, thereby reducing missed detections and false alarms and improving overall detection accuracy.
[0121] Third, task decoupling and optimization: Parallel classification and regression branches are adopted, which allows the model to optimize for two different sub-tasks (what and where) separately, avoiding mutual interference between tasks and helping to improve the convergence speed and final performance of the model.
[0122] Those skilled in the art will understand that the number of feature layers generated by the feature extraction module, the specific implementation method of the fusion path in the feature fusion module (e.g., connection method, convolution kernel size, etc.), and the specific network structure of the multi-branch prediction module can all be flexibly adjusted according to actual application scenarios and performance requirements. This embodiment of the invention does not impose specific limitations on these aspects. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this embodiment of the invention should be included within the protection scope of this embodiment of the invention.
[0123] In an optional implementation, the training of the neural network model further includes: after obtaining a training sample set containing labeled information, it further includes, using a task-aligned dynamic sample allocation strategy, assigning feature points (anchor points) on the feature map output by the feature fusion module as positive or negative samples; the dynamic sample allocation strategy specifically includes:
[0124] Obtain the predicted class score for each anchor point from the output of the classification branch. The intersection-union ratio (IUU) of the predicted bounding box and the ground truth bounding box for each anchor point output by the regression branch. ;
[0125] Based on the predicted category score and the crossover ratio value The task alignment metric of each anchor point relative to the ground truth bounding box is calculated using a higher-order alignment formula. The calculation formula is:
[0126] ,
[0127] in, and These are the preset hyperparameters used to adjust the classification weights and location weights, respectively.
[0128] For each ground truth bounding box, select the task alignment metric. The largest front Points marked as "positive samples" are treated as anchor points, while the remaining anchor points are treated as negative samples or ignored. The bounding box regression loss function is used to calculate the error between the predicted and ground truth bounding boxes for positive samples, and the model parameters are updated via backpropagation. For points labeled as "negative samples," their bounding box regression loss is not calculated; only the classification loss is calculated.
[0129] In traditional single-stage detectors, positive samples are typically selected solely based on an IoU (Intersection over Union) threshold. This leads to a common problem: a predicted bounding box generated for a feature point might have a high IoU (accurate localization) but low classification confidence; or high classification confidence (the model is very certain it's a defect), but the predicted bounding box is severely drifted. This misalignment between classification and localization results in inaccurate final output bounding boxes, or incorrectly retained boxes with inaccurate locations during post-processing. The above embodiment introduces... As a metric, positive samples must meet the "classification confidence level" requirement. ")" and "positioning accuracy" The model excels in both the "confidence detection" and "damage" dimensions. This ensures that the high-confidence detection boxes output by the model during the inference phase are always accompanied by high-precision location coordinates. In the rail inspection scenario, this means that as long as the system alarms to indicate damage, the marked location is accurate, greatly reducing the possibility of misinterpretation during manual review.
[0130] In one alternative implementation, the bounding box regression loss function is configured to be calculated based on the geometric relationship between the predicted box and the ground truth box.
[0131] Specifically, the bounding box regression loss function is:
[0132] ,
[0133] in, The bounding box regression loss function;
[0134] It is the ratio of the intersection area to the union area of the predicted bounding box and the ground truth bounding box;
[0135] γ is the dynamic weighting factor, configured to change negatively with the overlap value between the predicted box and the ground truth box;
[0136] For improvement The loss,
[0137] ,
[0138] in, For overlap loss term, = ;
[0139] For the center distance loss term, = , This represents the square of the Euclidean distance. and These represent the center points of the predicted bounding box and the ground truth bounding box, respectively. This represents the diagonal distance of the smallest closure region;
[0140] For size penalty items, = , This represents the square of the Euclidean distance. and These represent the widths of the predicted bounding box and the ground truth bounding box, respectively. and These represent the heights of the predicted bounding box and the ground truth bounding box, respectively. and These represent the width and height of the smallest closure region, respectively.
[0141] To facilitate understanding of the embodiments of the present invention, the technical terms involved are explained below.
[0142] Minimum closure region: refers to the smallest rectangular region that can simultaneously contain both the predicted bounding box and the ground truth bounding box. In loss function calculation, the diagonal distance of this region is often used as the normalization denominator to eliminate the influence of image scale on distance metrics.
[0143] Euclidean distance: the straight-line distance between two points in a Cartesian coordinate system. In this embodiment, it is used to measure the positional deviation between the center point of the predicted bounding box and the center point of the true bounding box.
[0144] Dynamic weighting factor: This refers to a coefficient that changes in real time during the loss calculation process based on the current predicted state of the sample (such as the IoU value). Its function is similar to the modulation coefficient in Focal Loss, used to change the magnitude of gradient propagation.
[0145] Normalization: refers to the process of mapping numerical values to a specific range (usually 0 to 1). Here, it is used to ensure that the center distance loss term and the size penalty term are consistent with the IoU loss term in terms of order of magnitude, so as to avoid one term dominating the direction of gradient descent.
[0146] As can be seen, in this embodiment, the bounding box regression loss function is associated with three penalty terms and one dynamic weighting factor. The three penalty terms are:
[0147] (1) Overlap loss term: calculated based on the ratio of the intersection to the union of the predicted box and the ground truth box;
[0148] (2) Center distance loss term: Based on the Euclidean distance between the center points of the two bounding boxes, and normalized using the diagonal distance of the minimum closure region;
[0149] (3) Size penalty: Directly calculate the difference between the predicted box and the actual box in the width and height dimensions.
[0150] Furthermore, the aforementioned overlap loss term is adjusted by a dynamic weighting factor, that is, the dynamic weighting factor is correlated with the overlap value and is used to control the contribution of samples of different quality to the total loss.
[0151] This embodiment introduces an improved focused and efficient Intersection over Union (IoU) loss function. By integrating "center point distance penalty" and "true difference in side length penalty" into the loss function, it solves the gradient vanishing problem of traditional IoU loss when the bounding boxes do not intersect. Simultaneously, by introducing a "dynamic weighting factor," it addresses the issues of sample imbalance and uneven distribution of easy and difficult samples in rail damage detection scenarios. This embodiment enables the model to simultaneously focus on the overlap area of the bounding boxes, center alignment, and the accuracy of length and width dimensions during regression, and adaptively adjusts the training focus.
[0152] The working mechanism of the loss function described in this embodiment is as follows:
[0153] First, when the neural network outputs a predicted bounding box, the model calculates the distance between the center of the predicted box and the center of the ground truth bounding box. By dividing this distance by the diagonal distance of the minimum closure region, the model can perceive directional shifts. Even when the predicted and ground truth bounding boxes do not overlap at all (IoU=0), the center distance loss term still provides effective gradient information, guiding the predicted box towards the ground truth bounding box, thus overcoming the "plateau" problem of traditional IoU loss in non-overlapping states.
[0154] Furthermore, unlike loss functions that only consider the aspect ratio (such as CIoU), the size penalty term in this embodiment calculates the differences in width and height independently. In rail damage detection, some defects (such as railhead core damage and bolt hole cracks) may have similar aspect ratios but drastically different dimensions. Directly regressing the width and height allows the model to decouple the gradient optimization paths for width and height, avoiding optimization stagnation caused by the same aspect ratio but different actual dimensions, and significantly improving the regression convergence speed.
[0155] Furthermore, the dynamic weighting factor is adaptively adjusted based on the overlap (IoU) between the predicted and ground truth bounding boxes. During training, for samples with low overlap (i.e., "hard examples" or small damaged targets where the model performs poorly), the dynamic weighting factor amplifies the resulting loss value, thereby increasing its gradient weight in backpropagation; for samples with high overlap (i.e., "easy examples"), the weight is relatively reduced.
[0156] Based on the above working mechanism, this embodiment can achieve the following technical effects:
[0157] First, improve the detection accuracy of small targets and long-tailed samples: Through dynamic weighting factors, the model can automatically focus on small target samples that are difficult to detect, such as screw hole cracks and early micro-nuclear damage, which effectively alleviates the problem of imbalance between positive and negative samples and easy and difficult samples in rail damage data.
[0158] Second, it accelerates model convergence speed: The combination of independent width and height penalty terms and center distance penalty terms provides a clearer direction of regression gradient, enabling the model to more quickly lock the approximate location and shape of the damage in the early stages of training.
[0159] Third, it enhances positioning accuracy: it avoids the positioning drift problem caused by using only IoU or aspect ratio, so that the final generated detection box can fit more closely to the real edge of the damage, thus improving the intersection-over-union (IoU) index of the detection results.
[0160] It should be understood that the specific formula of the loss function mentioned in the above embodiments is only an exemplary logical representation. In practical applications, those skilled in the art can adjust the combination of weight coefficients, normalization methods, or penalty terms in the formula according to the specific data distribution characteristics.
[0161] Furthermore, although this embodiment is named "bounding box regression loss function," it can also be called "localization loss," "coordinate error function," or "IoU variant loss" in different algorithm architectures. Any loss calculation logic that includes penalties based on center point distance and independent side length differences, combined with dynamic weighting based on sample quality, should be included within the scope of protection of this embodiment.
[0162] Furthermore, this loss function is not only applicable to the detection of ultrasonic B-mode images of rails, but also to other target detection tasks based on bounding box regression, such as visual image detection of rail surfaces and fastener status detection.
[0163] In one optional implementation, the rail damage types corresponding to the ultrasonic two-dimensional images obtained in step S1 include: bolt hole cracks, rail head core damage, conductor hole cracks, rail joint abnormalities, and rail bottom transverse cracks.
[0164] To facilitate understanding of the detection objects determined in the embodiments of the present invention, the key damage terms involved are explained below.
[0165] Bolt hole cracks: These are cracks that occur around the bolt holes in the web of the rail. This type of damage is usually caused by the concentration of shear and tensile stress when a train passes by, and appears as a linear abnormal echo signal radiating outward from the hole wall in ultrasonic imaging.
[0166] Railhead core damage: refers to fatigue damage originating inside the railhead, also known as "white spots" or internal transverse cracks. It is often caused by inclusions inside the material or wheel-rail contact fatigue, and is characterized by being invisible to the outside, highly concealed, and extremely likely to lead to brittle fracture of the rail.
[0167] Wire hole cracks: These are cracks that occur in the drilled holes used for installing signal wires or connectors on the web of the rail. Their formation mechanism and morphology are similar to those of bolt hole cracks, but because wire hole diameters are typically smaller, the crack characteristics are more subtle, making detection more difficult.
[0168] Rail joint anomalies refer to structural defects present at the rail connection points (joints), including but not limited to rail end breakage, misaligned joint teeth, or abnormal rail gap spacing. These anomalies directly increase wheel-rail impact and accelerate rail deterioration.
[0169] Transverse cracks at the rail base: These are cracks perpendicular to the longitudinal axis of the rail and located at the bottom of the rail. These cracks typically develop from scratches, rust pits, or casting defects on the rail base surface under alternating loads, and are considered transverse fracture sources that pose a significant threat to train safety.
[0170] The embodiments of the present invention can achieve the following technical effects:
[0171] First, the testing scope covers all key stress-bearing parts from the rail head, rail web to the rail bottom, eliminating blind spots that may exist with a single testing method and ensuring the overall safety of the rail structure.
[0172] Secondly, it can clearly distinguish between similar but different damages such as screw hole cracks and wire hole cracks, which helps maintenance personnel to take targeted repair strategies (such as rail replacement, clamp reinforcement or speed limit operation) based on the specific damage type, thus improving operation and maintenance efficiency.
[0173] Third, for track head core damage, which is highly concealed and extremely dangerous, the recall rate of such small-target high-risk damage is significantly improved by clearly defining the category and training the model accordingly, combined with the aforementioned Focal-EIoU loss function, effectively preventing the occurrence of track breakage accidents.
[0174] It should be understood that the five types of rail damage listed above are merely examples of embodiments of the present invention, intended to illustrate the method's ability to detect multiple complex damages in parallel. In practical applications, the types of rail damage to be detected may also include, but are not limited to, rail surface abrasion, rail head spalling, corrugated wear (waves), longitudinal cracks, and other forms of contact fatigue cracks.
[0175] In one optional implementation, the present invention further optimizes the training process of the neural network model, particularly the preprocessing stage of the training samples. The training of the neural network model also includes data augmentation. Data augmentation is a technique used in deep learning training that aims to artificially expand the size and diversity of the training dataset by applying various transformations to the original data. Without collecting additional data, it simulates different imaging conditions through algorithmic means, preventing the model from overfitting.
[0176] Specifically, the data augmentation operation is configured to transform the image and its corresponding ground truth bounding box using mosaic enhancement, hybrid enhancement, or copy-paste operations before inputting the training sample set into the model.
[0177] The mosaic enhancement operation is configured to randomly select multiple 2D ultrasound images, crop, scale, and stitch them into a single synthetic image. This method allows the model to "see" more image context during a batch of training, and the scaling operation increases the diversity and complexity of the detected target scale.
[0178] The hybrid enhancement operation is configured to perform pixel-level weighted superposition of two ultrasonic 2D images and their corresponding labels to generate new training samples. For example, two samples are randomly selected from the training set, and their pixel values and corresponding labels (one-hot encoded) are weighted and summed according to a certain ratio (usually following a Beta distribution). This allows the model to be exposed to non-discrete, continuously changing images and labels during training, which helps to enhance the model's resistance to adversarial examples. For example, an image containing a "cracked screw hole" is superimposed with an image of a "normal rail" at a ratio of 0.6:0.4. The input to the model is no longer a pure crack or background, but a mixture of the two. This requires the model's output prediction to no longer be an absolute 0 or 1, but a probability value close to 0.6, thus achieving a linear transition of the inter-class decision boundary in the feature space.
[0179] The copy-paste operation is configured to crop and paste an image region containing a damaged target into a background image region without damage, thereby increasing the number of positive samples containing damage. This is particularly suitable for datasets with long-tailed distributions. For example, a pixel region containing a specific target (such as a screw hole crack) can be "extracted" from the source image, and after possible geometric transformations, "pasted" to any position in another background image (such as the web region of an intact rail), while simultaneously updating the annotation information. This process directly changes the ratio of positive to negative samples in the training batch, ensuring that the model encounters sufficient damage features in each training iteration and preventing the model from being biased towards predicting "no damage" due to a lack of positive samples.
[0180] This embodiment introduces mosaic enhancement, hybrid enhancement, and copy-paste strategies to dynamically reconstruct the input data during the training phase. This approach aims to enrich the background semantic information of the training data, artificially increase the frequency of rare damaged targets, and smooth the distribution of category labels, thereby improving the generalization ability of the neural network model under complex conditions and its robustness to small target damage.
[0181] Those skilled in the art should understand that the data augmentation operations described in the embodiments of the present invention are not limited to the individual use of the three specific methods mentioned above, but can also be any combination of them or serialized applications.
[0182] In one optional implementation, the training of the neural network model further includes hyperparameter tuning and model evaluation; the hyperparameter tuning and model evaluation involves dynamically setting the batch size according to the memory capacity of the training device and updating the parameters using stochastic gradient descent or Adam optimizer; and during the training process, evaluating the model using an independent validation set, monitoring the mean accuracy index, and saving the model weights based on the best performance index on the validation set.
[0183] As can be seen, by dynamically adjusting the batch size based on hardware memory, this embodiment solves the problem of memory overflow or insufficient resource utilization that easily occurs in deep learning models with large-scale parameters. Simultaneously, by integrating specific optimization algorithms (e.g., SGD or Adam) and an "optimal model preservation" strategy based on the validation set mAP metric, a closed-loop training and monitoring system is established. This mechanism ensures that after long-term iterations, the model retains the weight parameters that perform best on unseen data, effectively avoiding overfitting and guaranteeing the reliability of the final rail damage detection model in practical applications.
[0184] Hyperparameter tuning refers to the process of setting parameters that cannot be directly learned from the data by the model before or during model training. In this embodiment, it specifically refers to the configuration of parameters such as batch size and learning rate optimization strategy. Mean Average Precision (mAP) is a commonly used evaluation metric in the field of object detection. It is the arithmetic mean of the average precision (AP) across all categories, and it comprehensively reflects the detection accuracy of the model at different recall rates.
[0185] The training optimization mechanism described in this embodiment is as follows:
[0186] During the training initialization phase, the hardware status of the current training environment (such as a server or workstation) is first read, specifically identifying the available GPU memory capacity. Based on a preset memory usage model, the memory overhead required by the current network architecture (such as the aforementioned feature extraction and fusion module) at a specific input resolution is automatically calculated, and the maximum allowable batch size is dynamically set. This process enables the software algorithm to adapt to the hardware environment;
[0187] During the training iteration phase, the set optimizer (SGD or Adam) is used to perform parameter updates. After the data is propagated forward through the network and the loss is calculated, the optimizer fine-tunes the weight parameters in the network based on the gradient information returned from backpropagation, combined with historical momentum or adaptive learning rate.
[0188] Simultaneously, the system performs parallel or periodic model evaluations. After completing a training epoch or a preset number of iterations, the system pauses parameter updates, switches the model to evaluation mode, and inputs independent validation set data. The validation set contains rail images that have not participated in gradient updates. The model's mAP metric on the validation set is calculated. If the current mAP value is higher than the highest value in history, the current model is determined to be the "optimal model," and the current weight file is overwritten and saved or saved as the optimal weights; if it is lower than the highest value in history, only the log is recorded without updating the optimal weights.
[0189] This embodiment can bring the following technical effects:
[0190] By dynamically adjusting the batch size based on GPU memory, the number of samples processed in parallel can be increased as much as possible without causing GPU memory overflow errors. A larger batch size helps improve the statistical stability of the batch normalization layer and speeds up training.
[0191] Employing SGD or Adam optimizers allows the model to select the most suitable descent path based on the data distribution characteristics. In particular, the Adam optimizer can quickly adapt to different parameter update frequencies for data with large differences in characteristic scales, such as rail damage, thus accelerating model convergence.
[0192] By monitoring the mAP on the validation set and saving the optimal weights, this mechanism effectively implements an "early stop" or "optimal checkpoint" strategy. This effectively prevents the model from degrading in actual detection performance due to overfitting to training set noise in the later stages of training, ensuring that the delivered rail inspection system has the best practical performance.
[0193] In an optional implementation, this embodiment of the invention optimizes the image input and inference process in step S2, making it particularly suitable for processing long-distance ultrasonic images of rails. The inference process of the neural network model employs an image block processing strategy, implemented through a sliding window processing strategy.
[0194] Specifically, the sliding window processing strategy is configured as follows:
[0195] Step S2 includes:
[0196] Determine whether the length of the ultrasonic two-dimensional image of the rail to be inspected exceeds a predetermined threshold;
[0197] If so, perform the following slide-down window operation:
[0198] According to the preset cutting size and step size, the ultrasonic two-dimensional image of the rail that exceeds the set threshold distance is divided into multiple local image blocks; the step size is less than or equal to the cutting size, so as to form an overlapping area between adjacent local image blocks to prevent damage features located at the edge from being truncated.
[0199] The local image patches are sequentially input into the neural network model for inference;
[0200] The detection results of all local image blocks are mapped back to the original image coordinate system based on the step size to obtain complete rail damage detection results;
[0201] If not, the ultrasonic two-dimensional image of the rail to be tested is directly input into the pre-trained neural network model, and the detection result containing the damage type and location coordinates is output.
[0202] To enable those skilled in the art to more clearly understand this embodiment, the key technical terms involved are explained below.
[0203] Image segmentation: refers to the process of decomposing an original image that is too high in resolution or too large in size into a series of smaller sub-images that meet the input requirements of a neural network model, according to specific rules.
[0204] Sliding window: An algorithmic mechanism for traversing an image. It defines a fixed-size window (cropping size) and moves it across the original large image at fixed distances (step sizes), cropping a local image patch with each move until the entire image is covered.
[0205] Threshold distance setting: This refers to the critical image length value that triggers block processing. When the length of the rail B-display image exceeds this value, the system determines that direct scaling will cause severe distortion of the longitudinal features, thereby activating the block processing flow.
[0206] Step size: refers to the pixel distance the sliding window moves relative to the previous cropping position when performing the next cropping operation. In this embodiment, the step size determines the sampling density of image blocks.
[0207] Overlapping region: refers to the region containing the same image content between two adjacent local image blocks. Its width is equal to the difference between the cropping size and the stride.
[0208] Coordinate mapping: refers to the mathematical transformation process of converting relative coordinates (e.g., x, y relative to the top left corner of the local image patch) into absolute coordinates (e.g., x, y relative to the starting point of the entire rail) in the original large image.
[0209] This invention introduces a sliding window segmentation strategy with an overlapping mechanism to resolve the contradiction between the input size limitations of deep learning models and the large aspect ratio characteristics of ultrasonic rail images. Instead of simply performing rigid image segmentation, this embodiment constructs a physical overlap buffer between adjacent image blocks by setting a step size smaller than the cropping size. This ensures that damage features truncated at the edges of a single image block are completely preserved in the central region of adjacent image blocks, thereby eliminating detection blind spots caused by image segmentation. Subsequently, coordinate mapping transformation restores the scattered local detection results to globally consistent detection results.
[0210] The image processing mechanism described in this embodiment of the invention will be further explained as follows:
[0211] In the preprocessing stage, the system first reads the dimensional information of the ultrasonic two-dimensional image of the rail to be inspected. Since rail flaw detection is usually a continuous operation, the generated B-mode images are often long and thin strips (e.g., several meters or even kilometers in length). Directly scaling such images to the commonly used square input size (e.g., 640x640 pixels) would result in significant aspect ratio distortion, making even minute crack features undetectable. Therefore, the system compares the image length with a set threshold distance and initiates sliding window logic for extremely long images.
[0212] During the block-based execution phase, the system generates a cutting grid based on a preset clipping size (e.g., 640 pixels wide) and a step size (e.g., 512 pixels). Because the step size is smaller than the clipping size, there is a 128-pixel overlap between adjacent image blocks. Suppose a screw hole crack is located precisely on the right edge of the first image block, for example, at pixel 630. It might be missed in the first block due to incomplete features; however, in the second image block, because the window has slid 512 pixels to the right, the crack will be located on the left side of the second block, for example, at pixel 118, thus presenting a complete morphological feature.
[0213] During the inference phase, the neural network model performs an independent forward propagation for each local image patch, outputting bounding boxes in the local coordinate system. Subsequently, it is transformed to global coordinates, restoring all detection boxes to the original large image.
[0214] Furthermore, the coordinate mapping process can also include deduplication logic. For example, when the same damaged target is detected simultaneously by two adjacent image blocks in an overlapping area, the system can use a weighted average to fuse the overlapping global prediction boxes, rather than simply listing them.
[0215] This embodiment ensures that any damaged target at any location is relatively intact within at least one local image block by constructing overlapping regions, effectively preventing missed damage detection due to physical image segmentation. Furthermore, it allows the model to process long-distance rail images without drastic scaling, preserving the subtle texture features of the original ultrasonic echoes, which is crucial for identifying minute early nuclear damage or wire hole cracks.
[0216] Those skilled in the art will understand that the sliding window parameter settings described in the embodiments of the present invention are flexible. The specific values of the cropping size and step size can be adaptively adjusted according to the input requirements of the neural network model and the size of the hardware memory. For example, the step size can be a fixed value or a dynamic value that changes according to the image content density.
[0217] In a preferred embodiment, the step of mapping the detection results of all local image blocks back to the original image coordinate system based on the step size to obtain complete rail damage detection results specifically includes the following steps:
[0218] Step A: Set the global coordinate system of the ultrasonic two-dimensional image (B-view image) of the rail as follows Its width is The height is Set the cutting size to... Step size is For the first Local image patch The starting x-coordinate of its upper left corner in the global coordinate system for:
[0219] in, ,and .
[0220] When the neural network is in a local image patch The first was detected in the middle When there are multiple damaged targets, the local relative coordinates output by the network are: Map it back to the original image's global coordinate system. The transformation formula is:
[0221] ,
[0222] in, and They are equal because they are divided only in the length direction (horizontal), while the height direction remains intact.
[0223] Step B: For overlapping areas between adjacent local image patches, if the same damaged target is detected in two adjacent local image patches, obtain the location information of the first prediction box respectively. The position information of the second prediction box ,in This represents a vector containing the x and y coordinates of the center point and its width and height. ; Calculate the distance from the center point of each predicted bounding box to the geometric center of its local image patch. and according to the distance Calculate fusion weights .
[0224] Calculate fusion weights The specific process is as follows:
[0225] First, determine the width of the local image patch as... and height are And obtain the coordinates of the center point of the prediction box in the local image patch coordinate system. ;
[0226] Secondly, the center point of the prediction box and the geometric center of the local image patch are calculated. Euclidean distance between The calculation formula is:
[0227] ,
[0228] Finally, based on the Euclidean distance The fusion weights are calculated according to the Gaussian decay model. :
[0229] ,
[0230] in, The preset standard deviation parameter is used to control the weights as a function of distance. The rate of increase and decrease.
[0231] Step C: Based on the fusion weights For the first prediction box Second prediction box The location information is weighted and fused to obtain the final corrected location of the damaged target. :
[0232] ,
[0233] The final category of the damaged target is determined by the first prediction box. Second prediction box The category to which those with higher confidence levels belong.
[0234] In an optional implementation, the present invention further includes a step S3 after step S2. Specifically, step S3 is configured to perform the following operations:
[0235] First, a confidence threshold filter is applied to all predicted bounding boxes output by the neural network model, specifically removing those predicted bounding boxes whose confidence values are lower than a preset value.
[0236] Secondly, non-maximum suppression (NMS) is performed on the remaining predicted bounding boxes after filtering. This operation involves calculating the intersection-over-union (IoU) ratio between overlapping predicted boxes, retaining the predicted box with the highest score, and suppressing the remaining predicted boxes whose overlap with the highest-scoring box exceeds a set threshold, in order to eliminate redundant detection of the same damaged target.
[0237] This embodiment constructs a post-processing pipeline that combines "coarse screening" and "fine deduplication." During the inference phase, neural network models typically generate a large number of candidate boxes with slightly different locations and varying confidence levels for the same target. This embodiment first filters out low-probability background noise using a confidence threshold, and then uses a non-maximum suppression algorithm to solve the problem of overlapping boxes, transforming a large number of original prediction results into sparse, accurate, and unique final detection results, ensuring that the output rail damage location information has practical engineering application value.
[0238] As mentioned above, the predicted bounding box refers to the set of coordinates output by the neural network model used to define the rectangular region of a potentially damaged target. It typically includes the center point coordinates, width, height, and corresponding confidence score. The confidence score is the probability assessment value of a specific type of damage within a given predicted bounding box. This value is usually between 0 and 1; a higher value indicates that the model believes there is a greater likelihood of damage in that region. Non-maximum suppression (NMS) is a post-processing algorithm in the field of object detection. Its core idea is to "suppress non-maximums and retain local maxima," that is, to retain only the highest-scoring detection box in the local neighborhood and remove other redundant boxes that severely overlap with it.
[0239] In practice, this embodiment includes the following two stages:
[0240] The first stage is probability-based noise removal: The neural network model generates tens of thousands of anchor boxes in its output layer. The vast majority of these boxes correspond to normal rail backgrounds, and their confidence scores are typically low. The system iterates through all predicted boxes, reads their confidence attributes, and compares them with preset values. Through this step, a large number of invalid background boxes are removed, leaving only a set of candidate boxes that may contain damage.
[0241] The second stage is deduplication optimization based on topological relationships:
[0242] (1) The system sorts the remaining candidate boxes from high to low confidence.
[0243] (2) Select the prediction box with the highest confidence as the "baseline box" and put it into the "final result set";
[0244] (3) Calculate the geometric overlap (i.e., intersection-union ratio IoU) between all remaining candidate boxes and the “reference box”.
[0245] (4) If the IoU between a candidate box and the reference box exceeds the set NMS threshold (e.g., 0.45), the candidate box is determined to be a duplicate detection of the same target and is "suppressed" (i.e. deleted); if the IoU is below the threshold, it is retained.
[0246] (5) In the remaining unprocessed candidate boxes, select the box with the highest confidence as the new reference box and repeat the above process until all candidate boxes have been processed.
[0247] This invention effectively blocks the model's misjudgment of fuzzy features or similar noise textures by using confidence threshold filtering, thereby reducing the false alarm rate of the system. By using non-maximum suppression, it solves the inherent "multi-box prediction" problem of deep learning models, ensuring that for each independent damage point on the rail, the system outputs only one bounding box with the most accurate location and the highest confidence, avoiding the interference of repeated alarms to maintenance personnel, reducing the amount of data output in the final output, and alleviating the data processing load for subsequent possible inspection report generation, data uploading, or manual review.
[0248] It should be understood that the non-maximum suppression (NMS) operation mentioned in the above embodiments is only an example of a typical deduplication strategy. In practical applications, to further improve the detection effect for dense damage (such as continuous screw hole cracks), variant algorithms such as soft-NMS, matrix NMS, or weighted NMS can also be used. These improved algorithms do not directly delete overlapping boxes, but rather attenuate their confidence, thereby preserving potential neighboring targets while deduplicating.
[0249] Furthermore, the confidence threshold and NMS threshold are not fixed and can be set according to different damage categories (for example, setting a lower filtering threshold for high-risk railhead nuclear damage to ensure recall). Any technical solution that filters and optimizes prediction results based on confidence ranking and geometric overlap should be included within the protection scope of this invention.
[0250] In one alternative implementation, the feature extraction module, the feature fusion module, and the multi-branch prediction module are configured to be built based on the YOLO v8 network architecture.
[0251] Specifically, the feature extraction module corresponds to the YOLO v8 backbone network, the feature fusion module corresponds to the neck network, and the multi-branch prediction module corresponds to the decoupled head.
[0252] By applying this architecture to ultrasonic image processing of rails, and leveraging its balance between inference speed and detection accuracy, this approach addresses the slow processing speed of traditional two-stage detection algorithms and the insufficient accuracy of early versions of YOLO in detecting small targets. Through efficient gradient flow design and multi-scale feature fusion, this architecture achieves rapid extraction and accurate regression of rail damage features.
[0253] Those skilled in the art should understand that although the embodiments of the present invention explicitly mention the YOLO v8 architecture, this should be understood as a reference to an efficient detection architecture with the characteristics of "single-stage, no anchor box, and decoupled head", rather than an absolute limitation on the software version number.
[0254] Secondly, the present invention also provides an embodiment of a rail damage detection system based on deep learning, the system comprising: a training module 10 and an inference module 20; wherein,
[0255] The reasoning module 20 includes:
[0256] Image acquisition unit 201 is used to acquire ultrasonic two-dimensional images of the rail to be inspected;
[0257] Detection unit 202 is used to input the ultrasonic two-dimensional image into a pre-trained neural network model and output a detection result including the damage type and location coordinates; and
[0258] The training module is used to train the neural network model;
[0259] The training of the neural network model includes: acquiring a training sample set containing labeled information, the labeled information including ground truth boxes representing the true location and size of the damage; inputting the training sample set into the neural network model for forward propagation, and having the neural network model output predicted boxes representing the predicted location and size of the damage; calculating the error between the predicted boxes and the ground truth boxes using a bounding box regression loss function, and updating the model parameters through backpropagation;
[0260] Furthermore, the bounding box regression loss function includes an overlap loss term, a center distance loss term, and a size penalty term, and is adjusted by a dynamic weighting factor;
[0261] The overlap loss term is configured to be negatively correlated with the ratio of the intersection area to the union area of the predicted box and the ground truth box, so as to characterize the degree of geometric overlap between the predicted box and the ground truth box.
[0262] The center distance loss term is configured to represent the Euclidean distance between the center point of the predicted box and the center point of the ground box, and is normalized using the diagonal distance of the minimum closure region.
[0263] The size penalty term is configured to calculate the distance between the predicted bounding box and the ground truth bounding box in the width and height dimensions;
[0264] The dynamic weighting factor is configured to change negatively with the overlap value between the predicted box and the ground truth box, so as to increase the weight of low overlap samples in error calculation.
[0265] The rail damage detection system based on deep learning is based on the same principle as the rail damage detection method based on deep learning described above. The rail damage detection method based on deep learning has already been described, and relevant details can be found in the foregoing description. This invention will not repeat the details here.
[0266] The invention will be further described below with reference to a specific implementation.
[0267] Step 1: First, use a flaw detection vehicle to conduct comprehensive data collection on the rails, focusing on typical damage types including bolt holes, guide wire holes, rail head defects, rail base defects, bolt hole cracks, rail joints and seams, while ensuring the inclusion of ultrasonic B-mode images of normal rail sections. This data should be collected under different vehicle speeds, gains, and environmental conditions to ensure sample diversity and comprehensiveness. Acquisition parameters and location information should also be included as metadata for subsequent processing and traceability. After collection, a preliminary quality check should be performed, removing images with excessive noise or missing signals to ensure the effectiveness and stability of the training data.
[0268] Step 2: After data collection, manually annotate the B-mode images using annotation tools such as LabelImg. Clearly distinguish different types of damage, such as bolt holes, cracks, joints, and railbed defects, from the normal rail structure and name them uniformly. Annotation must adhere to consistent standards to ensure the accuracy and completeness of the selected areas. Double-checking or sampling review will improve annotation quality. All annotated data is organized into a unified data format and divided into training, validation, and test sets, generally in a ratio of 7:2:1 or 7:1.5:1.5, to ensure the reliability and generalization ability of the model training. Simultaneously, to address the imbalanced distribution of sample classes, data augmentation can be used to expand the minority class samples, thereby constructing a more balanced sample library.
[0269] Step 3: During the model training phase, input the labeled dataset into the improved Ultralytics V8X network and set appropriate training parameters based on task characteristics. For example, adjust the batch size and number of epochs according to the data scale and GPU memory, set a suitable learning rate, and use optimizers such as SGD or Adam. During training, enable common data augmentation techniques such as mosaic, mixup, and copy-paste to enhance the model's ability to recognize complex scenes and small targets. Simultaneously, considering the long-tail characteristics of damage detection, improve the model's performance on the minority class by employing an improved loss function or a positive / negative sample balancing strategy. During training, continuously monitor the model's loss function convergence and metrics such as mAP, precision, and recall, and evaluate the model's accuracy and robustness using a validation set. Ultimately, a high-precision damage recognition model suitable for offline applications is obtained.
[0270] Step 4: After the model training is complete, it is applied to offline detection of rail B-scan images. During this process, the model can automatically infer from a large number of images and output the type, location, and confidence level of damage, thus significantly improving detection efficiency and consistency. To ensure detection stability, sliding window or tile processing methods can be combined during the inference stage to avoid missing small-sized defects. Post-processing methods (such as confidence threshold filtering and non-maximum suppression) further improve detection accuracy. The final detection results can be output as a standardized detection report, including the location, type, and possible risk level of each type of damage, for reference by subsequent maintenance and repair personnel. Furthermore, by comparing offline detection results with real-time on-site detection results, the dataset can be continuously improved, providing continuous feedback and model refinement, enabling the entire system to have dynamic evolution and self-optimization capabilities.
[0271] The dataset used in this embodiment consists of B-mode images obtained by a rail flaw detection vehicle running on a transportation rail line. As the vehicle travels a distance on the rail, the echoes from multiple probes can be displayed simultaneously. The horizontal axis represents mileage, and the vertical axis represents the height of the damage on the rail. The data includes various damage types and normal rail structures, including three types of damage: bolt hole damage, rail base damage, and rail head damage, as well as three normal rail structures: joints, guide holes, and bolt holes. This real-world rail ultrasonic B-mode dataset contains 1065 images, including 778 rail head damage samples, 552 joint samples, 405 guide hole samples, 2656 bolt hole samples, 758 bolt hole damage samples, and 41 rail base damage samples. The training, test, and validation sets were divided according to a ratio of 0.75:0.2:0.05 for subsequent experiments.
[0272] The improved Ultralytics V8X model was used for training and testing on the constructed dataset.
[0273] The improved model was tested on the defined test set, yielding detection results for each category: bolt holes, cracks, wire holes, joints, railhead damage, and railbed damage. These results included prediction accuracy and recall. To more closely approximate real-world real-time flaw detection, the test results of the UltralyticsV8X model with the improved loss function and the original model on the test set are shown in the table below:
[0274] Table 1 Comparison Experiment of Improved Offline Detection Model
[0275] .
[0276] Comparative experiment
[0277] To further verify the advantages of the comprehensive improved model, some classic two-stage target detection models were selected for comparative experiments, and the results are as follows:
[0278] Table 2 Comparison Experiments of Improved Offline Detection Models
[0279]
[0280] As shown in the table above, the offline damage detection model with improved loss function outperforms the classic two-stage target detection model and the original UltralyticsV8X model, significantly improving the damage detection performance.
[0281] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based method for detecting rail damage, characterized in that, The method includes the following steps: Step S1: Obtain an ultrasonic two-dimensional image of the rail to be inspected; Step S2: Input the ultrasonic two-dimensional image into a pre-trained neural network model and output the detection result containing the damage type and location coordinates; Step S2 includes: Determine whether the length of the ultrasonic two-dimensional image of the rail to be inspected exceeds a predetermined threshold; If so, perform the following slide-down window operation: According to the preset cutting size and step size, the ultrasonic two-dimensional image of the rail that exceeds the set threshold distance is divided into multiple local image blocks; the step size is less than or equal to the cutting size, so as to form an overlapping area between adjacent local image blocks to prevent damage features located at the edge from being truncated. The local image patches are sequentially input into the neural network model for inference; The detection results of all local image blocks are mapped back to the original image coordinate system based on the step size to obtain complete rail damage detection results; For overlapping areas between adjacent local image patches, if the same damaged target is detected in two adjacent local image patches, the location information of the first prediction box is obtained respectively. The position information of the second prediction box ,in, and These represent vectors containing the x and y coordinates of the center point and its width and height, respectively. Calculate the distance from the center point of each prediction box to the geometric center of its local image patch. and according to the distance Calculate fusion weights : , in, The preset standard deviation parameter is used to control the weights as a function of distance. The rate of increase and decrease; Based on the fusion weight For the first prediction box Second prediction box The location information is weighted and fused to obtain the final corrected location of the damaged target. : , The final category of the damaged target is determined by the first prediction box. Second prediction box The category to which those with higher confidence levels belong; If not, the ultrasonic two-dimensional image of the rail to be tested is directly input into the pre-trained neural network model, and the detection result containing the damage type and location coordinates is output. Furthermore, after the neural network model obtains a training sample set containing labeled information, it also includes: A task-alignment-based dynamic sample allocation strategy is adopted to assign feature points or anchor points on the feature map output by the feature fusion module as positive or negative samples; the dynamic sample allocation strategy specifically includes: Obtain the predicted class score for each anchor point from the output of the classification branch. And the intersection-union ratio (IUU) of the predicted bounding box and the ground truth bounding box for each anchor point output by the regression branch. ; Based on the predicted category score and the crossover ratio value The task alignment metric of each anchor point relative to the ground truth bounding box is calculated using a higher-order alignment formula. The calculation formula is: , in, and These are the preset hyperparameters used to adjust the classification weights and location weights, respectively. For each ground truth bounding box, select the task alignment metric. The largest front One anchor point is used as a positive sample, and the remaining anchor points are used as negative samples or ignored samples. The bounding box regression loss function is used to calculate the error between the predicted box and the true box of the positive sample and the model parameters are updated through backpropagation. For points that are labeled as negative samples, their bounding box regression loss is not calculated, and only the classification loss is calculated.
2. The rail damage detection method based on deep learning according to claim 1, characterized in that, The training of the neural network model includes: Obtain a training sample set containing annotation information, wherein the annotation information includes a ground truth bounding box representing the true location and size of the damage; The training sample set is input into the neural network model for forward propagation, and the neural network model outputs a prediction box representing the predicted location and size of the damage. The error between the predicted box and the ground truth box is calculated using the bounding box regression loss function, and the model parameters are updated through backpropagation. Furthermore, the bounding box regression loss function includes an overlap loss term, a center distance loss term, and a size penalty term, and is adjusted by a dynamic weighting factor; The overlap loss term is configured to be negatively correlated with the ratio of the intersection area to the union area of the predicted box and the ground truth box, so as to characterize the degree of geometric overlap between the predicted box and the ground truth box. The center distance loss term is configured to represent the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, and is then normalized. The size penalty term is configured to calculate the distance between the predicted bounding box and the ground truth bounding box in the width and height dimensions; The dynamic weighting factor is configured to change negatively with the overlap value between the predicted box and the ground truth box, so as to increase the weight of low overlap samples in error calculation.
3. The rail damage detection method based on deep learning according to claim 2, characterized in that, The bounding box regression loss function is: , in, The bounding box regression loss function; It is the ratio of the intersection area to the union area of the predicted bounding box and the ground truth bounding box; The dynamic weighting factor is configured to change in a negative correlation with the overlap value between the predicted bounding box and the ground truth bounding box; For improvement loss; , in, For overlap loss term, and ; Let be the center distance loss term, and: , in, This represents the square of the Euclidean distance. and These represent the center points of the predicted bounding box and the ground truth bounding box, respectively. The diagonal distance represents the minimum closure region; the minimum closure region is the smallest rectangular region that simultaneously contains the predicted bounding box and the ground truth bounding box. As a size penalty term, and: , in, and These represent the widths of the predicted bounding box and the ground truth bounding box, respectively. and These represent the heights of the predicted bounding box and the ground truth bounding box, respectively. and These represent the width and height of the minimum closure region, respectively.
4. The rail damage detection method based on deep learning according to claim 1, characterized in that, The types of rail damage corresponding to the ultrasonic two-dimensional images obtained in step S1 include: bolt hole cracks, rail head core damage, conductor hole cracks, rail joint abnormalities, and transverse cracks at the bottom of the rail.
5. The rail damage detection method based on deep learning according to claim 1, characterized in that, The training of the neural network model also includes data augmentation operations to increase the background complexity of the training samples and the distribution density of damage to small targets; The data augmentation operation involves transforming the image and its corresponding ground truth bounding box using mosaic enhancement, hybrid enhancement, or copy-paste operations before inputting the training sample set into the model. The mosaic enhancement operation involves randomly selecting multiple two-dimensional ultrasound images, cropping, scaling, and stitching them together into a single composite image to increase background complexity. The hybrid enhancement operation involves pixel-level weighted superposition of two ultrasonic two-dimensional images and their corresponding labels to generate new training samples. The copy-paste operation involves cropping the image region containing the damaged target and pasting it into the background image region that does not contain the damage, thereby increasing the number of positive samples containing the damage.
6. The rail damage detection method based on deep learning according to claim 1, characterized in that, The training of the neural network model also includes hyperparameter tuning and model evaluation; The hyperparameter tuning and model evaluation are performed by dynamically setting the batch size according to the memory capacity of the training device and updating the parameters using stochastic gradient descent or Adam optimizer. as well as During training, the model is evaluated using an independent validation set, the mean accuracy metric is monitored, and the model weights are saved based on the best performance metric on the validation set.
7. The rail damage detection method based on deep learning according to claim 1, characterized in that, The method further includes the following after step S2: A confidence threshold filter is applied to all predicted bounding boxes output by the neural network model to remove predicted boxes with a confidence level lower than a preset value. The filtered predicted bounding boxes are subjected to non-maximum suppression, and the cross-union ratio between overlapping predicted boxes is calculated. The predicted box with the highest score is retained and the remaining overlapping boxes are suppressed to eliminate redundant detection of the same damaged target.
8. The rail damage detection method based on deep learning according to claim 1, characterized in that, The neural network model includes a feature extraction module, a feature fusion module, and a multi-branch prediction module connected in sequence; wherein... The feature extraction module is configured to perform multi-level convolution and downsampling operations on the ultrasonic two-dimensional image to generate feature maps at least three different resolution levels. The feature fusion module is configured to perform cross-scale feature aggregation on feature maps at different resolution levels. The cross-scale feature aggregation process includes: upsampling the low-resolution feature map and fusing it with the high-resolution feature map through a top-down path; and downsampling the high-resolution feature map and fusing it with the low-resolution feature map through a bottom-down path. The multi-branch prediction module is configured to receive the feature map processed by the feature fusion module, and output the class probability and bounding box prediction location coordinates through parallel classification and regression branches, respectively.
9. The rail damage detection method based on deep learning according to claim 8, characterized in that, The feature extraction module, the feature fusion module, and the multi-branch prediction module are built based on the YOLO v8 network architecture.
10. A rail damage detection system based on deep learning, characterized in that, The system includes a training module and an inference module; among which, The reasoning module includes: The image acquisition unit is used to acquire ultrasonic two-dimensional images of the rail to be inspected; The detection unit is used to input the ultrasonic two-dimensional image into a pre-trained neural network model and output the detection result including the damage type and location coordinates; The detection unit is also used to perform the following processing: Determine whether the length of the ultrasonic two-dimensional image of the rail to be inspected exceeds a predetermined threshold; If so, perform the following slide-down window operation: According to the preset cutting size and step size, the ultrasonic two-dimensional image of the rail that exceeds the set threshold distance is divided into multiple local image blocks; the step size is less than or equal to the cutting size, so as to form an overlapping area between adjacent local image blocks to prevent damage features located at the edge from being truncated. The local image patches are sequentially input into the neural network model for inference; The detection results of all local image blocks are mapped back to the original image coordinate system based on the step size to obtain complete rail damage detection results; For overlapping areas between adjacent local image patches, if the same damaged target is detected in two adjacent local image patches, the location information of the first prediction box is obtained respectively. The position information of the second prediction box ,in, and These represent vectors containing the x and y coordinates of the center point and its width and height, respectively. Calculate the distance from the center point of each prediction box to the geometric center of its local image patch. and according to the distance Calculate fusion weights : , in, The preset standard deviation parameter is used to control the weights as a function of distance. The rate of increase and decrease; Based on the fusion weight For the first prediction box Second prediction box The location information is weighted and fused to obtain the final corrected location of the damaged target. : , The final category of the damaged target is determined by the first prediction box. Second prediction box The category to which those with higher confidence levels belong; If not, the ultrasonic two-dimensional image of the rail to be tested is directly input into the pre-trained neural network model, and the detection result containing the damage type and location coordinates is output. The training module is used to train the neural network model; Furthermore, after the neural network model acquires a training sample set containing labeled information, the training module is also used to employ a task-aligned dynamic sample allocation strategy to assign feature points or anchor points on the feature map output by the feature fusion module as positive or negative samples; the dynamic sample allocation strategy specifically includes: Obtain the predicted class score for each anchor point from the output of the classification branch. And the intersection-union ratio (IUU) of the predicted bounding box and the ground truth bounding box for each anchor point output by the regression branch. ; Based on the predicted category score and the crossover ratio value The task alignment metric of each anchor point relative to the ground truth bounding box is calculated using a higher-order alignment formula. The calculation formula is: , in, and These are the preset hyperparameters used to adjust the classification weights and location weights, respectively. For each ground truth bounding box, select the task alignment metric. The largest front One anchor point is used as a positive sample, and the remaining anchor points are used as negative samples or ignored samples. The bounding box regression loss function is used to calculate the error between the predicted box and the true box of the positive sample and the model parameters are updated through backpropagation. For points that are labeled as negative samples, their bounding box regression loss is not calculated, and only the classification loss is calculated.