Portable railroad track component inspection system
Patent Information
- Application Number
- PCT/US2025/010432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-01-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing railroad track inspection methods are manual, time-consuming, costly, and lack accuracy, leading to potential safety hazards and financial losses due to undetected track component issues.
An automated system using a neural network with predefined proposal templates and a matching mechanism for edge computing devices, capable of classifying fastening systems, detecting missing components, and segmenting rail and tie locations, employing an adaptively weighted loss (AWL) to enhance knowledge distillation between teacher and student models.
Enables accurate, real-time, and cost-effective detection of track components, improving performance by nearly 10% over native NanoDet models, suitable for edge devices with low computational load and fast processing speeds.
Smart Images

Figure US2025010432_02102025_PF_FP_ABST
Abstract
Description
Docket No.10780003.00112 PORTABLE RAILROAD TRACK COMPONENT INSPECTION SYSTEM STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0001] This invention was made with government support under Federal Railroad Administration Contract No. 693JJ621C000011. The government has certain rights in the disclosure. CROSS-REFERENCE TO RELATED APPLICATION
[0002] The present application claims priority to and the benefit of U.S. Provisional Patent Application No.63 / 561,898 filed March 6, 2024, the contents of which is hereby incorporated in its entirety. TECHNICAL FIELD
[0003] The subject matter disclosed herein is generally directed to a lightweight, power- efficient edge computing system, called the Autonomous Power-efficient Track Inspection System (APTIS), designed specifically for railroad track component inspection, and methods for making same, using an edge-computing device equipped with a custom computer vision model for meticulous track inspection that employs advanced deep learning and AI technologies, enabling automated detection and evaluation of the condition of various track components in concert with built-in cameras for surveillance. BACKGROUND
[0004] Railroad inspections to identify missing track components are crucial to railroad operational safety. Rail track inspections are performed on a routine basis to ensure the safety of railroad operations. As reported by the Federal Railroad Administration (FRA) safety database, FRA, Train accidents by cause form (form FRA F 6180.54 safetydata.fra.dot.gov / OfficeofSafety / publicsite / Query / inccaus.aspx, in 2022, more than 300(e.g., spikes or clips) due to a long time of use and the lack of maintenance, leading to more than $85 million financial losses. Track inspection is essential to identify these flaws in a timely manner, hence alleviating the 1 53600882 v1Docket No.10780003.00112 number and severity of accidents. However, most of the existing inspection methods are performed manually and heavily rely on the personnel's experience, and hence, are time-consuming, tedious, and costly. Accordingly, it is an object of the present disclosure to provide an automated, cost- effective, and accurate image-based system for real-time rail track inspection.
[0005] Citation or identification of any document in this application is not an admission that such a document is available as prior art to the present disclosure. SUMMARY
[0006] The above objectives are accomplished according to the present disclosure by providing in one instance a system for railroad component inspection. The system may include a railway track inspection device including at least one neural network; the at least one neural network employing at least one predefined proposal template, wherein the at least one predefined proposal template, via a matching mechanism, classifies at least one fastening system used on a railway track being inspected; and the at least one neural network employing the at least one predefined proposal template, after the at least one fastening system is classified via the matching mechanism, employs at least one predefined criteria tailored to the at least one fastening system to detect both a presence or an absence of at least one fastening component on a railway track as dictated by the at least predefined criteria tailored to the at least one fastening system. Further, the fastening component may comprise a clip, spike or bolt. Still, the at least one predefined proposal template may operate via at least one predefined proposal box being refined by at least one cascade box head. Additionally, the matching mechanism may employ the at least one predefined criteria tailored to the at least one fastening system determines if at least one fastening component on the railway track is correctly installed. Further again, the matching mechanism may employ the at least one predefined criteria tailored to the at least one fastening system determines if an empty installation hole indicates at least one missing spike or the installation hole is configured as empty as dictated by the at least one fastening system classified via the matching mechanism. Moreover, the at least one neural network may employ the at least one predefined proposal template to determine the presence or the absence of at least one fastening component on a railway track for a novel fastening system not included in a training data set for the system for railroad component inspection. Still further, the system may be integrated into an edge detection device for railway 2 53600882 v1Docket No.10780003.00112 inspection. Additionally still, the at least one neural network may employ the at least one predefined proposal template, via the matching mechanism, to classify the at least one fastening system used on the railway track being inspected via at least one specific layout of rail components associated with the at least one fastening system. Still yet again, the system may detect at least one: rail, tie, clip, missing clip, spike, bolt, installation hole or combinations of the above on the railway track. Even further, the system may detect and segment at least one rail and at least one tie to determine a location and a scale of the at least one fastening system.
[0007] In a further instance, the current disclosure may provide a method for railroad component inspection. The method may include providing at least one track inspection device including at least one neural network; the at least one neural network employing at least one predefined proposal template, wherein the at least one predefined proposal template, via a matching mechanism, classifies at least one fastening system used on a railway track being inspected; and the at least one neural network employing the at least one predefined proposal template, after the at least one fastening system is classified via the matching mechanism, employs at least one predefined criteria tailored to the at least one fastening system to detect both a presence or an absence of at least one fastening component on a railway track as dictated by the at least predefined criteria tailored to the at least one fastening system. Still, the fastening component being detected may include comprises a clip, spike or bolt. Moreover, the at least one predefined proposal template may operate via at least one predefined proposal box being refined by at least one cascade box head. Still again, the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening system may determine if at least one fastening component on the railway track is correctly installed. Additionally, the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening system may determine if an empty installation hole indicates at least one missing spike or the installation hole is configured as empty as dictated by the at least one fastening system classified via the matching mechanism. Furthermore, the at least one neural network employing the at least one predefined proposal template may determine the presence or the absence of at least one fastening component on a railway track for a novel fastening system not included in a training data set for the system for railroad component inspection. Still yet, the method may be integrated into an edge detection device for railway inspection. Yet further again, the at least one neural network employing the at 3 53600882 v1Docket No.10780003.00112 least one predefined proposal template, via the matching mechanism, may classify the at least one fastening system used on the railway track being inspected via at least one specific layout of rail components associated with the at least one fastening system. Further again, the method may detect at least one: rail, tie, clip, missing clip, spike, bolt, installation hole or combinations of the above on the railway track. Still yet again, the method may include detecting and segmenting at least one rail and at least one tie to determine a location and a scale of the at least one fastening system.
[0008] These and other aspects, objects, features, and advantages of the example embodiments will become apparent to those having ordinary skill in the art upon consideration of the following detailed description of example embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] An understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure may be utilized, and the accompanying drawings of which:
[0010] FIG.1 shows one illustration of a pipeline of a native NanoDet model.
[0011] FIG. 2 shows a histogram for the object location from the ground truth and from prediction by the teach model and by the student model based on the analysis of 64 sample images of the custom dataset.
[0012] FIG.3 shows one example of a pipeline of the adaptively weighted loss (AWL).
[0013] FIG.4 shows a graph of the relationship between weight wiand the teacher loss.
[0014] FIG.5 shows a photograph of one detection analysis.
[0015] FIG. 6 shows brightness of the image is adjusted for data augmentation at: (a) brightness reduced by 20%; (b) original image; and (c) brightness increased by 20%.
[0016] FIG.7 shows a comparison of the native NanoDet and the proposed AWL-NanoDet in P-R curves on the test dataset at: (a) spike class and (b) clip class.
[0017] FIG.8 shows detection of spikes and clips in (a) poor light conditions and (b) dense arrangement of objects.
[0018] FIG. 9 shows weight for the student loss vs. training epoch at: (a) 10 epochs; (b) 25 epochs; and (c) 100 epochs.
[0019] FIG.10 shows Table 1, Algorithm 1. 4 53600882 v1Docket No.10780003.00112
[0020] FIG.11 shows Table 2, Comparison of AWL-NanoDet with other lightweight models.
[0021] FIG.12 shows Table 3, a comparison of AWL-NanoDet with the SOTA large models.
[0022] FIG.13 shows photographs comparing concrete and wooden ties.
[0023] FIG. 14 shows Architectures of Cascade R-CNN and Cascade R-CNN predefined proposal templates (inference stage).
[0024] FIG.15 shows predefined proposal boxes of predefined proposal templates and process of creating predefined proposal boxes.
[0025] FIG.16 shows Table 4 – Criteria of each predefined proposal template.
[0026] FIGS.17A and 17B show implementation details of Cascade R-CNN with predefined proposal templates.
[0027] FIG. 18 shows photographs of field conditions and hardware configurations for the present disclosure.
[0028] FIG.19 shows Table 5 - Distribution of instances in training and validation sets.
[0029] FIG. 20 shows photographs of effect of contrast-limited adaptive histogram equalization.
[0030] FIG. 21 shows Table 6 - Performance of Cascade R-CNN with predefined proposal templates (CR-PPT), Cascade R-CNNs, and YOLOv8m on the validation set.
[0031] FIG. 22 shows images of example detection results (examples from top to bottom: Cascade R-CNN with predefined proposal templates [CR-PPT] (deep layer aggregation [DLA]- CenterNet), Cascade R-CNN (DLA-CenterNet), YOLOv8m).
[0032] FIG.23 shows Table 7 - Comparison between different finetuning methods.
[0033] FIGS.24A and 24B show images of performance of CR-PPT after different finetuning methods (examples from top to bottom: zero-shot, one-shot, few-shot, incremental learning (add one-shot images), incremental learning (add few-shot images)).
[0034] FIG.25 shows Table 8 - Inference speed (ms) on desktop and NVIDIA Jetson AGX Orin platforms (batch size = 2).
[0035] FIG.26 shows Table 9 - CR-PPT performance in missing component detection.
[0036] FIG. 27 images of CR-PPT missing detection results (examples from top to bottom: true-positive, true-negative, false-positive, false-negative. Each image includes a sub-image showing the matched template). 5 53600882 v1Docket No.10780003.00112
[0037] The figures herein are for illustrative purposes only and are not necessarily drawn to scale. DETAILED DESCRIPTION OF THE EXAMPLE EMBODIMENTS
[0038] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0039] Unless specifically stated, terms and phrases used in this document, and variations thereof, unless otherwise expressly stated, should be construed as open ended as opposed to limiting. Likewise, a group of items linked with the conjunction “and” should not be read as requiring that each and every one of those items be present in the grouping, but rather should be read as “and / or” unless expressly stated otherwise. Similarly, a group of items linked with the conjunction “or” should not be read as requiring mutual exclusivity among that group, but rather should also be read as “and / or” unless expressly stated otherwise.
[0040] Furthermore, although items, elements or components of the disclosure may be described or claimed in the singular, the plural is contemplated to be within the scope thereof unless limitation to the singular is explicitly stated. The presence of broadening words and phrases such as “one or more,” “at least,” “but not limited to” or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent.
[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.
[0042] All publications and patents cited in this specification are cited to disclose and describe the methods and / or materials in connection with which the publications are cited. All such publications and patents are herein incorporated by references as if each individual publication or patent were specifically and individually indicated to be incorporated by reference. Such 6 53600882 v1Docket No.10780003.00112 incorporation by reference is expressly limited to the methods and / or materials described in the cited publications and patents and does not extend to any lexicographical definitions from the cited publications and patents. Any lexicographical definition in the publications and patents cited that is not also expressly repeated in the instant application should not be treated as such and should not be read as defining any terms appearing in the accompanying claims. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.
[0043] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.
[0044] Where a range is expressed, a further embodiment includes from the one particular value and / or to the other particular value. The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within the respective ranges, as well as the recited endpoints. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. For example, where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, e.g., the phrase “x to y” includes the range from ‘x’ to ‘y’ as well as the range greater than ‘x’ and less than ‘y’. The range can also be expressed as an upper limit, e.g., ‘about x, y, z, or less’ and should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges 7 53600882 v1Docket No.10780003.00112 of ‘less than x’, less than y’, and ‘less than z’. Likewise, the phrase ‘about x, y, z, or greater’ should be interpreted to include the specific ranges of ‘about x’, ‘about y’, and ‘about z’ as well as the ranges of ‘greater than x’, greater than y’, and ‘greater than z’. In addition, the phrase “about ‘x’ to ‘y’”, where ‘x’ and ‘y’ are numerical values, includes “about ‘x’ to about ‘y’”.
[0045] It should be noted that ratios, concentrations, amounts, and other numerical data can be expressed herein in a range format. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms a further aspect. For example, if the value “about 10” is disclosed, then “10” is also disclosed.
[0046] It is to be understood that such a range format is used for convenience and brevity, and thus, should be interpreted in a flexible manner to include not only the numerical values explicitly recited as the limits of the range, but also to include all the individual numerical values or sub- ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. To illustrate, a numerical range of “about 0.1% to 5%” should be interpreted to include not only the explicitly recited values of about 0.1% to about 5%, but also include individual values (e.g., about 1%, about 2%, about 3%, and about 4%) and the sub-ranges (e.g., about 0.5% to about 1.1%; about 5% to about 2.4%; about 0.5% to about 3.2%, and about 0.5% to about 4.4%, and other possible sub-ranges) within the indicated range.
[0047] As used herein, the singular forms “a”, “an”, and “the” include both singular and plural referents unless the context clearly dictates otherwise.
[0048] As used herein, "about," "approximately," “substantially,” and the like, when used in connection with a measurable variable such as a parameter, an amount, a temporal duration, and the like, are meant to encompass variations of and from the specified value including those within experimental error (which can be determined by e.g., given data set, art accepted standard, and / or with e.g., a given confidence interval (e.g., 90%, 95%, or more confidence interval from the mean), 8 53600882 v1Docket No.10780003.00112 such as variations of + / -10% or less, + / -5% or less, + / -1% or less, and + / -0.1% or less of and from the specified value, insofar such variations are appropriate to perform in the disclosure. As used herein, the terms “about,” “approximate,” “at or about,” and “substantially” can mean that the amount or value in question can be the exact value or a value that provides equivalent results or effects as recited in the claims or taught herein. That is, it is understood that amounts, sizes, formulations, parameters, and other quantities and characteristics are not and need not be exact, but may be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of skill in the art such that equivalent results or effects are obtained. In some circumstances, the value that provides equivalent results or effects cannot be reasonably determined. In general, an amount, size, formulation, parameter or other quantity or characteristic is “about,” “approximate,” or “at or about” whether or not expressly stated to be such. It is understood that where “about,” “approximate,” or “at or about” is used before a quantitative value, the parameter also includes the specific quantitative value itself, unless specifically stated otherwise.
[0049] The term “optional” or “optionally” means that the subsequent described event, circumstance or substituent may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0050] As used interchangeably herein, the terms “sufficient” and “effective,” can refer to an amount (e.g., mass, volume, dosage, concentration, and / or time period) needed to achieve one or more desired and / or stated result(s). For example, a therapeutically effective amount refers to an amount needed to achieve one or more therapeutic effects.
[0051] As used herein, “tangible medium of expression” refers to a medium that is physically tangible or accessible and is not a mere abstract thought or an unrecorded spoken word. “Tangible medium of expression” includes, but is not limited to, words on a cellulosic or plastic material, or data stored in a suitable computer readable memory form. The data can be stored on a unit device, such as a flash memory or CD-ROM or on a server that can be accessed by a user via, e.g., a web interface.
[0052] Various embodiments are described hereinafter. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not 9 53600882 v1Docket No.10780003.00112 necessarily limited to that embodiment and can be practiced with any other embodiment(s). Reference throughout this specification to “one embodiment”, “an embodiment,” “an example embodiment,” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases “in one embodiment,” “in an embodiment,” or “an example embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure, in one or more embodiments. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the disclosure. For example, in the appended claims, any of the claimed embodiments can be used in any combination.
[0053] All patents, patent applications, published applications, and publications, databases, websites and other published materials cited herein are hereby incorporated by reference to the same extent as though each individual publication, published patent document, or patent application was specifically and individually indicated as being incorporated by reference. KITS
[0054] Any of the railroad track component inspection systems described herein can be presented as a combination kit. As used herein, the terms "combination kit" or "kit of parts" refers to the devices, computers, cameras, and any additional components that are used to package, sell, market, deliver, and / or provide the provide the combination of elements or a single element, such as the active ingredient, contained therein. Such additional components include, but are not limited to, sensors, train track applicable carrying systems, packaging, and the like. When one or more of the devices, computers, cameras, and any additional components described herein or a combination thereof (e.g., components contained in the kit are provided simultaneously, the combination kit can contain the railroad track component inspection systems in a single embodiment, such as a complete, ready-to-use system or in separate embodiments. When the devices, computers, cameras, and any additional components described herein or a combination thereof and / or kit components are not provided simultaneously, the combination kit can contain each devices, 10 53600882 v1Docket No.10780003.00112 computers, cameras, and any additional components in separate embodiments. The separate kit components can be contained in a single package or in separate packages within the kit.
[0055] In some embodiments, the combination kit also includes instructions printed on or otherwise contained in a tangible medium of expression. The instructions can provide information regarding the railroad track component inspection systems, safety information regarding the railroad track component inspection systems, assembly instructions, indications for use, and / or recommended maintenance regimen(s) for the railroad track component inspection systems contained therein. In some embodiments, the instructions can provide directions and protocols for providing the railroad track component inspection systems for use with a railroad section and the instructions can provide one or more embodiments of the methods for providing and using railroad track component inspection systems thereof such as any of the methods described in greater detail elsewhere herein.
[0056] This disclosure presents a new lightweight computer vision model on edge devices for accurate, real-time rail track inspection. It modifies the teacher-student guidance mechanism in NanoDet, see github.com / RangiLyu / nanodet, by introducing a new adaptively weighted loss (AWL) to the training process. The AWL evaluates the teacher and student model qualities, determines the weight of the student loss, and then balances their loss contributions on-the-fly, gearing the training process toward proper knowledge distillation and guidance. Compared to SOTA models, our AWL-NanoDet features a tiny model size of less than 10 MB and a computation cost of 1.52 G FLOPs and achieves an inference time of less than 14 ms per frame when tested on Nvidia’s AGX Orin. Relative to native NanoDet, it also notably improves the model's performance by nearly 10%, enabling highly accurate, real-time detection of track components.
[0057] Computer vision and object detection based on deep learning approaches has recently gained significant traction in various real-world applications, driven mainly by the tremendous success of Convolutional Neural Networks (CNN) based models. LeCun et al., see Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard and L. D. Jackel, "Backpropagation applied to handwritten zip code recognition," Neural computation, vol.1, 1989, proposed the CNN method and developed LeNet, see Y. LeCun, L. Bottou, Y. Bengio and P. Haffner, "Gradient- based learning applied to document recognition," Proceedings of the IEEE, vol.86, 1998, which 11 53600882 v1Docket No.10780003.00112 demonstrated the excellent accuracy and stability of the CNN-based models in classification tasks relative to the traditional approaches. The LeNet also confirmed that the CNN model was applicable to practical computer vision tasks. Later, Alex et al., see A. Krizhevsky, I. Sutskever and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," Communications of the ACM, vol.60, 2017, reported a more complex CNN structure, AlexNet, which was able to extract more complex features for image processing. Subsequently, Simonyan et al., see K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," arXiv preprint arXiv:1409.1556, 2014, proposed the VGG model, which employed a deep network to construct more comprehensive features. However, as the network's depth increased, the problems of overfitting and deterioration in the feature quality became more pronounced. To mitigate these issues, He et al., see K. He, X. Zhang, S. Ren and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, proposed the ResNet, which incorporated shortcut connections between inputs and outputs, allowing the CNN models to grow deeper without experiencing these problems.
[0058] These pioneering developments of CNNs laid a solid foundation for many object detection applications. Generally, the CNN-based object detection models can be classified into two categories: anchor-free and anchor-based models. Anchor-based methods use preconfigured anchors to match corresponding objects. The RCNN proposed by Girshick et al., see R. Girshick, J. Donahue, T. Darrell and J. Malik, "Rich feature hierarchies for accurate object detection and semantic segmentation," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, and YOLO proposed by Redmon et al., J. Redmon, S. Divvala and R. a. F. A. Girshick, "You only look once: Unified, real-time object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, are outstanding examples of such methods, demonstrating remarkable accuracy in reported results. However, a drawback of the anchor-based models is that the anchor matching step slows down the model inference. As a result, the anchor-free method was proposed. See, Law et al., H. Law and J. Deng, "Cornernet: Detecting objects as paired keypoints," in Proceedings of the European conference on computer vision (ECCV), 2018, developed a novel model structure called CornerNet, which could directly produce 12 53600882 v1Docket No.10780003.00112 the bounding box (BB) without the matching anchor. Meanwhile, the fully convolutional one-stage object detection (FCOS) proposed by Tian et al., see Z. Tian, C. Shen, H. Chen and T. He, "Fcos: Fully convolutional one-stage object detection," in Proceedings of the IEEE / CVF international conference on computer vision, 2019, could regress the object's centroid and compute the distance between the centroid to the boundary of the BB in parallel, achieving superior performance to CornerNet.
[0059] Nevertheless, the computational cost has increased dramatically as the model structures become deeper. When deployed on small form-factor edge devices, these complex models run too slow to achieve real-time detection. Although the Faster-CNN proposed by Ren et al., see S. Ren, K. He, R. Girshick and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," Advances in neural information processing systems, vol. 28, 2015, proved capable of processing a variety of tasks faster, the model still possesses 19 million trainable parameters, leading to high FLOPs that jeopardize its possibility for real-time use. Recently, the research communities have realized the significant implication of CNN-based models for deployable, real-time applications. Thus, a significant amount of research has been dedicated to developing lightweight models and approaches. One well-known lightweight model structure that serves as the backbone is MobileNet, developed by Howard et al., see A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto and H. Adam, "Mobilenets: Efficient convolutional neural networks for mobile vision applications," arXiv preprint arXiv:1704.04861, 2017. In MobileNet, depth-wise separable convolutions were designed to reduce the weight operation tremendously. Later, Szegedy et al., see C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke and A. Rabinovich, "Going deeper with convolutions," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, proposed InceptionNet, another lightweight model structure, which utilizes a new idea for reducing the number of trainable weights. Instead of exploiting deeper models, they adopted a wider model layer. Each layer has three different kernels to process the input features. Next, Iandola et al., see F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally and K. Keutzer, "SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size," arXiv preprint arXiv:1602.07360, 2016, developed the SqueezeNet, which uses a block called fire to reduce the number of parameters, and the fire uses 1*1 convolution filters to squeeze the filter size. 13 53600882 v1Docket No.10780003.00112 Then, it uses a group of 1*1 and 3*3 filters to expand the features. Subsequently, Zhang et al., see X. Zhang, X. Zhou, M. Lin and J. Sun, "Shufflenet: An extremely efficient convolutional neural network for mobile devices," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, developed the ShuffleNet, which is also based on depth-wise separable convolutions while creatively shuffling the input feature information. The feature combination between different channels could be more comprehensive. Moreover, feature representation is vastly enriched by mixing feature information. Recently, the NanoDet was proposed by RangiLyu et al., see RangiLyu. [Online]. Available: github.com / RangiLyu / nanodet, which uses the ShuffleNet as the backbone to improve feature representation and diversity.
[0060] Nevertheless, because of their simple structures, lightweight models also suffer from a common issue of worse learning capabilities. That is because the number of trainable parameters is very small, and the capacity for learning, generalizing, and converging from the data is undermined significantly. To alleviate these negative effects due to oversimplification, Hinton et al., see G. Hinton, O. Vinyals and J. Dean, "Distilling the knowledge in a neural network," arXiv preprint arXiv:1503.02531, 2015, proposed the method of knowledge distillation (KD). This technique employs a more complex model (teacher) with superior learning capabilities to guide a simplified model (student), facilitating the transfer of knowledge from the teacher to the student. One instance of this is the Label Assignment Distillation (LAD) proposed by Nguyen et al., see C. H. Nguyen, T. C. Nguyen, T. N. Tang and N. L. Phan, "Improving object detection by label assignment distillation," in Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, 2022, which uses the ResNet50 as a teacher. In LAD, the teacher model generates labels for the student model, instead of letting the student model learn directly from the ground truth. Usually, the labels assigned by the teacher model are easier to learn and more amenable to the student model. This approach is also utilized in NanoDet, where an Assign Guidance Module (AGM) is implemented as a teacher model that assigns learning labels to the student model's prediction head, enabling faster convergence. Note that in this disclosure, similarly, the AGM and prediction head will also be referred to as the teacher and student models.
[0061] The developments of new structures and methods have facilitated the creation of accurate and real-time models and processes for facility and infrastructure inspection. Guo et al., see F. Guo, Y. Qian and Y. Shi, "Real-time railroad track components inspection based on the 14 53600882 v1Docket No.10780003.00112 improved YOLOv4 framework," Automation in construction, vol.125, p.103596, 2021, extended YOLOv4 to inspecting track components of railroads, which achieved satisfactory performance in accuracy (mAP and F1 score). Additionally, Cha et al., see Y.-J. Cha, W. Choi, G. Suh, S. Mahmoudkhani and O. Bykzrk, "Autonomous structural visual inspection using region-based deep learning for detecting multiple damage types," Computer-Aided Civil and Infrastructure Engineering, vol.33, pp.731-747, 2018, introduced a damage detection system that utilizes Faster- RCNN to extract defective features from images of buildings or bridges. The CNN-based model was demonstrated to outperform traditional detection methods. In another study, Guo et al., see F. Guo, Y. Qian, Y. Wu, Z. Leng and H. Yu, "Automatic railroad track components inspection using real-time instance segmentation," Computer-Aided Civil and Infrastructure Engineering, vol.36, pp.362-377, 2021, developed a real-time inspection system for railroads using YOLACT, a pixel- level model that accurately segmented components, such as spikes and clips on the rail tracks. Zhang et al., C. Zhang, C.-c. Chang and M. Jamshidi, "Concrete bridge surface damage detection using a single-stage detector," Computer-Aided Civil and Infrastructure Engineering, vol.35, pp. 389-409, 2020, developed an improved model that exploits YOLOv3 to detect surface damage on the bridge, which showed highly accurate results. Li et al., see S. Li, X. Zhao and G. Zhou, "Automatic pixel-level multiple damage detection of concrete structure using fully convolutional network," Computer-Aided Civil and Infrastructure Engineering, vol. 34, pp. 616-634, 2019, proposed a fine-tuned DesNet-121 for detecting damages on concrete. Liang et al., see X. Liang, "Image-based post-disaster inspection of reinforced concrete bridge systems using deep learning with Bayesian optimization," Computer-Aided Civil and Infrastructure Engineering, vol.34, pp. 415-430, 2019, combined CNN models with Bayesian optimization, which achieved relatively high accuracy in inspecting bridge systems. Hao et al., see H. Feng, Z. Jiang, F. Xie, P. Yang, J. Shi and L. Chen, "Automatic fastener classification and defect detection in vision-based railway inspection systems," IEEE transactions on instrumentation and measurement, vol.63, pp.877-888, 2013, utilized Haar-like features to classify various types of fasteners. Although these seminal works demonstrated the great promise to achieve real-time performance (in both frame per second / FPS and latency) on powerful computing platforms, they are usually not well suited for field uses. On the other hand, research efforts focusing on lightweight models on edge devices to enable true mobile computing and field-deployable inspection are indeed scarce. As discussed 15 53600882 v1Docket No.10780003.00112 above, a concomitant issue of using lightweight models is the deteriorated learning capabilities and poor prediction accuracy, which could be more critical for infrastructure and facility applications.
[0062] Recognizing the urgent need, this disclosure presents a lightweight, edge-compatible computer vision model for real-time inspection of track components along the railroad by adapting the teacher-student guidance mechanism of the native NanoDet during the model training. The native NanoDet adopts an FCOS-like model and relies on the anchor-free approach (by regressing the centroid of the object), and has an incredibly straightforward model structure, resulting in low computational requirements and fast processing speed. For instance, its teacher model has four convolutional kernels, while the student model only has one. This simple structure gives rise to two issues: first, the teacher model is less capable or needs more time to learn from complex data, and second, there is a relatively notable gap in the learning ability between the teacher and the student model, which makes the original implementation of knowledge distillation less efficient because the quality of the knowledge from the teacher model may be poor, and the unqualified teacher may misguide the student model. In most of the previous efforts, the teacher's learning ability is not questioned, such as LAD, in which the knowledge from the teacher model is absorbed directly by the student model. Indeed, little research has investigated the qualities of the teacher and the student models during model training and their effect on model performance. To tackle this challenge, we present a novel strategy to evaluate the quality of the teacher model during model training, and a new adaptively weighted loss (AWL) is proposed to balance their loss contributions judiciously, hence gearing the learning process toward proper knowledge distillation and guidance. More specifically, by scoring the teacher’s knowledge / quality, our approach allows for weighing the teacher’s and the student’s contribution differently in the total loss during model training. For example, if the teacher model is not well-trained, the student model will learn less from the teacher model. Oppositely, more knowledge from the teacher will be passed to the student. The proposed AWL is straightforward to implement with salient explainability and is amenable to the mobile edge computing platform to enable true real-time processing speed and salient accuracy.
[0063] The key contributions of the present disclosure include: (1) different from the existing training practice for knowledge distillation models, our AWL-NanoDet considers the relationship 16 53600882 v1Docket No.10780003.00112 between the qualities of the teacher and the student during training and then selects adaptive weights to enable differential learning strategies that direct the learning process toward a favorable direction of knowledge distillation. It noticeably improves model performance over the native NanoDet. Meanwhile, our model remains the same lightweight structure and computation cost for online evaluation as the NanoDet; (2) the new AWL-NanoDet is thoroughly compared with the State-of-The-Art (SOTA) models in accuracy, computing load, and inference time. The performance gains achieved by AWL in the training stage and in the model evaluation stage are investigated and convincingly verified; and (3) the proposed AWL-NanoDet is implemented on Nvidia’s Jetson platform for edge computing and demonstrates automated, real-time, and accurate rail track component inspection.
[0064] Methodology
[0065] This section will first introduce the structure of the native NanoDet underlying the present work. Then the details of the proposed adaptively weighted loss (AWL) and models will be described.
[0066] Background: NanoDet
[0067] The structure of the native NanoDet 100, see github.com / RangiLyu / nanodet, is illustrated in FIG.1, which consists of three key modules:
[0068] 1. Backbone and Feature Pyramid Network (FPN) 102: It extracts multi-scale feature maps F out of input images 104.
[0069] 2. Assign Guidance Module (AGM) 106: It is also called the teacher model. As shown in FIG. 1, it takes feature maps F 108 and predicts bounding boxes (BA) and class (CA) 110 associated with the bounding boxes. Note that AGM 106 and the teacher model are used interchangeably in this disclosure. The prediction results by the teacher model will be compared to the bounding box ground truth (BG) and the class ground truth (CG) 112, yielding the teacher model loss LT 114. The teacher model results (BA and CA) 110 are the reassigned label (rather than BGand CG) for student model training. The method of generating reassigned labels by the teacher model is called Label Assignment Distillation (LAD).
[0070] 3. Prediction Head: It is also called the student model, as shown in FIG.1. It takes the same feature maps F 108 used by the teacher model and predicts the bounding boxes (BP) and their corresponding class labels (CP) 116. Different from the teacher model, the prediction made by the 17 53600882 v1Docket No.10780003.00112 student model will be compared with the reassigned labels produced by the teacher model, i.e., BA and CA (rather than the ground truth) 110 to calculate the student model loss (LS) 118, which is one of the foremost differences from traditional model training.
[0071] During training, the student model loss ^^ௌ118 and the teacher model loss ^^்114 will be combined as the total loss 120 for backpropagation, which distinguishes LAD-type models (such as NanoDet) from others. After the training process is completed, the teacher model will be removed, and only the student model will be retained for prediction. The following equations describe the steps of calculating the losses in NanoDet.
[0072] The feature map F 108, obtained from the backbone and FPN 102, is fed to both the teacher model (^^^) 106 and student model (^^^) 107 in NanoDet and ^^^are theclass output and the bounding box output made by the teacher model (^^^) while it is being trained. ^^^and ^^^are the classification output and bounding box output of the student model (^^^). The teacher loss (^^்) comprises two parts: the binary cross entropy ^^^^^^^^^^,^^ீ^ as the classification loss and the Generalized Intersection over Union ^^^^^^^^^^^^,^^ீ^ as the bounding box loss. Then ^^்is calculated as
[0075] (2)
[0076] where, again, ^^ீand ^^ீare the class and the bounding box ground truth. ^^^and ^^^are used as the reassigned labels for the student model. Thus, the student model could learn the true data patterns faster if the teacher model generates high-quality labels. To calculate the student loss (^^ௌ) in NanoDet, the classification output (^^^) and bounding box output (^^^) of the student model (^^^) are compared with the reassigned label (^^^,^^^), which is given by
[0077] (3)
[0078] (^^^) utilizes reassigned labels to facilitate backpropagation and learning rather than relying on the ground truth 18 53600882 v1Docket No.10780003.00112 alone. Additionally, the quality of the teacher model significantly impacts the performance ofNanoDet. Thus, the total loss for training the entire NanoDet would be ^^ௌ ^ ^^்.
[0079] Although architecturally similar to the FCOS model, the NanoDet as a highly lightweight model employing a teacher-student training strategy, which also makes it fundamentally different from the original FCOS model. Specifically, the structure of the prediction heads in NanoDet is significantly different from the larger FCOS model. In NanoDet, each prediction head is very simple and only consists of two depth-wise separable convolutions for classification and regression. As a comparison, FCOS has four groups of convolution layers with 256 channels and a 3*3 kernel size. In addition, the difference in network sizes leads to distinct convergence behavior. In FCOS, every pixel in the final feature map made by the prediction head will be directly compared to the ground truth, and the heads in FCOS are able to learn the data patterns directly and converge fast. Conversely, the parsimony of the prediction heads in NanoDet makes it ill-suited for direct learning and difficult to converge. Thus, NanoDet adopts a dynamic label assignment training technique to tackle this issue, which uses a teacher model consisting of four 3x3 convolutional kernels to guide the training of the student model. The rationale behind this model design is that the teacher model has a more complex architecture and is capable of better grasping the ground truth, making it valuable to enhance the student model's performance. However, the teacher model will not be present in online prediction.
[0080] Proposed Adaptively Weighted Loss (AWL) Method
[0081] This section will first discuss the issues associated with the native NanoDet and then describe the proposed AWL method in detail.
[0082] Issues with Native NanoDet
[0083] The teacher model in the native NanoDet was trained with the COCO dataset. See, "COCO," [Online]. Available: cocodataset.org / #home. However, in this work, we use a custom data set collected from railroad inspection, which is significantly different from COCO or ImageNet dataset. See, "ImageNet," [Online]. Available: image-net.org. Hence, the pre-trained teacher model needs to be retrained, which, however, causes another problem. As previously mentioned, the teacher model in NanoDet is intentionally kept small, comprising only four 3x3 convolutional kernels. As a result, during the initial stages of training, particularly the first few epochs, the teacher model may not be adequately trained and may incorrectly label the training 19 53600882 v1Docket No.10780003.00112 data for the student model. Furthermore, due to the nature of LAD, the student model training highly depends on the quality of the teacher model. Consequently, at the beginning of the training process, the risk of misguidance by the teacher model is notably high due to insufficient prior knowledge, compromising the quality of the student model and the learning process.
[0084] Within our custom dataset, there exists a possibility that the distribution of the output by the teacher model may deviate significantly from the actual distribution. FIG. 2 provides a visual representation of this problem. It compares the pixel-wise distribution of the detected objects that are output from the teacher model, student model, and ground truth on our dataset at the first training epoch. The x-axis denotes the pixel index / location of the image. When an object is detected at a specific pixel location, the object occurrence is counted and incremented. To generate FIG.2, we analyzed 64 example images and counted the number of all objects, including both spikes or clips, detected on each pixel. The y-axis represents the total number of objects occurrences at a specific pixel location, i.e., a histogram. In the figure, the blue line stands for the true distribution of the object location, while the orange and green lines correspond to the object location predicted by the teacher model and the student model. It is evident from the figure that both the teacher and student model outputs at the first epoch significantly deviate from the ground truth. In fact, the output of the teacher model is even worse. As a result, the student model could be misguided by the teacher model, which potentially jeopardizes the backpropagation of loss and slows down the training process. If the teacher model is trained excellently and can generate high- quality labels at the beginning, the student model training would benefit more.
[0085] Adaptively Weighted Loss (AWL)
[0086] In this context, we propose a new loss algorithm, termed the adaptively weighted loss (AWL), to address the issue. The underlying idea of AWL is to evaluate the quality of the teacher model and the student model in each training epoch and then assign a score / weight to the student loss to prevent the student model from learning from an unqualified teacher model while still allowing it to learn from a qualified teacher model when the teacher becomes qualified. Regarding the specific implementation, at the beginning of model training, the teacher loss should be weighted more while the student loss should be depreciated (as the reassigned labels by the teacher model are mostly incorrect), which, hence, accelerates the teacher model learning. Later, when the teacher is well-trained, and the accuracy of the reassigned labels is improved, the model will focus 20 53600882 v1Docket No.10780003.00112 more on student training. The whole pipeline of the proposed method 300 is presented in FIG.3, including three main parts: (1) in the first part, the algorithm will assess the quality of the teacher model (^^்^) 302 and student model (^^ௌ^) 304 in the current (ith) training epoch; (2) the second part will assign a quantitative score to the teacher model at the ^^௧^epoch (^^^) 306 and generate a weight (^^^) 308 for the student loss according to the current ^^்^302, ^^ௌ^304 and ^^^306; and (3) the last part will update the student loss ^^ௌ^310 using the weight ^^^308 obtained from the second part, yielding a weighted student loss ^^ᇱௌ^312. Eventually, the total loss for backpropagation duringtraining consists of the updated student’s loss and the teacher’s loss, which is given by ^^ᇱௌ^+ ^^்^. The details of each component will be discussed below.
[0087] Model Quality Assessment
[0088] The first step of our proposed method is to assess the quality of the teacher model ^^^and student model ^^^at the current (^^௧^) epoch. Their qualities will dictate the specific approaches used for assigning and updating the student loss ^^ௌ^. The model quality assessment requires apre-training process that uses the native NanoDet model without AWL. The pre-training produces the teacher and student models ^^^^^ ^^^^ and ^^^, from which their losses at each training epoch (^^^^^்^and ^^^^^ௌ^) are recorded. The mean and maximum of the teacher loss ^^ ^்^ೌ^ and ^^ ^்ೌ^ duringtraining process could also be obtained byprocess. ^^^^^is the total number of training epochs in pre-training. Again, all the parameter values in Eq. 4 are provided by the pre-trained model without AWL. In other words, they are treated as a kind 21 53600882 v1Docket No.10780003.00112 of prior knowledge prior to AWL-NanoDet training, while the prior knowledge is also generated by the dataset. Once ^^ ^்^ೌ^, ^^ௌ^^ೌ^ , ^^ ^்ೌ^ , and ^^ௌ^ೌ^ are available, they are used as thresholdingconstants to assess the of the teacher and the student model in AWL-enabled training, which is given by
[0091]
[0092] at the ^^௧^epoch will be determined, respectively, by their losses ^^்^and ^^ௌ^, relative to their mean values recorded during the pre-training process ^^ ^்^ೌ^and ^^ௌ^^ೌ^. When ^^்^ is smaller than ^^ ^்^ೌ^, viz.,^^்^ ൌ 0, it indicates that the teacher loss isand the teacher model is qualified forguiding the student model. Otherwise, ^^்^ is larger than ^^ ^்^ೌ^, i.e., ^^்^ ൌ 1, which means theteacher model is not qualified. The quality of the student model ^^^can be assigned in a similar manner.
[0093] Score and Weight Assignment
[0094] The reason for performing the quality assessment above is to differentially consider various scenarios that involve different combinations ^^^and ^^^qualities, and assign corresponding scores and weights to the student loss during model training. In Eq. Error! Reference source not found., there are four different combinations of ^^^and ^^^qualities, and the approaches to assign scores and weights for each of them are elucidated below.
[0095] Scenario i: ^^்^ ൌ 0 and ^^ௌ^ = 0
[0096] The first scenario considers a qualified teacher and a qualified student model, we propose ^^^will be assigned as 22 53600882 v1Docket No.10780003.00112 the updatedௌ^^^^can predict the ground truth well and generate high-quality ^^^, and ^^^outputs arealso close to reassigned labels from ^^^. Thus, the loss can used without any weighting for backpropagation.
[0099] Scenario ii: ^^்^ ൌ 1 and ^^ௌ^ = 1
[0100] The second scenario considers an unqualified ^^^and unqualified ^^^, and ^^^will be assigned asis away from being qualified. If ^^^is large, it means that ^^^is very unqualified. According to Eq. 5, for anunqualified teacher, ^^்^ െ ^^ ^்^ೌ^ is always greater than 0, and after normalization by ^^ ^்ೌ^, therange of ^^^is kept1. Once computed, ^^^is used as a penalty coefficient towith the teacher’s loss ^^்^in the second row of Eq. (7), where a constant hyperparameter number^^ is also employed as a gain factor for ^^^ ∗ ^^்^ for computing ^^^. Note that ^^^ is bounded between0 to 1. The first two rows of Eq. (7) also indicate that when ^^்^increases because of a poor match between the teacher and the ground truth, ^^^will rise, and ^^^will decrease correspondingly. It essentially implies that when ^^^is far from being qualified, it becomes pointless to let ^^^guide ^^^. Thus, a smaller ^^^is used to scale down ^^ௌ^and reduce its contribution to the total loss, which then allows for accelerating the teacher model's training (to make it qualified quickly).
[0103] Scenario iii: ^^்^ ൌ 0 and ^^ௌ^ = 123 53600882 v1Docket No.10780003.00112
[0104] In the third scenario, ^^^is qualified and ^^^is unqualified. We proposed ^^^will be assigned by(5) shows, when ^^^is^்^ೌ^, ^்^ೌ^ the absolute valueof ^^்^ െ ^^ ^்^ೌ^ guarantees The first two rows of Eq. (8) reveal that fora very low value of ^^்^, viz.,is highly qualified for generating accurate labels to guide the student ^^^, the score ^^^and weight ^^^are both significant. Note that ^^^ranges from 1 to 2, which will up the student loss ^^ௌ^increase its contribution to the total loss. As a result, thetraining process of ^^^is accelerated and ^^^learns faster toward ^^^.
[0107] Scenario iv: ^^்^ ൌ 1 and ^^ௌ^ = 0
[0108] The last scenario involves an unqualified ^^^and qualified ^^^, and its score and weight model are the same as Eq. (7). This is because when ^^^isand the teacher reassigns many wrong labels, the student loss is meaningless, and its contribution to the total loss should be reduced. Thus, the training process of the teacher model can be accelerated to make it qualified first.
[0109] After obtaining ^^ᇱௌ^in the four scenarios above, the total loss for training is simply ^^ᇱௌ^+ ^^்^. FIG. 4 illustrates the relationship between ^^்^.and ^^^. This nonlinear weight function and assignment algorithm is the foundation of the proposed AWL method. The blue and orange lines represent the relationship between ^^^and ^^்^for an unqualified and a qualified student model, respectively. In this figure, we choose ^^ ^்^ೌ^ൌ 0.8 as the criteria for qualifying a teacher model.A general pattern is that the weight declines nonlinearly as the teacher loss increases. We can also see that on the left side of the figure (i.e., when the teacher is qualified), a large value of ^^^is 24 53600882 v1Docket No.10780003.00112 assigned to the loss of the unqualified student (blue curve) to accelerate the learning process. However, for a qualified student (orange curve), the weight is kept at one, and the model is reduced to the native NanoDet. On the right side of the figure (i.e., the teacher model is not qualified, corresponding to teacher loss above 0.8), the student loss is meaningless whether the student is qualified or not. Therefore, a weight value lower than 1 is assigned to the student loss to accelerate the teacher model's training process preferentially.
[0110] Thus, the entire pipeline shown in FIG.3 can be summarized by the following equation the ^^௧^epoch.^^^^^^^^^^ a to assess ^^^^^^return ^^^at ^^௧^is the number of the total training epoch. Recall that ^^ௌ^is calculated by comparing the prediction from ^^^and ^^^, ^^ௌ^will be updated by multiplying ^^^. As mentioned above, this could prevent the from learning from the unqualified teacher model while still allowing it to learn from the qualified teacher model. Algorithm 1 is the pseudo-codes to illustrate the critical steps of the AWL-NanoDet model. Lines No. 2 and No. 3 assess the quality of the teacher ^^^and the student ^^^models at each training epoch using Eq. (5). Lines No.4 and 5, respectively, calculate the scores and the weights according to the qualities of both the teacher and the student models following Eq. (6) – Eq. (8). Lines 6 and 7 weigh the student’s loss and add the weighted student’s loss to the total loss using Eq. (9).
[0113] Algorithm 1, see FIG. 10, shows pseudo-codes of the proposed adaptively weighted loss (AWL) algorithm.
[0114] Results and Discussion
[0115] This section will present the dataset details and the results of the proposed AWL- NanoDet and the experiments. The model was trained using a dataset containing 1,895 images collected during railroad inspection. All models studied in this disclosure were trained on an NVIDIA RTX A5000 GPU and then test run on Nvidia’s AGX Orin, an edge computing device. The RTX A5000 has 8,192 CUDA cores, allowing 27.77 TFLOPS (FP 32) within its Ampere architecture, with a maximum power consumption of 230 W. On the other hand, the AGX Orin, 25 53600882 v1Docket No.10780003.00112 also based on the Ampere architecture, features a modest 2,048 CUDA cores and 5.3 TFLOPs (FP 32), with a 60W maximum power consumption. Accordingly, the computing power of AGX Orin is low relative to RTX A5000, making real-time detection more challenging. When deployed on AGX Origin, all models were implemented in the TensorRT form.
[0116] Dataset
[0117] The Track Component Imaging System (TCIS), see "TCIS," [Online]. Available: ensco.com / rail / track-component-imaging-system-tcis, was used for railroad safety inspections. Four TCIS cameras were firmly installed under the geometry car to ensure stable and high-quality image acquisition in the field. Each camera covered one side of a track, that is, the interior or the exterior sides of the left or the right track. The camera lens faced downward with a constant height relative to the ground, and the field of view was parallel to the tracks. As shown in FIG. 5, the camera system produces high-resolution grayscale images, and each image has a size of 416^416. In the present work, we focus on detecting and counting two essential components along the railroad, clips 502 and spikes 504, which are, respectively, illustrated by the blue and yellow bounding boxes in FIG. 5. Many missing clips 502 and spikes 504 during a railroad segment indicate a higher risk of railroad accidents. In total, 1,895 images were collected and split in a ratio of 80% : 20%, i.e., (1516 : 379) images, respectively, for model training and testing.
[0118] Because of the fixed distance and view angles between the TCIS and tracks under inspection, sophisticated data augmentation is unnecessary. In this work, only brightness augmentation is implemented to address the issues of varying light intensities during railroad inspection. FIG.6 shows the augmented images along with the original one, in which the value of brightness changes from -20% to +20%.
[0119] Performance Evaluation
[0120] Several commonly used metrics are employed to evaluate the performance of the proposed method and AWL-Nanodet model in terms of accuracy and speed, which are summarized as follows
[0121] 26 53600882 v1Docket No.10780003.00112 positive,the correctly detected objects (clips or spikes) relative to all detected and classified objects. The recall is the ratio of correctly detected objects to all the objects in the ground truth. quantify thean average derived from AP (Average Precision), and AP is the under-curve area of the P-R Curve (Precision-Recall curve). Each class will have its own AP. In Eq. (12), C is the total number of classes and is 2 in this disclosure, corresponding to clips and spikes. Consequently, mAP serves as a comprehensive measure of the AP values for all classes. In this disclosure, mAP@0.5 and mAP@0.5:0.95 are employed to assess our model. For the mAP@0.5 standard, only the predicted bounding boxes with an Intersection over Union (IoU) higher than 0.5 are considered positive predictions. For the mAP@0.5:0.95 standard, mAP values are calculated within the IoU range of 0.5 to 0.95 within an increment of 0.05, which are subsequently averaged as shown in Eq. (13) below.
[0126] (13)
[0127] is an essential measure for estimating the computational cost of models and defines the number of floating operations required for processing a single image with a given model.
[0128] In addition, in this disclosure, the inference time, the time for the model to process one image, is also used to evaluate its speed.
[0129] Results and Analysis
[0130] In this section, we will show the test results of the proposed AWL-NanoDet and assess its performance. The model will be first compared to several lightweight benchmark models 27 53600882 v1Docket No.10780003.00112 suitable for real-time image analysis, including the original native NanoDet, YOLOv4-tiny introduced by Alexey et al., see A. Bochkovskiy, C.-Y. Wang and H.-Y. M. Liao, "Yolov4: Optimal speed and accuracy of object detection," arXiv preprint arXiv:2004.10934, 2020, and EfficientDet-D0 proposed by Tan et al., see M. Tan, R. Pang and Q. V. Le, "Efficientdet: Scalable and efficient object detection," Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp.10781-10790, 2020, followed by comparison with the large SOTA models.
[0131] As described above, the native NanoDet is a highly compact model with a size of less than 10 Mb and a computational cost of 1.52G FLOPs. The YOLOv4-tiny has fewer backbone layers (29 convolutional layers) compared to the original YOLOv4 (53 convolutional layers), resulting in less computational cost (6.9G FLOPs). Furthermore, the primary distinction between EfficientDet-D0 and EfficientDet-Dx (where computational cost increases as the value of x increases) is that D0 has fewer FPN layers and channels (3 layers with 64 channels), while D5 or D7 uses more (7 layers with 228 channels and 8 layers with 384 channels, respectively). This makes EfficientDet-D0 more computationally efficient, whose computational cost is only 2.54G FLOPs. It can be observed that the FLOPs of the NanoDet, YOLOv4-tiny, and EfficientDet-D0 are on the same scale, making their comparison more relevant.
[0132] Table 2, FIG. 11, shows the quantitative results of the comparison in terms of precision, recall, mAP, FLOPS, and inference for different classes. It clearly shows that our AWL- NanoDet surpasses other models in the majority of performance metrics. In the spike class, NanoDet falls short in detecting many instances, achieving only a recall of 83.8%. EfficientDet- D0 performs even worse, with a mere 70.0% recall in the spike class. Similarly, YOLOv4-tiny still has a lower recall (88.7%) compared to our model. In this comparison, the recall of our AWL- NanoDet reaches 89.0%, resulting in a 5.2% improvement over the native NanoDet. Meanwhile, the precision of AWL-NanoDet reaches 93.2%, outperforming native NanoDet, YOLOv4-tiny, and EfficientDet-D0 by 2.9%, 5.8%, and 3.5%, respectively. Additionally, our model exceeds NanoDet by 3.2% in AP@0.5 and 10.1% in AP@0.5:0.9 (as shown in the columns of mAP@0.5 and mAP@0.5:0.9 in 2). When compared to YOLOv4-tiny and EfficientDet-D0, our model still maintains its advantages in the metrics of AP@0.5 and AP@0.5:0.9. Given the same model structures of the native NanoDet and our AWL-NanoDet, the performance enhancement of the latter in the spike class is salient and completely attributed to the AWL algorithm proposed in this 28 53600882 v1Docket No.10780003.00112 research. Interestingly, the improvement of classification for the clip class by AWL-NanoDet is minor compared to the native NanoDet. For example, recall of the former somewhat increases from 92.8% to 91.1% with only a slight decrease of 0.7% in precision. AWL-NanoDet exhibits 94.5% in AP@0.5 and 79.5 in AP@0.5:0.95, while NanoDet can also achieve 94.3% and 79.1%, respectively.
[0133] The difference in performance improvement between the spike class and the clip class is due to their different appearance and difficulty of detection. The small size of the spike class in our dataset makes them like the background and nearly indistinguishable from the ballast gravel. Consequently, the teacher model performs poor during the initial training process and may detect wrong objects. Our approach evaluates the teacher model and assigns a low weight to the student loss, preventing the student model from being misguided by the poor-performing teacher model and prioritizing the training of the teacher model first. However, the native NanoDet does not consider the teacher model's quality, which may lead to erroneous guidance for the student model and inadequate performance for difficult-to-detect classes, i.e., spikes. Thus, the advantage of the AWL algorithm can be fully utilized to effectively improve the model training and prediction accuracy. On the other hand, the clips look very different from the background and can be easily detected. For example, even by YOLOv4-tiny and EfficientDet-D0, a relatively good AP@0.5, viz., 91.0% and 89.2%, respectively, can be achieved. Therefore, even in the early epochs of the training process, the teacher model is already able to extract useful feature information and is qualified to guide the student model in the clip class, and the improvement by AWL is only marginal.
[0134] However, in general, AWL-NanoDet appreciably improves the overall performance with a 93.4% mAP@0.5 and a 69.9% mAP@0.5:0.95, which is higher than the NanoDet's performance with a 91.7% and a 64.9%. It is also prominently superior to the other two models, YOLOv4-tiny and EfficientDet-D0.
[0135] FIG. 7 compares the native NanoDet and AWL-NanoDet using the Precision-Recall (P-R) curves, which are, respectively, colored in orange and blue. We can see that including AWL in training notably improves the performance over the native NanoDet. The P-R curve of the spike class is shown in FIG.7 at (a), and the precision of the native NanoDet drops dramatically when its recall increases. In contrast, the precision drop in our AWL-NanoDet model is considerably 29 53600882 v1Docket No.10780003.00112 delayed and mild. Even if the recall is close to 1, the precision remains above 0.5. The salient performance agrees with the observation above, which confirms that our AWL method is especially useful for tackling hard-to-detect objects / classes. FIG.7 at (b) is the P-R curve of the clip class. Similarly, AWL-NanoDet slightly outperforms the NanoDet, and our method allows precision to decay more slowly when recall rises.
[0136] Since AWL notably boosts the detection performance, we next compare AWL- NanoDet with the State-of-The-Art (SOTA) object detection models of large sizes for a comprehensive evaluation, including YOLOv5-L and YOLOv5-S in terms of accuracy, FLOPs, and inference time. Usually, these models are much large compared to lightweight models above and require a large amount of data for training, while our dataset contains only 1,895 images. To address this issue, the initial weights of these models were pre-trained on the COCO dataset and retrained with our custom rail track dataset. It should be noted that large models like YOLOv5-L (or even YOLOv5-S) are capable of learning more complex data patterns and are generally more accurate than lightweight models. Therefore, their comparison in detection accuracy is somewhat unfavorable to AWL-NanoDet. Nevertheless, the AWL-NanoDet could achieve almost the same level of performance or even exceed these SOTA models for the clip class because, as discussed above, clips look more different from the background, and even the lightweight model can detect them well. For instance, the precision, recall, and AP@0.5:0.95 of AWL-NanoDet are 96.4%, 92.8%, and 79.5% (as shown in the mAP@0.5:0.9 columns in Table 2, see FIG.11), respectively, which are close to 97.8%, 97.0%, and 75.2% obtained by YOLOv5-L and 95.5%, 94.3% and 74.3% by YOLOv5-S. However, the spike class shows a more noticeable performance gap between AWL-NanoDet and the SOTA models, which, again, is caused by the close resemblance between the spikes and the rail track background. The SOTA models are more capable of learning complex features for distinction under this situation. However, AWL-NanoDet and SOTA models still stay at a similar level. For example, the precision, recall, and AP@0.5:0.95 of AWL-NanoDet are 93.2%, 89.0%, and 59.9%, respectively, which are 96.2%, 91.1%, and 65.0% for YOLOv5-L and 96.3%, 92.2%, and 64.7% for YOLOv5-S. For all classes, the mAP@0.5:0.95 of AWL- NanoDet, YOLOv5-L, and YOLOv5-S are 69.9%, 70.0%, and 69.5%, respectively. The indices of performance shown in Table 3, see FIG.12, are close to each other, confirming that the proposed 30 53600882 v1Docket No.10780003.00112 AWL approach could significantly improve the accuracy of the lightweight NanoDet model to a level comparable to SOTA models.
[0137] The primary driver of developing lightweight models like our AWL-NanoDet is its compact sizes and ultra-fast speed, enabling real-time processing on small-factor edge devices with limited computing power, such as the AGX Orin used in this work. The inference speed of AWL-NanoDet overwhelmingly surpasses all benchmark models presented in this disclosure. Table 2, see FIG.11, and Table 3, see FIG.12, that the average inference time for YOLOv5-L and YOLOv5s is 140.9 ms and 50.2 ms, and even YOLOv4-tiny needs 37.3 ms. Therefore, they might not be able to meet the requirements of real-time processing. On the other hand, the inference time of our AWL-NanoDet is only 4.2 ms, which is almost 34X faster than YOLOv5-L. Although the inference time of Efficientdet-D0 is comparable to our model, its mAP metrics are dramatically lower. The results generated from our model are shown in FIG. 8. By leveraging the AWL algorithm, our model can handle complex conditions, such as low light or clutter. In FIG.8 at (a), despite the relatively poor light condition, the accurate detection of both the spikes and clips is clearly observed. In FIG.8 at (b), while the arrangement of objects is dense, the model is still able to detect all the components.
[0138] Convergence Study
[0139] In this section, we will analyze the effect of the proposed AWL algorithm and the convergence of the weight of the student loss ^^ in the training process FIG.9 at (a)-(c) illustrate the changes of ^^ with the training epoch when 10, 25, and 100 epochs in total are used to train AWL-NanoDet. A similar trend is observed in all three subfigures. That is, the teacher model is unqualified at the beginning of the training process, and the guidance offered by the teacher is less useful. Thus, the learning of the teacher model should be accelerated, and that of the student should be given less priority because the reassigned labels from the teacher are not sufficiently accurate. This can be accomplished by assigning a low weight to the student loss at the first few epochs. FIG.9 at (a)-(c) all show that ^^ of the proposed model (orange line) is relatively lower than the native NanoDet (blue line representing a constant of 1 at all epochs).
[0140] After the teacher model becomes qualified, the student learning process needs to be expedited by AWL. As we can see, the weight ^^ grows above one after the first few epochs, which weighs the student loss more in the total loss, accelerating the training of the student model. After 31 53600882 v1Docket No.10780003.00112 the teacher and student models are both qualified, the weights would return to the constant 1, which is equivalent to the native NanoDet model, where student and teacher losses are treated equally. Our AWL algorithm does allow all the weights in the subfigures of FIG. 9 to converge to 1 eventually.
[0141] It is worth noting that although FIG.9 (a)-(c) exhibit a similar trend, their convergence rates, i.e., the number of epochs required to reach weight convergence, differ from each other. This is due to the use of a different number of epochs, namely 10, 25, and 100 in their corresponding pre-training processes. Recall that our AWL algorithm in Section 2 needs a pre-training step to extract the statistics of the loss (^^ ^்^ೌ^, ^^ௌ^^ೌ^, ^^ ^்ೌ^, and ^^ௌ^ೌ^), and the number of the pre-training epochs is set equal to ൌ ^^^^^). In FIG. 9 at (b) and(c), more epochs are used to our essentially increases the criterion of qualifying a teacher and a student model since the mean of the teacher and student loss in the pre-training decreases with more epochs. As a result, more epochs are needed to converge the weight. For each experiment using a different number of epochs, 15 runes are performed. The shade and the curve in FIG.9, respectively, represent the range and the mean value of the weights in these 15 runs. It clearly shows that the weight ^^^in the training process with fewer epochs and a lower criterion of the teacher and student qualification converges faster than that with more epochs and a higher teacher criterion. Therefore, its convergence is notably delayed when the number of training epochs increases, as shown in FIG.9 at (a), (b), and (c).
[0142] This disclosure provides a new lightweight computer vision model, AWL-NanoDet, for accurate, real-time rail track inspection on edge devices. The model introduces a new AWL strategy to modify the teacher-student guidance mechanism on-the-fly during the training process. Its foremost innovation is to examine the qualities of the teacher and student models during training and then adjust the weight of the student loss contribution in the total loss. Specifically, the AWL algorithm first evaluates the teacher quality by comparing its loss to a pre-training process and then assigns a weight to the student loss to tune its contribution, directing the learning process toward optimal knowledge distillation and guidance. AWL considering various combinations of the teacher and the student qualities is proposed. In short, when the teacher model has poor quality, the weight of the student loss should be assigned a low value to preferentially accelerate teacher model training. 32 53600882 v1Docket No.10780003.00112
[0143] The new AWL-NanoDet model is then tested with a custom dataset collected during railroad inspection, and a comprehensive comparison is made with respect to existing lightweight models (native NanoDet, YOLOv4-tiny, EfficientDet-D) and large SOTA models (YOLOv5-L and YOLOv5-S) in precision, recall, mAP, FLOPS, and inference speed. Our model demonstrates a precision of 94.8%, recall of 90.9%, mAP@0.5 of 93.4, and mAP@0.5:0.95 of 69.9 across all classes, outperforming the other lightweight models. AWL is especially effective in detecting classes, e.g., spikes that are generally difficult to identify in the complex background. Furthermore, although NanoDet has an extremely compact model structure, AWL enables the excellent potential to boost detection accuracy to a level comparable to the large SOTA models.
[0144] In terms of computing efficiency, same as native NanoDet, our AWL-NanoDet has a very low FLOPS of only 1.52 G that is dramatically smaller than the YOLOv5-L (109.1 G) and the YOLOv5-s (15.8 G). When tested on Nvidia’s Orin device, it achieves the fastest inference time of 4.2 ms, holding great promise for real-time image processing and inspection. The convergence of the weight in AWL is also thoroughly investigated for different numbers of training epochs. In all cases, AWL automatically adjusts the weight during model training without human intervention, and a similar trend is observed, that is, the weight is significantly less than one initially to prioritize the teacher model learning and then reaches beyond 1 to accelerate the student model training after the teacher becomes qualified. Eventually, the weight converges to one, i.e., the native NanoDet, where the contributions of the teacher and the student losses are equal.
[0145] In the field of railway infrastructure maintenance, timely and accurate detection of component anomalies is crucial for safety and efficiency. This disclosure presents the Cascade Region-based convolutional neural network with Predefined Proposal Templates (CR-PPT), an innovative method for railroad components inspection in complex railway infrastructure using edge-computing devices.
[0146] Unlike previous systems, CR-PPT employs a series of predefined templates that enable it to detect both the presence and missing elements within various fastening systems. Our experimental analysis pinpoints the most effective network configurations for CR-PPT. Furthermore, the current disclosure examines CR-PPT’s proficiency in zero-shot learning and fine- tuning, highlighting its adaptability to new fastening systems. We have developed an optimized 33 53600882 v1Docket No.10780003.00112 inference pipeline on NVIDIA Jetson AGX Orin, significantly enhancing its applicability for railway inspection practices. Field blind tests validate the model’s high precision and efficiency, greatly reducing the time and labor required for inspections. The findings highlight CR-PPT’s potential as an efficient and robust tool for track health assessment, marking a notable progression in the integration of AI and computer vision in rail track inspection.
[0147] As of 2024, the extensive railroad network in the United States, comprising 136,729 miles of freight railways and 21,124 miles of passenger railways, forms a crucial part of the nation’s transportation infrastructure (Robinson et al., 2023). This network is vital in transporting massive quantities of goods and numerous passengers daily.
[0148] The Federal Railroad Administration (FRA) underscores the importance of track inspections for railway safety and efficiency (FRA, 2018). Key track components like spikes, bolts, and clips are particularly prone to damage from loading and environmental factors. Not identifying these defects can lead to serious accidents and significant consequences (Yunpeng Wu et al., 2023). Traditional track inspections primarily rely on manual examinations or specialized inspection vehicles. While effective, these methods are often time-consuming, susceptible to human error, and require considerable expertise for accurate data interpretation. This underscores the need for meticulous maintenance and robust safety protocols within the railroad network (F. Guo et al., 2023; YunpengWu et al., 2022). Deep learning, a branch of machine learning, has seen significant advancements recently (Alam et al., 2020; Rafiei et al., 2022). These advancements open new possibilities for enhancing the efficiency and accuracy of structural health monitoring (Amezquita-Sanchez & Adeli, 2019; Javadinasab Hormozabad et al., 2021). They offer powerful tools for assessing the condition of critical structures like buildings (Perez-Ramirez et al., 2019), roads (E. Yang et al., 2023; A. A. Zhang et al., 2022), bridges (Chun et al., 2022; G.-Q. Zhang et al., 2022), railways (F. Guo, Qian, Rizos, et al., 2021; Yunpeng Wu et al., 2021), and marine structures (Pezeshki et al., 2023). By precisely identifying structural damage, these technologies play an essential role in improving the longevity and safety of such infrastructure.
[0149] Recent research has focused on utilizing advanced deep learning models for the detection and inspection of railway track components. For instance, Yilmazer and Karakose (2022, 2023) investigated the effectiveness of You Only Look Once version 5 (YOLOv5) and Mask Region-based Convolutional Neural Network (MaskR-CNN) in identifying normal and abnormal 34 53600882 v1Docket No.10780003.00112 clips. Tang and Qian (2024) focused on the high-performance deployment of YOLOv8 for fastener inspection using parallel and concurrent computation to significantly boost model inference speed. Jianwei Liu, Liu, et al. (2023) employed a Single Shot multibox Detector (SSD) model to quickly locate fastener regions, followed by a modified Faster R-CNNfor assessing the condition of detected fasteners. These methods excel in identifying abnormal components by learning from annotated track images, distinguishing between normal and defective states, such as broken, displaced, or missing clips, as well as loose spikes or bolts. They are particularly effective in modern railway systems as shown in FIG.13 at that utilize concrete ties and constantly use clips as fastening methods.
[0150] However, the US railway system, with its long history and legacy infrastructure, presents unique challenges for automatic inspection. A major issue is the detection of missing spikes and bolts. As shown in FIG.13 at b, in the United States, many rails are still mounted on wooden ties using a variety of fastening systems, and one rail may use multiple combinations of these fastening systems. These systems employ different types of baseplate and clips and offer varying numbers of holes for spike or bolt installation.
[0151] To reduce costs, not all installation holes are used, which complicates the task of distinguishing between unused holes and missing spikes or bolts. This complexity necessitates the development of a more sophisticated automatic track inspection system, specifically tailored to meet the unique requirements of US railway track inspection.
[0152] To address the specific needs of the US railway system, we have developed a novel approach, Cascade Region-based convolutional neural network with Predefined Proposal Templates (CR-PPT). This method evolved from our previous research (Tang et al., 2024) and is specifically tailored for detecting missing track components such as clips, spikes, and bolts. As illustrated in FIG.14 at a, CR-PPT operates by refining predefined proposal boxes using a cascade box head. This matching mechanism not only facilitates the detection of rail components but also classifies different types of fastening systems. Once the type of current fastening system currently is identified, missing components are detected based on predefined criteria tailored to that specific system. This tailored approach is particularly effective in enhancing track safety and maintenance by ensuring that all essential components are present and correctly installed. The key contributions of this innovative method are outlined as follows: 35 53600882 v1Docket No.10780003.00112
[0153] 1. Developed the CR-PPT method, specifically engineered for railroad track inspection, marking an advancement from previous methods that struggled with differentiating between missing spikes and bolts and their unused installation holes.
[0154] 2. Detailed exploration of CR-PPT’s adaptability to novel fastening systems that are not included in the training dataset. Common network transfer learning methods like zero-shot, one-shot, few-shot, and incremental learning are explored to evaluate the model’s flexibility and efficiency in handling new and diverse data.
[0155] 3. Development of a streamlined inference pipeline on the NVIDIA Jetson AGX Orin, integrating C++, Torch-Script, and oneTBB to optimize speed and enhance performance. This deployment on an edge device is aimed at improving the portability and affordability of railway inspection systems. This practice significantly reduces the operational burdens associated with railroad inspections, making the process more efficient and cost-effective.
[0156] 4. Comprehensive field blind testing for missing component detection, revealing the model’s high accuracy and significantly reducing labor and time expenditure for actual railway inspection tasks.
[0157] The remainder of the current disclosure is organized as follows: The Related Works section provides a comprehensive review of prior research focus on the detection and inspection of rail components, as well as an overview of the Cascade R-CNN and template matching, which are foundational to the proposed CR-PPT model. The Methodology section delves into the detailed architecture of the CR-PPT, outlining its unique features and the theoretical underpinnings that guide its design. In the Experiments section, we extensively discuss the training, evaluation, deployment, and field testing of the CR-PPT and its effectiveness.
[0158] RELATEDWORKS
[0159] Railroad components detection and inspection algorithms
[0160] Over the past decade, significant advancements have been made in systems for detecting and inspecting railway track components. This section comprehensively reviews the ongoing research in this field, and finally concludes the necessity of this disclosure.
[0161] Y. Li et al. (2014) introduced a mathematical approach for locating ties, tie plates, and anchors. Several deep learning-based models have been trained to enhance the accuracy of detecting railway track components. F. Guo, Qian, and Shi (2021) developed an enhanced 36 53600882 v1Docket No.10780003.00112 YOLOv4 network tailored for this purpose, experimenting with various activation functions to optimize performance. More recently, Gosiewska et al. (2023) investigated the capabilities of the YOLOv5 model in detecting rail components, utilizing a training dataset characterized by a low data volume. In a separate advancement, J. Guo et al. (2023) employed a knowledge distillation approach to develop an adaptively weighted loss function, which refined the accuracy of a compact student model in detecting these components. Additionally, Qi et al. (2020) introduced the MYOLOv3-Tiny network, specifically designed for clip detection. This model integrates depth- wise separable convolution and linear bottlenecks with inverted residuals to enhance detection efficiency. Despite these advancements, these methods primarily focus on the localization of rail components, and they still require subsequent postprocessing steps to evaluate the condition of the detected components and to identify any defects.
[0162] In addition to track component localization, research has increasingly focused on recognizing defective components, particularly worn and missing fasteners. Feng et al. (2014) developed an image processing method for locating and detecting clip defects without deep learning, achieving high accuracy. Deep learning-based models for detecting abnormalities typically fall into three categories: object detection, classification, and segmentation.
[0163] Object detection methods employ detectors such as YOLO or R-CNN and their enhanced variants to localize and classify railway components simultaneously. These methods are trained on railway images featuring both normal and defective components. Cao et al. (2023) enhanced the YOLOv5 model by integrating an attention block and a transformer-based prediction head, improving detection accuracy and speed. This model can differentiate between normal, displaced, and missing clips.
[0164] Similarly, X. Li et al. (2023) added an attention module, a weighted bidirectional feature pyramid network, and K-means++ refined anchor boxes to the YOLOv5s for increased efficiency. J. Hu et al. (2022) upgraded the YOLOX-Nano model with a Coordinate Attention (CA) attention mechanism and adaptive spatial feature fusion to detect normal, displaced, and broken clips. Jianwei Liu, Qiu, et al. (2023) proposed the Op-YOLOv4-Tiny network for rail components inspection, which enhances the YOLOv4-Tiny by substituting its Cross Stage Partial Block (CSPBlock) with ResBlock-N. Fu et al. (2022) applied YOLOv4 to detect normal and abnormal clips, showing that using the MobileNet backbone instead of the original Darknet 37 53600882 v1Docket No.10780003.00112 improves accuracy and speed. Beyond model enhancement, Xiao et al. (2023) implemented the YOLOv2 model on a Field Programmable Gate Array (FPGA) chip, achieving substantial efficiency gains in embedded systems. However, due to the presence of defective components being much less than normal components, object detection methods often struggle with inaccuracy due to dataset imbalance. To address this issue, Y. Gao et al. (2023) proposed a self-driven loss function to balance the weights between positive and negative samples.
[0165] To address the challenges inherent in object detection methods, classification approaches separate the processes of component localization and evaluation. Initially, these methods identify component regions using object detectors without immediate classification. Subsequently, additional computations are performed to assess the condition of the detected components. This two-stage process typically offers improved accuracy but requires more time for inference and greater computational resources. For example, Junbo Liu et al. (2022) and Aydın et al. (2022) combined non-deep learning techniques for initial localization with Convolutional Neural Networks (CNNs) for subsequent classification of fasteners. This approach allows the classifier to evaluate one component at a time, facilitating training on a more controlled and balanced dataset. While using object detectors for fastener localization gains more popularity, Aydin et al. (2023) utilized YOLOv4-Tiny for rapid localization of fasteners, followed by a lightweight CNN to classify defects. Jianwei Liu et al. (2021) employed a network akin to R-CNN, termed MSF-DDN, for locating fastener regions, and then introduced a Region Classification Network and Decision Tree (RCN&DT) algorithm for distinguishing between normal and abnormal fasteners. Bai et al. (2021) integrated detection and classification by using a modified Faster R-CNN to detect complete, broken, and missing fasteners, followed by a Support Vector Date Description (SVDD) classifier for assessing normal and displaced fasteners. In a similar vein, Zhuang et al. (2022) employed the Squeeze and Excitation participated YOLOv3 (SE-YOLOv3) model for the initial detection of rail components, followed by a Domain-Logic-based Hybrid Model (DLHM) to refine the classification results.
[0166] Given the complexity of fastener evaluation relative to its localization, Qiu et al. (2023) focused on fastener classification using a CNN model trained with center-triplet loss and a SupportVector Machine (SVM) classifier. Y. Liu and Song (2022) also deployed Shuffle-RepNet for fastener classification, leveraging a reparameterization method to reduce CNN parameters and 38 53600882 v1Docket No.10780003.00112 enhance inference speed. Additionally, Su et al. (2022) applied an image inpainting network to generate synthetic fastener images for training a more robust classifier.
[0167] Besides detection and classification, instance segmentation methods provide pixel-wise shape data of track components. This approach delineates the precise boundaries of each component, enabling fine-grained analysis of their conditions. Techniques like Mask R-CNN (He et al., 2017) or You Only Look At Coefficients (YOLACT) (Bolya et al., 2019) are often used in this context. Shape analyzing is computationally intensive, and segmentation methods are particularly effective in capturing subtle anomalies and variations in the shape and texture of railway components, which might be missed by object detection or classification methods alone. Notably, F. Guo, Qian,Wu, et al. (2021) and Wei et al. (2023) have implemented YOLACT for detecting and segmenting track components like spikes, bolts, and clips. Additionally, Su et al. (2023) introduced a shape guided network specifically for clip segmentation, enhancing the accuracy and utility of instance segmentation in railway maintenance and safety assessments.
[0168] The application of language-based models in rail track inspection is a promising area of research, especially with the recent advancements in large language models like Generative Pre- trained Transformer (GPT; Bubeck et al., 2023) and multi-modal models such as Contrastive Language-Image Pre-training (CLIP; Radford et al., 2021), which integrate textual and visual data. Wei et al. (2022) made an early effort in this domain by combining R-CNN and Long Short-Term Memory (LSTM), marking a significant step toward using language-based models to identify abnormal clips in rail tracks. However, the model of Wei et al. (2022) faced limitations due to their small scale and the lack of support from more advanced large language models, which restricted their effectiveness compared to purely image-based models.
[0169] Alternative approaches to image-based detection have been explored using non-visual technologies such as vibrometers and electromagnetic inductors for identifying defective clips. These methods rely on the distinct frequency and signal characteristics differentiating normal and abnormal fasteners. Chandran et al. (2022), Chen et al. (2022), C. Yang et al. (2023), and Yin et al. (2022) have all contributed to this field, demonstrating the effectiveness of these technologies in detecting anomalies in railway components.
[0170] This review of current research highlights the absence of a specialized rail inspection framework tailored for the US railway systems. Many studies effectively distinguish between 39 53600882 v1Docket No.10780003.00112 normal, broken, and missing clips and can address loose spikes and bolts. However, none are able to differentiate between a missing spike or bolt and a mere installation hole.
[0171] CR-PPT is an advanced model evolved from our prior research (Tang et al., 2024). Unlike our earlier work which utilized template matching during post-processing, CR-PPT incorporates this matching mechanism directly into the model framework. This integration significantly enhances the model’s effectiveness in detecting missing components.
[0172] Cascade R-CNN
[0173] The proposed CR-PPT is an extension of the well established Cascade R-CNN (Cai & Vasconcelos, 2018). Cascade R-CNN builds upon the Faster R-CNN (Ren et al., 2016) architecture, aiming to enhance object detection accuracy.
[0174] As illustrated in FIG.14 at b, Cascade R-CNN and its predecessor, Faster R-CNN, are typically two-stage detectors. The first stage of Faster R-CNN consists of a backbone and a Region Proposal Network (RPN). When an image is fed into this stage, the system generates a set of bounding boxes representing object proposals. These proposals are potential object-containing regions, though they have not yet been classified. In the second stage, the flexibility of Faster R- CNN becomes apparent, as it can incorporate a wide array of Region of Interest (RoI) heads tailored for distinct tasks. These tasks might include box refinement and classification (through a box head), object segmentation (via a mask head; He et al., 2017), as well as additional functionalities like keypoint detection (He et al., 2017) and tracking (Zhou et al, 2022). This modular framework makes the Faster R-CNN an ideal platform for modification and extension.
[0175] Faster R-CNN and Cascade R-CNN are two-stage detectors that incorporate an additional box refinement step, which is absent in one-stage detectors like YOLO. This step typically enhances accuracy but requires more computational resources. Cascade R-CNN emphasizes this approach by integrating a multi-stage box head into its framework, specifically designed to further boost accuracy.
[0176] Typically consisting of three stages, the box head of Cascade R-CNN refines the predictions made by the previous stage. This cascading refinement process progressively enhances the precision of object localization. As a result, it allows for increasingly accurate adjustments to bounding boxes at every stage, thereby achieving more precise and reliable object detection outcomes. However, this increased precision comes at the cost of additional computation. 40 53600882 v1Docket No.10780003.00112
[0177] In the realm of rail component inspection, one-stage detectors such as YOLO are increasingly favored for their rapid processing speeds and satisfactory accuracy (Gosiewska et al., 2023; Jianwei Liu, Qiu, et al., 2023; Qi et al., 2020). Consequently, despite the high precision of Cascade R-CNN, it is often relegated to a comparative role in academics due to its longer latency (Bai et al., 2021; Y. Gao et al., 2023; Yunpeng Wu et al., 2023).
[0178] The multi-stage box head, a core component of the proposed CR-PPT model, inevitably inherits the slow processing speed characteristic of Cascade R-CNN. To address this challenge, we have implemented the same inference pipeline as detailed in another study (Tang & Qian, 2024). This pipeline leverages parallel and concurrent computation techniques to enhance the model’s inference speed. We utilize these computational strategies to reduce inference latency, thereby optimizing the model for faster performance.
[0179] Template matching
[0180] In this disclosure, the predefined proposal boxes serve as core elements of Predefined Proposal Templates (PPT), which are utilized to detect components and classify fastening systems by leveraging human knowledge. This mechanism closely resembles template matching, a broad concept in the field of computer vision and image processing, and its basic definition is searching and finding the location of a template in an image (Brunelli, 2009).
[0181] For instance, A. Zhang et al. (2013) utilized Gaussian curves to design Two- Dimensional (2D) templates specifically for detecting pavement cracks, exemplifying the technique’s application in extracting targeted features from images. In Three-Dimensional (3D) reconstruction, template matching is crucial for creating detailed 3D scenes by aligning multiple 2D images. S. Hu et al. (2020) addressed the selection of optimal images for texture mapping in this context, highlighting the importance of precision and quality. Edge detectors, like the Canny and Sobel methods, play an essential role in template matching by identifying image boundaries and contours (Jähne, 2005). Advances in deep learning, particularly through CNNs, have significantly enhanced template-matching capabilities. CNNs, with their layered convolutional processes, excel in detecting and learning patterns at various scales and complexities, offering a robust tool for feature detection in dynamic environments (R. Zhang et al., 2018).
[0182] METHODOLOGY
[0183] Overall approach 41 53600882 v1Docket No.10780003.00112
[0184] We, supra, highlighted the difficulties in detecting missing components for railways on wooden ties. To overcome these challenges, we introduced the CRPPT method, which leverages the detection and inspection characteristics of rail components. As shown in FIG.13 at b, despite the variety of types available, each type has a specific layout. This uniformity implies that if a fastening system is classified, we can implement specific criteria and rules tailored to the identified fastening system to accurately detect missing components.
[0185] One natural thought is to develop a classifier for baseplates or fastening systems. This classifier would categorize each system. While conceptually sound, this method faces significant challenges. First, compiling a comprehensive dataset encompassing all varieties of fastening systems present in the market is a daunting task. Furthermore, the introduction of new fastening systems in the field presents another hurdle. Each new fastening system necessitates retraining the classifier, which is a labor-intensive process involving extensive data labeling and training.
[0186] In this disclosure, CR-PPT utilizes predefined proposal templates for rail components detection, fastening system classification and missing components identification. As shown in FIG. 14 at a, this process begins with the detection of the rail and a central tie. Once they are identified, CRPPT proceeds to refine the predefined proposal boxes of rail components. In these two steps, CR-PPT disables the RPN and directly inputs predefined proposal boxes into the box head for the detection of the rail, tie, and other components. From Steps 3 to 4, CR-PPT identifies the most appropriate predefined proposal template that matches the observed components. Through this process, the rail components are detected, and the specific type of fastening system currently in use is classified. In Step 5, any missing components are detected using the criteria specified by the chosen predefined proposal template.
[0187] Network output classes design
[0188] As shown in Step 4 of FIG.14 at a, the proposed model is designed to detect six classes of track components: rails, ties, clips, missing clips, spikes and bolts, and installation holes for spikes and bolts. For efficiency, the model categorizes various clip types into a single group. This simplifies the annotation process and reduces algorithm complexity. Similarly, spikes and bolts are combined into one category due to their similar appearance and function as baseplate anchors.
[0189] Similar to previous studies, the approach to inspecting clips in CR-PPT is straightforward, involving the labeling and detection of both normal and missing clips. However, 42 53600882 v1Docket No.10780003.00112 to effectively infer the absence of spikes and bolts, CR-PPT must also detect the installation holes for spikes and bolts.
[0190] CR-PPT detects and segments rails and ties to determine their precise location, aiding in the accurate location and scaling of fastening systems. Note that this model does not segment other components such as clips because the detection of missing components does not depend on shape analysis.
[0191] Predefined proposal templates
[0192] The proposed CR-PPT method enhances the efficiency of rail component detection and inspection using predefined proposal templates. These templates are equipped with predefined proposal boxes for component location and fastening system classification, as well as corresponding criteria for the identification of missing components.
[0193] Before employing the CR-PPT in rail system inspections, it is crucial to incorporate a comprehensive range of templates. However, given the extensive variety and ongoing evolution of fastening system designs, covering all existing baseplate types found in the field is a daunting task.
[0194] Fortunately, the flexibility of our approach is a significant advantage. It allows for the integration of new design types into our template library, ensuring that the system remains current and effective in inspecting a variety of fastening systems.
[0195] FIG.15 at a displays the predefined boxes of different templates, and Table 4, see FIG. 16, lists the corresponding criteria of each template type. For this research, we have compiled a collection of 15 distinct template types. Of these, the first 13 types have been incorporated into our training dataset. The remaining two types, which are novel, serve to evaluate the network efficacy in adapting to new fastening systems. It is important to note that some template types include multiple sets of predefined boxes. This occurs when various fastening systems, sharing similar layouts and the same criteria, are grouped under the same template type. Additionally, not all fastening systems exhibit lateral symmetry, such as the types 1, 6, and 10 fastening systems in FIG.15 at a. In real-world applications, they may be oriented at 0 or 180 degrees. To accommodate this variability, asymmetric fastening systems are represented by two sets of predefined boxes, one for a 0-degree orientation and another for a 180-degree orientation. 43 53600882 v1Docket No. 10780003.00112
[0196] FIG. 15 at b outlines the development process of predefined proposal boxes for a fastening system. The initial step involves selecting images with a centrally located tie, as the centered fastening system enhances the effectiveness of the template used in component inspection. Next, the rail, tie, and components on the central tie are labeled. The subsequent step is box scaling, which is essential due to the variable positions of the fastening systemin each image and the diverse heights from which these rail track images might be captured. This scaling necessitates recentering and normalizing each box to compensate for perspective and distance variations in the images. The scaling process of each predefined proposal box is as follows: left and bottom-right coordinatesof an original box, respectively, with the zero point at the top left of the image. [^^^^, ^^^^] denotes the insertion center of the rail and the center tie. The variable ^^ stands for the average pixel width of the rail. The rationale for using ^^ as a normalization factor lies in the constancy of the rail width, which ensures ^^ remains a consistent reference across different image scales.
[0199] The criteria established for each template in the CR-PPT system are essential for inferring missing components in rail fastening systems. These criteria, detailed in Table 4, see FIG. 16, define the acceptable installation configurations for each template type. The criteria system uses numerical coding to represent different states of component installation: “1” represents an installation hole for a spike or bolt, “2” indicates the presence of a spike or bolt, and “3” denotes the presence of a clip. It is important to note that these criteria codes apply to both the 0-degree and 180-degree versions of the predefined proposal boxes. However, the arrangement of these components within the template criteria differs based on the orientation. For the 0-degree version, the components are arranged from left to right and then top to bottom, following a natural, horizontal reading order. For the 180-degree version, the arrangement is from right to left and then bottom to top. 44 53600882 v1Docket No.10780003.00112
[0200] Predefined proposal templates like types 3, 5, and 12 have a single criterion and do not include the number “1.” This implies that they do not allow for any missing components, as the presence of “1” in a criterion permits the installation of an additional spike or bolt for enhanced security. For instance, the detection [2, 2, 2, 2, 2] would still be acceptable if the criterion is [2, 2, 1, 2, 2]. On the other hand, templates such as types 6, 8, and 9 feature multiple criteria, indicating that the installation methods for their fastening systems are more flexible.
[0201] CR-PPT
[0202] The CR-PPT enhances the Cascade R-CNN by making improvements in the inference stage while maintaining the same training process as the Cascade R-CNN.
[0203] Rail / tie detection using predefined proposal boxes
[0204] As shown in FIG.14 at a, the initial step of the process bypasses the RPN and uses a set of predefined proposal boxes to detect the rail and a centered tie. The rationale behind employing predefined proposal boxes for rail and tie detection is rooted in the consistent shooting angle of the camera used in this application. Since the camera is fixed, the rail typically appears along the central horizontal line of the image, while the ties are oriented vertically. This predictable positioning allows for the effective use of manually designed proposal boxes, enhancing detection efficiency and speed.
[0205] In the first step of FIG. 14 at a, these predefined boxes are specifically designed to target the areas of interest: five boxes are allocated for the tie and three for the rail. The predefined boxes are carefully sized, with the rail detection boxes measuring 512 × 128 pixels, and the tie detection boxes measuring 160 × 512 pixels. Additionally, these boxes have a stride of 32 pixels. This approach improves processing speed in two key aspects. First, by eliminating the need for the RPN, the method reduces inference time. Second, it contrasts with the conventional approach where an RPN generates around 1000 proposal boxes. By using only eight predefined boxes, this method enhances the speed of the box head and reduces the burden of the Non Maximum Suppression (NMS) process, thereby saving time.
[0206] In this disclosure, a “centered tie” refers to a tie whose bounding box touches the vertical center line of the image. This condition is met when the bounding box’s top-left ^^1 and bottom-right ^^2x-coordinates satisfy: 45 53600882 v1Docket No. 10780003.00112
[0207]
[0208] lacks a centered tie, the CR-PPT algorithm will proceed to analyze the subsequent frame.
[0209] Scaling component predefined proposal boxes
[0210] In the second step of FIG.14 at a, the focus is on accurately positioning the predefined proposal boxes on the centered tie. This rescaling procedure involves two key actions. First, the size of the boxes is adjusted to match the scale of the current image, based on the current rail width, denoted as ^^′. This adjustment serves as the de-normalizing factor for the predefined proposal boxes. Second, these boxes are placed on the centered tie using the insertion center coordinates of the rail and the centered tie, denoted as . This rescaling procedure can be viewed as the inverse of Equation (1):
[0211]
[0212] Where represents a predefined proposal box before rescaling.
[0213] Boxes refinement using cascade box head
[0214] In the third step of FIG. 14 at a, the cascade box head refines the predefined proposal boxes. This involves repositioning and scaling the boxes based on the features extracted from the image. The figure shows the transition of these boxes from their initial state in Step 2 to their refined state in Step 3. A matched box that has a proper initial position can align with a rail component, and thus achieve a high confidence score, indicated by green in the third step of FIG. 14 at a. Conversely, a mismatched box that is poorly positioned initially is less likely to correctly align with a rail component, which results in a low confidence score or a match with a rail 46 53600882 v1Docket No.10780003.00112 component of the incorrect class, denoted by red in this step. Additionally, a predefined box designated for clips can only match with clips or missing clips, while a spike and bolt box can only detect spikes and bolts or their installation holes.
[0215] In this disclosure, a box is considered successfully matched if it attains a confidence score above 0.1. Typically, the common threshold for confidence scores in object detection is set at 0.25 as observed in the box head of Faster R-CNN (Ren et al., 2016) and the detection head of YOLO (Jocher et al., 2023). This threshold, which ranges from 0 to 1, can be adjusted based on the specific challenges of the task. For instance, in the LVIS dataset challenge, the threshold may be set as low as 0.0 (Gupta et al., 2019). Given the challenge associated with refining rough predefined proposal boxes using a cascade box head, we have lowered the threshold to 0.1 to accommodate these difficulties.
[0216] Additionally, it is worth detailing the reason for using the cascade box head instead of the standard box head. FIG. 17A at a illustrates the third step of the process, showing the progressive evolution of boxes at each stage of the cascade box head and highlighting the importance of employing the cascade box head over the conventional one-stage box head. In contrast to the proposal boxes generated by the RPN, the boxes from predefined templates lack precision and accuracy. There is often a discrepancy in both the distance and size between a predefined box and the actual component it is meant to represent. This gap highlights why a one- stage box head fails to accurately shift and adjust the predefined box to align precisely with the component’s actual location. The employed cascade box head proves to be a more effective solution, offering multi-stage refinements in the localization of components.
[0217] Fastening system matching
[0218] FIG. 17B at b demonstrates the process of selecting the best-matched predefined proposal template for detection filtering and fastening system classification in the fourth step of FIG.14 at a. This process utilizes a combination of box matching scores and average confidence scores. The matching score for each template is computed by summing the values assigned to each component: “+1” for a matched component and “−1” for a mismatched component. For instance, if a predefined template contains four matched components and one mismatched component, the box matching for the template would be 3 (4 matched − 1 mismatched). 47 53600882 v1Docket No.10780003.00112
[0219] The template with the highest box matching score is initially considered the best match. However, in situations where multiple templates share the highest matching score, the process further refines the selection by comparing the average confidence scores of the boxes. Due to the powerful box refinement capacity of the cascade box head, in the last refinement stage, the score difference between templates may be minor. In such cases, we use the average confidence score of the boxes from the first stage for comparison.
[0220] Missing components detection
[0221] FIG.17B at c illustrates the final step of FIG.14 at a, which involves comparing the current detection result with predefined criteria to identify missing components. In a detection result, besides the installation hole of spike and bolt, the missing clip and low confidence component will be also marked as “1.” For this disclosure, the confidence threshold is set at 0.25.
[0222] In a criterion, a spike or bolt is represented by “2,” and a clip by “3.” If a component is not detected as expected, this indicates that the component is missing. FIG. 17B at c illustrates scenarios where a template has multiple predefined criteria. In such instances, the criterion with the fewest missing components is chosen to determine the missing detection result. If there are multiple criteria with an equal number of the fewest missing components, the first criterion listed is chosen.
[0223] EXPERIMENTS
[0224] Field data collecting and hardware configuration
[0225] FIG.18 at a illustrates two distinct sections at the Transportation Technology Center in Pueblo, Colorado, that were selected for data collection and field testing. The first section, extending over 2 miles, was specifically chosen for network training and validation. In contrast, the second section, covering a length of 500 feet, was utilized for conducting field testing. As depicted in FIG. 18 at b and c, in collaboration with ENSCO Inc., we integrated cameras and a computing system onto a rail inspection truck. The setup includes two Universal Serial Bus 3.0 (USB 3.0) industrial grade cameras, each boasting a resolution of 1024 × 1024 pixels and capable of capturing images at a maximum frame rate of 211 Frames Per Second (FPS). These cameras are mounted perpendicular to each rail on an aluminum extrusion rack located at the front of the truck. To support the computational demands, a portable computing platform has been installed. All devices on the platform are powered by a 12 V power supply, including batteries or a car’s 48 53600882 v1Docket No. 10780003.00112 cigarette lighter socket. The centerpiece of this platform is the NVIDIA Jetson AGX Orin. This edge computing device is equipped with a 12-core ARM CPU and boasts 2048 CUDA cores along with 64 tensor cores. Its normal and peak power consumption is 30 and 60W.
[0226] Evaluation metrics
[0227] Mean average precision (mAP) is used to evaluate the detection performance of CR- PPT. mAP is a common benchmark for object detection models, combining precision and recall at different thresholds. It gives a balanced score between detecting true positives and avoiding false positives: precision for class c. p(r) is ther.
[0230] FIG. 17B at c describes a specific aspect of CR-PPT. It highlights the necessity of selecting a correctly matched predefined template, which is crucial for both component detection and the recognition of missing components. The metric “Correctly Matched Number versus Image Number” (CMN / IN) is chosen as a key performance indicator, reflecting the proportion of validation images that are correctly matched to their templates.
[0231] To gauge the effectiveness of the proposed method in detecting missing components, recall, precision, and accuracy are employed as our evaluation metrics. In this disclosure, “positive samples” pertain to fastening systems exhibiting missing components, while “negative samples” refer to wholly intact ones. Recall assesses the proportion of true positive samples that are correctly identified, thereby indicating the defect detection rate of our proposed method, as exemplified in Equation (6). Precision, demonstrated in Equation (7), signifies the ratio of identified positive samples that have been accurately classified, thereby underlining the method’s resilience against noise. Accuracy, depicted in Equation (8), reflects the overall success rate of correct classification across all samples. 49 53600882 v1Docket No. 10780003.00112
[0232]
[0233] where TP, TN,negative, false- positive, and false-negative samples, respectively.
[0234] Network training and evaluation
[0235] CR-PPT is developed based on Cascade R-CNN, primarily modifying the inference process while retaining the main architecture of Cascade R-CNN. Our experiments involved a comparative analysis of two backbone networks and two RPNs within both the Cascade R-CNN and CR-PPT. For the backbone configuration, we selected Res2Net-101 (S. H. Gao et al., 2021) and Deep Layer Aggregation (DLA; Yu et al., 2018). Res2Net-101 is an enhancement over ResNet-101, aimed at higher accuracy but at the cost of slower processing speed. In contrast, DLA offers a balance of accuracy and lightness. Regarding the RPNs, we compared the original RPN (oriRPN) against CenterNet (Zhou et al., 2021). CenterNet is recognized for its precision in generating proposal boxes, boosting the overall performance of R-CNN. However, there is a possibility that CenterNet could reduce the learning intensity and, consequently, the robustness of the box head in the CR-PPT model. This aspect is crucial, as the box head robustness is integral to the effectiveness of CR-PPT. Thus, we need to determine whether CenterNet negatively impacts the box head functionality within the CR-PPT system.
[0236] Additionally, YOLOv8m (Jocher et al., 2023), a more recent state-of-the-art model, was chosen for a comparative disclosure, offering insights into the performance differences between the latest advanced models.
[0237] We collected over 60,000 images from the first section of FIG. 18 at a and selected 1200 images for annotation due to the large scale of these entire images. These 1200 images feature an approximately even distribution of each fastening system, types 1–13 as depicted in FIG. 15 at 50 53600882 v1Docket No.10780003.00112 a. We randomly divided the images into two groups: 900 for training and 300 for validation. Table 5, see FIG.19, presents the number of instances in each category, totaling 14,498 instances across the dataset. Creating such a comprehensive dataset for this specialized downstream task is markedly challenging. Our data indicate that annotating a single image requires at least 2 min, leading to a total annotation time exceeding 40 h for the 1200 images.
[0238] These networks were trained and evaluated on a system equipped with an Intel i9- 10920X CPU and an NVIDIA RTX A6000 GPU. Both Cascade R-CNN and CR-PPT were implemented using the Detectron2 framework (Yuxin Wu et al., 2019), while YOLOv8m utilized the official codebase from Ultralytics (Jocher et al., 2023). The models were trained using the Stochastic Gradient Descent (SGD) optimizer, following a stepped learning rate schedule of 0.01, 0.001, and 0.0001. Their batch size was set to 8. All models completed 90,000 iterations, which is approximately equivalent to 797 epochs. During the inference stage, the score and Intersection over Union (IoU) thresholds of Cascade R-CNN and YOLOv8 were set to 0.25 and 0.2, respectively.
[0239] To enhance inference speed, images were downsized to 512 × 512. Moreover, to effectively manage diverse lighting conditions and shadow effects, we implemented contrast limited adaptive histogram equalization (CLAHE; Reza, 2004) as a key data augmentation technique throughout the training and testing phases. During the training phase, there is a 50% chance that images will be processed with CLAHE. In contrast, during the model inferencing stage, CLAHE is consistently applied to all images.
[0240] CLAHE is critical for enhancing the model’s performance in varying environmental conditions. FIG.20 at a clearly demonstrates the effectiveness of CLAHE. In the original image, components under shadows are barely visible due to the strong sunlight. However, the equalized image processed with CLAHE shows significantly enhanced details thatwere previously obscured, thus eliminating the need for additional artificial lighting. Although CLAHE adds latency in the data preprocessing stage, we address optimization strategies for the inference pipeline infra.
[0241] Table 6, see FIG.21, details their validation performance metrics on the first validation dataset, along with their inference speeds.
[0242] Recent advancements in object detection have led to YOLOv8 significantly surpassing the R-CNN family in both accuracy and processing speed. According to Table 6, see FIG. 21, 51 53600882 v1Docket No.10780003.00112 YOLOv8 shows a 3.5 to 7.5 increase in mAP and a 2 to 6.5 times improvement in processing speed, compared to the R-CNNs. However, despite its efficiency, the YOLOv8 single-stage detection framework offers less customizability than the highly modular networks of the R-CNN family. Consequently, this disclosure continues to utilize Cascade R-CNN for the development of CRPPT, balancing the trade-off between performance and adaptability.
[0243] In each configuration, CR-PPT slightly reduces the mAP of Cascade R-CNN by approximately 0.31 to 1.38. This modest decrease is deemed acceptable, considering the predefined proposal boxes are not as accurate as those generated by the RPN, which contributes to a decline in mAP. Regarding processing speed, the disparity between Cascade R-CNN and CR- PPT is negligible. Although CRPPT bypasses the RPN and utilizes predefined proposal boxes for component detection, it goes through the cascade box head twice as depicted in FIG.14 at a (Steps 1 and 3), reducing its processing speed.
[0244] Regarding the backbone configuration, Res2Net-101 does not yield improvements in either speed or accuracy. R-CNNs utilizing Res2Net-101 are approximately 2.5 times slower than those employing DLA. In terms of accuracy, there is a reduction of 0.5%–2.1% in mAP when compared to DLA. Given the small size of the training set in this disclosure, 900 images, the diminished mAP could be attributed to larger backbones’ tendency to overfit on smaller datasets.
[0245] In terms of the RPN setting, experimental results indicate that a CenterNet-based RPN does not negatively impact and may even enhance the performance of CR-PPT. Specifically, CR- PPT (DLA-CenterNet) achieves a 70.69% mAP, which is 1.1% higher than CR-PPT (DLA- oriRPN) with the original RPN. The CMN / IN of CR-PPT (DLACenterNet) is 290 / 300, successfully matching four more predefined templates than CR-PPT (DLA-oriRPN). Consequently, CR-PPT (DLA-CenterNet) is selected for the training results of CR-PPT and will be utilized for further testing in subsequent sections.
[0246] FIG. 22 shows the example detection results of CRPPT (DLA-CenterNet), and its counterparts including Cascade R-CNN (DLA-CenterNet) and YOLOv8m. FIG.22 illustrates the strengths and weaknesses of CR-PPT. Conventional object detectors solely depend on image features for detection, making it difficult to completely avoid noise and misdetection. CR-PPT, on the other hand, leverages predefined proposal boxes and human knowledge to enhance robustness. First, it more effectively detects missing clips as evident in the first and second columns. Second, 52 53600882 v1Docket No.10780003.00112 it can infer components obscured by obstacles such as rocks and grass, as demonstrated in the third and fourth columns. However, a hidden normal component might be identified as missing due to low confidence, necessitating manual inspection. Third, CR-PPT can filter out noise resembling rail components as shown in the fifth and sixth columns. However, as indicated in the seventh and eighth columns, CR-PPT’s performance relies on both the accuracy of the predefined proposal boxes and the network’s robustness, leading to some inevitable errors and mismatches.
[0247] Performance on novel fastening systems
[0248] One critical aspect of CR-PPT’s performance is its effectiveness on novel fastening systems. As depicted in FIG. 15 at a, CR-PPT underwent testing on two additional fastening systems that were not included in its training.
[0249] The validation set for each new fastening system comprises 25 images. Table 7, see FIG.23, and FIGS. 24A and 24B initially present the zero-shot (no fine-tuning) performance of CR-PPT. Although CR-PPT demonstrates certain capabilities with these new images, the performance is not as robust as with the trained fastening systems. The mAP for these two new systems is recorded at 47.55% and 63.10%. Additionally, the CMN / IN stands at 13 / 25 and 20 / 25 for each system.
[0250] Despite the similarity in appearance of the clips, spikes, and bolts in the new systems to those in the trained systems, CR-PPT experiences a noticeable degradation in performance on these new types. This issue stems from the limited scope of the training set, which comprises only 900 images and is confined to a specific location in Colorado. As a result, to improve the zero- shot performance of CR-PPT, a more extensive and diverse training set is necessary. This set should encompass data from various locations and time periods to enhance its generalizability and robustness.
[0251] Creating a comprehensive training set to meet the zero-shot performance needs of CR- PPT is indeed challenging, requiring significant manpower and resources. Consequently, fine- tuning the network represents a more cost-effective method to enhance CR-PPT’s performance on new fastening systems. This approach allows for adjustments and improvements to be made to the model with less data and resource investment, adapting CR-PPT to recognize and work effectively with novel fastening types. 53 53600882 v1Docket No.10780003.00112
[0252] In this disclosure, we explore three network fine-tuning methods: one-shot, few-shot, and incremental learning. In one-shot fine-tuning, the network is trained on just one image per new fastening system, totaling two new images for training. Few-shot fine-tuning involves labeling and training the network on 10 images for each system, using a total of 20 new images. Incremental learning, on the other hand, integrates the new images into the existing training set, allowing the network to learn from both the old and new data, thereby fine-tuning its performance with a more comprehensive dataset.
[0253] In the one-shot and few-shot fine-tuning scenarios, the batch size and learning rate are configured to 2 and 1e-5, respectively, with the total number of iterations set to 1000. To prevent overfitting, given the small number of new images used for training, the network undergoes validation every 100 iterations, using the previous validation set as a benchmark. In contrast, for incremental learning, the batch size and learning rate are increased to 8 and 1e-4, respectively, with a longer training span of 10,000 iterations.
[0254] Table 7, see FIG.23, and FIGS.24A and 24B detail the fine-tuning results of CR-PPT, indicating that all methods significantly enhance performance on new fastening systems while maintaining robustness on the previous validation set. However, there are notable differences among the methods. Few-shot finetuning appears more prone to overfitting, with a slight decrease in the CMN / IN from 290 / 300 to 287 / 300. Incremental learning, when adding only one-shot images, is less effective, compared to other methods, showing the least improvement in mAP for the new fastening systems, with increases of only 7.66% and 0.86% for each type. Also, the CMN / IN metric saw only marginal gains of 7 / 25 and 3 / 25 for each system.
[0255] One-shot fine-tuning and incremental learning by adding few-shot images each present a distinct advantage. One-shot fine-tuning requires less time and effort for labeling and training, taking only about 11 min total for both processes. This makes it an efficient option when resources are limited. However, its improvements are moderate, with mAP increases of 10.05%, 0.94%, and 0.45% for the first and second new fastening systems and the previous validation set, respectively. The CMN / IN metrics show an increase of 9 / 25, 3 / 25, and 0 / 300 for each respective category.
[0256] On the other hand, incremental learning by adding few-shot images takes longer, approximately 63 min for labeling and training, but tends to yield more robust improvements. This method shows a higher boost in mAP of 14.50%, 3.26%, and 2.62% for the first and second new 54 53600882 v1Docket No.10780003.00112 fastening systems and the previous validation set, respectively. The corresponding CMN / IN increases are 12 / 25, 3 / 25, and 1 / 300.
[0257] Therefore, one-shot fine-tuning emerges as a time- and resource-efficient approach with moderate performance gains, while incremental learning with few-shot images demonstrates the higher potential for robust performance enhancements at the cost of increased time and effort. In practice, the selection between these methods depends on specific task requirements and constraints. Infra, we will utilize the fine-tuned model through incremental learning with few-shot images for further exploration and analysis.
[0258] Model deployment and speed optimization
[0259] This section details tests conducted on both the Jetson Orin and a desktop platform for the network training outlined in the previous section. The desktop platform is equipped with an Intel i9-10920X CPU and an NVIDIA RTX A6000 GPU. Notably, the A6000 GPU, with its 10,752 CUDA cores and 336 Tensor cores, provides substantial computational power and efficiency for complex tasks and processes. The batch size of the inference pipeline is set to two for the two cameras positioned on both rails.
[0260] Table 8, see FIG.25, lists the speed measurements on both platforms. C++, TorchScript and oneTBB have been adopted to optimize the speed of the CR-PPT.
[0261] The trained model is converted from PyTorch to Torch-Script, a serialized and optimized framework for network inference. TorchScript advantage lies in its ability to allow dynamic operations such as NMS and “if-else” statements, which are integral to the structure of CR-PPT. Despite the speed of TorchScript not matching that of TensorRT, which is optimized for NVIDIA’s hardware, its adoption is favored due to the challenges of applying dynamic operations in TensorRT. Table 8, see FIG.25, illustrates TorchScript’s efficiency on both platforms, reducing the model inference time from 32 to 21 ms for the desktop platform and from102 to 67 ms for Jetson Orin.
[0262] The entire pipeline is then converted from Python to C++ to decrease the overall inference time. C++ offers closer control over hardware and system resources, making it faster and more efficient than Python, which is a high-level interpreted language. This conversion decreases the overall inference time from 69 to 39 ms for the desktop framework and from 170 to 98 ms for the NVIDIA Jetson AGX Orin. 55 53600882 v1Docket No.10780003.00112
[0263] We have employed the C++ oneTBB library (Reinders, 2007) to develop a model inference pipeline, based on the producer-consumer model. This method leverages data parallelization and concurrent computation to significantly reduce latency brought by the pre- processing and post-processing stages. As a result of this optimization, the overall inference time has been reduced from 39 to 30 ms on desktop frameworks and from 98 to 71 ms on the NVIDIA Jetson AGX Orin. More comprehensive insights into this approach are discussed in Tang and Qian (2024).
[0264] In the United States, the spacing between railroad ties is approximately 24 inches. When processing each tie individually, the desktop system and the Jetson Orin can achieve the maximum processing speeds of 66.67 feet per second and 28.57 feet per second, respectively. This equates to speeds of 45.46 mph (73.16 km / h) for the desktop and 19.48 mph (31.35 km / h) for the Jetson Orin.
[0265] Filed blind testing for missing components detection
[0266] We conducted a blind field test in the second section depicted in FIG.18 at a, spanning a length of 500 feet. This involved a total of 283 railroad ties, from which we gathered 566 images, capturing both sides of each tie. Prior to the testing, the staff at ENSCO Inc. assisted in selectively removing or changing some components from various ties, ensuring a randomized element in the test setup. The changes were kept from us and shared only with the FRA.
[0267] Table 9, see FIG. 26, shows the missing detection results. Among these evaluated images, 30 displayed missing components, while 536 remained intact. Considering this extremely uneven distribution in the real world, recall emerges as the key metric for assessing performance. CR-PPT successfully identified 23 of the 30 images with anomalies in reality, achieving a recall rate of 76.66%. This result underscores CR-PPT’s capability in detecting missing components. Despite the imbalance, CR-PPT demonstrates impressive anti-noise effectiveness. It accurately classified most positive and negative samples, culminating in a robust accuracy rate of 97.70%.
[0268] This section concentrates on field testing for the detection of missing components, with results derived from manual statistics. Consequently, individual rail component annotations are not provided, and Table 9, see FIG. 26, does not include the mAP indicator, which is typically used to evaluate object detection accuracy. 56 53600882 v1Docket No.10780003.00112
[0269] The implementation of CR-PPT has markedly diminished the labor required for component inspection. Before its introduction, staff would be tasked with the meticulous examination of all 566 baseplate images, a process susceptible to human error. However, with CR- PPT, there were 23 true positives and six false positives identified, leading to a precision rate of 79.31%. This suggests that the majority of detections flagged by CR-PPT are indeed true positives. Consequently, personnel now only need to scrutinize these 29 highlighted images, greatly reducing the workload and time expenditure.
[0270] FIG.27 illustrates various detection results achieved using CR-PPT, highlighting both its strengths and limitations. The first two rows exhibit successful cases, demonstrating CR-PPT’s proficiency in component detection and identifying missing components. However, in complex real-world scenarios, CR-PPT, like other deep learning models, may not consistently maintain high performance, leading to potential errors in detection. The false-positive examples in the third row illustrate a specific challenge: Components obscured by obstacles such as stones and grass can be mistakenly identified as missing. This indicates a limitation in the model’s ability to discern between genuinely missing components and those merely hidden from view. Also, the false- negative samples point out a need for improvement in CR-PPT’s recognition of missing clips. This issue stems from the scarcity of missing clip examples in the training dataset. In real-world applications, data on missing clips are rare and difficult to collect, posing a challenge for training robust detection models for missing clips.
[0271] To address these challenges and enhance CR-PPT’s effectiveness, a focus on diversifying the datasets and increasing the variety of images used for training is crucial. Enriching the training data with a wider range of scenarios, especially those involving obstacles and missing clips, will improve CR-PPT’s accuracy and reliability in diverse and challenging real-world settings.
[0272] This disclosure introduces the innovative CR-PPT method for detecting missing components in railway infrastructure, marking an advancement from previous methods that struggled with differentiating between missing spikes and bolts and their unused installation holes. CR-PPT integrates predefined proposal templates designed to discern both existing and missing components within diverse fastening systems. The process initiates with the detection of rails and ties using these templates. Subsequent steps include scaling and positioning the predefined boxes 57 53600882 v1Docket No.10780003.00112 on centered ties, followed by the cascade box head refining these boxes based on image features. The method then employs template matching, combining box matching and confidence scores, to select the most suitable template corresponding to the observed components. Finally, CRPPT uses predefined criteria within the chosen template to determine the presence or absence of components of various fastening systems. This refined approach marks a significant enhancement in the tools available for assessing the health of railway infrastructure. The main findings are summarized as follows: 1. The proposed CR-PPT, using Res2Net-101 and DLA backbones and original RPN and CenterNet RPNs, was compared with YOLOv8m. Despite YOLOv8m’s superior accuracy and speed, CR-PPT’s adaptability, especially with DLA-CenterNet, effectively balances accuracy and speed. This setup provides a balanced performance, achieving an accuracy of 70.69% mAP with 290 / 300 CMN / IN while operating at a speed of 32.04 FPS. 2. Initially, CR-PPT demonstrated zero-shot performance for two new fastening systems, achieving mAP of 47.55% and 63.10% and CNM / IN of 13 / 25 and 20 / 25. Performance enhancements were explored through one-shot, few- shot, and incremental learning, showing varying degrees of improvement and highlighting trade- offs between time, resources, and performance gains.3. The system, incorporating two cameras, underwent optimization by converting from Python to C++ and employing TorchScript and oneTBB. These changes resulted in a substantial increase in processing speed. Specifically, the inference time was reduced from 69 to 30 ms on a desktop and from 170 to 71 ms on a Jetson Orin platform.4. In a 500-foot section field blind test involving 566 images, CR-PPT identified 23 out of 30 images with missing components, achieving a 76.66% recall, 79.31% precision, and 97.70% accuracy.
[0273] The following References are hereby incorporated by reference to the extent not inconsistent herewith:
[0274] Alam, K. M. R., Siddique, N., & Adeli, H. (2020). A dynamic ensemble learning algorithm for neural networks. Neural Computing and Applications, 32(12), 8675–8690.
[0275] Amezquita-Sanchez, J. P., & Adeli, H. (2019). Nonlinear measurements for feature extraction in structural health monitoring. Scientia Iranica, 26(6), 3051–3059.
[0276] Aydın, İ., Sevi, M., Salur, M. U., & Akın, E. (2022). Defect classification of railway fasteners using image preprocessing and a lightweight convolutional neural network. Turkish 58 53600882 v1Docket No.10780003.00112 Journal of Electrical Engineering and Computer Sciences, 30(3), 891–907. doi.org / 10.55730 / 1300-0632.3817.
[0277] Aydin, I., Sevi, M., Akin, E., Güçlü, E., Karaköse, M., & Aldarwich, H. (2023). A Deep learning-based hybrid approach to detect fastener defects in real-time. Tehnički vjesnik, 30(5), 1461–1468. doi.org / 10.17559 / TV-20221020152721.
[0278] Bai, T., Yang, J., Xu, G., & Yao, D. (2021). An optimized railway fastener detection method based on modified faster R-CNN. Measurement, 182, 109742. doi.org / 10.1016 / j.measurement.2021.109742
[0279] Brunelli, R. (2009). Template matching techniques in computer vision: Theory and practice. John Wiley & Sons.
[0280] Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., & Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv preprint arXiv:2303.12712.
[0281] Bolya, D., Zhou, C., Xiao, F., & Lee, Y. J. (2019). YOLACT: Real-time instance segmentation. Proceedings of the IEEE / CVF International Conference on Computer Vision, Seoul, South Korea (pp.9157–9166).
[0282] Cai, Z., & Vasconcelos, N. (2018). Cascade R-CNN: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6154–6162).
[0283] Cao, Y., Chen, Z., Wen, T., Roberts, C., Sun, Y., & Su, S. (2023). Rail fastener detection of heavy railway based on deep learning. Highspeed Railway, 1(1), 63–69. doi.org / 10.1016 / j.hspr.2022.11.001.
[0284] Chandran, P., Thiery, F., Odelius, J., Lind, H., & Rantatalo, M. (2022). Unsupervised machine learning for missing clamp detection from an in-service train using differential eddy current sensor. Sustainability, 14(2), 1035. doi.org / 10.3390 / su14021035.
[0285] Chen, M., Zhai, W., Zhu, S., Xu, L., & Sun, Y. (2022). Vibration based damage detection of rail fastener using fully convolutional networks. Vehicle System Dynamics, 60(7), 2191–2210.
[0286] Chun, P.-J., Yamane, T., & Maemura, Y. (2022).Adeep learning-based image captioning method to automatically generate comprehensive explanations of bridge damage. 59 53600882 v1Docket No.10780003.00112 Computer-Aided Civil and Infrastructure Engineering, 37(11), 1387–1401. doi.org / 10.1111 / mice.12793
[0287] Feng, H., Jiang, Z., Xie, F., Yang, P., Shi, J., & Chen, L. (2014). Automatic fastener classification and defect detection in vision-based railway inspection systems. IEEE Transactions on Instrumentation and Measurement, 63(4), 877–888. doi.org / 10.1109 / TIM.2013.2283741
[0288] FRA. (2018). Track and rail and infrastructure integrity compliance manual. railroads.dot.gov / sites / fra.dot.gov / files / fra_net / 17940 / CM%20Vol%20II%20Ch1%202018.pdf
[0289] Fu, J., Chen, X., & Lv, Z. (2022). Rail fastener status detection based on mobilenet- YOLOv4. Electronics, 11(22), 3677. doi.org / 10.3390 / electronics11223677
[0290] Gao, S.-H., Cheng, M.-M., Zhao, K., Zhang, X.-Y., Yang, M H., & Torr, P. (2021). Res2Net: A new multi-scale backbone architecture. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(2), 652–662. doi.org / 10.1109 / tpami.2019.2938758
[0291] Gao, Y., Cao, Z., Qin, Y., Ge, X., Lian, L., Bai, J., & Yu, H. (2023). Railway fastener anomaly detection via multi-sensor fusion and self driven loss reweighting. IEEE Sensors Journal, 24(2), 1812–1825. doi.org / 10.1109 / JSEN.2023.3336962
[0292] Gosiewska, A., Baran, Z., Baran, M., & Rutkowski, T. (2023). Seeking a sufficient data volume for railway infrastructure component detection with computer vision models. Sensors, 23(18), 7776. doi.org / 10.3390 / s23187776
[0293] Guo, F., Qian, Y., Rizos, D., Suo, Z., & Chen, X. (2021). Automatic rail surface defects inspection based on Mask R-CNN. Transportation Research Record, 2675(11), 655–668. doi.org / 10.1177 / 03611981211019034
[0294] Guo, F., Qian, Y., & Shi, Y. (2021). Real-time railroad track components inspection based on the improved yolov4 framework. Automation in Construction, 125, 103596. doi.org / 10.1016 / j.autcon.2021.103596
[0295] Guo, F., , Y., , Y., Leng, Z.,& ,H. (2021). Automatic railroad track components inspection using real-time instance segmentation.
[0296] Computer-Aided Civil and Infrastructure Engineering, 36(3), 362–377. doi.org / 10.1111 / mice.12625Guo, F., Qian, Y., & Yu, H. (2023). Automatic rail surface defect inspection using the pixelwise semantic segmentation model. IEEE Sensors Journal, 23(13), 15010–15018. doi.org / 10.1109 / JSEN.2023.3280117 60 53600882 v1Docket No.10780003.00112
[0297] Guo, J., Zhang, S., Qian, Y., & Wang, Y. (2023). An adaptively weighted loss-enabled lightweight teacher–student model for real time railroad inspection on edge devices. Neural Computing and Applications, 35(34), 24455–24472. doi.org / 10.1007 / s00521-023-09038-2
[0298] Gupta, A., Dollar, P., & Girshick, R. (2019). LVIS: A dataset for large vocabulary instance segmentation. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA (pp.5356–5364).
[0299] He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy (pp.2961–2969).
[0300] Hu, S., Li, Z., Wang, S., Ai, M., & Hu, Q. (2020). A texture selection approach for cultural artifact 3D reconstruction considering both geometry and radiation quality. Remote Sensing, 12(16), 2521.
[0301] Hu, J., Qiao, P., Lv, H., Yang, L., Ouyang, A., He, Y., & Liu, Y. (2022). High speed railway fastener defect detection by using improved YoLoX-Nano model. Sensors, 22(21), 8399. doi.org / 10.3390 / s22218399
[0302] Jähne, B. (2005). Digital image processing. Springer Science & Business Media.
[0303] Javadinasab Hormozabad, S., Gutierrez Soto, M., & Adeli, H. (2021). Integrating structural control, health monitoring, and energy harvesting for smart cities. Expert Systems, 38(8), e12845.
[0304] Jocher, G., Chaurasia, A., & Qiu, J. (2023). Ultralytics YOLOv8. github.com / ultralytics / ultralytics
[0305] Li, X., Wang, Q., Yang, X., Wang, K., & Zhang, H. (2023). Track fastener defect detection model based on improved YOLOv5s.
[0306] Sensors, 23(14). doi.org / 10.3390 / s23146457 Li, Y., Trinh, H., Haas, N., Otto, C., & Pankanti, S. (2014). Rail component detection, optimization, and assessment for automatic rail track inspection. IEEE Transactions on Intelligent Transportation Systems, 15(2), 760–770. doi.org / 10.1109 / TITS.2013.2287155
[0307] Liu, J., Huang, Y., Wang, S., Zhao, X., Zou, Q., & Zhang, X. (2022). Rail fastener defect inspection method for multi railways based on machine vision. Railway Sciences, 1(2), 210–223. doi.org / 10.1108 / RS-04-2022-0012 61 53600882 v1Docket No.10780003.00112
[0308] Liu, J., Liu, H., Chakraborty, C., Yu, K., Shao, X., & Ma, Z. (2023). Cascade learning embedded vision inspection of rail fastener by using a fault detection iot vehicle. IEEE Internet of Things Journal, 10(4), 3006–3017. doi.org / 10.1109 / JIOT.2021.3126875
[0309] Liu, J., Qiu, Y., Ni, X., Shi, B., & Liu, H. (2023). Fast detection of railway fastener using a new lightweight network Op-YOLOv4-tiny. IEEE Transactions on Intelligent Transportation Systems, 25(1), 133–143. doi.org / 10.1109 / TITS.2023.3305300.
[0310] Liu, J., Teng, Y., Shi, B., Ni, X., Xiao, W., Wang, C., & Liu, H. (2021). A hierarchical learning approach for railway fastener detection using imbalanced samples. Measurement, 186, 110240. doi.org / 10.1016 / j.measurement.2021.110240
[0311] Liu, Y., & Song, B. (2022). Shuffle-RepNet: A reparameterized convolutional neural network for rail clip state recognition.202241stChinese Control Conference (CCC), Hefei, China (pp.7343–7348). doi.org / 10.23919 / CCC55666.2022.9902547
[0312] Perez-Ramirez, C. A., Amezquita-Sanchez, J. P., Valtierra-Rodriguez, M., Adeli, H., Dominguez-Gonzalez, A., & Romero-Troncoso, R. J. (2019). Recurrent neural network model with Bayesian training and mutual information for response prediction of large buildings. Engineering Structures, 178, 603–615.
[0313] Pezeshki,H., Pavlou, D., Adeli, H., & Siriwardane, S. C. (2023). Modal analysis of offshore monopile wind turbine: An analytical solution. Journal of Offshore Mechanics and Arctic Engineering, 145(1), 010907.
[0314] Qi, H., Xu, T., Wang, G., Cheng, Y., & Chen, C. (2020). MYOLOv3-Tiny: A new convolutional neural network architecture for realtime detection of track fasteners. Computers in Industry, 123, 103303. doi.org / 10.1016 / j.compind.2020.103303
[0315] Qiu, Y., Liu, H., Shi, B., & Liu, J. (2023). Center-triplet loss for railway defective fastener detection. IEEE Sensors Journal, 24(3), 3180–3190. doi.org / 10.1109 / JSEN.2023.3339883.
[0316] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning. (pp.8748–8763). PMLR. 62 53600882 v1Docket No.10780003.00112
[0317] Rafiei, M. H., Gauthier, L. V., Adeli, H., & Takabi, D. (2022). Self-supervised learning for electroencephalography. IEEE Transactions and Learning Systems, 35(2), 1457–1471.
[0318] Reinders, J. (2007). Intel threading building blocks: Outfitting C++ for multi-core processor parallelism. O’Reilly Media, Inc.
[0319] Ren, S., He, K., Girshick, R., & Sun, J. (2016). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137–1149.
[0320] Reza, A. M. (2004). Realization of the contrast limited adaptive histogram equalization (CLAHE) for real-time image enhancement. Journal of VLSI Signal Processing Systems for Signal, Image and Video Technology, 38, 35–44.
[0321] Robinson, R., Nguyen, L., Moore, W. H., Culotta, K., Hocevar, H., Kimmel, S., Stacey, M., Bricka, S., Bronzini, M., Edmonds, J., Fang, B., Firestine, T., Fletcher, W., Greene, D., Kent, P., Pisarski, A., & Rick, C. (2023). Transportation statistics annual report 2023. Department of Transportation. Bureau of Transportation Statistics. doi.org / 10.21949 / 1529944
[0322] Su, S., Du, S., & Lu, X. (2022). Geometric constraint and image inpainting-based railway track fastener sample generation for improving defect inspection. IEEE Transactions on Intelligent Transportation Systems, 23(12), 23883–23895.
[0323] Su, S., Du, S., Wei, X., & Lu, X. (2023). RFS-Net: Railway track fastener segmentation network with shape guidance. IEEE Transactions on Circuits and Systems for Video Technology, 33(3), 1398–1412. doi.org / 10.1109 / TCSVT.2022.3212088
[0324] Tang, Y., & Qian, Y. (2024). High-speed railway track components inspection framework based on YOLOv8 with high-performance model deployment. High-speed Railway, 2(1), 42–50.
[0325] Tang, Y., Wang, Y., & Qian, Y. (2024). Edge-computing oriented real-time missing track components detection. Transportation Research Record, Advance online publication. doi.org / 10.1177 / 03611981241230546
[0326] Wei, D., Wei, X., & Jia, L. (2022). Automatic defect description of railway track line image based on dense captioning. Sensors, 22(17), 6419. doi.org / 10.3390 / s22176419 63 53600882 v1Docket No.10780003.00112
[0327] Wei, D., Wei, X., Tang, Q., Jia, L., Yin, X., & Ji, Y. (2023). RTLSeg: A novel multi- component inspection network for railway track line based on instance segmentation. Engineering Applications of Artificial Intelligence, 119, 105822. doi.org / 10.1016 / j.engappai.2023.105822
[0328] Wu, Y., Chen, P., Qin, Y., Qian, Y., Xu, F., & Jia, L. (2023). Automatic railroad track components inspection using hybrid deep learning framework. IEEE Transactions on Instrumentation and Measurement, 72, 5011415. doi.org / 10.1109 / TIM.2023.3265636
[0329] Wu, Y., Kirillov, A., Massa, F., Lo, W.-Y., & Girshick, R. (2019). Detectron2. github.com / facebookresearch / detectron2
[0330] Wu, Y., Qin, Y., Qian, Y., & Guo, F. (2021).Automatic detection of arbitrarily oriented fastener defect in high-speed railway. Automation in Construction, 131, 103913. doi.org / 10.1016 / j.autcon.2021.103913
[0331] Wu, Y., Qin, Y., Qian, Y., Guo, F., Wang, Z., & Jia, L. (2022). Hybrid deep learning architecture for rail surface segmentation and surface defect detection. Computer-Aided Civil and Infrastructure Engineering, 37(2), 227–244. doi.org / 10.1111 / mice.12710
[0332] Xiao, T., Xu, T., & Wang, G. (2023). Real-time detection of track fasteners based on object detection and FPGA. Microprocessors and Microsystems, 100, 104863. doi.org / 10.1016 / j.micpro.2023.104863
[0333] Yang, C., Kaynardag, K., & Salamone, S. (2023). Missing rail fastener detection based on laser Doppler vibrometer measurements. Journal of Nondestructive Evaluation, 42(3), 68. doi.org / 10.1007 / s10921-023-00981-7
[0334] Yang, E., Tang, Y., Zhang, A. A., Wang, K. C. P., & Qiu, Y. (2023). Policy gradient– based focal loss to reduce false negative errors of convolutional neural networks for pavement crack segmentation. Journal of Infrastructure Systems, 29(1), 04023002. doi.org / 10.1061 / JITSE4.ISENG-2157
[0335] Yilmazer, M., & Karakose, M. (2022). Mask R-CNN architecture based railway fastener fault detection approach. 2022 International Conference on Decision Aid Sciences and Applications (DASA), Chiangrai, Thailand (pp. 1363–1366). doi.org / 10.1109 / DASA54658.2022.9765024 64 53600882 v1Docket No.10780003.00112
[0336] Yilmazer, M., & Karakose, M. (2023). YOLOv5 based fault detection approach in railway components. 2023 27th International Conference on Information Technology (IT), Zabljak, Montenegro (pp.1–4). doi.org / 10.1109 / IT57431.2023.10078564
[0337] Yin, X., Wei, X., & Zheng, H. (2022). Railway track vibration analysis and intelligent recognition of fastener defects. Advanced Theory and Simulations, 5(10), 2200027. doi.org / 10.1002 / adts.202200027
[0338] Yu, F., Wang, D., Shelhamer, E., & Darrell, T. (2018).Deep layer aggregation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2403–2412).
[0339] Zhang, A., Li, Q., Wang, K. C., & Qiu, S. (2013).Matched filtering algorithm for pavement cracking detection. Transportation Research Record, 2367(1), 30–42.14678667, 2024,
[0340] Zhang, A. A., Wang, K. C. P., Liu, Y., Zhan, Y., Yang, G., Wang, G., Yang, E., Zhang, H., Dong, Z., He, A., Xu, J., & Shang, J. (2022). Intelligent pixel-level detection of multiple distresses and surface design features on asphalt pavements. Computer-Aided Civil and Infrastructure Engineering, 37(13), 1654–1673. doi.org / 10.1111 / mice.12909
[0341] Zhang, G.-Q., Wang, B., Li, J., & Xu, Y.-L. (2022). The application of deep learning in bridge health monitoring: A literature review. Advances in Bridge Engineering, 3(1), 22. doi.org / 10.1186 / s43251-022-00078-7
[0342] Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT (pp.586–595).
[0343] Zhou, X., Koltun, V., & Krähenbühl, P. (2021). Probabilistic two-stage detection. arXiv preprint arXiv:2103.07461.
[0344] Zhou, X., Yin, T., Koltun, V., & Krähenbühl, P. (2022). Global tracking transformers. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp.8771–8780).
[0345] Zhuang, L., Qi, H., Wang, T., & Zhang, Z. (2022). A deep-learning powered near-real- time detection of railway track major components: A two-stage computer-vision-based method. IEEE Internet of Things Journal, 9(19), 18806–18816. doi.org / 10.1109 / JIOT.2022.3162295 **** 65 53600882 v1Docket No.10780003.00112
[0346] Various modifications and variations of the described methods, compositions, and kits of the disclosure will be apparent to those skilled in the art without departing from the scope and spirit of the disclosure. Although the disclosure has been described in connection with specific embodiments, it will be understood that it is capable of further modifications and that the disclosure as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the disclosure that are obvious to those skilled in the art are intended to be within the scope of the disclosure. This application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the present disclosure come within known customary practice within the art to which the disclosure pertains and may be applied to the essential features herein before set forth. 66 53600882 v1
Claims
Docket No.10780003.00112 CLAIMS What is claimed is: What is claimed is:
1. A system for railroad component inspection comprising: a railway track inspection device including at least one neural network; the at least one neural network employing at least one predefined proposal template, wherein the at least one predefined proposal template, via a matching mechanism, classifies at least one fastening system used on a railway track being inspected; and the at least one neural network employing the at least one predefined proposal template, after the at least one fastening system is classified via the matching mechanism, employs at least one predefined criteria tailored to the at least one fastening system to detect both a presence or an absence of at least one fastening component on a railway track as dictated by the at least predefined criteria tailored to the at least one fastening system.
2. The system for railroad component inspection of claim 1, wherein the fastening component comprises a clip, spike or bolt.
3. The system for railroad component inspection of claim 1, wherein the at least one predefined proposal template operates via at least one predefined proposal box being refined by at least one cascade box head.
4. The system for railroad component inspection of claim 1, wherein the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening system determines if at least one fastening component on the railway track is correctly installed.
5. The system for railroad component inspection of claim 1, wherein the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening 67 53600882 v1Docket No.10780003.00112 system determines if an empty installation hole indicates at least one missing spike or the installation hole is configured as empty as dictated by the at least one fastening system classified via the matching mechanism.
6. The system for railroad component inspection of claim 1, wherein the at least one neural network employing the at least one predefined proposal template determines the presence or the absence of at least one fastening component on a railway track for a novel fastening system not included in a training data set for the system for railroad component inspection.
7. The system for railroad component inspection of claim 1, wherein the system is integrated into an edge detection device for railway inspection.
8. The system for railroad component inspection of claim 1, wherein the at least one neural network employing the at least one predefined proposal template, via the matching mechanism, classifies the at least one fastening system used on the railway track being inspected via at least one specific layout of rail components associated with the at least one fastening system.
9. The system for railroad component inspection of claim 1, wherein the system detects at least one: rail, tie, clip, missing clip, spike, bolt, installation hole or combinations of the above on the railway track.
10. The system for railroad component inspection of claim 1, wherein the system detects and segments at least one rail and at least one tie to determine a location and a scale of the at least one fastening system.
11. A method for railroad component inspection comprising: providing at least one track inspection device including at least one neural network; the at least one neural network employing at least one predefined proposal template, wherein 68 53600882 v1Docket No.10780003.00112 the at least one predefined proposal template, via a matching mechanism, classifies at least one fastening system used on a railway track being inspected; and the at least one neural network employing the at least one predefined proposal template, after the at least one fastening system is classified via the matching mechanism, employs at least one predefined criteria tailored to the at least one fastening system to detect both a presence or an absence of at least one fastening component on a railway track as dictated by the at least predefined criteria tailored to the at least one fastening system.
12. The method for railroad component inspection of claim 11, further comprising wherein the fastening component comprises a clip, spike or bolt.
13. The method for railroad component inspection of claim 11, further comprising operating the at least one predefined proposal template via at least one predefined proposal box being refined by at least one cascade box head.
14. The method for railroad component inspection of claim 11, further comprising determining via the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening system if at least one fastening component on the railway track is correctly installed.
15. The method for railroad component inspection of claim 11, further comprising determining via the matching mechanism employing the at least one predefined criteria tailored to the at least one fastening system if an empty installation hole indicates at least one missing spike or the installation hole is configured as empty as dictated by the at least one fastening system classified via the matching mechanism.
16. The method for railroad component inspection of claim 11, further comprising determining via the at least one neural network employing the at least one predefined proposal 69 53600882 v1Docket No.10780003.00112 template the presence or the absence of at least one fastening component on a railway track for a novel fastening system not included in a training data set for the system for railroad component inspection.
17. The method for railroad component inspection of claim 11, further comprising integrating the method into an edge detection device for railway inspection.
18. The method for railroad component inspection of claim 11, further comprising classifying via the at least one neural network employing the at least one predefined proposal template, via the matching mechanism, the at least one fastening system used on the railway track being inspected via at least one specific layout of rail components associated with the at least one fastening system.
19. The method for railroad component inspection of claim 11, further comprising detecting at least one: rail, tie, clip, missing clip, spike, bolt, installation hole or combinations of the above on the railway track.
20. The method for railroad component inspection of claim 11, further comprising detecting and segmenting at least one rail and at least one tie to determine a location and a scale of the at least one fastening system. 70 53600882 v1