Urban rail multi-category disease real-time detection method based on improved YOLO v8

By improving the YOLO v8 model, constructing a multi-category disease image dataset and adding a feature reconstruction module, and combining TensorRT optimization and multi-threading strategies, the problems of low efficiency and poor accuracy in urban rail transit disease detection were solved, achieving high-precision and high-efficiency disease identification.

CN121010955APending Publication Date: 2025-11-25BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511114739.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing urban rail transit track defect detection methods suffer from low detection efficiency, high cost, high false alarm and missed detection rates, and poor real-time identification performance. In particular, they are difficult to achieve high accuracy and efficiency when detecting multiple types of defects.

Method used

A real-time detection method for multiple types of urban rail transit defects based on an improved YOLO v8 is constructed. By building a multi-type defect image dataset, adding a feature reconstruction module, and optimizing the model with different loss functions, the detection capability and speed of the model are improved by combining TensorRT optimization strategy and multi-threaded parallel strategy.

Benefits of technology

It achieves high-precision and high-efficiency identification of urban rail transit defects, reduces the false alarm rate and missed alarm rate, meets the real-time requirements of urban rail transit inspection, and improves the model's generalization ability and computing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010955A_ABST
    Figure CN121010955A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of rail disease detection, in particular to an urban rail multi-category disease real-time detection method based on improved YOLO v8, which comprises the following steps: constructing a multi-category disease image data set of an urban rail; yOLO v8 is used as a basic network framework, a feature reconstruction module is added in a detection head of a YOLO v8 network, and a disease detection model is constructed; the feature reconstruction module is used for optimizing normal data and disease data by adopting different loss functions, and carrying out lossless reconstruction of normal characterization and directional deduction and strong feature reconstruction of disease characterization relative to a normal characterization clustering center; training a disease detection model based on the multi-category disease image data set; and reasoning the real-time orbit data based on the trained disease detection model. According to the invention, high-precision and high-efficiency identification of urban rail disease categories can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of track defect detection technology, and more specifically to a real-time detection method for multiple types of defects in urban rail transit based on an improved YOLO v8. Background Technology

[0002] As an effective carrier for the safe operation of vehicles in urban rail transit infrastructure, the subway track structure directly bears the train load and plays a guiding role in train operation. A good and healthy track condition is the foundation for ensuring the safe and stable operation of the line. However, as urban rail trains gradually accumulate service years, the condition of the line equipment foundation continues to change, and defects gradually increase, which has a significant impact on the safety, stability, resilience, efficiency, and service capacity of the urban integrated transportation system with the urban rail network as its main artery.

[0003] Current urban rail transit inspection methods mainly include three categories: manual inspection, methods based on traditional image processing, and methods based on deep learning.

[0004] Traditional manual inspection methods are inefficient, costly, and susceptible to subjective influence from technicians, making it difficult to guarantee the quality of test results.

[0005] Traditional image processing-based methods for detecting defects in rail transit facilities first extract key information from images of these facilities by manually designing features, and then combine prior knowledge with traditional image processing techniques to detect potential defects. This approach has limitations, as it cannot simultaneously detect and identify multiple types of track defects.

[0006] The core idea of ​​deep learning-based methods for detecting defects in rail transit facilities is to design deep neural networks based on a large number of rail transit image samples. Through sufficient model training and optimization, the visual semantic features of rail transit targets are automatically extracted, thereby accurately detecting and identifying defect areas. Two-stage detection networks, represented by Faster-RCNN, and single-stage detection networks, represented by YOLO, have been applied to rail transit defect detection tasks and have achieved good results. However, due to the continuous expansion of urban rail transit networks and the increasing years of operation and maintenance, the current rail inspection systems have revealed shortcomings in technical capabilities. These shortcomings mainly manifest as a small number of inspected defects with insufficient representativeness, high false negative and false positive rates, and poor real-time identification performance.

[0007] Therefore, how to expand the categories of disease images and improve the detection capabilities of the model, while reducing the false negative and false positive rates, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] In view of this, the present invention provides a real-time detection method for multiple types of urban rail transit defects based on an improved YOLO v8, which can achieve high-precision and high-efficiency identification of urban rail transit defect categories.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] A real-time detection method for multiple types of defects in urban rail transit based on an improved YOLO v8 includes the following steps:

[0011] Construct a multi-category image dataset of urban rail transit defects;

[0012] Based on the YOLO v8 network framework, a feature reconstruction module is added to the detection head of the YOLO v8 network to construct a disease detection model. The feature reconstruction module is optimized by using different loss functions for normal data and disease data respectively, to perform lossless reconstruction of normal representation and directional displacement and strong feature reconstruction of disease representation relative to the cluster center of normal representation.

[0013] The disease detection model is trained based on the multi-category disease image dataset;

[0014] The trained defect detection model is used to infer information from real-time track data.

[0015] Furthermore, the process of constructing the multi-category disease image dataset includes:

[0016] Track condition inspection equipment is installed at the bottom of the urban rail comprehensive inspection train. Five high-speed line array cameras are mounted on the track condition inspection equipment so that the imaging area covers the entire track bed.

[0017] Based on the normal images collected by the track condition inspection equipment, as well as images of six types of track facility defects, including foreign objects in the track bed, missing elastic clips, misaligned fasteners, loose bolts, reversed elastic clips, and broken elastic clips;

[0018] For a single or several types of disease image samples that are insufficient in number, geometric transformation and image enhancement methods are used to expand them, and after expansion, the final multi-category disease image dataset is obtained.

[0019] Furthermore, the disease detection model includes a backbone network, a neck network, and a detection network;

[0020] The backbone network uses deep convolutional networks and residual networks to extract visual representation information of images at different levels;

[0021] The neck network adopts the PANet network structure to propagate strong semantic information from the top layer to the bottom layer, and propagates the strong localization features of the bottom layer to the top layer. By fusing the features of different levels, feature maps of three scales, 20×20, 40×40 and 80×80, are output.

[0022] The detection network comprises three detection heads, each of which receives feature maps of three scales transmitted by the neck network and performs detection on large, medium and small targets respectively. All three detection heads use a decoupled head structure, and each detection head contains two branches, which respectively complete the classification task and the regression task.

[0023] Furthermore, each of the detection heads is equipped with the feature reconstruction module; the feature reconstruction module consists of four convolutional blocks responsible for feature extraction and an SE attention mechanism weighted according to the correlation of feature map channels.

[0024] Furthermore, the loss function L for normal data reconstruction normal Represented as:

[0025]

[0026] Where, x i This represents the value of the i-th normal data point before reconstruction. Let represent the value after reconstruction of the i-th normal data point, and n represent the sample size of the normal data.

[0027] Furthermore, the loss function L for disease data reconstruction anomaly The expression is:

[0028] L anomaly =L direction +L rec

[0029] Among them, L direction For directional loss, L rec For reconstruction loss; directional loss L direction The feature maps of diseased data before and after reconstruction, and the feature map of cluster centers of normal data before reconstruction, are unfolded into vectors. Cosine similarity is used to control the reconstruction of diseased data in the feature space along the direction that is significantly different from that of normal data; reconstruction loss L rec The degree of enhancement of disease characteristics is controlled by a preset intensity threshold.

[0030] Furthermore, the directional loss L direction The expression is:

[0031]

[0032] Where, x i' represents the value of the i-th outlier before reconstruction, x' i ' represents the reconstructed value of the i-th outlier, n' represents the sample size of the outlier, and c represents the cluster centers of the normal data; directional loss L direction It is used to reconstruct the original features of the non-abnormal areas in disease images, while enhancing the features of the abnormal areas.

[0033] Furthermore, reconstruction loss L rec The expression is:

[0034]

[0035] Where m is a hyperparameter, and L is the reconstruction loss. rec It is used to control the degree of enhancement of abnormal region features during the feature reconstruction of disease images.

[0036] Furthermore, when the disease detection model performs inference on real-time track data, it adopts a TensorRT optimization strategy and a multi-threaded parallel strategy.

[0037] Furthermore, when the disease detection model performs inference on real-time track data, the number of images detected in a single instance is 600.

[0038] As can be seen from the above technical solution, compared with the prior art, the present invention has the following beneficial effects:

[0039] This invention optimizes three aspects: detection of multiple types of diseases, detection accuracy, and detection speed.

[0040] In terms of multi-category defect detection, this invention constructs an image dataset containing six categories of track defects, including foreign objects in the track bed, missing spring clips, misaligned fasteners, loose bolts, and broken spring clips. Furthermore, it expands the number of defect category samples that are insufficient, enriching the number of defect images and providing more training samples for the model to accurately detect defect areas.

[0041] Regarding detection accuracy, this invention uses YOLO v8 as the basic network framework and adds a feature reconstruction module to the detection head. The feature reconstruction module increases the distance between normal and abnormal representations of the image in the potential feature space, enabling better identification of normal and defective areas in urban rail images. This greatly improves the model's detection capability and solves the problem of high false alarm and false alarm rates in current urban rail inspection models.

[0042] In terms of detection speed, this invention adopts Tensor RT optimization strategy and multi-threaded parallel inference to accelerate model inference, optimizes model network structure and computational accuracy, and makes full use of machine hardware resources for image data loading, preprocessing, inference and post-processing, thus solving the problem that the current model detection speed in urban rail transit inspection cannot meet the real-time requirements. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0044] Figure 1 A flowchart of the real-time detection method for multiple types of urban rail defects based on improved YOLO v8 provided by the present invention;

[0045] Figure 2 This is a schematic diagram of the track condition inspection equipment provided by the present invention;

[0046] Figure 3 This is a comparison image of urban rail transit before and after enhancement provided by the present invention;

[0047] Figure 4 This is a schematic diagram of the structure of the disease detection model provided by the present invention;

[0048] Figure 5 A schematic diagram of the feature reconstruction module provided by the present invention;

[0049] Figure 6 A schematic diagram illustrating the reconstruction process of image appearance features provided by the present invention in the feature reconstruction module;

[0050] Figure 7 A schematic diagram of the inference process using Tensor RT optimization strategy and multi-threaded parallel inference provided by the present invention;

[0051] Figure 8 This is a schematic diagram illustrating the relationship between the number of detected images and the detection speed provided by the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] like Figure 1 As shown in the figure, this invention discloses a real-time detection method for multiple types of defects in urban rail transit based on an improved YOLO v8, including the following steps:

[0054] S1. Construct a multi-category image dataset of urban rail transit defects;

[0055] S2. Based on the YOLO v8 network framework, a feature reconstruction module is added to the detection head of the YOLO v8 network to build a disease detection model. The feature reconstruction module uses different loss functions to optimize normal data and disease data respectively, and performs lossless reconstruction of normal representation and directional displacement and strong feature reconstruction of disease representation relative to the cluster center of normal representation.

[0056] S3. Train the disease detection model based on a multi-category disease image dataset;

[0057] S4. Inferring from real-time track data based on the trained disease detection model.

[0058] The following is a further explanation of each of the above steps.

[0059] S1. Construction of a multi-category disease image dataset, specifically including:

[0060] like Figure 2 As shown, a track condition inspection device is installed at the bottom of the urban rail comprehensive inspection train. The track condition inspection device is equipped with five high-speed line array cameras, so that the imaging area covers the entire track bed (including rails, fasteners, track bed surface, etc.). The five high-speed line array cameras on the track condition inspection device can acquire high-resolution grayscale images of 1024×1024 under extreme lighting changes, low contrast and image blur conditions. At a trigger frequency of 35kHz, it can achieve a distance of 35 meters per second under the condition of longitudinal 1.0mm / pixel, thus providing detailed and fast photographs of the tunnel structure.

[0061] Based on the normal images collected by the track condition inspection equipment, as well as images of six types of track facility defects, including foreign objects in the track bed, missing elastic clips, misaligned fasteners, loose bolts, reversed elastic clips, and broken elastic clips;

[0062] Analysis revealed a scarcity of labeled samples for four types of defects: misaligned fasteners, loose bolts, reversed spring clips, and broken spring clips. This makes it difficult for the model to learn effective visual representations from these scarce samples during training, thus affecting its robustness in distinguishing facility instances and different defects, and ultimately limiting the model's generalization ability. Therefore, the scarcity of defect images poses a potential obstacle to improving model performance, potentially leading to insufficient accuracy and reliability in practical applications.

[0063] This invention defines foreign objects in the track bed as objects that do not belong to urban rail transit and its supporting facilities and equipment. Although there is a significant imbalance in the number of labeled samples of foreign objects in the track bed compared to other defect categories, the visual characteristics of foreign objects in the track bed are significantly different from those of other facility defects. This difference helps to alleviate the performance degradation that may be caused by sample imbalance during model training. By emphasizing these differences in visual characteristics, the model can enhance its recognition ability to a certain extent when facing the detection of foreign objects in the track bed, thereby more effectively supporting the safe operation of urban rail transit.

[0064] To improve the generalization ability of the improved YOLO v8 detection model and solve the problem of imbalance of disease categories in the dataset, geometric transformation and image enhancement methods are used to expand the insufficient number of one or several disease image samples. After expansion, the final multi-class disease image dataset is obtained.

[0065] Data augmentation methods include geometric transformations such as flipping, rotating, cropping, translating, and scaling, as well as image enhancement operations such as brightness and contrast adjustment. Among these image enhancement methods, brightness and contrast adjustment effectively augment the dataset by altering the global brightness value of the original image and the brightness differences between different regions. These enhancement methods can simulate the varying lighting conditions that may occur between different urban rail lines, enabling the model to capture visual representations highly relevant to defect detection even under poor or uneven lighting conditions. This adaptability to lighting changes enhances the model's robustness, allowing it to maintain good detection performance in various complex environments. The comparison before and after image augmentation is shown below. Figure 3 As shown.

[0066] By combining geometric transformations and image enhancement methods, the diversity and representativeness of the dataset are effectively improved, providing more comprehensive and richer image samples for model training, thereby enhancing the model's generalization ability and accuracy in practical applications.

[0067] To ensure that the improved model can deeply learn the abnormal features of track facility defect areas during training and exhibit excellent generalization ability when faced with subtle differences between different defects, this invention meticulously constructed a track inspection dataset. This dataset originates from a large number of raw grayscale images collected by a high-speed linear array camera system installed on the underside of an urban rail integrated inspection train, containing images of multiple typical defect categories. For image samples of scarce defect categories, this invention also employs data augmentation techniques to expand the dataset, increasing its diversity and sample size, thereby helping the model to more comprehensively learn the characteristics of different defect categories.

[0068] S2. In the real and complex environment of urban rail transit sections, only a small amount of labeled data on facility and equipment defects and a large amount of unlabeled normal data can usually be obtained. Furthermore, some defective areas are small and have weak feature information, making it difficult for the defect detection model to learn the visual semantic features of various defects in multiple facilities from a very limited number of defect images. This poses a significant challenge to defect classification and localization. Therefore, this invention improves upon the inherently fast single-stage YOLO v8 model by adding a feature reconstruction module, effectively alleviating the difficulty of the model learning defect representations.

[0069] Specifically, such as Figure 4 As shown, the disease detection model includes a backbone network, a neck network, and a detection network;

[0070] The backbone network uses deep convolutional networks and residual networks to extract visual representation information at different levels of the image, providing rich details and semantic information for subsequent networks.

[0071] The neck network adopts the PANet network structure to propagate strong semantic information from the top to the bottom. PANet is different from the traditional Feature Pyramid Network (FPN). It propagates the strong localization features of the bottom layer to the top layer. By fusing the features of different levels, it outputs feature maps of three scales: 20×20, 40×40, and 80×80.

[0072] The detection network consists of three detection heads, each corresponding to a feature map of three scales transmitted by the neck network, which are used to detect large, medium and small targets respectively. All three detection heads use a decoupled-head structure, and each detection head contains two branches, which perform classification and regression tasks respectively. The two tasks are optimized using different loss functions, which is beneficial for the refined prediction of disease type and location.

[0073] This invention adjusts the overall network architecture by adding a feature reconstruction module to the network structure of the detection head. The improved model can further achieve accurate identification of normal and defective features. The final optimization strategy of the model combines the bounding box regression loss and classification loss of the original method with the feature reconstruction loss proposed in this invention. It is trained on the urban rail image dataset, so that the improved YOLO v8 target detection algorithm can be better applied to the track defect inspection task.

[0074] The detection of defects is usually based on the difference between the representation of normal data and defect data in the latent feature space. The feature reconstruction module used in this invention reconstructs the input data from the latent representation, making the visual representation information extracted by the network more robust. This helps the model learn the different data distributions of normal features and defect features, thereby improving the detection capability of urban rail defect areas.

[0075] Specifically, such as Figures 5-6 As shown, the feature reconstruction module consists of four convolutional blocks responsible for feature extraction and an SE (Squeeze-and-Excitation) attention mechanism weighted according to the correlation of feature map channels. The specific method of this module is to optimize different loss functions for normal data and disease data to achieve lossless reconstruction of normal representation and directional displacement and strong feature reconstruction of disease representation relative to the cluster center of normal representation.

[0076] Loss function L for normal data reconstruction normal Represented as:

[0077]

[0078] Where, x i This represents the value of the i-th normal data point before reconstruction. Let represent the value reconstructed from the i-th normal data point, and n represent the number of normal data samples. Under the constraint of this loss function, the output of the normal image representation after passing through the feature reconstruction module can remain consistent with the representation at the input. This ensures that the network can accurately recover the original features of the normal image, making it more robust in information extraction, able to resist potential noise and interference, and improving the model's performance and generalization ability in practical applications.

[0079] Loss function L for disease data reconstruction anomaly The expression is:

[0080] L anomaly =L direction +L rec

[0081] Among them, L direction For directional loss, L rec For reconstruction loss; directional loss L direction The feature maps of diseased data before and after reconstruction, and the feature map of cluster centers of normal data before reconstruction, are unfolded into vectors. Cosine similarity is used to control the reconstruction of diseased data in the feature space along the direction that is significantly different from that of normal data; reconstruction loss L rec The degree of enhancement of disease characteristics is controlled by a preset intensity threshold.

[0082] Among them, the directional loss L direction The expression is:

[0083]

[0084] Where, x i ' represents the value of the i-th outlier before reconstruction, x' i ' represents the reconstructed value of the i-th outlier, n' represents the sample size of the outlier, and c represents the cluster centers of the normal data; directional loss L direction It is used to reconstruct the original features of the non-abnormal areas in disease images, while enhancing the features of the abnormal areas.

[0085] Reconstruction loss L rec The expression is:

[0086]

[0087] Where m is a hyperparameter, and L is the reconstruction loss. rec It is used to control the degree of enhancement of abnormal region features during the feature reconstruction of disease images.

[0088] The feature reconstruction module proposed in this invention can not only restore the details of normal areas, but also highlight the important features of abnormal areas, thereby providing richer information for subsequent analysis and processing. Therefore, it significantly improves the recognition effect of disease images, effectively enhances the identifiability of abnormal area features, and improves the detection effect of the model.

[0089] S3. The disease detection model constructed in S2 is trained and validated using the multi-category disease image dataset from S1. During dataset partitioning, this invention rationally allocates image samples for each category, dividing the data into a training set and a validation set at a 9:1 ratio. This partitioning ensures that the model can fully engage with multiple disease features during training, thereby acquiring richer disease recognition capabilities. Simultaneously, independent testing on the validation set effectively validates the model's generalization performance and stability in practical applications.

[0090] The S4 urban rail comprehensive inspection train can reach a maximum speed of 400 km / h, which places stringent requirements on the real-time detection speed of the defect detection model. To meet the real-time requirements of inspection tasks, this invention provides an in-depth theoretical analysis of the optimization strategies and multi-threading technology of the TensorRT framework, elucidating their feasibility in accelerating model inference, and applying them to the improved model to achieve high-speed inference. Furthermore, this invention systematically analyzes the relationship between detection speed and the number of images detected per cycle, determining the range of image processing quantities that optimizes the defect detection model's detection speed.

[0091] 1) When the disease detection model performs inference on real-time track data, it adopts the TensorRT optimization strategy and the multi-threaded parallel strategy.

[0092] Specifically, the TensorRT deep learning optimization library accelerates the model inference process. Its main principle is to optimize the model computation graph and the method of quantizing model parameters, reducing computational overhead and memory access during inference, thus achieving accelerated model inference. This invention applies TensorRT optimization strategies to an improved YOLO v8 defect detection model. A high-performance runtime engine generated based on a track inspection vehicle server environment is serialized and stored on disk. During detection tasks, the engine is deserialized and loaded into memory for inference. The model's runtime engine employs layer fusion and half-precision (fp16) optimization strategies. These operations significantly reduce the computational load and memory consumption during the model's forward inference process, effectively improving the utilization of computing resources. Layer fusion combines multiple operations into a single execution, reducing unnecessary computational steps; while fp16 optimization reduces data storage and transmission while maintaining accuracy. These optimization strategies work together to make model inference more efficient, providing strong support for meeting real-time detection requirements under high-speed operating conditions.

[0093] Furthermore, this invention employs multi-threading technology to achieve parallel inference of the model. By separating the core code of the disease detection model's inference algorithm and rewriting and encapsulating it using multi-threading, the number of threads can be dynamically adjusted according to actual needs to flexibly cope with different load conditions. This not only improves the system's resource utilization efficiency but also significantly enhances inference performance, enabling the model to operate stably under high load conditions. Through this optimization, the improved model can better meet the real-time inference requirements of the inspection system, achieving accurate and rapid disease detection even during high-speed travel and complex environments, providing crucial assurance for the system's stability and reliability.

[0094] 2) Optimal range setting for the number of images detected in a single run.

[0095] In urban rail transit integrated detection systems, the number of images detected per batch directly impacts memory requirements and computational throughput, thus affecting detection efficiency. A larger number of images allows for better utilization of the GPU's parallel computing capabilities and enables the GPU to warm up to its optimal operating state, helping to reduce the average inference time per image. However, as the number of images increases, so does the usage of video memory. When the hardware load limit is reached, it can lead to memory swapping and competition for computational resources, slowing down inference speed and even causing system instability. Therefore, determining the number of images detected per batch is crucial. This invention was tested in an experimental system environment to determine the optimal balance between model performance and system load for the number of images processed per batch.

[0096] Next, a specific example will be used to further illustrate the method and performance of the present invention.

[0097] 1. Construction of a multi-category disease image dataset:

[0098] 1) Configuration of track condition inspection equipment

[0099] The image data used in the experiment came from five high-speed linear scan cameras mounted on the track inspection equipment. The image type was a high-resolution grayscale image of 1024×1024. Each of the five linear scan cameras could cover an area of ​​81° directly below the track, successfully achieving imaging coverage of the entire track bed (including rails, fasteners, track bed surface, etc.).

[0100] 2) Disease data analysis and expansion

[0101] This invention detects six categories of track facility defects: foreign objects in the track bed, missing spring clips, misaligned fasteners, loose bolts, reversed spring clips, and broken spring clips. The number of labels for each category is shown in Table 1.

[0102] Table 1 Number of labels for various diseases

[0103] Disease categories Foreign objects in the track bed Missing spring bar crooked fastener bolts came loose Reverse loading of the spring bar spring clip breaks Number of tags 1635 165 16 7 5 3

[0104] The number of disease image labels after expansion is shown in Table 2, using data augmentation techniques including geometric transformation and image enhancement methods.

[0105] Table 2. Number of labels for various diseases after enhancement

[0106] Disease categories Foreign objects in the track bed Missing spring bar crooked fastener bolts came loose Reverse loading of the spring bar spring clip breaks Number of tags 1840 165 16 149 5 109

[0107] Data augmentation methods expand the number of labels for rare diseases, providing more diverse and representative image samples for model training, thereby enhancing the model's generalization ability and accuracy in practical applications.

[0108] 3) Custom dataset construction

[0109] A custom dataset was constructed using raw grayscale images collected by an urban rail transit inspection train and images augmented with data augmentation techniques for scarce defect images. This dataset includes not only 2053 defect images but also 300 normal images, enabling the model to learn the significant differences between normal and abnormal appearances of facilities.

[0110] Before training begins, the labelImg tool is used to label the diseased areas in each image and save them as VOC format files. The VOC format is then converted into a dataset format that YOLO can recognize, which includes the location of the label center point and the width and height of the labeled area.

[0111] In constructing the dataset, this invention rationally divided the images for each category, using a 9:1 ratio to split the data into a training set and a validation set. This aims to ensure sufficient learning and reasonable validation of the model during training. The training set contains 2103 images, while the validation set contains 250 images. This ensures sample diversity during training and provides the necessary testing foundation for the subsequent validation phase, enabling the model to exhibit stronger adaptability and reliability in practical applications.

[0112] 2. Construction of disease detection model.

[0113] 1) Adjustment of overall network architecture and optimization strategies

[0114] By analyzing the network structure and optimization strategies of the existing YOLO v8 algorithm, and considering the needs of urban rail transit defect inspection tasks, a feature reconstruction module was added to the detection head. This allows the features of normal and defective urban rail transit images extracted by the backbone network, as well as the multi-level features fused by the neck network, to be fully applied to the task of identifying and locating abnormal areas.

[0115] The improved YOLO v8 network includes three optimization strategies: regression loss, classification loss, and feature reconstruction loss. The regression loss is primarily used to predict the location and size of the detected target. It calculates the positional error between the predicted bounding box and the ground truth bounding box to optimize the accuracy of the bounding box, ensuring that the predicted box better encloses the detected object. This invention uses IoU (Intersection over Union) as the regression loss. The classification loss is used to predict the category of the target, ensuring that the model can accurately identify the disease type.

[0116] This invention uses binary cross-entropy loss as the classification loss. The feature reconstruction loss is used to restore the representation of normal regions in the image and increase the distance between the representations of normal and abnormal regions, enabling the model to better identify and locate the areas where diseases occur. The feature reconstruction loss function used consists of the two objective functions mentioned above in this invention.

[0117] 2) Feature Reconstruction Module

[0118] This invention proposes a feature reconstruction module that combines a loss function with a self-attention mechanism. The core function of this module is to perform deep analysis and reconstruction of the input feature map to improve the model's ability to separate disease features. Specifically, the feature reconstruction module adaptively reconstructs the feature map according to different disease categories, enabling the model to more accurately distinguish between normal and diseased features, thereby improving detection performance.

[0119] During implementation, the feature reconstruction module introduces a self-attention mechanism to dynamically adjust feature weights. This not only helps the model effectively focus features across different disease categories but also reduces the overlap between disease and normal features. Combining the characteristics of the self-attention mechanism, the feature reconstruction module can significantly improve the response of diseased regions by focusing on salient features, while suppressing the response of normal regions, thus providing clearer feature support for subsequent classification and localization tasks.

[0120] Furthermore, the feature reconstruction module introduces additional constraints to the loss function, enhancing the model's ability to distinguish between normal and abnormal samples. In this way, the model can better classify and predict each type of disease, while also exhibiting stronger robustness and generalization ability when faced with subtle differences in disease features. This multi-layered feature reconstruction and optimization design makes the improved YOLO v8 algorithm more suitable for multi-category disease detection tasks in railway facilities, achieving more accurate and reliable detection results.

[0121] In the implementation process, this invention incorporates a feature reconstruction module into the detection head of the YOLO v8 algorithm model, forming an improved YOLO v8 algorithm. The experimental evaluation metrics, experimental environment, and experimental setup are described below. Finally, the effectiveness of the improved method is verified through analysis of the experimental results.

[0122] This invention uses precision and recall to evaluate the accuracy performance of the improved model across different categories of urban rail transit facility defects, and uses the comprehensive evaluation metrics F1 score and mean precision (mAP) to measure the overall performance of the model. Precision refers to the proportion of samples predicted as positive by the model that are actually positive; recall represents the proportion of all samples that were actually positive that the model correctly predicted as positive; the F1 score is the harmonic mean of precision and recall; and the mean precision is calculated by summing the average of the areas under the precision-recall curves (PR curves) for each predicted category at different recall rates, and then averaging the sums as the overall evaluation metric for the model's accuracy across all categories. The calculation formulas for each metric are as follows:

[0123]

[0124] In the formula, TP is the number of positive samples correctly detected; FP is the number of negative samples incorrectly detected as positive samples; FN is the number of positive samples incorrectly detected as negative samples; N is the number of classes to be detected; AP is the average precision across different classes, calculated as follows:

[0125]

[0126] For each recall value r, the corresponding precision value Precision(r) is calculated. The cumulative precision is obtained through integration and used as the average precision for the current class, represented as the area under the curve in the PR curve. The calculation of AP needs to consider the model's performance at different IOU thresholds, and IOU = 0.5 is usually selected for evaluation.

[0127] Regarding the experimental environment, the hardware configuration used in this invention's experiments consisted of an Nvidia GeForce RTX 3090 series graphics card, 24GB of memory, and an AMD EPYC 734316-Core CPU. The software operating system was Ubuntu 20.04, with the following main components: Python 3.8, PyTorch 1.12.1, CUDA 11.3, and CUDNN 8.0.

[0128] The experimental setup for this invention is as follows: The improved YOLO v8 s model is used for training, with 300 training epochs. To fully utilize hardware resources for parallel training, the batch size is 16 images per epoch. The optimization algorithm used in the model training process is stochastic gradient descent (SGD), with an initial learning rate of 0.01, a weight decay coefficient of 0.0005, and a momentum of 0.937 to accelerate the optimization process and reduce oscillations. Furthermore, early stopping is employed to avoid unnecessary training iterations, saving computational resources and time. This experiment sets the patience threshold for early stopping to 100, representing the maximum number of iterations allowed before mAP no longer improves; training will stop if this number is exceeded. The mosaic enhancement probability is set to 1, and four images are stitched together online for training. Mosaic enhancement is disabled in the last ten epochs of training.

[0129] Finally, this invention trains the model using both the pre- and post-enhancement datasets, and verifies the effectiveness of the proposed method by comparing the results before and after model improvement. The comparison of detection results is shown in Table 3.

[0130] Table 3 Comparison of detection results before and after data augmentation and model improvement.

[0131] Has the model been improved? Is data augmentation necessary? Recall mAP F1 score improve Before data augmentation 0.733 0.769 0.740 No improvement Before data augmentation 0.678 0.755 0.730 improve After data augmentation 0.955 0.951 0.910 No improvement After data augmentation 0.952 0.948 0.890

[0132] The data augmentation method employed in this invention significantly improves the performance of the improved model, achieving improvements of 17.0% to 22.2% in recall, mean precision, and F1 score compared to the un-augmented model. Furthermore, the improvement also reaches 16.0% to 27.4% in the un-augmented model. The results demonstrate that the geometric transformation and image augmentation methods employed in this invention can effectively simulate real-world urban rail transit scene data under different scales, regions, and lighting conditions, successfully mitigating the long-tail problem and enhancing the model's generalization ability across different defect categories. This method not only improves the model's robustness but also lays a solid foundation for addressing diverse practical application scenarios.

[0133] On the un-data-augmented dataset, the improved model achieved improvements of 1.0% to 5.5% in recall, mean precision, and F1 score compared to the un-augmented model. On the data-augmented dataset, the improved model also achieved improvements of 0.3% to 2.0%. These results demonstrate that the improved model is more capable of identifying visual feature differences between images of normal facilities and images of damaged facilities, thus effectively improving detection performance.

[0134] The dataset used in this invention covers six categories of facility defects. The improved model achieved improvements of 0.3% to 1.7% in recall and average precision across multiple categories. This indicates that the overall performance of the model can be significantly improved by employing the feature reconstruction module and optimization strategy proposed in this invention. Although the improved model showed a slight decrease in average precision for the bolt loosening category (by 0.6%), it still demonstrated better performance in other categories. Specific detection results are as follows: Figure 7 As shown.

[0135] The above results demonstrate that the method proposed in this invention enables the model to more accurately capture the characteristics of different defect categories in the task of classifying and locating defects in urban rail transit facilities, thereby improving the overall detection capability and providing a solid guarantee for the reliability of track inspection.

[0136] 3. Reasoning acceleration methods

[0137] 1) Analysis and Application of Tensor RT and Multithreading Technology

[0138] like Figure 7 As shown, by installing TensorRT, a high-performance deep learning inference library for NVIDIA GPUs provided by NVIDIA, the trained .pt format PyTorch model is exported to ONNX (Open Neural Network Exchange) format. ONNX is an open format that can convert models between different deep learning frameworks. Subsequently, the ONNX model is converted into a .engine format TensorRT engine file. This format of model file realizes the optimization of the model computation graph and the quantization of model parameters, which improves the utilization efficiency of computing resources.

[0139] Meanwhile, the core code of the model inference algorithm was modularized to facilitate subsequent multi-threaded rewriting. Next, multi-threading was used to distribute the various tasks in the inference process into different threads, enabling them to execute in parallel. This implementation allows the system to monitor and dynamically adjust the number of threads in real time based on the actual load, specifically increasing threads to improve processing capacity when the load is high and decreasing threads to conserve resources when the load is low.

[0140] This invention conducted detailed experimental comparisons to verify that the TensorRT framework optimization strategy and the fully multi-threaded parallel inference method effectively improve the model's speed performance, meeting the requirements of real-time inference for urban rail integrated trains. Specifically, this invention tested the models before and after the improvement, comparing the running speed of the models using the TensorRT method, the multi-threaded inference method, and a combination of both. The experimental results are shown in Table 4. The experimental results show that, with the combined approach, the improved model enhances detection performance while only taking 1ms longer to detect a single image compared to the original model, effectively meeting the requirements of real-time inference.

[0141] Meanwhile, ablation experiments were conducted on the two acceleration methods to further verify their effectiveness. Experimental results show that both the TensorRT method and the multi-threaded inference method can effectively improve the model's inference speed, providing a solid technical foundation for real-time monitoring and analysis of urban rail integrated trains.

[0142] Table 4. Results of TensorRT Framework and Multithreaded Accelerated Inference

[0143]

[0144] 2) Optimal range setting for the number of images detected in a single run

[0145] Determining the number of images to be detected in a single pass is crucial for detection speed. This invention was tested in an experimental system environment to determine the optimal balance between model performance and system load regarding the number of images processed per pass. The experimental results are as follows: Figure 8 As shown.

[0146] In this invention, the step size for the number of images detected per run is 200. Experimental results show that the number of images that the system can detect per second is closely related to the number of images processed per run. When the number of input images is small, system resources are not fully utilized, resulting in the model's detection speed not reaching its optimal performance. However, as the number of images processed gradually increases, the parallel computing power of the GPU can be better utilized, thereby significantly improving the inference speed. In the experiments, the model achieved its optimal inference speed when the number of images processed reached 600. This indicates that the system's performance is significantly optimized with an appropriate number of images.

[0147] However, when the number of images processed exceeds 600, excessive memory usage leads to a decrease in inference speed. This indicates that an excessively large input not only fails to further improve performance but may also cause an overall performance decrease due to memory overflow. These results suggest that in urban rail transit inspection, the method of this invention needs to adjust the number of images detected per cycle based on hardware resources and actual needs to achieve optimal detection efficiency and system stability.

[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0149] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A real-time detection method for multiple types of defects in urban rail transit based on an improved YOLO v8, characterized in that, Includes the following steps: Construct a multi-category image dataset of urban rail transit defects; Based on the YOLO v8 network framework, a feature reconstruction module is added to the detection head of the YOLO v8 network to construct a disease detection model. The feature reconstruction module is optimized by using different loss functions for normal data and disease data respectively, to perform lossless reconstruction of normal representation and directional displacement and strong feature reconstruction of disease representation relative to the cluster center of normal representation. The disease detection model is trained based on the multi-category disease image dataset; The trained defect detection model is used to infer information from real-time track data.

2. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, The process of constructing the multi-category disease image dataset includes: Track condition inspection equipment is installed at the bottom of the urban rail comprehensive inspection train. Five high-speed line array cameras are mounted on the track condition inspection equipment so that the imaging area covers the entire track bed. Based on the normal images collected by the track condition inspection equipment, as well as images of six types of track facility defects, including foreign objects in the track bed, missing elastic clips, misaligned fasteners, loose bolts, reversed elastic clips, and broken elastic clips; For a single or several types of disease image samples that are insufficient in number, geometric transformation and image enhancement methods are used to expand them, and after expansion, the final multi-category disease image dataset is obtained.

3. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, The disease detection model includes a backbone network, a neck network, and a detection network; The backbone network uses deep convolutional networks and residual networks to extract visual representation information of images at different levels; The neck network adopts the PANet network structure to propagate strong semantic information from the top layer to the bottom layer, and propagates the strong localization features of the bottom layer to the top layer. By fusing the features of different levels, feature maps of three scales, 20×20, 40×40 and 80×80, are output. The detection network comprises three detection heads, each of which receives feature maps of three scales transmitted by the neck network and performs detection on large, medium and small targets respectively. All three detection heads use a decoupled head structure, and each detection head contains two branches, which respectively complete the classification task and the regression task.

4. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 3, characterized in that, Each of the detection heads is equipped with the feature reconstruction module; The feature reconstruction module consists of four convolutional blocks responsible for feature extraction and an SE attention mechanism weighted according to the correlation of feature map channels.

5. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, Loss function L for normal data reconstruction normal Represented as: Where, x i This represents the value of the i-th normal data point before reconstruction. Let represent the value after reconstruction of the i-th normal data point, and n represent the sample size of the normal data.

6. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, Loss function L for disease data reconstruction anomaly The expression is: L anomaly =L direction +L rec Among them, L direction For directional loss, L rec For reconstruction loss; directional loss L direction The feature maps of diseased data before and after reconstruction, and the feature maps of cluster centers of normal data before reconstruction, are unfolded into vectors. Cosine similarity is used to control the reconstruction of diseased data in the feature space along directions that are significantly different from normal data; reconstruction loss L rec The degree of enhancement of disease characteristics is controlled by a preset intensity threshold.

7. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 6, characterized in that, Directional loss L direction The expression is: Where, x i ' represents the value of the i-th outlier before reconstruction, x' i ' represents the reconstructed value of the i-th outlier, n' represents the sample size of the outlier, and c represents the cluster centers of the normal data; directional loss L direction It is used to reconstruct the original features of the non-abnormal areas in disease images, while enhancing the features of the abnormal areas.

8. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 7, characterized in that, Reconstruction loss L rec The expression is: Where m is a hyperparameter, and L is the reconstruction loss. rec It is used to control the degree of enhancement of abnormal region features during the feature reconstruction of disease images.

9. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, The disease detection model employs a TensorRT optimization strategy and a multi-threaded parallel strategy when inferring from real-time track data.

10. The real-time detection method for multiple types of urban rail transit defects based on improved YOLO v8 according to claim 1, characterized in that, When the disease detection model infers from real-time track data, the number of images detected in a single test is 600.