DETR-based end-to-end catenary part defect detection method
Through the end-to-end contact network component defect detection method based on DETR, combined with Co-DETR, CFPT and QAT technologies, the existing detection methods have solved the problems of high computational complexity and poor robustness, and achieved efficient and real-time contact network defect detection, which significantly improved detection accuracy and resource efficiency.
Patent Information
- Application Number
- CN202510148664.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing contact network component defect detection methods have problems such as high computational complexity, large resource occupation, poor robustness and low detection accuracy of small targets, which are difficult to meet the needs of real-time and efficientness.
Using the end-to-end contact network component defect detection method based on DETR, image features are extracted through the Backbone module, the Encoder module encodes the status vector, and the Decoder module performs target query and image features to interact to achieve contact network defect detection. Co-DETR's multi-objective collaborative training, CFPT feature pyramid and QAT quantized perception training are introduced to improve the detection performance and resource efficiency of the model.
It realizes efficient and real-time contact network defect detection, improves the model's image understanding and target recognition capabilities, significantly improves the recall and detection accuracy of small targets, and meets the needs of embedded and real-time monitoring scenarios.
Smart Images

Figure CN119941708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of railway infrastructure equipment detection, and in particular to an end-to-end contact network component defect detection method based on DETR. Background Art
[0002] As an important railway infrastructure, the overhead contact network is the core guarantee for the safe operation of trains. Therefore, the daily inspection of key components of the overhead contact network is an important part of line operation and maintenance. Traditional overhead contact network component inspection usually relies on manual inspection, which is labor-intensive, inefficient, limited in accuracy and highly dangerous. Therefore, the overhead contact network component defect detection technology based on target detection has received widespread attention in the field of railway intelligent operation due to its characteristics of automation, unmanned operation and high detection efficiency. Traditional target detection methods, such as algorithms based on edge detection and morphological processing, have certain effects on specific types of defects (such as cracks and wear), but their detection performance depends on artificially designed features, making it difficult to cope with a variety of defect types and complex background interference, and their robustness in the face of complex scenarios is poor. With the rapid development of deep learning technology, target detection algorithms based on deep convolutional networks have been tried to be applied to the detection of defects in overhead contact network components. However, common general target detection algorithms based on deep convolutional networks have certain problems in the task of detecting defects in overhead contact network components. At present, commonly used target detection algorithms are divided into detection algorithms based on Faster R-CNN and detection algorithms based on YOLO. Faster R-CNN is a two-stage target detection algorithm. First, a series of regional candidate frames are extracted through the regional candidate network RPN, and then the candidate frames are positionally corrected. Therefore, its computational complexity and resource occupation are large, and it is not suitable for tasks with high real-time requirements such as contact network component defect detection. In addition, the Faster R-CNN algorithm uses RoI Pooling to replace the feature pyramid and maps multi-scale information to the same feature dimension, resulting in the loss of key information of small targets. The YOLO algorithm is a single-stage target detection algorithm with low resource consumption and fast inference speed. However, the feature map resolution is reduced by multiple downsampling in the YOLO network, resulting in low detection accuracy for small targets, and its resolution ability of complex scenes is insufficient, with many false detections and missed detections. Therefore, the present invention proposes an end-to-end contact network component defect detection method based on DETR to solve the problems existing in the prior art. Summary of the invention
[0003] In view of the above problems, the present invention proposes an end-to-end contact network component defect detection method based on DETR, which uses an end-to-end target detection method based on DETR to enhance the image understanding and target recognition capabilities of the model, and improves small target detection, so as to better achieve the goal of automated detection of contact network defects.
[0004] To achieve the purpose of the present invention, the present invention is implemented by the following technical solutions: a DETR-based end-to-end contact network component defect detection method, including a Backbone module, an Encoder module and a Decoder module, wherein the Backbone module, the Encoder module and the Decoder module are all based on the DETR detection model, and the Backbone module is used to extract texture and semantic features in the collected image and convert the input two-dimensional image into a feature vector; The Encoder module is the encoder of the network, which is used to encode the image features extracted by the Backbone module into the state vector required for decoding; The Decoder module uses the state vector extracted in the Encoder module as the initial state, interacts with the image features through a set of learnable object queries, and detects defects in contact network components.
[0005] A further improvement is that the Backbone module is an EfficientNet-B0 backbone network, and the EfficientNet-B0 backbone network is used to improve the frame rate of contact network component identification and the real-time performance of feature extraction.
[0006] A further improvement is that the Encoder module adopts a Transformer Encoder structure and a deformable attention, and the deformable attention is used to focus on features near key points of the recognition target.
[0007] A further improvement is that in the Decoder module, the target query Object Queries performs self-attention and cross-attention calculations through a multi-layer Transformer network and the global information in the state vector to learn the relationship between each query and the target position. After multi-layer decoding, the Decoder module outputs a set of prediction results, including the target category and bounding box coordinates, converting the target detection task into a set prediction problem. Through Hungarian matching, each query is controlled to uniquely correspond to the real target, thereby completing the detection of defective targets in the contact network.
[0008] A method for end-to-end contact network component defect detection based on DETR, comprising the following steps: S1: Collect contact network scene data, preprocess and enhance the data, and build a contact network component detection dataset; S2: Based on the DETR detection model, the model is pre-trained using the general target detection dataset, and the model is fine-tuned and trained using the contact network component detection dataset; S3: After completing model training, quantize the model parameters and convert the single-precision model used in training into an INT8 quantized model suitable for deployment on edge computing devices; S4: Deploy the INT8 quantization model on the contact network detection equipment, convert the data collected by the camera in real time into the format required by the model, input it into the model for inference, and input the detection results into the cloud for data analysis and real-time monitoring.
[0009] A further improvement is that: S1 comprises the following steps: Collect contact network scene data, including component images under various working conditions, use annotation tools to annotate the targets, and generate annotation files in COCO format, including category labels and target bounding boxes; Clean the collected data, remove images of poor quality or not meeting the requirements, and then standardize the data to ensure data consistency; Data enhancement is used to improve the sample distribution of the sample data set. First, rotation and translation transformations are used to increase the diversity of samples. Then the Copy-Paste data enhancement method is used to enhance the data diversity of small components and balance the detection performance of multi-category targets. Mosaic data enhancement is then used to provide a diverse visual background, exposing the defects of contact network components under various backgrounds. Thus, a contact network component detection data set is constructed.
[0010] A further improvement is that: S2 comprises the following steps: Based on the DETR detection model, the model is pre-trained using the COCO dataset and the Object 365 general object detection dataset to enable it to have general object detection capabilities. The pre-trained model is used as the basis for fine-tuning. During the training process, the loss function of the training process is monitored and the optimal model weights are saved. The contact network component detection dataset is used to fine-tune the model to improve its performance on the contact network component detection task. Then, the defect data in the contact network component detection dataset is used to train the model. The backbone network weights of the pre-trained model are retained, and only the classification head and some new model layers are updated to make the model adapt to the contact network component defect detection task.
[0011] Further improvements are as follows: In training, the multi-target collaborative training proposed by Co-DETR is introduced, combined with a universal auxiliary head with different one-to-many label assignment paradigms. The universal auxiliary head accepts supervision signals from different tasks, combines tasks with label assignment strategies, and enables the model to obtain effective feature representation in the early stage of training through multi-level and diversified supervision, accelerates the convergence of the training process, and matches each predicted box with multiple real boxes through one-to-many label assignment, providing the encoder with more information from different angles and generating more fine-grained features. At the same time, the CFPT feature pyramid is introduced into the model to allow the high-resolution layer to focus on the details of small targets to improve the recall rate of small targets.
[0012] A further improvement is that in S3, QAT quantization-aware training is used, quantization operations are performed during the training process, pseudo-quantization operations are inserted to simulate quantization errors, and fine-tuning training is performed to allow the model to adapt to the error, and finally an INT8 quantization model is obtained.
[0013] The beneficial effects of the present invention are: 1. The end-to-end target detection algorithm based on the DETR detection model of the present invention does not need to manually design the anchor box size, has good adaptability to targets of different scales, and the end-to-end design adopted does not require post-processing steps such as maximum suppression. The process is more concise and efficient, the model has high parallelism, and the image understanding and target recognition capabilities of the model are enhanced.
[0014] 2. The present invention introduces the multi-objective collaborative training proposed by Co-DETR, combines a universal auxiliary head with different one-to-many label assignment paradigms, combines the task with the label assignment strategy, and enables the model to obtain more effective feature representation in the early stage of training through multi-level and diversified supervision, accelerates the convergence of the training process, shortens the training cycle, and quickly achieves stable performance.
[0015] 3. The CFPT feature pyramid introduced in the present invention allows the high-resolution layer to focus on the details of small targets, significantly improves the recall rate of small targets, improves feature extraction and multi-target correlation, and significantly improves the accuracy and robustness of small target detection, thereby better achieving the goal of automated detection of contact network defects.
[0016] 4. The present invention uses QAT quantization-aware training, performs quantization operations during the training process, inserts pseudo-quantization operations to simulate quantization errors, performs fine-tuning training to adapt the model to the errors, and finally obtains the quantization model and deploys it to the target device to achieve efficient reasoning, faster speed, and less resource consumption, meeting the needs of embedded and real-time monitoring scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall architecture diagram of the present invention. DETAILED DESCRIPTION
[0018] In order to deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with examples. The examples are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0019] Embodiment 1 according to Figure 1 As shown, this embodiment proposes an end-to-end contact network component defect detection method based on DETR, including a Backbone module, an Encoder module and a Decoder module. The image input is converted into image features of different granularities by Backbone, encoded into a state vector by the Encoder, and finally the Decoder is used to complete the identification of the contact network component.
[0020] The Backbone module is used to extract texture and semantic features in the image and convert the input two-dimensional image into a feature vector. The backbone network used is EfficientNet-B0. EfficientNet is an efficient convolutional neural network architecture that aims to achieve higher performance and lower computational costs through a more effective model scaling method. Compared with the traditional ResNet network, EfficientNet achieves higher performance and efficiency while using fewer model parameters. The present invention uses EfficientNet as the Backbone, which helps to further improve the frame rate of contact network component recognition without reducing the overall recognition accuracy of the algorithm, thereby improving the real-time performance of feature extraction.
[0021] The Encoder module acts as the encoder of the network, further encoding the image features extracted by Backbone into the state vector required for decoding. The overall structure of the module adopts the Transformer Encoder structure, and uses deformable attention instead of the global attention used in DETR. The deformable attention network retains the global modeling capability of the self-attention network while having the sparse sampling characteristics of deformable convolution. By focusing on the features near the key points of the recognition target instead of the global attention of the original attention network, the computing power requirements of the network are greatly reduced, the processing delay is reduced, and the real-time performance of the encoding process is improved.
[0022] The Decoder module uses the state vector (state encoding) extracted from the Encoder module as the initial state, and interacts with the image features through a set of learnable object queries to detect defects in contact network components. The input target query is self-attention and cross-attention calculated with the global information in the state vector (state encoding) through a multi-layer Transformer network to learn the relationship between each query and the target position. Finally, after multiple layers of decoding, the Decoder outputs a set of prediction results, including the category and bounding box coordinates of the target. This process transforms the target detection task into a set prediction problem, and ensures that each query uniquely corresponds to the real target through Hungarian matching, thereby completing the detection of defective contact network targets. Similar to the Encoder, the Decoder also uses deformable attention to reduce network resource consumption and processing latency.
[0023] In the training, the multi-target collaborative training proposed by Co-DETR was introduced, combining universal auxiliary heads with different one-to-many label assignment paradigms (such as ATSS and Faster R-CNN). The universal auxiliary head is responsible for accepting supervision signals from different tasks, combining tasks with label assignment strategies, and enabling the model to obtain more effective feature representations at the beginning of training through multi-level and diversified supervision, thereby accelerating the convergence of the training process. One-to-many label assignment enables each predicted box to match multiple real boxes, providing more information to the encoder from different angles and generating more fine-grained features. For example, ATSS pays more attention to dynamic anchor box selection and reduces label noise, while Faster R-CNN focuses on ensuring accurate bounding box regression through fixed matching methods. Through these diverse label assignments, the model can handle more complex object detection tasks. Object detection tasks involve object classification and positioning. Through multi-target collaborative training, the loss function of each task will be merged into a total loss, which is ultimately used to update the parameters of the model. In the inference phase, these auxiliary heads are discarded, so the multi-target collaborative training method does not add additional parameters and computational costs to the original detector.
[0024] The target detection task of contact network components involves multi-scale targets (such as bolts, insulators, slings, etc.), and the size and distribution of components in the scene vary greatly. The original DETR relies on a single global feature modeling, and its ability to capture small targets and multi-scale features is insufficient, resulting in weak small target detection capabilities, insufficient multi-scale target characteristics, and loss of feature details. The introduction of the CFPT feature pyramid allows the high-resolution layer to focus on the details of small targets, significantly improving the recall rate of small targets. CFPT feature pyramid CFPT uses two carefully designed attention modules with linear computational complexity: cross-layer channel attention (CCA) and cross-layer spatial attention (CSA), which can enhance the model's ability to utilize global contextual information and cross-layer multi-scale information.
[0025] The model training is based on the DETR detection model. DETR (Detection Transformer) is a target detection model based on the attention mechanism and Transformer structure. It is an Anchor-Free end-to-end target detection algorithm that does not require manual design of the anchor box size. It has good adaptability to targets of different scales, and the end-to-end design does not require post-processing steps such as maximum suppression (NMS). The process is more concise and efficient, and the model has high parallelism. In addition, DETR can adaptively model global contextual information through the self-attention mechanism, capture long-distance dependencies between targets or between background and target, and help handle target detection tasks in complex scenes. Therefore, compared with target detection algorithms based on convolutional neural networks such as Faster-RCNN and YOLO, DETR has achieved significant performance improvements and has high flexibility and scalability.
[0026] Embodiment 2 according to Figure 1 As shown, this embodiment proposes an end-to-end contact network component defect detection method based on DETR, comprising the following steps: Data acquisition and processing Data collection and annotation: Collect contact network scene data, including component images under various working conditions, such as bolts, insulators, arm structures, string suspension devices, etc. Use professional tools (such as LabelImg) to annotate the targets and generate annotation files in COCO format, including category labels and target bounding boxes.
[0027] Data preprocessing: First, the collected images are cleaned to remove images of poor quality or that do not meet the requirements to ensure the quality of the data set. The images are then standardized to ensure data consistency and facilitate model training and reasoning.
[0028] Data enhancement: The contact network inspection images collected in the previous step have problems such as too few samples and imbalanced positive and negative samples. Therefore, before model training, a series of data enhancements are needed to improve the sample distribution of the sample data set, improve training efficiency, and avoid overfitting and underfitting. After increasing the diversity of samples with conventional rotation and translation transformations, the Copy-Paste data enhancement method is used to enhance the data diversity of small components, balance the detection performance of multi-category targets, and solve the problem of missed detection and false detection caused by category imbalance. Mosaic data enhancement is used to provide a variety of visual backgrounds, so that the defects of contact network components are exposed in various backgrounds, improve the generalization ability of the model, and construct a contact network component inspection data set.
[0029] Model Training Pre-training: Use the COCO dataset and Object 365 and other general object detection datasets to pre-train the model so that it has strong general object detection capabilities. Use the pre-trained model as the basis for fine-tuning. During the training process, monitor the loss function of the training process and save the best model weights.
[0030] Model fine-tuning: Use the contact network component detection dataset to fine-tune the model and improve the performance of the model in the contact network component detection task. The model is trained with the collected and labeled contact network component defect data, the backbone network weights of the pre-trained model are retained, and only the classification head and a small number of new model layers are updated to fine-tune the model to adapt it to the contact network component defect detection task.
[0031] Model Quantization In the actual target detection scenario of contact network components, the detection system needs to be deployed on edge computing devices with limited computing resources such as FPGA and ASCI, while ensuring a certain degree of real-time performance. The trained single-precision model has high computational complexity and low frame rate when running on resource-constrained access devices, which makes it difficult to meet real-time requirements. In order to reduce resource consumption and improve model throughput, it is necessary to quantize the model parameters after completing model training, and convert the single-precision model used in training into an INT8 quantized model suitable for deployment on edge computing devices. Traditional post-training quantization often leads to a decrease in model accuracy, so we use QAT (Quantization-Aware Training) quantization-aware training to perform quantization operations during training. Pseudo-quantization operations are inserted to simulate quantization errors, and fine-tuning training is performed to adapt the model to these errors. Finally, the quantized model is obtained and deployed on the target device to achieve efficient reasoning and operation.
[0032] Deployment and Inference The model is deployed on the contact network detection equipment, the data collected by the camera in real time is converted into the format required by the model, input into the model, perform model inference, and input the model's detection results into the cloud for data analysis and real-time monitoring.
[0033] The end-to-end target detection algorithm based on the DETR detection model of the present invention does not need to manually design the anchor frame size, has good adaptability to targets of different scales, and the adopted end-to-end design does not need post-processing steps such as maximum suppression. The process is more concise and efficient, the model has high parallelism, and the image understanding and target recognition capabilities of the model are enhanced. In addition, the present invention introduces the multi-target collaborative training proposed by Co-DETR, combines a universal auxiliary head with different one-to-many label allocation paradigms, combines the task with the label allocation strategy, and through multi-level and diversified supervision, enables the model to obtain a more effective feature representation in the early stage of training, accelerates the convergence of the training process, shortens the training cycle, and quickly achieves stable performance. At the same time, the present invention introduces the CFPT feature pyramid, which allows the high-resolution layer to focus on the details of small targets, significantly improves the recall rate of small targets, improves feature extraction and multi-target correlation, and significantly improves the accuracy and robustness of small target detection, thereby better achieving the goal of automated detection of contact network defects. In addition, the present invention uses QAT quantization-aware training, performs quantization operations during the training process, inserts pseudo-quantization operations to simulate quantization errors, performs fine-tuning training to allow the model to adapt to the errors, and finally obtains the quantization model and deploys it to the target device to achieve efficient reasoning, faster speed, and less resource consumption, meeting the needs of embedded and real-time monitoring scenarios.
[0034] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A DETR-based end-to-end contact network component defect detection method, comprising a Backbone module, an Encoder module and a Decoder module, characterized in that: The Backbone module, Encoder module and Decoder module are all based on the DETR detection model. The Backbone module is used to extract texture and semantic features in the collected image and convert the input two-dimensional image into a feature vector; The Encoder module is the encoder of the network, which is used to encode the image features extracted by the Backbone module into the state vector required for decoding; The Decoder module uses the state vector extracted in the Encoder module as the initial state, interacts with the image features through a set of learnable object queries, and detects defects in contact network components.
2. The end-to-end contact network component defect detection method based on DETR according to claim 1 is characterized in that: The Backbone module is the EfficientNet-B0 backbone network, and the EfficientNet-B0 backbone network is used to improve the frame rate of contact network component recognition and the real-time performance of feature extraction.
3. The end-to-end contact network component defect detection method based on DETR according to claim 2 is characterized in that: The Encoder module adopts a Transformer Encoder structure and a deformable attention DeformableAttention, and the deformable attention Deformable Attention is used to focus on features near key points of the recognition target.
4. The end-to-end contact network component defect detection method based on DETR according to claim 3 is characterized in that: In the Decoder module, the object query Object Queries performs self-attention and cross-attention calculations through a multi-layer Transformer network and the global information in the state vector to learn the relationship between each query and the target position. After multiple layers of decoding, the Decoder module outputs a set of prediction results, including the target category and bounding box coordinates, converting the target detection task into a set prediction problem. Through Hungarian matching, each query is controlled to uniquely correspond to the real target, thereby completing the detection of defective targets in the contact network.
5. The end-to-end contact network component defect detection method based on DETR according to claim 4 is characterized in that: The following steps are involved: S1: Collect contact network scene data, preprocess and enhance the data, and build a contact network component detection dataset; S2: Based on the DETR detection model, the model is pre-trained using the general target detection dataset, and the model is fine-tuned and trained using the contact network component detection dataset; S3: After completing model training, quantize the model parameters and convert the single-precision model used in training into an INT8 quantized model suitable for deployment on edge computing devices; S4: Deploy the INT8 quantization model on the contact network detection equipment, convert the data collected by the camera in real time into the format required by the model, input it into the model for inference, and input the detection results into the cloud for data analysis and real-time monitoring.
6. The end-to-end contact network component defect detection method based on DETR according to claim 5, characterized in that: The S1 comprises the following steps: Collect contact network scene data, including component images under various working conditions, use annotation tools to annotate the targets, and generate annotation files in COCO format, including category labels and target bounding boxes; Clean the collected data, remove images of poor quality or not meeting the requirements, and then standardize the data to ensure data consistency; Data enhancement is used to improve the sample distribution of the sample data set. First, rotation and translation transformations are used to increase the diversity of samples. Then the Copy-Paste data enhancement method is used to enhance the data diversity of small components and balance the detection performance of multi-category targets. Mosaic data enhancement is then used to provide a diverse visual background, exposing the defects of contact network components under various backgrounds. Thus, a contact network component detection data set is constructed.
7. The end-to-end contact network component defect detection method based on DETR according to claim 5, characterized in that: The S2 comprises the following steps: Based on the DETR detection model, the model is pre-trained using the COCO dataset and the Object 365 general object detection dataset to enable it to have general object detection capabilities. The pre-trained model is used as the basis for fine-tuning. During the training process, the loss function of the training process is monitored and the optimal model weights are saved. The contact network component detection dataset is used to fine-tune the model to improve its performance on the contact network component detection task. Then, the defect data in the contact network component detection dataset is used to train the model. The backbone network weights of the pre-trained model are retained, and only the classification head and some new model layers are updated to make the model adapt to the contact network component defect detection task.
8. The end-to-end contact network component defect detection method based on DETR according to claim 7, characterized in that: During training, the multi-target collaborative training proposed by Co-DETR was introduced, combined with a universal auxiliary head with different one-to-many label assignment paradigms. The universal auxiliary head accepts supervision signals from different tasks, combines tasks with label assignment strategies, and enables the model to obtain effective feature representation in the early stage of training through multi-level and diversified supervision, accelerates the convergence of the training process, and matches each predicted box with multiple real boxes through one-to-many label assignment, providing the encoder with more information from different angles and generating more fine-grained features. At the same time, the CFPT feature pyramid is introduced into the model to allow the high-resolution layer to focus on the details of small targets to improve the recall rate of small targets.
9. The end-to-end contact network component defect detection method based on DETR according to claim 5, characterized in that: In S3, QAT quantization-aware training is used, quantization operations are performed during the training process, pseudo-quantization operations are inserted to simulate quantization errors, and fine-tuning training is performed to allow the model to adapt to the error, and finally an INT8 quantization model is obtained.
Citation Information
Patent Citations
Small-size target detection method based on dynamic anchor frame and Transform
CN116403090A
River surface drifting object detection method and device based on improved DETR and related assembly
CN116630804A
Casting surface defect detection method and system based on improved DETR
CN117314837A
Railway bullet train bottom plate bolt loss fault detection method
CN117689873A
Unmanned aerial vehicle 3D target detection multi-modal fusion method based on Transform
CN118837875A