A method for detecting defects in cotter pins of U-shaped clamps for high-speed railway contact networks

By using deep learning technology to construct a detection method for the cotter pins of U-shaped clamps of high-speed rail contact networks, efficient and accurate automated detection is achieved, which solves the problems of low efficiency and insufficient accuracy of manual detection in existing technologies and improves the inspection efficiency of high-speed rail contact networks.

CN116958079BActive Publication Date: 2025-10-10INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310881126.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-17
Publication Date
2025-10-10
Estimated Expiration
2043-07-17

AI Technical Summary

Technical Problem

In the existing technology, the defect detection of the cotter pin of the U-shaped clamp of the high-speed railway contact network relies on manual judgment, which is inefficient and lacks accuracy, making it difficult to achieve intelligent detection.

Method used

Deep learning technology is used to construct a dataset. A multi-level positioning strategy and an improved recognition network model are used to accurately locate and detect defects in the cotter pins of U-shaped clamps. The CSPDarknet53 backbone network, the improved ASPP module and the PAN structure are combined to achieve multi-scale feature fusion and feature extraction, and the ResNet50 network is used for fault prediction.

Benefits of technology

It improves the accuracy and automation of detection, reduces human errors, significantly improves detection efficiency, simplifies work processes and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958079B_ABST
    Figure CN116958079B_ABST
Patent Text Reader

Abstract

The application discloses a high-speed rail contact network U-shaped hoop split pin defect detection method, and belongs to the target detection field. The method comprises the following steps: step 1: using high-definition pictures collected by a high-speed rail power supply safety detection monitoring system as a training set; step 2: the data in the training set is a picture containing a high-speed rail contact network, after the picture is preprocessed, a multi-stage positioning strategy is used for the target; step 3: the picture labeled in step 2 is used as a training set and is put into a recognition network model for training, and after training weights are obtained, the data set is detected; step 4: the split pin image data positioned in step 3 is input into a defect detection network (ResNet50 network) to realize fault detection of the split pin. The application has high accuracy and automation degree, reduces the complexity of the work flow and the labor cost, and improves the work efficiency of high-speed rail contact network inspection personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of defect detection based on deep learning image processing, and in particular relates to a method for detecting defects in a cotter pin of a U-shaped clamp of a high-speed railway contact network. Background Art

[0002] High-speed rail catenary systems are a crucial component of the high-speed railway system, carrying the crucial task of transmitting and transferring electrical energy during operation. U-shaped clamps are a key fastener in the support and suspension system of the catenary. Whether the cotter pins on the U-shaped clamps are in good working order is a common fault in the U-shaped clamps. Therefore, detecting missing cotter pins is crucial for the safe and stable operation of high-speed trains. Before the widespread interest in deep learning, traditional image-based methods struggled to achieve high precision and accuracy in the field. Consequently, manual identification of defects was common.

[0003] The image data collected by the system still needs to be manually judged whether there are defects. This process is time-consuming, labor-intensive, and inefficient, and fails to achieve intelligent detection. Summary of the Invention

[0004] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method for detecting defects in the cotter pins of the U-shaped clamp of the high-speed railway contact network, so as to reduce the complexity of the work process and labor costs, improve the accuracy and degree of automation of the detection, and ultimately improve the work efficiency of the high-speed railway contact network inspectors.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A method for detecting defects in a U-shaped clamp cotter pin of a high-speed railway contact network comprises the following steps:

[0007] Step 1: Construct a dataset containing images of various states of the cotter pins of the high-speed railway contact network U-shaped clamps.

[0008] Step 2: After pre-processing the image, a multi-level positioning strategy is used for the target;

[0009] Step 3: Use the labeled images in step 2 as the training set, put them into the recognition network model for training, and then test the data set after obtaining the training weights;

[0010] Step 4: Input the cotter pin image data located in step 3 into the defect detection network to realize fault detection.

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] Compared to traditional manual inspection, this method avoids errors caused by human factors such as fatigue and distraction, ensuring continuous and accurate measurement. Its accuracy far exceeds that of manual inspection. Furthermore, this method can complete a large number of inspection tasks in a short period of time, greatly improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a schematic flow chart of the method of the present invention.

[0014] Figure 2 The invention relates to an overall network structure for detecting targets of U-shaped clamps and split pins of a high-speed railway contact network.

[0015] Figure 3 It is a flow chart of the defect detection network of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] This invention addresses the problem of small cotter pins in high-speed rail contact network components, making them difficult to accurately locate and detect. By leveraging deep learning technology, it accurately locates cotter pins within complex images of contact network support devices and detects missing cotter pins with high reliability, demonstrating excellent noise immunity and robustness. Furthermore, the invention boasts a high degree of automation, reducing workflow complexity and labor costs, improving the efficiency of high-speed rail contact network inspectors and significantly facilitating inspections.

[0018] Figure 1 A flowchart of a method for detecting defects in a U-shaped clamp cotter pin of a high-speed railway contact network in a specific embodiment of the present invention is provided. The method mainly includes the following steps:

[0019] Step 1: Use high-definition images collected by the high-speed railway power supply safety detection and monitoring system to build a training set.

[0020] Through the high-speed railway power supply safety detection and monitoring system established in my country, the high-speed railway contact network was photographed in all directions and without blind spots, and pictures of various states of the U-shaped clamp cotter pins of the high-speed railway contact network were obtained and used as the data set.

[0021] The data in the training set of this invention was screened from high-definition images collected by the high-speed railway power supply safety detection and monitoring system. Since not all images can be used as a network training set, they need to be screened. Specific requirements for this screening are that the dataset must include various states of high-speed railway contact network cotter pins, such as normal, damaged, and partially worn. A balanced number of images with different backgrounds, structures, distances, and angles must also be obtained to improve the model's generalization capabilities.

[0022] Step 2: After preprocessing the images in the dataset, a multi-level positioning strategy is adopted for the target.

[0023] This step can be described in detail as follows:

[0024] S21: Performing an enhancement operation on the image data, where the enhancement operation includes but is not limited to horizontal flipping, vertical flipping, and Gaussian blurring.

[0025] S22: Label preprocessing, that is, the labeling of the image is divided into two steps. First, the U-shaped clamp is positioned and labeled using tools such as LabelImg, and the labeled image is divided into a training set and a test set in a ratio of 9:1. The same method is used to perform the same data processing on the cotter pin of the U-shaped clamp image, that is, positioning and labeling and data set division.

[0026] Step 3: Use the labeled images in step 2 as a training set and put them into the recognition network model for training. After obtaining the training weights, test the data set.

[0027] This step can be described in detail as follows:

[0028] Step 31: Perform size preprocessing on the image, that is, adjust the image to an appropriate size for input into the network.

[0029] Step 32: Use CSPDarknet53 or the like as a backbone network to extract the features of the input image in step 31; the backbone network of the present invention includes three feature layers for fusion. For example, the CSPDarknet53 adopted in the embodiment of the present invention is composed of modules one to ten connected in sequence, wherein module one is a Focus module, modules two, four, six and eight are CBS modules, modules three, five and seven are C3 modules, module nine is an SPP module, and module ten is a CSP module; the modules five, seven and ten are three feature layers for fusion.

[0030] Step 33: Construct a multi-scale fusion module and input the output features of the backbone network into the improved ASPP module to capture features of different scales and different levels of context information to achieve multi-scale fusion.

[0031] like Figure 2 As shown, the improved ASPP module includes modules 27 to 31 connected in series, where modules 27, 28, 29, and 31 are dilated convolution modules with different dilation rates. Each module contains 1×1 convolution and 3×3 dilated convolution, and the dilation rates of the dilated convolution are 1, 2, 3, and 5, respectively. Module 30 is a 1×1 convolution layer.

[0032] The final output feature P of the backbone network is used as input and sent to module 27. Next, the input and output of module 27 are concat-operated and spliced ​​as the input of module 28 (i.e., the next dilated convolution block). And so on. Each dilated convolution is performed according to the above operation, that is, the input and output of module 28 are concat-operated and spliced ​​as the input of module 29, and the input and output of module 29 are concat-operated and spliced ​​as the input of module 31. Finally, the feature maps output by the four dilated convolution blocks are concat-operated and spliced, and the number of channels is adjusted to the same size as the output feature of the backbone network through module 30. Finally, the output of module 30 and the output of the backbone network are added to achieve feature fusion.

[0033] Step 34: The feature information learned by the backbone network and the improved ASPP module is passed to the improved PAN structure. The high-level feature map is first upsampled and then element-wise added with the low-level feature map for further feature extraction and fusion.

[0034] like Figure 2 As shown in FIG, the improved PAN module includes modules 11 to 23 connected in series, wherein modules 11, 12 and 13 are all self-attention modules, modules 14, 17, 21 and 23 are all C3+concat modules, modules 15 and 18 are all upsampling modules, and modules 16, 19, 20 and 22 are all CBS modules.

[0035] In the order of fusion, the three feature layers (module five, module seven, module ten) fused in the Darknet53 network are respectively followed by the CBAM attention mechanism (i.e., module eleven, module twelve, and module thirteen) to adaptively adjust the feature weights of each feature layer in order to achieve the purpose of focusing on important features and improving feature expression capabilities. Then, the weighted feature layers (i.e., module eleven, module twelve, and module thirteen) are respectively ADDed with the intermediate feature layers (i.e., module fourteen, module sixteen, and module nineteen), which prevents the loss of feature information in the backbone network and makes up for the problem of missing feature information caused by the upsampling operation.

[0036] Step 35: Input the result obtained in step 34 into the classification prediction network.

[0037] In an embodiment of the present invention, the classification prediction network is composed of module twenty-four, module twenty-five and module twenty-six. Each module adopts a decoupling head structure, including three parts: Reg part, Obj part, and Cls part; the Reg part is used to represent the offset of the center point of the prediction box and the width and height information of the prediction box, the Obj part is used to represent the probability of containing an object in the prediction box, and the C1s part is used to represent the probability that the prediction box corresponds to a certain type of defect.

[0038] After the improved PAN module, modules 24, 25, and 26 each obtain three prediction results: Reg(H, W, 4), Obj(H, W, 1), and Cls(H, W, C). These three prediction results are stacked to form the prediction result (H, W, 4 + 1 + C). H and W are the length and width of the feature map, respectively; 4 is the length, width, and center coordinates of the predicted box; 1 is the confidence level of the predicted box; and C is the classification type. The feature information output by module 24 is 20*20*85, the feature information output by module 25 is 40*40*85, and the feature information output by module 24 is 80*80*85.

[0039] Step 36: Decode the predicted box, classification probability, and target confidence, and convert the output result into the final result of target detection.

[0040] The feature map is input to the decoupling head to generate a prediction result, which includes the coordinate information, category probability, and confidence score of the prediction box. After the prediction box is generated, the positive sample information is initially screened out through the center point and coordinate information. Then, the optimized label assignment strategy SimOTA is used to refine the selection of candidate boxes. The specific process of SimOTA is as follows:

[0041] (1) Calculate the IoU of all predicted boxes and the true target box, and select the 10 anchor boxes with the largest IoU for each open pin;

[0042] (2) Select candidate frames based on the cost value. The smaller the cost, the lower the learning cost of the network. Calculate the Iou values ​​of the 10 candidate frames and round them up to get the k value. Then, select k frames with the smallest cost values ​​from the 10 candidate frames as the candidate frames corresponding to the target frame.

[0043] (3) If there is a conflict in candidate allocation, the candidate box is assigned to the real box with the lowest cost.

[0044] The cost calculation formula is as follows:

[0045] Cost = L cls +λLreg

[0046] In the formula: L cls represents the target classification loss, L reg represents the target positioning loss, and λ represents the balancing coefficient of the positioning loss.

[0047] Step 4: input the step 3 processed positioning opening pin image data into a defect detection network (ResNet50 network) to realize fault detection of the opening pin;

[0048] The present application uses ResNet50 network for defect classification prediction. First, a pre-trained ResNet50 network model is initialized, which has been trained on a large-scale image dataset. The opening pin image data processed by step 3 is input into the ResNet50 network, and a probability vector is output. ResNet50 network automatically performs feature extraction and classification, and the last layer of the full connection layer outputs the probability prediction of the opening pin failure.

[0049] The step 3 processed opening pin image data is preprocessed, the image is adjusted to 224*224 size, and the preprocessed image is input into the ResNet50 network, as shown in Figure 3 The feature representation of the image data is obtained by forward propagation. ResNet50 network has multiple convolutional layers and pooling layers to automatically extract the features of the image. When the features are sent to the full connection layer, a series of linear and nonlinear transformations are performed, and finally a probability vector is output. The output of the full connection layer is normalized using the Softmax function to convert it into a probability distribution. The probability prediction result of each class (faulty, non-faulty) can be obtained, and then the probability is used to judge whether the high-speed railway catenary U-shaped clamp opening pin has a fault.

Claims

1. A method for detecting defects in a U-shaped clamp cotter pin of a high-speed railway contact network, comprising the following steps: Step 1: Construct a dataset containing images of various states of the cotter pins of the high-speed railway contact network U-shaped clamps. Step 2: After pre-processing the image, a multi-level positioning strategy is used for the target; Step 3: Use the labeled images in step 2 as the training set, put them into the recognition network model for training, and then test the data set after obtaining the training weights. The process is as follows: Step 31: Pre-process the image size; Step 32: extracting features of the input image in step 31 using a backbone network; the backbone network includes three feature fusion layers; Step 33: Input the output features of the backbone network into the improved ASPP module to capture features of different scales and different levels of context information and achieve multi-scale fusion; The improved ASPP module includes modules 27 to 31, wherein modules 27, 28, 29, and 31 are dilated convolution modules with different expansion rates, each module containing 1×1 convolution and 3×3 dilated convolution, and the dilation rates of dilated convolution are 1, 2, 3, and 5, respectively. Module 30 is a 1×1 convolution layer. The output features of the backbone network are input to module 27, and the input and output of module 27 are concat-operated and spliced ​​as the input of module 28, and the input and output of module 28 are concat-operated and spliced ​​as the input of module 29. The input and output of module 29 are concat-operated and spliced ​​as the input of module 31. The feature maps output by the four dilated convolution modules are concat-operated and spliced, and the number of channels is adjusted to the same size as the output features of the backbone network through module 30. The output of module 30 and the output of the backbone network are added to realize feature fusion. Step 34: The feature information learned by the backbone network and the improved ASPP module is passed to the improved PAN structure. The high-level feature map is first upsampled and then element-wise added with the low-level feature map for further feature extraction and fusion. The improved PAN structure includes modules 11 to 23, wherein modules 11, 12 and 13 are self-attention modules, modules 14, 17, 21 and 23 are C3+concat modules, modules 15 and 18 are upsampling modules, and modules 16, 19, 20 and 22 are CBS modules. According to the fusion order, the three fused feature layers are connected to module 11, module 12 and module 13 respectively, and the CBAM attention mechanism is added; Perform ADD operations on module 11, module 12, and module 13 with the intermediate feature layer, where the intermediate feature layer is module 14, module 16, and module 19; Step 35: Input the result obtained in step 34 into the classification prediction network; Step 36: Decode the predicted box, classification probability, and target confidence, and convert the output result into the final result of target detection; Step 4: Input the cotter pin image data located in step 3 into the defect detection network to realize fault detection.

2. The method for detecting defects in the cotter pins of the high-speed railway contact network U-shaped clamps according to claim 1 is characterized in that: In step 1, the data in the training set are screened from high-definition images collected by the high-speed railway power supply safety detection and monitoring system. The states of the cotter pins of the high-speed railway contact network U-shaped clamps include normal state, damaged state and partially worn state; the data in the training set include images with different backgrounds, structures, distances and angles.

3. The method for detecting defects in U-shaped clamp cotter pins of high-speed railway contact network according to claim 1, characterized in that: In step 2, the image is enhanced, wherein the enhancement operations include horizontal flipping, vertical flipping, and Gaussian blurring. The labeling preprocessing is to label the image twice. First, the U-shaped clamp is positioned and labeled using the LabelImg tool, and the labeled image is divided into a training set and a test set at a ratio of 9:

1. The same method is used to perform the same positioning and labeling and data set division on the cotter pin on the U-shaped clamp image.

4. The method for detecting defects in the cotter pins of a high-speed railway contact network U-shaped clamp according to claim 1, characterized in that: The backbone network is CSPDarknet53, which is composed of modules 1 to 10 connected in sequence, where module 1 is the Focus module, modules 2, 4, 6 and 8 are CBS modules, modules 3, 5 and 7 are C3 modules, module 9 is the SPP module and module 10 is the CSP module; modules 5, 7 and 10 are the three feature layers for fusion.

5. The method for detecting defects in the cotter pins of a high-speed railway contact network U-shaped clamp according to claim 1, characterized in that: The classification prediction network consists of modules 24, 25, and 26. Each module adopts a decoupled head structure, including three parts: Reg, Obj, and Cls. The Reg part is used to represent the offset of the center point of the prediction box and the width and height information of the prediction box. The Obj part is used to represent the probability that the prediction box contains an object. The Cls part is used to represent the probability that the prediction box corresponds to a certain type of defect. After the improved PAN module, module 24, module 25 and module 26 each obtain three prediction results Reg(H,W,4), Obj(H,W,1), Cls(H,W,C); the three prediction results are stacked to use the prediction result (H,W,4+1+C), where H and W are the length and width of the feature map respectively, 4 is the length and width of the prediction box and the coordinates of the center point, 1 is the confidence of the prediction box, and C is the classification type.

6. The method for detecting defects in the cotter pins of a high-speed railway contact network U-shaped clamp according to claim 1, characterized in that: In step 36, the decoupling head is used to generate a prediction result, which includes the coordinate information, category probability, and confidence score of the prediction box. After the prediction box is generated, the positive sample information is preliminarily screened out through the center point and coordinate information, and then the candidate box is refined through the optimized label assignment strategy SimOTA.

7. The method for detecting defects in the cotter pins of a high-speed railway contact network U-shaped clamp according to claim 1, characterized in that: In step 4, the ResNet50 network is used to perform defect classification prediction. The method is as follows: first, a pre-trained ResNet50 network model is initialized. The model has been trained on a large-scale image dataset. The cotter pin image data processed and located in step 3 is input into the ResNet50 network, and a probability vector is output. The ResNet50 network automatically extracts and classifies features, and the output of the last fully connected layer predicts the probability of cotter pin failure.

Citation Information

Patent Citations

  • Intelligent detection method for high-speed rail overhead line system cotter pin defects

    CN114202540A

  • Small target identification method and system for complex scene of power transmission line

    CN115294483A