Train bow net system fault detection method and system based on YOLOv11

By improving the YOLOv11 network structure, combining the wavelet convolution module, attention mechanism and adaptive threshold focus loss function, the accuracy and efficiency problems of fault detection of train bow network system in complex environments are solved, and efficient fault identification is achieved.

CN120495682APending Publication Date: 2025-08-15HUNAN FIRST NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510559100.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the fault detection of train bow network systems under complex environments, the detection efficiency is susceptible to natural environment interference, and the equipment operating conditions and complex background noise reduce defect recognition accuracy, making it difficult to meet real-time requirements.

Method used

The fault detection method based on YOLOv11 is adopted, and feature extraction is processed through the wavelet convolution module, the feature weight is adjusted using the attention mechanism module, and an adaptive threshold focus loss function is introduced to optimize the network structure to improve detection accuracy and robustness.

Benefits of technology

It improves the accuracy of fault detection in complex environments, enhances the robustness of the model to extreme lighting and dynamic range changes, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495682A_ABST
    Figure CN120495682A_ABST
Patent Text Reader

Abstract

The invention discloses a train pantograph-catenary system fault detection method and system based on YOLOv11, and relates to the image recognition technology, and the method comprises the steps: employing a wavelet convolution module to process feature extraction of different stages of a backbone network, carrying out the multi-scale decomposition of an input signal through wavelet transformation, and extracting high-frequency details and low-frequency contour features; adjusting weights of different brightness or color interval features by using an attention mechanism module, and simultaneously paying attention to local and global information by using parallel double-path attention; a loss function module is used for determining a proper threshold value for the sample, and the loss weight is automatically adjusted according to the characteristics of the sample and model output; the wavelet convolution module better retains image details, achieves the effect of picture deblurring, and effectively solves the problem of loss of feature problem information; the attention mechanism module reduces the influence of feature degradation caused by complex weather and background, and the loss function module enables the model to pay more attention to difficult samples and solves the problem of low detection precision of some samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a train pantograph-catenary system fault detection method and system based on YOLOv11. Background Art

[0002] Railway transportation has the advantages of large capacity and high speed. It has become the most convenient means of transportation for the masses, and the electrification system is one of its most important components. The pantograph-catenary system is developed to provide continuous power for high-speed trains, in which the pantograph provides power support for the train through the contact network, and the dropper is connected between the contact wire and the load-bearing cable. Since the dropper needs to withstand the impact of the pantograph on the contact network during train operation, and is also subject to damage from the natural environment, the dropper is prone to slack, which affects the normal operation of the train. In addition, when the train is running at high speed, factors such as body vibration, contact network height changes, and insufficient contact pressure will cause an air gap between the carbon slide and the contact wire. When the pantograph is separated from the contact network, arcing will occur. Frequent pantograph arcing will shorten the service life of the contact network and increase the risk of line failures and accidents. Dropper string relaxation and arcing phenomenon (reference Figure 2 ) brings many hazards to the stable and safe operation of trains, so it is very important to conduct regular monitoring of the pantograph system to ensure the stable operation of high-speed railways.

[0003] In recent years, the high-speed railway catenary automatic detection and monitoring system (abbreviated as 6C system, the system diagram is referenced Figure 8 ), real-time high-resolution image acquisition of the pantograph and catenary suspension system through a dynamic image acquisition module, and transmission of the image back to the main control for data analysis has become an important means of monitoring the pantograph-catenary system at home and abroad. This detection method significantly improves the efficiency and safety of pantograph-catenary system detection while reducing labor costs. However, with the continuous expansion of the scale of train lines, the data collected by on-board equipment has increased exponentially, which has increased the workload of relevant technical personnel. With the vigorous development of machine vision-related technologies, the application of visual algorithms combined with convolutional neural network models to pantograph-catenary system fault detection and the realization of automatic identification and quantitative analysis of elastic deformation parameters of suspension strings and abnormal discharge points has become an important research direction and a very challenging topic.

[0004] While these methods can achieve effective detection in controlled environments, deploying pantograph-catenary inspection systems in real-world railway scenarios presents multiple challenges. Detection performance is easily affected by natural environmental factors (such as heavy fog, rainfall, and dramatic changes in daytime and nighttime illumination). Furthermore, equipment operating conditions (including lens contamination, low-light conditions in tunnels, image blur caused by high-speed motion and liquid adhesion) and complex background noise can significantly reduce defect recognition accuracy. Therefore, the development of new inspection solutions is urgently needed to overcome the application bottlenecks of existing technologies in complex real-world conditions.

[0005] The evolution of existing algorithms continues to drive the development of object detection technology. Among mainstream methods based on convolutional neural networks (CNNs), two-stage detection algorithms (such as the R-CNN series) employ a two-stage architecture of "region proposal generation → refined classification and regression," achieving high-precision detection through a region proposal network (RPN). However, this cascaded process suffers from duplicate feature extraction and computational redundancy, resulting in detection frame rates (FPS) that struggle to meet real-time requirements. In contrast, single-stage algorithms (such as the YOLO series, SSD, and RetinaNet) perform object localization and classification prediction directly on a grid of feature maps, improving processing speed but still lagging behind two-stage approaches in accuracy. This performance difference between these two approaches provides important optimization directions for the development of object detection technology. Summary of the Invention

[0006] The present invention aims to provide a train catenary system fault detection method and system based on YOLOv11, so as to solve the problem that the detection efficiency is easily disturbed by the natural environment, and the equipment operating conditions and complex background noise will also reduce the accuracy of defect identification.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A first aspect of the present invention provides a train pantograph-catenary system fault detection method based on YOLOv11, comprising the following steps:

[0009] S1. Collect sample data sets and perform data enhancement;

[0010] S2. Build a YOLOv11 network framework, where the YOLOv11 network framework includes a backbone network, a neck network, and a head network.

[0011] S3, using the wavelet convolution module (C3k2WT) to process the feature extraction at different stages of the backbone network, performing multi-scale decomposition of the input signal through wavelet transform to extract high-frequency details and low-frequency contour features;

[0012] S4, using the attention mechanism module (DHPA) to adjust the weights of features in different brightness or color intervals, and using parallel dual-path attention to focus on local and global information at the same time;

[0013] S5. Use the loss function module (ATFL) to determine the appropriate threshold for each sample and automatically adjust the loss weight based on the characteristics of the sample and the model output;

[0014] S6. Use the data-enhanced dataset to train the constructed YOLOv11 framework to obtain a detection model, which is applied to train pantograph system fault detection.

[0015] Furthermore, in step S1, the data set is divided into a test set, a validation set, and a training set in a ratio of 1:1:8, and the data set is amplified using data amplification methods such as flipping, scaling, cropping, brightness adjustment, contrast adjustment, saturation adjustment, and adding noise.

[0016] Furthermore, the step S3 includes the following steps:

[0017] S301: The input undergoes a first-layer ordinary convolution transformation, and a cross-layer feature fusion network is designed. The low-frequency features extracted by the current layer are fused with the reconstruction results of the next layer after channel alignment. This effectively alleviates the feature loss problem of deep networks through multi-level information superposition:

[0018]

[0019] Among them, P represents the input feature map, P LL ,P LH ,P HL ,P HH Represent the output of the four filters respectively, WT represents wavelet transform, and i represents the i-th layer;

[0020] Then enter the wavelet transform, and its output resolution satisfies the relationship: Where W and H are the input width and height:

[0021]

[0022] Among them, F LL Represents a low-pass filter, F LH Represents the horizontal high-pass filter, F HL Represents a vertical high-pass filter, F HH represents a diagonal high-pass filter;

[0023] S302: Perform independent convolution and scaling on the high-frequency components through the convolution layer (wavelet_convs) and the scale scaling layer (wavelet_scale) to suppress the noise in the high-frequency components.

[0024] [P LL ,P LH ,P HL ,P HH ]=Conv([F LL ,F LH ,F HL ,F HH ],P)

[0025] Among them, P LL ,P LH , P HL , P HHRepresents the output of the four filters, F LL , F LH ,F HL ,F HH Represent four filters respectively;

[0026] S303, obtain the inverse wavelet transform (IWT) through transposed convolution. First, the low-frequency and high-frequency components after wavelet transform filtering need to be passed through a small convolution kernel:

[0027] Q = IWT(DW_Conv(W, WT(P)))

[0028] Where DW_Conv represents depthwise convolution, W is the weight tensor of the 3×3 depthwise kernel convolution, IWT represents inverse wavelet transform, WT represents wavelet transform, and P represents input features.

[0029] The high-frequency subband with enhanced features is combined with the low-frequency component in the complex domain to achieve the deblurring effect:

[0030] P=Conv T ([F LL ,F LH ,F HL ,F HH ],[P LL ,P LH ,P HL ,P HH ])

[0031] Among them, Conv T represents transposed convolution, P LL ,P LH , P HL ,P HH Represents the output of the four filters, F LL ,F LH ,F HL , F HH Represent four filters respectively;

[0032] S304. Using linear operations, the low frequencies obtained by the WT operation and its inverse operation are superimposed to combine different frequency outputs:

[0033]

[0034] Assume Y (i) It is the summary output starting from the i-th layer. The outputs of two convolutions of different sizes are summed as the output, thus obtaining the sum of convolutions at different levels.

[0035] Furthermore, step S4 includes the following steps:

[0036] S401, Dynamic Range Convolution: By performing dynamic range structured reorganization on the diagonal area of the input feature matrix, high and low intensity pixels are distributed in a regularized manner, thereby supporting the convolution kernel to perform cross-channel calculations. The reorganized features are then passed through depth-wise separable convolution:

[0037] X1,X2=Split(X),X1=Sort v (Sort h (X1))

[0038]

[0039] in, It is a 3×3 depth convolution, Conv 1×1 It is a 1×1 point-wise convolution, Concat is a concatenation operation along the channel, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension;

[0040] S402: The features after depthwise separable convolution enter the parallel dual-loop attention mechanism, which divides the output of the dynamic range convolution into value features and two pairs of query key values, and then passes them to the two loops. First, the value features V are sorted, and the query key values are sorted according to the index:

[0041] V,d=Sort(V)

[0042] Q1,K1=Split(Gather((X QK,1 ),d))

[0043] Q2, K2 = Split (Gather ((X QK,2 ), d))

[0044] Where Q, K, and V represent the query, key, and value matrices respectively, d is the index of the sort value, Gather represents a retrieval operation to search for elements in the tensor given an index value, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension.

[0045] The sorted V and two pairs of key values (Q1, K1 and Q2, K2) enter two types of reconstruction and attention mechanisms respectively, and we get:

[0046]

[0047] A=A B ⊙A F

[0048] Among them, k is the number of attention heads, Q, K, V represent the query, key, and value matrices respectively, and R B Represents the reshaping operation of BHR, RF represents the reshaping operation of FHR, A represents the obtained attention map, softmax represents the flexible maximum function, and ⊙ represents the dot product operation;

[0049] S403. The reordered features are sorted back to their original positions before the final output of the 1×1 point-wise convolution, and spatial dynamic weather degradation features are extracted through parallel bidirectional attention.

[0050] Furthermore, step S5 includes the following steps:

[0051] S501. Add the ATFL function to the loss function to assign weights to various sample types, so that the model pays more attention to small samples:

[0052]

[0053] Among them, λ is the parameter value with 0.5 as the limit, η is the adaptive modulation factor, γ is the modulation factor, P t is the average predicted probability value of this time;

[0054] S502. Adaptively improve TFL and mathematically model the model training speed:

[0055]

[0056] Among them, P t is the average predicted probability value, P i is the average predicted probability value of each training epoch, P c is the predicted value of the next epoch;

[0057] S503, finally get the adaptive threshold focus loss function expression:

[0058]

[0059] Among them, P t is the average predicted probability value this time, and λ is the parameter.

[0060] A second aspect of the present invention provides a train bow-net system fault detection system based on YOLOv11, which adopts a YOLOv11 network framework, including a backbone network (backbone), a neck network (neck) and a head network (head); the C3K2 module in the YOLOv11 network framework is adjusted to a wavelet convolution module (C3k2WT), the C2PSA module in the YOLOv11 network framework is adjusted to an attention mechanism module (DHPA), and a loss function module (ATFL) is introduced;

[0061] The backbone network is responsible for feature extraction;

[0062] The neck network is located between the backbone network and the head network and is responsible for feature fusion and enhancement;

[0063] The head network is the decision-making part of the target detection model and is responsible for generating the final detection results.

[0064] The wavelet convolution module (C3k2WT) is responsible for processing the feature extraction at different stages of the backbone;

[0065] The attention mechanism module (DHPA) enhances feature extraction capabilities through a multi-head attention mechanism and a feedforward neural network;

[0066] The loss function module (ATFL) automatically adjusts the loss weight according to the characteristics of the sample and the model output.

[0067] The principles and beneficial effects of this technical solution include at least:

[0068] 1. This invention primarily improves the backbone network, which employs a series of convolutional and deconvolutional layers, along with residual connections and a bottleneck architecture to reduce network size and improve performance. A wavelet convolution module (C3k2WT) handles feature extraction at different stages of the backbone. Wavelet transforms perform a multi-scale decomposition of the input signal, extracting high-frequency details (such as edges and textures) and low-frequency contour features. Noise in high-frequency components is specifically suppressed, while multi-stage reconstruction using the inverse transform restores lost high-frequency information, achieving a deblurring effect. Furthermore, a smaller 3×3 kernel allows for more efficient computation while preserving the model's ability to capture essential features in the image. The core of the backbone is the wavelet convolution module (C3k2WT), which optimizes information flow within the network by segmenting the feature map and applying a series of smaller kernel convolutions. This is faster and less computationally expensive than larger kernel convolutions. By processing smaller, independent feature maps and merging them after several convolutions, this improves feature representation using fewer parameters compared to the C2f module in YOLOv8.

[0069] 2. In addition to C3K2WT, the attention structure in YOLOv11 has been modified to adopt an attention mechanism module (DHPA). Based on the dynamic range histogram statistical feature distribution, the attention mechanism dynamically adjusts the weights of features in different brightness or color ranges. The use of parallel dual-path attention simultaneously focuses on local and global information, addressing the traditional attention mechanism's sensitivity to extreme values and enhancing the model's robustness to complex lighting or dynamic range changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1This is a schematic diagram of the overall structure of the YOLOv11 network framework;

[0071] Figure 2 This is a schematic diagram of a loose dropper string and arcing fault. The red dotted box represents a loose dropper string, and the green solid box represents a normal dropper string.

[0072] Figure 3 Schematic diagram of the C3K2WT structure;

[0073] Figure 4 Schematic diagram of receptive field mapping;

[0074] Figure 5 Schematic diagram of the attention mechanism module (DHPA);

[0075] Figure 6 Schematic diagram of the comparison of mAP50(%) curves between the baseline model and the improved model;

[0076] Figure 7 Schematic diagram of some test results;

[0077] Figure 8 This is a schematic diagram of the equipment for the automatic detection and monitoring system for high-speed railway catenary lines. DETAILED DESCRIPTION

[0078] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0079] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0080] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0081] It should be emphasized here that the step marks mentioned below do not limit the order of the steps, but it should be understood that the steps can be executed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be executed simultaneously.

[0082] Example 1

[0083] refer to Figure 1 , a train pantograph-catenary system fault detection method based on YOLOv11, comprising the following steps:

[0084] S1. Collect sample data sets and perform data enhancement.

[0085] A total of 1182 original sample data were collected using the high-speed railway bow-chain automatic detection and monitoring system (referred to as the 6C system) (arcing, normal string hanging, and loose string hanging three states, reference Figure 2 ), before training, the images in the dataset were annotated using the LableImg software, with arcing, normal hanging string, and relaxed hanging string labeled as 0, 1, and 2, respectively.

[0086] The dataset was divided into a test set, a validation set, and a training set in a ratio of 1:1:8. Due to insufficient data samples, overfitting is a common problem. To minimize the impact on the model's accuracy in detecting contact network targets, data augmentation was applied to the training set. Key methods include flipping, scaling, cropping, adjusting brightness, contrast, and saturation, and adding noise. These methods train the neural network to ignore background variations unrelated to the target, accurately extracting target features, increasing data diversity, improving the model's generalization ability, and alleviating sample imbalance. Using these data augmentation methods, the training set was expanded to 6,363 contact network images.

[0087] S2. Build a YOLOv11 network framework, which includes a backbone network (backbone), a neck network (neck), and a head network (head).

[0088] Among them, the backbone network is responsible for feature extraction; the neck network is located between the backbone network and the head network, and its function is to perform feature fusion and enhancement; the head network is the decision-making part of the target detection model, responsible for generating the final detection results.

[0089] S3, using wavelet convolution module (C3k2WT) to process feature extraction at different stages of the backbone network (reference Figure 3 ), the input signal is decomposed into multiple scales through wavelet transform to extract high-frequency details and low-frequency contour features.

[0090] This paper primarily improves the backbone network, which employs a series of convolutional and deconvolutional layers, along with residual connections and a bottleneck structure to reduce network size and improve performance. A wavelet convolution module (C3k2WT) handles feature extraction at different stages of the backbone. Wavelet transforms perform a multi-scale decomposition of the input signal, extracting high-frequency details (such as edges and textures) and low-frequency contour features. Noise in high-frequency components is specifically suppressed, while multi-stage reconstruction using the inverse transform restores lost high-frequency information, achieving a deblurring effect. A smaller 3x3 kernel allows for more efficient computation while preserving the model's ability to capture essential features in the image. The core of the backbone network is the wavelet convolution module (C3k2WT), which optimizes information flow within the network by segmenting the feature map and applying a series of smaller kernel convolutions (3x3). This is faster and less computationally expensive than larger kernel convolutions. By processing smaller, independent feature maps and merging them after several convolutions, the proposed method improves feature representation using fewer parameters compared to the C2f module in YOLOv8.

[0091] Specifically:

[0092] S301: The input undergoes a first-layer ordinary convolution transformation, and a cross-layer feature fusion network is designed. The low-frequency features extracted by the current layer are fused with the reconstruction results of the next layer after channel alignment. This effectively alleviates the feature loss problem of deep networks through multi-level information superposition:

[0093]

[0094] Among them, P represents the input feature map, P LL ,P LH ,P HL ,P HH Represent the output of the four filters respectively, WT represents wavelet transform, and i represents the i-th layer.

[0095] Then enter the wavelet transform to expand the receptive field of the model (refer to Figure 4 ), its output resolution strictly satisfies the relationship: (W, H are input width and height):

[0096]

[0097]

[0098] Among them, F LL Represents a low-pass filter, F LH Represents the horizontal high-pass filter, F HL Represents a vertical high-pass filter, F HH represents a diagonal high-pass filter.

[0099] S302: Subsequently, the high-frequency components are independently convolved and scaled through the convolution layer (wavelet_convs) and the scale scaling layer (wavelet_scale) to suppress the noise in the high-frequency components:

[0100] [P LL ,P LH ,P HL ,P HH ]=Conv([F LL ,F LH ,F HL ,F HH ],P)

[0101] Among them, P LL ,P LH , P HL , P HH Represents the output of the four filters, F LL , F LH ,F HL ,F HH Represent four filters respectively.

[0102] S303. Finally, the inverse wavelet transform (IWT) is obtained by transposed convolution. However, before this, the low-frequency and high-frequency components after the wavelet transform filtering need to be passed through a small convolution kernel:

[0103] Q = IWT(DW_Conv(W, WT(P)))

[0104] Among them, DW_Conv represents depth convolution, W is the weight tensor of 3×3 depth kernel convolution, IWT represents inverse wavelet transform, WT represents wavelet transform, and P represents input features.

[0105] Alleviate the surge in computational complexity, realize multi-scale feature reconstruction, and perform tensor synthesis of the feature-enhanced high-frequency subband and the low-frequency component in the complex domain to achieve a deblurring effect.

[0106] P=Conv T ([F LL ,F LH ,F HL ,F HH ],[P LL ,P LH ,P HL ,P HH ])

[0107] Among them, Conv T represents transposed convolution, P LL ,P LH , P HL ,P HHRepresents the output of the four filters, F LL ,F LH ,F HL , F HH Represent four filters respectively;

[0108] S304. Using linear operations, the low frequencies obtained by the WT operation and its inverse operation are superimposed to combine different frequency outputs:

[0109]

[0110] Assume Y (i) It is the summary output starting from the i-th layer. The outputs of two convolutions of different sizes are summed as the output, thus obtaining the sum of convolutions at different levels.

[0111] S4. Use the attention mechanism module (DHPA) to adjust the weights of features in different brightness or color intervals, and use parallel dual-path attention to focus on local and global information at the same time.

[0112] In addition to C3K2WT, the attention structure in YOLOv11 has been modified to adopt an attention mechanism module (DHPA). Based on the dynamic range histogram statistical feature distribution, the attention mechanism dynamically adjusts the weights of features in different brightness or color ranges. Parallel dual-path attention is used to simultaneously focus on local and global information, addressing the traditional attention mechanism's sensitivity to extreme values and enhancing the model's robustness to complex lighting or dynamic range changes.

[0113] Specifically:

[0114] S401, attention mechanism module (DHPA) consists of two parts (structure reference Figure 5 ), firstly, dynamic range convolution, which performs a structured dynamic range reorganization on the diagonal regions of the input feature matrix, regularizing the distribution of high and low intensity pixels, thereby enabling the convolution kernel to perform cross-channel calculations. The reorganized features are then passed through depthwise separable convolution.

[0115] X1,X2=Split(X),X1=Sort v (Sort h (X1))

[0116]

[0117] in, It is a 3×3 depth convolution, Conv 1×1 It is a 1×1 point-wise convolution, Concat is a concatenation operation along the channel, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension.

[0118] S402. The features after depthwise separable convolution enter the parallel dual-loop attention mechanism, which divides the output of the dynamic range convolution into value features and two pairs of query key values, and then passes them to the two loops. First, the value features V are sorted, and then the query key values are sorted according to the index.

[0119] V,d=Sort(V)

[0120] Q1,K1=Split(Gather((X QK,1 ),d))

[0121] Q2, K2 = Split (Gather ((X QK,2 ), d))

[0122] Among them, Q, K, and V represent the query, key, and value matrices respectively, d is the index of the sort value, Gather represents a retrieval operation to search for elements in the tensor given an index value, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension.

[0123] The second is a parallel dual-path attention mechanism: one path focuses on local structure, while the other captures global relationships, aggregating global and local dynamic features. Values are also shared, and both paths use the same sorted values to reduce computational complexity. The sorted V and the two key-value pairs (Q1, K1 and Q2, K2) are fed into two different types of reconstruction and attention mechanisms, yielding:

[0124]

[0125] A=A B ⊙A F

[0126] Among them, k is the number of attention heads, Q, K, V represent the query, key, and value matrices respectively, and R B Represents the reshaping operation of BHR, R F represents the reshaping operation of FHR, A represents the obtained attention map, softmax represents the flexible maximum function, and ⊙ represents the dot product operation;

[0127] S403. Finally, to maintain spatial consistency, the reordered features are sorted back to their original positions before the final output of the 1×1 point-wise convolution. Spatial dynamic weather degradation features are extracted through parallel bidirectional attention.

[0128] S5. Due to the class imbalance in the dataset, a loss function module (ATFL) is used to determine a suitable threshold for each sample and automatically adjust the loss weight according to the characteristics of the sample and the model output.

[0129] Adopting the Adaptive Threshold Focal Loss function, ATFL determines a suitable threshold for each sample and automatically adjusts the loss weight according to the characteristics of the sample and the model output, introducing a focal mechanism.

[0130] In the case of two categories, the predicted probabilities of different categories are p and 1-p. On this basis, the classic cross entropy loss function can be expressed as:

[0131]

[0132] Among them, y i Indicates the label of the sample, p i It represents the probability that sample i is predicted to be positive.

[0133] The ATFL function is added to the loss function to assign weights to various sample types, solve the problem of class label imbalance, and make the model pay more attention to small samples:

[0134]

[0135] Among them, λ is the parameter value with 0.5 as the limit, η is the adaptive modulation factor, γ is the modulation factor (the γ parameter value in this model is 1.5), P t is the average predicted probability value for this time.

[0136] Adaptive improvements are made based on TFL, and mathematical modeling of the model training speed is performed.

[0137]

[0138] Among them, P t is the average predicted probability value, P i is the average predicted probability value of each training epoch, P c It is the predicted value of the next epoch.

[0139] Finally, the adaptive threshold focus loss function expression can be obtained:

[0140]

[0141] Among them, P t is the average predicted probability value this time, and λ is the parameter.

[0142] S6. Use the data-enhanced dataset to train the constructed YOLOv11 framework to obtain a detection model, which is applied to train pantograph system fault detection.

[0143] Example 2

[0144] A YOLOv11-based train pantograph-network system fault detection system adopts the YOLOv11 network framework, including a backbone network (backbone), a neck network (neck), and a head network (head); the C3K2 module in the YOLOv11 network framework is adjusted to a wavelet convolution module (C3k2WT); the C2PSA module in the YOLOv11 network framework is adjusted to an attention mechanism module (DHPA); and a loss function module (ATFL) is introduced;

[0145] The backbone network is responsible for feature extraction;

[0146] The neck network is located between the backbone network and the head network and is responsible for feature fusion and enhancement;

[0147] The head network is the decision-making part of the target detection model and is responsible for generating the final detection results.

[0148] The wavelet convolution module (C3k2WT) is responsible for processing the feature extraction at different stages of the backbone;

[0149] The attention mechanism module (DHPA) enhances feature extraction capabilities through a multi-head attention mechanism and a feedforward neural network;

[0150] The loss function module (ATFL) automatically adjusts the loss weight according to the characteristics of the sample and the model output.

[0151] Experimental Example 1

[0152] In order to objectively evaluate the advantages of the algorithm proposed in this paper, the evaluation indicators selected include mAP@50 (mean average precision of 0.5), GFLOPs (floating point operations / second), speed (ms) and number of parameters (millions). The formulas involved are as follows:

[0153]

[0154] in:

[0155] TP (True Positive): The number of samples whose true category is positive and correctly judged as positive by the model;

[0156] FP (False Positive): The number of samples whose true category is negative but mistakenly judged as positive by the model;

[0157] FN (False Negative): The number of samples whose true category is positive but mistakenly classified as negative by the model;

[0158] N: the total number of target categories in the multi-classification task;

[0159] AP (Average Precision): The integrated area under the Precision-Recall curve for a single category, which quantifies the detection stability of the model for that category;

[0160] mAP (mean Average Precision): The mean of the AP values of all N categories, which comprehensively evaluates the global performance of the model.

[0161] To further validate the superiority of this algorithm, we compared it with other object detection algorithms, including earlier versions of the YOLO family, such as YOLOv8 and YOLOv5s, and YOLOv10n. These algorithms were trained and tested on the dataset and evaluated based on four metrics. The experimental results are shown in Table 1. These comparative tests show that the proposed algorithm achieved a detection accuracy of 97.1%, 6.3 GFLOPS, and 2.57M parameters on the dataset. Compared to YOLOv11n, this algorithm achieved a 1.3% improvement in detection accuracy, albeit with a slightly slower computational speed. Compared to the lighter algorithms, YOLOv5s and YOLOv8n, the improved algorithm achieved higher detection accuracy.

[0162] Table 1: Comparative experimental results

[0163]

[0164] To verify the effectiveness of each improved module, we conducted ablation experiments on the dataset, as shown in Table 2. Using YOLOv11 as the baseline model (mAP@0.5: 95.8%, Params: 2.6M), we gradually introduced improved modules: After adding the wavelet convolution module (C3k2WT), mAP@0.5 increased by 0.5% to 96.3%, and the number of parameters decreased to 2.56M, demonstrating its effectiveness in enhancing the features of blurred samples and improving computational efficiency. Further integration of the DHPA dynamic attention mechanism achieved mAP@0.5 of 96.9% (+0.6%), with the number of parameters reduced to 2.52M. Furthermore, the recall rate of string slack detection was significantly improved from 95.4% to 96.3%, verifying the module's robustness optimization in low-light scenes. Finally, when the ATFL loss function was introduced, mAP@0.5 increased by 0.2%, demonstrating its effectiveness in improving the class imbalance problem. When the full model (C3K2WT+DHPA+ATFL) worked together, mAP@0.5 reached 97.1% (ref. Figure 6 Compared with the baseline model (+1.3%), the number of parameters has decreased by 1.15% (2.57M), and the model is more efficient. Experiments show that the proposed method achieves the best balance between accuracy and efficiency, especially for fault detection under extreme lighting and complex backgrounds. Some results are shown in the figure below. Figure 7shown.

[0165] Table 2: Ablation experiment results

[0166]

[0167] The experimental setup consisted of a distributed cluster of 10 Intel(R) Xeon Platinum 8352V CPUs, an NVIDIA GeForce RTX 4090 GPU with CUDA 12.1, and a Python 3.12 programming environment. During training of the improved model, images were resized to 640×640, the batch size was set to 32, the number of training epochs was set to 120, the SGD optimizer was used, and the initial learning rate was set to 0.01. All other parameters were kept to their default values.

[0168] To address various issues encountered in pantograph-catenary system detection under complex working conditions, an improved algorithm based on YOLOv11 is proposed. This model enhances the backbone and loss function based on YOLOv10n. The main contributions of this invention are:

[0169] a) The C3K2 module in the network is replaced with a wavelet convolution module (C3k2WT). The Bottleneck module in the C3K2 module is replaced with a wavelet convolution module WT-Bneck, designed based on the wavelet transform and its inverse transform. This modification better preserves image details and achieves image deblurring, effectively addressing the loss of feature information. Experimental results show that detection accuracy improves by 0.5% on a custom dataset.

[0170] b) Drawing on the design principles of the multi-head attention mechanism, we propose a dynamic convolutional collaborative attention module, or DHPA, and integrate it into the C2PSA backbone network. By dynamically calculating the correlation weights between elements within a sequence, a parallel dual-loop attention mechanism is used to capture both global and local information, mitigating the impact of feature degradation caused by complex weather and background conditions. Experimental results show a 0.6% improvement in detection accuracy on a custom dataset.

[0171] c) The ATFL function is added to the loss function to assign weights to various sample types, allowing the model to focus more on difficult samples and address the problem of low detection accuracy for some samples. Experimental results show that the detection accuracy on the custom dataset is improved by 0.2%.

[0172] d) Finally, ablation experiments on a custom dataset show that the detection accuracy is improved by 1.3% compared to the original YOLOv11 model.

[0173] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0174] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0175] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0176] The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification may be used to interpret the content of the claims.

Claims

1. A train pantograph-catenary system fault detection method based on YOLOv11, characterized in that: The steps include: S1. Collect sample data sets and perform data enhancement; S2. Build a YOLOv11 network framework, where the YOLOv11 network framework includes a backbone network, a neck network, and a head network. S3, using the wavelet convolution module (C3k2WT) to process the feature extraction at different stages of the backbone network, performing multi-scale decomposition of the input signal through wavelet transform to extract high-frequency details and low-frequency contour features; S4, using the attention mechanism module (DHPA) to adjust the weights of features in different brightness or color intervals, and using parallel dual-path attention to focus on local and global information at the same time; S5. Use the loss function module (ATFL) to determine the appropriate threshold for each sample and automatically adjust the loss weight based on the characteristics of the sample and the model output; S6. Use the data-enhanced dataset to train the constructed YOLOv11 framework to obtain a detection model, which is applied to train pantograph system fault detection.

2. A train pantograph-catenary system fault detection method based on YOLOv11 according to claim 1, characterized in that: In step S1, the data set is divided into a test set, a validation set, and a training set in a ratio of 1:1:8, and the data set is amplified using data amplification methods such as flipping, scaling, cropping, brightness adjustment, contrast adjustment, saturation adjustment, and noise addition.

3. A train pantograph-catenary system fault detection method based on YOLOv11 according to claim 1, characterized in that: The step S3 includes the following steps: S301: The input undergoes a first-layer ordinary convolution transformation, and a cross-layer feature fusion network is designed. The low-frequency features extracted by the current layer are fused with the reconstruction results of the next layer after channel alignment. This effectively alleviates the feature loss problem of deep networks through multi-level information superposition: Among them, P represents the input feature map, P LL ,P LH ,P HL ,P HH Represent the output of the four filters respectively, WT represents wavelet transform, and i represents the i-th layer; Then enter the wavelet transform, and its output resolution satisfies the relationship: Where W and H are the input width and height: Among them, F LL Represents a low-pass filter, F LH Represents the horizontal high-pass filter, F HL Represents a vertical high-pass filter, F HH represents a diagonal high-pass filter; S302: Perform independent convolution and scaling on the high-frequency components through the convolution layer (wavelet_convs) and the scale scaling layer (wavelet_scale) to suppress the noise in the high-frequency components. [P LL ,P LH ,P HL ,P HH ]=Conv([F LL ,F LH ,F HL ,F HH ],P) Among them, P LL ,P LH , P HL , P HH Represents the output of the four filters, F LL , F LH ,F HL ,F HH Represent four filters respectively; S303, obtain the inverse wavelet transform (IWT) through transposed convolution. First, the low-frequency and high-frequency components after wavelet transform filtering need to be passed through a small convolution kernel: Q = IWT(DW_Conv(W, WT(P))) Where DW_Conv represents depthwise convolution, W is the weight tensor of the 3×3 depthwise kernel convolution, IWT represents inverse wavelet transform, WT represents wavelet transform, and P represents input features. The high-frequency subband with enhanced features is combined with the low-frequency component in the complex domain to achieve the deblurring effect: P=Conv T ([F LL ,F LH ,F HL ,F HH ],[P LL ,P LH ,P HL ,P HH ]) Among them, Conv T represents transposed convolution, P LL ,P LH , P HL ,P HH Represents the output of the four filters, F LL ,F LH ,F HL , F HH Represent four filters respectively; S304. Using linear operations, the low frequencies obtained by the WT operation and its inverse operation are superimposed to combine different frequency outputs: Assume Y (i) It is the summary output starting from the i-th layer. The outputs of two convolutions of different sizes are summed as the output, thus obtaining the sum of convolutions at different levels.

4. A train pantograph-catenary system fault detection method based on YOLOv11 according to claim 1, characterized in that: The step S4 comprises the following steps: S401, Dynamic Range Convolution: By performing dynamic range structured reorganization on the diagonal area of the input feature matrix, high and low intensity pixels are distributed in a regularized manner, thereby supporting the convolution kernel to perform cross-channel calculations. The reorganized features are then passed through depth-wise separable convolution: X1,X2=Split(X),X1=Sort v (Sort h (X1)) in, It is a 3×3 depth convolution, Conv 1×1 It is a 1×1 point-wise convolution, Concat is a concatenation operation along the channel, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension; S402: The features after depthwise separable convolution enter the parallel dual-loop attention mechanism, which divides the output of the dynamic range convolution into value features and two pairs of query key values, and then passes them to the two loops. First, the value features V are sorted, and the query key values are sorted according to the index: V,d=Sort(V) Q1,K1=Split(Gather((X QK,1 ),d)) Q2,K2=Split(Gather((X QK,2 ),d)) Where Q, K, and V represent the query, key, and value matrices respectively, d is the index of the sort value, Gather represents a retrieval operation to search for elements in the tensor given an index value, Sort represents a horizontal or vertical sorting operation, and Split represents an operation to split features along the channel dimension. The sorted V and two pairs of key values (Q1, K1 and Q2, K2) enter two types of reconstruction and attention mechanisms respectively, and we get: A=A B ⊙A F Among them, k is the number of attention heads, Q, K, V represent the query, key, and value matrices respectively, and R B Represents the reshaping operation of BHR, R F represents the reshaping operation of FHR, A represents the obtained attention map, softmax represents the flexible maximum function, and ⊙ represents the dot product operation; S403. The reordered features are sorted back to their original positions before the final output of the 1×1 point-wise convolution, and spatial dynamic weather degradation features are extracted through parallel bidirectional attention.

5. The train pantograph-catenary system fault detection method based on YOLOv11 according to claim 1 is characterized in that: The step S5 comprises the following steps: S501. Add the ATFL function to the loss function to assign weights to various sample types, so that the model pays more attention to small samples: Among them, λ is the parameter value with 0.5 as the limit, η is the adaptive modulation factor, γ is the modulation factor, P t is the average predicted probability value of this time; S502. Adaptively improve TFL and mathematically model the model training speed: Among them, P t is the average predicted probability value, P i is the average predicted probability value of each training epoch, P c is the predicted value of the next epoch; S503, finally get the adaptive threshold focus loss function expression: Among them, P t is the average predicted probability value this time, and λ is the parameter.

6. A train pantograph-catenary system fault detection system based on YOLOv11, characterized by: The YOLOv11 network framework is adopted, including a backbone network, a neck network, and a head network; the C3K2 module in the YOLOv11 network framework is adjusted to a wavelet convolution module (C3k2WT), the C2PSA module in the YOLOv11 network framework is adjusted to an attention mechanism module (DHPA), and a loss function module (ATFL) is introduced; The backbone network is responsible for feature extraction; The neck network is located between the backbone network and the head network and is responsible for feature fusion and enhancement; The head network is the decision-making part of the target detection model and is responsible for generating the final detection results. The wavelet convolution module (C3k2WT) is responsible for processing the feature extraction at different stages of the backbone; The attention mechanism module (DHPA) enhances feature extraction capabilities through a multi-head attention mechanism and a feedforward neural network; The loss function module (ATFL) automatically adjusts the loss weight according to the characteristics of the sample and the model output.

Citation Information

Cited By

  • Small target detection method based on convolution attention and adaptive loss optimization

    CN121074590A