High-speed rail dropper insulator integrity identification method and system based on improved YOLOv9

By improving the detection model built by YOLOv9 framework, the identification accuracy and speed problems of the high-iron hanging string insulator detection system are solved, and high-precision and real-time integrity recognition of hanging string insulators are achieved, which reduces the false detection and missed detection rates and improves the industrial application capabilities of the detection system.

CN120279382AActive Publication Date: 2025-07-08GUANGZHOU INST OF RAILWAY TECH

Patent Information

Application Number
CN202510643002.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing high-iron hanging string insulator integrity detection system has low recognition accuracy and is easily affected by the subjective judgment of the operator, resulting in frequent mis-checking and missed inspections, which cannot meet the high-precision needs of industrial production.

Method used

The detection model is constructed using the improved YOLOv9 framework. Through the reparameterization improvement of the RepNCSPELAN4 module, the coordinated attention mechanism embedded in the spatial channel between the backbone network and the neck network, the introduction of hollow convolution and effectiveSE attention mechanism optimization, combined with the high-speed comprehensive detection train to collect image data and enhance data, high-quality labeled data sets are generated, iterative training is carried out and deployed to edge computing devices, and the detection results are output in real time.

Benefits of technology

It significantly improves the accuracy and speed of the integrity recognition of the hanging string insulator, reduces the leakage detection rate and false alarm rate, realizes the high accuracy and real-time detection of high-iron hanging string insulators, and improves the intelligence level of contact network operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279382A_ABST
    Figure CN120279382A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and particularly relates to a high-speed rail dropper insulator integrity identification method and system based on improved YOLOv9, and the method comprises the following steps: collecting dropper insulator image data with preset definition, and generating a double-category annotation data set; performing negative sample screening on the double-category annotation data set, and dividing a training set, a verification set and a test set based on a preset proportion of 8: 1: 1; constructing an improved detection model based on a YOLOv9 framework; performing iterative training based on the training set, monitoring convergence based on the verification set, and outputting a trained improved detection model; inputting the test set into the trained improved detection model, judging whether the performance reaches the standard or not, and outputting the trained improved detection model; inputting a to-be-detected image in real time and outputting a detection result; and generating a defect alarm signal and a visual report based on the detection result. The method has the advantages of being high in high-speed rail dropper insulator integrity recognition accuracy and high in recognition speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of object detection, and particularly relates to a method and system for identifying the integrity of pantograph suspension insulators on high-speed railways based on improved YOLOv9. Background Art

[0002] In the current field of pantograph suspension insulator integrity detection on high-speed railways, there is generally a problem that the recognition accuracy is not ideal, and the detection results are easily affected by the subjective judgment of operators. This leads to a low accuracy rate during the recognition process, and false detections and missed detections frequently occur. Traditional detection methods, such as making minor hyperparameter adjustments to the object detection model or adding additional detection levels, cannot fully meet the high-precision requirements of industrial production for the integrity detection of pantograph suspension insulators. Therefore, if there are relatively serious recognition errors in the pantograph suspension insulator detection system, it is not suitable for deployment and use in the actual industrial production environment.

[0003] Therefore, improvement is needed. Summary of the Invention

[0004] To solve the technical problems such as low accuracy and recognition speed in the integrity recognition of pantograph suspension insulators on high-speed railways, this application provides a method and system for identifying the integrity of pantograph suspension insulators on high-speed railways based on improved YOLOv9.

[0005] The first invention object of this application is achieved through the following technical solutions:

[0006] A method for identifying the integrity of pantograph suspension insulators on high-speed railways based on improved YOLOv9 includes the following steps:

[0007] Collect image data of pantograph suspension insulators with a preset clarity through a high-speed comprehensive inspection train, label the insulators in the pantograph suspension insulator image data as "normal" or "missing" categories based on a labeling platform, and generate a dual-category labeled data set;

[0008] Perform negative sample screening on the dual-category labeled data set, divide the training set, validation set, and test set based on a preset ratio of 8:1:1, and perform data augmentation of geometric transformation and color space adjustment on the training set and validation set;

[0009] Construct an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network, and optimizing the SPPCSPC module by introducing dilated convolution and the EffectiveSE attention mechanism;

[0010] Input the training set and the validation set into the improved detection model in sequence. Based on the training set, perform iterative training, and based on the validation set, monitor the convergence, and output the trained improved detection model;

[0011] Input the test set into the trained improved detection model. Based on the test set, calculate the precision, recall, mean average precision, and intersection over union, determine whether the performance meets the standard, and output the improved detection model with training completed;

[0012] Deploy the trained improved detection model to the edge computing device, input the image to be detected in real time and output the detection results, where the detection results include the insulator bounding box coordinates, the integrity status classification results, and the classification confidence scores;

[0013] Generate a defect alarm signal and a visualization report based on the detection results, and transmit them to the monitoring center through wireless communication.

[0014] In a preferred embodiment, the improved detection model is constructed based on the YOLOv9 framework. The improved detection model includes the steps of reparameterizing and improving the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module, including:

[0015] The steps of reparameterizing and improving the RepNCSPELAN4 module include:

[0016] Construct a multi-branch structure including a 3×3 convolutional layer, a 1×1 convolutional layer, and an identity mapping branch during the training phase;

[0017] In the inference phase, merge the multi-branches into a single-path structure through parameter fusion technology, where the fusion formula is:

[0018] where, W i and b i are the convolutional kernel weights and biases of each branch, and α i is the fusion coefficient.

[0019] In a preferred embodiment, the improved detection model is constructed based on the YOLOv9 framework. The improved detection model includes the steps of reparameterizing and improving the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module, and further includes:

[0020] The steps of embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network include:

[0021] Extract multi-scale spatial features through a shareable multi-semantic space attention module, where the calculation formula is: s = σ(Σ k∈{3,5,7} Conv1D k (F));

[0022] Among them, F is the input feature map, Conv1D k is a 1D convolution with a kernel size of k, and σ is the Sigmoid function;

[0023] Generate a channel weight matrix through a progressive channel self-attention module, where the calculation formula is:

[0024]

[0025] Among them, GAP and GMP are global average pooling and max pooling respectively, is the concatenation operation;

[0026] Fuse spatial and channel weights and perform feature recalibration:

[0027] Among them, is the element-wise multiplication, and · is the channel dimension broadcast multiplication.

[0028] In a preferred embodiment, the improved detection model is constructed based on the YOLOv9 framework. The improved detection model includes steps of reparameterizing and improving the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module. It also includes:

[0029] The step of introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module includes:

[0030] Construct a three-level dilated convolution layer with dilation coefficients of 2, 4, and 6 respectively;

[0031] Parallel the EffectiveSE attention module with a channel compression ratio of 16:1. Among them, the calibration formula is: F calibrated = F · δ(W2 · ReLU(W1 · GAP(F)));

[0032] Among them, and are the parameters of the fully connected layer, and δ is the Sigmoid function.

[0033] In a preferred embodiment, the step of sequentially inputting the training set and the validation set into the improved detection model, performing iterative training based on the training set, monitoring convergence based on the validation set, and outputting the trained improved detection model includes:

[0034] Adopt the cosine annealing strategy to dynamically adjust the learning rate, set the initial learning rate η0 = 0.01, the total training period T = 100, and the learning rate update formula is:

[0035]

[0036] where t is the current training iteration number;

[0037] At the end of each training period, calculate the mean average precision (mAP) and the loss value of the validation set

[0038] If the relative decrease in the loss value of the validation set is less than the threshold ∈ = 1×10 -4 within 10 consecutive training periods, then trigger the early stopping mechanism;

[0039] Select the model parameters corresponding to the peak of the mean average precision of the validation set from the training history, and load them into the improved detection model through the parameter writing-back technology.

[0040] In a preferred embodiment, the step of inputting the test set into the trained improved detection model, calculating the precision, recall, mean average precision, and intersection over union based on the test set, determining whether the performance meets the standard, and outputting the trained improved detection model includes:

[0041] The calculation formula of the precision:

[0042] where TP is the number of correctly detected insulators, and FP is the number of misdetected non-insulator targets;

[0043] The calculation formula of the recall:

[0044] where FN is the number of truly missed detected insulators;

[0045] The calculation formula of the mean average precision:

[0046] where N = 2 is the number of categories, p c (r) is the recall-precision interpolation curve of category c;

[0047] The calculation formula of the intersection over union:

[0048] where B predis the predicted bounding box area, B gt is the true annotation bounding box area;

[0049] When Precision≥95% and Recall≥93%, it is determined that the classification performance meets the standard;

[0050] When the proportion of detection boxes with IoU≥0.8 in the test set≥95%, it is determined that the localization accuracy is qualified;

[0051] When the single-frame inference time≤35ms and the false detection rate≤3%, it is determined that the real-time requirement is met.

[0052] In a preferred embodiment, the step of deploying the trained improved detection model to the edge computing device, inputting the image to be detected in real time and outputting the detection result, where the detection result includes the insulator bounding box coordinates, the integrity status classification result, and the classification confidence score, includes:

[0053] Perform normalization processing on the image to be detected, where the normalization processing includes size normalization and pixel value normalization, and output the normalized image data;

[0054] Input the normalized image data into the trained improved detection model, and generate the original detection tensor including the bounding box coordinates, classification confidence, and class scores through forward propagation;

[0055] Perform the following processing on the original detection tensor and output the detection result:

[0056] Based on the anchor box mechanism, convert the relative coordinates to absolute coordinates, and the calculation formula is: x abs =(2σ(x pred )-0.5+c x )·s w y abs =(2σ(y pred )-0.5+c y )·s h ;

[0057] where, (c x ,c y ) is the grid center coordinate, s w ,s h is the feature map stride, and σ is the Sigmoid function;

[0058] Apply the Softmax function to the class scores and retain the results with confidence≥0.9;

[0059] Sort by confidence and remove the overlapping boxes with IoU≥0.5.

[0060] The second object of the invention of this application is achieved through the following technical solutions:

[0061] The first module: Collect image data of suspension insulator with preset clarity through a high-speed comprehensive detection train, label the insulators in the suspension insulator image data as "normal" or "missing" categories based on a labeling platform, and generate a dual-category labeled data set;

[0062] The second module: Screen negative samples from the dual-category labeled data set, divide the training set, validation set and test set based on a preset ratio of 8:1:1, and perform data augmentation on the training set and validation set by geometric transformation and color space adjustment;

[0063] The third module: Build an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network, and optimizing the SPPCSPC module by introducing dilated convolution and the EffectiveSE attention mechanism;

[0064] The fourth module: Input the training set and the validation set into the improved detection model in sequence. Based on the training set, perform iterative training, and based on the validation set, monitor the convergence, and output the trained improved detection model;

[0065] The fifth module: Input the test set into the trained improved detection model. Based on the test set, calculate the precision rate, recall rate, mean average precision and intersection over union, determine whether the performance meets the standard, and output the trained improved detection model;

[0066] The sixth module: Deploy the trained improved detection model to an edge computing device, input the image to be detected in real time and output the detection result. The detection result includes the insulator bounding box coordinates, integrity status classification result, and classification confidence score;

[0067] The seventh module: Generate a defect alarm signal and a visualization report based on the detection result, and transmit them to the monitoring center through wireless communication.

[0068] The third invention object of this application is achieved through the following technical solutions:

[0069] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above method for identifying the integrity of high-speed railway suspension insulators based on improved YOLOv9.

[0070] The fourth invention object of this application is achieved through the following technical solutions:

[0071] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9 are implemented.

[0072] In summary, the present application includes at least one of the following beneficial technical effects:

[0073] The RepNCSPELAN4 module of YOLOv9 is improved by using the reparameterization method. At the same time, a spatial-channel collaborative convolutional attention mechanism is introduced and inserted into the transmission between the backbone and neck networks of YOLOv9. The SPPCSPC is optimized by using dilated convolution and the EffectiveSE attention mechanism. Through the above improvements, the integrity of pantograph suspension insulators of high-speed railways can be identified more accurately, quickly, and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 FIG. is a flowchart of an implementation of an embodiment of the method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9 according to the present application;

[0075] Figure 2 FIG. is a network structure diagram of step S30 in an embodiment of the method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9 according to the present application;

[0076] Figure 3 FIG. is another network structure diagram of step S30 in an embodiment of the method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9 according to the present application;

[0077] Figure 4 FIG. is a network model diagram of step S30 in an embodiment of the method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9 according to the present application;

[0078] Figure 5 FIG. is a schematic block diagram of a computer device according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] The following further describes the present application in detail Figures 1-5 with reference to the accompanying drawings.

[0080] In one embodiment, as Figure 1 shown, the present application discloses a method for identifying the integrity of pantograph suspension insulators of high-speed railways based on improved YOLOv9, which specifically includes the following steps:

[0081] S10: Collect image data of pantograph suspension insulators with a preset clarity through a high-speed comprehensive inspection train, label the insulators in the pantograph suspension insulator image data as "normal" or "missing" categories based on a labeling platform, and generate a dual-category labeled data set;

[0082] S20: Screen the negative samples from the dual-class labeled dataset, divide the training set, validation set, and test set based on the preset ratio of 8:1:1, and perform data augmentation on the training set and validation set by geometric transformation and color space adjustment;

[0083] S30: Build an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module;

[0084] S40: Input the training set and the validation set into the improved detection model in sequence. Based on the training set, perform iterative training, and based on the validation set, monitor the convergence, and output the trained improved detection model;

[0085] S50: Input the test set into the trained improved detection model. Based on the test set, calculate the precision, recall, mean average precision, and intersection over union, determine whether the performance meets the standard, and output the trained improved detection model;

[0086] S60: Deploy the trained improved detection model to the edge computing device, input the image to be detected in real-time and output the detection results. The detection results include the insulator bounding box coordinates, the integrity status classification results, and the classification confidence scores;

[0087] S70: Generate a defect alarm signal and a visualization report based on the detection results, and transmit them to the monitoring center through wireless communication.

[0088] In this embodiment, high-definition images are collected by a high-speed integrated detection train (S10), and a dual-class labeled dataset is generated in combination with a professional annotation platform to provide high-quality supervision signals for model training; the dataset is divided in a ratio of 8:1:1 (S20). While maintaining the consistency of data distribution, data augmentation is implemented through geometric transformation and color space adjustment, enabling the model to have stronger robustness to scenarios such as illumination changes and perspective offsets, and effectively reducing the risk of overfitting. An improved detection model is built based on the YOLOv9 framework (S30), and performance breakthroughs are achieved through three core improvements: reparameterization improvement of the RepNCSPELAN4 module to improve the inference efficiency; embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network to enhance the feature representation ability; introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module to expand the receptive field and strengthen the channel dimension feature extraction. These improvements enable the model to achieve a better balance between detection accuracy and speed.

[0089] Iterative training is carried out using the training set, and the convergence is monitored in real time through the validation set (S40). Hyperparameters such as the learning rate are dynamically adjusted based on the change of the loss function to ensure that the model converges stably to the optimal state during training; finally, metrics such as precision, recall, mAP, and IoU are calculated through the test set (S50), and a quantitative performance evaluation system is established to ensure that the model meets the engineering application standards. The trained model is deployed to the edge computing device (S60), and lightweight deployment is achieved through model quantization and optimization techniques. On the premise of ensuring the detection accuracy, the single-frame inference time meets the real-time requirements; the structured detection results including the bounding box coordinates, classification results, and confidence scores are output to provide a complete information chain for defect location and identification. Based on the detection results, a defect alarm signal and a visualization report are automatically generated (S70) and transmitted to the monitoring center in real time through the wireless communication module to form a complete business closed-loop of "detection-identification-alarm-disposal". This system reduces the missed detection rate of insulator defect detection, controls the false alarm rate within the acceptable range of the project, and significantly improves the intelligent level of catenary operation and maintenance.

[0090] Figure 2 , step S30 includes:

[0091] S301: The steps for reparameterization improvement of the RepNCSPELAN4 module include:

[0092] Construct a multi-branch structure including a 3×3 convolutional layer, a 1×1 convolutional layer, and an identity mapping branch during the training phase;

[0093] S302: In the inference phase, the multi-branches are merged into a single-path structure through parameter fusion technology. The fusion formula is:

[0094] S303: Among them, W i and b i are the convolutional kernel weights and biases of each branch, and α i is the fusion coefficient.

[0095] In this embodiment, the reparameterization improvement of the RepNCSPELAN4 module uses RepConv. RepConv is an advanced model reparameterization method that significantly improves the efficiency and performance of the model by merging multiple computational units into one unit during the inference phase. The basic idea of RepConv is to use multi-path convolutional layers during training, and then, during inference, use reparameterization techniques to merge the parameters of these path layers into the main path. This method not only greatly reduces the computational requirements but also significantly reduces the memory usage.

[0096] Figure 3 , step S30 further includes:

[0097] S304: The steps of embedding the spatial channel collaborative attention mechanism between the backbone network and the neck network include:

[0098] Extract multi-scale spatial features through a shareable multi-semantic spatial attention module, where the calculation formula is: s = σ(∑ k∈{3,5,7} Conv1D k (F));

[0099] S305: Among them, F is the input feature map, Conv1D k is a 1D convolution with a kernel size of k, and σ is the Sigmoid function;

[0100] S306: Generate a channel weight matrix through a progressive channel self-attention module, where the calculation formula is:

[0101]

[0102] S307: Among them, GAP and GMP are global average pooling and max pooling respectively, is the concatenation operation;

[0103] S308: Fuse the spatial and channel weights and perform feature recalibration:

[0104] S309: Among them, is the element-wise multiplication, and · is the channel dimension broadcast multiplication.

[0105] In this embodiment, on the images taken by the China Railway High-Speed Comprehensive Inspection Train, using the spatial and channel collaborative attention mechanism can extract the attention regions of the feature maps, helping YOLOv9 distinguish the information of complex catenary insulators and enabling the network to focus on the targets to be detected. The Spatial and Channel Synergy Attention (SCSA) algorithm is an attention mechanism used in deep learning models, especially in image processing tasks such as object detection, image classification, and semantic segmentation. The principle of this algorithm is to enhance the model's attention to important features by considering both the spatial dimension and the channel dimension of the feature maps, thereby improving the performance of the model.

[0106] Specifically, the basic principle of the SCSA algorithm is as follows: Shareable Multi-Semantic Space Attention (SMSA). Principle: SMSA captures the multi-semantic space information of each feature channel through multi-scale and depth-shared 1D convolutions, thus effectively integrating global context dependencies and multi-semantic space priors; Progressive Channel Self-Attention (PCSA). Principle: PCSA uses the self-attention mechanism to adjust the mutual relationships between channels. By calculating the feature map in the channel dimension, where A is the channel attention weight, FC is the fully connected layer, Concat is the concatenation operation, and GAP and GMP are global average pooling and global max pooling respectively; Spatial and Channel Cooperative Attention (SCSA). Principle: SCSA combines SMSA and PCSA, considering both spatial and channel information. Through these three steps, the SCSA algorithm can effectively highlight the key spatial regions in the image and enhance the model's attention to specific features through adjustments in the channel dimension, thereby achieving better performance in visual tasks.

[0107] Figure 4 , step S30 further includes:

[0108] SA1: The step of introducing dilated convolution and EffectiveSE attention mechanism to optimize the SPPCSPC module includes:

[0109] Construct a three-level dilated convolution layer with dilation coefficients of 2, 4, and 6 respectively;

[0110] SA2: Parallel the EffectiveSE attention module with a channel compression ratio set to 16:1. Among them, the calibration formula is: F calibrated = F · δ(W2 · ReLU(W1 · GAP(F)));

[0111] SA3: Among them, and are the parameters of the fully connected layer, and δ is the Sigmoid function.

[0112] In this embodiment, Dilated Convolution is a technique used in the field of deep learning to expand the receptive field of a convolutional neural network, especially outstanding in tasks such as image processing. This technique is achieved by inserting "holes" or "gaps" between the elements of the convolutional kernel, which can increase the field of view of the convolutional operation without increasing the number of parameters or computational cost. The EffectiveSE (Effective Squeeze-and-Excitation) attention mechanism is an improved Squeeze-and-Excitation (SE) module used to enhance feature representation in deep learning models. The SE module improves the representational ability of the network by explicitly modeling the dependencies between channels. EffectiveSE is optimized on this basis to improve efficiency and performance.

[0113] Step S40 includes:

[0114] S401: Dynamically adjust the learning rate using the cosine annealing strategy, set the initial learning rate η0 = 0.01, the total number of training epochs T = 100, and the learning rate update formula is:

[0115]

[0116] S402: where t is the current training iteration number;

[0117] S403: At the end of each training epoch, calculate the mean average precision (mAP) and loss value of the validation set

[0118] S404: If the relative decrease in the loss value of the validation set is less than the threshold ∈ = 1×10 -4 within 10 consecutive training epochs, then trigger the early stopping mechanism;

[0119] S405: Select the model parameters corresponding to the peak of the mean average precision of the validation set from the training history and load them into the improved detection model through the parameter writing-back technique.

[0120] In this embodiment, during the model training stage (S401 - S405), the training process is dynamically optimized through the cosine annealing learning rate strategy, setting the initial learning rate η0 = 0.01 and the training epoch T = 100, and the learning rate decays periodically according to the formula (t is the current iteration number), which not only avoids the local optimal trap but also accelerates the model convergence; at the same time, combined with the intelligent early stopping mechanism, calculate the mean average precision and loss value of the validation set at the end of each training epoch If the loss decrease is less than the threshold ∈ = 1×10 within 10 consecutive epochs -4, the training will be automatically terminated and the parameter writing-back technology will be triggered, and the model parameters corresponding to the peak value of the average precision of the validation set in the training history will be loaded into the improved detection model.

[0121] Step S50 includes:

[0122] S501: The calculation formula of the precision rate:

[0123] S502: Among them, TP is the number of correctly detected insulators, and FP is the number of misdetected non-insulator targets;

[0124] S503: The calculation formula of the recall rate:

[0125] S504: Among them, FN is the number of real insulators that are missed;

[0126] S505: The calculation formula of the mean average precision:

[0127] S506: Among them, N = 2 is the number of categories, p c (r) is the recall-precision interpolation curve of class c;

[0128] S507: The calculation formula of the intersection over union:

[0129] S508: Among them, B pred is the predicted box area, and B gt is the real annotation box area;

[0130] S509: When Precision ≥ 95% and Recall ≥ 93%, it is determined that the classification performance meets the standard;

[0131] S510: When the proportion of detection boxes with IoU ≥ 0.8 ≥ 95%, it is determined that the localization accuracy is qualified;

[0132] S511: When the single-frame inference time ≤ 35ms and the false detection rate ≤ 3%, it is determined that the real-time requirement is met.

[0133] In this embodiment, through the binary evaluation system of precision (S501 - S502) and recall (S503 - S504), the recognition accuracy (TP / (TP + FP)) and coverage integrity (TP / (TP + FN)) of the model for positive samples are quantified respectively, forming the basic performance scale for the detection task. The mean average precision (mAP, S505 - S506) further integrates the multi - category detection results, draws the PR curve by interpolation method and calculates the area under the curve to achieve the standardized comparison of cross - category performance. The intersection over union (IoU, S507 - S508), as the core index of geometric matching degree, converts the spatial positioning accuracy into a computable continuous value through the ratio of the intersection / union of the predicted box and the ground - truth box (A∩B / A∪B), providing an objective evaluation basis for the quality of bounding box regression. Set double passing thresholds (S509 - S511), ensure the classification reliability when the precision ≥ 95% and the recall ≥ 93%; require that the proportion of detection boxes with IoU ≥ 0.8 exceeds 95% to guarantee the positioning accuracy; finally, through the real - time constraint of single - frame inference time ≤ 35ms and false - detection rate ≤ 3%, a full - chain evaluation standard from detection accuracy to computational efficiency is constructed.

[0134] Step S60 includes:

[0135] S601: Perform normalization processing on the image to be detected. The normalization processing includes size normalization and pixel - value normalization, and output the normalized image data;

[0136] S602: Input the normalized image data into the trained improved detection model, and generate the original detection tensor containing bounding box coordinates, classification confidence, and class scores through forward propagation;

[0137] S603: Perform the following processing on the original detection tensor and output the detection results:

[0138] S604: Based on the anchor box mechanism, convert the relative coordinates to absolute coordinates. The calculation formula is: x abs =(2σ(x pred ) - 0.5 + c x )·s w y abs =(2σ(y pred ) - 0.5 + c y )·s h ;

[0139] S605: Among them, (c x , c y ) is the grid center coordinate, s w , s h is the feature map stride, and σ is the Sigmoid function;

[0140] S606: Apply the Softmax function to the class scores and retain the results with confidence ≥ 0.9;

[0141] S607: Sort by confidence and remove overlapping boxes with IoU ≥ 0.5.

[0142] In this embodiment, the input image resolution is unified through size normalization (S601) to eliminate the feature extraction deviation caused by image scale differences; pixel value normalization maps the pixel values to a standard interval, accelerating the model training convergence and enhancing the generalization ability, providing a standardized input basis for subsequent detection. Feature extraction and decoding mechanism: The improved detection model (S602) generates an original tensor containing bounding box coordinates, classification confidence, and class scores through forward propagation. Its model optimization may involve a Feature Pyramid Network (FPN) or multi-scale feature fusion to enhance the feature expression ability for small target insulators; in the coordinate decoding stage (S604 - S605), the anchor box mechanism is adopted. Through the linear mapping of the grid center coordinates and the feature map stride, combined with the Sigmoid function, the relative coordinates are converted into absolute coordinates to achieve accurate spatial positioning of the detection box. The class scores are normalized by the Softmax function (S606) to generate a probability distribution vector and filter out low-confidence predictions, effectively suppressing the interference of background noise; Non-Maximum Suppression (NMS, S607) screens through confidence sorting and IoU thresholds to remove redundant detection boxes with an overlap degree exceeding the threshold, reducing the false detection rate to an acceptable range in engineering while ensuring the recall rate.

[0143] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0144] In one embodiment, a high-speed catenary insulator integrity recognition system based on improved YOLOv9 is provided. This high-speed catenary insulator integrity recognition system based on improved YOLOv9 corresponds to the above-mentioned high-speed catenary insulator integrity recognition method based on improved YOLOv9. This high-speed catenary insulator integrity recognition system based on improved YOLOv9 includes:

[0145] The first module: Collect catenary insulator image data with a preset clarity through a high-speed comprehensive inspection train, label the insulators in the catenary insulator image data as "normal" or "missing" categories based on a labeling platform, and generate a dual-category labeled data set;

[0146] The second module: Screen the negative samples in the dual-category labeled data set, divide the training set, validation set, and test set based on a preset ratio of 8:1:1, and perform data augmentation on the training set and validation set by geometric transformation and color space adjustment;

[0147] The third module: constructing an improved detection model based on the YOLOv9 framework, where the improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network, introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module;

[0148] The fourth module: sequentially inputting the training set and the validation set into the improved detection model, performing iterative training based on the training set, and monitoring the convergence based on the validation set, and outputting the trained improved detection model;

[0149] The fifth module: inputting the test set into the trained improved detection model, calculating the precision, recall, mean average precision, and intersection over union based on the test set, determining whether the performance meets the standard, and outputting the trained improved detection model;

[0150] The sixth module: deploying the trained improved detection model to an edge computing device, inputting the image to be detected in real time and outputting the detection result, where the detection result includes the insulator bounding box coordinates, the integrity status classification result, and the classification confidence score;

[0151] The seventh module: generating a defect alarm signal and a visualization report based on the detection result, and transmitting them to the monitoring center through wireless communication.

[0152] Optionally, it further includes:

[0153] The first included module: the steps of reparameterization improvement of the RepNCSPELAN4 module include:

[0154] Constructing a multi-branch structure including a 3×3 convolutional layer, a 1×1 convolutional layer, and an identity mapping branch during the training phase;

[0155] The first formula module: merging the multi-branches into a single-path structure through parameter fusion technology during the inference phase, where the fusion formula is:

[0156] The first where module: where, W i and b i are the convolutional kernel weights and biases of each branch, and α i is the fusion coefficient.

[0157] Optionally, it further includes:

[0158] The second included module: the steps of embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network include:

[0159] Extract multi-scale spatial features through a shareable multi-semantic space attention module, where the calculation formula is: S = σ(∑ k∈{3,5,7} Conv1D k (F));

[0160] Second "where" module: where F is the input feature map, Conv1D k is a 1D convolution with a kernel size of k, and σ is the Sigmoid function;

[0161] Second formula module: Generate a channel weight matrix through a progressive channel self-attention module, where the calculation formula is:

[0162]

[0163] Third "where" module: where GAP and GMP are global average pooling and max pooling respectively, is a concatenation operation;

[0164] First fusion module: Fuse spatial and channel weights and perform feature recalibration:

[0165] Fourth "where" module: where is element-wise multiplication, and · is channel dimension broadcast multiplication.

[0166] Optionally, it further includes:

[0167] Third "including" module: The steps of introducing dilated convolution and EffectiveSE attention mechanism to optimize the SPPCSPC module include:

[0168] Construct a three-level dilated convolution layer with dilation coefficients of 2, 4, and 6 respectively;

[0169] Fifth "where" module: Parallel the EffectiveSE attention module, and set its channel compression ratio to 16:1, where the calibration formula is: F calibrated = F · δ(W2 · ReLU(W1 · GAP(F)));

[0170] Sixth "where" module: where and are the parameters of the fully connected layer, and δ is the Sigmoid function.

[0171] Optionally, it further includes:

[0172] Third formula module: Adopt the cosine annealing strategy to dynamically adjust the learning rate, set the initial learning rate η0 = 0.01, the total training period T = 100, and the learning rate update formula is:

[0173]

[0174] Seventh "where" module: where t is the current training iteration number;

[0175] First calculation module: At the end of each training cycle, calculate the mean average precision (mAP) and the loss value L of the validation set val ;

[0176] First trigger module: If the relative decrease in the loss value of the validation set is less than the threshold ∈ = 1×10 within 10 consecutive training cycles -4 , then trigger the early stopping mechanism;

[0177] First selection module: Select the model parameters corresponding to the peak of the mean average precision of the validation set from the training history, and load them into the improved detection model through the parameter writing-back technology.

[0178] Optionally, it further includes:

[0179] Second calculation module: The formula for calculating the precision:

[0180] Eighth "where" module: where TP is the number of correctly detected insulators, and FP is the number of misdetected non-insulator targets;

[0181] Third calculation module: The formula for calculating the recall:

[0182] Ninth "where" module: where FN is the number of truly missed insulators;

[0183] Calculation module: The formula for calculating the mean average precision:

[0184] "where" module: where N = 2 is the number of categories, p c (r) is the recall-precision interpolation curve for category c;

[0185] Fifth calculation module: The formula for calculating the intersection over union:

[0186] Fifth "where" module: where B pred is the predicted box region, and B gt is the true annotation box region;

[0187] First determination module: When Precision ≥ 95% and Recall ≥ 93%, determine that the classification performance meets the standard;

[0188] Second determination module: When the proportion of detection boxes with IoU ≥ 0.8 ≥ 95%, determine that the localization accuracy is qualified;

[0189] Third determination module: When the single-frame inference time ≤ 35 ms and the false detection rate ≤ 3%, it is determined that the real-time requirement is met.

[0190] Optionally, it further includes:

[0191] Including a module to perform normalization processing on the image to be detected. The normalization processing includes size normalization and pixel value normalization, and outputs normalized image data;

[0192] Input module: Input the normalized image data into the trained improved detection model, and generate an original detection tensor containing bounding box coordinates, classification confidence, and class scores through forward propagation;

[0193] Output module: Perform the following processing on the original detection tensor and output the detection result:

[0194] Sixth calculation module: Based on the anchor box mechanism, convert the relative coordinates to absolute coordinates. The calculation formula is: x abs =(2σ(x pred ) - 0.5 + c x )·s w y abs =(2σ(y pred ) - 0.5 + c y )·s h ;

[0195] Sixth where module: Among them, (c x , c y ) is the grid center coordinate, s w , s h is the feature map stride, and σ is the Sigmoid function;

[0196] Retention module: Apply the Softmax function to the class scores and retain the results with a confidence ≥ 0.9;

[0197] Elimination module: Sort by confidence and eliminate the overlapping boxes with IoU ≥ 0.5.

[0198] For the specific limitations of a high-speed rail suspension string insulator integrity recognition system based on improved YOLOv9, reference can be made to the limitations of a high-speed rail suspension string insulator integrity recognition method based on improved YOLOv9 in the above text, which will not be elaborated here. Each module in the above high-speed rail suspension string insulator integrity recognition system based on improved YOLOv9 can be implemented in whole or in part through software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0199] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 5 . The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store catenary insulator image data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for identifying the integrity of high-speed railway catenary insulators based on improved YOLOv9.

[0200] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a method for identifying the integrity of high-speed railway catenary insulators based on improved YOLOv9.

[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it is a method for identifying the integrity of high-speed railway catenary insulators based on improved YOLOv9.

[0202] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0203] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

Claims

1. An integrity recognition method for pantograph suspension insulators of high-speed railways based on improved YOLOv9, characterized in that, It includes the following steps: Collect the image data of suspension insulator with preset clarity by a high-speed comprehensive detection train, label the insulators in the suspension insulator image data as "normal" or "missing" categories based on a labeling platform, and generate a dual-category labeled data set; Perform negative sample screening on the dual-category labeled data set, divide it into a training set, a validation set and a test set based on a preset ratio of 8:1:1, and perform data augmentation on the training set and the validation set by geometric transformation and color space adjustment; Build an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and EffectiveSE attention mechanism optimization for the SPPCSPC module; Input the training set and the validation set into the improved detection model in sequence. Based on the training set, perform iterative training, and based on the validation set, monitor the convergence, and output the trained improved detection model; Input the test set into the trained improved detection model. Based on the test set, calculate the precision, recall, mean average precision and intersection over union, determine whether the performance meets the standard, and output the trained improved detection model; Deploy the trained improved detection model to an edge computing device, input the image to be detected in real time and output the detection result. The detection result includes the insulator bounding box coordinates, the integrity status classification result, and the classification confidence score; Generate a defect alarm signal and a visualization report based on the detection result, and transmit them to the monitoring center through wireless communication.

2. The method according to claim 1, wherein The step of building an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and EffectiveSE attention mechanism optimization for the SPPCSPC module, includes: The step of reparameterization improvement of the RepNCSPELAN4 module, includes: Construct a multi-branch structure containing a 3×3 convolutional layer, a 1×1 convolutional layer and an identity mapping branch during the training phase; In the inference stage, the multi-branches are merged into a single-path structure through the parameter fusion technology, where the fusion formula is: Among them, W i and b i are the convolution kernel weights and biases of each branch, and α i is the fusion coefficient.

3. The method according to claim 1, wherein The step of building an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and EffectiveSE attention mechanism optimization for the SPPCSPC module, also includes: The step of embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, includes: Extract multi-scale spatial features through a shareable multi-semantic space attention module, where the calculation formula is: S = σ(Σ k∈{3,5,7} Conv1D k (F)); where F is the input feature map, Conv1D k is a 1D convolution with kernel size k, and σ is the Sigmoid function; Generate a channel weight matrix through a progressive channel self-attention module, where the calculation formula is: Among them, GAP and GMP are global average pooling and max pooling respectively, which is a concatenation operation; Fuse spatial and channel weights and perform feature recalibration: wherein, is element-wise multiplication, and · is channel dimension broadcast multiplication.

4. The method according to claim 1, wherein The improved detection model is constructed based on the YOLOv9 framework. The improved detection model includes the steps of reparameterizing and improving the RepNCSPELAN4 module, embedding a spatial-channel collaborative attention mechanism between the backbone network and the neck network, and optimizing the SPPCSPC module by introducing dilated convolution and the EffectiveSE attention mechanism. It also includes: The step of optimizing the SPPCSPC module by introducing dilated convolution and the EffectiveSE attention mechanism includes: Construct a three-level dilated convolution layer with dilation coefficients of 2, 4, and 6 respectively; Parallel Effective SE Attention Module with a channel compression ratio set to 16:1, where the calibration formula is: F calibrated = F · δ(W2 · ReLU(W1 · GAP(F))); Among them, and are the parameters of the fully connected layer, and δ is the Sigmoid function.

5. The method according to claim 1, characterized in that, The step of sequentially inputting the training set and the validation set into the improved detection model, performing iterative training based on the training set, monitoring the convergence based on the validation set, and outputting the trained improved detection model includes: Adopt the cosine annealing strategy to dynamically adjust the learning rate. Set the initial learning rate η0 = 0.01, the total training period T = 100, and the learning rate update formula is: where t is the current training iteration number; At the end of each training cycle, calculate the mean average precision (mAP) and the loss value of the validation set If the relative decrease in the validation set loss value within 10 consecutive training cycles is less than the threshold ∈ = 1×10 -4 , then the early stopping mechanism is triggered; Select the model parameters corresponding to the peak value of the average precision mean of the validation set from the training history and load them into the improved detection model through the parameter writing-back technology.

6. The method according to claim 1, characterized in that, The step of inputting the test set into the trained improved detection model, calculating the precision, recall, mean average precision, and intersection over union based on the test set, determining whether the performance meets the standard, and outputting the trained improved detection model includes: The calculation formula for the precision rate: where TP is the number of correctly detected insulators, and FP is the number of misdetected non-insulator targets; The calculation formula for the recall rate: where FN is the number of real insulators missed; The calculation formula for the mean average precision: where N = 2 is the number of classes, and p c (r) is the recall-precision interpolation curve for class c; The calculation formula of the intersection over union: Among them, B pred is the predicted bounding box area, and B gt is the ground truth bounding box area; When Precision ≥ 95% and Recall ≥ 93%, it is determined that the classification performance meets the standard; When the proportion of detection boxes with IoU ≥ 0.8 in the test set ≥ 95%, it is determined that the localization accuracy is qualified; When the single-frame inference time ≤ 35ms and the false detection rate ≤ 3%, it is determined that the real-time requirement is met.

7. The method according to claim 1, wherein The step of deploying the trained improved detection model to the edge computing device, inputting the image to be detected in real time and outputting the detection result, where the detection result includes the insulator bounding box coordinates, the integrity status classification result, and the classification confidence score includes: Perform normalization processing on the image to be detected. The normalization processing includes size normalization and pixel value normalization, and output the normalized image data; Input the normalized image data into the trained improved detection model, and generate an original detection tensor containing the bounding box coordinates, classification confidence, and class scores through forward propagation; Perform the following processing on the original detection tensor and output the detection result: Convert the relative coordinates to absolute coordinates based on the anchor box mechanism. The calculation formula is: x abs =(2σ(x pred ) - 0.5 + c x )·s w y abs =(2σ(y pred ) - 0.5 + c y )·s h ; Among them, (c x , c y ) is the grid center coordinate, s w , s h is the feature map stride, and σ is the Sigmoid function; Apply the Softmax function to the class scores and retain the results with confidence ≥ 0.9; Sort by confidence and remove the overlapping boxes with IoU ≥ 0.

5.

8. An integrity recognition system for pantograph suspension insulators of high-speed railways based on improved YOLOv9, characterized in that, It includes the following steps: The first module: Collect the catenary insulator image data with a preset clarity through a high-speed comprehensive detection train, label the insulators in the catenary insulator image data as "normal" or "missing" categories based on the annotation platform, and generate a dual-category annotation data set; The second module: perform negative sample screening on the dual-category labeled dataset, divide the training set, validation set, and test set based on the preset ratio of 8:1:1, and perform data augmentation of geometric transformation and color space adjustment on the training set and validation set; The third module constructs an improved detection model based on the YOLOv9 framework. The improved detection model includes reparameterization improvement of the RepNCSPELAN4 module, embedding a spatial channel collaborative attention mechanism between the backbone network and the neck network, and introducing dilated convolution and the EffectiveSE attention mechanism to optimize the SPPCSPC module; The fourth module: sequentially input the training set and the validation set into the improved detection model, perform iterative training based on the training set, and monitor the convergence based on the validation set, and output the trained improved detection model; The fifth module: input the test set into the trained improved detection model, calculate the precision rate, recall rate, mean average precision, and intersection over union based on the test set, determine whether the performance meets the standard, and output the trained improved detection model; The sixth module: deploy the trained improved detection model to an edge computing device, input the image to be detected in real time and output the detection result. The detection result includes the insulator bounding box coordinates, the integrity status classification result, and the classification confidence score; The seventh module: generate a defect alarm signal and a visualization report based on the detection result, and transmit them to the monitoring center through wireless communication.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of an integrity recognition method for high-speed rail suspension insulator based on improved YOLOv9 as claimed in claims 1-7 are implemented.

10. A computer-readable storage medium storing a computer program, which when executed by a processor, implements the steps of an integrity recognition method for high-speed rail suspension insulator based on improved YOLOv9 as claimed in claims 1-7.

Citation Information

Patent Citations

  • Insulator defect detection method in foggy day scene based on improved YOLOv7 algorithm

    CN116843636A

  • Insulator defect identification method based on improved YOLOv8 model and cosine annealing learning rate attenuation method

    CN117197530A

  • Rainy day insulator defect detection method based on improved MS-PReNet and GAM-YOLOv7

    CN117422689A

  • Unmanned aerial vehicle image all-weather detection method and system under urban low-altitude complex background

    CN118115952A

  • Glass bead defect detection method based on improved spatial pyramid pooling

    CN118485631A

Cited By

  • Breaker switch intelligent detection method and system based on deep learning, medium and equipment

    CN120932036A

  • A circuit breaker switch intelligent detection method, system, medium and device based on deep learning

    CN120932036B

  • Optical fiber distributed sensing signal identification method and system with low false detection rate

    CN120995053A

  • Detection method, device and equipment for safety of tower crane and storage medium

    CN122156620A