Pantograph slide plate abnormity optimization detection method and system

By using a fusion technology of gated network and hybrid model in pantograph skateboard abnormality detection, the problem of low detection accuracy in the prior art is solved, and efficient detection and accurate identification of various types of abnormalities are achieved.

CN120070379APending Publication Date: 2025-05-30LI CHUANG ZHI HENG ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150796.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has low detection accuracy when detecting pantograph skateboard abnormalities, especially when dealing with various types of abnormalities, the recognition accuracy is insufficient and the adaptability is poor.

Method used

A pantograph skateboard abnormality optimization detection method is adopted. By obtaining the image to be processed and inputting it into the gated network to obtain weights, the images are then inputted into the crack, block drop and foreign object detection models in the mixed model, the detection results are respectively output, and the final fusion detection results are output based on the weights.

Benefits of technology

It improves the accuracy and adaptability of pantograph skateboard abnormal detection, can effectively deal with various types of abnormalities, and improves processing speed and detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070379A_ABST
    Figure CN120070379A_ABST
Patent Text Reader

Abstract

The invention provides a pantograph slide plate abnormity optimization detection method and system, and the method comprises the steps: obtaining a to-be-processed image, inputting the to-be-processed image to a gating network, outputting a weight through the gating network, inputting the to-be-processed image to a hybrid model, and outputting a first detection result through a crack detection model, a second detection result is output through the chipping detection model, a third detection result is output through the foreign matter detection model, and the first detection result, the second detection result and the third detection result comprise crack, chipping and foreign matter detection results; and based on the weight, fusing the first detection result, the second detection result and the third detection result, and outputting a fused detection result. According to the method, different detection models are set for different types of anomalies, and the detection models are matched through the gating network, so that the processing speed can be improved while the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a method and system for optimizing the detection of pantograph slider anomalies. Background Art

[0002] The pantograph slider of a train is the part installed on the pantograph of an electric locomotive or multiple unit train, which directly contacts the overhead catenary, and is responsible for conducting current from the catenary to the electrical system of the train to provide power for the train. The working conditions of the pantograph slider are relatively harsh, and in actual use, the slider will show wear and other phenomena during use.

[0003] To detect the pantograph slider, manual inspection or image recognition technology based on a single detection model can be used. Manual inspection has a high work intensity, low efficiency, and is easily affected by human factors. The method based on a single model faces problems of insufficient recognition accuracy and poor adaptability in dealing with different types of anomalies, such as crack, chunk loss, and foreign object detection.

[0004] It is also possible to use a target detection method based on deep learning. For example, YOLO (You Only Look Once) is an algorithm for target detection. Among them, YOLOv8 has excellent real-time target detection capabilities, but only focuses on a certain type of anomaly and cannot perform joint detection of multiple types of anomalies, resulting in relatively low detection accuracy. Summary of the Invention

[0005] The present application provides a method and system for optimizing the detection of pantograph slider anomalies to solve the problem of relatively low detection accuracy.

[0006] In a first aspect, the present application provides a method for optimizing the detection of pantograph slider anomalies, including:

[0007] Obtain an image to be processed;

[0008] Input the image to be processed into a gating network to output weights through the gating network;

[0009] Input the image to be processed into a hybrid model to output a first detection result through a crack detection model, a second detection result through a chunk loss detection model, and a third detection result through a foreign object detection model. The first detection result, the second detection result, and the third detection result include the detection results of cracks, chunk loss, and foreign objects;

[0010] Based on the weights, fuse the first detection result, the second detection result, and the third detection result to output a fused detection result.

[0011] In some feasible embodiments, the step of inputting the image to be processed into a gating network to output weights through the gating network includes:

[0012] Input the image to be processed into the convolutional neural network of the gating network to extract global features;

[0013] Through linear transformation, calculate the first score, the second score, and the third score based on the global features. The first score is the score for the crack detection task, the second score is the score for the chipping detection task, and the third score is the score for the foreign object detection task;

[0014] Use an activation function to convert the first score into a first weight, convert the second score into a second weight, and convert the third score into a third weight.

[0015] In some feasible embodiments, based on the weights, fuse the first detection result, the second detection result, and the third detection result, and output a fused detection result, including:

[0016] Perform weighted summation on the first detection result and the first weight, and output a first fused result;

[0017] Perform weighted summation on the second detection result and the second weight, and output a second fused result;

[0018] Perform weighted summation on the third detection result and the third weight, and output a third fused result;

[0019] Based on the first fused result, the second fused result, and the third fused result, output a fused detection result, which is the result obtained by performing element-wise addition on the first fused result, the second fused result, and the third fused result.

[0020] In some feasible embodiments, input the image to be processed into the hybrid model to output a first detection result through the crack detection model, including:

[0021] Input the image to be processed into the first convolutional layer of the crack detection model to adjust the number of channels through the first convolutional layer, and split the image to be processed into a first sub-image and a second sub-image;

[0022] Input the second sub-image into the serpentine convolutional module of the crack detection model to perform feature extraction through the serpentine convolutional module, and output a feature extraction map;

[0023] Concatenate the first sub-image and the feature extraction map, and perform channel number adjustment on the concatenated image to output a first detection result.

[0024] In some feasible embodiments, output a second detection result through the chipping detection model, including:

[0025] Input the image to be processed into the backbone network of the block loss detection model to generate multi-scale feature maps;

[0026] Input the feature maps into the neck network to output the fused feature maps;

[0027] Input the fused feature maps into the head network to output the second detection result, where the head network includes the NWD loss function.

[0028] In some feasible embodiments, the output of the third detection result by the foreign object detection model includes:

[0029] Input the image to be processed into the foreign object detection model to perform the first average pooling on the image to be processed by the foreign object detection model and obtain a first pooled image, where the width of the first pooled image is 1;

[0030] Perform the second average pooling on the image to be processed to obtain a second pooled image, where the height of the second pooled image is 1;

[0031] Combine the first pooled image and the second pooled image to generate a feature layer, where the height of the feature layer is 1 and the width is the sum of the height of the first pooled image and the width of the second pooled image;

[0032] Output the third detection result based on the feature layer.

[0033] In some feasible embodiments, the output of the third detection result based on the feature layer includes:

[0034] Separate the feature layer to output a first sub-feature layer and a second sub-feature layer, where the height of the first sub-feature layer is 1 and the width is the height of the image to be processed, and the width of the second sub-feature layer is 1 and the height is the width of the image to be processed;

[0035] Transpose the first sub-feature layer to output a third sub-feature layer, where the height of the third sub-feature layer is the height of the image to be processed and the width is 1;

[0036] Transpose the second sub-feature layer to output a fourth sub-feature layer, where the height of the fourth sub-feature layer is 1 and the width is the width of the image to be processed;

[0037] Perform channel adjustment on the third sub-feature layer through a second convolutional layer to obtain a fifth sub-feature layer, and perform feature adjustment on the fourth sub-feature layer through a third convolutional layer to obtain a sixth sub-feature layer;

[0038] Apply an activation function to the fifth sub-feature layer to output a first attention score map, and apply an activation function to the sixth sub-feature layer to output a second attention score map;

[0039] Multiply the first attention score map by the image to be processed to output a first modulated feature map, and multiply the second attention score map by the image to be processed to output a second modulated feature map;

[0040] Output a third detection result based on the first modulated feature map and the second modulated feature map.

[0041] In some feasible embodiments, before inputting the image to be processed into the hybrid model to output a first detection result through the crack detection model and a second detection result through the spalling detection model, it includes:

[0042] Perform preprocessing on the image to be processed to obtain a preprocessed image;

[0043] Remove the background of the preprocessed image based on a segmentation algorithm to extract an image of the carbon skateboard area;

[0044] Based on the image of the carbon skateboard area, output a first detection result and a second detection result through the hybrid model.

[0045] In some feasible embodiments, the first detection result includes a first detection area detected and a confidence score of the first detection area, the second detection result includes a second detection area and a confidence score of the second detection area, and the third detection result includes a third detection area and a confidence score of the third detection area;

[0046] The fusing the first detection result, the second detection result, and the third detection result based on the weights to output a fused detection result includes:

[0047] Calculate the center point distances, which include a first center point distance, a second center point distance, and a third center point distance. The first center point distance is the Euclidean distance of the center point of the first detection area, the second center point distance is the Euclidean distance of the center point of the second detection area, and the third center point distance is the Euclidean distance of the center point of the third detection area;

[0048] Set a distance threshold;

[0049] If at least two center point distances are less than the distance threshold, merge the detection results corresponding to the center point distances;

[0050] Compare the confidence scores of the detection results to output a detection result, and the detection result is the detection result with a higher confidence score.

[0051] In a second aspect, the present application provides a pantograph slider anomaly optimization detection system, including:

[0052] An acquisition module for acquiring an image to be processed;

[0053] A weight processing module for inputting the image to be processed into a gated network to output weights through the gated network;

[0054] A detection module for inputting the image to be processed into a hybrid model to output a first detection result through a crack detection model, a second detection result through a chunk loss detection model, and a third detection result through a foreign object detection model, where the first detection result, the second detection result, and the third detection result include detection results of cracks, chunk losses, and foreign objects;

[0055] A fusion processing module for fusing the first detection result, the second detection result, and the third detection result based on the weights to output a fusion detection result.

[0056] As can be seen from the above technical solutions, the present application provides a pantograph slider anomaly optimization detection method and system. The method includes: acquiring an image to be processed, inputting the image to be processed into a gated network to output weights through the gated network, and inputting the image to be processed into a hybrid model to output a first detection result through a crack detection model, a second detection result through a chunk loss detection model, and a third detection result through a foreign object detection model, where the first detection result, the second detection result, and the third detection result include detection results of cracks, chunk losses, and foreign objects; then fusing the first detection result, the second detection result, and the third detection result based on the weights to output a fusion detection result. The method sets different detection models for different types of anomalies and matches the detection models through a gated network, which can improve the accuracy and processing speed at the same time. Description of the Drawings

[0057] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0058] Figure 1 Schematic flow chart of the pantograph slider anomaly optimization detection method provided by the embodiment of the present application;

[0059] Figure 2 Schematic diagram of the preprocessing process provided by the embodiment of the present application;

[0060] Figure 3Schematic diagram of the DSC_C2f module provided by the embodiment of the present application;

[0061] Figure 4 Schematic diagram of the DSC_Bottleneck unit provided by the embodiment of the present application;

[0062] Figure 5 Schematic diagram of the CA_C2f module provided by the embodiment of the present application;

[0063] Figure 6 Schematic diagram of the CA_Bottleneck module provided by the embodiment of the present application. Detailed implementation manners

[0064] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all the implementation manners consistent with the present application. They are only examples of the systems and methods consistent with some aspects of the present application detailed in the claims.

[0065] The pantograph current collector slide is used for stable and efficient power transmission between the pantograph and the catenary to ensure the normal operation of the vehicle. It is a component for the train to obtain electric energy from the catenary, ensuring the continuity and stability of the train traction power supply system. The pantograph carbon slide is the current-carrying part used for contacting the catenary in the pantograph. It is made of materials with good electrical conductivity, such as carbon slide or copper-based composite materials. The carbon slide can enable the stable contact between the pantograph and the catenary, thereby effectively transmitting current and providing power for electric locomotives or trams.

[0066] When an electric locomotive or tram is running, the pantograph rises and contacts the catenary. At this time, the carbon slide is in close contact with the catenary, forming a current path. The current flows into the pantograph through the carbon slide and is then short-circuited through the current-carrying wire inside the pantograph and transmitted to the vehicle electrical system to provide power for the vehicle.

[0067] To ensure the safety of train operation, it is necessary to detect the pantograph carbon slide of the train. For example, through manual inspection. However, the pantograph carbon slide is located on the top of the train and moves continuously with the running of the train. Manual inspection requires frequent climbing up and down for detailed inspection, with a relatively high working intensity. For trains running over long distances and at high frequencies, it increases the frequency and time cost of manual inspection.

[0068] Manual inspection requires checking each skateboard one by one, which is time-consuming. During the peak train operation period, manual inspection may not be completed in time, resulting in potential safety hazards not being discovered in a timely manner. Moreover, the experience, skill level, work attitude, etc. of the inspection personnel will all affect the accuracy and reliability of the inspection results. For some subtle abnormalities, such as tiny cracks or chips, it may be difficult for the inspection personnel to detect them.

[0069] For another example, the image recognition technology based on a single detection model can only detect specific types of abnormalities, such as cracks or chips. For different types of abnormalities, the recognition accuracy of the model may vary, resulting in some abnormalities not being accurately recognized. Due to the limitations of image recognition technology, such as the influence of factors like lighting conditions and image quality, the recognition accuracy of the model may decrease in actual applications.

[0070] Furthermore, the image recognition technology based on a single detection model can only process specific types of images or scenarios. When the train operation environment changes, such as changes in weather conditions, lighting intensity, etc., the adaptability of the model may be affected. For some complex abnormal situations, such as foreign object detection, that is, when a foreign object enters the pantograph skateboard area, a single model may not be able to accurately identify and process it.

[0071] To solve the above problems, in some embodiments, the object detection method based on deep learning can be used for pantograph skateboard abnormality detection. For example, target detection can be performed through YOLO. However, small targets such as pantograph skateboards occupy fewer pixels in the image, resulting in low resolution and difficulty in capturing sufficient features. Small targets also lack sufficient semantic information, making classification and recognition difficult. Moreover, small targets are easily integrated with the background and difficult to distinguish. The scale of the pantograph skateboard varies greatly in different scenarios, increasing the complexity of detection and resulting in lower detection accuracy.

[0072] The complex background may contain a large amount of irrelevant information, such as noise, texture, shadow, etc. These factors will interfere with the accuracy of target detection. The target objects in the complex background may have diverse scale, shape, and appearance features. The machine vision system needs to have the ability to extract multi-scale features to ensure the comprehensive capture and accurate recognition of the target. The sizes and shapes of different models of pantographs are different, and the installation positions, working fields of view, and working distances of image sensors on different vehicles are different. This leads to an increase in the variation range of the size of the pantograph skateboard target in the image under different conditions, requiring the detection model to have strong generalization ability.

[0073] Although deep learning models such as YOLOv8 have made progress in the field of object detection, they still require a large amount of computing resources. In real-time applications, this can lead to a decrease in processing speed, thus failing to meet the real-time requirements. Moreover, to use deep learning models in real-time applications, the models need to be optimized and deployed, including model pruning, quantization, using GPU acceleration, etc. However, these optimization measures affect the detection accuracy and robustness of the models to a certain extent.

[0074] In summary, the accuracy of object detection methods based on deep learning in pantograph slider anomaly detection is relatively low.

[0075] To solve the problem of relatively low detection accuracy, some embodiments of this application provide an optimized detection method for pantograph slider anomalies. The method sets different detection models for different types of anomalies and matches the detection models through a gating network, which can improve the accuracy while enhancing the processing speed.

[0076] As Figure 1 shown, the method includes the following steps:

[0077] S100: Obtain the image to be processed.

[0078] The image to be processed is an image of the pantograph slider, which can be an image of the pantograph slider whose anomaly situation needs to be detected collected by a camera. The image can cover all possible abnormal areas. The image contains pixel data, and each pixel has three channel values of red (R), green (G), and blue (B), forming a color image in RGB format.

[0079] To improve the accuracy of pantograph slider image detection, before detecting the anomaly situation, preprocessing steps can be performed on the image, which helps to enhance the effect of feature extraction, reduce noise interference, and make the data input into the model have higher quality.

[0080] In some embodiments, first adjust the size of the image to be processed, and then sequentially perform preprocessing steps such as normalization, color space conversion, denoising, contrast enhancement, edge enhancement, geometric correction, segmentation, illumination normalization, data augmentation, etc.

[0081] Among them, adjusting the size is to ensure that the image to be processed has the same resolution, which is convenient for model processing. The image can be adjusted to be the same as the fixed size used during model training. For example, select a standard size suitable for the object detection task.

[0082] Normalization can make the pixel value distributions of images from different sources or under different conditions more unified, accelerate model convergence, and improve generalization ability. For example, map the image pixel values from the range of 0 - 255 to the range of 0 - 1.

[0083] Some models may perform better in specific color spaces, such as grayscale, HSV (Hue, Saturation, Value), etc. According to actual needs, RGB images can be converted to other color spaces. For example, grayscale images can reduce the computational load and highlight structural information, while HSV helps to separate luminance and color information.

[0084] Denoising is to remove random noise in the image to prevent it from interfering with feature extraction. For example, techniques such as Gaussian filtering, bilateral filtering, or non-local means filtering are used to smooth the image while trying to keep the edge information intact.

[0085] Contrast enhancement can improve the visual effect of the image, making subtle defects easier to identify. For example, methods such as histogram equalization and adaptive histogram equalization can be used to improve the image contrast.

[0086] Edge enhancement can strengthen the information of object boundaries in the image, especially for abnormal situations such as cracks or chippings. For example, operators such as Sobel and Canny are used for edge detection, and the intensity of the edge region is appropriately enhanced.

[0087] Geometric correction can correct perspective distortion caused by the shooting angle or equipment to ensure that the true proportions of objects in the image remain unchanged. Exemplarily, the image is corrected through affine transformation, perspective transformation, etc., to make it more consistent with the actual physical size.

[0088] Cropping and segmentation can focus on the region of interest, exclude unnecessary background information, and reduce the computational burden. For example, manually or automatically crop the part containing the pantograph slider, or apply a semantic segmentation algorithm to retain only the relevant regions.

[0089] Illumination normalization is used to solve the problem of brightness differences caused by light source changes to ensure the consistency of the image under different lighting conditions. For example, gamma correction, log transformation, or other illumination compensation algorithms are adopted to balance the overall brightness of the image.

[0090] For data augmentation, during the model training phase, data augmentation can be used to increase the diversity of the training set and improve the generalization ability and robustness of the model. Although mainly applied in the training phase, performing data augmentation during testing can also help the model better handle various actual situations. For example, operations such as slight rotation, flipping, and scaling.

[0091] It can be understood that in the subsequent processes, that is, in the processes of S200, S300, and S400, the image to be processed can be the preprocessed image to be processed or the image to be processed without preprocessing.

[0092] S200: Input the image to be processed into the gating network to output weights through the gating network.

[0093] The gating network is a lightweight network architecture of Mobile Net, which can reduce the amount of computation and the number of parameters, and improve the real-time processing ability. The gating network judges whether the image to be processed needs to be detected for cracks, spalls, and foreign objects according to the global features of the image, and assigns the tasks to the corresponding detection models. The gating network extracts the global features of the image through a convolutional neural network and outputs the activation probability, that is, the weight, of each task to dynamically allocate computing resources.

[0094] Mobile Net is a lightweight convolutional neural network architecture designed for mobile and embedded vision applications, which can significantly reduce the computational complexity and model size while maintaining high performance. Mobile Net is achieved by introducing depthwise separable convolutions.

[0095] In this embodiment, the gating network includes a feature extraction layer, a fully connected layer with activation functions, and a Softmax layer. Among them, the feature extraction layer includes a convolutional neural network and global pooling. The convolutional neural network is used to extract the global features of the input image and can represent the compact feature representation of the image content; global pooling can further compress the feature dimension and capture the overall information of the image, generating a fixed-length feature vector h(x), that is, the global feature vector.

[0096] Then, the extracted global feature vector is input into the fully connected layer for linear transformation to calculate the score s j (x) of each task. In this process, in some cases, activation functions such as ReLU can also be considered to be added to enhance the expressive ability.

[0097] To ensure that the sum of the weights of all tasks is 1, the gating network will convert the above scores into a weight vector g(x) in the form of a probability distribution through the softmax function to more intuitively understand the possibility of each task being selected. Then, according to the generated weight vector, the predicted outputs of each detection model are weighted and summed to obtain the final detection result.

[0098] During the training process of the gating network, the goal is to make its weight distribution g(x) as close as possible to the distribution of the true task label y. Therefore, the cross-entropy loss function is used as the optimization objective, as shown in the following formula:

[0099]

[0100] where N is the total number of training samples, y ij is the true label (using one-hot encoding) of the jth task corresponding to the ith image, and g(x i ) j is the jth value of the weight distribution generated by the gating network for the ith image.

[0101] In the separate training phase, the gating network adjusts the parameters W based on this loss function g and b g , so as to be able to accurately assign task weights and provide effective support for the cooperation of multiple detection models.

[0102] In some embodiments, the image to be processed is input into the convolutional neural network of the gating network to extract global features, and then through a linear transformation, the first score, the second score, and the third score are calculated based on the global features. Using an activation function, the first score is converted into the first weight, the second score is converted into the second weight, and the third score is converted into the third weight. Among them, the first score is the score of the crack detection task, the second score is the score of the chipping detection task, and the third score is the score of the foreign object detection task.

[0103] Taking the first score as an example, it is calculated by the following formula:

[0104] s j (x) = W g ·h(x) + b g ;

[0105] where x is the image to be processed, W g is the weight matrix, h(x) is the global feature vector output by the feature extraction layer, and b g is the bias term.

[0106] The score is converted into a weight vector g(x) through the soft max activation function, and it is ensured that the sum of the weights of all tasks is 1.

[0107] S300: Input the image to be processed into the hybrid model, so as to output the first detection result through the crack detection model, output the second detection result through the chipping detection model, and output the third detection result through the foreign object detection model.

[0108] Under the framework of the Mixture of Experts (MoE), multiple detection models can be included, and each model can perform better on specific tasks or data subsets. MoE dynamically selects and combines the outputs of the detection models through a gating network to improve the performance and efficiency of the overall model.

[0109] The hybrid model includes a crack detection model, a chunk detection model, and a foreign object detection model. The crack detection model is a neural network specifically optimized based on the YOLOv8 framework, which is used to efficiently detect and identify cracks on the pantograph slider. The crack detection model enhances the ability to capture slender or curved structural features by replacing the ordinary convolution operation in the C2f module with a Dynamic Snake Convolution (DSC) module.

[0110] For the first detection result and the second detection result, that is, in the detection of cracks and chunks, to improve the detection accuracy and efficiency, by removing the background and focusing on the area of the pantograph slider, that is, the carbon slider area, interference factors can be reduced, enabling the model to focus more on the features within the target area, thereby enhancing the detection effect.

[0111] As Figure 2 shown, in some embodiments, preprocessing is performed on the image to be processed to obtain a preprocessed image, and then the background of the preprocessed image is removed based on a segmentation algorithm to extract the carbon slider area image. Based on the carbon slider area image, the hybrid model outputs the first detection result and the second detection result.

[0112] Among them, the preprocessing process can be performed based on the preprocessing process described in S100, which will not be elaborated here.

[0113] The segmentation algorithm can adopt, for example, semantic segmentation, instance segmentation, or edge detection combined with morphological operations. Using a pre-trained segmentation model, such as U-Net, Mask R-CNN, etc., to segment the preprocessed image to generate a binary mask map, marking the carbon slider area. Then, the original image is cropped or masked using the generated mask map, retaining the carbon slider area to form a carbon slider area image. Then, the extracted carbon slider area image is respectively input into the crack detection model and the chunk detection model, and each detection model will output corresponding detection results, including bounding boxes, confidence scores, and class labels.

[0114] The specific detection process can be referred to the following content. It can be understood that for the first detection result and the second detection result, among them, the image to be processed can be the processed carbon slider area image.

[0115] By setting the image to be processed as the processed carbon slider area image, background interference can be reduced, detection accuracy can be improved, computing resources can be optimized, and robustness can be enhanced. Specifically, the complex background around the pantograph slider, such as the sky, power poles, other devices, etc., may introduce noise, affecting the model's ability to identify cracks and chunks. By removing the background through a segmentation algorithm, these interferences can be significantly reduced.

[0116] It can also process the details in the carbon skateboard area more intensively. Especially for small or irregularly shaped cracks and chips, it helps to improve the accuracy and sensitivity of detection. Only processing the carbon skateboard area can reduce unnecessary computational workload and speed up the detection. After removing the background, the model is more robust to factors such as illumination changes and shooting angles.

[0117] As Figure 3 、 Figure 4 shown, the crack detection expert model includes an input layer, an improved backbone network, a neck network, and a head network. The input layer is used to receive the image to be processed. The improved backbone network includes multiple stages of downsampling operations to generate multi-scale feature maps. And, the C2f module is replaced by the DSC_C2f module. The DSC_C2f module consists of a series of DSC_Bottleneck units, and each unit contains two dynamic snake-shaped convolutional kernels.

[0118] In the DSC_C2f module, first, it passes through a convolutional layer, which can adjust the size or number of channels of the feature map. The feature map processed by the convolutional layer is divided into multiple parts, and these parts will enter different DSC_Bottleneck modules for processing respectively. Each divided part is processed by the DSC_Bottleneck module. The DSC_Bottleneck module can better adapt to the geometric structure of the target by dynamically adjusting the shape of the convolutional kernel, and while retaining the original features through residual connections, it enhances important feature information.

[0119] The DSC_Bottleneck module can be repeated multiple times (n times) to further refine and enhance the feature extraction ability. All the feature map parts processed by DSC_Bottleneck are concatenated together to form a new feature map. The new feature map then passes through a convolutional layer to adjust the size or number of channels of the output feature map and finally outputs.

[0120] Among them, the C2f module is composed of a series of Bottleneck structures, and each Bottleneck contains two 3x3 convolutional layers. In this embodiment, the 3x3 convolutional kernel is replaced by a dynamic snake-shaped convolutional kernel. Each DSC_Bottleneck unit contains two dynamic snake-shaped convolutional kernels (DSC_Conv), allowing the convolutional kernel to dynamically adjust its geometric shape according to the crack shape.

[0121] The neck network is used for cross-scale information exchange and feature fusion, such as structures like PANet, etc., to help further enhance the feature expressiveness. The head network includes functions such as bounding box prediction, confidence score, and category classification, and is used for the final crack detection output.

[0122] In some embodiments, the image to be processed is input into the first convolutional layer of the crack detection model to adjust the number of channels through the first convolutional layer, and the image to be processed is split into a first sub-image and a second sub-image. The second sub-image is input into the serpentine convolutional module of the crack detection model to perform feature extraction through the serpentine convolutional module, and a feature extraction map is output. Then, the first sub-image and the feature extraction map are concatenated, and the number of channels of the concatenated image is adjusted to output a first detection result.

[0123] The first convolutional layer is used to adjust the number of channels. After receiving the image to be processed at the input layer, the number of channels is adjusted through the first convolutional layer to facilitate matching with the DSC_Bottleneck unit. Then, the image to be processed with the adjusted number of channels is segmented. Part of the feature map remains unchanged and is directly passed to the next layer as part of the residual connection, and the other part of the feature map enters multiple DSC_Bottleneck units for deep feature extraction.

[0124] By introducing a bias, the dynamic serpentine convolution kernel enables the convolution kernel to swing freely in the x-axis and y-axis directions, as shown in the following formula:

[0125]

[0126] where c represents the parameter controlling the movement of the convolution kernel along the x-axis, and Δx i represents the bias distance of the convolution kernel in the y-axis direction.

[0127] In this way, the DSC module can better capture the features of slender or curved structures.

[0128] Then, the feature maps from the direct transmission path and the feature extraction path are concatenated together. The concatenated feature map combines the original features and the enhanced deep features. Finally, a convolutional layer is used to adjust the concatenated feature map to the required number of output channels and input it into the neck network and the head network.

[0129] After being processed by the head network, the first detection result includes a first detection region and a confidence score for the first detection region. Among them, the first detection region can be a detection region related to cracks, spalls, and foreign objects, including a bounding box. The bounding box is one or more of the detected crack regions, spall regions, and foreign object regions. The bounding box can be one or more, and each bounding box is defined by four coordinate values, representing the position of the crack region in the image to be processed. For example, the coordinates of the upper left corner and the lower right corner.

[0130] For each bounding box, a confidence score is also marked. The confidence score can be a value between 0 and 1, where 1 represents absolute certainty. The confidence score reflects the confidence level of the detection model in the existence of abnormalities within the bounding box, which is beneficial for screening out the first detection results with high credibility, such as the detection results of cracks; for another example, the detection results of cracks and spalls.

[0131] If the detection model is designed for multi-class detection, a class label will be assigned to each detected object, such as being divided into cracks and non-cracks, or more fine-grained classification according to specific crack types, to further identify by clearly marking the class of the detected target.

[0132] The DSC module improves the ability to capture crack features of different scales and shapes, especially having better performance for slender or curved cracks. Due to the ability to dynamically adjust the convolutional kernel, the model can more accurately locate the position of cracks and reduce the false alarm rate. The improved structure makes the model more robust to factors such as illumination changes and shooting angles.

[0133] The spall detection model is specifically optimized based on the YOLOv8 framework, enhancing the detection ability for small targets, such as spalls. The CompleteIoU (CIoU) loss function is replaced by the Normalized Wasserstein Distance (NWD) loss function to improve the sensitivity to the position deviation of small targets and reduce the missed detection situation.

[0134] NWD is a similarity measurement method based on the Gaussian distribution. By modeling the bounding box as a Gaussian distribution and using the Wasserstein distance to measure the similarity between two distributions, it is suitable for small target detection. NWD does not rely on the overlap between bounding boxes and can provide an optimization signal even when the bounding boxes do not overlap at all.

[0135] The crack detection expert model includes an input layer, a backbone network, a neck network, and a head network. The input layer is used to receive the image to be processed. The backbone network includes multiple stages of downsampling operations to generate multi-scale feature maps, extracting rich feature information through a series of convolutional layers and residual blocks. The neck network is used for cross-scale information exchange and feature fusion, such as structures like PANet, which can further enhance the feature expressiveness, especially for small target detection.

[0136] The head network includes functions such as bounding box prediction, confidence scoring, and class classification, outputting the detection results of cracks, spalls, and foreign objects. In the head network, the CIoU loss function is replaced by the NWD (Normalized Wasserstein Distance) loss function to specifically optimize the small target detection performance.

[0137] The NWD loss function is based on the Wasserstein distance and measures the difference between two probability distributions. Different from IoU (Intersection over Union) or its variants (such as GIoU, DIoU, CIoU), NWD takes into account the position and shape differences of bounding boxes and is more sensitive to small objects.

[0138] NWD first converts the bounding box into a Gaussian distribution, where the center of the bounding box has the highest weight and the weight decreases from the center to the boundary. The center position (x, y) and size (w, h) of the target bounding box are converted into a two-dimensional Gaussian distribution. For two bounding boxes a and b, the Gaussian distributions corresponding to bounding box a are respectively The Gaussian distribution corresponding to bounding box b is The mean μ = (x, y), representing the center position of the bounding box, and σ = (w, h), representing the width and height of the bounding box. NWD is calculated as follows:

[0139]

[0140] where, W 2 represents the Wasserstein distance, C is a constant closely related to the dataset, used to adjust the influence of the distance metric on the model output, exp is the exponential function, and the Wasserstein distance is decayed through the exponential function exp, which helps to smooth the loss function and provides better gradient propagation during training. The Wasserstein distance is as follows:

[0141]

[0142] Here, μ a and μ b are the means of the two Gaussian distributions respectively, representing the center positions of the bounding boxes; σ a and σ b are the standard deviations, related to the size of the bounding box.

[0143] In this way, NWD can more accurately measure the position deviation of small objects and improve the sensitivity of the detection algorithm to small objects.

[0144] To make the loss value have better convergence, the Wasserstein distance is normalized using the dataset-related constant C, which can make the loss value more convergent. The calculation of NWD is as follows:

[0145]

[0146] Among them, C is used to balance the weight of the Wasserstein distance in model optimization, ensuring the applicability of the loss function to the small target detection task.

[0147] In some embodiments, the image to be processed is input into the backbone network of the block dropout detection model to generate multi-scale feature maps, and then the feature maps are input into the neck module to output the fused feature maps. Finally, the fused feature maps are input into the head network to output the second detection result. Among them, the loss function of the head network is the NWD loss function.

[0148] After being processed by the head network, the second detection result includes the second detection region and the confidence score of the second detection region. The second detection region is the detection region related to cracks, block dropouts, and foreign objects, including the bounding box. The bounding box is the detected region of cracks, block dropouts, and foreign objects. The bounding box, confidence score, and category are the same as those of the first detection result and will not be elaborated here.

[0149] Integrate the NWD loss into the training process of YOLOv8 to replace the CIoU loss. By calculating the NWD loss between the positive samples of the model and the ground truth labels and incorporating it into the total loss of object detection, the bounding box regression ability of the model is optimized. Through this improvement, the block dropout detection model can more accurately measure the position deviation of small targets and improve the sensitivity of the detection algorithm to small targets.

[0150] The foreign object detection model is based on the YOLOv8 framework and is optimized, especially focusing on detecting foreign objects in complex backgrounds. The foreign object detection model enhances the ability to identify foreign object features by introducing CA (Coordinate Attention, an attention mechanism), enabling the model to better capture the position and shape features of foreign objects.

[0151] In the YOLOv8 framework, a C2f module is set in the backbone network. The C2f module is used to enhance feature extraction and reduce computational complexity. The C2f module enables more effective fusion of features at different scales by introducing cross-stage partial connections, thereby improving the performance of the model.

[0152] Exemplarily, the C2f module is located in the middle layer of the backbone network. These layers are used to generate multi-scale feature maps. Specifically, after the downsampling operation, the C2f module processes the feature maps from the previous layer and outputs richer feature representations for subsequent layers to use.

[0153] Such as Figure 5As shown, in some embodiments, the foreign object detection model includes an input layer, a backbone network, a neck network, and a head network. The input layer is used to receive the image to be processed. The backbone network includes multiple stages of downsampling operations to generate multi-scale feature maps. Moreover, based on the C2f module, a Channel Attention (CA) mechanism (CA_Bottleneck), namely CA_C2f, is introduced. Through steps such as global average pooling, feature map merging, capturing dimensional relationships, separable transposition, adjusting the number of channels, and applying attention scores, the model's understanding of the spatial structure of the input data is strengthened.

[0154] As Figure 6 shown, CA_Bottleneck includes two convolutional layers. After passing through the first convolutional layer and the second convolutional layer, it enters the CA module for channel attention calculation. For example, a global average pooling operation is performed, that is, a global average pooling operation is performed on each channel of the feature map to obtain a vector, where each element corresponds to the global average value of a channel. One or more fully connected layers are used to learn the dependencies between channels. If two fully connected layers are used, the first fully connected layer maps the vector to a smaller dimension, and the second fully connected layer then maps it back to the original number of channels. The weights of each channel are obtained through the Sigmoid function. These weights are between 0 and 1. The obtained weights are multiplied by the original feature map to obtain a weighted feature map. The weighted feature map and the original input feature map are added together, and the processed feature map is used as the output of the CA_Bottleneck module.

[0155] The CA attention mechanism adaptively adjusts the weights of different channels through steps such as global average pooling, fully connected layers, and Sigmoid activation, thereby enhancing the model's ability to capture important features. Through the above steps, it is possible to better focus on key information and improve the overall performance.

[0156] After the image to be processed is input to the input layer and passes through the C2f module of the backbone network, in some embodiments, the foreign object detection model performs a first average pooling on the image to be processed and obtains a first pooled image with a width of 1. Then, a second average pooling is performed on the image to be processed to obtain a second pooled image with a height of 1. It can be understood that the image to be processed is the image output after passing through the C2f module of the backbone network.

[0157] The first pooled image and the second pooled image are merged to generate a feature layer with a height of 1 and a width equal to the sum of the height of the first pooled image and the width of the second pooled image. Based on the feature layer, a third detection result is output.

[0158] Exemplarily, if the size of the image output after the C2f module of the backbone network is [C, H, W], first perform global average pooling on the image in the width direction to obtain a feature map [C, H, 1], which is the first pooling map; then perform pooling on it in the height direction to obtain a feature map [C, 1, W], which is the second pooling map.

[0159] Merge the first pooling map and the second pooling map into a feature layer with a shape of [C, 1, H + W], and through 1x1 convolution and an activation function, such as ReLU processing, to obtain a preliminary feature representation. Perform a convolution operation on the merged feature layer to capture the relationship between the width and height dimensions, and apply normalization and activation functions to further process the features.

[0160] For the feature layer, in some embodiments, separate the feature layer to output a first sub-feature layer and a second sub-feature layer. The height of the first sub-feature layer is 1, and the width is the height of the image to be processed. The width of the second sub-feature layer is 1, and the height is the width of the image to be processed;

[0161] Then transpose the first sub-feature layer to output a third sub-feature layer, where the height of the third sub-feature layer is the height of the image to be processed and the width is 1; and transpose the second sub-feature layer to output a fourth sub-feature layer, where the height of the fourth sub-feature layer is 1 and the width is the width of the image to be processed;

[0162] Perform channel adjustment on the third sub-feature layer through a second convolutional layer to obtain a fifth sub-feature layer, and perform feature adjustment on the fourth sub-feature layer through a third convolutional layer to obtain a sixth sub-feature layer;

[0163] Apply an activation function to the fifth sub-feature layer to output a first attention score map, and apply an activation function to the sixth sub-feature layer to output a second attention score map;

[0164] Multiply the first attention score map with the image to be processed to output a first modulated feature map, and multiply the second attention score map with the image to be processed to output a second modulated feature map;

[0165] Output a third detection result through the first modulated feature map and the second modulated feature map.

[0166] Exemplarily, separate the features in the width and height directions from the feature layer. The width direction is [C, 1, H], and the height direction is [C, 1, W]. Then perform a transpose operation on the two separated feature layers to restore the width and height dimensions, obtaining two feature layers as [C, H, 1], which is the third sub-feature layer, and [C, W, 1], which is the fifth sub-feature layer.

[0167] Apply 1×1 convolutions to the feature maps of [C, H, 1] and [C, W, 1] respectively, namely the second convolutional layer, to adjust the number of channels to adapt to attention calculation. Then apply the Sigmoid activation function to obtain the attention scores in the width and height dimensions.

[0168] Multiply the original input feature map, that is, the image output after passing through the C2f module of the backbone network, by the calculated attention scores to obtain the output feature maps that enhance the spatial features, namely the first modulated feature map and the second modulated feature map. Then input the first modulated feature map and the second modulated feature map into the neck network and the head network, and output the third detection result.

[0169] The neck network is used for cross-scale information exchange and feature fusion, such as structures like PANet, to help further enhance the feature expressiveness. The head network includes bounding box prediction, confidence scoring, and class classification, etc., and outputs the final crack, chip, and foreign object detection results, the third detection result.

[0170] After being processed by the head network, the third detection result includes the third detection region and the confidence score of the third detection region. The third detection region is the detection region related to cracks, chips, and foreign objects, including the bounding box, and the bounding box is the detected region of cracks, chips, and foreign objects. The bounding box, confidence score, and class are the same as those of the first detection result and will not be elaborated here.

[0171] The CA attention mechanism can enhance the understanding of the spatial structure of the input data by the foreign object detection model by introducing coordinate information, enabling the model to more accurately capture the position and shape features of foreign objects. Especially in the foreign object detection task with a complex background, the CA attention mechanism helps to highlight the target region and reduce background interference, thereby improving the detection accuracy. The improved structure makes the foreign object detection model more robust to factors such as illumination changes and shooting angles.

[0172] It can be understood that for the first detection result, the second detection result, and the third detection result, the detection results of three types of defects, namely cracks, chips, and foreign objects, will be output. However, optimization is carried out on the defect detection that the model is good at, so that the detection accuracy of this specific defect is higher.

[0173] S400: Based on weights, fuse the first detection result, the second detection result, and the third detection result, and output the fused detection result.

[0174] For different detection tasks, the gating network assigns weights to each detection model, and performs weighted averaging on the prediction results of the detection models according to these weights. For example, when there are more crack defects in the image to be detected, the gating network will assign more weights to the crack detection model, which means that the detection result of the crack detection model will contribute more to the final fused detection result.

[0175] In some embodiments, a weighted sum is performed on the first detection result and the first weight to output a first fusion result; a weighted sum is performed on the second detection result and the second weight to output a second fusion result; a weighted sum is performed on the third detection result and the third weight to output a third fusion result, and then, based on the first fusion result, the second fusion result, and the third fusion result, a fusion detection result is output, where the fusion detection result is the result obtained by performing element-wise addition on the first fusion result, the second fusion result, and the third fusion result.

[0176] Refer to the following formula to perform a weighted average on the prediction results of the detection model to calculate the final detection result y:

[0177]

[0178] where E j (x) is the predicted output of the j-th detection model for the input x, and g(x) j is the weight assigned by the gating network to the j-th detection model. In this way, the gating network can dynamically adjust the contribution degrees of the detection models according to the characteristics of the input image, thereby improving the detection efficiency and accuracy.

[0179] For example, if the image to be processed is a case of crack detection, then the gating network will assign a higher weight to the detection model that is good at crack detection. For another example, for an image containing complex background noise, the gating network can increase the weight of the detection model that performs better in anti-noise performance.

[0180] In this way, not only can the computing resources be efficiently utilized, but also the overall detection accuracy and robustness can be improved. Since each detection model is optimized for a specific type of task, when they are reasonably combined, they can cover a wider range of scenarios and provide more reliable detection results.

[0181] As can be seen from the above, the first detection result includes the detected crack area and the confidence score of the crack area, the second detection result includes the spalled area and the confidence score of the spalled area, and the third detection result includes the foreign object area and the confidence score of the foreign object area.

[0182] In multi-task detection, multiple detection models may simultaneously detect the same target or different types of anomalies. To improve the accuracy of the output detection results, a result fusion strategy can be adopted, and the result fusion strategy includes spatial position fusion and type fusion.

[0183] In some embodiments, the central point distances are calculated, where the central point distances include a first central point distance, a second central point distance, and a third central point distance. The first central point distance is the Euclidean distance of the central point of the crack region, the second central point distance is the Euclidean distance of the central point of the chip removal region, and the third central point distance is the Euclidean distance of the central point of the foreign object region. Then, a distance threshold is set. If at least two central point distances are less than the distance threshold, the detection results corresponding to the central point distances are merged. The confidence scores of the detection results are compared to output the detection result, and the detection result is the detection result with a high confidence score.

[0184] For spatial location fusion, when multiple detection models detect the same region, multiple bounding boxes, i.e., different regions, may be generated. These bounding boxes may correspond to the same actual target. To merge these redundant detection boxes, a spatial location fusion strategy based on the distance between the central points of the detection boxes is adopted.

[0185] Exemplarily, for each bounding box, the central point coordinates are calculated, and then a distance threshold is preset. The distance threshold can be adjusted according to the specific application scenario and the target size. If the Euclidean distance between the central points of two bounding boxes is less than the distance threshold, it indicates the same target. For example, if the Euclidean distance between the central points of the A crack region and the A chip removal region is less than the distance threshold, it indicates that the A crack region and the A chip removal region are the same target. At this time, the position information of these bounding boxes, such as the average value of the central point coordinates, can be used to determine the position of the merged detection box. Non-maximum suppression technology can also be used to select the bounding box with the highest confidence as the representative and suppress other overlapping bounding boxes.

[0186] In this way, the number of redundant detection boxes can be effectively reduced, so that each actual target has only one corresponding detection result, thereby improving the clarity and accuracy of the detection.

[0187] For type fusion, in some cases, different detection models may detect different types of targets or anomalies in the same region. At this time, the type with a high confidence after weighted fusion is used as the final detection result.

[0188] Exemplarily, the image to be processed is input into the gating network, and the output weight vector is (0.7, 0.2, 0.1), corresponding to the weights of the crack, chip removal, and foreign object experts respectively. The image to be processed is input into three detection networks respectively to obtain the first detection result, the second detection result, and the third detection result. Among them, these three detection results all include the bounding box and the confidence scores of the three types of defects. After determination, that is, the intersection over union is greater than the threshold, the bounding box B1 in the same region in the image to be processed. The confidence levels output by the three detection models for this bounding box are as follows in the table:

[0189] Chunk loss Crack Foreign object Crack expert (0.7) 0.1 0.8 0.1 Chunk loss expert (0.2) 0.3 0.4 0.3 Foreign object expert (0.1) 0.2 0.5 0.3

[0190] Perform weighted summation to obtain the final confidence (0.15, 0.69, 0.16). The calculation process is shown in the following table:

[0191]

[0192] According to the set type threshold, for example, 0.4, only the final results greater than 0.4 can be retained. Since 0.69 > 0.4, the type result of this detection box is crack. If there are two detection results both greater than 0.4, select the larger type result as the final detection result.

[0193] If the crack expert detects a bounding box B2, and the other two detection models do not detect this bounding box or the intersection over union ratio of position fusion is less than the threshold, then the confidence outputs of each type of abnormality for this bounding box by the other two detection models can be regarded as 0. See the following table:

[0194] Chunk loss Crack Foreign object Crack expert (0.7) 0.1 0.8 0.1 Chunk loss expert (0.2) 0 0 0 Foreign object expert (0.1) 0 0 0

[0195] The final confidence result is calculated by the following formula:

[0196] 0.1 * 0.7 = 0.07, 0.8 * 0.7 = 0.56, 0.1 * 0.7 = 0.07;

[0197] Among them, since 0.56 > 0.4 (type threshold), the result is retained, and the detection type of the bounding box B2 is crack. If there is no result greater than the type threshold in the confidence results, then this bounding box is not retained.

[0198] Through the process of type fusion, the best choice can be made among multiple detection models to ensure the accuracy and reliability of the final detection result.

[0199] Based on the above abnormal optimization detection method for the pantograph slider, some embodiments of the present application further provide an abnormal optimization detection system for the pantograph slider, including:

[0200] An acquisition module, configured to acquire an image to be processed;

[0201] A weight processing module, configured to input the image to be processed into a gated network to output weights through the gated network;

[0202] A detection module, configured to input the image to be processed into a hybrid model to output a first detection result through a crack detection model, output a second detection result through a chunk missing detection model, and output a third detection result through a foreign object detection model. The first detection result, the second detection result, and the third detection result include the detection results of cracks, chunk missing, and foreign objects;

[0203] A fusion processing module, configured to fuse the first detection result, the second detection result, and the third detection result based on the weight, and output a fusion detection result.

[0204] For the effects during the operation of the above system embodiment, reference may be made to the effects of the above method embodiment, which will not be elaborated here.

[0205] This application provides a method and a system for abnormal optimization detection of a pantograph slide plate. The method includes: obtaining an image to be processed, inputting the image to be processed into a gated network to output a weight through the gated network, and inputting the image to be processed into a hybrid model to output a first detection result through a crack detection model, output a second detection result through a chunk loss detection model, and output a third detection result through a foreign object detection model, where the first detection result, the second detection result, and the third detection result include detection results of cracks, chunk losses, and foreign objects; and then fusing the first detection result, the second detection result, and the third detection result based on the weight to output a fusion detection result. The method sets different detection models for different types of abnormalities and matches the detection models through a gated network, which can improve the accuracy and the processing speed at the same time.

[0206] For the similar parts between the embodiments provided in this application, reference may be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other implementation manner extended based on the solution of this application without creative efforts belongs to the protection scope of this application.

Claims

1. A pantograph slide abnormality optimization detection method, characterized in that: include: Get the image to be processed; Inputting the image to be processed into a gating network to output weights through the gating network; Inputting the image to be processed into the hybrid model to output a first detection result through a crack detection model, a second detection result through a chipping detection model, and a third detection result through a foreign matter detection model, wherein the first detection result, the second detection result, and the third detection result include detection results of cracks, chipping, and foreign matter; Based on the weight, the first detection result, the second detection result and the third detection result are fused, and a fused detection result is output.

2. The pantograph slide abnormality optimization detection method according to claim 1 is characterized in that: The step of inputting the image to be processed into a gating network to output weights through the gating network comprises: Inputting the image to be processed into the convolutional neural network of the gating network to extract global features; By linear change, a first score, a second score and a third score are calculated based on the global feature, wherein the first score is the score of the crack detection task, the second score is the score of the block drop detection task, and the third score is the score of the foreign body detection task; Using an activation function, the first score is converted into a first weight, the second score is converted into a second weight, and the third score is converted into a third weight.

3. The pantograph slide abnormality optimization detection method according to claim 2 is characterized in that: The step of fusing the first detection result, the second detection result, and the third detection result based on the weight, and outputting the fused detection result includes: Performing weighted summation on the first detection result and the first weight, and outputting a first fusion result; Performing weighted summation on the second detection result and the second weight, and outputting a second fusion result; Performing weighted summation on the third detection result and the third weight, and outputting a third fusion result; Based on the first fusion result, the second fusion result and the third fusion result, a fusion detection result is output, where the fusion detection result is a result obtained by performing element-by-element addition of the first fusion result, the second fusion result and the third fusion result.

4. The pantograph slide abnormality optimization detection method according to claim 1, characterized in that: The step of inputting the image to be processed into the hybrid model to output a first detection result through a crack detection model includes: Inputting the image to be processed into a first convolutional layer of the crack detection model, so as to adjust the number of channels through the first convolutional layer, and splitting the image to be processed into a first sub-image and a second sub-image; Inputting the second sub-image into the serpentine convolution module of the crack detection model to perform feature extraction through the serpentine convolution module and output a feature extraction map; The first sub-image and the feature extraction image are spliced ​​together, and the number of channels is adjusted on the spliced ​​image to output a first detection result.

5. The pantograph slide abnormality optimization detection method according to claim 1, characterized in that: The outputting a second detection result through the block drop detection model includes: Inputting the image to be processed into the backbone network of the block loss detection model to generate a multi-scale feature map; Inputting the feature map into the neck network to output a fused feature map; The fused feature map is input into a head network to output a second detection result, wherein the head network includes an NWD loss function.

6. The pantograph slide abnormality optimization detection method according to claim 1, characterized in that: The outputting a third detection result through the foreign body detection model includes: Inputting the image to be processed into a foreign body detection model, so as to perform a first average pooling on the image to be processed through the foreign body detection model, and obtaining a first pooling map, wherein the width of the first pooling map is 1; Performing a second average pooling on the image to be processed to obtain a second pooling map, where the height of the second pooling map is 1; Merging the first pooling map and the second pooling map to generate a feature layer, wherein the height of the feature layer is 1, and the width of the feature layer is the sum of the height of the first pooling map and the width of the second pooling map; Based on the feature layer, a third detection result is output.

7. The pantograph slide abnormality optimization detection method according to claim 6, characterized in that: The outputting a third detection result based on the feature layer includes: Separating the feature layer to output a first sub-feature layer and a second sub-feature layer, wherein the first sub-feature layer has a height of 1 and a width of the image to be processed, and the second sub-feature layer has a width of 1 and a height of the image to be processed; Transposing the first sub-feature layer to output a third sub-feature layer, wherein the height of the third sub-feature layer is the height of the image to be processed, and the width is 1; Transposing the second sub-feature layer to output a fourth sub-feature layer, wherein the height of the fourth sub-feature layer is 1 and the width is the width of the image to be processed; Performing channel adjustment on the third sub-feature layer through the second convolution layer to obtain a fifth sub-feature layer, and performing feature adjustment on the fourth sub-feature layer through the third convolution layer to obtain a sixth sub-feature layer; Applying an activation function on the fifth sub-feature layer to output a first attention score map, and applying an activation function on the sixth sub-feature layer to output a second attention score map; Multiplying the first attention score map with the image to be processed to output a first modulation feature map, and multiplying the second attention score map with the image to be processed to output a second modulation feature map; A third detection result is outputted through the first modulation characteristic graph and the second modulation characteristic graph.

8. The pantograph slide abnormality optimization detection method according to claim 1, characterized in that: The step of inputting the image to be processed into the hybrid model to output a first detection result through a crack detection model and outputting a second detection result through a chip drop detection model comprises: Performing preprocessing on the image to be processed to obtain a preprocessed image; Removing the background of the preprocessed image based on a segmentation algorithm to extract a carbon slide plate area image; Based on the carbon slide plate area image, a first detection result and a second detection result are outputted through the hybrid model.

9. The pantograph slide abnormality optimization detection method according to claim 1, characterized in that: The first detection result includes a first detection area and a confidence score of the first detection area, the second detection result includes a second detection area and a confidence score of the second detection area, and the third detection result includes a third detection area and a confidence score of the third detection area; The step of fusing the first detection result, the second detection result, and the third detection result based on the weight, and outputting the fused detection result includes: Calculate center point distances, where the center point distances include a first center point distance, a second center point distance, and a third center point distance, where the first center point distance is the Euclidean distance of the center points of the first detection area, the second center point distance is the Euclidean distance of the center points of the second detection area, and the third center point distance is the Euclidean distance of the center points of the third detection area; Set distance threshold; If the distance between at least two center points is less than the distance threshold, merging the detection results corresponding to the center point distances; The confidence scores of the detection results are compared to output a detection result, wherein the detection result is a detection result with a high confidence score.

10. A pantograph slide abnormality optimization detection system, characterized in that: include: An acquisition module, used for acquiring an image to be processed; A weight processing module, used for inputting the image to be processed into a gating network to output weights through the gating network; A detection module, used for inputting the image to be processed into a hybrid model, so as to output a first detection result through a crack detection model, a second detection result through a chipping detection model, and a third detection result through a foreign matter detection model, wherein the first detection result, the second detection result and the third detection result include detection results of cracks, chipping and foreign matter; A fusion processing module is used to fuse the first detection result, the second detection result and the third detection result based on the weight, and output a fusion detection result.