Pantograph image-based anomaly detection method and system

By calculating the occlusion area and illumination uniformity of the pantograph image and identifying its components, the problem of inaccurate detection of pantograph abnormalities in the prior art is solved, and more accurate and reliable detection results are achieved.

CN120182280AActive Publication Date: 2025-06-20CHENGDU YUNDA TECH CO LTD +1

Patent Information

Application Number
CN202510664472.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect abnormal situations of pantographs when the locomotive is running outdoors, resulting in inaccurate and reliable detection results.

Method used

By acquiring the image of the pantograph area, calculating the occlusion area and illumination uniformity of the image, sifting out images that do not meet the preset conditions, and identifying the components of the pantograph in the remaining images to determine whether there is any abnormality.

Benefits of technology

A more accurate and reliable pantograph abnormal detection results are achieved, which can effectively avoid misjudgment caused by factors such as strong light reflection and rain and snow occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182280A_ABST
    Figure CN120182280A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, and discloses an anomaly detection method and system based on a pantograph image, and the method comprises the steps: obtaining a series of first images of a pantograph region, calculating the shielding area and illumination uniformity of the first images, and obtaining the first images of the pantograph region; if any one of the shielding area or the illumination uniformity does not meet a preset condition, screening out the corresponding first image; and identifying each part of the pantograph in the first image, judging whether each part is abnormal or not according to a preset condition, and outputting a first judgment result. Therefore, due to the fact that the part, not meeting the requirement, of the first images in the series of first images is screened out, all the components are recognized, and a more accurate and reliable pantograph anomaly detection result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and particularly to an abnormal detection method and system based on a pantograph image. Background Art

[0002] With the development of image recognition technology, it is possible to install high-definition cameras along the train operation route or on the train to collect image data of the pantograph in real time. Then, image recognition algorithms are used to analyze and process these images, extract the characteristic information of the pantograph, and compare it with the preset normal state characteristics to determine whether there is an abnormality in the pantograph.

[0003] For example, in the patent with the publication number CN119251755A, a pantograph image is obtained based on an object recognition algorithm, and then the recognized image is processed through a specific algorithm, and data enhancement processing is performed on the image based on a GAN network to obtain a pantograph abnormality detection result.

[0004] However, when the locomotive runs outdoors, due to its complex operation conditions and high image complexity, it is often difficult to distinguish pantograph abnormalities; specifically, the existing pantograph image detection means only use different methods to enhance image processing capabilities, and they cannot adapt to the complex conditions of the above-mentioned pantograph, such as the technical solutions for realizing pantograph image detection by means of enhancing processing based on images; therefore, various pantograph image detection means in the prior art cannot obtain accurate and reliable pantograph detection results. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art, provide an abnormal detection method based on a pantograph image, and propose a new technical route for realizing abnormal detection of a pantograph image, which can adapt to the complex conditions when the locomotive runs outdoors and obtain more accurate and reliable detection results.

[0006] The purpose of the present invention is achieved through the following technical solutions: In a first aspect, an abnormal detection method based on a pantograph image, a series of first images of the pantograph area are obtained, the occlusion area and the illumination uniformity of the first images are calculated, and if either the occlusion area or the illumination uniformity does not meet the preset conditions, the corresponding first image is screened out; each component of the pantograph is recognized in the first image, and whether each component has an abnormality is determined according to the preset conditions, and a first determination result is output.

[0007] The beneficial effects are as follows: Since a series of first images that do not meet the requirements in the first images are screened out, and then each component is identified, a more accurate and reliable pantograph anomaly detection result can be obtained. For example, by calculating the illumination uniformity, an image with severely uneven illumination distribution caused by strong light reflection can be accurately identified. If the illumination uniformity of the first image does not meet the preset condition, it is promptly screened out; by calculating the occlusion area, the occlusion degree of the image by rain, snow, and the gantry can be quantified. When the occlusion area of the first image exceeds the preset threshold, it indicates that the occlusion by rain, snow, and / or the gantry has affected the usability of the image, and it is then screened out.

[0008] Preferably, identifying each component of the pantograph in the first image includes: carbon slide plate identification; inputting the first image into a semantic segmentation model to generate a mask image of the first image, and the mask image divides the carbon slide plate area and the non-carbon slide plate area to identify the carbon slide plate; wherein, the semantic segmentation model includes: an encoder and a decoder, and the decoder adopts a nested skip connection structure and integrates a CBAM attention module.

[0009] Preferably, the first image with the attached mask image is input into a first large model, and the first large model generates a second determination result.

[0010] Preferably, determining whether each component has an anomaly according to preset conditions and outputting the first determination result includes: determining whether there is a component loss in the first image according to the preset topological dependency relationship between each component; identifying the dynamic brightness threshold in the area around the catenary in the first image and using the optical flow tracking technology to determine whether there is a spark phenomenon; calculating whether the proportion of the fracture length of the carbon slide plate in the first image exceeds the threshold, calculating whether the inclination angle of the carbon slide plate in the first image exceeds the threshold, and if the proportion of the fracture length or the inclination angle exceeds the threshold, it is determined that the carbon slide plate is abnormal; calculating whether the relative position between the pantograph and the catenary in the first image exceeds the threshold, and if it exceeds the threshold, it is determined to be abnormal.

[0011] Preferably, the topological dependency relationship is the topological dependency relationship among the base, the bow angle, and the carbon slide plate: if the base is not detected, it is determined that the whole pantograph is lost; if the base is detected but the bow angle is not detected, it is determined that the bow angle is lost; if the bow angle is detected but the carbon slide plate is not detected, it is determined that the carbon slide plate has fallen off.

[0012] Preferably, select the first images determined as spark phenomena and the first images that do not meet the illumination uniformity in the first determination results, and mark them as controversial images; each of the controversial images is attached with a component bounding box obtained through object recognition, and each of the controversial images is also attached with a confidence level; input the controversial images with the bounding boxes and the confidence level into a second large model, and the second large model outputs a classification probability as the third determination result.

[0013] Preferably, obtain a semantic consistency score according to the detection category of the first determination result and the semantic category of the second large model, and assign a classification consistency weight; assign a confidence weight to the confidence level and a probability judgment weight to the classification probability; input the confidence level, confidence weight, classification probability, probability judgment weight, classification consistency weight, and semantic consistency score into a fusion decision algorithm to obtain a comprehensive score; determine whether the comprehensive score exceeds a threshold and output a fourth determination result.

[0014] Preferably, calculating the occlusion area and illumination uniformity of the first image, and screening out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions includes: segmenting the occluder based on the Mask R-CNN model, calculating whether the proportion of the occlusion area of the occluder in the first image exceeds a threshold, and screening out if it exceeds; calculating whether the variance of the brightness component analyzed in the HSV space exceeds a threshold, and screening out if it exceeds.

[0015] Preferably, calculating the occlusion area and illumination uniformity of the first image, and screening out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions further includes: determining whether the clarity of the first image exceeds a threshold.

[0016] In a second aspect, an abnormal detection system based on pantograph images includes a storage and a processor, where the storage stores a computer-readable medium for the above abnormal detection method based on pantograph images, and the processor is configured to execute the computer-readable medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic flowchart of an abnormal detection method based on pantograph images according to an embodiment of the present application; Figure 2 It is a schematic flowchart of component loss detection according to an embodiment of the present application; Figure 3 It is a schematic diagram of semantic segmentation model construction according to an embodiment of the present application; Figure 4 It is a schematic diagram of the second model architecture according to an embodiment of the present application; Figure 5Schematic flowchart of another abnormal detection method based on pantograph images according to an embodiment of the present application; Figure 6 Schematic flowchart of yet another abnormal detection method based on pantograph images according to an embodiment of the present application. Detailed implementation manners

[0018] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] When the locomotive runs outdoors, the pantograph image is easily affected by strong light reflection, rain and snow occlusion, gantry shadow and other interferences, resulting in a high false detection rate of traditional object detection and segmentation models; even if enhanced technologies of image processing are applied, such as data enhancement processing of images based on the GAN network, it is only a means to enhance the image and cannot solve the problem that it is difficult to obtain accurate and reliable results for pantograph images.

[0020] Refer to Figures 1 - 4 , the present invention provides a technical solution: An abnormal detection method based on pantograph images, which collects images of the pantograph area in real time through a high-definition camera, and hereinafter is defined as the first image for the convenience of description; calculates the occlusion area and illumination uniformity of the first image, and screens out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions.

[0021] Specifically, in the process of collecting and analyzing pantograph images, factors such as strong light reflection, rain and snow occlusion, and gantry shadow will seriously affect the image quality, thereby interfering with the accuracy and reliability of subsequent abnormal detection.

[0022] Strong light reflection is that when sunlight or other strong light sources directly irradiate the surface of the pantograph, strong reflected light will be generated, causing local or large-area over-bright areas in the pantograph area of the first image collected by the high-definition camera. These over-bright areas will cause abnormal pixel values of the image and lose the detailed features of the pantograph itself. For example, the edges of the key components of the pantograph may become blurred due to strong light reflection, or even completely covered by the over-bright areas.

[0023] In this embodiment, by calculating the illumination uniformity, an image with severely uneven illumination distribution caused by strong light reflection can be accurately identified. If the illumination uniformity of the first image does not meet the preset conditions and is screened out in time, this kind of image containing a large amount of invalid information can be prevented from being used for subsequent abnormal detection, thereby preventing misjudgment caused by strong light reflection.

[0024] Rain and snow occlusion can affect image quality. Under harsh weather conditions, raindrops or snowflakes falling on the pantograph surface can partially or completely obscure key parts of the pantograph, resulting in incomplete characteristic information of the pantograph in the first captured image. For example, rain and snow may cover the carbon slide of the pantograph, making it impossible to accurately determine the wear degree of the carbon slide or whether there are abnormalities such as cracks. Similarly, in the railway operating environment, the gantry will cast shadows or directly obscure the pantograph at different times.

[0025] In this embodiment, by calculating the occlusion area, the occlusion degree of rain, snow, and the gantry on the image can be quantified. When the occlusion area of the first image exceeds the preset threshold, it indicates that the occlusion of rain, snow, and / or the gantry has affected the usability of the image. At this time, screening it out can effectively avoid abnormal detection errors caused by rain and snow occlusion.

[0026] Furthermore, each component of the pantograph is identified in the first image, and it is determined whether each component is abnormal according to the preset conditions, and the first determination result is output. In this way, since a series of first images that do not meet the requirements in the first images are screened out and then each component is identified, a more accurate and reliable pantograph abnormal detection result can be obtained.

[0027] Next, with reference to Figure 1 , a specific embodiment will be described.

[0028] An abnormal detection method based on pantograph images includes the following steps: Obtain a series of first images of the pantograph area, calculate the occlusion area and illumination uniformity of the first image, and screen out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions. Specifically, it includes: S100. Screen the obtained first images to filter out images that do not meet the preset conditions. Among them, the preset conditions include: S110. Clarity evaluation can adopt an improved Brenner gradient algorithm (Brenner gradient algorithm), and the formula of the algorithm is: , and the threshold is set to S≥120. In the formula, S is the clarity score, which is the total image gradient energy calculated by the improved Brenner gradient algorithm and reflects the sharpness of the image edges. x is the column coordinate (horizontal direction) of the image pixel. y is the row coordinate (vertical direction) of the image pixel. I(x, y) is the pixel intensity value of the image at the coordinate (x, y) . In a grayscale image, it represents brightness (0 - 255). For an RGB image, it needs to be converted to grayscale first.

[0029] S120. Occlusion detection: Based on the Mask R-CNN model, segment obstacles such as gantries, rain, and snow, calculate the proportion of the occlusion area Q1 in the total image area Q, and set the threshold as the occlusion area Q1 ≤ 15% of Q.

[0030] S130. Illumination uniformity: Calculate the variance S of the luminance component (V channel) through HSV space analysis 2 , and set the threshold as S 2 ≤ 500. If the value of S 2 is greater than the threshold, it indicates that the image is overexposed.

[0031] If any of the above indicators exceeds the limit, discard the current frame and trigger image re-acquisition; otherwise, enter the target detection module. Of course, the above S110 - S130 can be performed sequentially or synchronously. In this way, it can be ensured that the gradient is calculated only for the pantograph area (ROI) to exclude background interference.

[0032] S200. Input the first images that meet the preset conditions into the target detection module one by one. The target detection module performs adaptive anchor boxes on each component of the pantograph, so as to identify each component of the pantograph in the first image.

[0033] Specifically, the target detection module can adopt a target detection network. For example, it can use basic structures such as cnn convolution, depthwise separable convolution, and dilated convolution combined with the SPPF module (Spatial Pyramid Pooling), PAA module (Label Smoothing Assignment Technique), and PAN module (Path Aggregation Network Technique) to construct a deep learning network to enhance the feature extraction ability and the detection effect of scale targets. The size of the first image input can be 640×640.

[0034] S300. According to the target region image obtained in S200, determine whether each component is abnormal according to the preset conditions, and output the first determination result, including: S310. Determine whether a component is missing in the first image according to the preset topological dependency relationship between each component.

[0035] Specifically, define the key components, including: the base, the fixed structure where the pantograph is connected to the roof. The bow angle, the bent or upturned structure at both ends of the pantograph carbon slide. The carbon slide, the conductive slide that directly contacts the catenary.

[0036] Combined with Figure 2 As shown, according to the preset topological dependency rules, it is determined that: if the base is not detected, it is determined that the entire pantograph is missing. If the base is detected but the bow angle is not detected, it is determined that the bow angle is missing. If the bow angle is detected but the carbon slide is not detected, it is determined that the carbon slide has fallen off. And, it is only when the loss of the same component is detected in 3 consecutive frames (i.e., three first images) that the final alarm is triggered to avoid false detection.

[0037] In this way, the determination method based on the logical relationship between components avoids misjudgment or missed judgment that may occur in the detection of a single component.

[0038] S320. Identify the dynamic brightness threshold of the area around the catenary in the first image (such as the area where the pixel value in the image exceeds 220), and use the optical flow tracking technology. If the high-brightness area is tracked in 3 or more consecutive frames of images, it is determined as a real spark, and it is determined that the spark occurs around the catenary, thereby determining that a spark phenomenon appears.

[0039] S330. Determine whether the carbon slide is abnormal, including: Calculate whether the proportion of the fracture length of the carbon slide in the first image exceeds the threshold, and calculate whether the inclination angle of the carbon slide in the first image exceeds the threshold. If the proportion of the fracture length or the inclination angle exceeds the threshold, it is determined that the carbon slide is abnormal.

[0040] Specifically, perform fracture detection and calculate the proportion of the fracture length of the carbon slide mask:

[0041] In the formula, L 断裂 is the fracture length, and L 总 is the total length of the carbon slide.

[0042] The threshold is set to 5%, and if it exceeds 5%, it is determined as fractured.

[0043] Perform inclination detection, use the RANSAC algorithm to fit the straight line of the carbon slide edge, calculate the deviation from the reference angle, and the threshold is set to 3°. If it exceeds the threshold, it is determined as inclination abnormality.

[0044] Generally, fracture and inclination usually occur independently. Therefore, if either the fracture or the inclination condition exceeds the threshold, it is determined that the carbon slide is abnormal. In this way, by judging whether the carbon slide is abnormal from two dimensions, a reliable judgment result can be obtained.

[0045] S340. Calculate whether the relative position between the pantograph and the catenary in the first image exceeds the threshold. If it exceeds the threshold, it is determined as abnormal. Specifically, the calculation formula is:

[0046] In the formula, H is the pantograph height, that is, the vertical distance from the apex of the pantograph angle to the track plane, which can be calibrated through the image, H t is the pantograph height of the current frame; C is the catenary height, which can be the pre-calibrated static height value.

[0047] The above S310 - S340 can be carried out synchronously or sequentially, and the execution order can be different.

[0048] It is understandable that by making detailed and targeted judgments on aspects such as component integrity, spark phenomenon, carbon slide plate status, and the relative position of the pantograph and the contact network, the operating status of the pantograph can be fully and accurately grasped, and various abnormal situations can be discovered in time, thereby providing accurate and reliable results for the maintenance and management of the pantograph, ensuring the safety and stability of locomotive operation.

[0049] In some embodiments, identifying the carbon slide plate in the pantograph in the first image includes: The first image is input into the semantic segmentation model to generate a mask image of the first image, and the mask image divides the carbon skateboard area and the non-carbon skateboard area to identify the carbon skateboard; wherein the semantic segmentation model includes: an encoder and a decoder, the decoder adopts a nested jump connection structure and integrates a CBAM attention module.

[0050] For example, an improved U-Net++ architecture is used to enhance the fine segmentation capability of the carbon skateboard area through multi-level feature fusion and attention mechanism, as follows: refer to Figure 3 Understand that ResNet-34 is used as the encoder backbone network, and the weights pre-trained on the ImageNet dataset can be loaded to accelerate convergence and perform multi-scale feature extraction. For example, the five stages of ResNet-34 (Stage 1 to Stage 5) generate feature maps of different scales, with sizes of 512×512 (Stage 1), 256×256 (Stage 2), 128×128 (Stage 3), 64×64 (Stage 4), and 32×32 (Stage 5), respectively, and the number of channels is 64, 128, 256, 512, and 512. The output feature map of each stage is passed to the corresponding level of the decoder through a jump connection, retaining spatial information of different granularities to output a multi-scale feature map.

[0051] The decoder uses nested skip connections. For example, the decoder adopts the dense nested connection design of U-Net++, and fuses the same-level encoder features and deep decoder features through horizontal connections. In each level of skip connection, CBAM (Convolutional Block Attention Module) is integrated, which contains two branches of channel attention and spatial attention: Channel attention dynamically adjusts the channel weights through global average pooling and fully connected layers to focus on key feature channels. Spatial attention uses convolutional layers to generate spatial weight maps to highlight the spatial positions of the carbon skateboard areas. Based on the upsampling strategy, such as using bilinear interpolation for 2x upsampling, the refined features combined with skip connections are gradually restored to the original resolution. Finally, the decoder output is activated by convolution and activation functions to generate a 512×512 pixel-level binary mask, where the white area (pixel value = 1) represents the carbon skateboard area, and the black area (pixel value = 0) is the non-carbon skateboard area.

[0052] In this example, during the training of the semantic segmentation model, a composite loss function is adopted, that is, Dice Loss and Focal Loss are weighted and summed to balance the class imbalance problem between the carbon skateboard area (foreground) and the background. And data augmentation is carried out, randomly rotating by ±30, the probability of random horizontal / vertical flipping is 0.5, the filling mode is mirror reflection, and the brightness jitter is ±20% to simulate light changes and simulate rain and snow noise. For example, superimpose semi-transparent stripe noise (random direction, density 5% - 15%) and Gaussian snowflakes (density 3% - 8%).

[0053] In this way, the method shows significant application advantages in the actual detection under complex working conditions. Facing the problem of texture blur caused by edge wear, local fracture or surface oxidation layer attachment of the carbon skateboard during the high-speed operation of the pantograph, the nested skip connection structure can capture the macroscopic contour of the overall shape of the carbon skateboard and finely locate the edges of fine cracks or notches by fusing multi-scale features. The integrated CBAM attention mechanism effectively suppresses the influence of dynamic interferences such as catenary arc spots and rain and snow splashes on the segmentation results. It enhances the recognition of the metal reflection characteristics of the carbon skateboard conductive coating through channel weight recalibration, and at the same time, the spatial attention focuses on high-risk areas such as the connection between the skateboard and the pantograph horn. For the scenarios where there may be oil attachment or ice and snow coverage on the surface of the carbon skateboard, the mirror reflection filling and noise simulation mechanisms introduced during model training enhance the robustness of the segmentation algorithm to non-uniform occlusion, ensuring that even if part of the skateboard area is covered by semi-transparent foreign objects, the complete contour can be inferred through context features. This high-precision pixel-level segmentation ability provides a reliable basis for subsequent calculations of key indicators such as the proportion of fracture length and tilt angle. Especially for the accurate delineation of the serrated fracture edge of the carbon skateboard, it avoids the fracture misjudgment caused by uneven illumination in the traditional threshold segmentation method, making the damage quantification assessment based on image analysis closer to the real physical state.

[0054] Based on the above examples, although targeted detections are carried out for different abnormal conditions of the pantograph, there may still be problems with inaccurate judgments for image samples such as subtle abnormalities (such as microcracks in the carbon slide plate) and controversial samples (such as the distinction between sparks and reflections).

[0055] Therefore, two embodiments will be disclosed later to respectively identify microcracks in the carbon slide plate and distinguish between sparks and reflections.

[0056] The first embodiment, referring to Figure 5 As shown, after S300, continue with S400A, input the first image with the masked image into the first large model, and the first large model generates a second determination result.

[0057] Specifically, the first large model is an image large model, and the first image input into the first large model is the first image identified as containing the carbon slide plate and being...

[0058] Exemplarily, the first large model selects BLIP-2 as the basic large model, and inputs the candidate image with the masked picture carrying the masked label into the vision-language large model constructed based on the BLIP-2 architecture for secondary verification; among them, the model realizes domain adaptation through the LoRA (Low-Rank Adaptation) lightweight fine-tuning method: insert a trainable rank decomposition matrix (such as the rank dimension is set to 64) between the original Q-Former and the vision Transformer module of BLIP-2, freeze 95% of the parameters of the original model and only update the weights of the newly added adaptation layer, so as to inject prior knowledge while retaining the general visual representation ability. The multi-task joint loss function is adopted in the fine-tuning stage, including cross-entropy classification loss, contrastive learning loss and image-text matching loss.

[0059] Thus, when inputting, perform pixel-level superposition of the candidate image and the carbon slide plate mask, crop out the high-confidence abnormal area and enlarge it to a resolution of 1024×1024. Subsequently, input the processed image patch and the preset prompt template; for example, "Does this image show real microcracks in the carbon slide plate? Possible options: A. Yes, B. No, C. Uncertain", and then jointly input them into the large model, and generate a fine-grained analysis result through vision-language joint reasoning. The final alarm determination includes: If the number of first images recognized by the second determination result contains the first determination result, output the second determination result. If there is an intersection between the output results of the second determination result and the first determination result or the first determination result contains the second determination result, output the second determination result.

[0060] The second embodiment, referring to Figure 6 As shown, after S300, continue with S400B: S410B. Select the first images determined to be spark phenomena and the first images that do not meet the illumination uniformity (image reflection) in the first determination results, and mark these two types of first images as controversial images.

[0061] S420B. Based on the object recognition algorithm described in the above embodiments or any object recognition method, each controversial image can be attached with a part bounding box obtained through object recognition. Moreover, each controversial image is also attached with a confidence level. This confidence level can be further analyzed during object recognition for the controversial image (i.e., the first image that may have a spark phenomenon or image reflection), and the confidence level is calculated based on a classification model. The model outputs two probability values, and the larger probability value is used as the confidence level for the corresponding category (spark phenomenon or image reflection) of the controversial image. For example, if the model outputs a probability of 0.8 for a certain controversial image being a spark phenomenon and a probability of 0.2 for image reflection, then 0.8 is used as the confidence level for the image being a spark phenomenon, and 0.2 is used as the confidence level for it being image reflection. In this way, by performing the above processing on each controversial image, the confidence levels of each image that may be a spark phenomenon or image reflection can be obtained.

[0062] S430B. Input the controversial image with the bounding box and the confidence level into the second large model, and the second large model outputs a classification probability as the third determination result.

[0063] Exemplarily, the second large model is also an image large model. Refer to Figure 4 Understand its architecture, including a vision-language joint framework composed of a ViT encoder, a LoRA fine-tuning module, and an OPT generation module. After inputting the controversial image (such as a sample suspected of a spark phenomenon or image reflection) with the attached part bounding box and the primary detection confidence level into this model, the model is guided to perform fine-grained analysis through a preset prompt template "Judge whether the [region marked by the bounding box] in the image is [the true abnormal type], and it is necessary to distinguish [normal interference factors] and refer to the primary detection confidence level {confidence level value}", and finally output the classification probability result for the controversial category (for example, presented in a quantitative form such as "True spark probability: 0.92, Reflection interference probability: 0.08").

[0064] Of course, the above S400A and S400B can also be executed simultaneously or successively to output the final result and obtain a more accurate pantograph abnormal detection result.

[0065] In some embodiments, according to the above embodiments, it is also possible to perform an evaluation through a fusion algorithm to overcome the problem that the model generalization performance is insufficient and a single detection module is difficult to cover the dynamically changing outdoor scenarios.

[0066] Specifically, perform S500 after S400B. Continue to refer toFigure 6 As shown: Based on the parameters in the above embodiments, to enhance the accuracy of determination, a final alarm decision is made; the fusion algorithm is as follows:

[0067] In the formula, scorei is the comprehensive score, ranging from 0 to 1, and an anomaly with a score higher than 0.9 is considered a correct anomaly; si is the confidence of primary detection, including the confidence of the spark phenomenon and the reflection phenomenon in the foregoing examples; α is the weight of the primary detection confidence, obtained through adjustment or training. Pi is the determination probability of the second largest model, β is the weight of the determination probability of the large model, obtained through adjustment or training; is the semantic consistency score obtained based on the determination category of the first determination result and the semantic category output by the second largest model, is the classification consistency weight. Among them, α, β, λ is determined by grid search and cross-validation, for example, the optimal value: α = 0.5, β = 0.3, λ = 0.2 . The semantic consistency score is calculated by comparing the cosine similarity between the category of the first determination result ci and the category of the semantic determination large model ei .

[0068] The constraint condition is: α + β + λ = 1 .

[0069] An abnormal detection system for a pantograph image according to an embodiment of the present application includes a storage and a processor. The storage stores a computer-readable medium for implementing the abnormal detection method of the pantograph image in the above embodiment, and the processor is used to execute the computer-readable medium. The above storage and processor can be set in any computer system or an integrated system constructed by multiple computers.

[0070] The above is only the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope conceived herein through the above teachings or the technology or knowledge in related fields. And any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. An abnormal detection method based on pantograph images, characterized in that: Obtain a series of first images of the pantograph area, calculate the occlusion area and illumination uniformity of the first images, and screen out the corresponding first images if either the occlusion area or the illumination uniformity does not meet the preset conditions; Identify each component of the pantograph in the first image, determine whether each component is abnormal according to the preset conditions, and output a first determination result.

2. The abnormal detection method based on pantograph images according to claim 1, characterized in that, Identifying each component of the pantograph in the first image includes: carbon slide identification; Input the first image into a semantic segmentation model to generate a mask image of the first image, and the mask image divides the carbon slide area and the non-carbon slide area to identify the carbon slide; Among them, the semantic segmentation model includes: an encoder and a decoder, the decoder adopts a nested skip connection structure, and integrates a CBAM attention module.

3. The abnormal detection method based on pantograph images according to claim 2, characterized in that, Input the first image with the mask image into a first large model, and the first large model generates a second determination result.

4. The abnormal detection method based on pantograph images according to claim 1 or 2, characterized in that, Determine whether each component is abnormal according to the preset conditions, and output the first determination result, including: Determine whether there is a missing component in the first image according to the preset topological dependency relationship between the components; Identify the dynamic brightness threshold in the catenary surrounding area of the first image and use the optical flow tracking technology to determine whether there is a spark phenomenon; Calculate whether the fracture length ratio of the carbon slide in the first image exceeds the threshold, calculate whether the tilt angle of the carbon slide in the first image exceeds the threshold, and if the fracture length ratio or the tilt angle exceeds the threshold, it is determined that the carbon slide is abnormal; Calculate whether the relative position between the pantograph and the catenary in the first image exceeds the threshold, and if it exceeds the threshold, it is determined to be abnormal.

5. The abnormal detection method based on pantograph images according to claim 4, characterized in that, The topological dependency relationship is the topological dependency relationship between the base, bow angle, and carbon slide: If the base is not detected, it is determined that the entire pantograph is missing; If the base is detected but the bow angle is not detected, it is determined that the bow angle is missing; If the bow angle is detected but the carbon slide is not detected, it is determined that the carbon slide has fallen off.

6. The abnormal detection method based on pantograph images according to claim 4, characterized in that, Select the first images determined to have a spark phenomenon in the first determination result, and the first images that do not meet the illumination uniformity, and mark them as controversial images; Each of the controversial images is attached with a component bounding box obtained by object recognition, and each of the controversial images is also attached with a confidence level; Input the controversial images with the bounding boxes and the confidence levels into a second large model, and the second large model outputs a classification probability as a third determination result.

7. The abnormal detection method based on pantograph images according to claim 6, characterized in that, Obtain a semantic consistency score according to the detection category of the first determination result and the semantic category of the second large model, and assign a classification consistency weight; assign a confidence weight to the confidence level, and assign a probability judgment weight to the classification probability; Input the confidence level, confidence weight, classification probability, probability judgment weight, classification consistency weight, and semantic consistency score into a fusion decision algorithm to obtain a comprehensive score; Determine whether the comprehensive score exceeds the threshold, and output a fourth determination result.

8. The abnormal detection method based on pantograph images according to claim 1, characterized in that, Calculating the occlusion area and illumination uniformity of the first image, and screening out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions includes: Segmenting the occluder based on the Mask R-CNN model, calculating whether the proportion of the occlusion area of the occluder in the first image exceeds a threshold, and screening out if it exceeds; calculating whether the variance of the brightness component analyzed in the HSV space exceeds a threshold, and screening out if it exceeds.

9. The abnormal detection method based on pantograph images according to claim 1 or 8, characterized in that, The calculating the occlusion area and illumination uniformity of the first image, and screening out the corresponding first image if either the occlusion area or the illumination uniformity does not meet the preset conditions further includes: Determining whether the clarity of the first image exceeds a threshold.

10. An abnormal detection system based on pantograph images, characterized in that, It includes a memory and a processor. The memory stores a computer-readable medium for implementing the abnormal detection method based on the pantograph image described in any one of claims 1-9, and the processor is used to execute the computer-readable medium.

Citation Information

Patent Citations

  • Analysis processing method and device of three-dimensional model

    CN107679562A

  • Neural network model prediction confidence calibration method and system and storage medium

    CN112183751A

  • Vehicle-mounted image recognition method for pantograph fault

    CN113436157A

  • Semi-supervised sketch image retrieval method based on pseudo labels and reordering

    CN114168773A

  • Lane line detection method based on traffic off-site law enforcement scene

    CN115620259A

Cited By

  • Defect analysis method for electrified railway power supply safety detection and monitoring system 5C

    CN122134722A

  • Defect Analysis Method for 5C Power Supply Safety Detection and Monitoring System in Electrified Railways

    CN122134722B