Visual detection method and system for urban rail rigid contact network ground wire loose strand disease

By combining the Mask R-CNN network and the loose strand defect classifier, the automatic, real-time and efficient detection of loose strand defects in the ground wire of the rigid contact network of urban rail transit was realized. This solved the inconsistency and environmental adaptability problems of traditional manual detection, and improved the accuracy of detection and the reliability of the system.

CN121010571APending Publication Date: 2025-11-25BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511111042.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional manual inspection methods are difficult to meet the high-efficiency and accurate requirements for detecting loose strands in the ground wire of urban rail rigid contact networks, and are easily affected by environmental and human factors, leading to inconsistent inspection results and safety issues.

Method used

The Mask R-CNN network, combined with RoI pooling and a scattered disease classifier, is used to achieve accurate segmentation and disease classification of the ground wire area through high-definition image acquisition, multi-stage training and attention mechanism. Real-time detection is carried out in conjunction with the vehicle-mounted processing system and the ground control center.

Benefits of technology

It enables automated, real-time, and efficient detection of ground wire strand defects, improving the accuracy and adaptability of detection, adapting to different line and environmental conditions, reducing manual intervention, and ensuring the all-weather reliability and safety of the detection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010571A_ABST
    Figure CN121010571A_ABST
Patent Text Reader

Abstract

The invention discloses a visual detection method and system for urban rail rigid contact net ground wire loose strand diseases. The method comprises the steps that high-definition image data of a rigid contact net ground wire are collected and preprocessed; inputting the preprocessed ground wire image data into the trained Mask R-CNN network, and segmenting a grounding jumper and a ground wire area based on a cascade segmentation strategy; through RoI pooling operation, high-dimensional feature representation of a ground wire area is extracted from the deep convolutional feature map of the Mask R-CNN network, and a multi-scale deep feature map is output; based on a Mask R-CNN deep feature trained loose strand disease classifier, performing classification prediction and category judgment on normal ground lines and different degrees of loose strand diseases by adopting an attention mechanism and multi-scale feature fusion; according to the invention, automatic identification and positioning of the ground wire loose strands are realized through the deep learning algorithm, the detection efficiency and accuracy are improved, and the labor cost and the safety risk are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of rail transit, more particularly to a visual detection method and system for ground wire loose strands of rigid catenary of urban rail transit. BACKGROUND

[0002] The rigid catenary of urban rail transit is a key component of the power supply system of urban rail transit, and the ground wire as an important safety protection device is crucial to ensure the safety of train operation.

[0003] Ground wire loose strands are a common disease caused by long-term mechanical stress, environmental corrosion and other factors, and if not discovered and treated in time, it may lead to ground wire breakage and even cause serious safety accidents.

[0004] Therefore, timely and accurate detection of ground wire loose strand disease is of great significance to maintain the safe operation of urban rail transit systems.

[0005] With the rapid expansion of urban rail transit networks, traditional manual inspection methods have been difficult to meet the growing detection needs.

[0006] Manual detection not only takes a long time and is low in efficiency, but also is easily affected by subjective factors of the detector, leading to inconsistency of the detection results; in addition, under complex environmental conditions such as insufficient light or high-altitude work, the accuracy and safety of manual detection also face challenges.

[0007] Therefore, how to develop an automatic, efficient and accurate ground wire loose strand disease detection method is a problem that needs to be solved by those skilled in the art. SUMMARY

[0008] Therefore, the present application provides a visual detection method and system for ground wire loose strands of rigid catenary of urban rail transit to solve some of the technical problems mentioned in the background art.

[0009] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0010] A visual detection method for ground wire loose strands of rigid catenary of urban rail transit, comprising the following steps:

[0011] S1. Collecting high-definition image data of the rigid catenary ground wire and performing pretreatment;

[0012] S2. Inputting the pretreated ground wire image data into the trained Mask R-CNN network, and based on the cascaded segmentation strategy, segmenting the ground jump wire and ground wire area;

[0013] S3. Extracting high-dimensional feature representation of the ground wire region from the deep convolutional feature map of the Mask R-CNN network through the RoI pooling operation, and outputting a multi-scale deep feature map;

[0014] S4. Inputting the multi-scale deep feature map of the ground wire region into the stock disease classifier trained based on the deep features of the Mask R-CNN, and adopting attention mechanism and multi-scale feature fusion technology to perform classification prediction and category determination of normal ground wire and different degrees of stock disease.

[0015] Preferably, the visual detection method for the stock disease of the urban rail rigid catenary ground wire further comprises smoothing and post-processing fusion of continuous multiple frames of results, generating detection results and early warning information, and transmitting the detection results and early warning information to the ground control center in real time.

[0016] Preferably, the pre-processing content includes but is not limited to image correction, denoising, contrast enhancement, and blocking, so as to improve the image quality and the recognizability of the features.

[0017] Preferably, the specific content of training the Mask R-CNN network is as follows:

[0018] S21. Collecting high-definition image data of the rigid catenary ground wire under different time, weather conditions and line environments, and pre-processing the collected original images;

[0019] S22. Data labeling on the pre-processed image data of the rigid catenary ground wire;

[0020] S23. Using a multi-stage training strategy, training the Mask R-CNN network using the labeled ground jump wire data, optimizing the network parameters to accurately segment the ground jump wire, and based on the segmentation result of the ground jump wire, further training the Mask R-CNN network to accurately segment the ground wire; Specifically, using the transfer learning technology, using the pre-trained model weight on a large-scale general data set to initialize the network, introducing the feature pyramid network FPN structure on the basis of the Mask R-CNN, and performing multi-scale feature fusion.

[0021] Preferably, in step S22, the specific content of data labeling is as follows:

[0022] S21. Using a pre-trained target detection model to preliminarily label the image, and automatically marking the positions of the ground jump wire and the ground wire;

[0023] S22. An expert reviews and corrects the labeling result to ensure the accuracy and integrity of the labeling, and for the labeling of the stock disease, marks the position, degree and type of the stock.

[0024] Preferably, the trained Mask R-CNN network comprises an input layer, a backbone feature extraction network, a region proposal network, a RolAlign module, a multi-task branch, and an output layer.

[0025] The input layer is used to input the collected high-definition ground wire image of the urban rail rigid contact network, and the image contains complete ground wires and possible broken strand disease areas.

[0026] The backbone feature extraction network extracts multi-scale ground wire and broken strand disease features through ResNet-50 or ResNet-101 combined with a feature pyramid network (FPN).

[0027] The region proposal network inputs the feature map output by the backbone feature extraction network to generate a series of candidate regions for positioning the ground wire area where the disease may exist.

[0028] The RolAlign module precisely aligns the candidate regions generated by the region proposal network to ensure the spatial information consistency of the subsequent branches.

[0029] The multi-task branch includes a classification branch to determine whether each candidate region is a broken strand disease and its category, a bounding box regression branch to accurately regress the bounding box coordinates of the disease area, and a mask branch to generate a pixel-level segmentation mask for each positive sample region to accurately outline the outline of the broken strand disease.

[0030] The output layer outputs the category, bounding box coordinates, and segmentation mask of each detected disease area.

[0031] Preferably, the specific content of step S3 is as follows:

[0032] S31. Obtain the bounding box coordinates (x, y, width, height) of the ground wire area from the Mask R-CNN network output of step S2.

[0033] S32. Select multiple deep convolutional layers in the Mask R-CNN network as feature extraction sources.

[0034] S33. Perform RoIPooling operation on each selected feature level, including coordinate mapping, feature region cropping, and fixed-size pooling.

[0035] S34. Perform multi-scale feature fusion on the features extracted from different levels, including feature alignment and feature splicing.

[0036] S35. Perform feature enhancement processing on the fused features to output multi-scale deep feature maps.

[0037] Preferably, the stock disease classifier trained based on the Mask R-CNN deep features comprises an input feature layer, a feature fusion module, an attention mechanism module, a fully connected classification head, and an output result layer.

[0038] By fusing features of different convolution layers, global structure information and local detail information of the ground wire are captured; by the attention mechanism, the most recognizable feature area is focused on; by the fully connected classification head, the probability distribution of each category of normal ground wire and different degrees of stock disease is output; and the category with the highest probability is selected as the final determination result.

[0039] A visual detection system for stock disease of a rigid catenary ground wire of urban rail transit, comprising a roof camera system, a vehicle-mounted processing system, and a ground control center.

[0040] The vehicle-mounted processing system comprises an image acquisition unit, a GPU computing device, and a high-speed wireless communication module.

[0041] The roof camera system is used to acquire high-definition image data of the rigid catenary ground wire.

[0042] The image acquisition unit comprises an image signal processor, which is used to segment the acquired original image data into multiple overlapping small blocks and perform preprocessing.

[0043] The GPU computing device deploys a trained deep learning model, which comprises a Mask R-CNN network, a RoI pooling module, a stock disease classifier, and a post-processing and result output module.

[0044] The Mask R-CNN network is used to input the preprocessed ground wire image data, and based on a cascaded segmentation strategy, the ground wire and the grounding jumper are segmented.

[0045] The RoI pooling module is used to extract high-dimensional feature representations of the ground wire area from deep convolution feature maps of the Mask R-CNN network, and output multi-scale deep feature maps.

[0046] The stock disease classifier is used to input the multi-scale deep feature maps of the ground wire area, and adopts an attention mechanism and a multi-scale feature fusion technology to perform classification prediction and category determination of normal ground wire and different degrees of stock disease.

[0047] The post-processing and result output module is used to perform smoothing processing and post-processing fusion on continuous multiple frames of results, and generate detection results and warning information.

[0048] The high-speed wireless communication module is used to transmit the generated detection results and warning information to the ground control center in real time.

[0049] Preferably, the vehicle-mounted processing system further comprises an intelligent power management device for power management of the image acquisition unit, the GPU computing device and the high-speed wireless communication module, and adopts an intelligent dynamic power consumption management mechanism to automatically adjust the processor frequency and the GPU usage according to the current processing load and the confidence of the detection result.

[0050] Via the technical solution, compared with the prior art, the application provides a visual detection method and system for the strand breakage disease of the rigid catenary ground wire of urban rail, and has the following beneficial effects:

[0051] The application realizes accurate identification and positioning of the strand breakage disease of the ground wire by adopting super-high-definition image acquisition and multi-stage deep learning algorithms; compared with the traditional method, the application significantly improves the precision and recall rates of detection, and can identify early strand breakage signs, thereby providing the possibility for preventive maintenance.

[0052] The application realizes full automation of the detection of the strand breakage disease of the ground wire, and can complete a large-scale and high-frequency detection task without manual intervention; through a series of optimization strategies, real-time detection is realized on a vehicle-mounted low-power server, thereby providing a guarantee for timely discovery and processing of potential diseases.

[0053] By adopting an advanced image enhancement algorithm, high-quality ground wire images can be obtained, thereby ensuring the all-weather reliability of the detection system; in addition, the application has good generalization ability and can adapt to detection requirements under different lines and environmental conditions. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0055] Figure 1 A visual detection method for the strand breakage disease of the rigid catenary ground wire of urban rail provided by the application is shown in the figure.

[0056] Figure 2 A Mask R-CNN network provided by the application is shown in the figure.

[0057] Figure 3 A strand breakage disease classifier training diagram provided by the application is shown in the figure.

[0058] Figure 4 A visual detection system for the strand breakage disease of the rigid catenary ground wire of urban rail provided by the application is shown in the figure. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0060] The embodiments of the present application disclose a visual detection method for strand disease of urban rail rigid catenary ground wire, which comprises the following steps: Figure 1 , comprising the following steps:

[0061] S1. Collecting high-definition image data of the rigid catenary ground wire and performing preprocessing;

[0062] S2. Inputting the preprocessed ground wire image data into a trained Mask R-CNN network, and based on a cascaded segmentation strategy, segmenting out the ground jump wire and ground wire regions;

[0063] S3. Extracting high-dimensional feature representation of the ground wire region from a deep convolution feature map of the Mask R-CNN network through a RoI pooling operation, and outputting a multi-scale deep feature map;

[0064] S4. Inputting the multi-scale deep feature map of the ground wire region into a strand disease classifier trained based on the deep features of the Mask R-CNN, and adopting an attention mechanism and a multi-scale feature fusion technology to perform classification prediction and category determination of normal ground wire and strand diseases of different degrees.

[0065] In order to further implement the above technical solutions, the visual detection method for strand disease of urban rail rigid catenary ground wire further comprises smoothing and post-processing fusion on continuous multiple frames of results, generating detection results and early warning information, and transmitting the detection results and early warning information to a ground control center in real time.

[0066] In order to further implement the above technical solutions, the preprocessing content includes but is not limited to image correction, denoising, contrast enhancement, blocking, so as to improve the image quality and the recognizable of the features.

[0067] In order to further implement the above technical solutions, the specific content for training the Mask R-CNN network is:

[0068] S21. Collecting high-definition image data of the rigid catenary ground wire under different time, weather conditions and line environments and preprocessing the collected original images;

[0069] S22. Data labeling on the preprocessed image data of the rigid catenary ground wire;

[0070] S23. Using the multi-stage training strategy, the labeled grounding jumper data is used to train the Mask R-CNN network, and the network parameters are optimized to accurately segment the grounding jumper. Based on the segmentation result of the grounding jumper, the Mask R-CNN network is further trained to accurately segment the ground wire; specifically, the transfer learning technology is used, the model weight pre-trained on a large-scale general data set is used to initialize the network, and the feature pyramid network FPN structure is introduced into the Mask R-CNN to perform multi-scale feature fusion.

[0071] In this embodiment, the segmentation strategy of the cascaded Mask R-CNN network fully utilizes the prior knowledge of the overhead contact system structure, and significantly improves the accuracy and robustness of the ground wire positioning.

[0072] The classification method based on deep features has stronger expression and generalization ability than the traditional method based on shallow features, and can more accurately identify various types and degrees of loose stock diseases

[0073] In order to further implement the above technical solutions, in step S22, the specific content of data labeling is:

[0074] S21. Using a pre-trained target detection model to preliminarily label the image, automatically marking the positions of the grounding jumper and the ground wire;

[0075] S22. The labeling results are reviewed and corrected by experts to ensure the accuracy and completeness of the labeling. For the labeling of loose stock diseases, mark the position, degree and type of loose stock.

[0076] In this embodiment, during the training process of the Mask R-CNN network, the transfer learning technology is used, the model weight pre-trained on a large-scale general data set is used to initialize the network, the training process is accelerated and the generalization ability of the model is improved; in order to improve the sensitivity of the model to small targets and detailed features, the feature pyramid network FPN structure is introduced into the Mask R-CNN to realize effective fusion of multi-scale features.

[0077] In order to further implement the above technical solutions, as Figure 2 , the trained Mask R-CNN network includes an input layer, a backbone feature extraction network, a region proposal network, a RolAlign module, a multi-task branch and an output layer;

[0078] The input layer is used to input the collected high-definition ground wire image of the urban rail rigid overhead contact system, and the image contains complete ground wire and possible loose stock disease area;

[0079] The backbone feature extraction network extracts multi-scale ground wire and loose stock disease features through ResNet-50 or ResNet-101 combined with the feature pyramid network FPN.

[0080] Region Proposal Network, input the feature map output by the backbone feature extraction network, generate a series of candidate regions for positioning the ground wire area where the disease may exist;

[0081] RolAlign module, accurate alignment of the candidate regions generated by the region proposal network, ensure the spatial information consistency of the subsequent branches;

[0082] Multi-task branch, including: classification branch, judge whether each candidate region is loose stock disease and its category; bounding box regression branch, accurately regress the bounding box coordinates of the disease area; mask branch, generate a pixel-level segmentation mask for each positive sample region, accurately outline the outline of the loose stock disease;

[0083] Output layer, output the category, bounding box coordinates and segmentation mask of each detected disease area.

[0084] In this embodiment, the cascaded segmentation strategy makes full use of the prior knowledge of the catenary structure, and significantly improves the accuracy and robustness of the ground wire positioning.

[0085] In order to further implement the above technical scheme, step S3, through RoIpooling operation, extract high-dimensional feature representation of ground wire area from deep convolution feature map of Mask R-CNN network, the specific method of outputting multi-scale deep feature map is:

[0086] S31. Obtain the bounding box coordinates (x, y, width, height) of the ground wire area from the Mask R-CNN network output of step S2;

[0087] S32. Select multiple deep convolution layers in the Mask R-CNN network as feature extraction sources;

[0088] In this embodiment, usually select feature maps of different resolutions, such as: select high-resolution feature maps to retain more detailed information, select medium-resolution feature maps to balance detailed and semantic information, and select low-resolution feature maps to contain more semantic information;

[0089] S33. Perform RoIPooling operation on each selected feature layer, including coordinate mapping, feature region cropping and fixed size pooling;

[0090] Coordinate mapping: considering the scaling ratio of the feature map relative to the original image, map the ground wire area bounding box coordinates from the original image coordinate system to the coordinate system of the corresponding feature map

[0091] Feature region cropping: according to the mapped coordinates, crop the corresponding ground wire area features from the feature map, handle the case that the bounding box exceeds the feature map boundary;

[0092] Fixed-size pooling: using max pooling or average pooling operation to unify the ground line area features of different sizes into fixed size (such as 7x7 or 14x14);

[0093] S34. Multi-scale feature fusion is performed on the features extracted at different levels, including feature alignment and feature splicing;

[0094] Feature alignment: aligning the features extracted at different levels to the same spatial size, using upsampling or downsampling operation for size alignment;

[0095] Feature splicing: splicing the features of multiple scales in the channel dimension to form a high-dimensional feature vector containing multi-scale information;

[0096] S35. Feature enhancement processing is performed on the fused features, and a multi-scale deep feature map is output;

[0097] Feature enhancement processing includes feature normalization and dimension reduction processing; feature normalization: standardizing the extracted features to ensure that the numerical ranges of different scale features are consistent; dimension reduction processing: using 1x1 convolution or fully connected layer for feature dimension reduction to reduce computational complexity and extract more compact feature representation;

[0098] The final output multi-scale deep feature map contains a high-dimensional feature representation of the ground line area multi-scale information, and the feature map contains detailed and semantic information extracted from different levels, providing rich feature input for the subsequent scattered stock disease classifier.

[0099] In order to further implement the above technical solutions, such as Figure 3 The scattered stock disease classifier based on Mask R-CNN deep feature training includes: input feature layer, feature fusion module, attention mechanism module, fully connected classification head and output result layer;

[0100] By fusing features of different convolution layers, global structure information and local detail information of the ground line can be captured; by attention mechanism, the most recognizable feature area is focused; by fully connected classification head, the probability distribution of normal ground line and different degrees of scattered stock disease categories is output; the class with the highest probability is selected as the final determination result;

[0101] Specifically:

[0102] The input feature layer is used for multi-scale deep feature map input and feature preprocessing, wherein the feature preprocessing includes feature normalization, dimension checking and data format conversion;

[0103] The feature fusion module is used for multi-scale feature alignment, feature fusion strategy and fused feature optimization. The multi-scale feature alignment is specifically: projecting features of different scales to the same feature space, adjusting the number of channels using a 1x1 convolution layer, and ensuring that all features have the same spatial resolution; the feature fusion strategy is specifically: assigning adaptive weights to features of different scales, determining the weights through learnable parameters, connecting features of different scales in order, learning the interaction between features through a convolution layer, and maintaining the hierarchical structure information of the features; the fused feature optimization is specifically: normalizing the fused features, applying an activation function, and adding a skip connection to alleviate the gradient vanishing problem.

[0104] The attention mechanism module includes a spatial attention mechanism, a channel attention mechanism and a multi-head attention mechanism.

[0105] The spatial attention mechanism includes: calculating spatial attention weights for the fused feature map, obtaining channel descriptors using global average pooling and maximum pooling, generating a spatial attention map through a convolution layer, multiplying the attention weights with the original features, highlighting important spatial regions and suppressing irrelevant background information; the channel attention mechanism includes: obtaining channel statistical information using global average pooling, learning the dependency between channels through a fully connected layer, generating channel attention weights, multiplying the channel weights with the features, enhancing the feature representation of important channels, and weakening the influence of irrelevant channels; the multi-head attention mechanism includes: dividing the features into multiple heads, independently calculating attention for each head, and finally concatenating the outputs of all heads; the features processed by the attention mechanism contain richer semantic information, providing more discriminative features for the classification task.

[0106] The fully connected classification head includes a feature dimension reduction layer, an intermediate classification layer and a final classification layer; the feature dimension reduction layer includes: converting the spatial feature map into a feature vector, reducing the number of parameters, maintaining the global information of the features, using a fully connected layer to reduce the feature dimension from 1024 to 512, applying Dropout to prevent overfitting, and using a ReLU activation function to introduce nonlinearity; the intermediate classification layer includes: mapping the 512-dimensional features to 256-dimensional features, further extracting high-level semantic features, applying batch normalization and Dropout, adding a residual connection, using a LeakyReLU activation function, and enhancing the expression ability of the features; the final classification layer includes: mapping the 256-dimensional features to the dimension of the number of categories, the number of categories is determined according to the number of shares of the disease (such as: normal, mild, moderate, severe), and using a Softmax activation function to output a probability distribution;

[0107] The output result layer includes a category probability output, a confidence evaluation, and a loose stock disease category determination; the category probability output is an output of a probability value of each category, and the sum of the probabilities is 1; the confidence evaluation is a prediction confidence calculated based on the highest probability value, and a confidence threshold is set, and a low confidence prediction can trigger manual review; and the loose stock disease category determination is to select the category with the highest probability as the final prediction, apply post-processing rules (such as smoothing filtering), and output the final loose stock disease category label.

[0108] In the present embodiment, during the training of the loose stock disease classifier, cross-validation and early stopping strategies are used to prevent overfitting; at the same time, through data enhancement techniques such as random rotation, scaling, brightness adjustment, etc., the diversity of the training samples is increased, and the robustness of the model is improved; in order to solve the problem of sample imbalance in practical application, the focal loss (FocalLoss) function is used, which effectively improves the recognition ability of the model for rare categories (such as severe loose stock).

[0109] A visual detection system for loose stock disease of a rigid catenary ground wire of a city rail, based on a visual detection method for loose stock disease of a rigid catenary ground wire of a city rail, such as Figure 4 , comprising a roof camera system, a vehicle-mounted processing system and a ground control center;

[0110] The vehicle-mounted processing system comprises an image acquisition unit, a GPU computing device and a high-speed wireless communication module;

[0111] The roof camera system is used to acquire high-definition image data of the rigid catenary ground wire;

[0112] A high-definition camera system is installed on the roof of the detection vehicle, and each camera can shoot a super-high-definition image of 5120x5120 pixels. The installation position and angle of the camera are accurately calculated and adjusted to cover the omnidirectional view of the rigid catenary, including the connection with the ground jumpers. The high-resolution image is transmitted in real time to the high-performance computing unit in the vehicle through a high-speed data transmission line for processing and storage in the GPU.

[0113] The image acquisition unit comprises an image signal processor for converting the acquired original video signal into a digital signal and performing preprocessing, including image correction, enhancement and blocking;

[0114] The preprocessed ground wire image data is transmitted to the high-performance GPU computing device in the vehicle through a PCIE high-speed bus;

[0115] The GPU computing device deploys a trained deep learning model, which comprises a Mask R-CNN network, a RoI pooling module, a loose stock disease classifier and a post-processing and result output module;

[0116] The Mask R-CNN network is used for inputting the pre-processed ground wire image data, and based on a cascaded segmentation strategy, the ground wire and the ground wire region are segmented out.

[0117] The RoI pooling module is used for extracting the high-dimensional feature representation of the ground wire region from the deep convolution feature map of the Mask R-CNN network, and outputting a multi-scale deep feature map.

[0118] The stock disease classifier is used for inputting the multi-scale deep feature map of the ground wire region, adopting an attention mechanism and a multi-scale feature fusion technology, and performing classification prediction and category determination of normal ground wires and different degrees of stock diseases.

[0119] The post-processing and result output module is used for smoothing and post-processing fusion of continuous multiple frames of results, and generating detection results and warning information.

[0120] The high-speed wireless communication module is used for transmitting the generated detection results and warning information to the ground control center in real time.

[0121] In actual application, for the deployment and execution of the system on hardware, in order to optimize the performance of the model in the actual environment, model quantization and pruning technology is adopted to reduce the parameter quantity and calculation complexity of the model, while maintaining high detection accuracy; in addition, model acceleration based on TensorRT is also realized, which fully utilizes the parallel computing capability of GPU, and significantly improves the inference speed.

[0122] In order to further implement the above technical scheme, the vehicle-mounted processing system further comprises an intelligent power management device for power management of the image acquisition unit, the GPU computing device and the high-speed wireless communication module, and adopts an intelligent dynamic power management mechanism to automatically adjust the processor frequency and GPU usage according to the current processing load and the confidence of the detection result.

[0123] Further, the GPU computing device comprises a plurality of high-performance GPUs, a large-capacity memory and a solid state disk array.

[0124] In the embodiment, the application adopts a series of optimization strategies in real-time aspect to ensure efficient image processing and analysis on the vehicle-mounted low-power server, specifically:

[0125] Firstly, a lightweight Mask R-CNN variant is adopted to significantly reduce the parameter quantity and calculation complexity of the model while maintaining high accuracy through model compression and knowledge distillation technology; secondly, a GPU-based parallel computing framework is realized to fully utilize the graphics processing unit of the vehicle-mounted server, and the inference speed of the deep learning model is greatly improved.

[0126] On the image processing flow, the streaming data processing technology is adopted, high-definition images are divided into multiple overlapping small blocks, the small blocks are loaded into the GPU memory in sequence for processing, in this way, the method effectively solves the memory limitation problem in high-resolution image processing, and realizes real-time analysis of large-size images;In order to adapt to the low-power requirement of the vehicle-mounted environment, the application realizes an intelligent dynamic power management mechanism, the system will automatically adjust the processor frequency and GPU usage according to the current processing load and the confidence of the detection result, when processing a simple scene, the system will reduce the power consumption to prolong the device use time;When encountering a complex scene or potential disease, the system will quickly improve the processing capacity to ensure the accuracy and real-time performance of the detection.

[0127] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between various embodiments can be referred to each other.

[0128] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual detection method for the disease of loose strands of the overhead contact system ground wire of urban rail transit rigid contact system, characterized in that, The method comprises the following steps: S1. Collecting high-definition image data of rigid catenary ground wire and pre-processing; S2. Inputting the pre-processed ground wire image data into the trained Mask R-CNN network, and segmenting the ground wire area and the grounding jumper based on a cascaded segmentation strategy; S3. Extracting high-dimensional feature representation of the ground wire area from the deep convolution feature map of the Mask R-CNN network through RoIpooling operation, and outputting a multi-scale deep feature map; S4. Inputting the multi-scale deep feature map of the ground wire area into a stock disease classifier trained based on the deep features of the Mask R-CNN, and adopting attention mechanism and multi-scale feature fusion technology to perform classification prediction and category determination of normal ground wire and different degrees of stock disease.

2. The visual detection method for the disheveling disease of the overhead line of the rigid catenary system for urban rail transit according to claim 1, characterized in that, It also includes smoothing processing and post-processing fusion of continuous multiple frames of results, generation of detection results and early warning information, and real-time transmission to the ground control center.

3. The visual detection method for the broken strand disease of the overhead contact system ground wire according to claim 1, characterized in that, The pre-processing content includes but is not limited to image correction, denoising, contrast enhancement, blocking, to improve the image quality and the recognizability of the features.

4. The visual detection method for the broken strand disease of the overhead contact system ground wire according to claim 1, characterized in that, The specific content of training the Mask R-CNN network is: S21. Collecting high-definition image data of rigid catenary ground wire under different time, weather conditions and line environments and pre-processing the collected original images; S22. Data labeling on the pre-processed image data of the rigid catenary ground wire; S23. Using a multi-stage training strategy, using the labeled grounding jumper data to train the Mask R-CNN network, optimizing the network parameters to accurately segment the grounding jumper, and based on the segmentation result of the grounding jumper, further training the Mask R-CNN network to accurately segment the ground wire; Specifically, using the transfer learning technology, using the pre-trained model weight on the large-scale general data set to initialize the network, introducing the feature pyramid network FPN structure on the basis of Mask R-CNN, and performing multi-scale feature fusion.

5. The visual inspection method for the broken strand disease of the overhead contact system ground wire according to claim 4, characterized in that, In step S22, the specific content of data labeling is: S21. Using a pre-trained target detection model to preliminarily label the image, and automatically marking the positions of the grounding jumper and the ground wire; S22. The labeling results are audited and corrected by experts to ensure the accuracy and integrity of the labeling, and for the labeling of stock disease, the position, degree and type of stock are marked.

6. The visual inspection method for the broken strand disease of the overhead contact system ground wire according to claim 1, characterized in that, The trained Mask R-CNN network includes an input layer, a backbone feature extraction network, a region proposal network, a RolAlign module, a multi-task branch and an output layer; The input layer is used for inputting the collected high-definition ground wire image of the urban rail rigid catenary, and the image contains complete ground wire and possible stock disease area; The backbone feature extraction network extracts multi-scale ground wire and stock disease features through ResNet-50 or ResNet-101 combined with the feature pyramid network FPN; The region proposal network inputs the feature map output by the backbone feature extraction network, generates a series of candidate regions, and is used for positioning the ground wire area where the disease may exist; The RolAlign module accurately aligns the candidate regions generated by the region proposal network to ensure the spatial information consistency of the subsequent branches; The multi-task branch includes: a classification branch for determining whether each candidate region is a loose stock disease and a category thereof; a bounding box regression branch for accurately regressing the bounding box coordinates of the disease region; and a mask branch for generating a pixel-level segmentation mask for each positive sample region to accurately outline the outline of the loose stock disease. An output layer outputs the category, bounding box coordinates, and segmentation mask of each detected disease region.

7. The visual inspection method for the broken strand disease of the overhead contact system ground wire according to claim 1, characterized in that, The specific content of step S3 is: S31. Obtain the bounding box coordinates (x, y, width, height) of the ground wire region from the Mask R-CNN network output of step S2; S32. Select multiple deep convolutional layers in the Mask R-CNN network as feature extraction sources; S33. Perform RoIPooling operations on each selected feature level, including coordinate mapping, feature region cropping, and fixed-size pooling; S34. Perform multi-scale feature fusion on the features extracted from different levels, including feature alignment and feature splicing; S35. Perform feature enhancement processing on the fused features to output multi-scale deep feature maps.

8. The visual inspection method for the broken strand disease of the overhead contact system ground wire according to claim 1, characterized in that, The loose stock disease classifier trained based on the Mask R-CNN deep features includes: an input feature layer, a feature fusion module, an attention mechanism module, a fully connected classification head, and an output result layer; By fusing features of different convolutional layers, global structural information and local detail information of the ground wire are captured; by using the attention mechanism, the most recognizable feature region is focused on; by using the fully connected classification head, the probability distribution of normal ground wire and different degrees of loose stock disease categories is output; and the category with the highest probability is selected as the final determination result.

9. A visual inspection system for loose strand defects in the ground wire of a rigid contact wire in urban rail transit, characterized in that, The visual detection method for loose stock disease of a rigid catenary ground wire according to any one of claims 1-8 comprises: a roof camera system, a vehicle-mounted processing system, and a ground control center; The vehicle-mounted processing system comprises: an image acquisition unit, a GPU computing device, and a high-speed wireless communication module; The roof camera system is used to acquire high-definition image data of the rigid catenary ground wire; The image acquisition unit comprises an image signal processor and is used to divide the acquired original image data into multiple overlapping small blocks and perform preprocessing; The GPU computing device deploys a trained deep learning model, which comprises a Mask R-CNN network, an RoI pooling module, a loose stock disease classifier, and a post-processing and result output module; The Mask R-CNN network is used to input the preprocessed ground wire image data and segment the ground wire and the ground wire region based on a cascaded segmentation strategy; The RoI pooling module is used to extract high-dimensional feature representations of the ground wire region from deep convolutional feature maps of the Mask R-CNN network and output multi-scale deep feature maps; The loose stock disease classifier is used to input the multi-scale deep feature maps of the ground wire region, adopt an attention mechanism and a multi-scale feature fusion technology, and perform classification prediction and category determination of normal ground wire and different degrees of loose stock disease; The post-processing and result output module is used to perform smoothing processing and post-processing fusion on continuous multiple frames of results to generate detection results and warning information; A high-speed wireless communication module is used to transmit the generated detection results and early warning information to the ground control center in real time.

10. The visual inspection system for detecting broken strands of an overhead contact system for urban rail transit according to claim 9, characterized in that, The vehicle-mounted processing system further comprises an intelligent power management device for power management of the image acquisition unit, the GPU computing device and the high-speed wireless communication module, and adopts an intelligent dynamic power consumption management mechanism to automatically adjust the processor frequency and GPU usage according to the current processing load and the confidence of the detection result.