Complex environment bow net positioning method based on deep learning
Through the DETR model based on deep learning, the problem of difficulty in identifying bow net detection in complex environments is solved, the stability and accurate identification of bow net is achieved, and the robustness and accuracy of the detection system are improved.
Patent Information
- Application Number
- CN202510290473.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
When the existing non-contact detection technology of bow nets is faced with overexposure images caused by arcing or mutations in the contact position of the bow nets at the anchor joint, it is difficult to identify feature information, which increases the risk of target loss and leads to the failure of target detection methods.
Using a complex environmental bow network positioning method based on deep learning, the DETR model is constructed, and the convolutional neural network and Transformer structure are used to detect and identify the target area of the bow network. The method includes data preprocessing, model initialization, training and verification, and is able to process bow net images in complex environments.
It significantly improves the robustness and accuracy of the bow net detection system, and can stably and accurately identify the bow net in complex scenarios, ensuring the continuity and reliability of the detection process.
Smart Images

Figure CN120220023A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pantograph-catenary detection for rail locomotives, and more specifically, to a pantograph-catenary positioning method based on deep learning in a complex environment. Background Art
[0002] Most rail locomotives in urban rail transit obtain power through the pantograph-catenary system to operate. The relationship between the pantograph and the catenary can clearly reflect the current collection relationship of the pantograph, and its working quality directly determines whether electrical energy can be normally supplied to the running vehicle. Therefore, it is very necessary to perform accurate and rapid target detection on the pantograph-catenary.
[0003] The detection of the pantograph-catenary system mainly adopts manual inspection and contact sensor methods. With the rapid development of computer vision technology, non-contact image detection technology has gradually been applied to the detection of the pantograph-catenary system. Compared with the traditional contact measurement method, the non-contact pantograph-catenary system detection using image processing technology has the characteristics of rapid response, rich image information, strong intuitiveness, and comprehensive functions of the data processing system.
[0004] The existing non-contact detection technology for the pantograph-catenary mainly relies on extracting feature information from images and monitoring the state of the pantograph based on the positions of pixel points in these feature information. However, when arcing occurs in the pantograph-catenary, resulting in overexposure of the image, or when the pantograph passes through an anchor section joint, causing a sudden change in the pantograph-catenary contact position, the recognition of feature information becomes difficult, which increases the risk of target loss. In these cases, traditional target detection methods may fail because they cannot effectively process overexposed pictures and rapid changes in the joint area, resulting in instability in target tracking. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a pantograph-catenary positioning method based on deep learning in a complex environment, which can improve the robustness and accuracy of the pantograph-catenary detection system.
[0006] The technical solution adopted by the present invention to solve its technical problems is: constructing a pantograph-catenary positioning method based on deep learning, including the following steps:
[0007] Step S1: Obtain train video data containing complex situations of arcing anchor section joints;
[0008] Step S2: Calibrate the pantograph-catenary contact area, and jointly form a training set and a validation set with images in normal and complex situations;
[0009] Step S3: Design and initialize a pantograph-catenary target area detection model;
[0010] Step S4: Train the convolutional neural network DETR model with the obtained training set to obtain a trained pantograph-catenary target area detection model;
[0011] Step S5: Input the validation set data into the trained pantograph-catenary target area detection model to verify the recognition effect.
[0012] According to the above scheme, in the step S1, the train video data of the entire section is selected to ensure that the data set covers various situations, including the images of arcing in the pantograph-catenary contact area, the alternation of the catenary, and the images of special contact areas.
[0013] According to the above scheme, the construction process of the training set and the validation set in the step S2 includes the following steps:
[0014] Preprocess the train operation video images to confirm the number of video frames and resolution parameters;
[0015] During the process of calibrating the target area, for the complex situation areas that appear in the video, a manual division strategy is adopted, and necessary manual annotations are made in the case of the arcing bright spot obscuring the pantograph-catenary area and the complex situation of the sudden change in the contact position of the pantograph passing through the anchor section.
[0016] Use the python script to remove the parking intervals outside the running time frame by frame, so as to generate the corresponding data set, and then generate the training set and the validation set according to the ratio.
[0017] According to the above scheme, in the step S3, the trained DETR network model includes a backbone network, an encoding layer, a decoding layer, and an output head. Initialize it and set the output format; and add positional encoding to the input feature map to help the model perceive the spatial position information of the target.
[0018] According to the above scheme, in the step S3, the backbone network of the model uses ResNet with 101 layers, and the final output is (batch_size, 768, 256); the encoding layer has 6 layers, that is, iterate 6 times, and the final output of the Encoder is (batch_size, 768, 256); the decoding layer has 6 layers, the input is 100 learning vectors, the vector channel number is 256, and these 100 vectors will also perform attention mechanism calculations with the output K and V vectors provided by the encoding layer. The final output of the decoding layer is (batch_size, 100, 256).
[0019] According to the above scheme, in the step S4, the method for model training includes:
[0020] Input the enhanced training images and the corresponding pantograph-catenary target area annotation data;
[0021] The backbone network extracts the feature map, the encoder performs global modeling on the feature map, outputs the K and V vectors, and the decoder interacts with the K and V output by the encoder through the query vector to predict the category and bounding box of the target.
[0022] Obtain the calculated loss function and optimize it. Calculate the loss function based on the Hungarian matching result and the true label, and use the AdamW optimizer for gradient descent.
[0023] According to the above scheme, in the step S4, the method of Hungarian matching includes:
[0024] Hungarian matching optimization based on pantograph-catenary dynamics, deeply integrate the loss calculation of the DETR model with the physical parameters of the pantograph-catenary system, expand the original predicted categories to the categories of key pantograph-catenary components, and introduce the vertical displacement and lateral offset of the catenary as additional features;
[0025] The predicted bounding box value is defined as a vector containing geometric parameters:
[0026] b i =[x1,y1,Δh,Δθ]
[0027] where x1 and y1 are the coordinates of the pantograph-catenary contact point, Δh is the height change at the contact position, and Δθ is the tilt change of the pantograph
[0028] Calculate the L1 and GIOU losses for the predicted bounding box results and the true annotation box results to generate a predicted bounding box loss matrix;
[0029] Form the final loss matrix by weighted summation of the predicted class loss matrix and the predicted bounding box loss matrix; the Hungarian algorithm completes the optimal matching of the predicted boxes and the true boxes:
[0030]
[0031] In the formula: is the predicted loss of the true class, is the predicted loss between the true pantograph-catenary contact geometric parameters and the predicted pantograph-catenary geometric parameters, l is the weight, N is all classes, is the result after Hungarian matching.
[0032] According to the above scheme, in the step S4, the method of calculating the final loss includes:
[0033] Initialize a vector with 100 elements, all of which are background classes; according to the Hungarian matching result, replace the corresponding positions of the vector with the true classes of the predicted boxes;
[0034] Calculate the cross-entropy loss for this vector and the prediction result to obtain the final class loss;
[0035] According to the matching results, the corresponding bounding box prediction values are obtained from 100 prediction results, and the L1 and GIOU losses are calculated with the true annotation boxes to obtain the final bounding box loss. The formula is as follows:
[0036]
[0038] According to the above scheme, in the step S5, in order to more intuitively display the prediction ability of the model for the pantograph-catenary contact area, two types of interfering pantograph-catenary images are selected for detection. One is the case where the pantograph-catenary contact area is blocked by arcing, and the other is the case where catenary alternation appears in the image.
[0039] Implementing the complex environment pantograph-catenary positioning method based on deep learning of the present invention has the following beneficial effects:
[0040] 1. Through in-depth training of the DETR (Detection Transformer) neural network model of the present invention, the image detection robustness of the model in the face of complex situations is significantly improved. This optimization enables the model to achieve stable and accurate identification of the pantograph-catenary even in complex scenarios such as overexposed pictures caused by arcing or sudden changes in the pantograph-catenary contact position at the anchor section joint during pantograph-catenary detection, thus ensuring the continuity and reliability of the detection process.
[0041] 2. The pantograph-catenary positioning method provided by the present invention can quickly and effectively process overexposed pictures and pantograph-catenary pictures related to anchor section joints, and is suitable for object detection tasks in complex environments. This method can achieve real-time monitoring of the pantograph-catenary state during train operation, showing high accuracy and robustness, and providing strong technical support for the safe operation of trains. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0043] Figure 1 is the flow chart of the complex environment pantograph-catenary positioning method based on deep learning of the present invention;
[0044] Figure 2 is the schematic diagram of the training process of the pantograph-catenary target area detection model;
[0045] Figure 3 is the schematic diagram of the change curve of the loss function accuracy;
[0046] Figure 4 is the schematic diagram of the verification result. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific implementation manners of the present invention will now be described in detail with reference to the accompanying drawings.
[0048] As Figure 1 shown, the pantograph-catenary positioning method based on deep learning of the present invention includes the following steps:
[0049] Step S1: Obtain train video data containing complex situations such as arcing anchor section joints.
[0050] Step S2: Calibrate the pantograph-catenary contact area, and jointly form a training set and a validation set with the images in normal and complex situations.
[0051] Step S3: Design and initialize the pantograph-catenary target area detection model.
[0052] Step S4: Train the convolutional neural network DETR model with the obtained training set to obtain the trained pantograph-catenary target area detection model.
[0053] Step S5: Input the validation set data into the trained pantograph-catenary target area detection model to verify the recognition effect.
[0054] In step S1, the images of this data set are sourced from the actual operation video of a certain subway. A video of a complete operation section is selected to ensure that the data set covers various situations, including special contact area images such as arcing images in the pantograph-catenary contact area and alternating catenary.
[0055] In the process of constructing the training set and the validation set in step S2, the following steps are included: preprocess the train operation video images to confirm parameters such as the number of video frames and resolution; during the process of calibrating the target area, for the complex situation areas that appear in the video, adopt a manual division strategy, and perform necessary manual annotations in complex situations such as the arcing bright spot covering the pantograph-catenary area and the sudden change in the position where the pantograph passes through the anchor section contact; then use a python script to remove the parking intervals at non-operation times frame by frame to generate the corresponding data set, and then generate the training set and the validation set according to the ratio of 8:2.
[0056] In the model pre - design process of step S3, the DETR network model trained this time includes a backbone network, an encoding layer, a decoding layer, an output head, etc. Initialize it and set the output format; and add positional encoding to the input feature map to help the model perceive the spatial position information of the target. In this model, the backbone network uses ResNet with 101 layers, and the final output is (batch_size, 768, 256); the encoding layer has 6 layers, that is, it iterates 6 times, and the final output of the Encoder is (batch_size, 768, 256); the decoding layer also has 6 layers, the input is 100 learning vectors, and the vector channel number is 256. These 100 vectors will also perform attention mechanism calculations with the output K and V vectors provided by the encoding layer. Finally, the output of the decoding layer is (batch_size, 100, 256).
[0057] In the model training process of step S4, the MMDetetion module in the OpenMMLab framework is used as the experimental environment in this stage. The GPU is RTX 3060 12GB, the development language is Python3.7, the Pytorch version is 1.10, and the CUDA version is 11.3. The overall training process is as Figure 2 shown. It mainly includes the training images after input enhancement and the corresponding labeled data of the pantograph - catenary target area; the backbone network extracts the feature map, the encoder performs global modeling on the feature map, outputs the K and V vectors, and the decoder interacts with the K and V output by the encoder through the query vector to predict the category and bounding box of the target. That is, the backbone network extracts features, and the encoder captures the global context information. In the case of over - exposure or under - exposure, the pixel values in the local area may be completely lost, but the positional encoding retains the spatial arrangement information to assist the model in inferring the spatial structure of the target; calculate the loss function and optimize it. Calculate the loss function based on the Hungarian matching result and the ground truth label, and use the AdamW optimizer for gradient descent. The loss calculation is mainly divided into two steps: Hungarian matching and loss calculation, Figure 3 which is the result of the training loss of the network model. After 200 rounds of training, the loss function has tended to converge; for the evaluation index, when the DETR model is used for pantograph - catenary contact area detection, its accuracy can reach 89.9%.
[0058] The first step is Hungarian matching.
[0059] Based on the Hungarian matching optimization of pantograph - catenary dynamics, aiming at the particularity of pantograph - catenary contact area detection, the loss calculation of the DETR model is deeply integrated with the physical parameters of the pantograph - catenary system. The original predicted category is extended to the key component categories of the pantograph - catenary, and the pantograph - catenary dynamic parameters (vertical displacement of the contact wire, lateral offset) are introduced as additional features.
[0060] The predicted value of the bounding box is defined as a vector containing geometric parameters:
[0061] b i = [x1, y1, Δh, Δθ]
[0062] Wherein, x1 and y1 are the coordinates of the pantograph-catenary contact point, Δh is the height change at the contact position, and Δθ is the tilt change of the pantograph.
[0063] Meanwhile, the L1 and GIOU losses are calculated for the predicted bounding boxes results and the true annotation box results to generate a predicted bounding box loss matrix. Finally, the final loss matrix is formed by taking the weighted sum of the predicted class loss matrix and the predicted bounding box loss matrix. The Hungarian algorithm completes the optimal matching of the predicted boxes and the true boxes.
[0064]
[0065] In the formula: is the predicted loss of the true class, is the predicted loss between the true pantograph-catenary contact geometric parameters and the predicted pantograph-catenary geometric parameters, l is the weight, N is all classes, is the result after Hungarian matching.
[0066] Step 2: Calculate the final loss.
[0067] Initialize a vector containing 100 elements with all values being the background class. According to the matching result of the first step, replace the corresponding positions of this vector with the true classes of the predicted boxes. For example, replace the 14th position with the true class 1 of the first predicted box. Then, calculate the cross-entropy loss between this vector and the predicted results to obtain the final class loss. Similarly, obtain the corresponding bounding box prediction values from the 100 predicted results according to the matching result, and calculate the L1 and GIOU losses with the true annotation boxes to obtain the final bounding box loss. The formula is as follows:
[0068]
[0069] In step S5, input the validation set data into the trained pantograph-catenary target area detection model to verify the recognition effect. To more intuitively display the prediction ability of the model for the pantograph-catenary contact area, two types of interfering pantograph-catenary images are selected for detection. One is the case where the pantograph-catenary contact area is blocked by arcing, and the other is the case where catenary alternation appears in the image. The detection results are as Figure 4 shown. The content marked by the red box is the result of the pantograph-catenary contact area detected by the model. It can be found from the figure that whether it is a normal pantograph-catenary contact image, or an image where the pantograph-catenary contact area is blocked by arcing or there are multiple catenary phenomena, the model can accurately predict the target, indicating that the model has good accuracy and robustness.
[0070] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for pantograph positioning in complex environments based on deep learning, characterized in that: The following steps are involved: Step S1: Acquire train video data containing complex conditions of arcing anchor segment joints; Step S2: Calibrate the pantograph-net contact area, and use the images in normal and complex situations to form a training set and a verification set; Step S3: Design and initialize the pantograph target area detection model; Step S4: training the convolutional neural network DETR model using the obtained training set to obtain a trained bow-net target area detection model; Step S5: input the verification set data into the trained bow-net target area detection model to verify the recognition effect.
2. The complex environment bow-net positioning method based on deep learning according to claim 1 is characterized in that: In step S1, the train video data selects the entire section to ensure that the data set covers various situations, including arcing images in the pantograph-catenary contact area, contact network alternation and special contact area images.
3. The complex environment pantograph positioning method based on deep learning according to claim 1 is characterized in that: The process of constructing the training set and the validation set in step S2 includes the following steps: Pre-process the train running video images and confirm the video frame number and resolution parameters; In the process of calibrating the target area, a manual segmentation strategy is adopted for the complex areas appearing in the video, and necessary manual annotation is performed in the complex situation where the arcing bright spot blocks the pantograph-net area and the pantograph passes through the anchor section contact position mutation; Use the python script to remove the parking intervals during non-operating hours frame by frame to generate the corresponding data set, and then generate the training set and verification set according to the proportion of the data set.
4. The complex environment pantograph positioning method based on deep learning according to claim 1 is characterized in that: In step S3, the trained DETR network model includes a backbone network, an encoding layer, a decoding layer and an output head, which are initialized and the output format is set; and position encoding is added to the input feature map to help the model perceive the spatial position information of the target.
5. The complex environment bow-net positioning method based on deep learning according to claim 1 is characterized in that: In the step S3, the model uses ResNet as the backbone network with 101 layers, and the final output is (batch_size, 768, 256); the encoding layer uses 6 layers, that is, iterates 6 times, and the final Encoder output is (batch_size, 768, 256); the decoding layer uses 6 layers, the input is 100 learning vectors, and the vector channel number is 256. These 100 vectors will also be calculated together with the output K and V vectors provided by the encoding layer for the attention mechanism, and the final decoding layer output is (batch_size, 100, 256).
6. The complex environment bow-net positioning method based on deep learning according to claim 1 is characterized in that: In step S4, the model training method includes: Input the enhanced training image and the corresponding bow-net target area annotation data; The backbone network extracts feature maps, the encoder performs global modeling on the feature maps, and outputs K and V vectors. The decoder interacts with K and V output by the encoder through the query vector to predict the category and bounding box of the target. Obtain and optimize the loss function, calculate the loss function based on the Hungarian matching results and the true labels, and use the AdamW optimizer for gradient descent.
7. The method for complex environment pantograph positioning based on deep learning according to claim 6, characterized in that: In step S4, the Hungarian matching method includes: Based on the Hungarian matching optimization of the pantograph-catenary dynamics, the loss calculation of the DETR model is deeply integrated with the physical parameters of the pantograph-catenary system, the original prediction category is expanded to the category of key components of the pantograph-catenary, and the vertical displacement and lateral offset of the contact line are introduced as additional features; The bounding box prediction is defined as a vector containing the geometric parameters: b i =[x1,y1,Δh,Δθ] Among them, x1, y1 are the coordinates of the pantograph-catenary contact point, Δh is the change in contact position height, and Δθ is the change in pantograph inclination. Perform L1 and GIOU loss calculations on the predicted bounding box results and the true annotation box results to generate a predicted bounding box loss matrix; The final loss matrix is formed by weighted summing the predicted category loss matrix and the predicted bounding box loss matrix; the Hungarian algorithm achieves the optimal match between the predicted box and the true box: Where: is the prediction loss of the true category, is the prediction loss between the real bow-net contact geometry parameters and the predicted bow-net geometry parameters, l is the weight, N is all categories, The result after matching Hungary.
8. The complex environment pantograph positioning method based on deep learning according to claim 7 is characterized in that: In step S4, the method for calculating the final loss includes: Initialize a vector containing 100 elements, all of which are background categories. According to the Hungarian matching results, replace the corresponding position of the vector with the real category of the predicted box. Calculate the cross entropy loss between this vector and the prediction result to get the final category loss; The corresponding bounding box prediction value is obtained from the 100 prediction results through matching results, and the L1 and GIOU losses are calculated with the real annotation box to obtain the final bounding box loss. The formula is as follows:
9. The complex environment pantograph positioning method based on deep learning according to claim 7, characterized in that: In step S5, in order to more intuitively demonstrate the model's prediction ability for the bow-catenary contact area, two interfering bow-catenary images are selected for detection, one is the case where the bow-catenary contact area is blocked by the burning arc, and the other is the case where the contact network appears alternately in the image.