Train passing detection method based on tunnel monitoring video
The tunnel monitoring video is classified through deep learning models and combined with the Kalman filtering algorithm to estimate the state, which solves the problems of low accuracy and high cost of monitoring of trains in the prior art, and achieves higher accuracy and wider applicability.
Patent Information
- Application Number
- CN202311548238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, when the train passes in a monitoring tunnel, there are problems such as high cost, low accuracy and sensitivity to image quality, especially when multiple trains pass at the same time or trains stop.
By constructing the training set, each frame of the tunnel monitoring video is classified using a deep learning model to determine whether a train has passed, and a state estimation is performed in combination with the Kalman filtering algorithm to extract the video segment containing the train.
It improves the accuracy of train passing detection, reduces the rate of misjudgment, is suitable for various types of tunnel surveillance video, saves costs and has wide applicability.
Smart Images

Figure CN120020905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of video anomaly detection and subway operation guarantee, and specifically relates to a train passing detection method based on tunnel monitoring video. Background Art
[0002] Judging whether a train passes through by means of tunnel monitoring video can help monitor the running conditions of the train, and can timely discover possible dangerous situations, thereby effectively improving the safety of the train. In addition, through the monitoring video, the passing time period of the train can be better understood. On the one hand, the passing time of the train can be analyzed to improve the running efficiency of the train. On the other hand, the information can be processed and provided to users to facilitate the lives of residents.
[0003] At present, the technologies for monitoring the passing of trains in tunnels include contact monitoring technologies and non-contact detection technologies. Among them, the contact monitoring technologies include installing track circuits, or installing electronic tags on trains and identifying them through ground detection devices. Both of these methods require direct contact transformation of the track or the train, which is relatively difficult to construct, has a high cost, and requires repeated confirmation of safety.
[0004] For non-contact monitoring technologies, there is currently a detection technology for video image processing. For example, in Chinese Patent CN105303575B, a method of using video image processing is used for train passing detection. The collected images are analyzed for tracks and optical flow, and it is identified that there are moving objects passing through the track. This method only obtains whether a train passes through through the image processing of tunnel monitoring video, starting from the gray value, and conducts train passing detection. However, this method relies on high-precision image recognition technology. If the image quality is not high or is interfered, such as by light, shadow, dust, etc., it may affect the accuracy of the recognition result. If the situation when the train passes through is relatively complex, such as multiple trains passing through simultaneously, or the train stopping on the track, etc., it may increase the difficulty and error rate of recognition. In Chinese Patent CN114253842A, simulation data is obtained by simulating various obstacles and passing situations in rail transit, and then real data is obtained through real train operation videos and records. Finally, the performance parameters of the train intelligent detection system are obtained through the method of deep learning image recognition. However, this method not only requires train video records, but also requires simulation experiments, and the required cost is relatively high. Summary of the Invention
[0005] The purpose of the present invention is to propose a more convenient, cost-saving, and higher-precision train passing detection method according to the above-mentioned deficiencies. The present invention transforms the train passing detection method into a binary classification problem of whether there is a train in each frame of video.
[0006] To achieve the above object, the present invention provides a train passing detection method based on tunnel monitoring videos, including:
[0007] Construct a training set: Collect historical videos, manually annotate each frame in the collected historical videos, extract frames from the historical videos through code, where each frame of the video indicates the presence or absence of a train, obtaining multiple training video frames and forming a training set;
[0008] Select hyperparameters and train a deep learning model: Input all the video frames in the constructed training set into multiple deep learning models, select hyperparameters and train the multiple deep learning models, obtain and save the deep learning model with the highest accuracy, where the accuracy is the comprehensive evaluation F-Score of accuracy rate and recall rate;
[0009] Perform target video classification and state estimation: Obtain a target video, successively put the first frame extracted per second of the target video into the saved deep learning model for classification, obtaining a classification result and classification probability, and perform state estimation using the Kalman filter algorithm;
[0010] Output the result: According to the content of the Kalman filter algorithm, the video segment containing the train is the video segment of the train passing, and the video segment is multiple consecutive video frames.
[0011] In one embodiment, the manually annotating each frame in the collected historical videos, extracting frames from the historical videos through code to obtain multiple training video frames and form a training set includes:
[0012] Judge whether there is a train passing in each frame of the historical video. If there is no train passing, assign a label of 0 to this frame. If there is a train passing, assign a label of 1 to this frame;
[0013] Extract frames from the historical videos through code to form a binary classification dataset with labels of 0 and 1.
[0014] In one embodiment, the manual annotation process is stored and managed through a CSV file, and the CSV file is used for frame extraction, reading, and writing.
[0015] In one embodiment, the inputting all the video frames in the constructed training set into multiple deep learning models for training and saving the deep learning model with the highest accuracy includes:
[0016] Input all the video frames in the constructed training set into each deep learning model, respectively judge whether each video frame contains a train, obtaining a classification result and classification probability;
[0017] Calculate the accuracy of each deep learning model based on the obtained classification results and classification probabilities, and save the deep learning model with the highest accuracy.
[0018] In one embodiment, the precision and recall are calculated based on the values in the confusion matrix. The calculation methods for accuracy and recall are as follows:
[0019] Precision = TP / (TP + FP), Recall = TP / (TP + FN);
[0020] Among them, the precision is the proportion of correctly classified positive samples TP among all samples classified as positive, and FP is the negative samples predicted as positive;
[0021] The recall is the proportion of correctly classified positive samples TP among all actually positive samples, and FN is the positive samples predicted as negative.
[0022] In one embodiment, the comprehensive evaluation F-Score is calculated based on the precision and recall. The calculation method for the comprehensive evaluation F-Score is as follows:
[0023] F-Score = (2 * Precision * Recall) / (Precision + Recall).
[0024] In one embodiment, the deep learning model with the highest accuracy is the Residual Neural Network ResNet34 model. The accuracy of the Residual Neural Network ResNet34 model is 0.942. The structure of the Residual Neural Network ResNet34 model includes:
[0025] The first part: The first part performs a convolution operation with a convolution kernel of 7×7 and a stride of 2×2 on the input video frames, and performs a max pooling operation, and outputs the obtained data;
[0026] The second part, connected to the first part, includes:
[0027] The first sub-part: The first sub-part receives the data output by the first part. The first sub-part includes 3 residual blocks. The number of convolution kernels in the first sub-part is 64. It performs a convolution operation with a convolution kernel of 3×3 on the received data and outputs the obtained data;
[0028] The second sub-part: The second sub-part receives the data output by the first sub-part. The second sub-part includes 4 residual blocks. The number of convolution kernels in the second sub-part is 128. It performs a convolution operation with a convolution kernel of 3×3 on the received data and outputs the obtained data;
[0029] The third sub - part receives the data output by the second sub - part. The third sub - part includes 6 residual blocks, and the number of convolution kernels in the third sub - part is 256. It performs a convolution operation with a 3×3 convolution kernel on the received data and outputs the obtained data.
[0030] The fourth sub - part receives the data output by the third sub - part. The fourth sub - part includes 3 residual blocks, and the number of convolution kernels in the fourth sub - part is 512. It performs a convolution operation with a 3×3 convolution kernel on the received data and outputs the obtained data.
[0031] The third part is connected to the second part. The third part receives the data output by the fourth sub - part, performs average pooling on the received data, passes through a flattening layer, and finally performs a fully connected operation to output classification probabilities.
[0032] In one embodiment, downsampling operations are respectively performed on the first residual blocks of the second sub - part, the third sub - part, and the fourth sub - part through a 1×1 convolution operation of the convolution kernel, so that the number of channels of the second sub - part, the third sub - part, and the fourth sub - part is consistent with the number of channels of its previous sub - part.
[0033] In one embodiment, the Kalman filter algorithm is used for state estimation as follows: The classification results and classification probabilities obtained per second are input into the Kalman filter algorithm to perform state estimation on the target video.
[0034] Among them, the establishment of the Kalman filter algorithm includes the following steps:
[0035] Step 1: Establish the state variable equation of the Kalman filter:
[0036]
[0037] Among them, x k+1 is the posterior state estimation value of predicting whether to pass the train at the (k + 1)-th second of the video, A is the state transition coefficient, and B is the coefficient that converts the input into the state.
[0038] Step 2: Predict the estimated covariance equation and calculate the prior probability:
[0039] P k+1 = P k + Q
[0040] Among them, P k+1 is the uncertainty, and Q is the covariance of the system process.
[0041] Step 3: Update the state of the system according to the observation value:
[0042] x k = x k + K k (zk -Hx k )
[0043] where z k is the classification probability of the first frame at k seconds, H is the conversion coefficient from the state variable to the observation quantity, set to 1, K k is the Kalman gain at k seconds, specifically:
[0044]
[0045] where R is the measurement noise;
[0046] Step 4: Estimate the future state of the system based on the updated state, where the updated estimated covariance equation formula is:
[0047] P k =(1 - K k )P k .
[0048] In one embodiment, A is set to 1, B is set to 0, Q is set to 0.1, H is set to 1, and R is set to 0.5.
[0049] The train passing detection method based on tunnel monitoring video of the present invention has the following beneficial effects:
[0050] Compared with the detection means of traditional image processing methods, the present invention has a higher accuracy. Compared with the advantages of pure video deep learning, it can not only classify video frame images through deep learning technology, improve the classification accuracy, and reduce the misjudgment rate, but also use the Kalman filter algorithm to process video segments, and can more accurately extract the time period containing the train. In addition, the present invention can be applied to various types of tunnel monitoring videos, saving costs and having wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flow chart of the train passing detection method based on tunnel monitoring video according to an embodiment of the present invention;
[0052] Figure 2 is a schematic diagram of the confusion matrix of the training model in the train passing detection method based on tunnel monitoring video according to an embodiment of the present invention;
[0053] Figure 3 is a schematic diagram of the structure of the ResNet34 model used in the train passing detection method based on tunnel monitoring video according to an embodiment of the present invention;
[0054] Figure 3 In, 1 is the first part, 2 is the second part, 3 is the third part, 21 is the first sub - part, 22 is the second sub - part, and 23 is the third sub - part. Detailed implementation manners
[0055] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and are not used to limit the invention.
[0056] The present invention provides a train passing detection method based on tunnel monitoring videos, including:
[0057] Constructing a training set: Collecting historical videos, manually annotating each frame in the collected historical videos, and extracting frames from the historical videos through code. Each frame of the video indicates the presence or absence of a train, so as to obtain a plurality of training video frames and form a training set;
[0058] Selecting hyperparameters and training a deep learning model: Inputting all the video frames in the constructed training set into a plurality of deep learning models, selecting appropriate hyperparameters and training the plurality of deep learning models, obtaining and saving the deep learning model with the highest accuracy. The accuracy is the comprehensive evaluation F-Score of the accuracy rate and the recall rate;
[0059] Performing target video classification and state estimation: Obtaining a target video, successively putting the first frame extracted from the target video per second into the saved deep learning model for classification, obtaining a classification result and a classification probability, and performing state estimation using the Kalman filtering algorithm;
[0060] Outputting a result: According to the content of the Kalman filtering algorithm, the video segment containing the train is the video segment of the train passing. The video segment is a plurality of consecutive video frames.
[0061] Further, manually annotating each frame in the collected historical videos, and extracting frames from the historical videos through code, so as to obtain a plurality of training video frames and form a training set, including:
[0062] Judging whether there is a train passing in each frame of the historical video. If there is no train passing, assign a label of 0 to this frame. If there is a train passing, assign a label of 1 to this frame;
[0063] Extracting frames from the historical videos through code to form a binary classification data set with labels of 0 and 1.
[0064] It should be understood that the classification in the present application is: Assigning a label of 0 to the video frame without a train passing, and assigning a label of 1 to the video frame with a train passing.
[0065] Furthermore, the manual annotation process is stored and managed through a CSV file, which is used for frame extraction, reading, and writing. A CSV (Comma-Separated Values) file is a common text file format used to store structured data, where the data is separated by commas and stored in text form. CSV files are typically used to export and import tabular data to different computer programs. Manual annotation using a CSV file: The CSV file contains label information for video frames, which are usually 0 or 1, indicating whether certain specific attributes, objects, or features appear in each video frame. In this embodiment, the CSV file is created by a manual annotator, who manually adds corresponding labels to each frame of the video. Write code to read this CSV file in order to load the label information into the computer's memory. A video processing library (such as OpenCV) can be used to open the video file and extract video frames from it. These frames can be extracted as image files or image data for subsequent processing. During video frame extraction, the label data in the CSV file will be used to assign corresponding labels to each extracted video frame. This usually involves associating the labels with each frame and retaining these labels in the processed dataset. Code can also be used to save the frames with labels as image files for use in subsequent tasks, such as training a deep learning model. In summary, use the label information in the CSV file, associate it with the video frames, and perform frame extraction, reading, and writing operations to construct a labeled dataset. This dataset can be used to train a machine learning model to identify specific objects, scenes, or features in the video, depending on the meaning of the labels and the requirements of the task.
[0066] Furthermore, input all the video frames in the constructed training set into multiple deep learning models for training, and save the deep learning model with the highest accuracy, including:
[0067] Input all the video frames in the constructed training set into each deep learning model, and respectively determine whether each video frame contains a train to obtain the classification result and classification probability;
[0068] Based on the obtained classification result and classification probability, calculate the accuracy of each deep learning model, and save the deep learning model with the highest accuracy.
[0069] Furthermore, the precision and recall are calculated based on the values in the Confusion Matrix, as shown in Figure 2Predicted labels represent the labels or classifications assigned by a machine learning or classification model to a set of data points. These labels are the model's predictions for each data point in the dataset. In a classification task, predicted labels are typically assigned based on the characteristics and patterns learned by the model from the training data, considering the nature of the input data points. In the binary classification of the present invention, the predicted labels can be 0 or 1, representing the categories of whether a train passes or not, respectively. The accuracy and performance of a machine learning model are usually evaluated by comparing the predicted labels with the true labels in a validation or test dataset. True labels refer to the actual or real classification labels of data points in a machine learning or classification task. In this embodiment, these labels are determined by manual annotation. The calculation methods for precision and recall are as follows:
[0070] Precision = TP / (TP + FP), Recall = TP / (TP + FN);
[0071] Among them, the precision is the proportion of correctly classified positive samples TP among all samples classified as positive. FP is the negative samples predicted as positive;
[0072] The recall is the proportion of correctly classified positive samples TP among all samples that are actually positive. FN is the positive samples predicted as negative.
[0073] In one embodiment, the comprehensive evaluation F-Score is calculated based on precision and recall. The calculation method for the comprehensive evaluation F-Score is as follows:
[0074] F-Score = (2 * Precision * Recall) / (Precision + Recall).
[0075] In this embodiment, preferably, the deep learning model with the highest accuracy is the Residual Neural Network ResNet34 model, and its accuracy is 0.942. As Figure 3 shown, the structure of the Residual Neural Network ResNet34 model includes:
[0076] The first part 1 performs a convolution operation with a convolution kernel of 7×7 and a stride of 2×2 on the input video frame, and then performs a max-pooling operation to output the obtained data.
[0077] The second part 2, connected to the first part 1, includes: a first sub - part 21, a second sub - part 22, a third sub - part 23, and a fourth sub - part 24. The first sub - part 21 receives the data output by the first part 1. The first sub - part 21 includes 3 residual blocks, and the number of convolutional kernels in the first sub - part 21 is 64. It performs a convolution operation with a 3×3 convolutional kernel on the received data and outputs the obtained data. The second sub - part 22 receives the data output by the first sub - part 21. The second sub - part 22 includes 4 residual blocks, and the number of convolutional kernels in the second sub - part 22 is 128. It performs a convolution operation with a 3×3 convolutional kernel on the received data and outputs the obtained data. The third sub - part 23 receives the data output by the second sub - part 22. The third sub - part 23 includes 6 residual blocks, and the number of convolutional kernels in the third sub - part 23 is 256. It performs a convolution operation with a 3×3 convolutional kernel on the received data and outputs the obtained data. The fourth sub - part 24 receives the data output by the third sub - part 23. The fourth sub - part 24 includes 3 residual blocks, and the number of convolutional kernels in the fourth sub - part 24 is 512. It performs a convolution operation with a 3×3 convolutional kernel on the received data and outputs the obtained data.
[0078] The third part 3, connected to the second part 2, receives the data output by the fourth sub - part 24, performs average pooling on the received data, passes through a flattening layer, and finally performs a fully - connected operation to output classification probabilities.
[0079] Furthermore, for the input of the first residual block of each part (i.e., Figure 3 each dotted - line part in
[0080] ), since the number of channels in the previous part is inconsistent with the number of channels in this part, a 1×1 convolution operation is implicitly used to form downsampling, increase the number of channels, so that the number of channels of the input of this part is the same as the number of channels of this part, so that pixel superposition on the same channels can be performed. Specifically, downsampling operations are performed on the first residual blocks of the second sub - part, the third sub - part, and the fourth sub - part respectively through a 1×1 convolution operation with a convolutional kernel, so that the number of channels of the second sub - part, the third sub - part, and the fourth sub - part is the same as the number of channels of its previous sub - part respectively.Furthermore, the Kalman filter algorithm is used to continuously update the estimation of the system state to adapt to the observed data and provide a more accurate state estimation considering uncertainties. In this specific application, the Kalman filter is used to predict whether the target video contains a train by estimating the state variables and updating the covariance equation. The state estimation using the Kalman filter algorithm is as follows: The classification results and classification probabilities obtained per second are input into the Kalman filter algorithm for state estimation of the target video. The Kalman filter algorithm will use these inputs to continuously update the state estimation to adjust the predicted value of the state according to the latest information. In this embodiment, the specific steps are as follows:
[0081] Step 1: Establish the state variable equation of the Kalman filter:
[0082]
[0083] where x k+1 is the posterior state estimation value for predicting whether a train passes at the (k + 1)-th second of the video, A is the state transition coefficient, and B is the coefficient for converting the input into the state;
[0084] Step 2: Predict the estimated covariance equation and calculate the prior probability:
[0085] P k+1 = P k + Q
[0086] where P k+1 is the uncertainty and Q is the covariance of the system process;
[0087] Step 3: Update the state of the system according to the observed value
[0088] x k = x k + K k (z k - Hx k )
[0089] where z k is the classification probability of the first frame at the k-th second, H is the conversion coefficient from the state variable to the observed quantity, assumed to be 1, and K k is the Kalman gain at the k-th second, specifically:
[0090]
[0091] where R is the measurement noise;
[0092] Step 4: Estimate the future state of the system according to the updated state. The formula for updating the estimated covariance equation is:
[0093] P k = (1 - Kk )P k 。
[0094] In this embodiment, it is preferred to take the first frame of the first and second seconds as the starting point for recursive calculation. The Kalman filter algorithm represents the state as whether the target video contains a train. By continuously updating the state estimate, the Kalman filter algorithm integrates information from the classification model to provide a more accurate estimate of the train appearance state. Therefore, the Kalman filter is used here to estimate the state of the target video based on the classification results and classification probabilities, that is, whether there is a train in the video and the state change over time.
[0095] In this embodiment, it is preferred that A is set to 1, B is set to 0, Q is set to 0.1, H is set to 1, and R is set to 0.5.
[0096] The train passing detection method based on tunnel monitoring video of the present invention has the following beneficial effects:
[0097] Compared with the detection means of traditional image processing methods, the present invention has a higher accuracy rate. Compared with the advantages of pure video deep learning, it can not only classify video frame images through deep learning technology, improve the classification accuracy rate, and reduce the misjudgment rate, but also use the Kalman filter algorithm to process video segments, and can more accurately extract the time period containing trains. In addition, the present invention can be applied to various types of tunnel monitoring videos, saving costs and having wide applicability.
[0098] In the description of the embodiments of the present application, it should be noted that unless otherwise stated and limited, the term "connection" should be understood in a broad sense. For example, it can be an electrical connection, or the connection inside two components. It can be directly connected, or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meaning of the above terms can be understood according to specific situations.
[0099] It should be noted that the terms "first", "second", "third", and "fourth" involved in the embodiments of the present application are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first", "second", "third", and "fourth" can be interchanged in a specific order or sequence under allowable circumstances. It should be understood that the objects distinguished by "first", "second", "third", and "fourth" can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here.
[0100] Although the above methods are illustrated and described as a series of acts for simplicity of explanation, it should be understood and appreciated that the methods are not limited by the order of the acts, since according to one or more embodiments, some acts may occur in a different order and / or concurrently with other acts that are illustrated and described herein or that are not illustrated and described herein but would be understood by those of ordinary skill in the art.
[0101] The above-described embodiments are only further illustrations of the present invention and do not limit the present invention in any other form. The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding modifications and changes according to the present invention, but these corresponding modifications and changes should all fall within the protection scope of the present invention.
Claims
1. A train passing detection method based on tunnel monitoring video, characterized in that: include: Construct a training set: collect historical videos, manually annotate each frame in the collected historical videos, extract frames from the historical videos through code, and each frame of the video indicates whether there is a train or not. Multiple training video frames are obtained to form a training set; Select hyperparameters and train deep learning models: input all video frames in the constructed training set into multiple deep learning models, select hyperparameters and train the multiple deep learning models, and obtain and save the deep learning model with the highest accuracy, where the accuracy is the comprehensive evaluation F-Score of the accuracy and recall rate; Perform target video classification and state estimation: Obtain the target video, and put the first frame extracted from the target video every second into the saved deep learning model for classification, obtain the classification result and classification probability, and use the Kalman filter algorithm for state estimation; Output result: According to the content of the Kalman filter algorithm, the video segment containing the train is obtained, that is, the video segment where the train passes, and the video segment is a plurality of continuous video frames.
2. The train passing detection method based on tunnel monitoring video according to claim 1 is characterized in that: The method of manually labeling each frame of the collected historical video and extracting frames from the historical video through code to obtain multiple training video frames and form a training set includes: Determine whether a train passes through each frame in the historical video. If no train passes through, assign a label of 0 to the frame. If a train passes through, assign a label of 1 to the frame. The code extracts frames from historical videos to form a binary classification dataset with labels of 0 and 1.
3. The train passing detection method based on tunnel monitoring video according to claim 1 is characterized in that: The manual annotation process is stored and managed through a CSV file, and the CSV file is used for frame extraction, reading and writing.
4. The train passing detection method based on tunnel monitoring video according to claim 1 is characterized in that: The method of inputting all video frames in the constructed training set into multiple deep learning models for training, and saving the deep learning model with the highest accuracy, includes: Input all the video frames in the constructed training set into each deep learning model, determine whether each video frame contains a train, and obtain the classification result and classification probability; According to the obtained classification results and classification probabilities, the accuracy of each deep learning model is calculated, and the deep learning model with the highest accuracy is saved.
5. The train passing detection method based on tunnel monitoring video according to claim 1 is characterized in that: Precision and recall are calculated based on the values in the confusion matrix. The calculation methods of precision and recall are as follows: Precision = TP / (TP+FP), Recall = TP / (TP+FN); The accuracy is the ratio of correctly classified positive samples TP to all samples classified as positive, and FP is the negative samples predicted as positive. The recall rate is the ratio of correctly classified positive samples TP to all actually positive samples, and FN is the positive samples predicted to be negative.
6. The train passing detection method based on tunnel monitoring video according to claim 5 is characterized in that: The comprehensive evaluation F-Score is calculated based on precision and recall. The calculation method of the comprehensive evaluation F-Score is as follows: F-Score = (2*precision*recall) / (precision+recall).
7. The train passing detection method based on tunnel monitoring video according to claim 1 is characterized in that: The deep learning model with the highest accuracy is the residual neural network ResNet34 model. The accuracy of the residual neural network ResNet34 model is 0.
942. The structure of the residual neural network ResNet34 model includes: The first part performs a convolution operation with a kernel of 7×7 and a stride of 2×2 on the input video frame, and performs a maximum pooling operation, and outputs the obtained data; The second part, connected with the first part, includes: The first sub-part receives the data output by the first part, includes 3 residual blocks, has 64 convolution kernels, performs a convolution operation with a convolution kernel of 3×3 on the received data, and outputs the obtained data; The second sub-part receives the data output by the first sub-part, includes 4 residual blocks, has 128 convolution kernels, performs a convolution operation with a convolution kernel of 3×3 on the received data, and outputs the obtained data; The third sub-part receives the data output by the second sub-part, includes 6 residual blocks, has 256 convolution kernels, performs a convolution operation with a convolution kernel of 3×3 on the received data, and outputs the obtained data; The fourth sub-part receives the data output by the third sub-part, includes 3 residual blocks, has 512 convolution kernels, performs a convolution operation with a convolution kernel of 3×3 on the received data, and outputs the obtained data; The third part is connected to the second part. The third part receives the data output by the fourth sub-part, performs average pooling on the received data, passes through the straightening layer, and finally performs a full connection to output the classification probability.
8. The method for detecting train passing through a tunnel based on tunnel monitoring video according to claim 7, characterized in that: The first residual blocks of the second sub-part, the third sub-part and the fourth sub-part are respectively downsampled by a convolution operation with a convolution kernel of 1×1, so that the number of channels of the second sub-part, the third sub-part and the fourth sub-part is consistent with the number of channels of the previous sub-part.
9. The method for detecting train passing through a tunnel based on tunnel monitoring video according to claim 1, characterized in that: The state estimation using Kalman filter algorithm is as follows: the classification results and classification probabilities obtained every second are input into the Kalman filter algorithm to estimate the state of the target video; Among them, the establishment of the Kalman filter algorithm includes the following steps: Step 1: Establish the state variable equation of Kalman filter: Among them, x k+1 is the posterior state estimate of whether the video passes the train at the k+1 second, A is the state transfer coefficient, and B is the coefficient that converts the input into the state; Step 2: Predict the estimated covariance equation and calculate the prior probability: P k+1 =P k +Q Among them, P k+1 is the uncertainty, Q is the covariance of the system process; Step 3: Update the state of the system based on the observed value: x k =x k +K k (z k -Hx k ) Among them, z k is the classification probability of the first frame at k seconds, H is the conversion coefficient from state variables to observations, set to 1, K k is the Kalman gain at k seconds, specifically: Where R is the measurement noise; Step 4: Estimate the future state of the system based on the updated state, where the updated estimated covariance equation formula is: P k =(1-K k )P k 。 10. The method for detecting train passing through tunnel based on tunnel monitoring video according to claim 9, characterized in that: A is set to 1, B is set to 0, Q is set to 0.1, H is set to 1, and R is set to 0.5.
Citation Information
Patent Citations
Video detection method for identifying rail train passage
CN105303575B
Method and device for testing train intelligent detection system based on image recognition
CN114253842A