Single turnout status detection system, detection method, and semi-automatic data annotation method

Through a single-opening switch state detection system, combined with deep learning and traditional image processing technology, the problems of large calculation, difficulty in identification and high hardware cost are solved, and accurate switch state recognition and reduced data labeling workload are achieved to ensure the safe operation of the train.

CN115601558BActive Publication Date: 2025-07-18JILIN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211305472.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-07-18
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In the prior art, single switch status detection is large, difficult to identify, high hardware cost and large data labeling workload, resulting in inaccurate identification of switch status and inability to provide effective line information for train drivers.

Method used

A single-open switch state detection system is adopted, including an image acquisition module, a data preprocessing module and a switch state discrimination module. Combined with deep learning network and traditional image processing technology, track lines are extracted and screened through regional growth strategies to reduce the search space. The traditional method is used to verify it again after the initial detection of the deep learning network, and screen and identify it in combination with the track area where the train is located.

Benefits of technology

It improves the accuracy of switch status recognition, reduces hardware cost and calculation amount, reduces data labeling workload, and provides accurate switch status information to ensure the safe operation of the train.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601558B_ABST
    Figure CN115601558B_ABST
Patent Text Reader

Abstract

The object of the present invention is to provide a single turnout state detection system, a detection method, and a semi-automatic data annotation method, including: an image acquisition module for providing raw data for the test process; a turnout data set construction module for providing raw data for the training and verification processes; a data preprocessing module for identifying the track area of the raw sample data, further reducing the target search space, and providing clues for subsequent screening of the turnout targets on the track where the train is located; and a turnout state discrimination module for improving the accuracy of turnout state recognition. The present invention solves the problems in the prior art such as large computational amount for single turnout state detection, difficult turnout state recognition, high cost of recognition hardware, and large workload of data annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rail transit, and relates to a single turnout state detection system and method, and a semi-automatic data annotation method. Background Art

[0002] With the rapid development of rail transit, the operation safety of rail vehicles has received extensive attention. As a key component of railways, turnouts control the intersection and connection between different tracks. Incorrect turnout position opening or failure of the dispatching system will cause serious safety hazards. Therefore, turnout state recognition is of great significance for the safe operation of trains. Specifically: (1) Turnout state recognition can provide important clues for the train's driving route recognition in the turnout area, and even avoid rail traffic accidents in the event of a failure of the dispatching system. (2) The turnout area is a high-risk area for track damage. The switch rail structure of the turnout often bears high impact loads. In addition, due to problems such as non-optimal contact between the switch rail and the wheel flange, and the angular clearance between the switch rail and the stock rail, it is extremely easy to cause wear and plastic deformation in the turnout area. Turnout state recognition can also extract the corresponding turnout area for track fault detection tasks. (3) Turnouts are important reference objects for track lines. Turnout recognition can assist trains in positioning. The selective positioning of rail vehicles on tracks is an unsolved task. The positioning accuracy based on the navigation system cannot meet all requirements, and the navigation system signal is not always available. By combining turnout recognition and data Figure 1 used together, the self-positioning quality of rail vehicles can be improved.

[0003] Currently, the turnout scene recognition method is to use a single-rail segmentation model to extract the single-rail area with a bifurcation point, and further detect the turnout center point to achieve the recognition of the rail turnout. However, this method can only recognize the turnout area and cannot identify the turnout state. There is also a turnout recognition method based on traditional image processing. This method extracts track features through grayscale conversion, median filtering, and edge detection methods. Then, through the Hough transform, the main track and the siding are extracted using features such as inclination angle and length. Finally, linear correlation analysis is used to achieve turnout recognition. However, this method requires manual design and extraction of track features, and the algorithm has a large computational amount. In addition, this method only classifies the turnout types and does not detect the real-time state of the turnout, and cannot provide completely effective line information for train drivers.

[0004] The existing turnout state recognition device uses a turnout sensor to identify the position of the switch rail and judge the turnout state, uses a wheel sensor to monitor the driving signal, and finally designs a warning strategy by combining the above two pieces of information. This method directly uses sensors to identify the turnout state, ensuring the reliability of the data. However, this method requires the installation of corresponding recognition devices in each turnout area to be recognized, resulting in high implementation costs and poor applicability.

[0005] The turnout structure in the railway environment mainly consists of single - turnout switches. In addition, the annotation workload of large - scale turnout datasets is relatively large, posing great challenges to the training and testing of related algorithms. Therefore, the present invention proposes a single - turnout switch status detection system, a detection method, and a semi - automatic data annotation method. Summary of the Invention

[0006] The purpose of the present invention is to provide a single - turnout switch status detection system, which solves the problems of high hardware cost for identification and inaccurate detection results in the prior art.

[0007] The present invention also provides a detection method for a single - turnout switch status detection system, which solves the problems of large computational workload for single - turnout switch status detection, difficulty in identifying turnout status, and inaccurate extraction of the track area in the prior art.

[0008] The present invention also provides a semi - automatic data annotation method for a single - turnout switch status detection system, which solves the problem of large data annotation workload.

[0009] The technical solution adopted by the present invention is a single - turnout switch status detection system, including: an image acquisition module for providing original image data for the test process; a turnout dataset construction module for providing original data for the training and verification processes; a data pre - processing module for identifying the track area of the original sample data, further reducing the target search space, and providing clues for subsequent screening of the turnout target on the track where the train is located; a turnout status discrimination module for improving the accuracy of turnout status identification.

[0010] A detection method for a single - turnout switch status detection system includes the following steps:

[0011] S1, making a turnout dataset;

[0012] S2, performing image pre - processing on the dataset images;

[0013] S3, constructing a track turnout status discrimination module;

[0014] S4, using the pre - processed turnout images and corresponding labels as input data, and using the constructed turnout status discrimination model as a detection tool to obtain turnout detection results;

[0015] S5, combining the track where the train is located and the turnout status discrimination result to screen and identify the turnout on the track where the train is located, and output the final turnout status identification result.

[0016] A semi - automatic data annotation method for a single - turnout switch status detection system specifically includes the following steps:

[0017] Step 1: Use the single turnout state detection system to collect turnout images on different lines and under different weather conditions, and divide the collected images into two parts: a manual annotation set and an automatic annotation set. Use the VIA annotation tool to annotate the manual annotation set, and randomly divide the manual annotation set into three sub-datasets of train, trainval, and test in a ratio of 8:1:1. The manual annotation set has at least 8,000 frames of images and ensures that there are no less than 10,000 turnout state instances in each category.

[0018] Step 2: Use the S31 turnout state discrimination deep learning model to train the labeled data set to obtain the optimal turnout recognition model of the model; the specific process is as follows:

[0019] The turnout state discrimination model in the single turnout state detection system is trained and tested using the manually labeled set in step 1. The number of training times is recorded as R. i , when R i When the accuracy and recall rate of turnout state detection are both above 0.85, the model parameters of the turn are used as available model parameters, and the model parameters of the turn and the test results are packaged together as the R i The available turnout state discrimination model of the first round is stored in the available model list; otherwise, the model parameters of this round are discarded, and the model is trained and tested for the next round. When the available model list is not empty and there is no optimization effect in the results of three consecutive rounds of tests, the operation is stopped, and the best performance of the available turnout state discrimination model is selected from the available model list as the optimal turnout state discrimination model; when the available model list is empty and the number of training times reaches the threshold R i阈 When R i阈 The default value is 300. At this time, the single turnout status detection system cannot meet the detection requirements, so the calculation process is stopped and the system is optimized;

[0020] Step 3: randomly shuffle the images in the automatic annotation set and group them into units of 1000 images. Groups with less than 1000 images are counted as 1000. Randomly select a group of data and use the optimal turnout state discrimination model obtained in step 2 to perform turnout state detection to obtain the label information of the images in the automatic annotation set. The selected grouped data is removed from the automatic annotation data set.

[0021] Step 4: manually check whether the automatically generated annotation files meet the annotation requirements, optimize the unreasonable annotations, remove the images without annotation information, and add the qualified images to the manual annotation set;

[0022] Step 5, select another group of images from the remaining group of images in the automatic annotation set and repeat steps 2, 3, and 4 until the automatic annotation set becomes empty and all images in the automatic annotation set are annotated, thus ending the annotation process.

[0023] The beneficial effects of the present invention are as follows:

[0024] 1. A method for extracting track regions is proposed. By designing a region growing strategy to extract track lines and screening and pairing the track lines, the accuracy of track region extraction is improved, and the search space of the deep learning network is reduced.

[0025] 2. A turnout state discrimination model combining a deep learning network and traditional image processing techniques is proposed. First, the deep learning network is used to identify the turnout state, and then the traditional image processing method is used to perform secondary detection on difficult samples with low confidence, establishing a turnout state recognition mechanism. Compared with a single detection scheme, the detection mechanism is more transparent and the result reliability is higher.

[0026] 3. A method for screening turnouts on the track where a train is located is proposed. By combining the track region where the train is located and the predicted turnout target box information, it is possible to distinguish the turnout targets on the track where the train is located from the turnout targets on the adjacent track, and identify the turnout targets in different regions according to their importance.

[0027] 4. A semi-automatic data annotation method is proposed. This method can perform initial annotation on images through a deep learning network, and then be corrected manually, avoiding direct manual annotation and reducing the annotation workload. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 is the flowchart of the single turnout state detection system in the embodiment of the present invention.

[0030] Figure 2a is the schematic diagram of the left-turn turnout state in the embodiment of the present invention.

[0031] Figure 2b is the schematic diagram of the right-turn turnout state in the embodiment of the present invention.

[0032] Figure 3 is the flowchart of the data preprocessing in the embodiment of the present invention.

[0033] Figure 4 is the effect diagram of the data preprocessing in the embodiment of the present invention.

[0034] Figure 5 is the schematic diagram of the turnout discrimination model in the embodiment of the present invention.

[0035] Figure 6 It is a schematic structural diagram of the TRCSP_N module in the turnout discrimination model of the present invention.

[0036] Figure 7a It is an effect diagram of the detection result of the turnout by the YOLOv5 algorithm.

[0037] Figure 7b It is an effect diagram of the detection result of the turnout by the algorithm of the present invention.

[0038] Figure 8 It is a flowchart of semi-automatic dataset annotation in an embodiment of the present invention.

[0039] Figure 9 It is a comparison diagram of the effects of automatically generated annotation frames and manually annotated frames in an embodiment of the present invention. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] As Figure 1 shown, the present invention provides a single turnout status detection system, including:

[0042] An image acquisition module, used to provide raw data for the test process;

[0043] A turnout dataset construction module, used to provide raw data for the training and verification processes;

[0044] A data preprocessing module, used to identify the track area of the raw sample data, further reduce the target search space, and provide clues for subsequent screening of the turnout targets on the track where the train is located;

[0045] A turnout status discrimination module, used to improve the accuracy of turnout status recognition.

[0046] The present invention also provides a detection method for a single turnout status detection system, and the specific steps are as follows:

[0047] S1: Make a turnout dataset. Since there is a lack of publicly available turnout datasets, it is necessary to make a turnout dataset for trains in real operating scenarios.

[0048] Further, S1 includes the following sub-steps:

[0049] S11: Install the camera at the center position of the train cab, ensure that the optical axis of the camera is consistent with the positive direction of train operation, and ensure that there are no obstacles blocking within the camera view.

[0050] Preferably, a Hikvision DS-2DC4223IW-D camera is selected. The maximum detection distance of this type of camera is 200 m, which can meet the requirements for turnout status detection. The size of the collected images is uniformly 1080×1920 for easy post-processing.

[0051] S12: Record the video of the track scene collected by the camera under different line and weather conditions during the train operation process, then disassemble the video into frames, and manually screen out the images containing the recognizable single turnout status as the original images of the data set.

[0052] Preferably, to ensure that the data set has sufficient representativeness, data collection should be carried out under different lighting conditions and different railway lines. At least 8000 frames of original images are collected and the number of turnout status instances in each category is not less than 10000.

[0053] S13: Label the selected images using the VIA annotation tool. The annotation categories are divided into two states: left-turn turnout and right-turn turnout, as Figure 2a - Figure 2b shown. The discrimination of the turnout status is judged by identifying the contact relationship between the switch rail and the stock rail. If there is no contact between the left switch rail and the static stock rail, the turnout is considered a left-turn turnout. Similarly, if there is no contact between the right switch rail and the static stock rail, the turnout is considered a right-turn turnout.

[0054] S14: Randomly divide the data set into three sub-data sets: train, trainval, and test, according to the ratio of 8:1:1. The train data set and the trainval data set are used for the training and validation process, and the test data set is used for the test and inference process.

[0055] Since there is a large time correlation and similarity between adjacent turnout images obtained after video collection, in order to ensure the randomness of data division, the labeled image sequence needs to be randomly shuffled before data set division.

[0056] S2: As Figure 3 shown, perform image preprocessing on the data set images.

[0057] Furthermore, S2 includes the following sub-steps:

[0058] S21: Convert the collected RGB images in the data set to grayscale. The RGB image pixel unit consists of three channels: R, G, and B. The grayscale pixel gray is obtained by taking the average value of the three channel values of the RGB image, that is, gray = (R + G + B) / 3. Where R, G, and B are the three channel values of the RGB image respectively.

[0059] S22: Extract the track line information using the region growing method. The region growing algorithm utilizes the significant gray level difference between the track and the track base plane in the switch image, merges regions with similar gray levels, and the finally remaining region is the track line information. The extracted track line is fragmented.

[0060] Furthermore, the specific steps of S22 are as follows:

[0061] S221: Obtain the seed starting point for region growing. The seed starting point is the growth starting point that needs to be predefined in the region growing process. Since the switch position in the switch image often lies in the lower half of the image, to eliminate the redundant trackless regions and ensure the retention of all track regions, the region growing method only processes the region within the range of the y-axis coordinate of the switch image in [1 / 3h, h]. Set 10 seed starting points, which are set in the image using the uniform distribution method. The image positions of the first 5 seed starting points are respectively (w×i / 5, 5 / 9×h), and the image positions of the last 5 seed starting points are respectively (w×(i - 5) / 5, 7 / 9×h), where h is the height of the original switch image, w is the width of the original switch image, and i represents the i-th seed starting point.

[0062] S222: Denote the coordinates of the i-th seed starting point as (sdx i , sdy i ), and obtain the pixel coordinates within the 8-neighborhood of the seed starting point position. The neighborhood pixel coordinates are denoted as (X 邻 , Y 邻 ). It can be known that (X 邻 , Y 邻 ) = (sdx i , sdy i ) + R, where R represents the 8-neighborhood relative coordinate array, and R = [(-1, -1), (-1, 0), (-1, 1), (1, 1), (0, 1), (1, 0), (1, -1), (0, -1)].

[0063] S223: Compare the gray level difference between the i-th seed starting point position and the 8-neighborhood position. When the gray level difference D((sdx i , sdy i )) ≤ gray t , it is considered that the seed starting point position and the corresponding neighborhood position belong to the same gray region. Set the gray levels of the seed starting point position and its corresponding neighborhood position to 255 uniformly, and set the neighborhood position as the new seed starting point. gray t is the gray level difference threshold, and the default value is set to 6. The expression of the gray level difference is as follows:

[0064] D((sdx i , sdy i)) = gray((sdx i , sdy i )) - gray((X 邻 , Y 邻 ))

[0065] Among them, D((sdx i , sdy i )) represents the difference between the gray value of the i-th seed starting point and the corresponding neighborhood, gray((sdx i , sdy i )) represents the gray value at the position of the seed starting point, and gray((X 邻 , Y 邻 ) represents the gray value at the neighborhood position.

[0066] S224: Repeat S223 until there are no available seed starting points, indicating that the corresponding area has completed the "growing process". After the area growth, the remaining area surrounded by the 255 gray value is the track line area.

[0067] S23: Screen and pair the track lines to determine the track area.

[0068] The track lines obtained after S224 often have a lot of noise and false track lines that need to be removed. The Hough transform algorithm is used to extract the track lines. The obtained track lines are represented as follows: The j-th track line is represented by the two endpoint values of the line segment, denoted as (xh j , yh j ), (xt j , yt j ), where (xh j , yh j ) represents the endpoint closer to the x-axis in the track line segment, that is, yh j < yt j . From the endpoint coordinates of the track line, the expression of the j-th track line where y is the y-axis coordinate of the track line and x is the x-axis coordinate of the track line; the slope of the track line The length of the track line in the pixel coordinate system where k j is the slope of the j-th track line and L j is the length of the j-th track line.

[0069] Furthermore, the specific steps of S23 are as follows:

[0070] S231: Remove false track lines: When k j ∈[-∞, -1] ∪ [1, ∞] and L j > 60 pixels, keep the track line, otherwise delete the track line. Where k jThe value range of is the optimal value obtained through manual adjustment and repeated testing; L j The value range of is the empirical value obtained after processing a large number of track images. When L j > 60 pixels, a better track line removal effect can be obtained on the vast majority of track images.

[0071] S232: Group the track lines according to the track line slope.

[0072] First, calculate the slope of the remaining track lines, sort all the track lines in descending order according to the absolute value of the slope, and then select n track lines as the pairing criteria by equidistant sampling from the sorted track lines. The default value of n is of the total number of all track lines. Denote the slopes of the n tracks as k n , traverse all the remaining retained track lines, and denote the slope of the track line in each traversal as k g . When ||k n |-|k g || < k t , add the track line corresponding to k g to the track line group represented by k n . k t is the threshold of the absolute value difference of the slopes, and the default value is set to 0.1.

[0073] S233: Repeat S232 for the remaining ungrouped track lines until all track lines are added to the groups, end the grouping process, and directly discard the groups with only one track line.

[0074] S234: Merge different track line segments of the same track, and perform track merging and pairing.

[0075] Divide the track lines within the group into two categories according to the positive and negative of the slope, and denote them as category a and category b groups respectively. Do not merge the groups with only one track line. Denote the number of track lines in each group that meets the merging requirements as N (N≥2), and perform track merging and track line pairing respectively according to the following steps:

[0076] The specific process of track merging is as follows:

[0077] Process the track lines of category a and category b groups respectively as follows: Calculate the average slope k 均 of all the track lines within the track line category group, fine-tune the slopes of the N track lines to be unified as k 均 . For the N track lines, referring to the track line expression in S23, transform the track line expression into: y - k 均 x = y h - k 均 ·xh , where y h is the y-axis coordinate of a certain point on the track line, and x h is the x-axis coordinate corresponding to the y-axis coordinate of a certain point on the track line; let C = y h - k 均 ·x h , where C represents the intercept of the corresponding track line expression; sort the N track lines in descending order according to the C value, and calculate the distance between every two adjacent track lines where p ∈ N - 1; in the above formula, D p is the distance between the P-th track line after descending order sorting and the (P + 1)-th track line sorted behind it; C p is the intercept of the expression of the P-th track line; C p+1 is the intercept of the expression of the (P + 1)-th track line; k 均 is the slope of the track line after fine-tuning; N is the number of track lines. When D p < D t , it indicates that the track lines P and P + 1 belong to the same track, and these two track lines are merged. D t is the track line merging distance threshold, and the default value is set to 6.

[0078] Specifically, the merging process is as follows: Extract the two end points of the P-th track line as (xh p , yh p ), (xt p , yt p ); Extract the two end points of the (P + 1)-th track line as (xh p+1 , yh p+1 ), (xt p+1 , yt p+1 ); Set new end points for the merged track line, denoted as H and T. The coordinates of the new end points are determined according to the coordinates of the original track line end points. When yh p> yh p+1 , H = (xh p+1 , yh p+1 ), otherwise H = (xh p , yh p ). Similarly, when yt p> yt p+1 , T = (xt p , yt p ), otherwise T = (xt p+1 , yt p+1 ).

[0079] Track line pairing: The truly paired track lines are parallel. However, in the turnout image, when the paired track lines are on one side of the image center line, the paired track lines are still parallel, that is, they have the same slope. But when the paired track lines are on both sides of the image center line, the magnitudes of the slopes of the two track lines are similar and the signs are opposite.

[0080] Within the a and b category groups, the processed track lines are sorted in descending order again according to the C value. Within the a and b groupings, the sorted track lines are paired up two by two in sequence. When there are still remaining unpaired track lines in the a and b category groups after pairing, the remaining track lines in the a category group and the remaining track lines in the b category group are paired separately, and this paired track line is specially marked as LinepairC.

[0081] Specifically, when the train is running on a curve or the camera has a position offset during data acquisition and other factors, there will be no LinepairC paired track lines. Under normal operating conditions, the LinepairC paired track lines are the boundary lines of the track area where the train is located.

[0082] S235: Connect the four endpoints of the paired track lines in sequence to form the track area.

[0083] S236: Repeat S234 and S235 to complete the classification, merging, and pairing of the track lines in all groups, and extract all the track areas in the turnout image.

[0084] S24: Screen all the track areas according to the track area positions to determine the track where the train is located. When there is a track area formed by the LinepairC paired track lines, this area is the track where the train is located. Otherwise, screen the track area where the train is located according to the following process: Specifically, according to the centroid calculation formula, screen all the track areas, and the specific process is as follows: Calculate the abscissa of the center point of the e-th track area, denoted as xz e , and calculate the track line spacing D of the e-th track area according to the track line spacing calculation formula in S234 e . Since the camera is located at the center of the train cab, the track where the train is located should be in the area closest to the center of the track in the image, and the track spacing imaging is the largest. Design the evaluation function as G according to the differences and characteristics between the side track area and the track area where the train is located e = 0.3×D e 2 - 0.7×(xz e - w / 2) 2 , where G eis the evaluation function value of the e-th track region, and w is the width of the original turnout image; one item of the evaluation function considers the track line spacing De in the track region. According to the characteristics of track image acquisition, since the track region where the train is located is often in the middle region of the image, the track line spacing De in the track region where the train is located should be the largest; another item of this evaluation function uses the horizontal coordinate of the center point of the track region and the distance from the center of the image as the evaluation index. The smaller the distance between the horizontal coordinate of the center point of the track region and the center of the image, the greater the possibility that this track region is the track region where the train is located. Finally, for the track line spacing item De 2 and the distance item between the horizontal coordinate of the center point of the track region and the center of the image (xz e -w / 2) 2 are respectively given different weights to form this evaluation function.

[0085] Calculate the G value of each track region according to the evaluation function, and take the track region corresponding to the largest G value as the track region where the train is located.

[0086] S25: Cut out the smallest circumscribed rectangle containing all track regions from the original image as the output image of data preprocessing.

[0087] Compare the corner points of all track regions, and take out the maximum and minimum values of the horizontal and vertical coordinates among all the corner points of the track regions. Take the minimum values of the horizontal and vertical coordinates as the upper left corner point of the smallest circumscribed rectangle, and take the maximum values of the horizontal and vertical coordinates as the lower right corner point of the smallest circumscribed rectangle. The smallest circumscribed rectangle is denoted as [x min , y min , x max , y max .

[0088] S26: Adjust the original image label so that it corresponds to the cropped image. Suppose the original image label is [x1, y1, x2, y2], then the label of the cropped image is [x1 - x min , y1 - y min , x2 - x min , y2 - y min .

[0089] As Figure 4 shown, after performing data preprocessing operations on the original turnout image, the output image only contains the track region of the original image. Compared with the original image, the invalid pixels in the image are reduced. Compared with the input of the traditional deep learning network, the size of the network input image is smaller, reducing the search space of the deep learning network.

[0090] S3: As Figure 5As shown in the figure, a track turnout state discrimination module is constructed, which is divided into two parts: a deep learning model and a traditional image processing technology model. The deep learning model is used to detect all turnout samples, while the traditional image processing technology model re-detects difficult samples on the basis of the deep learning model processing to improve the detection accuracy.

[0091] Furthermore, S3 includes the following sub-steps:

[0092] S31: Build a deep learning model.

[0093] Furthermore, the specific steps of S31 are as follows:

[0094] S311: Build the backbone network. The backbone network draws on the CSPDarknet53 structure and improves the CSP module in the CSPDarknet53 structure. The backbone network is divided into 11 network modules, namely Focus module, CONV module, TRCSP_1 module, CONV module, TRCSP_3 module, CONV module, TRCSP_3 module, CONV module, TRCSP_1 module, SPP module, and SE module.

[0095] The Focus module is a slicing method that retains image features as much as possible and quickly downsamples the image. Specifically, the image after S25 data preprocessing is sliced, and the width and height of the image after data preprocessing are set to W1 and H1, respectively. The slice size is W1 / 2 and H1 / 2. Then, the sliced feature map is stacked in the channel direction, and then the stacked feature map is downsampled using a 1×1 convolution kernel. Finally, the number of channels of the output feature map is adjusted to 64.

[0096] The CONV module performs convolution on the feature map, downsamples the obtained feature map by a factor of two, and adjusts the number of output channels.

[0097] The TRCSP_1 module and TRCSP_3 module are feature extraction modules that integrate the Transform structure. Their structure diagrams are as follows: Figure 6 As shown in the figure, TRCSP_1 indicates that N=1 in the structure diagram, that is, it contains 1 Transform component; similarly, TRCSP_3 indicates that N=3 in the structure diagram, that is, it contains 3 Transform modules. The introduction of the Transform module structure in the TRCSP_1 module and the TRCSP_3 module can improve the feature extraction capability of the network, and the CSP network structure itself can avoid a large number of repeated gradient calculations and retain the feature information of the upper network by merging the single convolution branch and the Transform module branch information in parallel.

[0098] Specifically, the Transform module is an excellent attention network module that can utilize the self-attention mechanism to improve the ability to capture different local features. The LayerNorm and Dropout layers in the Transform module can regularize the network, which is beneficial to the convergence of the network training process and avoids overfitting of the network.

[0099] Among them, the SPP module implements spatial pyramid pooling of image features. In the SPP module, first, the feature map input from the upper layer is convolved, then three parallel max-pooling layers are used to reduce the model parameters, and finally, the pooling results and the results after convolution are concatenated and convolved again to further extract image features.

[0100] Compared with the traditional CSPDarknet53 module, after the SPP module in the backbone network, the SE module is introduced to further extract image features. The SE module is a channel attention module that can automatically learn the importance of different channels of the feature map, focusing the attention of the network more on useful target features rather than image noise.

[0101] S312: Fuse the image features extracted in S311 to construct a detection head network.

[0102] Furthermore, the specific steps of S312 are as follows:

[0103] S3121: Specifically, the feature maps output by the first and second TRCSP_3 modules in S311 are denoted as C1 and C2 respectively, and the feature map output by the last SE module of the S311 backbone network is denoted as C3. The feature maps input to the first and second TRCSP_3 modules are denoted as C1 p 、C2 p , and the feature map input to the last SE module of the S311 backbone network is denoted as C3 p .

[0104] S3122: Convolve C3 to obtain a feature map denoted as D30, upsample C3, concatenate it with the result of C2 to obtain a feature map denoted as D20; upsample D20 and concatenate it with the result of C1 to obtain a feature map denoted as D10.

[0105] S3123: Perform atrous pyramid pooling operations on the three output feature maps D10, D20, and D30 respectively, and then feedback them back to the input ends of the first and second TRCSP_3 modules and the SE module respectively.

[0106] S3124: During the calculation process of the first and second TRCSP_3 modules and the SE module, the input ends are C1 p 、C2 p 、C3 pBased on the feedback in S3123, repeat S3121 and recalculate C1, C2, and C3.

[0107] S3125: Repeat S3122, S3123, and S3124 twice in sequence to perform cyclic extraction and fusion of image features, and obtain D12, D22, and D32. Concatenate D12, D22, D32 and D10, D20, D30 in the channel direction to obtain D1, D2, and D3 respectively as the detection head feature maps for turnout status detection.

[0108] S313: Decode the feature maps obtained in S3125 to obtain the predicted box positions, the object categories and category confidence levels within the predicted boxes. Specifically, the decoder selects the yolov5 decoder model.

[0109] S32: Construct a turnout status discrimination model using traditional image processing techniques.

[0110] Furthermore, the specific steps of S32 are as follows:

[0111] S321: Obtain the predicted box information with a confidence level lower than 0.5 in S313, and obtain the predicted box position information in the original image according to the downsampling ratio, and cut out the corresponding predicted box area in the original image, denoted as R 预 ;

[0112] S322: Grayscale the area of R 预 . Then extract the edge features of the image through the canny operator, and then use the Hough transform method to extract the track line features, and store the extracted line segment endpoint coordinates in a list.

[0113] S323: First, convert the track lines into the following expression form using the extracted track line endpoint information: y = Kx + C, where k is the intercept, and record the slopes and intercepts of each track line respectively. The slope and intercept of the t-th track line are denoted as Kt and Ct respectively.

[0114] Group the track lines according to the track line slopes. First, sort all the track lines in descending order according to the Kt value, and then select nt track lines as the pairing criteria by equidistant sampling from the sorted track lines. Denote the slopes of the nt tracks as kt nt , traverse all the remaining retained track lines, and denote the slope of the track line in each traversal as kt g , when |kt nt - kt g | < kt t , add the track line corresponding to kt g to the track line group represented by kt n . kt tIt is the slope difference threshold of the switch track line, and the default value is set to 0.2.

[0115] Merge the grouped track lines. Assume that each group has N t track lines. If there is only one track line in the group, no merging operation is required. Take the average value of the slopes of all track lines in the group and denote it as K 均1 , and sort them in descending order according to the intercept values. Calculate the distance between every two adjacent track lines where q ∈ N t -1. In the above formula, Dq is the track line distance between the q-th track line after descending order sorting and the (q + 1)-th track line sorted behind it; Cq is the intercept of the expression of the q-th track line; C q+1 is the intercept of the expression of the (q + 1)-th track line; N t is the number of track lines. When Dq < D1 t , it indicates that the track lines q and q + 1 belong to the same track, and merge these two track lines. D1 t is the switch track line merging distance threshold, and the default value is set to 4.

[0116] Specifically, extract the two end points of the q-th track line as (xh q , yh q ), (xt q , yt q ); extract the two end points of the (q + 1)-th track line as (xh q+1 , yh q+1 ), (xt q+1 , yt q+1 ); set new end points for the merged track line, denoted as H1 and T1. The coordinates of the new end points are determined according to the coordinates of the end points of the original track line. When yh q> yh q+1 , H1 = (xh q+1 , yh q+1 ), otherwise H1 = (xh q , yh q ). Similarly, when yt q> yt q+1 , T1 = (xt q , yt q ), otherwise T1 = (xt q+1 , yt q+1 ).

[0117] S324: When the number of the screened track lines is not equal to 4, it indicates that there is a missed detection or false detection in the switch track line detection. Adjust the hyperparameters for track line extraction, and repeat S322 and S323 until 4 track lines are detected.

[0118] S325: Calculate the midpoint coordinates of each track line and denote them as (xc n , yc n ), where n takes 1, 2, 3, 4. Sort the center points of each track line in ascending order according to the magnitude of xc n . The sorted track lines are denoted as L 左1 , L 左2 , L 右1 , L 右2 ; among them, the gap between L 左1 and L 左2 is the left gap of the switch, and the gap between L 右1 and L 右2 is the right gap of the switch.

[0119] S326: Calculate the intersection point between L 左1 and L 左2 and denote it as P 左 . Calculate the intersection point between L 右1 and L 右2 and denote it as P 右 . Determine whether the intersection point is within the corresponding track segment range. If the intersection point is within the corresponding track segment range, then consider this gap line as a false gap line and there is no gap when the two track lines intersect; if the intersection point is outside the corresponding track segment range, then consider this boundary line as a true gap line.

[0120] Specifically, if P 左 is within the corresponding track segment range and P 右 is not within the corresponding track segment range, then this switch is in the right-turn switch state;

[0121] If P 左 is not within the corresponding track segment range and P 右 is within the corresponding track segment range, then this switch is in the left-turn switch state;

[0122] If P 左 , P 右 are both not within the corresponding track segment range or P 左 , P 右 are both within the corresponding track segment range, it indicates that the state of this switch cannot be determined by the traditional image processing method, and directly output the recognition result of the deep learning model.

[0123] S4: Use the switch image and the corresponding label preprocessed by S25 and S26 as input data, and use the switch state discrimination model constructed by S31 and S32 as the detection tool to obtain the switch detection result.

[0124] Furthermore, S4 includes the following sub-steps:

[0125] S41: The turnout images preprocessed by S25 and S26 and the corresponding labels are sent to the network deep model of S31 for initial detection of the turnout state to obtain the predicted target box and category confidence.

[0126] S42: Judge the category confidence. When the confidence is greater than the threshold, directly output the turnout state judgment result. When the confidence is less than the judgment threshold, map the predicted target box back to the original image, and use the turnout state judgment model of S32 traditional image processing technology to judge the turnout state again. The default value of the confidence threshold is set to 0.5. When the two judgment results are consistent, the judgment result is directly output. When the two judgment results are inconsistent, the result of the traditional processing technology turnout model is output, but the predicted target box is specially marked. When the traditional processing technology turnout model cannot determine the result, the recognition result of the deep learning network model is used as the final output result, and special marking is also performed. Specifically, the special marking is to make a special distinction or special text description of the color of the prediction box during the visualization of the detection results to indicate the difference between the prediction box and other prediction boxes that directly judge the turnout state.

[0127] S5: Combine the output results of S24 and S42 to screen and identify the turnouts on the track where the train is located, and output the final turnout status recognition result. Compare the train track area image output by S24 with the turnout prediction frame output by S42, and calculate the intersection and union ratio J of all the recognized prediction frames with the track area image one by one. iou , let the image of the track area where the train is located be A, and the prediction box be B.

[0128]

[0129] In the above formula, area(A) represents the track area where the train is located, and area(B) represents the area where the prediction box is located.

[0130] When the intersection-merge ratio is greater than the threshold, the predicted box is considered to belong to the track area where the train is located, which is most important for the safety of train operation. When the intersection-merge ratio is less than the threshold, the predicted box is considered to be a side track switch, which has no direct impact on the safety of train operation, but can provide important references for tasks such as switch area identification and train positioning, so it is also retained.

[0131] In particular, in order for the driver to reasonably use the turnout detection results, the turnout status when the intersection ratio is greater than the threshold is directly displayed in a prominent red color, and the turnout status when the intersection ratio is less than the threshold is displayed in green for distinction.

[0132] The detection results of turnouts by the YOLOv5 algorithm and the detection results of turnouts by the algorithm of the present invention are compared. Figure 7a - Figure 7b As shown, Figure 7a The red ellipse in the middle indicates that the YOLOv5 algorithm misses detection.Figure 7b The single - turnout status detection system of the present invention has successfully detected the turnout targets at the corresponding positions and correctly identified the corresponding turnout status. The detection method based on the single - turnout status detection system of the present invention has accuracy and reliability.

[0133] The present invention also provides a semi - automatic data annotation method for a single - turnout status detection system, as Figure 8 shown, and the specific steps are as follows:

[0134] Step 1: Through the single - turnout status detection system, collect turnout images under different lines and different weather conditions, and divide the collected images into two parts: an artificial annotation set and an automatic annotation set. Use the VIA annotation tool to annotate the artificial annotation set, and randomly divide the artificial annotation set into three sub - datasets: train, trainval, and test according to the ratio of 8:1:1. The train dataset and the trainval dataset are used for the training and validation process, and the test dataset is used for the test and inference process. In particular, to ensure the accuracy of subsequent automatic annotation, the artificial annotation set images are at least 8000 frames and ensure that the number of turnout status instances in each category is not less than 10000.

[0135] In particular, the images in the artificial annotation set need to be manually screened, only select the images containing clear turnout status and perform accurate annotation, while the automatic annotation set can directly select the original video sequence images without manual screening, greatly reducing the workload of making the dataset.

[0136] Step 2: Train the annotated dataset through the turnout status discrimination deep - learning model of S31 to obtain the best turnout recognition model of the model.

[0137] First, use the artificial annotation set in Step 1 to train and test the turnout status discrimination model in the single - turnout status detection system, and record the number of training times as R i , when the turnout status detection accuracy and recall rate in the R i th round both reach above 0.85, take the model parameters of this round as available model parameters, and package the model parameters and test results of this round together as the available turnout status discrimination model of the R i th round and store it in the available model list. Otherwise, discard the model parameters of this round and continue to train and test the model. When the available model list is not empty and the test results of three consecutive rounds have no optimization effect, stop the operation, and select the group of available turnout status discrimination models with the best performance in the available model list as the optimal turnout status discrimination model. When the available model list is empty and the number of training times reaches the threshold R i阈 , the default value of R i阈 is taken as 300. At this time, the single - turnout status detection system cannot meet the detection requirements, stop the calculation process, and optimize the system.

[0138] Step 3: Randomly shuffle the images in the automatic annotation set and group them in units of 1000 images. Images in groups with less than 1000 are counted as 1000. Randomly select one group of data and use the optimal turnout status discrimination model obtained in Step 2 to detect the turnout status, obtain the label information of the images in the automatic annotation set, and remove the selected grouped data from the automatic annotation dataset.

[0139] Step 4: Manually check whether the automatically generated annotation files meet the annotation requirements, optimize the corresponding unreasonable annotations, remove the images without generated annotation information, and add the images with qualified annotations to the manual annotation set.

[0140] Step 5: Randomly select another group of images from the remaining small groups of images in the automatic annotation set and repeat Steps 2, 3, and 4 until the automatic annotation set becomes empty and all the images in the automatic annotation set are completely annotated, ending the annotation process.

[0141] As Figure 9 shown, the box with a label name is the automatic annotation box, and the box without a label is the manual annotation box. The label quality generated by this semi-automatic label annotation method for the dataset can meet the training requirements, while greatly improving the annotation efficiency and significantly reducing the annotation workload.

[0142] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A detection method for a single - open turnout state detection system, characterized in that, It includes the following steps: S1. Make a turnout data set; S2. Perform image preprocessing on the data set images; S3. Construct an orbital turnout status discrimination module; S4. Use the preprocessed turnout images and corresponding labels as input data, and use the constructed turnout status discrimination model as a detection tool to obtain turnout detection results; S5. Combine the track where the train is located and the turnout status discrimination result to screen and identify the turnouts on the track where the train is located, and output the final turnout status identification result; In S2, performing image preprocessing on the data set images includes the following steps: S21. Grayscale the RGB images collected in the data set; S22. Use the region growing method to extract track line information; S23. Screen and pair the track lines to determine the track area; S24. Screen all the track areas according to the track area position to determine the track where the train is located; S25, intercept the minimum bounding rectangle containing all track regions from the original image as the output image of data preprocessing; compare the corner points of all track regions, take out the maximum and minimum values of the horizontal and vertical coordinates among all the corner points of the track regions, use the minimum value of the horizontal and vertical coordinates as the upper left corner point of the minimum bounding rectangle, and use the maximum value of the horizontal and vertical coordinates as the lower right corner point of the minimum bounding rectangle. The minimum bounding rectangle is denoted as [x min , y min , x max , y max ; S26. Adjust the original image labels so that they correspond to the cropped image. Let the original image label be [x1, y1, x2, y2], then the label of the cropped image is [x1 - x min , y1 - y min , x2 - x min , y2 - y min .

2. The detection method of the single turnout status detection system according to claim 1, characterized in that In S22, using the region growing method to extract track line information includes the following steps: S221. Obtain the seed starting points for region growing. Set 10 seed starting points, which are set in the image using a uniform distribution method. The image positions of the first 5 seed starting points are respectively (w×i / 5, 5 / 9×h), and the image positions of the last 5 seed starting points are respectively (w×(i - 5) / 5, 7 / 9×h), where h is the height of the original turnout image, w is the width of the original turnout image, and i represents the i-th seed starting point; S222, record the coordinates of the i-th seed starting point as (sdx i , sdy i ), and obtain the coordinates of the pixel points within the 8-neighborhood of the seed starting point. Denote the coordinates of the neighborhood pixel points as (X 邻 , Y 邻 ). It can be known that (X 邻 , Y 邻 ) = (sdx i , sdy i ) + R, where R represents the 8-neighborhood relative coordinate array, and R = [(-1, -1), (-1, 0), (-1, 1), (1, 1), (0, 1), (1, 0), (1, -1), (0, -1)]; S223. Compare the gray - level difference between the starting position of the \(i\) - th seed and the positions in its 8 - neighborhood. When the gray - level difference \(D((sdx i , sdy i ))\leq gray t , the starting position of the seed and the corresponding neighborhood positions belong to the same gray - level region. Set the gray - level values of the starting position of the seed and its corresponding neighborhood positions to 255 uniformly, and set the neighborhood position as the new starting position of the seed. \(gray t is the gray - level difference threshold, and the default value is set to 6. The expression for the gray - level difference is as follows: D((sdx i ,sdy i )) = gray((sdx i ,sdy i )) - gray((X 邻 ,Y 邻 )) In the above formula, D((sdx i , sdy i )) represents the difference between the grayscale values of the i-th seed starting point and the corresponding neighborhood, gray((sdx i , sdy i )) represents the grayscale value at the position of the seed starting point, and gray((X 邻 , Y 邻 ) represents the grayscale value at the neighborhood position; S224. Repeat S223 until there are no available seed starting points, indicating that the corresponding area has completed the "growing process". After region growing, the remaining area surrounded by the 255 grayscale value is the track line area.

3. The detection method of the single-opening turnout state detection system according to claim 1, characterized in that In S23, screening and pairing the track lines to determine the track area includes the following steps: S231. Eliminate false track lines; The Hough transform algorithm is selected to extract the track line, and the obtained track line is represented as follows: The j-th track line is represented by the two endpoint values of the line segment, denoted as (xh j , yh j ), (xt j , yt j ), where (xh j , yh j ) represents the endpoint closer to the x-axis in the track line segment, that is, yh j < yt j ; From the endpoint coordinates of the track line, the expression of the j-th track line where y is the y-axis coordinate of the track line and x is the x-axis coordinate of the track line; The slope of the track line The length of the track line in the pixel coordinate system where k j is the slope of the j-th track line, and L j is the length of the j-th track line; When k j ∈ [-∞, -1] ∪ [1, ∞] and L j > 60 pixels, keep the track line, otherwise delete the track line; S232. Group the track lines according to the track line slope; Calculate the slopes of the remaining orbital lines, sort all the orbital lines in descending order according to the absolute values of the slopes, and select n orbital lines as the pairing criteria by equidistant sampling from the sorted orbital lines. The default value of n is Let the slopes of the n orbital lines be k n , traverse all the remaining retained orbital lines, and record the slope of the orbital line in each traversal as k g . When ||k n |-|k g || < k t , add the orbital line corresponding to k g to the orbital line group represented by k n . k t is the threshold of the absolute value difference of the slopes, and the default value is set to 0.1; S233. Repeat S232 for the remaining ungrouped track lines until all the track lines are grouped, and end the grouping process. Discard the groups with only one track line directly; S234. Merge different track line segments of the same track to perform track merging and pairing; Divide the track lines inside the group into two categories according to the positive and negative of the slope, and record them as category groups a and b respectively. Do not merge the category groups with only one track line. The number of track lines in each group that meets the merging requirements is recorded as N, N≥2. The following processes are respectively carried out for track merging and track line pairing: The specific process of track merging is as follows: For the track lines of category groups a and b, the following processing is performed respectively: calculate the mean slope k of all track lines within the track line category group 均 , Perform a slope fine-tuning on N track lines to unify them to k 均 ; For N track lines, transform the track line expression in S231 to: , where y h is the y-axis coordinate of a certain point on the track line, and x h is the x-axis coordinate corresponding to the y-axis coordinate of a certain point on the track line; Let , where C represents the intercept of the corresponding track line expression; Sort the N track lines in descending order according to the C value, and calculate the track distance between every two adjacent track lines where p ∈ N - 1; In the above formula, D p is the track distance between the P-th track line after descending order sorting and the (P + 1)-th track line sorted behind it; C p is the intercept of the P-th track line expression; C p+1 is the intercept of the (P + 1)-th track line expression; k 均 is the slope of the fine-tuned track line; N is the number of track lines. When D p < D t , it indicates that the track lines P and P + 1 belong to the same track, and these two track lines are merged. D t is the track line merging distance threshold, and the default value is set to 6; Extract the two end points of the P-th track line as (xh p , yh p ), (xt p , yt p ); extract the two end points of the (P + 1)-th track line as (xh p+1 , yh p+1 ), (xt p+1 , yt p+1 ); set new end points for the merged track line, denoted as H and T; the coordinates of the new end points are determined according to the coordinates of the end points of the original track lines. When yh p> yh p+1 , H = (xh p+1 , yh p+1 ), otherwise H = (xh p , yh p ); similarly, when yt p> yt p+1 , T = (xt p , yt p ), otherwise T = (xt p+1 , yt p+1 ); The specific process of track line pairing is as follows: Inside the category groups a and b, respectively sort the processed track lines again in descending order according to the C value. Inside the a and b groups, respectively pair the sorted track lines in sequence two by two. When there are still remaining unpaired track lines in the a and b category groups after pairing, pair the remaining track lines in the a category group and the remaining track lines in the b category group separately, and make a special mark for this paired track line, denoted as LinepairC; S235. Connect the four endpoints of the paired track lines in sequence as the track area; S236. Repeat S234 and S235 to complete the classification, merging, and pairing of the track lines for all groups, and extract all the track regions in the turnout image.

4. The detection method of the single turnout state detection system according to claim 1, characterized in that In S24, screen all the track regions according to the track region positions to determine the track where the train is located, including the following steps: When there is a track region formed by the paired track lines of LinepairC, this region is the track where the train is located; otherwise, screen the track region where the train is located according to the following process. The specific process is as follows: According to the centroid calculation formula, all track regions are screened: calculate the abscissa of the center point of the e-th track region, denoted as xz e , calculate the track line spacing D of the e-th track region according to the track line spacing calculation formula e , design the evaluation function as G according to the differences and characteristics between the side track region and the track region where the train is located e = 0.3×D e 2 - 0.7×(xz e - w / 2) 2 , where G e is the evaluation function value of the e-th track region, and w is the width of the original switch image; calculate the G values of each track region according to the evaluation function, and take the track region corresponding to the maximum G value as the track region where the train is located.

5. The detection method of the single turnout status detection system according to claim 1, characterized in that In S3, construct a track turnout state discrimination module, including the following steps: S31. Construct a deep learning model. The specific steps are as follows: S311. Construct a backbone network. The backbone network is divided into 11 network modules, which are successively a Focus module, a CONV module, a TRCSP_1 module, a CONV module, a TRCSP_3 module, a CONV module, a TRCSP_3 module, a CONV module, a TRCSP_1 module, an SPP module, and an SE module; The Focus module slices the pre-processed image of the data. Assume the width of the pre-processed image of the data is W1 and the height is H1, and the slicing size is W1 / 2, H1 / 2. Then stack the sliced feature maps in the channel direction, and then perform downsampling on the stacked feature maps using a 1×1 convolution kernel. Finally, adjust the number of channels of the output feature map to 64; The CONV module performs two-fold downsampling on the feature map output by the Focus module and adjusts the number of output channels; The TRCSP_1 module and the TRCSP_3 module are feature extraction modules that integrate the Transform structure. The TRCSP_1 module contains 1 Transform component, and the TRCSP_3 module contains 3 Transform modules; The Transform module is an attention network module. The LayerNorm and Dropout layers in the Transform module can regularize the network; The SPP module realizes spatial pyramid pooling of the image features. The SPP module first performs convolution on the feature map input from the upper layer, then uses 3 parallel maximum pooling layers to reduce the model parameters, and finally splices the pooling result and the result after convolution, and performs convolution again to further extract the image features; S312. Fuse the image features extracted in S311 to construct a detection head network. The specific steps are as follows: S3121, denote the feature maps output by the first and second TRCSP_3 modules in S311 as C1 and C2 respectively, denote the feature map output by the last SE module of the backbone network in S311 as C3, and denote the feature maps input to the first and second TRCSP_3 modules as C1 p , C2 p , and denote the feature map input to the last SE module of the backbone network in S311 as C3 p ; S3122. Perform convolution on C3 to obtain a feature map denoted as D30; Upsample C3 and splice it with the result of C2 to obtain a feature map denoted as D20; Upsample D20 and splice it with the result of C1 to obtain a feature map denoted as D10; S3123. Perform atrous pyramid pooling operations on the three output feature maps of D10, D20, and D30 respectively, and then feedback them back to the input ends of the first and second TRCSP_3 modules and the SE module respectively; S3124. During the calculation processes of the first and second TRCSP_3 modules and the SE module, the input end introduces the feedback in S3123 based on C1 p , C2 p , C3 p , repeats S3121, and recalculates C1, C2, and C3; S3125. Repeat S3122, S3123, and S3124 twice in sequence to perform cyclic extraction and fusion of image features, and obtain D12, D22, and D32. Concatenate D12, D22, D32 with D10, D20, and D30 in the channel direction to obtain D1, D2, and D3 respectively as the detection head feature maps for turnout state detection. S313. Decode the feature maps obtained in S3125 to obtain the predicted box positions, object categories within the predicted boxes, and category confidence levels. S32. Construct a turnout state discrimination model using traditional image processing techniques. The specific steps are as follows: S321, obtain the information of the prediction boxes with a confidence level lower than 0.5 in S313, obtain the position information of the prediction boxes in the original image according to the downsampling ratio, and intercept the corresponding prediction box regions in the original image, denoted as R 预 ; S322, for R 预 Perform grayscale conversion on the area, extract the edge features of the image through the Canny operator, then use the Hough transform method to extract the track line features, and store the coordinates of the line segment endpoints extracted in a list; S323. Use the extracted track line endpoint information to convert the track lines into the following expression form: y = Kx + C, where k is the intercept. Record the slopes and intercepts of each track line respectively. The slope and intercept of the t-th track line are denoted as Kt and Ct respectively. Group the track lines according to the slope of the track lines. The specific process is as follows: Sort all the track lines in descending order of the Kt value, and then select nt track lines from the sorted track lines by equidistant sampling as the pairing criteria. Denote the slopes of the nt tracks as kt nt , traverse all the remaining retained track lines, and denote the slope of the track line for each traversal as kt g , when |kt nt -kt g |<kt t , add the track line corresponding to kt g to the track line group represented by kt n , where kt t is the threshold for the slope difference of the turnout track lines, and the default value is set to 0.2; Merge the grouped track lines. Assume that each group has N t track lines. If there is only one track line in the group, no merging operation is required; take the average of the slopes of all track lines in the group, denoted as K 均1 , and sort them in descending order according to the intercept values. Calculate the distance between every two adjacent track lines where q ∈ N t -1. In the above formula, Dq is the track line distance between the q-th track line sorted in descending order and the (q + 1)-th track line sorted behind it; Cq is the intercept of the expression of the q-th track line; C q+1 is the intercept of the expression of the (q + 1)-th track line; N t is the number of track lines; when Dq < D1 t , it indicates that the track lines q and q + 1 belong to the same track, and these two track lines are merged. D1 t is the threshold for the merging distance of the turnout track lines, and the default value is set to 4; Extract the two end points of the q-th track line as (xh q , yh q ), (xt q , yt q ); extract the two end points of the (q + 1)-th track line as (xh q+1 , yh q+1 ), (xt q+1 , yt q+1 ); set new end points for the merged track line, denoted as H1 and T1; the new end point coordinates are determined according to the original track line end point coordinates. When yh q> yh q+1 , H1 = (xh q+1 , yh q+1 ), otherwise H1 = (xh q , yh q ); similarly, when yt q> yt q+1 , T1 = (xt q , yt q ), otherwise T1 = (xt q+1 , yt q+1 ); S324. When the number of filtered track lines is not equal to 4, it indicates that there are missed detections or misdetections in the turnout track line detection. Adjust the track line extraction hyperparameters and repeat S322 and S323 until 4 track lines are detected. S325, obtain the midpoint coordinates of each track line respectively, denoted as (xc n , yc n ), where n takes 1, 2, 3, 4; sort the center points of each track line in ascending order according to the magnitude of xc n . The sorted track lines are denoted as L 左1 , L 左2 , L 右1 , L 右2 respectively; the gap between L 左1 and L 左2 is the left gap of the turnout, and the gap between L 右1 and L 右2 is the right gap of the turnout; S326, calculate L 左1 and the intersection point between L 左2 is denoted as P 左 , calculate the intersection point between L 右1 and L 右2 and denote it as P 右 ; Determine whether the intersection point is within the range of the corresponding track segment. If the intersection point is within the range of the corresponding track segment, it is considered that the gap line is a false gap line and there is no gap at the intersection of the two track lines; if the intersection point is outside the range of the corresponding track segment, it is considered that the boundary line is a true gap line; If P 左 is within the corresponding track segment range, and P 右 is not within the corresponding track segment range, then the turnout is in the right-turn turnout state; If P 左 is not within the corresponding track segment range, and P 右 is within the corresponding track segment range, then the turnout is in the left-turn turnout state; If P 左 , P 右 are both not within the corresponding track segment range or P 左 , P 右 are both located within the corresponding track segment range, it indicates that the turnout state cannot be determined by traditional image processing methods, and the recognition result of the deep learning model is directly output.

6. The detection method of the single-opening turnout state detection system according to claim 1, characterized in that In S4, use the preprocessed turnout image and the corresponding labels as input data, and use the constructed turnout state discrimination model as the detection tool to obtain the turnout detection results. It includes the following steps: S41. Send the preprocessed turnout image and the corresponding labels into the deep network model for the initial detection of the turnout state to obtain the predicted target boxes and category confidence levels. S42. Judge the category confidence level. When the confidence level is greater than the threshold, directly output the turnout state discrimination result. When the confidence level is less than the judgment threshold, map the predicted target box back to the original image, and use the turnout state discrimination model of traditional image processing techniques to perform secondary discrimination on the turnout state. The default value of the confidence level threshold is set to 0.

5. When the two discrimination results are the same, directly output the discrimination result. When the two discrimination results are different, output the result of the traditional processing technique turnout model, but specially mark the predicted target box. When the traditional processing technique turnout model cannot determine the result, use the recognition result of the deep learning network model as the final output result and also perform special marking.

7. The detection method of the single turnout state detection system according to claim 1, characterized in that, In S5, combine the track where the train is located and the turnout state discrimination result to screen and identify the turnouts on the track where the train is located, and output the final turnout state recognition result. The specific process is as follows: Compare the output train track area image with the output turnout prediction box, and calculate the intersection over union J of each recognized prediction box with the track area image one by one iou , denote the train track area image as A and the prediction box as B In the above formula, area(A) represents the area of the track where the train is located, and area(B) represents the area of the predicted box. When the intersection over union is greater than the threshold, the predicted box belongs to the area of the track where the train is located. When the intersection over union is less than the threshold, the predicted box is a turnout on the side track. Display the turnout detection results in different colors respectively.

8. A semi-automatic data annotation method for a single turnout status detection system, characterized in that, Specifically, it includes the following steps: Step 1: Through the single turnout status detection system, collect turnout images under different lines and different weather conditions, and divide the collected images into two parts: an artificial annotation set and an automatic annotation set. Use the VIA annotation tool to annotate the artificial annotation set, and randomly divide the artificial annotation set into three sub-datasets: train, trainval, and test, in the ratio of 8:1:1 by quantity. The number of images in the artificial annotation set is at least 8000 frames, and ensure that the number of turnout status instances in each category is not less than 10000; Step 2: Train the annotated dataset through the turnout status discrimination deep learning model of S31 to obtain the best turnout recognition model; the specific process is as follows: Train and test the switch state discrimination model in the single turnout state detection system using the manually labeled set in step 1, and record the number of training times as R i , when the turnout state detection accuracy and recall rate in the R i -th round both reach above 0.85, take the model parameters of this round as available model parameters, and package the model parameters and test results of this round together as the R i -th round of available switch state discrimination model and store it in the available model list; Otherwise, discard the model parameters of this round and continue the next round of training and testing of the model. When the list of available models is not empty and there is no optimization effect in the test results for three consecutive rounds, stop the operation and select the group of available turnout state discrimination models with the best performance in the list of available models as the optimal turnout state discrimination model; when the list of available models is empty and the number of training times reaches the threshold R i阈 At this time, R i阈 The default value is 300. At this time, the single turnout state detection system cannot meet the detection requirements, stop the calculation process, and optimize the system; Step 3: Randomly shuffle the images in the automatic annotation set, and group them in units of 1000 images. Images in groups with less than 1000 are counted as 1000. Randomly select 1 group of data and use the optimal turnout status discrimination model obtained in Step 2 to perform turnout status detection to obtain the label information of the images in the automatic annotation set, and remove the selected grouped data from the automatic annotation dataset; Step 4: Manually check whether the automatically generated annotation files meet the annotation requirements, optimize the corresponding unreasonable annotations, remove the images without generated annotation information, and add the images with qualified annotations to the artificial annotation set; Step 5: Randomly select another group of images from the remaining small groups of images in the automatic annotation set and repeat Steps 2, 3, and 4; until the automatic annotation set becomes empty, all the images in the automatic annotation set are completely annotated, and the annotation process ends.

Citation Information

Patent Citations

  • Turnout and non-turnout rail fastener positioning method based on deep learning

    CN113506269A