Track turnout visual identification method based on bidirectional feature network
By collecting image data and point cloud data in various track environments on the data acquisition vehicle, a switch identification method based on a two-way feature network is established, which solves the problems of large amount of turntable detection and low recognition accuracy in the prior art, and realizes high-precision switch identification on embedded devices.
Patent Information
- Application Number
- CN202510169605.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, turntwitch detection and calculation volume are large, and it is difficult to use in embedded equipment. It has low recognition accuracy in complex environments, so it cannot meet the needs of railway transportation.
The visual recognition method of track switches based on two-way feature network is adopted. By installing point cloud radars and cameras on the data acquisition vehicle, image data and point cloud data in various track environments are collected, environmental image data sets and turntable image training sets are established, and the input data is processed using backbone networks, two-way feature networks and detection heads to obtain the turntable position and category prediction results.
It significantly reduces the amount of model parameters and improves the lightweighting of the model. It is suitable for resource-constrained embedded tasks and can achieve high-precision switch identification under different track environments and lighting conditions.
Smart Images

Figure CN120107918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of train intelligent perception, and in particular to a track turnout visual recognition method based on a bidirectional feature network. Background Art
[0002] Railway inspection is an important function to ensure the safe and efficient operation of trains. Regular railway inspection is the prerequisite for the normal operation of trains and the basic guarantee for the safe and efficient operation of trains. Currently, most of the existing train inspection technologies are carried out by manpower or by equipment beside the track, and their effects are limited. If the trackside equipment is damaged in the train travel section, the lack of autonomous operation and inspection capabilities of the train can easily lead to accidents. At present, the problem of high-precision positioning in the research on autonomous train inspection needs to be solved urgently, so it is of great significance to locate the train section through turnout identification.
[0003] There are currently two main solutions for train track switch identification: traditional stake target recognition and pixel tracking method. Due to factors such as dense switchboards in stations and short track sections, trains may pass through multiple target points in a short period of time during operation, but fewer measurement points are collected, resulting in positioning confusion or inability to quickly identify trains occupying tracks, which cannot meet the needs of railway transportation. The pixel tracking method has obvious limitations on visual recognition accuracy in complex operating environments and is easily affected by transmission methods and mechanisms. The granularity of perceived information varies greatly, making it difficult to provide real-time and dynamic reference for train operation and dispatching command, thereby increasing the difficulty of positioning.
[0004] The Chinese invention patent application with publication number CN118149727A provides a railway turnout detection method based on 3D point cloud, combining graying and deep learning algorithms to detect turnout key points and structural features. This method has high accuracy and the ability to adapt to complex scenes. However, 3D point cloud acquisition requires stable environmental conditions. For example, scenes with large vibrations (such as high-speed train scenes) may affect the acquisition quality. In addition, laser scanners are usually expensive, and high-precision equipment can only obtain sufficiently accurate point cloud data, which may increase the equipment cost burden for small and medium-sized railway maintenance units. Summary of the invention
[0005] In view of the above problems, the present invention provides a track turnout visual recognition method based on a bidirectional feature network, which solves the problem in the prior art that turnout detection has a large amount of calculation and is difficult to use in embedded devices.
[0006] The present invention provides a track turnout visual recognition method based on a bidirectional feature network, comprising the following steps:
[0007] Step S1, installing a point cloud radar and a camera on a data collection vehicle, and adjusting the internal and external parameters of the point cloud radar and the camera;
[0008] Step S2, deploying the data collection vehicle to the train track, collecting multiple image data and point cloud data containing turnouts in various track environments by the point cloud radar and camera on the data collection vehicle, and establishing an environmental image data set; marking the turnouts in the environmental image data, and establishing a turnout image training set;
[0009] Step S3, establishing a turnout recognition network, wherein the input data of the turnout recognition network is processed by the backbone network, the bidirectional feature network and the detection head in sequence to obtain the turnout position and category prediction results; the network depth of the bidirectional feature network increases linearly, and the network width increases exponentially;
[0010] Step S4, training the turnout recognition network based on the turnout image training set to obtain a trained turnout recognition network;
[0011] Step S5: Use the trained turnout recognition network to process the environmental image data collected in real time to obtain the turnout position and category prediction results.
[0012] Preferably, step S1 specifically includes:
[0013] Step S1-1, setting a camera bracket on the data collection vehicle, and fixing a camera and a point cloud radar on the camera bracket at the same time;
[0014] Step S1-2, calibrating the camera internal parameters to correct the distortion of the image captured by the camera;
[0015] Step S1-3: calibrate the external parameters of the point cloud radar and the camera so that the data collected by the point cloud radar and the camera match each other.
[0016] Preferably, step S2 specifically includes:
[0017] Step S2-1, deploying a data collection vehicle on the train track, and using the calibrated point cloud radar and camera to collect multiple camera-photographed images and point cloud data including turnouts in various track environments;
[0018] Step S2-2, converting the point cloud data into a grayscale image, taking an image captured by a camera shooting the same scene and a grayscale image converted from the corresponding point cloud data as samples, and taking a set formed by a plurality of the samples after data enhancement as the environmental image data set;
[0019] Step S2-3, annotate the turnouts in the camera-captured images and the corresponding grayscale images converted from the point cloud data in the samples in the environmental image data set to obtain the annotation results; use the camera-captured images, the grayscale images converted from the point cloud data and the corresponding annotation results as a training sample, and use a set formed by multiple training samples as the turnout image training set.
[0020] Preferably, the data enhancement step in step S2-2 specifically includes: rotating, cropping and color-adjusting the grayscale image converted from the camera image and the corresponding point cloud data to form a plurality of new samples;
[0021] The step of obtaining the annotation result in step S2-3 specifically includes: annotating the precise position of the turnout and the opening direction of the turnout for the grayscale image converted from the camera image and the corresponding point cloud data, wherein the annotation result includes a rectangular frame of the position of the turnout in the image and the opening state of the turnout, wherein the opening state includes left or right.
[0022] Preferably, the step S3 specifically includes:
[0023] Step S3-1, using the pre-trained EfficientNets as the backbone network, wherein the width and depth of the backbone network are set to the width and depth of the network variants of EfficientNet from B0 to B6;
[0024] Step S3-2, establishing a bidirectional feature network, the bidirectional feature network is composed of a plurality of bidirectional feature network layers connected sequentially, the output feature maps of the 3rd to 7th layers of the backbone network are linearly amplified, and then sequentially subjected to weighted fusion processing of the plurality of bidirectional feature network layers, and finally a multi-scale fusion feature is obtained; the number of network layers of the bidirectional feature network increases linearly, and the number of convolution kernels in each layer of the network increases exponentially;
[0025] Step S3-3, establish a detection head, input the multi-scale fusion features output by the bidirectional feature network into the detection head, and process them through the classification head and regression head respectively to obtain the turnout position and category prediction results.
[0026] Preferably, in step S3-2, the step of linearly amplifying the output feature maps of the 3rd to 7th layers of the backbone network comprises: adjusting the growth parameter The size of determines the base size of the output feature map of the 3rd to 7th layers after enlargement, and the length and width of the output feature map are enlarged in proportion to the base size. The expression of the base size is:
[0027]
[0028] Among them, R input is the base size of the feature map after enlargement, Represents the network scale magnification composite coefficient;
[0029] The expression of the weighted fusion processing of the bidirectional feature network layer is:
[0030]
[0031] Among them, I i represents the i-th input feature map, ω i represents the learnable weight corresponding to the i-th input feature map, ∈ is a stable constant, ∑ j ω j Represents the sum of the learnable weights corresponding to all input feature maps;
[0032] The expression of the number of network layers of the bidirectional feature network and the number of convolution kernels in each layer of the network is:
[0033]
[0034] Among them, W 双向特征网络 It represents the width of the bidirectional feature network, that is, the number of convolution kernels in each layer of the bidirectional feature network, D 双向特征网络 represents the depth of the bidirectional feature network, that is, the number of layers of the bidirectional feature network, It represents the network scale magnification compound coefficient, and α is the base of width growth.
[0035] Preferably, step S3-3 specifically includes:
[0036] A detection head is established. The detection head includes two branches, namely, a classification head and a regression head. The processing expression of the classification head is:
[0037]
[0038] Among them, p c Represents the category confidence of the opening direction of the turnout, w c represents the weight of the classification head, F represents the multi-scale fusion feature of the input, and Sigmoid(·) represents the Sigmoid function;
[0039] The processing expression of the regression head is:
[0040] b=Regression(F)
[0041] Where b represents the predicted turnout position bounding box parameter, b = [x min ,y min ,x max ,y max ], x min ,y min ,x max ,ymax Represent the horizontal and vertical coordinates of the upper left corner and the lower right corner of the bounding box respectively, and Regression(·) represents the regression process.
[0042] Preferably, step S4 specifically includes:
[0043] Step S4-1, determining the loss function of the turnout recognition network;
[0044] Step S4-2, keeping the network weights of the backbone network unchanged, based on the training samples of the turnout image training set, taking minimization of the loss function as the optimization goal, and only adjusting the network weights of the training detection head and the bidirectional feature network;
[0045] Step S4-3, making the network weight of the backbone network variable, based on the training samples of the turnout image training set, taking minimizing the loss function as the optimization goal, adjusting the network weights of the backbone network, the bidirectional feature network and the training detection head, and finally obtaining a trained turnout recognition network.
[0046] Preferably, in step S4-1, the loss function expression of the turnout recognition network is:
[0047] FL(p t )=-α t (1-p t ) γ log(p t )
[0048] Among them, p t Represents the predicted category probability, FL(p t ) represents the class probability p t The corresponding network loss, α t represents the weight of balancing positive and negative samples, and γ is the focus adjustment factor.
[0049] Compared with the prior art, the present invention has at least the following beneficial effects:
[0050] (1) The present invention ensures the accuracy and reliability of data collection by rationally arranging point cloud radars and cameras on the data collection vehicle and performing strict internal and external parameter calibration. At the same time, a data collection strategy under multiple scenes and multiple environments is adopted to establish a high-quality environmental image data set and a turnout image training set, providing a solid data foundation for subsequent model training. This not only improves the generalization ability of the model, but also enables the system to adapt to different track environments and lighting conditions.
[0051] (2) The present invention establishes a turnout recognition network for 2D image detection, and significantly reduces the number of model parameters through efficient network structure design. Compared with the traditional detection model, the present invention reduces the number of parameters by about 4-8 times while maintaining the same detection performance.
[0052] (3) The present invention adopts a weighted fusion mechanism in the bidirectional feature network, which further improves the lightweight degree of the model by reducing redundant calculations, and is particularly suitable for processing resource-constrained embedded tasks. At the same time, the present invention provides multiple network variants of the backbone network, and the appropriate model configuration can be flexibly selected according to the actual hardware resource limitations. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings are only for the purpose of illustrating particular embodiments and are not to be construed as limiting the invention.
[0054] Figure 1 A flow chart of a track turnout visual recognition method based on a bidirectional feature network provided by the present invention;
[0055] Figure 2 This is a schematic diagram of the network processing process of turnout identification provided by the present invention.
[0056] Figure 3 A schematic diagram of the bidirectional feature network layer provided by the present invention.
[0057] Figure 4 A schematic diagram of the environment for collecting train turnout data provided by the present invention.
[0058] Figure 5 The present invention provides an overall implementation flow chart of the track turnout visual recognition method based on a bidirectional feature network. DETAILED DESCRIPTION
[0059] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein, and therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0060] In order to illustrate the effectiveness of the method proposed by the present invention, the above technical solution of the present invention is described in detail below through a specific embodiment. Figure 1 , Figure 5 As shown, a track turnout visual recognition method based on a bidirectional feature network is disclosed, and the specific implementation steps are as follows:
[0061] Step S1: Install a point cloud radar and a camera on a data collection vehicle, and adjust the internal and external parameters of the point cloud radar and the camera.
[0062] The data collection vehicle of the present invention is provided with a camera bracket, and the camera and the point cloud radar are fixed on the camera bracket at the same time. The angle and position of the device can be adjusted to ensure that it can cover the track area required for collection.
[0063] After the camera and point cloud radar are installed, the camera's internal parameters are calibrated to achieve camera distortion correction. The internal parameters may include parameters such as the camera's focal length, principal point coordinates, and distortion coefficient.
[0064] In some embodiments, a standard checkerboard calibration plate can be used to collect multiple sets of images at different angles and distances, and the Zhang calibration method can be used to calibrate the camera's internal parameters to obtain the camera's focal length, principal point coordinates, distortion coefficient and other internal parameters. After completing the internal parameter calibration, the images collected by the camera can achieve distortion correction and meet the accuracy requirements of subsequent image processing.
[0065] Afterwards, the external parameters of the point cloud radar and the camera are calibrated so that the data collected by the point cloud radar and the camera can be fused in the same coordinate system. The external parameters may include parameters such as the rotation matrix and the translation vector between the point cloud radar and the camera.
[0066] In some embodiments, multiple feature calibration plates can be arranged in the calibration site to collect camera images and laser point cloud data. The rotation matrix and translation vector between the point cloud radar and the camera can be solved by matching feature points, thereby achieving accurate alignment of the point cloud data and image data.
[0067] Step S2: deploy the data collection vehicle to the train track, and use the point cloud radar and camera on the data collection vehicle to collect multiple image data and point cloud data containing turnouts in various track environments to establish an environmental image data set; mark the turnouts in the environmental image data to establish a turnout image training set.
[0068] In this step, the data collection vehicle is deployed on the train track for dynamic data collection. The calibrated point cloud radar and camera on the data collection vehicle collect multiple camera images and point cloud data of turnouts in various track environments. Figure 4 shown.
[0069] In some embodiments, the multiple track environments may include straight track sections, curved sections, uphill sections, downhill sections, etc. containing turnouts. To ensure the integrity and representativeness of the data, multiple acquisitions may be performed at different times and under different weather conditions to obtain multiple camera-captured images and point cloud data containing turnout areas. The camera-captured images and point cloud data are a set of paired data used to represent the same shooting scene.
[0070] In some embodiments, the collected data can be screened and preprocessed to select representative camera images and corresponding point cloud data from a variety of track environments, and to eliminate low-quality camera images and corresponding point cloud data caused by insufficient lighting, blurring, and jitter.
[0071] In some embodiments, the image data can be enhanced to expand the diversity of samples. The data enhancement can include rotation transformation (including multiple angles such as 0°, 90°, 180°, 270°, etc.), extracting local features of the image using a random cropping method, and performing color adjustment on the image (including adjustment of brightness, contrast, and saturation) to increase the diversity of samples.
[0072] The point cloud data collected by the point cloud radar is converted into a grayscale image. The specific process is: project the three-dimensional point cloud data onto a two-dimensional plane to generate a depth map. The maximum-minimum normalization method is used to map the depth values in the depth map to the range of [0, 255] to form a grayscale image. This conversion method not only retains the geometric and reflectivity information of the point cloud, but also matches the data format with the input requirements of the bidirectional feature network, which is conducive to subsequent feature extraction and fusion processing.
[0073] An image captured by a camera and a grayscale image converted from corresponding point cloud data of the same scene are taken as samples, and a set formed by a plurality of the samples in a variety of track environments is taken as the environmental image data set.
[0074] For the environmental image data set, the turnouts in the camera-photographed images and the corresponding grayscale images converted from the point cloud data in the samples are annotated to obtain the annotation results.
[0075] In some embodiments, the annotation result includes the precise position of the turnout and the opening direction of the turnout. Specifically, a rectangular frame can be used to annotate the spatial range of the turnout in the image, and the left or right opening state of the turnout can be recorded as the opening direction annotation result of the turnout.
[0076] In some embodiments, the labeling process of the present invention can be implemented by full manual labeling or by first performing rough machine labeling and then manual fine labeling and verification.
[0077] For the environmental image dataset, the camera images, grayscale images converted from point cloud data and corresponding annotation results in the samples are paired as a training sample, and a set formed by multiple training samples is used as the turnout image training set.
[0078] Step S3, establishing a turnout recognition network, the input data of the turnout recognition network is processed by the backbone network, the bidirectional feature network and the detection head in sequence to obtain the turnout position and category prediction results; the network depth of the bidirectional feature network increases linearly, and the network width increases exponentially.
[0079] like Figure 2 As shown in the figure, the turnout recognition network of the present invention is composed of a backbone network, a bidirectional feature network and a detection head. The camera image and the grayscale image converted from the corresponding point cloud data are used as inputs, and are processed by the backbone network, the bidirectional feature network and the detection head for the first time to obtain the turnout category prediction result. The specific processing process of the turnout recognition network is:
[0080] (1) EfficientNets pre-trained on the large-scale visualization database ImageNet is used as the backbone network of the present invention, and the backbone network is used to extract multi-scale features of the image. Compared with YOLO or RetinaNet commonly used in the field of object detection, the backbone network is more efficient and can achieve higher accuracy with fewer parameters and floating point operations per second (FLOPs).
[0081] The depth of the network refers to the number of layers in the network, which may include convolutional layers, pooling layers, and fully connected layers; the width of the network refers to the number of convolution kernels in each layer of the network.
[0082] The width and depth of the backbone network of the present invention follow the width and depth of the network variants from B0 to B6 of the EfficientNet series. These variants are designed with a unified width / depth scaling factor, where B0 is the base network and B1 to B6 are larger networks that are gradually expanded through compound scaling, so that their ImageNet pre-trained weights can be reused.
[0083] After the input image is processed by the backbone network, multi-scale features are obtained, and the multi-scale features are specifically the output features of the 3rd to 7th layers of the backbone network.
[0084] (2) The output features of the 3rd to 7th layers of the backbone network are input into the bidirectional feature network for fusion to obtain multi-scale fusion features. When fusing feature maps of different resolutions, they are usually first integrated into a uniform resolution and then superimposed. All previous works treat all input features equally without distinction. Since different input features have different resolutions, they often make different contributions to the output features. In order to solve this problem, the present invention proposes to add an additional weight to each input so that the network learns the importance of each input feature.
[0085] The weighted fusion expression of the bidirectional feature network provided by the present invention is:
[0086]
[0087] Among them, I i represents the i-th input feature map, ω i represents the learnable weight corresponding to the i-th input feature map, ∈ is a stable constant, ∈=0.0001 is a very small value to avoid the situation where the denominator is 0 and the numerical instability occurs, ∑ j ω j Represents the sum of the learnable weights corresponding to all input feature maps.
[0088] Through the above weighted fusion process, different weights are assigned to each feature map according to its importance, and these weights are automatically adjusted in a learnable way. Weighted fusion can help the network better utilize features of different resolutions or scales. The fast fusion method is more efficient and can speed up training on the GPU.
[0089] The bidirectional feature network of the present invention is composed of a plurality of bidirectional feature network layers connected in sequence. The structure of the bidirectional feature network layer is as follows: Figure 3 As shown in Figure 1, the output features of the 3rd to 7th layers of the backbone network are processed by multiple bidirectional feature network layers in sequence, and finally multi-scale fusion features are obtained.
[0090] For the output of the 3rd to 7th layers of the backbone network, the resolution is reduced in sequence. The present invention first linearly enlarges the resolution of the output feature map of the 3rd to 7th layers, and the expression is:
[0091]
[0092] Among them, R input is the base size of the feature map after enlargement, Represents the network scale magnification composite coefficient.
[0093] By using the above expression, we can adjust the growth parameter The size of determines the base size of the output feature map of the 3rd to 7th layers after enlargement, and the length and width of the output feature map are enlarged in equal proportion with the base size. The enlarged feature map is input into the subsequent bidirectional feature network.
[0094] The width and depth of the bidirectional feature network of the present invention satisfy that the depth increases linearly and the width increases exponentially, and the expression is:
[0095]
[0096] Among them, W 双向特征网络 It represents the width of the bidirectional feature network, that is, the number of convolution kernels in each layer of the bidirectional feature network, D 双向特征网络 represents the depth of the bidirectional feature network, that is, the number of layers of the bidirectional feature network, It represents the network scale magnification compound coefficient, and α is the base of width growth.
[0097] In some embodiments, the width growth base α of the present invention is obtained by a hyperparameter optimization method such as a grid search method. Preferably, the value of α is α=1.35.
[0098] The width and depth of the bidirectional feature network of the present invention are designed to meet the strategy of linear increase in depth and exponential increase in width, aiming to effectively improve the representation ability and computational efficiency of the network. The linear increase in depth ensures that the network can gradually capture more complex feature levels, while the exponential increase in width allows the network to process more feature channels at each layer, thereby improving the diversity and expression of features. The above approach balances computational complexity and performance improvement, and achieves higher detection accuracy using limited computing resources.
[0099] (3) The multi-scale fusion features output by the bidirectional feature network are input into the detection head to obtain the turnout position and category prediction results.
[0100] The detection head of the present invention includes two branches, namely, a classification head and a regression head. The processing expression of the classification head is:
[0101]
[0102] Among them, p c Represents the category confidence of the opening direction of the turnout, w c represents the weight of the classification head, F represents the multi-scale fusion feature of the input, and Sigmoid(·) represents the Sigmoid function.
[0103] The processing expression of the regression head is:
[0104] b=Regression(F)
[0105] Where b represents the predicted turnout position bounding box parameter, b = [x min ,y min ,x max ,y max ], x min ,y min ,x max ,y max Represent the horizontal and vertical coordinates of the upper left corner and the lower right corner of the bounding box respectively, and Regression(·) represents the regression process.
[0106] The present invention establishes a turnout recognition network including a backbone network, a bidirectional feature network and a detection head through the above process, and will subsequently be trained.
[0107] Step S4, training the turnout recognition network based on the turnout image training set to obtain a trained turnout recognition network;
[0108] In this step, the loss function of the turnout recognition network of the present invention is set, and the expression is:
[0109] FL(p t )=-α t (1-p t ) γ log(p t )
[0110] Among them, p t Represents the predicted category probability, FL(p t ) represents the class probability p t The corresponding network loss, α t represents the weight of balancing positive and negative samples, and γ is the focus adjustment factor, which is set to 2 in the present invention.
[0111] The present invention adopts a freeze and unfreeze training method for the training process of the turnout recognition network, which can efficiently utilize the characteristics of the pre-trained model backbone network, and specifically includes two stages: the first stage freezes the backbone network, that is, keeps the network weights of the backbone network unchanged, and only trains the network weights of the detection head and the bidirectional feature network; the second stage unfreezes the backbone network, trains the network weights of the backbone network, the bidirectional feature network and the training detection head, and finally obtains a trained turnout recognition network.
[0112] In some embodiments, during the freezing and unfreezing training process, an error back propagation algorithm such as a gradient descent method may be used to minimize the loss function of the turnout recognition network as an optimization goal to complete the optimization adjustment of the network weights.
[0113] Step S5: Use the trained turnout recognition network to process the environmental image data collected in real time to obtain the turnout position and category prediction results.
[0114] In this step, the trained turnout recognition network is deployed in an embedded device, and the camera images and the corresponding grayscale images converted from point cloud data are collected in real time by the data acquisition equipment deployed on the train. The camera images and the grayscale images are input into the trained turnout recognition network, and finally the real-time prediction results of the turnout position and category are obtained.
[0115] Although the specific embodiments of the present invention depict various actions or steps in a specific order, this should be understood as requiring such actions or steps to be performed in the specific order shown or in a sequential order, or requiring all illustrated actions or steps to be performed to obtain the desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of a separate embodiment can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation can also be implemented in multiple implementations individually or in any suitable sub-combination. The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
[0116] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A track turnout visual recognition method based on a bidirectional feature network, characterized in that: The following steps are involved: Step S1, installing a point cloud radar and a camera on a data collection vehicle, and adjusting the internal and external parameters of the point cloud radar and the camera; Step S2, deploying the data collection vehicle to the train track, collecting multiple image data and point cloud data containing turnouts in various track environments by the point cloud radar and camera on the data collection vehicle, and establishing an environmental image data set; marking the turnouts in the environmental image data, and establishing a turnout image training set; Step S3, establishing a turnout recognition network, wherein the input data of the turnout recognition network is processed by the backbone network, the bidirectional feature network and the detection head in sequence to obtain the turnout position and category prediction results; the network depth of the bidirectional feature network increases linearly, and the network width increases exponentially; Step S4, training the turnout recognition network based on the turnout image training set to obtain a trained turnout recognition network; Step S5: Use the trained turnout recognition network to process the environmental image data collected in real time to obtain the turnout position and category prediction results.
2. The track turnout visual recognition method based on a bidirectional feature network according to claim 1 is characterized in that: The step S1 specifically includes: Step S1-1, setting a camera bracket on the data collection vehicle, and fixing a camera and a point cloud radar on the camera bracket at the same time; Step S1-2, calibrating the camera internal parameters to correct the distortion of the image captured by the camera; Step S1-3: calibrate the external parameters of the point cloud radar and the camera so that the data collected by the point cloud radar and the camera match each other.
3. The track turnout visual recognition method based on a bidirectional feature network according to claim 2 is characterized in that: The step S2 specifically includes: Step S2-1, deploying a data collection vehicle on the train track, and using the calibrated point cloud radar and camera to collect multiple camera-photographed images and point cloud data including turnouts in various track environments; Step S2-2, converting the point cloud data into a grayscale image, taking an image captured by a camera shooting the same scene and a grayscale image converted from the corresponding point cloud data as samples, and taking a set formed by a plurality of the samples after data enhancement as the environmental image data set; Step S2-3, annotate the turnouts in the camera-captured images and the corresponding grayscale images converted from the point cloud data in the samples in the environmental image data set to obtain the annotation results; use the camera-captured images, the grayscale images converted from the point cloud data and the corresponding annotation results as a training sample, and use a set formed by multiple training samples as the turnout image training set.
4. The track turnout visual recognition method based on a bidirectional feature network according to claim 3 is characterized in that: The data enhancement step in step S2-2 specifically includes: Rotate, crop and adjust the color of the grayscale image converted from the camera image and the corresponding point cloud data to form multiple new samples; The step of obtaining the labeling result in step S2-3 specifically includes: The camera image and the corresponding grayscale image converted from the point cloud data are marked with the precise position of the turnout and the opening direction of the turnout. The marking result includes a rectangular frame of the position of the turnout in the image and the opening state of the turnout, and the opening state includes left or right.
5. The track turnout visual recognition method based on a bidirectional feature network according to claim 4 is characterized in that: The step S3 specifically includes: Step S3-1, using the pre-trained EfficientNets as the backbone network, wherein the width and depth of the backbone network are set to the width and depth of the network variants of EfficientNet from B0 to B6; Step S3-2, establishing a bidirectional feature network, the bidirectional feature network is composed of a plurality of bidirectional feature network layers connected sequentially, the output feature maps of the 3rd to 7th layers of the backbone network are linearly amplified, and then sequentially subjected to weighted fusion processing of the plurality of bidirectional feature network layers, and finally a multi-scale fusion feature is obtained; the number of network layers of the bidirectional feature network increases linearly, and the number of convolution kernels in each layer of the network increases exponentially; Step S3-3, establish a detection head, input the multi-scale fusion features output by the bidirectional feature network into the detection head, and process them through the classification head and regression head respectively to obtain the turnout position and category prediction results.
6. The track turnout visual recognition method based on a bidirectional feature network according to claim 5 is characterized in that: In step S3-2, the step of linearly amplifying the output feature map of the 3rd to 7th layers of the backbone network includes: adjusting the growth parameter The size of determines the base size of the output feature map of the 3rd to 7th layers after enlargement, and the length and width of the output feature map are enlarged in proportion to the base size. The expression of the base size is: Among them, R input is the base size of the feature map after enlargement, Represents the network scale magnification composite coefficient; The expression of the weighted fusion processing of the bidirectional feature network layer is: Among them, I i represents the i-th input feature map, ω i represents the learnable weight corresponding to the i-th input feature map, ∈ is a stable constant, ∑ j ω j Represents the sum of the learnable weights corresponding to all input feature maps; The expression of the number of network layers of the bidirectional feature network and the number of convolution kernels in each layer of the network is: Among them, W 双向特征网络 It represents the width of the bidirectional feature network, that is, the number of convolution kernels in each layer of the bidirectional feature network, D 双向特征网络 represents the depth of the bidirectional feature network, that is, the number of layers of the bidirectional feature network, It represents the network scale magnification compound coefficient, and α is the base of width growth.
7. The track turnout visual recognition method based on a bidirectional feature network according to claim 6 is characterized in that: The step S3-3 specifically includes: A detection head is established. The detection head includes two branches, namely, a classification head and a regression head. The processing expression of the classification head is: Among them, p c Represents the category confidence of the opening direction of the turnout, w c represents the weight of the classification head, F represents the multi-scale fusion feature of the input, and Sigmoid(·) represents the Sigmoid function; The processing expression of the regression head is: b=Regression(F) Where b represents the predicted turnout position bounding box parameter, b = [x min ,y min ,x max ,y max ], x min ,y min ,x max ,y max Represent the horizontal and vertical coordinates of the upper left corner and the lower right corner of the bounding box respectively, and Regression(·) represents the regression process.
8. The track turnout visual recognition method based on a bidirectional feature network according to claim 7 is characterized in that: The step S4 specifically includes: Step S4-1, determining the loss function of the turnout recognition network; Step S4-2, keeping the network weights of the backbone network unchanged, based on the training samples of the turnout image training set, taking minimization of the loss function as the optimization goal, and only adjusting the network weights of the training detection head and the bidirectional feature network; Step S4-3, making the network weight of the backbone network variable, based on the training samples of the turnout image training set, taking minimizing the loss function as the optimization goal, adjusting the network weights of the backbone network, the bidirectional feature network and the training detection head, and finally obtaining a trained turnout recognition network.
9. The track turnout visual recognition method based on a bidirectional feature network according to claim 8, characterized in that: In step S4-1, the loss function expression of the turnout recognition network is: FL(p t )=-a t (1-p t ) γ log(p t ) Among them, p t Represents the predicted category probability, FL(p t ) represents the class probability p t The corresponding network loss, α t represents the weight of balancing positive and negative samples, and γ is the focus adjustment factor.
Citation Information
Patent Citations
Method and system for detecting railway turnout track structure based on 3D point cloud
CN118149727A