A method for dividing track lines for high-precision detection of the end of a track
By designing the track line segmentation network RTEW-Net, the full attention calculation, global response normalization and intelligent weight maintenance modules are used to solve the problem that traditional methods are difficult to achieve accurate segmentation when the characteristics at the end of the track line disappear or are unclear, and high-precision detection and precise segmentation of the end of the track line are achieved, which improves the safety of train operation.
Patent Information
- Application Number
- CN202411654579.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-19
AI Technical Summary
The traditional track line segmentation method is difficult to achieve precise segmentation when the characteristics at the end of the track line disappear or are unclear, resulting in the safety of train operation.
A track line segmentation network RTEW-Net is designed, adopting an encoding-decoding structure, including the full attention calculation module FTM, the global response normalization module GRN and the intelligent weight maintenance module WWM. Through these modules, the characteristic information at the end of the track line is extracted and retained to achieve accurate segmentation.
Effectively extract and deeply maintain the characteristic information at the end of the track line, realize the precise segmentation of the end of the track line, improve the safety of train operation, and have good generalization capabilities.
Smart Images

Figure CN119540555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of track line segmentation, and particularly to a track line segmentation method for high-precision detection of the track end. Background Art
[0002] Rail transit is an important part of the modern transportation system, and the high-precision segmentation of tracks is a key technology to ensure the safe automatic driving of intelligent trains. The running speed of the train is relatively fast, and the end distance within the field of vision reaches in a short time. To ensure the safety of train operation, the high-precision and accurate detection of the railway track end is particularly important.
[0003] Currently, limited by the performance of data acquisition devices, the features at the track end often degenerate, resulting in a decrease in detection accuracy. In complex actual track operation scenarios, such as the entrance and exit areas of tunnels, due to the sharp change in light, the difficulty of feature extraction at the track end is further increased. This feature loss makes the existing track segmentation methods have poor detection effects at the track end, seriously affecting the safety of intelligent trains.
[0004] Traditional track segmentation methods have achieved certain results in the field of track line segmentation, but many studies focus on the overall visual performance of the track, showing great limitations in the detection effect of the end in complex scenarios. Therefore, although the existing track segmentation methods perform well in common scenarios, their performance still has great room for improvement in the complex scenarios at the track end. To address these challenges, it is particularly important to develop a segmentation method specifically for track end detection to further improve the safety of train operation. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a track line segmentation method for high-precision detection of the track end, aiming to solve the problem that it is difficult for traditional track line segmentation methods to achieve accurate segmentation of the track line end with disappearing or unclear features, and to achieve accurate segmentation of the track line, thereby improving the safety of train operation.
[0006] To achieve the above object, the present invention provides a track line segmentation method for high-precision detection of the track end, including:
[0007] Constructing the image data of the track line to be segmented;
[0008] Inputting the image data of the track line to be segmented into the track line segmentation network RTEW-Net for processing to obtain the segmentation result of the train running track line;
[0009] Among them, the rail line segmentation network RTEW-Net is used to extract the preliminary features of the rail line end from the train operation rail line feature map, maintain the feature details of the preliminary features, and retain the rail line end features in combination with the train operation rail line feature map; the rail line segmentation network RTEW-Net is obtained by training with a training set, and the training set is rail line images and corresponding annotations.
[0010] Preferably, constructing the training set includes:
[0011] Collect the initial rail line images in the real train operation scenario, perform manual annotation processing on the initial rail line images and then screen them, and combine the screened images with the initial rail line images to construct the RailMixed2024 dataset, that is, the training set; among them, the real train operation scenario includes image data with the disappearance or ambiguity of the rail line end features caused by entering and exiting tunnels and extremely bad weather conditions.
[0012] Preferably, the structure of the rail line segmentation network RTEW-Net is an encoder-decoder structure. The encoder structure includes a full attention calculation module FTM, a global response normalization module GRN, and an intelligent weight maintenance module WWM connected in sequence, and the decoder is a Uperhead decoder.
[0013] Preferably, inputting the rail line image data to be segmented into the rail line segmentation network RTEW-Net for processing includes:
[0014] Performing feature dimension elevation on the rail line image data to be segmented through the Patch Embedding method to obtain a feature map, and performing normalization processing on the size of the feature map;
[0015] Performing attention calculation on the normalized feature map through the full attention calculation module FTM, aggregating global features of the feature map through the global response normalization module GRN, and performing recalibration to output a normalized feature map;
[0016] Combining the feature map processed by the GRN with the normalized feature map through the intelligent weight maintenance module WWM to output the train operation rail line segmentation result.
[0017] Preferably, performing feature dimension elevation on the rail line image data to be segmented through the Patch Embedding method to obtain a feature map includes:
[0018] Based on the above-mentioned rail line segmentation network RTEW-Net, the rail line image data to be segmented is subjected to deep feature downsampling through PatchEmbedding operation and Full Transformer Block calculation, and downsampling is performed through PatchMerging. This process is repeated iteratively until the feature map is output;
[0019] Among them, each downsampling operation is combined with Block calculation to form an operation stage. During the Block calculation process, LN operation, FTM module calculation, GELU activation, and MLP classifier calculation are sequentially executed.
[0020] Preferably, attention calculation is performed through the full attention calculation module FTM, including: Base-patch calculation and pad-patch calculation;
[0021] The method of Base-patch calculation is:
[0022] FTM 1 (x) = ReverseSeparate(attn(Separate(x)))
[0023] The method of pad-patch calculation is:
[0024] FTM 2 (x) = RmPad(FTM 1 (Pad(x)))
[0025] In the formula, Pad and RmPad respectively represent the operations of expanding and removing padding according to the patch shape, Separate and ReverseSeparate respectively represent the operations of splitting the feature map into patches and restoring, attn represents performing attention calculation, FTM 1 (x) is the result of Base-patch calculation, FTM 2 (x) is the result of pad-patch calculation, and x is the input feature Figure 1 ;
[0026] During one Block calculation, Base-patch calculation is first performed, and then pad-patch calculation is performed.
[0027] Preferably, the method for the global response normalization module GRN to aggregate global features of the feature map and perform recalibration is:
[0028] X i+1 = γ * N(X i ) + β + X i
[0029] Wherein, γ and β are trainable parameters, and X i+1 is the output feature Figure 1 , N(X i ) is a regularization method, and X i is the input feature map.
[0030] Preferably, the intelligent weight maintenance module WWM retains the information of the normalized feature map and the processed feature map through an adaptively adjusted weight maintenance method, and sets an adaptive parameter participating in global optimization to organically combine the information of the normalized feature map and the processed feature map, so as to realize the adaptive retention and adjustment of the detailed features in the feature map by the model.
[0031] Preferably, training the track line segmentation network RTEW-Net includes:
[0032] Optimizing the parameters of RTEW-Net through the cross-entropy loss function and the adamW optimizer, and performing an intelligent residual connection operation to obtain the trained track line segmentation network RTEW-Net.
[0033] Preferably, the method for performing the intelligent residual connection operation is to add trainable weights to the calculated feature map and optimize them during the network training process. Specifically:
[0034] x out = ω 1 *x 1 + ω 2 *x 2
[0035] Wherein, x out is the output feature Figure 2 , ω 1 and ω 2 are respectively trainable parameters independently participating in global optimization, and x 1 and x 2 are respectively the feature map information processed by different modules.
[0036] Compared with the prior art, the present invention has the following advantages and technical effects:
[0037] The present invention realizes the effective extraction of the disappearing or unclear features at the end of the track line, and maintains the depth, completes the accurate segmentation of the end of the track line, and has good generalization, so as to complete the complete and accurate perception of the train operation track and ensure the train operation safety.
[0038] The present invention designs a full attention calculation module, namely Full-Transformer Module (FTM), which improves the calculation effect of traditional attention on subtle features and realizes the preliminary retention of the feature information at the end of the track line. The global response normalization module, namely Global Response Normalization (GRN), is introduced to aggregate global features and recalibrate them, achieving normalized feature output, improving the maintenance of feature details, and coping with complex track scenarios and drastic illumination changes. The intelligent weight maintenance module, namely Wise Weigh Maintain (WWM), is designed to adaptively adjust the intensity of feature learning for each layer, realize the deep retention of unclear feature details at the end of the track line, and deepen the detection ability of the model for the end of the track line. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0040] Figure 1 It is a display diagram of the track image and its label of the RM2024 dataset in the embodiment of the present invention;
[0041] Figure 2 It is a display diagram of the track image and its label of the RS19 dataset in the embodiment of the present invention;
[0042] Figure 3 It is a flowchart of the track line segmentation of the RTEW-Net in the embodiment of the present invention;
[0043] Figure 4 It is the overall structure diagram of the RTEW-Net in the embodiment of the present invention;
[0044] Figure 5 It is a calculation flowchart of the full attention calculation module in the embodiment of the present invention;
[0045] Figure 6 It is a visualization schematic diagram of the prediction results of the RTEW-Net and the SOTA model on the RM2024 dataset in the embodiment of the present invention;
[0046] Figure 7 It is a visualization schematic diagram of the prediction results of the RTEW-Net and the SOTA model on the RS19 dataset in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0048] Note that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] The present invention proposes a track line segmentation method for high-precision detection of the track end, including:
[0050] Construct the track line image data to be segmented;
[0051] Input the track line image data to be segmented into the track line segmentation network RTEW-Net for processing to obtain the train operation track line segmentation result;
[0052] Among them, the track line segmentation network RTEW-Net is used to extract the preliminary features of the track end from the train operation track line feature map, maintain the feature details of the preliminary features, and retain the track end features in combination with the train operation track line feature map; the track line segmentation network RTEW-Net is obtained by training with a training set, and the training set is track line images and corresponding annotations.
[0053] Furthermore, the structure of the track line segmentation network RTEW-Net is an encoder-decoder structure. The encoder structure includes a full attention calculation module FTM, a global response normalization module GRN, and an intelligent weight maintenance module WWM connected in sequence, and the decoder is a Uperhead decoder.
[0054] Specifically, this embodiment proposes a track line segmentation network (RTEW-Net) for high-precision detection of the track end. The network model sequentially integrates a full attention calculation module (FTM), a global response normalization module (GRN), and an intelligent weight maintenance module (WWM). The three modules are effective in series through a linear calculation process, effectively improving the performance of the algorithm, enabling the method to adapt to complex track environments, and particularly having a good recognition effect on the track end, achieving precise segmentation of the track line and ensuring the safety guarantee of train operation.
[0055] Furthermore, construct a training set, including:
[0056] Collect the initial track line images in the real train operation scenario, perform manual annotation processing on the initial track line images and then screen them, and combine the screened images with the initial track line images to construct the RailMixed2024 (RM2024) dataset, that is, the training set; among them, the real train operation scenario includes image data with the track end features disappearing or being unclear due to entering and exiting tunnels and extreme bad weather conditions.
[0057] Specifically, refer toFigure 1 , in this embodiment, the track line image data in the real train operation scenario is collected. The scenario is complex and rich, including the image data with the disappearance or ambiguity of the end characteristics of the track line caused by factors such as entering and leaving the track and extremely bad weather. The collected images are manually labeled. Referring to Figure 2 the rs19 dataset shown, some of the images are selected from it, and after processing, they are combined with the collected images to construct the RailMixed2024 dataset. There are a total of 2142 images collected on site, and 800 images are selected from the RS19 dataset, with a total of 2942 track images. They are divided according to the ratio of 9:1 for the training set and the test set.
[0058] Furthermore, the track line image data to be segmented is input into the track line segmentation network RTEW-Net for processing, including:
[0059] The feature dimension of the track line image dataset is increased through the Patch Embedding method to obtain a feature map, and the size of the feature map is normalized.
[0060] The normalized feature map is subjected to attention calculation through the full attention calculation module FTM, and the global features of the feature map are aggregated through the global response normalization module GRN, and recalibration is performed to output the normalized feature map.
[0061] The feature map processed by the GRN and the normalized feature map are combined through the intelligent weight maintenance module WWM to output the segmentation result of the train operation track line.
[0062] Specifically, as Figure 3 , in this embodiment, RTEW-Net is proposed. The overall structure of this network is an encoder-decoder structure. In this embodiment, Uperhead is selected as the decoder. Uperhead is a general decoder that can cooperate with a specific encoder to achieve feature fusion and result prediction. The encoder structure is comprehensively designed through FTM, GRN, and WWM to achieve efficient feature extraction of the track line. In the encoder part, the image is downsampled to obtain a high-dimensional feature map. Before inputting into the encoder, the image undergoes a reconstruction operation, and the pixel values are adjusted to be between 0 and 1 and conform to the normal distribution state.
[0063] Furthermore, the feature dimension of the track line image dataset is increased through the Patch Embedding method to obtain a feature map, including:
[0064] Based on the above-mentioned Rail Track Line Segmentation Network (RTEW-Net), the rail track line image data to be segmented is subjected to deep feature downsampling through PatchEmbedding operation and Full Transformer Block (FTBlock) calculation, and downsampling is performed through Patch Merging. This process is iteratively repeated until the feature map is output.
[0065] Among them, each downsampling operation combined with the Block calculation constitutes an operation stage. During the Block calculation process, the LN operation, FTM module calculation, GELU activation, and MLP classifier calculation are sequentially executed.
[0066] Specifically, referring to Figure 4 , the encoder of the RTEW-Net proposed in this embodiment is a hierarchical structure design as a whole. In the first calculation, the image is subjected to deep feature downsampling through Patch Embedding operation and Full Transformer Block calculation, and then downsampling is performed through Patch Merging. This process is iterated four times. Each downsampling operation combined with the Block calculation constitutes an operation stage. During the Block calculation process, the LN operation, FTM module calculation, GELU activation, and MLP classifier calculation are sequentially executed. The MLP module is derived from ViT and Swin, and makes up for the inherent local calculation problem of the attention mechanism through a fully connected operation. During a single Block calculation process, the operation of the feature map is as follows:
[0067] x stage+1 = FTBlock(DownSampling(x stage ))
[0068] In the formula, x stage+1 is the output feature after a single Block calculation Figure 3 , DownSampling is the downsampling operation, which is the Patch Embedding operation in the first layer of Block calculation, and the Patch Merging operation otherwise. x stage is the input feature Figure 2 ;
[0069] For the FTBlock, a complete calculation process is as follows:
[0070] x i+1 = ω 1 * GRN(GELU(FTM 1 (LN(x i )))) + ω 2 * x i
[0071] x i+2 = ω 3 *MLP(LN(x i+1 )) + ω 4 *x i+1
[0072] x i+3 = ω 5 *GRN(GELU(FTM 2 (LN(x i+2 )))) + ω 6 *x i+2
[0073] x i+4 = ω 7 *MLP(LN(x i+3 )) + ω 8 *x i+3
[0074] In the formula, x i is the input feature Figure 3 , x i+1 , x i+2 , x i+3 are the output feature maps of their respective layers and the input feature maps for the calculation of the next layer. From ω 1 to ω 8 are trainable parameters independently participating in global optimization. FTM 1 As shown in Figure 5 , is the mathematical implementation of the Base-Patch module. FTM 2 As shown in Figure 5 , is the mathematical implementation of the Pad-Patch module. GRN is the global response normalization module, MLP is the fully connected calculation method, GELU is a common activation function for neural networks, and LN is a common regularization operation for neural networks.
[0075] For an input image with a size of H×W×3, the Patch embedding operation first changes its shape to After each stage of operation, the height and width of the feature map are halved, and the number of channels is doubled. The final size is 8×C.
[0076] Furthermore, attention calculation is performed through the full attention calculation module FTM, including: Base-patch calculation and pad-patch calculation; to maintain and settle the symbol consistency and facilitate understanding of the structures of base-patch and pad-patch.
[0077] The method of Base-patch calculation is:
[0078] FTM 1(x) = ReverseSeparate(attn(Separate(x)))
[0079] The method for pad - patch calculation is as follows:
[0080] FTM 2 (x) = RmPad(FTM 1 (Pad(x)))
[0081] In the formula, Pad and RmPad respectively represent the operations of expanding and removing padding according to the patch shape, Separate and ReverseSeparate respectively represent the operations of splitting the feature map into patches and restoring, attn represents performing attention calculation, and FTM 1 (x) is the Base - patch calculation result, and FTM 2 (x) is the pad - patch calculation result;
[0082] In one Block calculation, first perform Base - patch calculation, and then perform pad - patch calculation.
[0083] Specifically, referring to Figure 5 , the FTM module is the core calculation module of the feature extraction network, in which the main feature extraction operations are implemented. The design inspiration of this module comes from ViT, Swin Transformer, MobileViT, LPT and their variants. The basic goal is to better extract the overall features of the target and achieve global attention interaction at low computational cost.
[0084] For Figure 5 the Base - patch calculation method in, initially divide the feature map with patches as the basic calculation unit according to preset parameters and perform attention calculation. After the depth connection of the divided and calculated patches, each individual patch is regarded as a sequence unit similar to the NLP task and input into the Transformer for calculation. After the calculation is completed, the feature map is restored according to the reverse steps of the segmentation.
[0085] Next is the pad - patch calculation method. Based on the size of the divided patches, the feature map is expanded by half of the patch size in four directions. The expanded image is re - divided according to the patch size and the subsequent attention calculation is performed as in Base - patch. After the Transformer calculation is completed, the feature map is restored to the original shape and the padding elements in four directions are removed to obtain the restored feature map.
[0086] Furthermore, the method for the Global Response Normalization module (GRN) to aggregate global features of the feature map and perform recalibration is as follows:
[0087] X i+1 = γ * N(X i ) + β + X i
[0088] In the formula, γ and β are trainable parameters, and X i+1 is the output feature Figure 1 , X i is the input feature map, and N(x) is a regularization method, which is defined as follows:
[0089]
[0090] In the formula, H and W respectively represent the height and width of the feature map, and ∈ is a small real number, and x ij represents the pixel at the coordinate (i, j) of the feature map.
[0091] Specifically, the purpose of the Global Response Normalization (GRN) module is to enhance channel contrast and selectivity. In this embodiment, GRN is used for normalizing the feature map after attention calculation, not before the MLP calculation, but applying it after the FTM module calculation.
[0092] Furthermore, the Intelligent Weight Maintenance module (WWM) retains the information of the normalized feature map and the processed feature map through an adaptive weight maintenance method, and sets adaptive parameters participating in global optimization to organically combine the information of the normalized feature map and the processed feature map, so as to enable the model to adaptively retain and adjust the detailed features in the feature map.
[0093] Specifically, the Intelligent Weight Maintenance module (WWM) designs an adaptive weight maintenance method to retain the information of the feature map before the first calculation and the processed feature map, and sets adaptive parameters participating in global optimization to organically combine the two feature maps, thereby allowing the model to adaptively retain and adjust the detailed features in the feature map, deepen the sensitivity to subtle features, and achieve accurate segmentation of the end of the track line.
[0094] Furthermore, training the Rail Track Endpoint Segmentation Network (RTEW-Net) includes:
[0095] Optimizing the parameters of RTEW-Net through the cross-entropy loss function and the adamW optimizer, and performing intelligent residual connection operations to obtain the trained Rail Track Endpoint Segmentation Network (RTEW-Net).
[0096] The method for performing the intelligent residual connection operation is to add trainable weights to the calculated feature map and optimize them during the network training process. Specifically:
[0097] x out =ω 1 *x 1 +ω 2 *x 2
[0098] In the formula, x out is the output feature Figure 2 ,ω 1 and ω 2 are trainable parameters independently participating in global optimization respectively, and x 1 and x 2 are the feature map information processed by different modules respectively.
[0099] Specifically, the purpose of the intelligent residual connection operation is to improve the model's ability to retain sparse features, thereby enhancing the extraction effect of the disappearing or unclear features at the end of the track line. The traditional residual connection operation inevitably reduces the model's feature learning ability to some extent. This reduction occurs when each newly learned feature is added to the previously learned feature, resulting in a weakening of the intensity of the newly learned feature.
[0100] To solve this problem and enhance the model's learning ability, in this embodiment, trainable weights are added to the feature map involved in the calculation and optimized during the network training process.
[0101] This operation enables the model to adaptively adjust the intensity of feature learning for each layer, thereby improving the model's learning ability at the end of the track while maintaining the stability advantage of the original residual connection.
[0102] Referring to Figure 6 , this embodiment shows the visual comparison of the method of the present invention with other SOTA methods on the RM2024 dataset. The comparison methods include Swin Transfrormer, Vision Transformer, Lraspp, BiFormer, ConvNeXtV2, and the RTEW-Net of the present invention.
[0103] In the railway images of the first four rows, the railway endpoints are affected by environmental interference to varying degrees: the first row is affected by the highlight at the tunnel exit, the fourth row is affected by the shadow in the middle of the tunnel, and the second and third rows are affected by lens contamination. These factors lead to the degradation of the railway endpoint features. RTEW-Net effectively reconstructs these features using the FTM module, achieving better railway endpoint segmentation. In contrast, the Swin model has difficulties in segmenting the railway endpoint on the right side of the first row, has problems in segmenting large-scale targets on the right side and the railway endpoint on the left side of the second row, and performs poorly in segmenting the railway endpoint on the left side of the third row. ViT and Lraspp perform well on some images but poorly in complex scenarios. BiFormer and ConvNeXt accurately segment and locate the railway position, but the overall segmentation accuracy is low, and they almost completely fail to segment the endpoints. The fifth row shows a typical train operation scenario. Swin still performs poorly in segmenting the railway endpoint on the left side, while only RTEW-Net accurately segments the endpoint, demonstrating the effectiveness of the RTEW-Net module design.
[0104] The specific comparison metrics are shown in Table 1, presenting the main test targets and mIoU data of rs19. These targets are labeled using an improved semi-automatic method, and the training results are better than other categories. As shown in Table 1, RTEW-Net achieves the best results on the railway-related targets of rs19. The IoU of the most critical target, rail track, is improved by 3.2 compared to the second-best model. The IoUs of rail raised and rail embedded are improved by 4.5 and 6.6 respectively. The overall mIoU is improved by 3.6, showing a significant advantage over other models. This indicates that RTEW-Net has superior feature extraction ability for railway-related targets, verifying the effectiveness of the proposed algorithm.
[0105] Refer to Figure 7 , presenting the performance of different models in railway endpoint segmentation. In Figure 7Among them, the first five lines show the inference results of the test set, and the last four lines show the results of the training set. The first column is the originally collected image, the second column is the ground truth, the third column is the result of RTEW-Net, and the remaining columns are the inference results of the comparison models. Visually, RTEW-Net provides the best segmentation result, almost completely segmenting out all railway targets. In the railway images of the first four lines, the railway endpoints are affected by environmental interference to varying degrees: the first line is affected by the high brightness at the tunnel exit, the fourth line is affected by the shadow in the middle section of the tunnel, and the second and third lines are affected by lens contamination. These factors lead to the degradation of the railway endpoint features. RTEW-Net effectively reconstructs these features using the FTM module, achieving better railway endpoint segmentation. On the contrary, the Swin model has difficulties in segmenting the railway endpoint on the right side of the first line, has problems in segmenting large-scale targets on the right side and the railway endpoint on the left side of the second line, and performs poorly in segmenting the railway endpoint on the left side of the third line. ViT and Lraspp perform well on some images but poorly in complex scenarios. BiFormer and ConvNeXt accurately segment and locate the railway position, but the overall segmentation accuracy is low, and they almost completely fail to segment the endpoints. The fifth line shows a typical train operation scenario. Swin still performs poorly in segmenting the railway endpoint on the left side, while only RTEW-Net accurately segments the endpoint, demonstrating the effectiveness of the RTEW-Net module design.
[0106] The specific comparison metrics for the rs19 dataset are shown in Table 2, which presents the main test targets and mIoU data of rs19. RTEW-Net achieved the best results for railway-related targets in rs19. The IoU of the most critical target, rail track, was improved by 3.2 compared to the second-best model. The IoUs of rail raised and rail embedded were improved by 4.5 and 6.6 respectively. The overall mIoU was improved by 3.6, showing a significant advantage over other models. This indicates that RTEW-Net has superior feature extraction capabilities for railway-related targets, verifying the effectiveness of the proposed algorithm.
[0107] Table 1
[0108]
[0109] Table 2
[0110]
[0111] Based on the comprehensive performance of the modules of the present invention, it can adapt to complex track operation scenarios, effectively extract new information about the track ends with unclear features in extreme cases, thereby achieving accurate segmentation of the track line, ensuring the safety of train travel, and having high practical value.
[0112] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A track line segmentation method for high-precision detection of track ends, characterized in that: include: Constructing the track line image data to be segmented; The track line image data to be segmented is input into the track line segmentation network RTEW-Net for processing to obtain the train running track line segmentation result; The track segmentation network RTEW-Net is used to extract preliminary features of the track end from the train running track feature map, maintain feature details of the preliminary features, and retain the track end features in combination with the train running track feature map; the track segmentation network RTEW-Net is obtained by training with a training set, and the training set is a track image and corresponding annotations; Inputting the track line image data to be segmented into the track line segmentation network RTEW-Net for processing includes: The track line image data to be segmented is subjected to feature dimension upgrading through the Patch Embedding method to obtain a feature map, and the size of the feature map is normalized; The normalized feature map is passed through the full attention calculation module FTM for attention calculation, and the feature map is aggregated and recalibrated through the global response normalization module GRN to output the normalized feature map; The feature map processed by the GRN is combined with the normalized feature map through an intelligent weight maintenance module WWM to output a train track segmentation result; The track line image data to be segmented is subjected to feature dimension upgrading by using a Patch Embedding method to obtain a feature map, including: Based on the track line segmentation network RTEW-Net, the track line image data to be segmented is downsampled by deep features through Patch Embedding operation and Full Transformer Block calculation, and downsampled through Patch Merging, and this process is repeated and iterated until the feature map is output; Among them, each downsampling operation combined with Block calculation constitutes an operation stage. During the Block calculation process, LN operation, FTM module calculation, GELU activation and MLP classifier calculation are performed in sequence; Attention calculation is performed through the full attention calculation module FTM, including: Base-patch calculation and pad-patch calculation; The method for calculating Base-patch is: , The pad-patch calculation method is: In the formula, and They represent the operations of expanding and removing the padding according to the patch shape, respectively. and They represent the operations of splitting the feature map into patches and restoring it. represents the execution of attention calculation, Base-patch calculation results, is the pad-patch calculation result, is the input feature map 1; In a block calculation, the base-patch calculation is performed first, and the pad-patch calculation is performed second; The method for the global response normalization module GRN to aggregate global features of the feature map and recalibrate is as follows: In the formula, and is a trainable parameter, is the output feature map 1, is the regularization method, is the input feature map; The intelligent weight maintenance module WWM retains the normalized feature map information and the processed feature map information through an adaptive weight maintenance method, and sets the adaptive parameters participating in the global optimization to organically combine the normalized feature map information and the processed feature map information, so as to realize the model to adaptively retain and adjust the detail features in the feature map; Training the track segmentation network RTEW-Net includes: The parameters of RTEW-Net are optimized through the cross entropy loss function and the adamW optimizer, and the intelligent residual link operation is performed to obtain the trained track segmentation network RTEW-Net.
2. The track line segmentation method for high-precision detection of track end according to claim 1, characterized in that: Constructing the training set includes: The initial track line images in the real train operation scenario are collected, the initial track line images are manually annotated and screened, and the screened images are combined with the initial track line images to construct the RailMixed2024 dataset, i.e., the training set; wherein the real train operation scenario includes image data in which the end features of the track line disappear or are unclear due to entering and exiting tunnels and extremely severe weather conditions.
3. The track line segmentation method for high-precision detection of track end according to claim 1, characterized in that: The structure of the track segmentation network RTEW-Net is an encoding-decoding structure, the encoder structure includes a full attention calculation module FTM, a global response normalization module GRN and an intelligent weight maintenance module WWM connected in sequence, and the decoder is an Uperhead decoder.
4. The track line segmentation method for high-precision detection of track end according to claim 1, characterized in that: The method of performing the intelligent residual link operation is to add trainable weights to the calculated feature map and optimize it during the network training process, specifically: In the formula, To output feature map 2, and are trainable parameters that participate in global optimization independently. and Feature map information after processing by different modules.
Citation Information
Patent Citations
Track obstacle detection method based on combination of deep learning and target detection
CN116524451A
Lane line detection method based on lightweight semantic segmentation network
CN118609091A