A method for identifying the state of track fasteners based on a three-stage parallel network architecture
Through the three-stage parallel network architecture combined with CNN and ViT, the precise identification and fine-grained analysis of track fasteners are achieved, which solves the problem of insufficient identification accuracy and real-time in the existing technology, improves the robustness and generalization capabilities of the model, and meets the real-time monitoring needs of railway operations.
Patent Information
- Application Number
- CN202411593618.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing track fastener state recognition algorithms are difficult to distinguish fasteners from environmental disturbances in complex environments, and cannot be described in fine-grained manner. The model generalization ability is weak, resulting in missed and missed detection, large calculation amount, and difficult to meet real-time requirements.
Using a three-stage parallel network architecture, combined with convolutional neural network (CNN) and Vision Transformer (ViT), multi-branch feature extraction and fusion, including local feature branches, global feature branches and mask generation branches, multi-scale feature extraction and instance segmentation of track fasteners are performed to generate accurate fastener masks, real-time identification and real-time analysis.
It improves the accuracy and robustness of track fastener identification, can effectively distinguish fastener from interference in complex environments, provides detailed damage analysis, meets real-time inspection needs, and improves the generalization ability and processing efficiency of the model.
Smart Images

Figure CN119478524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for identifying the state of track fasteners based on a three-stage parallel network architecture. Background Art
[0002] Fasteners are key connection components in the track system and play a crucial role in ensuring the stability, reliability, and safety of the track. The main function of the fasteners is to fix the rails on the sleepers, prevent the track from shifting laterally or longitudinally, and at the same time maintain the stable spacing between the tracks and reduce the impact force between the rails and the sleepers. With the rapid expansion of China's railway network and the increase in train operation speed, higher requirements are put forward for the safety and integrity of the fastener system. If a large area of fasteners is missing or damaged, it may cause the rails to deform or collapse, increasing the risk of train derailment and seriously threatening the safety of railway operation. Therefore, accurately and efficiently identifying the state of track fasteners is of great significance for ensuring the safe operation of railway lines.
[0003] Existing deep learning-based algorithms for identifying the state of track fasteners usually adopt the method of object detection. Although they can solve the real-time problem of track fastener inspection to a certain extent, they cannot describe the state of fasteners in a fine-grained manner, and it is difficult to distinguish fasteners from interference objects in the environment in a complex environment, resulting in missed detection or false detection. In addition, the generalization ability of existing methods in different track scenarios is weak, and the model is prone to failure when the scenario changes, unable to meet the high-efficiency identification requirements in a large range and complex environment.
[0004] The method for identifying the state of track fasteners disclosed in the patent (CN116721263A, a method for identifying the state of track fasteners based on real-time instance segmentation) uses a track fastener real-time instance segmentation model based on YOLACT. After feature extraction and feature fusion by the Res2Net backbone network, the state of track fasteners is identified. However, this patent solution has the following problems:
[0005] 1. The YOLACT algorithm is not suitable for small object detection, and the detection effect for small objects is not fine. Small objects will be ignored when using the YOLACT algorithm. Moreover, the YOLACT algorithm cannot establish the association between objects according to the scene. Taking the rail as an example, the structure of the rail is symmetric left and right, but the YOLACT algorithm cannot consider the potential connection between objects, resulting in inaccurate detection effects.
[0006] 2. The structure of the Res2Net backbone network is complex and the overall computational amount is large. Summary of the Invention
[0007] The purpose of the present invention aims to solve at least one of the above technical defects.
[0008] To this end, an object of the present invention is to propose a method for identifying the state of track fasteners based on a three-stage parallel network architecture to solve the problems mentioned in the background art and overcome the deficiencies existing in the prior art.
[0009] To achieve the above object, an embodiment of the present invention provides a method for identifying the state of track fasteners based on a three-stage parallel network architecture, including:
[0010] Step S1, obtaining a track inspection image, and preprocessing the track inspection image, including: enhancing the contrast of the track fastener image in the track inspection image;
[0011] Step S2, constructing a convolutional neural network CNN backbone network, and using the CNN backbone network to perform multi-scale feature extraction on the image of the track fastener, and stage-by-stage extracting the low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fastener;
[0012] Step S3, using the local feature branch in the multi-level feature processing network to perform convolutional processing and upsampling processing on the medium-scale fastener features and medium-high-level fastener features of the track fastener, extracting the detailed information of the track fastener, and obtaining the local fastener feature F local-final ;
[0013] Step S4, using the global feature branch in the multi-level feature processing network to perform global average pooling and convolutional processing on the high-level fastener features of the track fastener, extracting the global semantic information of the track fastener, and obtaining the global fastener feature F global-final ;
[0014] Step S5, using the mask generation branch based on the Vision Transformer architecture to process the medium-high-level fastener features and high-level fastener features of the track fastener, generating the prototype mask of the track fastener, and obtaining the masked fastener feature M(x, y, c);
[0015] Step S6, performing three-stage parallel feature fusion on the local fastener feature F local-final , the global fastener feature F global-final and the masked fastener feature M(x, y, c), and performing convolutional processing on the fused fastener feature map;
[0016] Step S7, performing non-maximum suppression and mask generation on the generated fastener feature map, realizing instance segmentation recognition of the track fastener, calculating the mask score of each instance and filtering, eliminating overlapping regions through non-maximum suppression, and retaining the masked instance M NMS (x, y);
[0017] Step S8, analyze the health status of the track fasteners according to the mask instance M NMS (x, y) to generate a health report of the track fasteners, where the health report of the track fasteners includes: the damaged area and the degree of damage of the track fasteners.
[0018] Preferably, in the step S1, enhancing the contrast of the track fastener image in the track inspection image includes: defining a local window W i,j , and this local window W i,j is used to calculate the local contrast enhancement parameter, and calculate the average pixel value μ of this local window i,j as:
[0019]
[0020] where I(x, y) is the track inspection image, (x, y) is the pixel coordinate in the image, and |W i,j | represents the number of track fastener pixels within the window;
[0021] Calculate the standard deviation σ of the scanned local window of the track fastener i,j ;
[0022] For each track fastener pixel I(x, y), define an adaptive contrast enhancement coefficient C i,j ;
[0023] Calculate the enhanced image I′(x, y) of the track fastener as:
[0024] I'(x, y) = C i,j ·(I(x, y) - μ i,j ) + μ global ;
[0025] where μ global is the average pixel value of the global image.
[0026] Preferably, in the step S2, four stages are adopted to extract the low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fasteners in stages, where
[0027] Low-level stage: Adopt 3 3×3 convolutional layers and one pooling layer, with a stride of 1, to extract low-level fastener features:
[0028] Medium-level stage: Adopt 4 3×3 convolutional layers and one pooling layer, with a stride of 2, to extract medium-scale fastener features;
[0029] Medium-high-level stage: Adopt 4 3×3 convolutional layers and one pooling layer, with a stride of 2, to extract medium-high-level fastener features;
[0030] High - level stage: Use 6 convolutional layers of 3×3 and one pooling layer with a stride of 2 to extract high - level fastener features.
[0031] Preferably, according to any of the above - mentioned solutions, the low - level fastener features include the edge information of the track fastener; the high - level fastener features are global fastener features, including the fastener shape and position of the track fastener.
[0032] Preferably, according to any of the above - mentioned solutions, in the step S3,
[0033] Perform convolution processing on the medium - scale fastener features and the medium - high - level fastener features respectively;
[0034] Adjust the resolution of the convolution - processed medium - high - level fastener features to the same as that of the convolution - processed medium - scale fastener features through up - sampling operation;
[0035] Fuse the convolution - processed medium - scale fastener features and the up - sampled medium - high - level fastener features;
[0036] Compress the channels of the fused fastener features through convolution to generate local fastener feature F local-final .
[0037] Preferably, according to any of the above - mentioned solutions, in the step S4,
[0038] First, perform global average pooling on the high - level fastener feature F4(x, y, c) to obtain the globally pooled fastener feature F global which is:
[0039]
[0040] where H and W are the height and width of the fastener feature map respectively;
[0041] Then process the globally pooled fastener feature F global through two convolutional layers to generate the global fastener feature F global-final .
[0042] Preferably, according to any of the above - mentioned solutions, in the step S5,
[0043] Cut the fastener feature maps of the medium - high - level fastener features and the high - level fastener features into image blocks of a fixed size. Suppose the size of each block is p×p, then the fastener feature map F(x, y, c) is divided into blocks, and each block is flattened and mapped to a vector
[0044]
[0045] Input each vector into the Transformer layer, and calculate the attention weights through the self-attention mechanism:
[0046]
[0047] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively, and d k is the dimension of the key;
[0048] After passing through the Transformer decoder layer, output the fastener features for generating the mask:
[0049]
[0050] Then remap the above fastener features back to the spatial dimension through the Transformer decoder layer:
[0051]
[0052] Finally, obtain the track fastener mask prediction fastener feature map M(x, y, c).
[0053] Preferably, in the step S6, according to any of the above solutions
[0054] First, reduce the dimension of the local fastener feature F local-final and the global fastener feature F global-final through 1×1 convolution:
[0055] F local-reduced = ReLU(W9 * F local-final + b9)
[0056] F global-reduced = ReLU(W 10 * F global-final + b 10 )
[0057] Then, perform weighted fusion on the dimension-reduced fastener features and the fastener feature map of the mask generation branch M(x, y, c):
[0058] F fusion = α·F local-reduced + β·F global-reduced + γ·M(x, y, c)
[0059] Among them, α, β, and γ are fusion weight coefficients;
[0060] Finally, process the fused feature map through a 3×3 convolutional layer:
[0061] F final = ReLU(W 11 * F fusion + b11 )。
[0062] Preferably, in step S7 of any of the above solutions, calculate the mask score of each instance and filter it through a threshold τ:
[0063]
[0064] Eliminate overlapping regions through non-maximum suppression and retain the instance with the highest score:
[0065] M NMS (x,y) = NMS(M filtered (x,y)).
[0066] Preferably, in step S8 of any of the above solutions,
[0067] calculate the area A of each track fastener:
[0068]
[0069] Judge the health status of the track fastener according to the shape, area and position information of the mask:
[0070] Health Status = f(A, Shape, Position)
[0071] where Shape is the shape of the track fastener and Position is the area of the track fastener.
[0072] Compared with the prior art, the advantages and beneficial effects of the present invention are:
[0073] 1. Precise instance segmentation and fine-grained recognition: Most of the prior art uses object detection methods, which can only roughly identify the status of track fasteners and cannot provide fine-grained analysis. The present invention combines CNN and ViT. Through instance segmentation technology, it can not only accurately distinguish track fasteners from environmental interferences, identify the positions of track fasteners, but also achieve fine-grained recognition of the fastener status, effectively distinguish fasteners from interferences in the environment, thus significantly improving the recognition accuracy. Especially in complex environments, it can effectively solve the problems of false detection and missed detection in the prior art. This method has stronger robustness and generalization ability in complex environments.
[0074] The present invention uses instance segmentation technology to describe the fastener status in fine-grained manner, and can distinguish the fastener status with different degrees and positions of damage. The innovation of this fine-grained analysis method ensures the accuracy of inspection. The present invention combines convolutional neural network (CNN) and Vision Transformer (ViT), and proposes a new three-stage parallel network architecture to achieve accurate identification and status analysis of track fasteners through instance segmentation, improving the recognition accuracy in complex environments.
[0075] 2. Multi-branch feature extraction and fusion enhance the robustness of the model: Traditional technologies are prone to detection failures or insufficient generalization ability in different scenarios or environments. The present invention adopts a multi-branch architecture design, including a local feature branch, a global feature branch and a mask generation branch, and through a feature fusion mechanism, effectively integrates feature information at different levels. This technology not only improves the robustness and generalization ability of the model, but also ensures efficient recognition in complex track scenarios, solving the defect of unstable performance of existing technologies when the environment changes.
[0076] The present invention uses ViT to generate a mask and fuses it with local and global features, which improves the segmentation accuracy of the model. This multi-branch feature fusion mechanism enhances the generalization ability of the model in different scenarios, enabling the model to adapt to different track environments and work stably in both simple and complex scenarios, significantly improving the accuracy and robustness of track fastener recognition.
[0077] 3. Strong real-time processing ability, meeting the needs of large-scale inspection: The processing efficiency of existing technologies is low and it is difficult to meet the real-time requirements. The present invention combines a lightweight CNN backbone network with ViT to optimize the inference speed of the model, enabling real-time recognition of track fastener status while ensuring recognition accuracy, adapting to large-scale inspection tasks, thus significantly solving the deficiencies of existing technologies in terms of processing speed and real-time performance.
[0078] Compared with traditional deep learning models, the present invention can achieve real-time status recognition of track inspection, and greatly improves the inspection efficiency without sacrificing the recognition accuracy.
[0079] 4. The present invention can provide a more detailed analysis report of the fastener status. Especially in the scenario where the fastener is damaged or missing, it can accurately identify the specific location and degree of the problem, greatly improving the accuracy and maintenance efficiency of track inspection.
[0080] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings
[0081] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:
[0082] Figure 1 It is a flowchart of a method for identifying the status of track fasteners based on a three-stage parallel network architecture implemented according to the present invention;
[0083] Figure 2 It is a schematic diagram of the model structure of a method for identifying the status of track fasteners based on a three-stage parallel network architecture implemented according to the present invention;
[0084] Figure 3 It is a schematic diagram of the generation of a fastener mask implemented according to the present invention;
[0085] Figure 4 It is a schematic diagram of local feature fusion implemented according to the present invention;
[0086] Figure 5 It is a schematic diagram of global feature fusion implemented according to the present invention;
[0087] Figure 6 It is a schematic diagram of multi-scale feature fusion implemented according to the present invention;
[0088] Figure 7a and Figure 7b They are test examples provided according to the present invention. Detailed implementation manners
[0089] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0090] First, the technical terms related to the present invention will be described below:
[0091] Inference speed: Measured by the number of frames processed per second (FPS), that is, the average number of iterations of the model per second, which can show the speed at which the model processes the input. The higher this value, the faster the inference speed and the better the model performance.
[0092] As Figure 1 and Figure 2 shown, the method for identifying the status of track fasteners based on a three-stage parallel network architecture in the embodiments of the present invention includes:
[0093] Step S1, obtaining a track inspection image and preprocessing the track inspection image, including: enhancing the contrast of the track fastener image in the track inspection image.
[0094] Specifically, an adaptive contrast enhancement algorithm for track fasteners (TFACE) is designed. Using this algorithm, the contrast is adaptively adjusted according to the local standard deviation of the image, and the contrast of the acquired track inspection images is enhanced.
[0095] Enhancing the contrast of the track fastener images in the track inspection images includes:
[0096] First, define a local window W i,j , and this local window W i,j is used to calculate the local contrast enhancement parameter, and the window size is m×m. Calculate the average pixel value μ i,j of this local window as:
[0097]
[0098] where I(x, y) is the track inspection image, (x, y) is the pixel coordinate in the image, and |W i,j | represents the number of track fastener pixels within the window.
[0099] Then, calculate the standard deviation σ i,j of the scanned local window of the track fastener as:
[0100]
[0101] Secondly, for each track fastener pixel I(x, y), define an adaptive contrast enhancement coefficient C i,j as follows:
[0102]
[0103] Finally, calculate the enhanced image I′(x, y) of the track fastener as:
[0104] I′(x, y) = C i,j · (I(x, y) - μ i,j ) + μ global (4)
[0105] where μ global is the average pixel value of the global image, which is used to maintain the consistency of the global brightness.
[0106] Step S2, construct a convolutional neural network CNN backbone network, and use the CNN backbone network to perform multi-scale feature extraction on the images of the track fasteners, and stagewise extract the low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fasteners.
[0107] Specifically, a backbone network is constructed using a convolutional neural network (CNN). Four stages are adopted to extract low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fasteners in stages. Each stage consists of several convolutional layers and pooling layers, and the network structure gradually extracts high-level fastener features from low-level fastener features.
[0108] (1) Low-level stage: Three 3×3 convolutional layers and one pooling layer with a stride of 1 are used to extract low-level fastener features. Among them, the low-level fastener features include the edge information of the track fasteners.
[0109] (2) Medium-level stage: Four 3×3 convolutional layers and one pooling layer with a stride of 2 are used to extract medium-scale fastener features.
[0110] (3) Medium-high-level stage: Four 3×3 convolutional layers and one pooling layer with a stride of 2 are used to extract medium-high-level fastener features.
[0111] (4) High-level stage: Six 3×3 convolutional layers and one pooling layer with a stride of 2 are used to extract high-level fastener features. Among them, the high-level fastener features are global fastener features, including the fastener shape and position of the track fasteners.
[0112] Step S3: Use the local feature branch in the multi-level feature processing network to perform convolutional processing and upsampling processing on the medium-scale fastener features and medium-high-level fastener features of the track fasteners, extract the detailed information of the track fasteners, and obtain the local fastener feature F local-final 。
[0113] Specifically, as Figure 4 shown, the local feature branch receives the fastener feature maps of the medium-scale fastener features and medium-high-level fastener features, and extracts the detailed information of the track fasteners through multiple convolutional operations. The specific convolutional operation process is as follows:
[0114] First, perform convolutional processing on the medium-scale fastener features and medium-high-level fastener features respectively.
[0115] The fastener feature map F2(x, y, c) of the medium-scale fastener features passes through three convolutional layers, and the output of each layer is:
[0116]
[0117] where * represents the convolutional operation, and W i and b i are the weights and biases of the i-th layer respectively.
[0118] The fastener feature map F3(x, y, c) of the medium-high-level fastener features passes through two convolutional layers:
[0119]
[0120] Secondly, the medium and high-level fastener features after convolution processing are adjusted to the same resolution as the medium-scale fastener features after convolution through upsampling operation as follows:
[0121]
[0122] Then, the medium-scale fastener features after convolution and the medium and high-level fastener features after upsampling are fused as follows:
[0123]
[0124] Finally, the fused fastener features are compressed in channels through 1×1 convolution to generate local fastener features F local-final , as follows:
[0125] F local-final = ReLU(W6 * F local + b6)(9)
[0126] Step S4, the global fastener feature branch is used to extract the global semantic information of the track fastener. The global semantic information of the track fastener is extracted by performing global average pooling (Global Average Pooling, GAP) and convolution processing on the high-level fastener features in the global feature branch of the multi-level feature processing network, and the global fastener feature F global-final is obtained.
[0127] As Figure 5 shown, first, global average pooling is performed on the fastener feature map F4(x, y, c) of the high-level fastener features to obtain the fastener features F global after global pooling as:
[0128]
[0129] where H and W are the height and width of the fastener feature map respectively.
[0130] Then, the fastener features F global after global pooling are processed through two convolutional layers to generate the global fastener feature F global-final as:
[0131]
[0132] Step S5, the mask generation branch based on the Vision Transformer (ViT) architecture is used to process the medium and high-level fastener features and high-level fastener features of the track fastener to generate a prototype mask of the track fastener, and the masked fastener feature M(x, y, c) is obtained.
[0133] Specifically, as Figure 3 shown, the mask generation branch adopts the Vision Transformer (ViT) architecture and takes the fastener feature maps of the middle and high-level fasteners and the high-level fastener features as inputs.
[0134] First, the input fastener feature map is sliced into image patches of a fixed size. Assuming the size of each patch is p×p, the fastener feature map F(x, y, c) is divided into patches, and each patch is flattened and mapped to a vector
[0135]
[0136] Then, the above vectors are input into the Transformer layer, and the attention weights are calculated through the Transformer self-attention mechanism:
[0137]
[0138] where Q, K, and V are the query matrix, key matrix, and value matrix respectively, and d k is the dimension of the key.
[0139] Subsequently, through the Transformer decoder layer, the fastener features for generating the mask are output:
[0140]
[0141] Secondly, the above fastener features are remapped back to the spatial dimension through the Transformer decoder layer:
[0142]
[0143] Finally, the fastener feature map M(x, y, c) for predicting the track fastener mask is obtained.
[0144] Step S6, perform three-stage parallel feature fusion on the local fastener feature F local-final , the global fastener feature F global-final and the mask fastener feature M(x, y, c) to complete the fastener feature fusion operation that can refine the track fastener recognition, and perform convolution processing on the fused fastener feature map.
[0145] As Figure 6 shown, first, the local fastener feature F local-final and the global fastener feature F global-final are dimension-reduced through 1×1 convolution:
[0146]
[0147] Then, the dimension-reduced fastener features are weighted and fused with the fastener feature map of the mask generation branch M(x, y, c):
[0148] F fusion = α·F local-reduced + β·F global-reduced + γ·M(x, y, c)(17)
[0149] where α, β, and γ are fusion weight coefficients.
[0150] Finally, the fused feature map is processed through a 3×3 convolutional layer as follows:
[0151] F final = ReLU(W 11 * F fusion + b 11 )(18)
[0152] The present invention adopts a multi-branch structure, in which the local feature branch extracts the edge details of the fastener, the global feature branch extracts the overall semantic information of the fastener, the mask generation branch accurately locates the fastener area through ViT, and finally the three are weighted and fused through feature fusion, so as to ensure the accuracy and robustness of recognition.
[0153] Step S7: Perform non-maximum suppression and mask generation on the generated fastener feature map to realize instance segmentation recognition of track fasteners, calculate the mask score of each instance and perform filtering, eliminate overlapping regions through non-maximum suppression, and retain the mask instance M NMS (x, y) with the highest score.
[0154] Specifically, after mask generation, non-maximum suppression (NMS) is used to process overlapping detection results. Calculate the mask score of each instance and filter it through a threshold τ:
[0155]
[0156] Eliminate overlapping regions through non-maximum suppression and retain the instance with the highest score:
[0157] M NMS (x, y) = NMS(M filtered (x, y))(20)
[0158] The present invention adopts a model architecture that combines a custom convolutional neural network (CNN) backbone network with a Vision Transformer (ViT). The CNN is used to extract multi-scale features of track fasteners, and through local feature branches, global feature branches, and mask generation branches in a parallel architecture, the local details, global semantic information, and precise masks of the fasteners are extracted and fused. At the same time, the ViT is used to generate precise masks of the fasteners through the self-attention mechanism, enhancing the segmentation accuracy in complex environments.
[0159] Step S8, according to the mask instance M NMS (x, y) analyze the health status of the track fastener to generate a health report of the track fastener, where the health report of the track fastener includes: the damaged area and the degree of damage of the track fastener.
[0160] In this step, analyze the track fastener, and combine the area, shape, and position information to judge the health status of the fastener.
[0161] Specifically, based on the generated mask M NMS (x, y), analyze the status of the track fastener. Calculate the area A of each fastener:
[0162] A = ∑ x,y M NMS (x, y)(21)
[0163] According to the information such as the shape, area, and position of the mask, combined with prior knowledge, judge the health status of the fastener:
[0164] Health Status = f(A, Shape, Position)(22)
[0165] where Shape is the shape of the track fastener and Position is the area of the track fastener.
[0166] Finally, generate a health report of the track fastener, indicating the possible damaged area and degree.
[0167] The present invention performs fine-grained segmentation and analysis of the fastener status through instance segmentation. The model combines local feature branches and global feature branches, and can accurately identify the status of fastener damage, missing, deformation, etc., and output fine information such as the shape and position of the fastener.
[0168] In addition, the present invention can also adopt the following method to extract features of track fasteners:
[0169] (1) Use the CNN backbone network to separately extract the original state feature map and the changed state feature map from the acquired track fastener image.
[0170] (2) Then, feature extraction is performed on the original state feature map and the changed state feature map through semantic marking, respectively obtaining the original state semantic label and the changed state semantic label.
[0171] (3) The original state semantic label and the changed state semantic label are subjected to local feature transformation encoding to obtain a semantic label with local feature information.
[0172] (4) The semantic label with local feature information is segmented into multiple local labels, and corresponding decoding methods are used for decoding to obtain the decoded local feature labels. The multiple local feature labels are weighted and fused to obtain the local feature information of the track fastener.
[0173] (5) The original state semantic label and the changed state semantic label are subjected to global feature transformation encoding to obtain a semantic label with global feature information.
[0174] (6) The semantic label with global feature information is segmented into multiple global labels, and corresponding decoding methods are used for decoding to obtain the decoded global feature labels. The multiple global feature labels are weighted and fused to obtain the global feature information of the track fastener.
[0175] (7) The mask generation branch of the architecture based on Vision Transformer (ViT) is used to process the fastener features of the track fastener, generate the prototype mask of the track fastener, and obtain the mask feature information.
[0176] (8) By fusing the local feature information, global feature information, and mask feature information, the identification of the health state of the track fastener is realized.
[0177] Figure 7a and Figure 7b The test examples provided according to the embodiments of the present invention. Table 1 shows the accuracy and speed performance of the solutions provided by the present invention and the advanced models on the test set.
[0178] Table 1 Accuracy and speed performance of the present invention and advanced models on the test set
[0179]
[0180]
[0181] The method for identifying the state of track fasteners based on a three - segment parallel network architecture proposed by the present invention combines convolutional neural network (CNN) and Vision Transformer (ViT) technologies to achieve precise identification and state analysis of track fasteners. This method can not only effectively distinguish track fasteners from interference objects in the environment, but also perform fine - grained segmentation and identification of fastener states. By adopting a custom CNN backbone network and multiple feature extraction branches, the model can automatically learn local and global features of track fasteners, enhancing the robustness of the model. At the same time, by leveraging Vision Transformer to generate accurate fastener masks, it ensures that the fastener state can still be accurately identified in complex environments. In addition, the present invention has good real - time performance and can efficiently process a large number of track inspection images without affecting the recognition accuracy, meeting the requirements for real - time monitoring of fastener states in railway operations.
[0182] The technical effects of convolutional neural network (CNN) and Vision Transformer (ViT) technologies in the present invention are described as follows:
[0183] 1. Global context capture ability
[0184] The self - attention mechanism of Transformer can handle global features in fastener detection tasks, especially for long - distance and continuous track scenarios. Traditional CNNs may only detect the state of fasteners in local areas and are difficult to effectively capture the distribution and state of fasteners at a long distance. Transformer can learn global context, thus enhancing the overall understanding of fastener states.
[0185] This ability is particularly important when detecting problems such as wear, looseness, and offset of fasteners, because these problems not only depend on the local state of fasteners, but may also be related to the overall stress distribution and dynamic environment of the track.
[0186] 2. Complex background and occlusion handling
[0187] The railway environment is usually complex, and fasteners are subject to various noise interferences, such as deformation of the rails themselves, debris, and shadows. CNN + Transformer can better distinguish fasteners in complex backgrounds by combining the local detail capture of CNN and the global information processing of Transformer, especially when there are occlusions or perspective changes. Transformer can ensure that the detection results are not affected by noise and background interference by modeling longer context information.
[0188] 3. Multi - scale feature fusion
[0189] The Transformer can capture features at different scales, while the CNN is good at extracting detailed information locally. Railway fasteners usually vary in size and may exhibit different scale features at different shooting angles. The combination of CNN + Transformer can handle these different scale features simultaneously, effectively coping with changes in fastener size, shape, and perspective, thereby improving the robustness and accuracy of detection.
[0190] This multi-scale processing ability is very important for detecting minor damages to fasteners, such as cracks and bolt looseness, and can ensure that small defects are not missed.
[0191] 4. Modeling Long-Term Dependencies
[0192] The long-range dependency modeling ability of the Transformer can capture the relationships between fasteners and the surrounding environment and other components. This means that the system can not only detect the status of individual fasteners but also understand the role and impact of fasteners in the entire track system. For example, when some fasteners are loose, it may cause abnormalities in other fasteners. Modeling these upstream and downstream relationships is crucial for detecting abnormal patterns.
[0193] In fastener detection, this global perspective helps to discover systematic problems rather than just local failures.
[0194] 5. Adapting to Dynamically Changing Scenarios
[0195] The railway environment changes dynamically, such as vibrations when trains pass by and weather changes. These can all affect the results of fastener detection. Through the dynamic feature modeling ability of the Transformer, CNN + Transformer can better adapt to these changes and thus maintain high detection performance in different environments. The CNN part is responsible for extracting local static details to ensure that key information is not lost during minor changes.
[0196] This ability to adapt to dynamic environments improves the robustness of the model in actual scenarios, enabling it to cope with complex external conditions.
[0197] 6. Handling Large Amounts of Data
[0198] Railway fastener detection usually involves large-scale image data, which may include continuous track images. The Transformer has a natural advantage in processing large-scale data and long sequence data. It can use the self-attention mechanism to effectively filter out key information and improve detection efficiency. In addition, the CNN can quickly extract local features. Combining the global optimization ability of the Transformer makes the system more efficient in processing large-scale images.
[0199] 7. High Detection Accuracy
[0200] The combination of CNN and Transformer can accurately locate the edge details of fasteners, which is especially important for detecting subtle problems such as damage and cracks in fasteners. Compared with traditional convolutional networks, the addition of Transformer enables the model to have better context awareness and can more accurately distinguish different states of fasteners.
[0201] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0202] It is not difficult for those skilled in the art to understand that the present invention includes any combination of the above-mentioned invention content and specific implementation parts in the specification and each part shown in the drawings. Due to space limitations and to make the specification concise, the various solutions formed by these combinations are not described one by one. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0203] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, replacements, and variations to the above embodiments within the scope of the present invention without departing from the principle and purpose of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying the state of track fasteners based on a three-stage parallel network architecture, characterized in that, It includes the following steps: Step S1: Obtain track inspection images and preprocess the track inspection images, including enhancing the contrast of the track fastener images in the track inspection images; Step S2: Construct a convolutional neural network CNN backbone network, and use the CNN backbone network to perform multi-scale feature extraction on the images of track fasteners, and stagewise extract low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fasteners; Four stages are adopted to stagewise extract low-level fastener features, medium-scale fastener features, medium-high-level fastener features, and high-level fastener features of the track fasteners, where Low-level stage: Use 3 3×3 convolutional layers and one pooling layer with a stride of 1 to extract low-level fastener features: Medium-level stage: Use 4 3×3 convolutional layers and one pooling layer with a stride of 2 to extract medium-scale fastener features; Medium-high-level stage: Use 4 3×3 convolutional layers and one pooling layer with a stride of 2 to extract medium-high-level fastener features; High-level stage: Use 6 3×3 convolutional layers and one pooling layer with a stride of 2 to extract high-level fastener features; Step S3, using the local feature branch in the multi-level feature processing network to perform convolution processing and upsampling processing on the mesoscale fastener features and the middle and high-level fastener features of the track fastener, extracting the detailed information of the track fastener, and obtaining the local fastener feature F local-final ; Perform convolutional processing on the medium-high-level fastener features; Adjust the resolution of the convolution-processed medium-high-level fastener features to the same as that of the convolution-processed medium-scale fastener features through upsampling operation; Fuse the convolution-processed medium-scale fastener features and the upsampled medium-high-level fastener features; Channel compression is performed on the fused fastener features through convolution to generate local fastener features F local-final ; F local-final = ReLU(W6 * F local + b6) Step S4, perform global average pooling and convolution processing on the high-level fastener features of the track fastener using the global feature branch in the multi-level feature processing network to extract the global semantic information of the track fastener, and obtain the global fastener feature F global-final ; Step S5: Use a mask generation branch based on the Vision Transformer architecture to process the medium-high-level fastener features and high-level fastener features of the track fasteners, generate the prototype mask of the track fasteners, and obtain the masked fastener feature M(x, y, c); The fastener feature maps of the medium-high level fastener features and the high-level fastener features are sliced into image blocks of a fixed size. Let the size of each block be p×p, then the fastener feature map F(x, y, c) is divided into blocks, and each block is flattened and mapped to a vector Input each vector into the Transformer layer and calculate the attention weights through the self-attention mechanism: Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively, and d k is the dimension of the key; Pass through the Transformer decoder layer and output the fastener features for mask generation: Then remap the above fastener features back to the spatial dimension through the Transformer decoder layer: Finally, obtain the track fastener mask prediction fastener feature map M(x, y, c); Step S6, perform three-stage parallel feature fusion on the local fastener feature F local-final , the global fastener feature F global-final and the mask fastener feature M(x, y, c), and perform convolution processing on the fused fastener feature map; Step S7, perform non-maximum suppression and mask generation on the generated fastener feature map to achieve instance segmentation recognition of track fasteners, calculate the mask score of each instance and perform filtering, eliminate overlapping regions through non-maximum suppression, and retain the mask instance M with the highest score NMS (x,y); Step S8, analyze the health status of the track fastener based on the mask instance M NMS (x, y), generate a health report of the track fastener, wherein the health report of the track fastener includes: the damaged area and the degree of damage of the track fastener.
2. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 1, characterized in that, In the step S1, enhancing the contrast of the track fastener image in the track inspection image includes: defining a local window W i,j , and this local window W i,j is used to calculate the local contrast enhancement parameter, and calculate the average pixel value μ of this local window i,j as: Among them, I(x, y) is the track inspection image, (x, y) is the pixel coordinates in the image, and |W i,j | represents the number of track fastener pixels within the window; Calculate the standard deviation σ of the local window of the scanning track fastener i,j ; For each track fastener pixel I(x, y), an adaptive contrast enhancement coefficient C is defined i,j ; Calculate the track fastener enhanced image I'(x, y) as: I'(x,y) = C i,j ·(I(x,y) - μ i,j ) + μ global ; where μ global is the average pixel value of the global image.
3. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 2, wherein, The low-level fastener features include the edge information of the track fasteners; the high-level fastener features are global fastener features, including the fastener shape and position of the track fasteners.
4. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 1, wherein In step S4, First, perform global average pooling on the high-level fastener feature F4(x, y, c) to obtain the fastener feature F after global pooling global It is as follows: where H and W are the height and width of the fastener feature map respectively; Then, the fastener feature F after global pooling global is processed through two convolutional layers to generate the global fastener feature F global-final .
5. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 1, characterized in that, In step S6, First, reduce the dimensionality of the local fastener feature F local-final and the global fastener feature F global-final through 1×1 convolution: F local-reduced = ReLU(W9 * F local-final + b9) F global-reduced = ReLU(W 10 * F global-final + b 10 ) Then perform weighted fusion on the dimension-reduced fastener features and the fastener feature map of the mask generation branch M(x, y, c): F fusion = α·F local-reduced + β·F global-reduced + γ·M(x,y,c) where α, β, and γ are fusion weight coefficients; Finally, process the fused feature map through a 3×3 convolutional layer: F final = ReLU(W 11 * F fusion + b 11 )。 6. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 1, wherein In step S7, calculate the mask score of each instance and filter it through the threshold τ: Eliminate the overlapping regions through non-maximum suppression and retain the instance with the highest score: M NMS (x,y) = NMS(M filtered (x,y)).
7. The method for identifying the state of track fasteners based on a three-stage parallel network architecture according to claim 1, wherein In step S8, Calculate the area A of each track fastener: Judge the health status of the track fasteners according to the shape, area, and position information of the mask: Health Status Health Status = f(A, Shape, Position) Wherein, Shape is the shape of the track fastener, and Position is the area of the track fastener.
Citation Information
Patent Citations
Cross-view gait recognition method based on staged multistage pyramid
CN115050093A
Track fastener state identification method based on real-time instance segmentation
CN116721263A