Railway steel rail abrasion intelligent detection method
By combining multimodal image fusion and progressive recognition technology with game theory optimization and adaptive parameter adjustment, the problems of low accuracy and excessive computational complexity in railway rail wear detection have been solved, achieving efficient and real-time wear detection and assessment.
Patent Information
- Application Number
- CN202511492328.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-03-13
AI Technical Summary
Existing methods for detecting rail wear have limitations such as low accuracy and excessive computational complexity, making real-time detection difficult, especially in identifying complex wear patterns and detecting minor wear in its early stages.
Visible light images and infrared thermal images are acquired using multimodal image acquisition equipment. Feature extraction and fusion are performed through multi-scale feature extraction and deep convolutional neural networks. Combined with a progressive wear recognition model and a game theory optimization model, the wear region is accurately identified and its trend is predicted. Wear status is assessed through adaptive parameter adjustment and an improved K-means clustering algorithm.
It improves the accuracy and efficiency of wear detection, can accurately identify wear areas in complex environments, achieves real-time detection and efficient wear condition assessment, and provides reliable maintenance recommendations.
Smart Images

Figure CN121658878A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of railway rail wear detection technology, and specifically relates to an intelligent detection method for railway rail wear. Background Technology
[0002] Rail wear detection is a key technology for ensuring railway transportation safety. Traditional detection methods mainly employ contact measuring equipment to assess rail surface wear, including profilometry, ultrasonic testing, and manual visual inspection. These methods are widely used in railway maintenance operations and can obtain basic information on rail wear for maintenance planning. However, traditional detection methods suffer from drawbacks such as low efficiency, limited accuracy, and inability to achieve continuous monitoring. They are particularly inadequate in identifying complex wear patterns and detecting minor wear in its early stages, failing to meet the demands of modern high-speed railways for precise wear assessment. In the current trend of intelligent railway development, traditional detection technologies lack multi-dimensional information fusion capabilities and intelligent analysis methods, making it impossible to achieve efficient real-time processing while maintaining detection accuracy. Existing technologies suffer from low rail wear detection accuracy and excessive computational complexity, hindering real-time detection. Summary of the Invention
[0003] In view of this, the present invention provides an intelligent detection method for railway rail wear, which can solve the technical problems of low detection accuracy and excessive computational complexity in the prior art, making it difficult to achieve real-time detection.
[0004] This invention is implemented as follows: It provides an intelligent detection method for railway rail wear. A multimodal image acquisition device scans the rail surface to obtain visible light and infrared thermal images, establishing a multimodal image dataset and constructing an initial wear matrix. The multimodal image dataset is preprocessed, and a multi-scale feature extraction algorithm is used to extract features from the preprocessed images, establishing a wear probability matrix. A deep convolutional neural network is used to perform first-level probability recognition on the wear probability matrix, identifying potential wear areas. A multilayer perceptron model is used for feature fusion to establish a wear probability gain matrix. A progressive wear recognition model is used to perform second-level fine recognition on the wear probability gain matrix, identifying wear boundaries through adaptive threshold segmentation and edge detection algorithms, and combining time series analysis to establish a wear development trend prediction model. A game theory optimization model is used to verify and optimize the wear recognition results. Cluster analysis is performed on the optimized wear recognition results, and an improved K-means clustering algorithm is used to establish a wear probability clustering matrix. A rail wear status assessment report is established based on the wear probability clustering matrix.
[0005] Specifically, the construction step of the initial wear matrix is to form a two-dimensional data structure by numerically encoding the spatial coordinates and attribute information of each pixel in the multimodal image dataset. The rows of the matrix represent the vertical pixel positions of the image, the columns represent the horizontal pixel positions of the image, and each matrix element contains the three-dimensional spatial coordinates, gray value and infrared temperature value of the corresponding pixel.
[0006] The preprocessing specifically includes image registration, noise filtering, and brightness normalization. The wear probability matrix is a probability distribution matrix calculated based on statistical methods, used to quantify the probability that each pixel belongs to the wear region. The matrix construction process includes feature vector extraction, probability density function fitting, and Bayesian inference calculation.
[0007] Specifically, the wear probability gain matrix is an enhancement matrix obtained by performing nonlinear transformation and signal amplification on the wear probability matrix. Its main function is to highlight the wear characteristic signal and suppress background noise interference. The gain function adopts the sigmoid function form, and the gain coefficient is adaptively adjusted according to the image quality parameter and the wear severity parameter.
[0008] The progressive wear recognition model is specifically based on a graph convolutional network architecture, which includes an input layer, three graph convolutional layers, two pooling layers, and an output layer. Each graph convolutional layer contains 128 neurons to process the spatial topological relationships of the wear region. The message passing mechanism in the model adopts a neighbor aggregation method.
[0009] The message passing steps of the progressive wear recognition model are dynamically adjusted by an adaptive parameter adjustment function based on the wear region complexity coefficient and the image resolution parameter. The wear region complexity coefficient is obtained by calculating the curvature change rate of the wear boundary and the irregularity of the wear shape. The image resolution parameter represents the pixel density of the input image, in pixels per inch (dpi).
[0010] Specifically, the step of establishing the training dataset for the progressive wear recognition model involves collecting multimodal image samples of rails under different service years and operating conditions. Each sample contains the original multimodal image dataset and manually annotated wear area masks. At the same time, laser point cloud data is used as manually annotated auxiliary data to verify the accuracy of wear area boundaries and supplement three-dimensional geometric information.
[0011] Specifically, the training steps of the progressive wear recognition model involve using the Adam optimization algorithm, with the initial learning rate set to 0.001, the learning rate decreasing to 0.8 times the original value every 30 training cycles, the batch size set to 32, the total number of training cycles being 300, and the loss function being a weighted combination of cross-entropy loss and graph structure loss with a weight ratio of 2 to 1.
[0012] Specifically, the game theory optimization model includes an upper-level model that aims to maximize wear detection accuracy and a lower-level model that aims to minimize computational complexity. The two models achieve collaborative optimization through coupling terms. The objective function of the upper-level model is used to maximize wear detection accuracy, and the objective function of the lower-level model is used to minimize computational complexity.
[0013] Specifically, the adaptive parameter adjustment function is used to adjust the message passing steps of the progressive wear recognition model. It is calculated based on the wear region complexity coefficient, image resolution parameter, image quality parameter, and wear severity parameter to obtain a passing adjustment value. The corresponding message passing steps and neighbor aggregation strategy are set according to different ranges of the passing adjustment value.
[0014] The objective function of the upper-layer model takes as input the number of true positives, false positives, false negatives, wear-prone region coverage, and computational time complexity, and outputs the optimized detection accuracy score. The objective function is expressed as the detection accuracy equal to the ratio of the number of true positives to the sum of the number of true positives, false positives, and false negatives, multiplied by the square root of the wear-prone region coverage, and then subtracted from the logarithm of the computational time complexity. The objective function of the lower-layer model takes as input image processing time, memory usage, CPU utilization, GPU utilization, and data transmission bandwidth, and outputs a system resource consumption assessment. The objective function is expressed as the product of computation time and memory usage, plus an exponential function of the number of algorithm iterations, multiplied by the reciprocal of hardware resource utilization. The coupling term specifically represents the mutual influence between the upper-layer and lower-layer models. Increasing detection accuracy increases computational complexity, while decreasing computational complexity affects detection accuracy. The mathematical expression of the coupling term is the correlation coefficient between detection accuracy and computational complexity multiplied by a weight adjustment factor.
[0015] Specifically, the wear probability clustering matrix is a classification matrix formed by grouping the identified wear regions according to similarity features. Each row of the matrix represents a wear cluster, each column represents a wear feature dimension, and the matrix elements represent the statistical characteristics of the clusters on the corresponding feature dimensions, which are used to support the classification of wear severity and the formulation of maintenance decisions.
[0016] Specifically, the improved K-means clustering algorithm groups similar wear features and calculates the severity level of wear in each group. Image quality parameters are obtained by calculating the signal-to-noise ratio and contrast of the image, and wear severity parameters are obtained by analyzing the ratio of wear depth to wear area.
[0017] The rail wear status assessment report specifically includes wear location coordinates, wear depth, wear area, and estimated remaining service life. Optionally, it also includes maintenance suggestions and early warning information. The detection accuracy score is used for cluster analysis weight adjustment, and the system resource consumption assessment value is used for maintenance suggestion generation.
[0018] This invention employs an intelligent detection method combining multimodal image fusion and progressive recognition. It quantifies the wear probability of each pixel by constructing a wear probability matrix and improves the accuracy of wear feature extraction using a hierarchical recognition mechanism of deep convolutional neural networks and graph convolutional networks, thus overcoming the technical shortcomings of traditional methods in terms of low detection accuracy. Furthermore, this invention achieves synergistic optimization of detection accuracy and computational complexity through a game theory optimization model. An adaptive parameter adjustment function dynamically adjusts model parameters based on the complexity of the wear region and image quality, effectively reducing the system's computational burden while maintaining detection performance, overcoming the limitation of excessive computational complexity in traditional methods. In summary, this invention solves the technical problems mentioned in the background art, namely, low accuracy and excessive computational complexity in rail wear detection, which hinder real-time detection. Attached Figure Description
[0019] Figure 1 This is a flowchart of the method of the present invention.
[0020] Figure 2 This is a scatter plot showing the distribution of rail wear detection areas in the embodiment.
[0021] Figure 3 This is a bar chart showing the feature distribution of the wear probability clustering matrix in the embodiment.
[0022] Figure 4 This is a curve showing the predicted wear development trend in the embodiments. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0024] like Figure 1 The diagram shown is a flowchart of an intelligent detection method for railway rail wear provided by the present invention. This method includes the following steps:
[0025] S01. A multimodal image acquisition device is used to scan the surface of the rail to obtain high-resolution visible light images and infrared thermal imaging images. A multimodal image dataset is established and an initial wear matrix is constructed. The initial wear matrix contains the three-dimensional coordinate information and gray value information corresponding to each pixel.
[0026] S02. Preprocess the multimodal image dataset, including image registration, noise filtering and brightness normalization. Then, use a multi-scale feature extraction algorithm to extract features from the preprocessed image and establish a wear probability matrix. The wear probability matrix is used to represent the initial probability distribution of wear at each pixel.
[0027] S03. Based on a deep convolutional neural network, the wear probability matrix is subjected to first-level probability recognition to identify potential wear regions. A multilayer perceptron model is used for feature fusion to establish a wear probability gain matrix. The wear probability gain matrix is used to amplify the wear feature signal and suppress background noise.
[0028] S04. A progressive wear recognition model is used to perform a second-level fine recognition of the wear probability gain matrix. Wear boundaries are identified through adaptive threshold segmentation and edge detection algorithms. A wear development trend prediction model is established by combining time series analysis. The message passing steps of the progressive wear recognition model are dynamically adjusted according to the wear area complexity coefficient and image resolution parameters through an adaptive parameter adjustment function.
[0029] S05. The wear identification results are verified and optimized using a game theory optimization model, including an upper-level model that aims to maximize the wear detection accuracy and a lower-level model that aims to minimize the computational complexity. The two models achieve synergistic optimization through coupling terms.
[0030] S06. Perform cluster analysis on the optimized wear identification results, use the improved K-means clustering algorithm to establish a wear probability clustering matrix, group similar wear features and calculate the severity level of wear in each group.
[0031] S07. Establish a rail wear status assessment report based on the wear probability clustering matrix, including wear location coordinates, wear depth, wear area and expected remaining service life.
[0032] The initial wear matrix is a two-dimensional data structure formed by numerically encoding the spatial coordinates and attribute information of each pixel in the multimodal image dataset. Rows represent the vertical pixel positions of the image, and columns represent the horizontal pixel positions. Each matrix element contains the three-dimensional spatial coordinates, grayscale value, and infrared temperature value of the corresponding pixel. The wear probability matrix is a probability distribution matrix calculated using statistical methods, used to quantify the probability that each pixel belongs to a wear region. The matrix elements range from 0 to 1; the closer the value is to 1, the higher the probability of wear at that pixel. The matrix construction process includes feature vector extraction, probability density function fitting, and Bayesian inference calculation. The wear probability gain matrix is an enhancement matrix obtained by performing nonlinear transformation and signal amplification on the wear probability matrix. Its main function is to highlight wear feature signals and suppress background noise interference. The gain function adopts the sigmoid function form, and the gain coefficient is adaptively adjusted according to image quality parameters and wear severity parameters. The wear probability clustering matrix is a classification matrix formed by grouping the identified wear regions according to similarity features. Each row of the matrix represents a wear cluster, each column represents a wear feature dimension, and the matrix elements represent the statistical characteristics of the cluster in the corresponding feature dimension, used to support the classification of wear severity and maintenance decision-making. The wear region complexity coefficient is obtained by calculating the rate of curvature change of the wear boundary and the irregularity of the wear shape, with a value range of 0.1 to 1.0. When the wear region shape is regular and the boundary is smooth, the wear region complexity coefficient is close to 0.1; when the wear region shape is irregular and the boundary is complex, the wear region complexity coefficient is close to 1.0. The image resolution parameter represents the pixel density of the input image, in pixels per inch (dpi), obtained from the multimodal image acquisition device in step S01. The image quality parameter is obtained by calculating the signal-to-noise ratio and contrast of the image, with a value range of 0 to 100, obtained from the preprocessing results in step S02. The wear severity parameter is calculated by analyzing the ratio of wear depth to wear area, with a value range of 1 to 10, and is obtained from the wear identification results in step S04.
[0033] The progressive wear detection model is based on a graph convolutional network architecture, comprising an input layer, three graph convolutional layers, two pooling layers, and an output layer. Each graph convolutional layer contains 128 neurons to process the spatial topological relationships of the wear region. The message passing mechanism in the model adopts a neighbor aggregation approach. The number of message passing steps is dynamically determined based on the wear region complexity coefficient and image resolution parameters. When the wear region complexity coefficient is less than 0.4 and the image resolution parameter is less than 300 dpi, the number of message passing steps is set to 2. When the wear region complexity coefficient is in the range of 0.4 to 0.7 and the image resolution parameter is in the range of 300 dpi to 600 dpi, the number of message passing steps is set to 3. When the wear region complexity coefficient is greater than 0.7 and the image resolution parameter is greater than 600 dpi, the number of message passing steps is set to 4. The steps for establishing the training dataset of the progressive wear recognition model specifically include collecting multimodal image samples of rails under different service years and operating conditions. Each sample contains the original multimodal image dataset and manually annotated wear area masks. At the same time, laser point cloud data is used as auxiliary data for manually annotated data to verify the accuracy of wear area boundaries and supplement three-dimensional geometric information. The training dataset contains 12,000 normal rail samples, 8,000 light wear samples, 6,000 moderate wear samples, and 4,000 heavy wear samples. All samples have undergone data augmentation processing, including rotation, scaling, brightness adjustment, and noise addition. The final training dataset contains 150,000 annotated images. The training steps of the progressive wear recognition model specifically include using the Adam optimization algorithm, setting the initial learning rate to 0.001, decreasing the learning rate to 0.8 times every 30 training cycles, setting the batch size to 32, and setting the total number of training cycles to 300. The loss function is a weighted combination of cross-entropy loss and graph structure loss with a weight ratio of 2:1. Training stops when the model achieves an accuracy of over 96% on the validation set. An early stopping mechanism is used during training to prevent overfitting.
[0034] The objective function of the upper-level model in the game theory optimization model is used to maximize the wear detection accuracy. Inputs include the number of true positives, false positives, false negatives, wear region coverage, and computational time complexity. The output is the optimized detection accuracy score. The objective function is expressed as the detection accuracy equal to the ratio of the number of true positives to the sum of the number of true positives and false negatives, multiplied by the square root of the wear region coverage, and then subtracted from the logarithm of the computational time complexity. The objective function of the lower-level model is used to minimize computational complexity. Inputs include image processing time, memory usage, CPU utilization, GPU utilization, and data transfer bandwidth. The output is a system resource consumption assessment. The objective function is expressed as the product of computation time and memory usage, plus an exponential function of the number of algorithm iterations, multiplied by the reciprocal of hardware resource utilization. The coupling term represents the mutual influence between the upper-level and lower-level models. Increasing detection accuracy increases computational complexity, while decreasing computational complexity affects detection accuracy. The mathematical expression of the coupling term is the correlation coefficient between detection accuracy and computational complexity multiplied by a weight adjustment factor. The number of true positives is obtained from the wear recognition result in step S04, representing the number of correctly identified wear pixels. The number of false positives is obtained from the wear recognition result in step S04, representing the number of normal pixels incorrectly identified as wear. The number of false negatives is obtained from the wear recognition result in step S04, representing the number of wear pixels that were not identified. The wear region coverage is obtained from the wear recognition result in step S04, representing the proportion of identified wear regions to the total wear regions. The computational time complexity is obtained from the processing time statistics in step S04, representing the time cost required for algorithm execution. The image processing time is obtained from the preprocessing operation in step S02, representing the time consumed by image preprocessing. The memory usage is obtained from the system monitoring module, representing the amount of memory occupied during algorithm operation. The CPU utilization is obtained from the system monitoring module, representing the percentage of CPU utilization. The GPU utilization is obtained from the system monitoring module, representing the percentage of GPU utilization. The data transmission bandwidth is obtained from the network monitoring module, representing the data transmission rate. The detection accuracy score is used for cluster analysis weight adjustment in step S06. The system resource consumption assessment value is used for the maintenance recommendation generation in step S07.
[0035] The adaptive parameter adjustment function is used to adjust the message passing steps of the progressive wear detection model. This function calculates a passing adjustment value based on the wear region complexity coefficient, image resolution parameter, image quality parameter, and wear severity parameter. When the passing adjustment value is less than 0.25, it sets the message passing steps to 2 and uses a simplified neighbor aggregation strategy to improve processing speed. When the passing adjustment value is between 0.25 and 0.5, it sets the message passing steps to 3 and uses a standard neighbor aggregation strategy to balance accuracy and efficiency. When the passing adjustment value is between 0.5 and 0.75, it sets the message passing steps to 4 and uses an enhanced neighbor aggregation strategy to improve detection accuracy. When the passing adjustment value is greater than 0.75, it sets the message passing steps to 5 and uses a multi-layer neighbor aggregation strategy to handle complex wear patterns. This passing adjustment value is used for parameter configuration of the progressive wear detection model in step S04.
[0036] The technical effect of dynamically adjusting the message passing steps using an adaptive parameter adjustment function is that it significantly improves the adaptability and accuracy of wear detection. When dealing with minor wear with regular shapes and clear boundaries, the adaptive parameter adjustment function avoids overcomputation and feature overfitting by setting fewer message passing steps, effectively shortening processing time and reducing computational resource consumption. When encountering severe wear with complex shapes and blurred boundaries, the adaptive parameter adjustment function increases the message passing steps, enabling the graph convolutional network to fully utilize neighborhood information for multi-level feature aggregation, thereby accurately capturing subtle feature changes and complex spatial topological relationships in the wear area. At the same time, by comprehensively considering four key factors—wear area complexity coefficient, image resolution parameter, image quality parameter, and wear severity parameter—the detection algorithm achieves intelligent adaptation to different wear types and image conditions, avoiding the problems of decreased detection accuracy or low computational efficiency that occur in traditional fixed parameter methods when dealing with diverse wear patterns. The four-stage adjustment strategy ensures that the algorithm maintains optimal performance under various complex working conditions, providing a more reliable and efficient technical solution for railway rail wear detection.
[0037] The specific implementation methods of the above steps are described in detail below.
[0038] The specific implementation of step S01 involves synchronously scanning and acquiring images of the rail surface using a multimodal image acquisition device. First, a visible light image is acquired using a high-resolution CCD camera, with an imaging resolution set to 1920×1080 pixels and an exposure time controlled within the range of 5–15 milliseconds to ensure image clarity. Simultaneously, an infrared thermal imager is used to acquire images of the rail surface temperature distribution. The thermal imager's temperature detection range is set to -10℃ to 80℃, with a thermal sensitivity of no less than 0.05℃. During image acquisition, the distance between the device and the rail surface is maintained at 0.5–1.0 meters, and the acquisition angle is perpendicular to the rail surface to ensure the integrity and accuracy of the image information. After time-stamping synchronization processing, the acquired multimodal image data is used to establish a multimodal image dataset. Each dataset contains the visible light image and infrared thermal image at the corresponding time. Based on the multimodal image dataset, an initial wear matrix is constructed. This matrix uses a three-dimensional coordinate system to spatially locate each pixel, where the X-axis and Y-axis correspond to the horizontal and vertical pixel coordinates of the image, respectively, and the Z-axis corresponds to the depth information of the pixel in actual space. Each matrix element contains the three-dimensional spatial coordinates of the corresponding pixel, the grayscale value of the visible light image, and the temperature value of the infrared image, forming a comprehensive numerical data structure. The purpose of this step is to establish the original data foundation for the surface condition of the rail, providing a multi-dimensional information source for subsequent wear analysis.
[0039] The specific implementation of step S02 involves systematic preprocessing and feature extraction of the multimodal image dataset. The image registration process employs a feature-point-based registration algorithm. By detecting corresponding feature points in the visible light and infrared images, the geometric transformation relationship between the two images is calculated, aiming for sub-pixel accuracy with a registration error controlled within 0.5 pixels. Noise filtering utilizes an adaptive median filtering algorithm, with the filter window size dynamically adjusted according to the image noise level, ranging from 3×3 to 9×9 pixels. Brightness normalization employs a histogram equalization algorithm to adjust the image's grayscale value distribution to the standard range of 0–255, improving image contrast and visual quality. After preprocessing, a multi-scale feature extraction algorithm is used for image feature analysis. This algorithm constructs multiple scale levels based on the Gaussian pyramid principle, with each scale level corresponding to a different resolution level. The scale factor is set to 0.5, and the number of pyramid layers is set to 5. At each scale level, the Sobel edge detection operator and Gabor filter are applied to extract texture features. The Sobel operator is used to detect edge information in the image, and the Gabor filter is used to extract directional texture features. A wear probability matrix is established based on the extracted multi-scale features. This matrix calculates the probability value of each pixel belonging to the wear region using statistical methods. The probability calculation is based on Bayesian inference principles, and the probability distribution is updated by combining prior knowledge and observation data. The purpose of this step is to eliminate noise interference in the image, extract effective wear feature information, and provide high-quality input data for subsequent intelligent recognition.
[0040] The specific implementation of step S03 involves using a deep convolutional neural network (DCNN) to perform the first-level probability identification of the wear probability matrix. The DCNN uses an improved version of the LeNet architecture, comprising four convolutional layers and two fully connected layers. Each convolutional layer uses a 3×3 kernel with a stride of 1 and zero padding. The first convolutional layer contains 32 kernels, the second contains 64 kernels, the third contains 128 kernels, and the fourth contains 256 kernels. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function to accelerate the training process and enhance the model's non-linear expressive power. The network extracts local feature patterns from the wear probability matrix through convolutional operations to identify potential wear regions. A confidence threshold of 0.7 is set during the identification process; regions with a wear probability exceeding this threshold are marked as potential wear regions. The recognition results are input into a multilayer perceptron model for feature fusion. The multilayer perceptron contains three hidden layers, each with 512 neurons, and the tanh activation function is used. The feature fusion process weights the feature maps from different convolutional layers, with the weights dynamically adjusted through an attention mechanism. A wear probability gain matrix is then established based on the fused features. This matrix uses a sigmoid gain function to non-linearly transform the original probability values. The gain coefficient is adaptively adjusted according to image quality parameters and wear severity parameters, with a value ranging from 1.2 to 3.5. The purpose of this step is to automatically identify wear features using deep learning methods, improving the accuracy and robustness of detection.
[0041] The specific implementation of step S04 involves using a progressive wear recognition model to perform a second-level fine-tuning of the wear probability gain matrix. The progressive wear recognition model is based on a graph convolutional network architecture, modeling the wear region as a graph structure, where each pixel is a node in the graph, and the relationships between adjacent pixels are edges. The model includes an input layer, three graph convolutional layers, two pooling layers, and an output layer. Each graph convolutional layer contains 128 neurons. The graph convolution operation updates the feature representation of the current node by aggregating information from neighboring nodes, using average pooling as the aggregation function. The number of message passing steps is dynamically determined using an adaptive parameter adjustment function based on the wear region complexity coefficient and image resolution parameters. When the complexity coefficient is less than 0.4 and the resolution is below 300 dpi, the number of message passing steps is set to 2; when the complexity coefficient is in the range of 0.4 to 0.7 and the resolution is in the range of 300 to 600 dpi, the number of message passing steps is set to 3; and when the complexity coefficient is greater than 0.7 and the resolution is greater than 600 dpi, the number of message passing steps is set to 4. The fine-grained identification process employs an adaptive threshold segmentation algorithm. The threshold is dynamically calculated based on the local statistical characteristics of the image, ranging from 0.3 to 0.8. Edge detection uses the Canny operator, with a high threshold set to 0.8 and a low threshold set to 0.4, to accurately locate the boundaries of the wear-affected areas. A wear-affected development trend prediction model is established by combining time series analysis. This model uses an autoregressive moving average model to predict future wear-affected development trends by analyzing historical wear-affected data. The purpose of this step is to achieve precise location and boundary identification of the wear-affected areas, providing accurate wear-affected geometric information.
[0042] The specific implementation of step S05 involves using a game theory optimization model to verify and optimize the wear detection results. The game theory model comprises two layers: an upper-level model and a lower-level model, constituting a two-layer optimization problem. The upper-level model aims to maximize the wear detection accuracy. Input parameters include the number of true positives, false positives, false negatives, wear region coverage, and computational time complexity. The objective function is defined as the product of detection accuracy and the square root of wear region coverage, minus the logarithm of computational time complexity. The lower-level model aims to minimize computational complexity. Input parameters include image processing time, memory usage, CPU utilization, GPU utilization, and data transfer bandwidth. The objective function is defined as the product of computation time and memory usage, plus an exponential function of the number of algorithm iterations, multiplied by the reciprocal of hardware resource utilization. The two models achieve collaborative optimization through a coupling term, which represents the trade-off between improved detection accuracy and increased computational complexity. The weight adjustment factor is set to 0.6. The optimization process employs a genetic algorithm with a population size of 50, a crossover probability of 0.8, a mutation probability of 0.1, and 100 iterations. The aim of this step is to optimize the algorithm's computational efficiency while maintaining detection accuracy, achieving a balance between performance and efficiency.
[0043] The specific implementation of step S06 involves clustering analysis of the optimized wear identification results. An improved K-means clustering algorithm is used to group the wear regions. Clustering features include wear depth, wear area, wear shape complexity, and wear location distribution. The number of clusters K is determined according to the elbow rule, ranging from 3 to 8. The improved K-means algorithm introduces a density weighting mechanism, adjusting the update weight of cluster centers based on the local density of the wear region. The density calculation uses a Gaussian kernel function with a kernel parameter set to 0.5. The convergence condition for the clustering process is set to a change in cluster center position of less than 0.01 or an iteration count of 200. A wear probability clustering matrix is established based on the clustering results. Each row of the matrix represents a wear cluster, each column represents a wear feature dimension, and the matrix elements represent the statistical characteristics of the cluster in the corresponding feature dimension, including mean, standard deviation, and coefficient of variation. The severity level of wear in each group is calculated based on the clustering matrix. A fuzzy evaluation method is used to classify the severity of wear into four levels: mild, moderate, severe, and critical. The threshold values for each level are set to 2.0, 4.5, 7.0, and 9.0, respectively. The purpose of this step is to classify the identified wear areas according to their similarity, providing a basis for developing targeted maintenance strategies.
[0044] The specific implementation of step S07 is to establish a rail wear status assessment report based on a wear probability clustering matrix. The assessment report includes key information such as wear location coordinates, wear depth, wear area, and estimated remaining service life. Wear location coordinates are obtained through image coordinate to actual coordinate conversion, with the conversion coefficient determined based on the calibration parameters of the image acquisition equipment. Wear depth is calculated by comparing the height difference between the wear area and the normal area, with a depth measurement accuracy of 0.1 mm. Wear area is calculated through pixel counting and scale conversion, with an area measurement accuracy of 1 square millimeter. The estimated remaining service life is calculated based on a wear development trend prediction model, considering the impact of wear rate, usage intensity, and environmental factors. Optionally, maintenance recommendations are also included, generated based on the wear severity level and location distribution, including maintenance priority, maintenance method, and maintenance cycle. A three-level warning mechanism is used for early warning information: a green warning indicates normal wear, a yellow warning indicates wear requires attention, and a red warning indicates wear requires immediate action. The warning thresholds are determined based on a comprehensive score of wear depth and area: 3.0 for green, 6.0 for yellow, and 8.5 for red. This step aims to provide comprehensive and accurate technical support for maintenance decisions, ensuring the safety and reliability of rail operation.
[0045] The detailed structure of the progressive wear recognition model is based on a graph convolutional network architecture, which abstracts the rail wear region as a graph structure for processing. The input layer receives the wear probability gain matrix, treating each pixel as a node in the graph, with the spatial relationships between adjacent pixels forming the edges. The first graph convolutional layer contains 128 neurons, using graph convolution operations to aggregate the neighbor information of each node. The aggregation function uses a weighted average method, with weights dynamically calculated based on the similarity between nodes. The first pooling layer uses graph pooling operations to merge similar nodes through hierarchical clustering, reducing computational complexity. The second graph convolutional layer also contains 128 neurons, further extracting high-level features from the graph structure. The second pooling layer continues the graph coarsening operation. The third graph convolutional layer contains 128 neurons, responsible for capturing the global features of the wear pattern. The output layer maps the graph features to the wear recognition result through a fully connected layer, outputting the probability value of each pixel belonging to the wear region.
[0046] The detailed steps for establishing the training dataset for the progressive wear identification model include three stages: data collection, annotation, and data augmentation. In the data collection stage, multimodal image acquisition equipment was deployed on different railway lines to collect rail image samples covering different years of use, operating conditions, and wear levels. The collection process lasted 18 months, covering all four seasons to ensure data representativeness and diversity. In the annotation stage, experienced railway maintenance experts manually annotated the collected images, including precise boundaries of wear areas, wear severity levels, and wear type classifications. To improve annotation accuracy, laser point cloud data was introduced as an auxiliary verification method, using 3D geometric information to verify the accuracy of wear area boundaries. In the data augmentation stage, various transformation operations were performed on the annotated samples, including random rotation, scaling, brightness adjustment, and Gaussian noise addition. The rotation angle range was -15 degrees to 15 degrees, the scaling ratio range was 0.8 to 1.2, the brightness adjustment range was ±20%, and the noise intensity was controlled at a signal-to-noise ratio above 30 dB. The final training dataset contains 12,000 normal rail samples, 8,000 lightly worn samples, 6,000 moderately worn samples, and 4,000 heavily worn samples, totaling 150,000 labeled images after data augmentation.
[0047] It's important to note that traditional convolutional neural network (CNN)-based methods treat images as regular grid structures, failing to fully utilize the spatial topological relationships and local connectivity of wear regions. In contrast, the progressive wear recognition model models wear regions as graph structures using graph convolutional networks (GCNNs). Each pixel acts as a node in the graph, and the relationships between adjacent pixels are represented as edges, enabling better capture of irregular shapes and complex boundary features in wear regions. Compared to traditional fully convolutional neural networks (WCNNs), GCNNs possess stronger geometric invariance and topological adaptability, handling wear regions of arbitrary shapes. The progressive processing mechanism refines feature representations step-by-step through multiple layers of graph convolution operations, and this progressive recognition process from local features to global patterns aligns better with human visual cognition. Compared to traditional machine learning methods based on support vector machines (SVMs), the progressive wear recognition model exhibits stronger feature learning capabilities and generalization performance, automatically learning effective feature representations from raw data and avoiding the limitations of manual feature design. The model's adaptive parameter adjustment mechanism dynamically adjusts the processing strategy based on the complexity of the wear region, optimizing computational efficiency while maintaining detection accuracy, achieving a good balance between precision and efficiency.
[0048] The key technical concepts of this invention include multimodal image fusion technology, progressive graph convolutional recognition technology, game theory optimization technology, and adaptive parameter adjustment technology. Multimodal image fusion technology, by simultaneously utilizing complementary information from visible light images and infrared thermal imaging images, provides a more comprehensive and accurate representation of wear features compared to traditional single-modal methods. Visible light images provide rich texture and shape information, while infrared images provide temperature distribution information. The fusion of these two technologies effectively distinguishes worn areas from normal areas, improving detection accuracy and robustness. Progressive graph convolutional recognition technology models the worn area as a graph structure and employs a progressive processing strategy. Compared to traditional grid-based convolutional methods, it has stronger geometric adaptability and topological representation capabilities, effectively handling irregular shapes and complex boundaries in worn areas. Through multi-level feature extraction and a progressive recognition process, it achieves accurate capture of wear patterns. Game theory optimization technology constructs a two-layer optimization model to balance the relationship between detection accuracy and computational efficiency. Compared with traditional single-objective optimization methods, it can find the optimal solution under multiple constraints. The upper-layer model aims to maximize detection accuracy, while the lower-layer model aims to minimize computational complexity. Coordination between the two objectives is achieved through coupling terms. Adaptive parameter adjustment technology dynamically adjusts model parameters according to the characteristics of the wear region. Compared with traditional methods with fixed parameters, it has stronger adaptability and flexibility, and can select the optimal processing strategy for wear conditions with different complexities and resolutions. The synergistic effect of these four key technical ideas forms a complete intelligent detection system. Multimodal fusion provides high-quality input data, progressive graph convolution achieves accurate feature recognition, game theory optimization ensures the global optimality of system performance, and adaptive adjustment ensures the flexible adaptation of the algorithm. Compared with existing traditional methods based on a single technical route, this invention achieves a comprehensive improvement in detection accuracy, computational efficiency, and adaptability through the synergy of multiple technologies, providing a more reliable and practical technical solution for intelligent detection of rail wear.
[0049] It should be noted that this invention also solves the following technical problems: Existing rail wear detection methods cannot effectively handle complex environmental interference and variable wear patterns. Traditional detection methods are prone to misjudgment when faced with environmental interference such as different lighting conditions, temperature changes, and surface stains. Furthermore, their ability to recognize complex wear patterns, such as irregular wear patterns and minor wear with blurred boundaries, is limited. This invention, through multimodal image fusion technology combining spatial detail information from visible light images and temperature distribution characteristics from infrared thermal imaging, can maintain stable detection performance under complex environmental conditions. The wear probability gain matrix uses a sigmoid gain function to effectively suppress background noise interference. The graph convolutional network in the progressive wear recognition model can handle the complex spatial topological relationships of the wear region. The adaptive parameter adjustment function dynamically adjusts the message passing steps according to the complexity coefficient of the wear region, ensuring accurate identification of various wear patterns. Existing rail wear detection methods lack the technical problem of predicting wear development trends and supporting intelligent maintenance decisions. Traditional detection methods can only provide wear status information at the current moment and cannot predict wear development trends and remaining service life, making it difficult to support preventative maintenance decisions. This invention establishes a wear development trend prediction model through time series analysis, combines historical wear data and current inspection results to predict the remaining service life of rails, and improves the K-means clustering algorithm to group similar wear features and calculate severity levels, providing a quantitative basis for maintenance decisions. The generated rail wear status assessment report not only includes current wear information but also provides maintenance suggestions and early warning information, realizing a technological upgrade from passive detection to proactive predictive maintenance.
[0050] Specifically, the principle of this invention is as follows: This invention solves the technical problems of low accuracy and excessive computational complexity in rail wear detection. Its fundamental principle lies in employing a multi-level progressive recognition architecture and a game-theoretic optimization strategy. First, visible light and infrared thermal imaging data are acquired through multi-modal image acquisition to establish an initial wear matrix containing three-dimensional coordinates and grayscale temperature information. This provides a rich multi-dimensional feature foundation for subsequent accurate identification, capturing more comprehensive wear feature information compared to single-modal detection. Second, a wear probability matrix is constructed using statistical methods. The wear probability distribution of each pixel is calculated through Bayesian inference, transforming wear identification from traditional binary judgment to probability quantification evaluation, significantly improving the accuracy and reliability of detection. Third, a two-level recognition architecture based on deep convolutional neural networks and multilayer perceptrons is designed. The first level performs coarse screening to identify potential wear areas, while the second level uses a progressive wear recognition model for fine identification. This hierarchical processing strategy ensures detection accuracy while avoiding high global computational complexity. Finally, a game theory optimization model is introduced to construct a two-layer optimization framework with the upper-level goal of maximizing accuracy and the lower-level goal of minimizing complexity. The two goals are balanced through coupling terms, which ensures that computational complexity is effectively controlled while maintaining high detection accuracy, so that the system can meet the performance requirements of real-time detection.
[0051] The following provides a specific embodiment 1 of the present invention, and the specific implementation of each step in this embodiment 1 is described in detail below.
[0052] The specific implementation of step S01 involves scanning the rail surface using a multimodal image acquisition device to obtain high-resolution visible light images and infrared thermal imaging images, establishing a multimodal image dataset, and constructing an initial wear matrix. Initial Wear Matrix The specific construction is represented as follows:
[0053] ;
[0054] In the formula, Position in the initial wear matrix The element vector at position; For pixels The corresponding horizontal spatial coordinates; For pixels The corresponding vertical space coordinates; For pixels The corresponding depth coordinates; For pixels The grayscale value of a visible light image; For pixels The infrared image temperature value. The spatial coordinates are calculated as follows:
[0055] ;
[0056] ;
[0057] ;
[0058] In the formula, This is a horizontal pixel size calibration coefficient, with a value ranging from 0.2 to 0.8 mm per pixel; This is a vertical pixel size calibration coefficient, with a value ranging from 0.2 to 0.8 mm per pixel; This represents the horizontal coordinate offset, in millimeters. This represents the vertical coordinate offset, in millimeters. This refers to the focal length parameter of the image acquisition device, in millimeters. This represents the baseline distance for stereo vision, in millimeters. For pixels The disparity value, in pixels. Note that the initial wear matrix... Each component has different physical dimensions, among which the coordinate components... The unit is millimeters, grayscale value The dimensionless numerical range is 0–255, and the temperature value is... The unit is Celsius, and in practical applications, normalization is required to ensure the stability of numerical calculations.
[0059] The specific implementation of step S02 involves preprocessing the multimodal image dataset, including image registration, noise filtering, and brightness normalization. Then, a multi-scale feature extraction algorithm is used to extract features from the preprocessed image to establish a wear probability matrix. Wear probability matrix The establishment of is specifically represented as follows:
[0060] ;
[0061] In the formula, For pixels There is a probability of wear and tear; For pixels Multi-scale feature vectors; This is the mean vector of the wear region features; This is the mean vector of the features of the normal region; Variance represents the characteristics of the wear zone; This represents the variance of the features in the normal region. The calculation of the multi-scale feature vector is expressed as follows:
[0062] ;
[0063] ;
[0064] In the formula, For the first Eigenvalues of each scale layer; For the first Window radius of each scale layer; For the first Gaussian kernel function for each scale layer; For position The grayscale value of the image at that location; The number of feature dimensions.
[0065] The specific implementation of step S03 is as follows: First-level probability identification of the wear probability matrix is performed based on a deep convolutional neural network to identify potential wear regions. Then, a multilayer perceptron model is used for feature fusion to establish a wear probability gain matrix. Wear probability gain matrix The establishment of is specifically represented as follows:
[0066] ;
[0067] In the formula, For pixels The wear probability gain value; For position Adaptive gain coefficient at the location; For position Offset parameter at that location. The calculation is expressed as follows:
[0068] ;
[0069] In the formula, This is the base offset parameter, with a default value of 0.5; This is the image quality offset coefficient, with a value of 0.15; This is the wear severity offset coefficient, with a value of 0.25; This represents the global mean of the image quality parameters; This represents the global mean of the wear severity parameter. The gain coefficient is calculated as follows:
[0070] ;
[0071] In the formula, This is the base gain coefficient, with a default value of 2.5. This is the image quality impact coefficient, with a value ranging from 0.1 to 0.3. The coefficient representing the influence of wear severity ranges from 0.2 to 0.5. For position Image quality parameters at the location; For position The degree of wear at the location. Image quality parameters. The calculation is expressed as follows:
[0072] ;
[0073] ;
[0074] ;
[0075] In the formula, This is the signal-to-noise ratio weighting coefficient, with a value of 0.6; This is the contrast weighting coefficient, with a value of 0.4; For position Signal-to-noise ratio at the location; For position Contrast at that location; For position Signal power at the location; For position Noise power at the location; For position The maximum gray value within the neighborhood; For position The minimum grayscale value within the neighborhood. A parameter indicating the severity of wear. The calculation is expressed as follows:
[0076] ;
[0077] In the formula, For position The wear depth at the point, in millimeters; For position The wear area at the point, in square millimeters; The depth of influence coefficient is 2.5.
[0078] The specific implementation of step S04 involves using a progressive wear identification model to perform a second-level fine-grained identification of the wear probability gain matrix. Wear boundaries are identified through adaptive threshold segmentation and edge detection algorithms, and a wear development trend prediction model is established by combining time series analysis. The wear region complexity coefficient is also considered. The calculation is expressed as follows:
[0079] ;
[0080] In the formula, The wear area complexity coefficient; The rate of change of curvature at the wear boundary; For wear shape irregularity; This is the curvature weighting coefficient, with a value of 0.6; This is the irregularity weighting coefficient, with a value of 0.4; This is the error term for complexity calculation, ranging from 0.01 to 0.05. Wear shape irregularity. The calculation is expressed as follows:
[0081] ;
[0082] In the formula, This represents the actual area of the wear zone, in square millimeters. The perimeter of the wear region is expressed in millimeters. The rate of change of curvature is calculated as follows:
[0083] ;
[0084] In the formula, The number of sampling points at the wear boundary; For the first The curvature value at each sampling point. Adaptive parameter adjustment function. The calculation is expressed as follows:
[0085] ;
[0086] In the formula, To transmit the adjustment value; This is a parameter for normalizing image resolution; For image quality normalization parameters; This is a parameter normalized to the severity of wear. Message passing steps. The calculation is expressed as follows:
[0087] ;
[0088] In the formula, The number of message passing steps for the progressive wear identification model.
[0089] The specific implementation of step S05 involves using a game theory optimization model to verify and optimize the wear identification results. This includes an upper-level model aimed at maximizing wear detection accuracy and a lower-level model aimed at minimizing computational complexity. The objective function of the upper-level model is as follows: The representation is as follows:
[0090] ;
[0091] In the formula, The number of true cases; The number of false positives; The number of false negatives; This refers to the coverage rate of the wear area; To calculate the time complexity. The objective function of the lower-level model. The representation is as follows:
[0092] ;
[0093] In the formula, Image processing time; This refers to memory usage. This represents the number of algorithm iterations. For hardware resource utilization. Coupling terms. The representation is as follows:
[0094] ;
[0095] In the formula, The correlation coefficient between detection accuracy and computational complexity; This is the weighting adjustment factor, with a value of 0.6.
[0096] The specific implementation of step S06 involves performing cluster analysis on the optimized wear identification results, and establishing a wear probability clustering matrix using an improved K-means clustering algorithm. Wear probability clustering matrix The establishment is represented as follows:
[0097] ;
[0098] In the formula, For the first The cluster in the th order of ... Statistical characteristics across each feature dimension; For the first Each cluster contains a set of pixels; For the first The number of pixels in each cluster; For position First Values for each feature dimension. Wear severity level. The calculation is expressed as follows:
[0099] ;
[0100] In the formula, The severity level of wear; The wear depth is expressed in millimeters. The wear area is expressed in square millimeters. This is the depth weighting coefficient, with a value of 3.2, in millimeters. ; This is the area weighting coefficient, with a value of 0.8, and the unit is millimeters. ; The severity assessment error term ranges from 0.1 to 0.3.
[0101] The specific implementation method of step S07 is the same as described above, and will not be repeated in detail here.
[0102] Regarding the methods for obtaining each parameter, The image quality parameters were obtained experimentally, including: Step 1: Calculating the signal-to-noise ratio (SNR) of the image using the ratio of signal power to noise power; Step 2: Calculating the contrast ratio of the image using the ratio of the difference between the maximum and minimum gray values to the sum of the differences; Step 3: Obtaining the image quality parameters by weighted averaging of the SNR and contrast ratio. The method used was experimental, including step 1: measuring the depth distribution of the wear area using a laser displacement sensor; step 2: calculating the area of the wear area; and step 3: calculating the wear severity parameter by the ratio of depth to area. The calculation method involves dividing the original resolution parameter by the standard resolution of 600 dpi and then normalizing it, specifically expressed as:
[0103] ;
[0104] In the formula, This is the resolution parameter of the input image, in dpi. and By respectively and It is obtained by dividing by its corresponding maximum value and then normalizing. , , The results were obtained by comparing the identification results with manually labeled real labels. It is obtained by recording the start and end time difference of image processing using the system clock. The system memory monitoring module obtains the peak memory usage during algorithm execution in real time. It is obtained by calculating the ratio of the number of pixels in the identified wear area to the number of pixels in the actual wear area. It is obtained through the iteration counter during the algorithm execution process. The average utilization of CPU and GPU is obtained through the system monitoring module.
[0105] The principle of the initial wear matrix construction formula is based on the multimodal information fusion theory. By spatially corresponding the texture information of the visible light image with the temperature information of the infrared image, a multidimensional quantitative characterization of the rail surface condition is achieved. Compared with the traditional single-modal method, it can provide more comprehensive wear feature information and effectively improve the data foundation quality of subsequent recognition algorithms.
[0106] The formula for establishing the wear probability matrix adopts the Bayesian inference principle, and its core form is as follows:
[0107] ;
[0108] This formula calculates the wear probability by comparing the similarity of pixel features with the feature distributions of worn and normal regions. Compared with traditional threshold segmentation methods, it has stronger statistical theoretical support and better noise robustness, and can effectively handle uncertainties and ambiguities in images.
[0109] The formula for establishing the wear probability gain matrix is based on the nonlinear amplification principle in signal processing, and its sigmoid function form is:
[0110] ;
[0111] This function amplifies useful wear feature signals while suppressing background noise by performing a nonlinear transformation on the original probability. Compared with linear enhancement methods, it can better preserve the dynamic range and detail information of the signal.
[0112] The formula for calculating the complexity coefficient of the wear region combines geometric shape analysis and topological characteristics, and its curvature calculation is based on the principles of discrete geometry:
[0113] ;
[0114] This formula provides a reliable basis for subsequent adaptive parameter adjustment by quantifying the curvature change and shape irregularity of the wear boundary, and has stronger adaptability than the traditional method with fixed parameters.
[0115] The adaptive parameter tuning function achieves dynamic optimization of model parameters by comprehensively considering multiple factors such as complexity, resolution, quality, and severity. Compared with tuning methods driven by a single factor, it can better balance detection accuracy and computational efficiency.
[0116] The objective function design of the game theory optimization model is based on multi-objective optimization theory, and its upper-level model adopts an accuracy maximization strategy.
[0117] ;
[0118] The lower-level model adopts a complexity minimization strategy, and achieves synergistic optimization of detection accuracy and computational complexity by constructing a game structure between the upper and lower levels. Compared with the traditional single-objective optimization method, it can find the global optimal solution under constraints, effectively solving the contradiction between accuracy and efficiency.
[0119] The formula for calculating the severity of wear is based on the physical meaning of the geometric characteristics of wear. It achieves a quantitative assessment of the degree of wear risk through a comprehensive evaluation of depth and area. Compared with evaluation methods based on a single indicator, it has stronger engineering practicality and decision-making guidance value.
[0120] To better understand and implement this invention, the following is a specific application scenario of this invention, Example 2:
[0121] A technical team needs to conduct rail wear condition inspections on a 5.8km long high-speed railway line. This line has been in operation for 7 years, with an average of 126 trains passing daily and a maximum operating speed of 350km / h. Due to long-term high-intensity operation, some rail sections have shown varying degrees of wear, requiring intelligent detection methods to accurately identify the wear areas and assess their severity.
[0122] The technical team first configured the multimodal image acquisition equipment according to the requirements of step S01. A high-resolution CCD camera, a Sony IMX455, was used, with an imaging resolution of 1920×1080 pixels and an exposure time controlled within 8ms. An infrared thermal imager, a FLIR A655sc, was selected, with a temperature detection range of -10℃ to 80℃ and a thermal sensitivity of 0.03℃. The equipment was installed on the inspection vehicle, maintaining a distance of 0.75m from the rail surface, and the acquisition angle was strictly perpendicular to the rail surface. During the three-day field acquisition process, the technical team acquired a total of 8960 visible light images and corresponding 8960 infrared thermal images, establishing a complete multimodal image dataset. Based on this dataset, the initial wear matrix was constructed with dimensions of 1920×1080×5, where the first two dimensions correspond to the image pixel coordinates, and the third dimension contains the X, Y, and Z three-dimensional spatial coordinates, as well as grayscale and temperature values.
[0123] In the preprocessing stage of step S02, the technical team systematically processed the acquired multimodal images. The image registration process employed the SIFT feature point detection algorithm, detecting an average of 487 feature points in each image pair, achieving a registration accuracy of 0.3 pixels. Noise filtering utilized a 5×5 window adaptive median filter, effectively removing random noise from the acquisition process. Brightness normalization adjusted the image grayscale values to a uniform range of 0–255, improving image contrast. A multi-scale feature extraction algorithm constructed a 5-layer Gaussian pyramid with a scale factor of 0.5, extracting edge and texture features at each scale layer. In the wear probability matrix established based on Bayesian inference, the probability values exhibited a clear spatial clustering characteristic, with the probability values of potential wear regions generally exceeding 0.6.
[0124] The deep convolutional neural network in step S03 adopted an improved LeNet architecture, containing four convolutional layers and two fully connected layers. The first convolutional layer used 32 3×3 convolutional kernels, the second used 64 kernels, the third used 128 kernels, and the fourth used 256 kernels. During network training, the technical team set the learning rate to 0.001, the batch size to 16, and trained for 150 epochs. The confidence threshold was set to 0.7, and a total of 328 potential wear regions were identified. The multilayer perceptron model contained three hidden layers, each with 512 neurons, and feature fusion was achieved through an attention mechanism. The established wear probability gain matrix adopted the sigmoid gain function, and the gain coefficient was dynamically adjusted according to the image quality, with an average gain coefficient of 2.3.
[0125] In the fine-grained identification stage of step S04, the technical team deployed a progressive wear recognition model based on graph convolutional networks. This model models the rail wear region as a graph structure, including an input layer, three graph convolutional layers, two pooling layers, and an output layer. The distribution of the transfer adjustment values calculated by the adaptive parameter adjustment function based on the wear region complexity coefficient and image resolution parameters is shown in Table 1.
[0126] Table 1. Distribution of transmission regulation values in different regions
[0127]
[0128] As shown in Table 1, the transmission moderating values differ significantly across regions. Region 4 has the highest transmission moderating value at 0.82, corresponding to the most complex wear pattern, requiring a 5-step message passing and multi-layer neighbor aggregation strategy. The adaptive threshold segmentation algorithm has a threshold range of 0.35–0.76, while the high threshold for Canny edge detection is set to 0.8 and the low threshold to 0.4. The wear development trend prediction model established through time series analysis uses an ARIMA(2,1,2) structure, achieving a prediction accuracy of 94.2%.
[0129] The game theory optimization model in step S05 consists of an upper-level model and a lower-level model. The upper-level model's input parameters include 1847 true positives, 126 false positives, 89 false negatives, a wear region coverage rate of 95.4%, and a computational time complexity of O(n log n). This operation involved the following steps. The input parameters for the lower-level model included an image processing time of 3.2 seconds, memory usage of 2.1 GB, CPU utilization of 76%, GPU utilization of 82%, and data transfer bandwidth of 45 MB / s. The weight adjustment factor for the coupling term was set to 0.6. After optimization using a genetic algorithm, the detection accuracy score improved to 96.8%, and the system resource consumption assessment value decreased to 78.5.
[0130] In the cluster analysis of step S06, the technical team used an improved K-means clustering algorithm, determining the number of clusters K=5 based on the elbow rule. Clustering features included wear depth, wear area, wear shape complexity, and wear location distribution. The density weighting mechanism used a Gaussian kernel function with a kernel parameter set to 0.5. The clustering process converged after 127 iterations, with the change in cluster center position less than 0.01. Figure 2 As shown, the scatter plot of rail wear detection areas clearly displays the spatial distribution characteristics of different wear levels along the 5.8km track. Light wear areas are mainly concentrated in the early section of the track, moderate wear areas are distributed in the middle section, and heavy wear areas mainly appear in certain locations under high load operation. Figure 3 As shown in the bar chart, the wear probability cluster matrix feature distribution reveals the characteristic distribution patterns of different clusters. The third cluster has the largest mean wear depth (2.3 mm) and a standard deviation of 0.4 mm, while the fifth cluster has a wear area of 287.3 mm. The complexity coefficient is 0.85. The distribution of wear severity levels calculated according to the fuzzy evaluation method is shown in Table 2.
[0131] Table 2. Distribution of Wear Severity Levels for Each Cluster
[0132]
[0133] As shown in Table 2, the five clusters exhibit a clear gradient distribution of wear severity, with cluster 5 showing the most severe wear, with an average depth of 4.2 mm, requiring immediate maintenance.
[0134] The rail wear condition assessment report established in step S07 includes detailed technical parameters. Wear location coordinates are converted using calibration parameters with a conversion factor of 0.52 mm / pixel. The wear depth measurement accuracy reaches 0.1 mm, and the area measurement accuracy reaches 1 mm. The estimated remaining service life is calculated based on a wear trend prediction model, taking into account the combined effects of an average wear rate of 0.15 mm / year, a service intensity coefficient of 1.8, and an environmental factor influence coefficient of 1.2. Figure 4As shown in the figure, the wear development trend prediction curve illustrates the wear depth development trend of four typical areas over 36 months. Area 4 shows the most rapid wear development, expected to reach the red warning threshold after 24 months. Area 1 shows relatively slow wear development, remaining within the green warning range for 36 months. Maintenance recommendations are generated based on severity levels: routine inspections are recommended after 6 months for areas with mild wear, intensive monitoring within 3 months for areas with moderate wear, maintenance within 1 month for areas with severe wear, and immediate replacement for areas with critical wear. A three-level warning mechanism is used, with a green warning threshold of 3.0, a yellow warning threshold of 6.0, and a red warning threshold of 8.5. A total of 156 green warnings, 73 yellow warnings, and 18 red warnings were issued during this inspection.
[0135] The technical team established a training dataset containing 150,000 labeled images during the model training phase, including 12,000 normal rail samples, 8,000 samples of light wear, 6,000 samples of moderate wear, and 4,000 samples of heavy wear. The training process employed the Adam optimization algorithm, with an initial learning rate of 0.001, decaying to 0.8 times the original rate every 30 epochs, a batch size of 32, and a total of 300 training epochs. The loss function used was a 2:1 weighted combination of cross-entropy loss and graph structure loss. Training stopped when the model achieved 96.3% accuracy on the validation set. This early stopping mechanism effectively prevented overfitting.
[0136] The progressive wear detection model's graph convolutional network architecture fully utilizes the spatial topological relationships of the wear region, with each pixel serving as a graph node and the relationships between adjacent pixels as graph edges. The first graph convolutional layer aggregates neighbor information using a weighted average method, with weights dynamically calculated based on node similarity. Graph pooling merges similar nodes through hierarchical clustering; the first pooling layer reduces the number of nodes from 1920×1080 to 960×540, and the second pooling layer further reduces it to 480×270. The output layer converts graph features into wear detection results through fully connected mapping, with the wear probability value for each pixel ranging from 0 to 1.
[0137] The entire testing process lasted 5 days, including 3 days for data acquisition and 2 days for data processing and analysis. The technical team successfully identified 328 wear areas, with a total wear area of [missing information]. The maximum wear depth was 4.2 mm, and the average wear depth was 1.8 mm. The consistency between the test results and manual visual inspection reached 97.1%, and the measurement error with the laser measuring equipment was controlled within ±0.05 mm.
[0138] This invention represents a significant technological advancement over traditional rail wear detection methods. Traditional manual visual inspection relies on the experience and judgment of inspectors, making it highly subjective and susceptible to the influence of lighting conditions and fatigue levels. In contrast, this invention utilizes multimodal image acquisition equipment to obtain visible light and infrared thermal imaging information, providing a more objective and comprehensive data foundation. Traditional contact-based measurement methods require interrupting train operation, resulting in low efficiency and high safety risks. This invention employs non-contact image detection technology, enabling continuous inspection without disrupting normal operations. Traditional methods typically only detect obvious wear defects, with limited ability to identify early, minor wear. This invention, through deep learning algorithms and probability matrix analysis, can identify minute wear signs imperceptible to the human eye, achieving early warning functionality. Traditional single-feature analysis methods are easily affected by noise interference and environmental changes. This invention, through multi-scale feature extraction and spatial topology modeling using graph convolutional networks, enhances the ability to identify complex wear patterns and improves its anti-interference capabilities. Traditional detection methods typically only provide current status information and lack predictions of future trends. However, this invention combines time series analysis to establish a wear development trend prediction model, which can estimate the wear development rate and remaining service life, providing forward-looking support for maintenance decisions.
[0139] It should be noted that the variables involved in this invention are explained in detail in Tables 3 and 4.
[0140] Table 3. Variable Explanation Table (Part 1)
[0141]
[0142] Table 4. Variable Explanation Table (Part Two)
[0143]
[0144] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for intelligent detection of railway rail wear, characterized in that, A multimodal image acquisition device was used to scan the rail surface to obtain visible light images and infrared thermal images, establishing a multimodal image dataset and constructing an initial wear matrix. The multimodal image dataset was preprocessed, and a multi-scale feature extraction algorithm was used to extract features from the preprocessed images to establish a wear probability matrix. A deep convolutional neural network was used to perform first-level probability recognition on the wear probability matrix to identify potential wear regions, and a multilayer perceptron model was used for feature fusion to establish a wear probability gain matrix. A progressive wear recognition model was used to perform second-level fine recognition on the wear probability gain matrix, identifying wear boundaries through adaptive threshold segmentation and edge detection algorithms, and a wear development trend prediction model was established by combining time series analysis. A game theory optimization model was used to verify and optimize the wear identification results; cluster analysis was performed on the optimized wear identification results, and an improved K-means clustering algorithm was used to establish a wear probability clustering matrix; a rail wear status assessment report was established based on the wear probability clustering matrix.
2. The intelligent detection method for railway rail wear according to claim 1, characterized in that, The construction steps of the initial wear matrix are specifically to form a two-dimensional data structure by numerically encoding the spatial coordinates and attribute information of each pixel in the multimodal image dataset. The rows of the matrix represent the vertical pixel positions of the image, the columns represent the horizontal pixel positions of the image, and each matrix element contains the three-dimensional spatial coordinates, grayscale value, and infrared temperature value of the corresponding pixel.
3. The intelligent detection method for railway rail wear according to claim 2, characterized in that, The preprocessing specifically includes image registration, noise filtering, and brightness normalization. The wear probability matrix is a probability distribution matrix calculated based on statistical methods, used to quantify the probability that each pixel belongs to the wear region. The matrix construction process includes feature vector extraction, probability density function fitting, and Bayesian inference calculation.
4. The intelligent detection method for railway rail wear according to claim 3, characterized in that, The wear probability gain matrix is specifically an enhancement matrix obtained by performing nonlinear transformation and signal amplification on the wear probability matrix. Its main function is to highlight the wear characteristic signal and suppress background noise interference. The gain function adopts the sigmoid function form, and the gain coefficient is adaptively adjusted according to the image quality parameters and wear severity parameters.
5. The intelligent detection method for railway rail wear according to claim 4, characterized in that, The progressive wear recognition model is based on a graph convolutional network architecture, which includes an input layer, three graph convolutional layers, two pooling layers, and an output layer. Each graph convolutional layer contains 128 neurons to process the spatial topological relationships of the wear region. The message passing mechanism in the model adopts a neighbor aggregation method.
6. The intelligent detection method for railway rail wear according to claim 5, characterized in that, The message passing steps of the progressive wear recognition model are dynamically adjusted through an adaptive parameter adjustment function based on the wear region complexity coefficient and the image resolution parameter. The wear region complexity coefficient is obtained by calculating the curvature change rate of the wear boundary and the irregularity of the wear shape. The image resolution parameter represents the pixel density of the input image, in pixels per inch (dpi).
7. The intelligent detection method for railway rail wear according to claim 6, characterized in that, The steps for establishing the training dataset of the progressive wear recognition model are as follows: Collect multimodal image samples of rails under different service years and operating conditions. Each sample contains the original multimodal image dataset and manually annotated wear area masks. At the same time, laser point cloud data is used as manually annotated auxiliary data to verify the accuracy of wear area boundaries and supplement three-dimensional geometric information.
8. The intelligent detection method for railway rail wear according to claim 7, characterized in that, The training steps of the progressive wear recognition model are as follows: the Adam optimization algorithm is used, the initial learning rate is set to 0.001, the learning rate is reduced to 0.8 times the original value every 30 training cycles, the batch size is set to 32, the total number of training cycles is 300, and the loss function is a weighted combination of cross-entropy loss and graph structure loss with a weight ratio of 2 to 1.
9. The intelligent detection method for railway rail wear according to claim 8, characterized in that, The game theory optimization model specifically includes an upper-level model that aims to maximize wear detection accuracy and a lower-level model that aims to minimize computational complexity. The two models achieve collaborative optimization through coupling terms. The objective function of the upper-level model is used to maximize wear detection accuracy, while the objective function of the lower-level model is used to minimize computational complexity.
10. The intelligent detection method for railway rail wear according to claim 9, characterized in that, The adaptive parameter adjustment function is specifically used to adjust the message passing steps of the progressive wear recognition model. It is calculated based on the wear region complexity coefficient, image resolution parameter, image quality parameter, and wear severity parameter to obtain a passing adjustment value. The corresponding message passing steps and neighbor aggregation strategy are set according to different ranges of the passing adjustment value.