Railway track intelligent detection system based on image fusion
By integrating high-resolution cameras, infrared thermal imagers and lidar sensors in the track detection system, combining convolutional neural networks and fast regional convolutional neural networks, high-precision and real-time detection of railway tracks are achieved, solving the problems of inefficiency of existing detection methods and inconsistent detection results, and ensuring the safety of railway transportation.
Patent Information
- Application Number
- CN202510478853.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-15
AI Technical Summary
The existing railway track detection methods are inefficient, the detection results are inconsistent, and it is difficult to find subtle defects. In addition, single image detection is prone to miss important information, affecting the safety of train operation.
High-resolution cameras, infrared thermal imagers and lidar sensors are used to collect track information, extract features through convolutional neural networks and perform wavelet transformation fusion, combine with fast regional convolutional neural networks for defect detection, and use 5G modules to transmit data to the monitoring center in real time.
It realizes high-precision, real-time detection and fault diagnosis of railway track status, and can identify subtle cracks, temperature abnormalities and geometric structure changes on the surface of the track, improves detection accuracy and efficiency, and ensures train safety.
Smart Images

Figure CN120495178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of railway track detection, and in particular to an intelligent railway track detection system based on image fusion. Background Art
[0002] As a key component of the nation's transportation infrastructure, railways play a vital role in economic and social development. In freight transport, they efficiently transport all kinds of production and living materials, ensuring industrial production and market supply, and promoting regional economic exchange. In passenger transport, their convenience and affordability make them the preferred choice for public transportation, facilitating the rapid and orderly flow of people and promoting cultural exchange. As the core load-bearing component of railways, the safety of tracks is directly related to the safety of train operations. Trains operate at high frequencies daily, subjecting tracks to immense pressure and friction, while also coping with a complex natural environment. Even the slightest anomaly can pose a serious safety risk.
[0003] For a long time, railway track inspection relied primarily on manual patrols. Inspectors walked along the tracks, using their eyes and simple tools to inspect. However, this method has numerous drawbacks. First, it requires a high level of manpower. With the continuous growth of railway operating mileage and the continuous addition of new lines, the labor costs required for comprehensive and regular inspections have skyrocketed. Furthermore, inspections are slow, with inspectors having to walk or use simple transportation, and the track length they can inspect each day is limited. Given the ever-expanding scale of railway operations, this inefficiency has become increasingly prominent. Furthermore, manual inspections are subject to significant subjective factors. Inspectors vary in their experience, commitment, and work attitude, leading to varying standards for determining track defects and their severity. Experienced inspectors can keenly detect minor wear, while those with less experience or a lack of focus are more likely to overlook them. This results in inconsistent and inaccurate inspection results, making it difficult to generate reliable and unified track safety assessment data. Furthermore, manual inspections struggle to detect subtle track defects, such as microcracks and initial wear. These subtle defects leave extremely subtle marks on the track surface, making them difficult to detect with the naked eye. However, the tracks are subjected to long-term pressure and vibration from trains, and untreated minor defects will gradually expand. Tiny cracks several millimeters long will expand to several centimeters or even longer over a long period of time, causing track breakage and seriously threatening train operation safety and railway transportation order.
[0004] With the development of computer vision technology, image-based track inspection methods are gradually being applied. However, existing single-image inspection methods have limitations. The track information they obtain is limited, reflecting only local conditions and easily missing important information. For example, when inspecting track fasteners, a single image often cannot cover all fasteners. Moreover, inspection accuracy is restricted by various factors. Poor image quality, improper shooting angles, complex lighting conditions, etc., can make it difficult for the inspection algorithm to accurately identify track defects and achieve ideal inspection results. Therefore, it is urgent to develop an intelligent railway track inspection system that can break through technical bottlenecks. This is of great significance to ensuring railway transportation safety and improving operational efficiency. It will provide more reliable and accurate technical means for railway track safety inspection and promote the intelligent and safe development of the railway transportation industry. Summary of the Invention
[0005] In order to solve the above problems, the purpose of the present invention is to provide an intelligent railway track detection system based on image fusion. By installing a high-resolution camera, an infrared thermal imager and a lidar sensor in front of the track inspection vehicle, the surface texture, temperature distribution and three-dimensional structure information of the track are collected respectively, and the collected information is transmitted to the control circuit board. The high-resolution camera is used to capture fine cracks and wear on the track surface, the infrared thermal imager is used to detect abnormal temperature points on the track, and the lidar sensor generates point cloud data to provide three-dimensional information such as track gauge and levelness; after the collected data is pre-processed by denoising, image enhancement, registration and normalization, a convolutional neural network (CNN) is used to extract features of multi-source images, and feature fusion is performed through a wavelet transform algorithm to generate a comprehensive feature image; a fast regional convolutional neural network (Faster The system uses R-CNN (Reverberation-based Convolutional Neural Network) to detect defects in comprehensive feature images, combines LiDAR point cloud data to measure track geometric parameters, and detects temperature anomalies based on infrared thermal imager data. The system controls 5G communication technology via a control circuit board to transmit data to a monitoring center in real time. When a track anomaly is detected, the system generates an alarm and alerts staff. The detection results and alarm information are stored in a distributed database for subsequent data analysis and historical record query, enabling high-precision, real-time detection and fault diagnosis of railway track conditions.
[0006] In order to achieve the above objectives, the present invention provides an intelligent railway track detection system based on image fusion, which is implemented as follows:
[0007] An intelligent railway track detection system based on image fusion includes a high-resolution camera, an infrared thermal imager, a lidar sensor, and a control box. The high-resolution camera, infrared thermal imager, and lidar sensor are installed at the front lower part of the track inspection vehicle, and the control box is installed on the center console of the track inspection vehicle. The control box is equipped with a control circuit board and a 5G module. The high-resolution camera collects surface texture image information of the track to capture fine cracks and wear on the track surface. The infrared thermal imager collects temperature distribution map information of the track to detect abnormal temperature points of the track. The lidar sensor collects three-dimensional structure map information of the track. The lidar sensor generates point cloud data to provide three-dimensional image information such as track gauge and levelness. The image information collected by the high-resolution camera, infrared thermal imager, and lidar sensor is transmitted to the control circuit board. After denoising, image enhancement, registration, and normalization preprocessing, a convolutional neural network (CNN) is used to extract the features of the surface texture image information, temperature distribution map information, and three-dimensional structure map information of the track, and feature fusion is performed through a wavelet transform algorithm to generate a comprehensive feature image. A fast regional convolutional neural network (Faster The system uses R-CNN (Reverberation-based Convolutional Neural Network) to detect defects in comprehensive feature images, combines lidar point cloud data to measure track geometric parameters, and detects temperature anomalies based on infrared thermal imager data. The control circuit board controls the 5G module to transmit data to the monitoring center in real time. When a track anomaly is detected, the system generates an alarm and alerts staff. The detection results and alarm information are stored in a distributed database for subsequent data analysis and historical record query, enabling high-precision, real-time detection and fault diagnosis of railway track status.
[0008] The present invention performs denoising, image enhancement, registration and normalization preprocessing on the surface texture image information, temperature distribution map information and three-dimensional structure map information of the track collected in the control circuit board, uses a Gaussian filtering algorithm to remove noise interference in the data, uses a histogram equalization algorithm to redistribute the image grayscale values, expand the dynamic range of the image grayscale, enhance the image contrast, and facilitate subsequent analysis, uses a scale-invariant feature transformation algorithm to detect feature points in the image, calculates the descriptors of the feature points, and then matches the feature points between different images to achieve accurate image registration, and uses a linear normalization algorithm to linearly map the image pixel values collected by the high-resolution camera, infrared thermal imager and lidar sensor to the [0,1] interval, eliminating the numerical differences caused by different data sources and achieving normalization processing.
[0009] The present invention adopts the convolutional neural network (CNN) to extract the characteristics of the track surface texture image information, temperature distribution map information, and three-dimensional structure map information as follows:
[0010] S1. Data input adjustment:
[0011] For surface texture image information, the preprocessed two-dimensional image data is input into the model input layer. For temperature distribution map information, it is also processed into a two-dimensional image format and input into the input layer. For three-dimensional structure map information, the point cloud data generated by the lidar is converted into a voxel grid image suitable for convolutional neural network (CNN) processing and then input into the model input layer.
[0012] S2. Convolutional layer feature extraction:
[0013] The data enters the convolution layer composed of 5 convolution blocks. Each convolution block contains 2 convolution layers with different sizes of convolution kernels (such as 3×3 and 5×5). Different combinations of convolution kernels process different information source data: Let the input image be I(x, y) and the convolution kernel be K(m, n), where (x, y) is the image pixel coordinate and (m, n) is the convolution kernel element coordinate. The mathematical formula of the convolution operation is:
[0014] (I*K)(i,j)=∑ m ∑ n I(i+m,j+n)K(m,n) (1)
[0015] In formula (1), (i, j) is the coordinate of the output feature map;
[0016] When processing surface texture image information, small convolution kernels extract subtle texture details, while large convolution kernels focus on overall geometric shape features. When processing temperature distribution map information, the convolution operation identifies features by analyzing the distribution of temperature values in the image. When the small convolution kernel slides on the temperature distribution map, it can perform a detailed analysis of the edge of the temperature anomaly area and detect the boundary where the temperature change is more drastic, thereby preliminarily outlining the general shape of the temperature anomaly area. The large convolution kernel, with its larger receptive field, comprehensively considers the larger area in the temperature distribution map, and can more accurately determine the size of the temperature anomaly area, judge the proportion of the area in the entire temperature distribution map, and judge its relationship with the surrounding normal temperature area. For three-dimensional structure map information, the convolution operation is committed to extracting the geometric structure features of the track. When the small convolution kernel acts on the voxel grid image, it can analyze the local details in the track geometry. Assuming the voxel value in the voxel grid image is V(x,y,z), the convolution kernel is U(m,n,p), and the three-dimensional convolution operation formula is:
[0017] (V*U)(i,j,k)=∑ m ∑ n ∑ p I(i+m,j+n,k+p)U(m,n,p) (2)
[0018] In formula (2), (i, j, k) is the coordinate of the output 3D feature map.
[0019] S3. Pooling layer data dimensionality reduction and feature retention:
[0020] The data processed by the convolution layer enters the pooling layer. The pooling layer adopts a combination of maximum pooling and average pooling. Maximum pooling retains the most significant features of the image, and average pooling obtains the overall features of the region. Suppose the input feature map is F(x,y). The maximum pooling operation selects the maximum value in the a×a window, which is expressed as:
[0021]
[0022] In formula (3), (i, j) is the coordinate of the output maximum pooling feature map, and the average pooling operation calculates the average value in the a×a window, which is expressed as:
[0023]
[0024] By combining maximum pooling and average pooling, the data dimension is reduced while retaining key information, improving the computational efficiency and generalization ability of the model.
[0025] S4.Fully connected layer feature integration:
[0026] The data after the pooling layer enters the fully connected layer, which integrates the feature vectors after convolution and pooling, and outputs a fixed-length feature representation to provide a basis for subsequent feature fusion. Suppose the input feature vector is x=(x1,x1,...,x n ), the weight matrix of the fully connected layer is W=(w ij ), the bias vector is b=(b1,b1,...,b m ), then the fully connected layer outputs y=(y1,y1,...,y n )for:
[0027] y=Wx+b (5)
[0028] S5. Model training optimization:
[0029] Prepare a large amount of labeled sample data containing normal track status and various defective statuses, input the sample data into the model for training, and use the cross entropy loss function to measure the difference between the model prediction result and the true label. Suppose the true label of the sample is t=(t1,t1,...,t c ), c is the number of categories, and the model prediction probability distribution is p=(p1,p1,...,p c ), the cross entropy loss function formula is:
[0030]
[0031] Use the stochastic gradient descent algorithm to optimize the model parameters. Let the model parameters be θ and the learning rate be α. In each iteration, the parameter update formula is:
[0032]
[0033] Where, It is the gradient of the cross entropy loss function L with respect to the parameter θ. Through continuous iterative optimization, the model can accurately extract features related to the track state.
[0034] The present invention adopts the wavelet transform algorithm to perform feature fusion scheme as follows:
[0035] S1. Wavelet decomposition:
[0036] The track surface texture image features, temperature distribution map features, and 3D structure map features extracted by convolutional neural networks (CNN) were subjected to wavelet decomposition. The Haar wavelet basis function was used to decompose each type of feature image into different frequency sub-bands, obtaining low-frequency approximate components and high-frequency detail components. The low-frequency approximate components reflect the overall outline and main features of the image, while the high-frequency detail components contain detailed information such as the image's edges and texture.
[0037] S2. Low-frequency approximate component fusion:
[0038] The low-frequency approximate components of different source images are fused by weighted averaging. The weights are determined according to the importance of different source images in track detection. Through analysis and experimental verification of a large amount of sample data, three appropriate weight coefficients ω1, ω2, and ω3 are determined, which correspond to the low-frequency approximate components of the track's surface texture image characteristics, temperature distribution map characteristics, and three-dimensional structure map characteristics, respectively. The fusion formula is:
[0039] L f =ω1×L1+ω2×L2+ω3×L3 (8)
[0040] In formula (1), L f is the low-frequency approximate component after fusion, L1 is the low-frequency approximate component of the track surface texture image, L2 is the low-frequency approximate component of the track temperature distribution map, and L3 is the low-frequency approximate component of the track three-dimensional structure map feature.
[0041] S3. High-frequency detail component fusion:
[0042] For high-frequency detail components, a fusion strategy based on regional energy is adopted. The track surface texture image is decomposed by wavelet to obtain the high-frequency sub-band map H1 of the surface texture image, the track temperature distribution map is decomposed by wavelet to obtain the high-frequency sub-band map H2 of the temperature distribution map, and the track three-dimensional structure map is decomposed by wavelet to obtain the high-frequency sub-band map H3 of the three-dimensional structure map. Each high-frequency sub-band image is evenly divided into M×N non-overlapping local small areas of equal size, denoted as Where k represents the source image type, k=1 represents the track surface texture image, k=2 represents the temperature distribution image, k=3 represents the three-dimensional structure image, i=1,2,...,M,,i=1,2,...,N, for each local small area Calculate its regional energy In pixel value Represents the pixels in the area, and the regional energy calculation formula is:
[0043]
[0044] At each corresponding local small area, compare the energy values from the track surface texture image, temperature distribution map, and three-dimensional structure map. For a small area at a specific position (i, j), if and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the track surface texture image. and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the position corresponding to the temperature distribution map. If and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the three-dimensional structure image;
[0045] S4. Reconstruct the fused image by inverse wavelet transform:
[0046] The fused low-frequency approximate component and high-frequency detail component are subjected to inverse wavelet transform. Through inverse wavelet transform, the fused different frequency sub-band information is recombined to reconstruct the comprehensive feature image.
[0047] The present invention adopts the fast regional convolutional neural network (Faster R-CNN) to detect defects in the comprehensive feature image as follows:
[0048] S1. Input preprocessing:
[0049] The initial size of the comprehensive feature image is (M, N). After preprocessing operations such as scaling and normalization, it is adjusted to (m, n) size and then input into the network to make the image meet the network requirements and improve the subsequent feature extraction effect;
[0050] S2. Feature extraction:
[0051] The input image enters the feature extraction layer consisting of 13 convolutional layers, 13 ReLU activation function layers, and 4 pooling layers. The convolutional layer uses different convolution kernels to extract image features by weighted summation of local pixel values. Small convolution kernels capture local features such as subtle textures, while large convolution kernels capture overall structural features. The ReLU activation function introduces nonlinearity to help the network learn complex patterns. The pooling layer uses maximum pooling to reduce data dimensionality, prevent overfitting, retain key features, and finally output a feature map.
[0052] S3. Region Proposal Network Processing:
[0053] The feature map is fed into the region proposal network layer, where it is further processed by a 3×3 convolution. It then passes through two parallel 1×1 convolutional layers. The first layer, after reshaping and applying a softmax activation function, outputs 18 values, which are used to determine the probability that nine different anchor boxes are foreground objects. The second layer, after reshaping, outputs 36 values, which predict four regression offset parameters for each anchor box relative to the ground-truth object box. Based on these values, candidate regions containing track defects are generated, and image information is used to assist in subsequent processing.
[0054] S4. Pooling operation of region of interest:
[0055] The candidate regions generated by the region proposal network and the features output by the feature extraction layer Figure 1 As for the input region of interest pooling layer, since the candidate regions vary in size and position, and the subsequent fully connected layer requires a fixed-size input, the region of interest pooling layer divides the feature map area corresponding to the candidate region into a fixed number of sub-regions for maximum pooling, and uniformly pools them into a fixed-dimensional feature vector for processing by the fully connected layer;
[0056] S5. Classification and Bounding Box Regression:
[0057] The feature vector after pooling of the region of interest is divided into the classification branch and the bounding box regression branch. The classification branch passes through five fully connected layers and then the Softmax activation function to output the probability of each category and determine the defect category in the candidate area. The bounding box regression branch outputs the bounding box prediction value through the fully connected layer, fine-tunes the position and size of the candidate area, and accurately frames the defect.
[0058] S6. Model training and inference:
[0059] Training: Prepare a large dataset of annotated, comprehensive feature images of normal and defective rails. Divide the dataset into training, validation, and test sets. Input the training data into the network, calculate the classification and regression losses of the region proposal network layer, as well as the corresponding losses of the classification and bounding box regression layers. Optimize network parameters using algorithms such as backpropagation and stochastic gradient descent. Use the validation set to adjust hyperparameters to prevent overfitting.
[0060] Inference: The comprehensive feature image to be detected is input into the trained network, processed through each layer in sequence, candidate regions are generated and classified and located, and non-maximum suppression is used to remove overlapping candidate frames, marking the location and category of the track defects on the image.
[0061] The Faster R-CNN uses a region proposal network to generate candidate regions and combines convolutional features for classification and bounding box regression, enabling efficient and accurate detection of various track surface defects. It also combines LiDAR point cloud data to measure track geometry and infrared thermal imager data to detect temperature anomalies, providing a comprehensive assessment of track condition.
[0062] The present invention controls the 5G module through the control circuit board to transmit the processed data to the monitoring center in real time. When a track abnormality is detected, the system automatically generates an alarm message, and reminds the staff of the monitoring center to deal with it in time. The detection results and alarm information are stored in a distributed database to facilitate subsequent data analysis and historical record query, providing data support for track maintenance and management.
[0063] The present invention uses a convolutional neural network to extract features, a wavelet transform algorithm to fuse features, and a fast regional convolutional neural network to detect defects. It also combines high-resolution cameras, lidar point cloud data, and infrared thermal imager data to assess track status. It also uses a 5G module to transmit data to a monitoring center in real time, and stores detection results and alarm information in a distributed database. This can achieve the following beneficial effects:
[0064] 1. Multi-source sensors collect rich information. Different convolution kernels extract subtle and macro features from different information source data in the convolutional neural network. Wavelet transform fuses the image features of each source. The fast regional convolutional neural network combines multiple aspects of information to accurately detect defects. It can identify subtle cracks on the track surface, temperature anomaly boundaries, and details of geometric structure changes, significantly improving detection accuracy.
[0065] 2. By integrating track surface texture, temperature distribution, and 3D structural information, it can not only detect track surface defects, but also measure track geometric parameters based on LiDAR point cloud data and detect temperature anomalies based on infrared thermal imager data, achieving a comprehensive assessment of track status. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 This is a schematic diagram of the installation structure of a railway track intelligent detection system based on image fusion according to the present invention;
[0067] Figure 2 This is a structural schematic diagram of a control box of a railway track intelligent detection system based on image fusion according to the present invention;
[0068] Figure 3This is a flowchart of multi-source image information preprocessing of a railway track intelligent detection system based on image fusion according to the present invention;
[0069] Figure 4 This is a flow chart of a multi-source image feature extraction solution for a railway track intelligent detection system based on image fusion according to the present invention;
[0070] Figure 5 This is a network structure diagram of image feature extraction for an intelligent railway track detection system based on image fusion according to the present invention;
[0071] Figure 6 This is a structural principle diagram of a railway track intelligent detection system based on image fusion and feature fusion based on a wavelet transform algorithm according to the present invention;
[0072] Figure 7 This is a structural principle diagram of a railway track intelligent detection system based on image fusion in the present invention.
[0073] Description of main component symbols.
[0074] High-resolution camera 1 LiDAR sensor 2 Infrared thermal imager 3 control box 4 Metal aluminum box 5 5G module 6 Control circuit board 7 DETAILED DESCRIPTION
[0075] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings.
[0076] See also Figures 1 to 7 The figure shows an image fusion-based intelligent railway track detection system in the present invention, which includes a high-resolution camera 1, an infrared thermal imager 3, a lidar sensor 2, and a control box 4.
[0077] like Figure 1 and Figure 2As shown, the high-resolution camera 1, infrared thermal imager 3, and lidar sensor 2 are installed at the front lower part of the track inspection vehicle, and the control box 4 is installed on the center console of the track inspection vehicle, and the control circuit board 7 and 5G module 6 are installed in the control box 4. The high-resolution camera 1 collects the surface texture image information of the track to capture the fine cracks and wear on the track surface, the infrared thermal imager 3 collects the temperature distribution map information of the track to detect abnormal temperature points of the track, and the lidar sensor 2 collects the three-dimensional structure map information of the track. The lidar sensor 2 generates point cloud data to provide three-dimensional image information such as track gauge and levelness; the image information collected by the high-resolution camera 1, infrared thermal imager 3, and lidar sensor 2 is transmitted to the control circuit board 7, and is pre-processed by denoising, image enhancement, registration and normalization. After processing, a convolutional neural network (CNN) is used to extract the features of the track's surface texture image information, temperature distribution map information, and three-dimensional structure map information, and the wavelet transform algorithm is used to perform feature fusion to generate a comprehensive feature image; a fast regional convolutional neural network (FasterR-CNN) is used to detect defects in the comprehensive feature image, and the track geometric parameters are measured in combination with the lidar point cloud data, and temperature anomaly information is detected based on the infrared thermal imager 3 data; the 5G module 6 is controlled by the control circuit board 7 to transmit data to the monitoring center in real time. When a track anomaly is detected, the system generates an alarm message and reminds the staff. At the same time, the detection results and alarm information are stored in a distributed database to facilitate subsequent data analysis and historical record query, thereby achieving high-precision, real-time detection and fault diagnosis of the railway track status.
[0078] like Figure 3 As shown, the control circuit board 7 performs denoising, image enhancement, registration and normalization preprocessing on the collected surface texture image information, temperature distribution map information and three-dimensional structure map information of the track, and uses the Gaussian filtering algorithm to remove noise interference in the data. The Gaussian filtering performs weighted averaging on the image pixels based on the Gaussian function. While effectively smoothing the image and removing noise, it can better retain key information such as the image edge, significantly improving the data quality. The histogram equalization algorithm is used to redistribute the image grayscale value, expand the dynamic range of the image grayscale, enhance the image contrast, and make the detailed features of the track surface, such as subtle Cracks, wear marks, etc. are clearer, image enhancement is achieved, and subsequent analysis is facilitated. The scale-invariant feature transformation algorithm is used to detect feature points in the image, calculate the descriptors of the feature points, and then match the feature points between different images to achieve accurate image registration, providing a basis for the subsequent fusion of different sensor data. The linear normalization algorithm is used to linearly map the image pixel values collected by the high-resolution camera 1, infrared thermal imager 3, and lidar sensor 2 to the [0,1] interval, eliminating the numerical differences caused by different data sources, achieving normalization processing, and providing a good data foundation for subsequent feature extraction and fusion.
[0079] like Figure 4 and Figure 5 As shown, the present invention uses a convolutional neural network (CNN) to extract the features of the track's surface texture image information, temperature distribution map information, and three-dimensional structure map information:
[0080] S1. Data input adjustment:
[0081] For surface texture image information, the preprocessed two-dimensional image data is input into the model input layer. For temperature distribution map information, it is also processed into a two-dimensional image format and input into the input layer. For three-dimensional structure map information, the point cloud data generated by the lidar is converted into a voxel grid image suitable for convolutional neural network (CNN) processing and then input into the model input layer.
[0082] S2. Convolutional layer feature extraction:
[0083] The data enters the convolution layer composed of 5 convolution blocks. Each convolution block contains 2 convolution layers with different sizes of convolution kernels (such as 3×3 and 5×5). Different combinations of convolution kernels process different information source data: Let the input image be I(x, y) and the convolution kernel be K(m, n), where (x, y) is the image pixel coordinate and (m, n) is the convolution kernel element coordinate. The mathematical formula of the convolution operation is:
[0084] (I*K)(i,j)=∑ m ∑ n I(i+m,j+n)K(m,n) (1)
[0085] In formula (1), (i, j) is the coordinate of the output feature map;
[0086] When processing surface texture image information, small convolution kernels extract subtle texture details, while large convolution kernels focus on overall geometric shape features. The receptive field of a small-sized 3×3 convolution kernel is relatively small. When scanning surface texture images, it can focus on local areas in the image. Since subtle texture details such as tiny cracks occupy a small area in the image, the small convolution kernel can keenly capture information such as the edges and grayscale changes of these subtle features by performing convolution operations such as weighted summation of local pixel values, thereby accurately extracting features such as tiny cracks on the track surface. The large 5×5 convolution kernel has a larger receptive field. When processing surface texture images, it focuses on a relatively large area in the image, which enables the large convolution kernel to comprehensively consider the information of the five local areas, thereby more effectively extracting the overall geometric shape features of the track, such as the straight segments, curved segments and other macroscopic morphological features of the track. Processing temperature distribution When the image is used for information, the convolution operation identifies features by analyzing the distribution of temperature values in the image. When the small convolution kernel slides on the temperature distribution map, it can perform a detailed analysis of the edge of the temperature anomaly area and detect the boundary where the temperature changes more dramatically, thereby preliminarily outlining the general shape of the temperature anomaly area. The large convolution kernel, with its larger receptive field, comprehensively considers the larger area in the temperature distribution map, and can more accurately determine the size of the temperature anomaly area, judge the proportion of the area in the entire temperature distribution map, and its relationship with the surrounding normal temperature area. For three-dimensional structural image information, the convolution operation is committed to extracting the geometric structure features of the track. When the small convolution kernel acts on the voxel grid image, it can analyze the local details in the track geometry. Assuming the voxel value in the voxel grid image is V(x,y,z), the convolution kernel is U(m,n,p), and the three-dimensional convolution operation formula is:
[0087] (V*U)(i,j,k)=∑ m ∑ n ∑ p I(i+m,j+n,k+p)U(m,n,p) (2)
[0088] In formula (2), (i, j, k) is the coordinate of the output 3D feature map.
[0089] In areas where track gauge varies, a small convolution kernel can capture the location and trend of subtle gauge changes by analyzing the properties of local voxels in the voxel grid. Large convolution kernels, due to their larger receptive field, can analyze the track's geometry from a more macroscopic perspective when processing 3D structural information. For example, by integrating information from four voxels, it can accurately extract the curvature of the track, including key information such as the direction and magnitude of the curvature.
[0090] S3. Pooling layer data dimensionality reduction and feature retention:
[0091] The data processed by the convolution layer enters the pooling layer. The pooling layer uses a combination of maximum pooling and average pooling. Maximum pooling retains the most significant features of the image, and average pooling obtains the overall characteristics of the region. The combination of the two reduces the data dimension and retains key information, improving the model's computational efficiency and generalization ability. Suppose the input feature map is F(x,y). The maximum pooling operation selects the maximum value in the a×a window, which is expressed as:
[0092]
[0093] In formula (3), (i, j) is the coordinate of the output maximum pooling feature map, and the average pooling operation calculates the average value in the a×a window, which is expressed as:
[0094]
[0095] By combining maximum pooling and average pooling, the data dimension is reduced while retaining key information, improving the computational efficiency and generalization ability of the model.
[0096] S4.Fully connected layer feature integration:
[0097] The data after the pooling layer enters the fully connected layer, which integrates the feature vectors after convolution and pooling. Suppose the input feature vector is x=(x1,x1,...,x n ), the weight matrix of the fully connected layer is W=(w ij ), the bias vector is b=(b1,b1,...,b m ), then the fully connected layer outputs y=(y1,y1,...,y n )for:
[0098] y=Wx+b (5)
[0099] S5. Model training optimization:
[0100] Prepare a large amount of labeled sample data containing normal track status and various defective statuses, input the sample data into the model for training, and use the cross entropy loss function to measure the difference between the model prediction result and the true label. Suppose the true label of the sample is t=(t1,t1,...,t c ), c is the number of categories, and the model prediction probability distribution is p=(p1,p1,...,p c ), the cross entropy loss function formula is:
[0101]
[0102] Use the stochastic gradient descent algorithm to optimize the model parameters. Let the model parameters be θ and the learning rate be α. In each iteration, the parameter update formula is:
[0103]
[0104] Where, It is the gradient of the cross entropy loss function L with respect to the parameter θ. Through continuous iterative optimization, the model can accurately extract features related to the track state.
[0105] like Figure 6 As shown, the present invention adopts the wavelet transform algorithm to perform feature fusion scheme as follows:
[0106] S1. Wavelet decomposition:
[0107] Wavelet decomposition is performed on the track surface texture image features, temperature distribution map features, and three-dimensional structure map features extracted by convolutional neural networks (CNN). The Haar wavelet basis function is used to decompose each type of feature image into different frequency sub-bands, obtaining low-frequency approximate components and high-frequency detail components. The low-frequency approximate components reflect the overall outline and main features of the image, while the high-frequency detail components contain detailed information such as the edges and texture of the image. For the track surface texture image features, the high-frequency detail components can highlight edge information such as fine cracks and wear marks; the high-frequency detail components of the temperature distribution map features help to detect the boundaries of temperature anomaly areas; and the high-frequency detail components of the three-dimensional structure map features can better reflect the details of the track geometric structure changes, such as the characteristics of small changes in track gauge.
[0108] S2. Low-frequency approximate component fusion:
[0109] The low-frequency approximate components of different source images are fused using a weighted average method. The weights are determined based on the importance of different source images in track detection. When detecting the macroscopic geometry of the track, the low-frequency approximate components of the 3D structural image features have a larger weight; when detecting track surface defects, the low-frequency approximate components of the surface texture image features have a relatively higher weight. Through analysis and experimental verification of a large amount of sample data, appropriate weight coefficients ω1, ω2, and ω3 are determined, corresponding to the low-frequency approximate components of the track's surface texture image features, temperature distribution map features, and 3D structural map features, respectively. The fusion formula is:
[0110] L f =ω1×L1+ω2×L2+ω3×L3 (8)
[0111] In formula (1), L f is the low-frequency approximate component after fusion, L1 is the low-frequency approximate component of the track surface texture image, L2 is the low-frequency approximate component of the track temperature distribution map, and L3 is the low-frequency approximate component of the track three-dimensional structure map feature.
[0112] S3. High-frequency detail component fusion:
[0113] For high-frequency detail components, a fusion strategy based on regional energy is adopted. The track surface texture image is decomposed by wavelet to obtain the high-frequency sub-band map H1 of the surface texture image, the track temperature distribution map is decomposed by wavelet to obtain the high-frequency sub-band map H2 of the temperature distribution map, and the track three-dimensional structure map is decomposed by wavelet to obtain the high-frequency sub-band map H3 of the three-dimensional structure map. Each high-frequency sub-band image is evenly divided into M×N non-overlapping local small areas of equal size, denoted as Where k represents the source image type, k=1 represents the track surface texture image, k=2 represents the temperature distribution image, k=3 represents the three-dimensional structure image, i=1,2,...,M,,i=1,2,...,N, for each local small area Calculate its regional energy In pixel value Represents the pixels in the area, and the regional energy calculation formula is:
[0114]
[0115] At each corresponding local small area, compare the energy values from the track surface texture image, temperature distribution map, and three-dimensional structure map. For a small area at a specific position (i, j), if and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the track surface texture image. and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the position corresponding to the temperature distribution map. If and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the three-dimensional structure image;
[0116] S4. Reconstruct the fused image by inverse wavelet transform:
[0117] The fused low-frequency approximate components and high-frequency detail components are then subjected to an inverse wavelet transform. This inverse wavelet transform then recombines the fused frequency subbands to reconstruct a comprehensive feature image. This comprehensive feature image incorporates multiple characteristic features, including track surface texture, temperature distribution, and three-dimensional structure.
[0118] like Figure 7 As shown in FIG, the present invention adopts the fast regional convolutional neural network (Faster R-CNN) to detect defects in the comprehensive feature image as follows:
[0119] S1. Input preprocessing:
[0120] The initial size of the comprehensive feature image is (M, N). This image integrates multiple sources of information, such as the track's surface texture, temperature distribution, and three-dimensional structure. Using algorithms such as bilinear interpolation, the image is scaled and resized proportionally to make its length and width fit the target size. This ensures that the image meets the size specifications for subsequent network processing without losing key information. The image pixel values are normalized to a specific interval and resized to (m, n) before being input into the network. This ensures that the image meets the network requirements and improves the subsequent feature extraction effect.
[0121] S2. Feature extraction:
[0122] The input image enters the feature extraction layer consisting of 13 convolutional layers, 13 ReLU activation function layers, and 4 pooling layers. The convolutional layer uses different convolution kernels to extract image features by weighted summation of local pixel values. Small convolution kernels capture local features such as subtle textures, while large convolution kernels capture overall structural features. The ReLU activation function introduces nonlinearity to help the network learn complex patterns. The pooling layer uses maximum pooling to reduce data dimensionality, prevent overfitting, retain key features, and finally output a feature map.
[0123] S3. Region Proposal Network Processing:
[0124] The feature map is fed into the region proposal network layer, where it is further processed by a 3×3 convolution. It then passes through two parallel 1×1 convolutional layers. The first layer, after reshaping and applying a softmax activation function, outputs 18 values, which are used to determine the probability that nine different anchor boxes are foreground objects. The second layer, after reshaping, outputs 36 values, which predict four regression offset parameters for each anchor box relative to the ground-truth object box. Based on these values, candidate regions containing track defects are generated, and image information is used to assist in subsequent processing.
[0125] S4. Region of Interest Pooling Operation
[0126] The candidate regions generated by the region proposal network and the features output by the feature extraction layer Figure 1 As for the input region of interest pooling layer, since the candidate regions vary in size and position, and the subsequent fully connected layer requires a fixed-size input, the region of interest pooling layer divides the feature map area corresponding to the candidate region into a fixed number of sub-regions for maximum pooling, and uniformly pools them into a fixed-dimensional feature vector for processing by the fully connected layer;
[0127] S5. Classification and Bounding Box Regression:
[0128] The feature vector after pooling of the region of interest is divided into the classification branch and the bounding box regression branch. The classification branch passes through five fully connected layers and then the Softmax activation function to output the probability of each category and determine the defect category in the candidate area. The bounding box regression branch outputs the bounding box prediction value through the fully connected layer, fine-tunes the position and size of the candidate area, and accurately frames the defect.
[0129] S6. Model training and inference:
[0130] Training: Prepare a large dataset of annotated, comprehensive feature images of normal and defective rails. Divide the dataset into training, validation, and test sets. Input the training data into the network, calculate the classification and regression losses of the region proposal network layer, as well as the corresponding losses of the classification and bounding box regression layers. Optimize network parameters using algorithms such as backpropagation and stochastic gradient descent. Use the validation set to adjust hyperparameters to prevent overfitting.
[0131] Inference: The comprehensive feature image to be detected is input into the trained network, processed through each layer in sequence, candidate regions are generated and classified and located, and non-maximum suppression is used to remove overlapping candidate frames, marking the location and category of the track defects on the image.
[0132] The Faster R-CNN uses a region proposal network to generate candidate regions and combines convolutional features for classification and bounding box regression, enabling efficient and accurate detection of various track surface defects. Furthermore, it combines LiDAR point cloud data to measure track geometry and uses infrared thermal imager data to detect temperature anomalies, providing a comprehensive assessment of track condition.
[0133] The present invention controls the 5G module 6 via the control circuit board 7 to transmit the processed data to the monitoring center in real time. When the system detects a track anomaly, it automatically generates an alarm message, which is used by the monitoring center to remind staff to handle it in a timely manner through various means such as pop-up windows, text messages, and voice prompts. At the same time, the test results and alarm information are stored in a distributed database. Distributed databases have advantages such as high reliability and high scalability, which facilitate subsequent data analysis. For example, data analysis can be used to mine the occurrence patterns of track defects, the distribution of defect types in different areas, and other information, providing strong data support for track maintenance and management, helping to formulate more reasonable maintenance plans and decisions, and ensuring the safe operation of railway tracks.
[0134] The working principle and working process of the present invention are as follows:
[0135] A high-resolution camera 1, an infrared thermal imager 3, and a lidar sensor 2 are installed at the front and lower part of the track inspection vehicle to collect information on the track surface texture, temperature distribution, and three-dimensional structure. The data is then transmitted to a control circuit board 7 in the control box 4. Within the control circuit board 7, Gaussian filtering is used for denoising, histogram equalization is used to enhance the image, a scale-invariant feature transformation algorithm is used for registration, and linear normalization is performed. Subsequently, a convolutional neural network (CNN) is used to extract features from different information source data using convolution kernels of varying sizes. Nonlinearity is introduced using the ReLU activation function, the pooling layer reduces the dimensionality to retain features, and the fully connected layer integrates the feature vectors. The model is trained using a cross-entropy loss function and a stochastic gradient descent algorithm to accurately extract features related to the track state. Next, the features extracted by the CNN are first decomposed using a wavelet transform algorithm. The low-frequency approximate components are then weighted averaged and fused, respectively, and the high-frequency detail components are fused based on regional energy. Finally, a comprehensive feature image is reconstructed using an inverse wavelet transform. Next, a Faster R-CNN (Fast R-CNN) is used to perform defect detection on the integrated feature image. The image input is preprocessed, and then a feature extraction layer undergoes multi-layer processing to output a feature map. The region proposal network generates candidate regions through convolution and branching operations. The region of interest pooling layer pools the feature map regions corresponding to the candidate regions into fixed-dimensional feature vectors. The classification and bounding box regression branches respectively determine the defect category and fine-tune the location and size of the candidate regions. The model is trained using the training set and optimized using the validation set. During inference, the images to be tested are processed sequentially, and overlapping candidate boxes are removed using non-maximum suppression. The defect locations and categories are then marked. Simultaneously, track geometry parameters are measured using lidar point cloud data, and temperature anomalies are detected based on data from the infrared thermal imager 3, providing a comprehensive assessment of the track condition. Finally, the control circuit board 7 controls the 5G module 6 to transmit the processed data in real time to the monitoring center. When an anomaly is detected, an alarm message is automatically generated to alert staff. The detection results and alarm messages are stored in a distributed database for subsequent data analysis, providing data support for track maintenance and management.
Claims
1. An intelligent railway track detection system based on image fusion, characterized by: The system includes a high-resolution camera, an infrared thermal imager, a lidar sensor, and a control box. The high-resolution camera, infrared thermal imager, and lidar sensor are installed at the front lower part of the track inspection vehicle, and the control box is installed on the center console of the track inspection vehicle. The control box is equipped with a control circuit board and a 5G module. The high-resolution camera collects the surface texture image information of the track to capture the fine cracks and wear on the track surface. The infrared thermal imager collects the temperature distribution map information of the track to detect abnormal temperature points of the track. The lidar sensor collects the three-dimensional structure map information of the track. The lidar sensor generates point cloud data to provide three-dimensional image information such as track gauge and levelness. The image information collected by the high-resolution camera, infrared thermal imager, and lidar sensor is transmitted to the control circuit board. After denoising, image enhancement, registration, and normalization preprocessing, a convolutional neural network (CNN) is used to extract the features of the surface texture image information, temperature distribution map information, and three-dimensional structure map information of the track, and the wavelet transform algorithm is used to perform feature fusion to generate a comprehensive feature image. The fast regional convolutional neural network (Faster regional convolutional neural network) is used. R-CNN) performs defect detection on the comprehensive feature image, combines the laser radar point cloud data to measure the track geometric parameters, and detects temperature anomaly information based on the infrared thermal imager data; the control circuit board controls the 5G module to transmit data to the monitoring center in real time. When a track anomaly is detected, the system generates an alarm message and reminds the staff. At the same time, the detection results and alarm information are stored in a distributed database for subsequent data analysis and historical record query, realizing high-precision, real-time detection and fault diagnosis of the railway track status; the control circuit board performs denoising and image enhancement on the collected track surface texture image information, temperature distribution map information, and three-dimensional structure map information. , registration and normalization preprocessing, use Gaussian filtering algorithm to remove noise interference in the data, use histogram equalization algorithm to redistribute the image grayscale value, expand the dynamic range of image grayscale, enhance the image contrast, and facilitate subsequent analysis, use scale-invariant feature transformation algorithm to detect feature points in the image, calculate the descriptor of feature points, and then match feature points between different images to achieve accurate image registration, use linear normalization algorithm to linearly map the image pixel values collected by high-resolution cameras, infrared thermal imagers, and lidar sensors to the [0,1] interval, eliminate the numerical differences caused by different data sources, and achieve normalization processing.
2. The railway track intelligent detection system based on image fusion according to claim 1 is characterized in that: The scheme for extracting the features of the track's surface texture image information, temperature distribution map information, and three-dimensional structure map information using a convolutional neural network (CNN) is as follows: S1. Data input adjustment: For surface texture image information, the pre-processed 2D image data is input into the model input layer. For temperature distribution information, it is also processed into a 2D image format and then input into the input layer. For 3D structure information, the point cloud data generated by the lidar is converted into a voxel grid image suitable for convolutional neural network (CNN) processing and then input into the model input layer. S2. Convolutional layer feature extraction: The data enters the convolution layer composed of 5 convolution blocks. Each convolution block contains 2 convolution layers with different sizes of convolution kernels (such as 3×3 and 5×5). Different combinations of convolution kernels process different information source data: Let the input image be I(x, y) and the convolution kernel be K(m, n), where (x, y) is the image pixel coordinate and (m, n) is the convolution kernel element coordinate. The mathematical formula of the convolution operation is: (I*K)(i,j)=∑ m ∑ n I(i+m,j+n)K(m,n) (1) In formula (1), (i,j) is the coordinate of the output feature map; When processing surface texture image information, small convolution kernels extract subtle texture details, while large convolution kernels focus on overall geometric shape features. When processing temperature distribution map information, the convolution operation identifies features by analyzing the distribution of temperature values in the image. When the small convolution kernel slides on the temperature distribution map, it can perform a detailed analysis of the edges of the temperature anomaly area and detect the boundaries where the temperature changes more drastically, thereby preliminarily outlining the general shape of the temperature anomaly area. The large convolution kernel, with its larger receptive field, comprehensively considers larger areas in the temperature distribution map, and can more accurately determine the size of the temperature anomaly area, judge the proportion of the area in the entire temperature distribution map, and judge its relationship with the surrounding normal temperature areas. For the three-dimensional structure information, the convolution operation is dedicated to extracting the geometric structure features of the track. When a small convolution kernel acts on the voxel grid image, it can analyze the local details in the track geometry. Assuming the voxel value in the voxel grid image is V(x, y, z), the convolution kernel is U(m, n, p), and the three-dimensional convolution operation formula is: (V*U)(i,j,k)=∑ m ∑ n ∑ p I(i+m,j+n,k+p)U(m,n,p) (2)In formula (2), (i,j,k) is the coordinate of the output three-dimensional feature map; S3. Pooling layer data dimensionality reduction and feature retention: The data processed by the convolution layer enters the pooling layer. The pooling layer adopts a combination of maximum pooling and average pooling. Maximum pooling retains the most significant features of the image, and average pooling obtains the overall features of the region. Suppose the input feature map is F(x,y). The maximum pooling operation selects the maximum value in the a×a window, which is expressed as: In formula (3), (i, j) is the coordinate of the output maximum pooling feature map, and the average pooling operation calculates the average value in the a×a window, which is expressed as: By combining maximum pooling and average pooling, the data dimension is reduced while retaining key information, improving the computational efficiency and generalization ability of the model. S4.Fully connected layer feature integration: The data after the pooling layer enters the fully connected layer, which integrates the feature vectors after convolution and pooling, and outputs a fixed-length feature representation to provide a basis for subsequent feature fusion. Suppose the input feature vector is x=(x1,x1,...,x n ), the weight matrix of the fully connected layer is W=(w ij ), the bias vector is b=(b1,b1,...,b m ), then the fully connected layer outputs y=(y1,y1,...,y n )for: y=Wx+b (5) S5. Model training optimization: Prepare a large amount of labeled sample data containing normal track status and various defective statuses, input the sample data into the model for training, and use the cross entropy loss function to measure the difference between the model prediction result and the true label. Suppose the true label of the sample is t=(t1,t1,...,t c ), c is the number of categories, and the model prediction probability distribution is p=(p1,p1,...,p c ), the cross entropy loss function formula is: Use the stochastic gradient descent algorithm to optimize the model parameters. Let the model parameters be θ and the learning rate be α. In each iteration, the parameter update formula is: Where, It is the gradient of the cross entropy loss function L with respect to the parameter θ. Through continuous iterative optimization, the model can accurately extract features related to the track state.
3. The railway track intelligent detection system based on image fusion according to claim 1 is characterized in that: The scheme for feature fusion using wavelet transform algorithm is: S1. Wavelet decomposition: The track surface texture image features, temperature distribution map features, and 3D structure map features extracted by convolutional neural networks (CNN) were subjected to wavelet decomposition. The Haar wavelet basis function was used to decompose each type of feature image into different frequency sub-bands, obtaining low-frequency approximate components and high-frequency detail components. The low-frequency approximate components reflect the overall outline and main features of the image, while the high-frequency detail components contain detailed information such as the image's edges and texture. S2. Low-frequency approximate component fusion: The low-frequency approximate components of different source images are fused by weighted averaging. The weights are determined according to the importance of different source images in track detection. Through analysis and experimental verification of a large amount of sample data, three appropriate weight coefficients ω1, ω2, and ω3 are determined, which correspond to the low-frequency approximate components of the track's surface texture image characteristics, temperature distribution map characteristics, and three-dimensional structure map characteristics, respectively. The fusion formula is: L f =ω1×L1+ω2×L2+ω3×L3 (8) In formula (1), L f is the low-frequency approximate component after fusion, L1 is the low-frequency approximate component of the track surface texture image, L2 is the low-frequency approximate component of the track temperature distribution map, and L3 is the low-frequency approximate component of the track three-dimensional structure map feature. S3. High-frequency detail component fusion: For high-frequency detail components, a fusion strategy based on regional energy is adopted. The track surface texture image is decomposed by wavelet to obtain the high-frequency sub-band map H1 of the surface texture image, the track temperature distribution map is decomposed by wavelet to obtain the high-frequency sub-band map H2 of the temperature distribution map, and the track three-dimensional structure map is decomposed by wavelet to obtain the high-frequency sub-band map H3 of the three-dimensional structure map. Each high-frequency sub-band image is evenly divided into M×N non-overlapping local small areas of equal size, denoted as Where k represents the source image type, k=1 represents the track surface texture image, k=2 represents the temperature distribution image, k=3 represents the three-dimensional structure image, i=1,2,...,M,,i=1,2,...,N, for each local small area Calculate its regional energy In pixel value Represents the pixels in the area, and the regional energy calculation formula is: At each corresponding local small area, compare the energy values from the track surface texture image, temperature distribution map, and three-dimensional structure map. For a small area at a specific position (i, j), if and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the track surface texture image. and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the position corresponding to the temperature distribution map. If and The high-frequency detail component of the small area at this position after fusion adopts the high-frequency detail component of the corresponding position of the three-dimensional structure image; S4. Reconstruct the fused image by inverse wavelet transform: The fused low-frequency approximate component and high-frequency detail component are subjected to inverse wavelet transform. Through inverse wavelet transform, the fused different frequency sub-band information is recombined to reconstruct the comprehensive feature image.
4. The railway track intelligent detection system based on image fusion according to claim 1 is characterized in that: The scheme for defect detection using Faster Regional Convolutional Neural Network (FasterR-CNN) on comprehensive feature images is as follows: S1. Input preprocessing: The initial size of the comprehensive feature image is (M, N). After preprocessing operations such as scaling and normalization, it is adjusted to (m, n) size and then input into the network to make the image meet the network requirements and improve the subsequent feature extraction effect; S2. Feature extraction: The input image enters the feature extraction layer consisting of 13 convolutional layers, 13 ReLU activation function layers, and 4 pooling layers. The convolutional layer uses different convolution kernels to extract image features by weighted summation of local pixel values. Small convolution kernels capture local features such as subtle textures, while large convolution kernels capture overall structural features. The ReLU activation function introduces nonlinearity to help the network learn complex patterns; the pooling layer uses maximum pooling to reduce the data dimension, prevent overfitting, retain key features, and finally output feature maps; S3. Region Proposal Network Processing: The feature map is fed into the region proposal network layer, where it is further processed by a 3×3 convolution. It then passes through two parallel 1×1 convolutional layers. The first layer, after reshaping and applying a softmax activation function, outputs 18 values, which are used to determine the probability that nine different anchor boxes are foreground objects. The second layer, after reshaping, outputs 36 values, which predict four regression offset parameters for each anchor box relative to the ground-truth object box. Based on these values, candidate regions containing track defects are generated, and image information is used to assist in subsequent processing. S4. Pooling operation of region of interest: The candidate regions generated by the region proposal network are input into the region of interest pooling layer together with the feature map output by the feature extraction layer. Since the candidate regions vary in size and position, and the subsequent fully connected layer requires a fixed-size input, the region of interest pooling layer divides the feature map area corresponding to the candidate region into a fixed number of sub-regions for maximum pooling, and uniformly pools them into a fixed-dimensional feature vector for processing by the fully connected layer. S5. Classification and Bounding Box Regression: The feature vector after pooling of the region of interest is divided into the classification branch and the bounding box regression branch. The classification branch passes through 5 fully connected layers and then the Softmax activation function to output the probability of each category and determine the defect category in the candidate region; The bounding box regression branch outputs the bounding box prediction value through the fully connected layer, fine-tunes the position and size of the candidate region, and accurately frames the defect; S6. Model training and inference: Training: Prepare a large dataset of annotated, comprehensive feature images of normal and defective rails. Divide the dataset into training, validation, and test sets. Input the training data into the network, calculate the classification and regression losses of the region proposal network layer, as well as the corresponding losses of the classification and bounding box regression layers. Optimize network parameters using algorithms such as backpropagation and stochastic gradient descent. Use the validation set to adjust hyperparameters to prevent overfitting. Inference: The comprehensive feature image to be detected is input into the trained network. After being processed through each layer, candidate regions are generated and classified and located. Non-maximum suppression is used to remove overlapping candidate frames, and the location and category of the track defects are marked on the image. The Faster R-CNN (Fast R-CNN) generates candidate regions through a region proposal network and combines convolutional features for classification and bounding box regression. It can efficiently and accurately detect various defects on the track surface. At the same time, it combines lidar point cloud data to measure track geometric parameters and detects temperature anomalies based on infrared thermal imager data to comprehensively assess the track status.