High-speed railway multi-professional data unified registration method and device based on visual features
By using a facility matching and consistency optimization method based on visual features, the problem of unified registration of multi-disciplinary inspection data in high-speed railways under complex operating conditions was solved, achieving high-precision data alignment, supporting joint analysis and trend judgment, and improving operation and maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ACADEMY OF RAILWAY SCI CORP LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to achieve high-precision, robust, and unified registration of multi-disciplinary inspection data for high-speed railways under complex operating conditions. In particular, the synchronization of track inspection and network inspection data suffers from system delays, dynamic distortions, and sparse anchor point dependencies, making it difficult to stably register data in a unified coordinate system.
By extracting robust facility features from images and performing feature-level or mileage-level matching and constraint correction with typical waveform features of track inspection and network inspection systems, the mapping relationship from mileage to image pixels is fitted using a random sample consistency algorithm, and consistency verification is performed through sequence consistency optimization and prior knowledge base, achieving high-precision registration of multi-disciplinary data in a unified coordinate system.
It effectively eliminates mileage deviations caused by system delays and dynamic distortions, achieves high-precision registration of multi-disciplinary data in a unified coordinate system, supports cross-domain joint analysis and trend judgment, and improves the accuracy and efficiency of railway infrastructure operation and maintenance.
Smart Images

Figure CN121921345A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of railway infrastructure operation and maintenance and inspection technology, and in particular to a method and device for unified registration of multi-disciplinary data of high-speed railway based on visual features. Background Technology
[0002] In the field of operation and maintenance and inspection of high-speed railway infrastructure, the core of ensuring train operation safety and facility health lies in the accurate analysis and correlation diagnosis of dynamic inspection data from multiple disciplines such as track inspection and network inspection. However, the existing technical system faces a series of long-standing technical challenges in achieving unified registration of multi-source data, which seriously restricts the effectiveness and reliability of joint analysis.
[0003] First, traditional sensor calibration-based data synchronization methods have inherent limitations. Currently, dynamic detection equipment such as track inspection and network inspection generally rely on onboard mileage pulses, IMU or GNSS (Inertial Measurement Unit or Global Navigation Satellite System) and multi-source timestamps for data calibration. Although integrated mileage identification is used, due to various factors such as sensor response delay, clock asynchrony, acquisition link and communication buffer delay, speed fluctuations, and pulse quantization errors, there are still non-negligible residual deviations between different detection systems even if they receive the same mileage signal. More complexly, slippage, wheel diameter differences, temperature drift, and changes in track conditions during train operation can cause dynamic and non-linear distortions in the mapping relationship from the time domain to the mileage domain, and even local non-monotonic phenomena. This makes it difficult to achieve stable and accurate registration of track inspection waveform data and network inspection waveform data under a unified mileage coordinate system, thus greatly increasing the difficulty of cross-disciplinary joint analysis and fault correlation diagnosis.
[0004] Secondly, registration methods relying on external anchor points face significant challenges in practical applications. To correct the aforementioned deviations, existing technologies often rely on GNSS signals, transponders, or artificially placed ground markers as sparse anchor points. However, in areas where GNSS signals are missing or attenuated, such as tunnels and canyons, and over long distances, such sparse anchor points are unavailable or insufficient in density, failing to effectively suppress the accumulation and drift of registration errors. Furthermore, traditional correction methods based on waveform cross-correlation or template matching are highly susceptible to mismatches and exhibit poor robustness under complex conditions with high data noise and strong non-stationary interference.
[0005] Furthermore, while machine vision-based alternatives hold promise, their direct application in line inspection scenarios still faces technical bottlenecks. Machine vision offers the intuitiveness and repeatability of "what you see is what you get," and long strip images acquired by line scan cameras can stably present the spatial distribution and geometric features of lines and line-side facilities (such as overhead contact line supports and turnouts) in a high-resolution, uniformly accumulated manner, serving as an ideal common reference system. However, line scan images suffer from problems such as densely packed small targets, frequent reflections, and occlusions, leading to persistently high rates of missed and false detections in target detection.
[0006] In summary, existing technologies are insufficient to achieve high-precision and robust unified registration of multi-disciplinary testing data under complex working conditions. There is an urgent need for an innovative solution that can integrate multi-source information, overcome dynamic distortion, and possess self-verification capabilities. Summary of the Invention
[0007] To address the problems existing in the prior art, this invention proposes a method and device for unified registration of multi-disciplinary data in high-speed railways based on visual features. This invention extracts robust facility features from images and performs feature-level or mileage-level matching and constraint correction with typical waveform features from track inspection, network inspection, and other systems. This effectively eliminates mileage deviations introduced by system delays and dynamic distortions, achieving high-precision registration of multi-disciplinary data in a unified coordinate system, thereby supporting cross-domain joint analysis, condition assessment, and trend analysis.
[0008] In a first aspect of this invention, a unified registration method for multi-disciplinary data of high-speed railways based on visual features is proposed, the method comprising:
[0009] Acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and preprocess the multi-source data;
[0010] Feature points of track inspection facilities and feature points of catenary inspection facilities are extracted from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively. After identification and processing, the facility category and mileage location are output.
[0011] Image facility feature points are detected from preprocessed linear scan camera image data. After detection, facility category and image pixel coordinates are output. A unified coordinate system is constructed based on the image pixel coordinates.
[0012] Grouped by facility type, the feature points of track inspection facilities are matched with image facility points, and the network inspection facility points are matched with image facility points to obtain matching results;
[0013] Based on the matching results, mileage location, and image pixels, a random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency.
[0014] The consistency of the matching results is checked using a prior knowledge base. Missing points are filled and erroneous points are removed through a multi-source mutual verification mechanism to obtain the verified matching pairs.
[0015] Piecewise linear tiling is performed between adjacent matching pairs. A monotonic mileage-to-image pixel mapping function is generated based on the mileage-to-image pixel mapping relationship. The multi-source data is then mapped to the unified coordinate system to obtain the registration result.
[0016] In a second aspect of the present invention, a high-speed railway multi-disciplinary data unified registration device based on visual features is proposed, the device comprising:
[0017] The data acquisition module is used to acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and to preprocess the multi-source data.
[0018] The facility feature point advance module is used to extract track inspection facility feature points and catenary inspection facility feature points from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively, and output the facility category and mileage location after identification processing; it detects facility feature points from the preprocessed linear scan camera image data, outputs the facility category and image pixel coordinates after detection, and constructs a unified coordinate system based on the image pixel coordinates;
[0019] The cross-modal matching module is used to group facilities according to facility type and match feature points of track inspection facilities with image facility points and network inspection facility points with image facility points to obtain matching results. Based on the matching results, mileage location and image pixels, the random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency.
[0020] The verification module uses a prior knowledge base to perform consistency checks on the matching results, and fills in missing points and removes erroneous points through a multi-source mutual verification mechanism to obtain verified matching pairs.
[0021] The registration module is used to perform piecewise linear expansion between adjacent matching pairs, generate a monotonic mileage-to-image pixel mapping function based on the mileage-to-image pixel mapping relationship, and map the multi-source data to the unified coordinate system to obtain the registration result.
[0022] In a third aspect of the present invention, a computer device is proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a unified registration method for multi-disciplinary data of high-speed railway based on visual features.
[0023] In a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for unified registration of multi-disciplinary data for high-speed railways based on visual features.
[0024] In a fifth aspect of the present invention, a computer program product is proposed, the computer program product comprising a computer program that, when executed by a processor, implements a method for unified registration of multi-disciplinary data for high-speed railways based on visual features.
[0025] This invention proposes a unified registration method and device for multi-disciplinary data in high-speed railways based on visual features. This constructs an automated and robust registration process using visual features as a common benchmark, fundamentally solving the problem of difficulty in unifying and aligning multi-source detection data in high-speed railways caused by system delays, dynamic distortion, and sparse anchor point dependence. This invention accurately extracts facility feature points from one-dimensional track inspection and network inspection data and linear array images, respectively, and achieves precise cross-modal association between waveform facilities and image facilities based on classification matching and random sample consistency algorithms. Furthermore, by introducing a multi-source mutual verification mechanism driven by sequence consistency optimization and prior knowledge base, it effectively suppresses mismatches and missed matches, ensuring the global monotonicity and stability of the mapping from mileage to image pixels. With a piecewise linear unfolding strategy, it generates a continuous and high-precision unified coordinate mapping function, seamlessly aligning previously isolated multi-disciplinary data to the same visual coordinate system. This provides a strong data foundation for joint defect location, root cause analysis, and trend judgment, significantly improving the accuracy and efficiency of railway infrastructure operation and maintenance analysis. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of a method for unified registration of multi-disciplinary data in high-speed railways based on visual features, according to an embodiment of the present invention.
[0028] Figure 2 This is a schematic diagram of a method for unified registration of multi-disciplinary data in high-speed railways based on visual features, according to another embodiment of the present invention.
[0029] Figure 3 This is a schematic diagram illustrating the steps of a visual feature-based multi-disciplinary data unified registration method for high-speed railways according to a specific embodiment of the present invention.
[0030] Figure 4This is a schematic diagram of the orbital geometry one-dimensional data of the feature points of the facility to be extracted according to a specific embodiment of the present invention.
[0031] Figure 5 This is a schematic diagram of the architecture of a 1D ResNet deep model according to a specific embodiment of the present invention.
[0032] Figure 6 This is a schematic diagram of one-dimensional catenary data of the facility feature points to be extracted according to a specific embodiment of the present invention.
[0033] Figure 7 This is a schematic diagram of the data processing flow according to a specific embodiment of the present invention.
[0034] Figure 8 This is a schematic diagram of the architecture of an improved RT-DETR model according to a specific embodiment of the present invention.
[0035] Figure 9 This is a schematic diagram of the facility matching and sequence consistency optimization process according to a specific embodiment of the present invention.
[0036] Figure 10 This is a schematic diagram of a turnout facility for one-dimensional data recognition of track geometry according to a specific embodiment of the present invention.
[0037] Figure 11 This is a schematic diagram of a support pillar for one-dimensional data recognition of overhead contact lines according to a specific embodiment of the present invention.
[0038] Figure 12 This is a schematic diagram of the architecture of a high-speed railway multi-disciplinary data unified registration device based on visual features, according to an embodiment of the present invention.
[0039] Figure 13 This is a schematic diagram of a computer device structure according to an embodiment of the present invention. Detailed Implementation
[0040] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0041] Those skilled in the art will recognize that embodiments of the present invention can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0042] According to an embodiment of the present invention, a method and apparatus for unified registration of multi-disciplinary data for high-speed railways based on visual features are proposed, which relates to the field of railway infrastructure operation and maintenance and inspection technology.
[0043] This invention addresses the challenges of stable registration in a unified coordinate system for dynamic detection systems such as track inspection and network inspection, which suffer from residual deviations at the same mileage due to link delays, clock asynchrony, and speed fluctuations; nonlinear or even locally non-monotonic time-to-mileage mapping; the uncontrollable drift in long-distance registration caused by sparse or unavailable GNSS, transponders, or artificial landmarks; the susceptibility of traditional waveform correlation and template matching to mismatches under noise and non-stationary interference; the difficulty of high false or missed detection rates due to dense small targets and numerous reflections in long-span images from linear array cameras; and the difficulty in constructing a monotonic, segmented, high-precision "mileage-pixel" mapping from discrete facility correspondences, lack of self-verification for false or missed detections, prior completion and multi-source mutual verification mechanisms, lack of comprehensive quality assessment and online threshold adaptive or backoff strategies, and inability to achieve online drift correction. It proposes a unified registration method that can achieve robust cross-source anchor point matching and sequence consistency optimization under complex operating conditions.
[0044] The principles and spirit of the present invention will be explained in detail below with reference to several representative embodiments.
[0045] Figure 1 This is a schematic diagram of a method for unified registration of multi-disciplinary data in high-speed railways based on visual features, according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0046] S101, acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and preprocess the multi-source data;
[0047] S102: Extract feature points of track inspection facilities and feature points of catenary inspection facilities from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively, and output the facility category and mileage location after identification processing;
[0048] S103, detects image facility feature points from the preprocessed line scan camera image data, outputs facility category and image pixel coordinates after detection, and constructs a unified coordinate system based on the image pixel coordinates;
[0049] S104. Group the facilities according to their categories, match the feature points of track inspection facilities with image facility points, and the network inspection facility points with image facility points to obtain the matching results.
[0050] S105. Based on the matching results, mileage location and image pixels, the random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency.
[0051] S106, use the prior knowledge base to check the consistency of the matching results, and use the multi-source mutual verification mechanism to fill in missing points and remove erroneous points to obtain the checked matching pairs;
[0052] S107, segmented linear expansion is performed between adjacent matching pairs, and a monotonic mileage-to-image pixel mapping function is generated based on the mileage-to-image pixel mapping relationship to map the multi-source data to the unified coordinate system and obtain the registration result.
[0053] To provide a clearer explanation of the above-mentioned method for unified registration of multi-disciplinary data in high-speed railways based on visual features, each step will be explained in detail below.
[0054] In one embodiment, for S101, multi-source data is acquired, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and the multi-source data is preprocessed.
[0055] The specific methods for preprocessing the multi-source data include:
[0056] The one-dimensional track geometry data is denoised by digital filtering, bandpass filtering, or wavelet transform, and abnormal segments are verified in conjunction with odometer data.
[0057] The one-dimensional data of the overhead contact system is subjected to amplitude normalization, baseline drift correction, and robust filtering.
[0058] The image data from the line scan camera is subjected to distortion correction, illumination equalization, and image stitching processing.
[0059] In one specific embodiment, the one-dimensional track geometry data preprocessing involves using methods such as digital filtering, bandpass filtering, or wavelet transform to denoise and enhance the original data, and combining this with odometer data for verification and removal of abnormal segments to ensure data quality.
[0060] One-dimensional data preprocessing of overhead contact lines: noise and outliers are suppressed by robust filtering algorithms, amplitude normalization is performed on the data, and baseline drift is corrected to improve data comparability.
[0061] Linear scan camera image preprocessing: distortion correction is performed on the original image to eliminate uneven lighting and reflection interference, and image stitching technology is used to achieve consistent connection of the banner images.
[0062] In one embodiment, for S102, feature points of track inspection facilities and feature points of catenary inspection facilities are extracted from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively, and the facility category and mileage location are output after identification processing.
[0063] For one-dimensional orbital geometry data, the specific identification and processing flow includes:
[0064] Feature points of track inspection facilities are extracted from the preprocessed one-dimensional track geometry data, and an end-to-end 1DResNet depth model is used for identification, outputting facility category, mileage location and confidence level;
[0065] The 1DResNet deep model includes a Stem layer, multiple residual stages, and a classification head.
[0066] The Stem layer is a one-dimensional convolutional layer equipped with a BatchNorm layer and the SiLU activation function;
[0067] The residual stage includes narrow spike feature extraction, preset scale feature extraction, wide spike feature extraction, and contextual feature aggregation, which use parallel convolution kernels and attention mechanisms, respectively.
[0068] The classification head outputs the probability of facility point existence, and FocalLoss is used as the loss function.
[0069] For one-dimensional data of overhead contact lines, the specific identification and processing procedures include:
[0070] Feature points of the overhead contact system are extracted from the preprocessed one-dimensional data of the overhead contact system. The identification is performed based on feature engineering and XGBoost classifier, and the facility category, mileage location and confidence score are output.
[0071] The feature engineering process generates candidate inflection points based on the zero crossover of the first derivative and the extrema of the second derivative, and extracts time-domain, frequency-domain, morphological, and prior features.
[0072] The XGBoost classifier is trained using the XGBoost model for binary classification, and the evaluation metrics include at least PR-AUC and recall.
[0073] In one embodiment, for S103, image facility feature points are detected from the preprocessed linear scan camera image data, and the facility category and image pixel coordinates are output after detection. A unified coordinate system is constructed based on the image pixel coordinates.
[0074] The specific processing flow includes:
[0075] Facility feature points are detected from preprocessed linear scan camera image data. An improved RT-DETR model is used for end-to-end detection, and the facility category, image pixel coordinates, and confidence score are output.
[0076] The improved RT-DETR model includes: a Backbone feature extraction network, a hybrid editor, and a decoder;
[0077] The improved RT-DETR model employs a Backbone feature extraction network combined with multi-scale feature fusion; a hybrid editor combined with a deformable attention mechanism aggregates feature information at different scales; the decoder uses a matching algorithm and a decoupled detection head to jointly optimize classification and regression tasks; the training strategy of the improved RT-DETR model uses a multi-task loss function and adopts the AdamW optimizer and Cosine annealing learning rate scheduling.
[0078] In one specific embodiment, track inspection facility point extraction: using waveform morphology analysis and time-frequency feature detection methods, template matching or deep learning classifiers, characteristic facility points such as turnouts and insulation joints are identified from the one-dimensional track geometry data, and facility category and mileage location information are output.
[0079] Catenary inspection facility point extraction: Based on peak and valley detection, feature facility points such as supports and anchor joints are extracted from one-dimensional catenary data to determine their category and mileage location.
[0080] Image facility detection: Using object detection or semantic segmentation algorithms (such as anchor-based or anchor-free architecture), the location of facilities is detected from the linear scan camera image, and a mapping relationship between facility category and image coordinates is established.
[0081] In one embodiment, for S104, the feature points of track inspection facilities are grouped according to facility type, and the feature points of track inspection facilities are matched with image facility points, and the feature points of network inspection facilities are matched with image facility points to obtain matching results.
[0082] A cross-modal matching strategy is constructed, in which candidate pairs are grouped by facility category and selected based on mileage prior constraints and neighborhood gating mechanism. The candidate pairs are scored according to multidimensional scoring, including neighboring Gaussian kernel, order relation consistency, context distance ratio and confidence weighting. Candidate pairs that meet the preset confidence requirements are selected based on the score and soft threshold.
[0083] In one embodiment, for S105, based on the matching results, mileage location and image pixels, a random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency.
[0084] Cross-modal matching strategies mainly include:
[0085] Track inspection and image matching: Candidate pairs are constructed by grouping according to facility category, and scores are made based on proximity, order relation and context consistency. The Random Sample Consensus Algorithm (RANSAC algorithm) is used to fit the mapping relationship from mileage to pixel.
[0086] Network inspection and image matching: The processing method is the same as that of "track inspection and image matching", and it is completed independently to obtain the mapping relationship from mileage to pixel.
[0087] Sequence consistency optimization: Dynamic time warping (DTW) and the longest consistent subsequence algorithm are applied to eliminate skip points and cross-matches, ensuring the monotonicity and stability of the mapping.
[0088] In one embodiment, for S106, the consistency of the matching results is checked using a prior knowledge base, and missing points are filled and erroneous points are removed through a multi-source mutual verification mechanism to obtain a verified matching pair.
[0089] The prior knowledge base includes facility pairing rules, typical spacing ranges, and relative topological order;
[0090] The multi-source mutual verification mechanism includes cross-verification using the matching results of track inspection and network inspection.
[0091] In one specific embodiment, the prior knowledge base is a prior constraint library that establishes rules for the occurrence of facilities in pairs or groups (e.g., support-cantilever, insulation joint-anchor segment, etc.), typical spacing ranges, and relative topological order.
[0092] Consistency check: For cases where "image exists but waveform does not" or "waveform exists but image does not", verification is performed based on prior knowledge. Strong prior missing points are imputed, and weak consistency points are downweighted.
[0093] Multi-source mutual verification mechanism: Cross-validation is performed by using the matching results of track inspection and network inspection with the image. The confidence level is improved by the intersection, potential missed detections are found by the union, and screening is carried out by combining priors.
[0094] In one embodiment, for S107, piecewise linear tiling is performed between adjacent matching pairs, and a monotonic mileage-to-image pixel mapping function is generated based on the mileage-to-image pixel mapping relationship to map the multi-source data to the unified coordinate system and obtain the registration result.
[0095] Piecewise linear mapping: A linear interpolation method is used to evenly spread the matching pairs between adjacent pairs, achieving a continuous mapping from mileage to image coordinates. Let the matching pair be (m... i ,x i ), m i Let x be the mileage of the i-th known point. i For the i-th facility image pixel, in the interval [m i ,m i+1 Linear interpolation is used within the interval to achieve "uniform spreading" and obtain the pixels corresponding to the data points within the interval.
[0096] When performing piecewise linear unfolding, the calculation formula used is:
[0097]
[0098] In the formula, f(m) is the output result at the input point, representing the image pixel coordinates of the facility point's mileage m; x i Let x represent the image pixel coordinates of the i-th known point. i With m i The coordinates of the i-th known point (m) i ,x i );m i The distance to the i-th known point is represented by m; the distance to the facility point is represented by x. i+1 m i+1 Let represent the image pixel coordinates and mileage of the (i+1)th known point, respectively.
[0099] Smoothing and Constraints: Piecewise affine transformation or spline regularization is used to ensure the monotonicity and boundary continuity of the mapping, and the nearest matching pair is used to extrapolate for missing segments.
[0100] Output results: Generate track inspection mileage, image coordinates and network inspection mileage, and image coordinate calibration function to achieve accurate alignment of multi-disciplinary data in a unified coordinate system.
[0101] Further reference Figure 2 The method also includes:
[0102] S108, construct a fallback strategy, wherein when the quality index of the registration result does not meet the standard, the registration effect is optimized by relaxing the candidate threshold, increasing the prior weight, or switching to sparse anchor point alignment; when the registration quality index does not meet the standard, the candidate threshold, prior weight, or sparse anchor point alignment is adjusted for optimization.
[0103] Specifically, the fallback strategy is to optimize the system when quality metrics fail to meet the standards by relaxing the candidate threshold, increasing prior weights, or switching to sparse anchor alignment.
[0104] Application output: Supports the linked display of multimodal data, realizes joint defect localization and root cause analysis, and provides a data foundation for trend judgment.
[0105] It should be noted that although the operation of the method of the present invention has been described in a specific order in the above embodiments and figures, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0106] The following description uses a specific embodiment as an example.
[0107] refer to Figure 3This diagram illustrates the steps of a visual feature-based unified registration method for multi-disciplinary data in high-speed railways, according to a specific embodiment of the present invention. The method progressively processes, matches, and optimizes raw, disorganized multi-source data to output a high-precision unified registration result. Using a high-resolution, high-stability linear array image as a unified spatial scale, precise direct registration is achieved by matching track inspection data and network inspection data onto the image separately. Figure 3 As shown, the specific steps include:
[0108] Layer 1, Data Input and Preprocessing Layer:
[0109] This layer processes raw data from three different sources. Its primary function is to eliminate interference from equipment and the environment, providing high-quality, standardized data for subsequent feature extraction.
[0110] One-dimensional track geometry data: after denoising and enhancement, and verification of mileage accuracy, abnormal sections are removed.
[0111] One-dimensional data of the overhead contact line: After filtering, amplitude normalization and baseline drift correction, the signal characteristics are more obvious.
[0112] Line scan camera image: After distortion correction, illumination suppression, and strip stitching, a complete and clear panoramic image of the line is formed.
[0113] Layer 2, Facility Feature Extraction Layer:
[0114] Representative facility feature points, such as turnouts, insulation joints, and supports, are identified from multi-source data. This layer transforms the continuous data stream and images into a series of discrete feature points with category and location labels, preparing for subsequent matching.
[0115] Track inspection facility point extraction: Using a 1D end-to-end depth model, the category of the facility point and its precise mileage location are directly identified, and the facility category and mileage location are output.
[0116] Network inspection facility point extraction: Combining peak and valley detection, time and frequency feature analysis and other signal processing methods, the facility event points and their mileage are identified, and the facility category and mileage location are output.
[0117] Image Facility Detection: Using a deep learning object detection model, the facility is identified and its pixel coordinates on the image are obtained. The facility category and image coordinates are output.
[0118] Layer 3, Cross-modal matching strategy layer:
[0119] This layer mainly implements track inspection and image matching, and network inspection and image matching.
[0120] Track inspection and image matching: Candidate pairs are constructed by grouping according to facility category, and scores are made based on proximity, order relation and context consistency. The RANSAC algorithm is used to fit the mapping relationship from track inspection mileage to pixel.
[0121] Network inspection and image matching: Candidate pairs are constructed by grouping according to facility category, and scores are made based on proximity, order relation and context consistency. The RANSAC algorithm is used to fit the mapping relationship from network inspection mileage to pixels. The two matching processes are the same and are completed independently.
[0122] Sequence consistency optimization: Dynamic time warping (DTW) and the longest consistent subsequence algorithm are applied to eliminate skip points and cross-matches, ensuring the monotonicity and stability of the mapping.
[0123] In this layer, grouping and constraints are employed: grouping is done by facility category, and then a comprehensive score is calculated using proximity, order (who comes first, who comes last), and contextual consistency.
[0124] Robust fitting: Using robust algorithms such as RANSAC, a preliminary mapping relationship from "mileage" to "image pixels" is fitted, and then optimized using algorithms such as DTW (Dynamic Time Warping) to eliminate local mismatches.
[0125] Layer 4, False Alarm or Missed Alarm Handling and Prior Constraints Layer:
[0126] This layer uses prior knowledge to check and correct the matching results, which greatly improves the robustness and fault tolerance of the system and ensures that the results remain reliable even when some data is of poor quality.
[0127] Prior knowledge base: Establish a prior constraint base for the rules of the occurrence of facilities in pairs or groups, typical spacing ranges and relative topological order.
[0128] Consistency check: This checks whether the matching results conform to prior knowledge. This includes determining whether an image or waveform is present or absent, strong prior interpolation, and weak consistency reduction. For example, if a facility is clearly present in the image but not found in the waveform data, the interpolation mechanism will be activated.
[0129] Multi-source verification: The matching results of track inspection and network inspection are cross-validated. Points that are jointly confirmed have higher confidence, while points that are complementary to each other indicate possible missed detections.
[0130] Layer 5, Unified Coordinate System and Expanded Mapping Layer:
[0131] This layer transforms sparse matching points into continuous, dense mapping relationships, so that any mileage value can find its unique corresponding image location.
[0132] Piecewise linear mapping: A linear interpolation method is used to evenly spread the mileage between adjacent matching pairs to achieve a continuous mapping from mileage to image coordinates.
[0133] Smoothing and Constraints: Piecewise affine transformation or spline regularization is used to ensure the monotonicity and boundary continuity of the mapping, and the nearest matching pair is used to extrapolate for missing segments.
[0134] Piecewise linear mapping is used for interpolation between optimized and corrected matching point pairs. Through smoothing and constraints, the final generated mileage-pixel mapping function is ensured to be monotonic and continuous, avoiding intersections or reversals.
[0135] Output results: Generate track inspection mileage and image coordinates, network inspection mileage and image coordinates, and achieve precise alignment of multi-disciplinary data in a unified coordinate system through a precise alignment calibration function.
[0136] Layer 6, Rollback Strategy and Application Output Layer:
[0137] Rollback strategy: If the matching quality is not up to standard, a rollback strategy is initiated, such as relaxing the conditions or relying more on prior knowledge, to ensure that a usable result is output.
[0138] Application output: Supports multimodal interactive display, joint defect localization and root cause analysis, greatly improving operation and maintenance efficiency.
[0139] This invention uses long strip images formed by a line scan camera as a unified coordinate system carrier. It extracts facility feature points from one-dimensional track geometry data (track inspection) and one-dimensional catenary data (catenary inspection), respectively. Then, it detects the spatial positions of similar facilities in the images, constructing a dual-channel matching system for waveform facilities and image facilities. After matching, prior rules are used to reduce errors and fill in missing data. Finally, segmented linear expansion is performed between adjacent facility matching pairs, mapping all one-dimensional data to the unified image coordinate system, achieving precise cross-disciplinary registration.
[0140] The following section details the algorithms for identifying facility feature points from one-dimensional track geometry data, one-dimensional overhead contact line data, detection algorithms from track camera images, and the process of facility matching and sequence consistency optimization.
[0141] I. Algorithm for facility feature point identification based on one-dimensional track geometry data:
[0142] refer to Figure 4 This is a schematic diagram of the one-dimensional orbital geometry data of feature points of the facility to be extracted, according to a specific embodiment of the present invention. Figure 4 As shown, an example is presented, illustrating the one-dimensional orbital geometry data of the facility feature points to be extracted, with the spiked region representing the facility feature area.
[0143] To accurately identify and extract feature facility points, the following identification algorithm was specifically designed; this invention employs a 1D ResNet deep model structure for facility point extraction from one-dimensional track geometry data, referencing... Figure 5 This is a schematic diagram of the architecture of the 1DResNet deep model according to a specific embodiment of the present invention. The overall architecture is a phased feature extraction + classification architecture, specifically adapted to the spike-like facility features (such as turnouts, insulating joints, etc.) in track inspection data;
[0144] like Figure 5 As shown, the identification of "spiky facility points" in a one-dimensional track geometry waveform is modeled as a point event detection problem on a sequence. The input is a sequence of uniformly sampled mileage. sm∈R^(1×L); L=512~2048 (uniform mileage sampling). The output is a set of events. Each point includes mileage Confidence level Attributes such as amplitude or width are used for subsequent cross-disciplinary registration.
[0145] Labeling and Sample Construction: Positive samples are generated using known facility locations or manually labeled results. Multi-scale sliding windows are then used to extract positive examples through center alignment. Negative examples are uniformly collected from line sections far from facilities and with representative noise levels. Considering the issue of highly imbalanced class data, sampling is performed at a ratio of 1:10 to 1:50, and difficult negative examples are recorded simultaneously to optimize the model's generalization ability. During the dataset partitioning phase, stratified sampling by line or interval is used to divide the dataset into training, validation, and test sets, effectively avoiding data leakage and ensuring the accuracy of model training and evaluation.
[0146] Feature region identification employs an end-to-end 1DResNet deep model: using multi-scale feature extraction and multi-task learning strategies to achieve accurate detection and localization of spike-like facility features in one-dimensional track geometry data. Progressive feature extraction ensures both performance and computational efficiency.
[0147] Stem layer:
[0148] The Stem layer uses a one-dimensional convolutional layer as the network entry point, specifically employing a large convolutional kernel (k=7) for initial feature extraction. At the same time, the output channels are expanded to 64 to enhance feature expression capabilities. It is also equipped with a BatchNorm layer and the SiLU activation function, laying a solid foundation for feature extraction in subsequent stages while maintaining the sequence length unchanged.
[0149] Structure: Conv1D(64,k=7,s=1) (1D convolution, 64 channels, kernel size 7, stride 1) + BatchNorm (batch normalization) + SiLU (activation function).
[0150] Stage 1, Narrow spike feature extraction (small receptive field, capturing narrow spikes):
[0151] This stage is designed for narrow and sharp facility features, including two residual blocks while maintaining 64 channels. Each residual block uses parallel small convolutional kernels (k=3,5) to capture fine features, while introducing an SE attention mechanism to enhance feature selection capabilities. Residual connections ensure stable gradient backpropagation and efficient feature reuse, thereby improving the recognition accuracy of narrow and sharp features.
[0152] Residual blocks (1-1, 1-2), each with the structure: Conv1D (k=3), Conv1D (k=5), parallel convolution, simultaneously extracting multi-scale narrow features using both k=3 and k=5 Conv1D. Concat to Conv1D (1×1), feature fusion, concatenation of parallel results, and then channel compression using 1×1 convolution. Attention enhancement: Incorporating SE attention mechanism (channel attention, SE Attention C / 16 to C) to enhance effective features. Residual connections + BN + SiLU activation. The number of channels remains at 64 to accommodate fine features of narrow spikes.
[0153] Phase 2, Medium-scale feature extraction (medium receptive field, multi-scale fusion):
[0154] To balance local details with medium-range contextual information, Stage2 expands the number of channels to 96 to increase feature capacity, uses medium-scale parallel convolution kernels (k=3,7) to extract features, and introduces dilated convolution (d=1,2) to expand the receptive field; while keeping the feature map size unchanged, it effectively enhances the detection capability of medium-width facilities.
[0155] Two residual blocks (2-1, 2-2) are used, each residual block is used with parallel dilated convolution (k=3,7, dilation rate d=1,2); the advantages of dilated convolution are: expanding the receptive field without increasing the number of parameters, adapting to facility features of medium width; the number of channels is increased to 96, enhancing the feature representation capability.
[0156] Features of dilated convolution: Conv1D (k=3, d=1), Conv1D (k=7, d=2), parallel convolution kernels (k=3,7) extract features, and dilated convolution (d=1,2) is introduced; the receptive field is expanded while maintaining parameter efficiency, and it is suitable for medium-width spike features; 1×1 convolution is used to compress channels and residual connections are activated.
[0157] Stage 3, Wide Spike Feature Extraction (Large Receptive Field, Wide Spike Detection):
[0158] To address the detection requirements for wider facility features, this stage further expands the number of channels to 128, uses large convolutional kernels (k=5,9) to accurately capture wide features, and combines them with a large dilation rate (d=2,4) to obtain a wider range of contextual information. Through hierarchical design, the ability to express large-scale facility features is enhanced, adapting to the recognition scenarios of wide spike features.
[0159] Two residual blocks (3-1, 3-2) with a large receptive field design: Conv1D (k=5, d=2) and Conv1D (k=9, d=4), employing parallel convolutions with larger kernels and higher dilation rates (k=5, 9, d=2, 4); the receptive field is expanded to 20-30+ sampling points to capture wide spikes and contextual information, covering the complete features of wide spikes and their surrounding context; the number of channels is increased to 128, enhancing global feature representation.
[0160] Phase 4, Contextual Feature Aggregation:
[0161] This stage employs an ASPP (Spatial Pyramid Pooling) structure, using parallel multi-scale dilated convolutions (parallel dilated convolutions, d ∈ {1,2,4,8}) to extract features while maintaining 128 channels to preserve feature richness and global context aggregation. Through the synergistic effect of multi-scale dilated convolutions, feature fusion of the global receptive field is achieved, significantly enhancing the multi-scale expressive power of features and providing comprehensive contextual support for subsequent classification.
[0162] Category Header:
[0163] The classification head is used to predict the probability of facility locations. The intermediate layer uses a 64-channel convolutional layer to process features, and the BatchNorm layer and SiLU activation function are used to ensure feature quality. Finally, a single-channel probability map is output through 1×1 convolution, and the output result is constrained to the range [0,1] using the Sigmoid activation function to accurately reflect the probability of facility locations.
[0164] The structure is: Conv1D(128,k=3), BN, SiLU, Conv1D(1,k=1), Sigmoid activation; the final output is a binary classification result (whether it is a facility point or not).
[0165] Loss function:
[0166] To address the challenges of sparse facility point distribution and severe imbalance between positive and negative samples in one-dimensional orbital geometry data, an improved FocalLoss loss function is employed as the classification loss function. This loss function dynamically adjusts sample weights, adaptively focusing on difficult samples, thus effectively improving detection performance.
[0167] ;
[0168] in, y is the loss function; p̂ is the probability of the facility point predicted by the model; y is the true label (0 or 1); α is the positive sample weighting factor; γ is the focusing parameter.
[0169] Through the above design, the weight ratio of positive and negative samples will be dynamically adjusted during the model training process. This loss function can effectively solve the problems of class imbalance and hard sample detection faced by facility point detection in one-dimensional track geometry data, and significantly improve the detection performance and generalization ability of the model.
[0170] II. Algorithm for identifying facility feature points in one-dimensional catenary data:
[0171] refer to Figure 6 This is a schematic diagram of one-dimensional catenary data for extracting feature points of a facility according to a specific embodiment of the present invention. Figure 6 As shown, unlike the spike feature points in the one-dimensional data of track geometry, the location of the overhead contact system facilities in the one-dimensional data is usually at the turning point. Therefore, a separate facility feature point recognition algorithm was designed for the one-dimensional data of the overhead contact system.
[0172] refer to Figure 7 This is a schematic diagram of a data processing flow according to a specific embodiment of the present invention. Figure 7 As shown, the specific process includes:
[0173] Phase 1, Data Preprocessing:
[0174] The original one-dimensional catenary data is standardized, including mileage-equidistant resampling, outlier detection, and masking. Baseline drift is estimated using rolling median or Huber spline methods, and residual sequences are constructed to highlight facility features. Savitzky-Golay filtering is used to obtain the first and second derivatives of the zero-phase data, and the MAD noise scale within the sliding window is calculated, providing a robust data foundation for subsequent feature extraction.
[0175] Mileage resampling: Equidistant sampling (interval 0.1~0.2m), detecting and removing outliers, and converting the time dimension to the mileage dimension.
[0176] Baseline estimation: Fit the baseline using rolling median or Huber spline, calculate the residuals between the data and the baseline, and correct for data drift.
[0177] Filtering and Derivatives: Smooth the data using Savitzky-Golay filtering and calculate the first and second derivatives (reflecting the trend of data changes).
[0178] Noise estimation: The local noise scale is obtained by calculating the MAD (median absolute deviation) using a sliding window, and the SNR (signal-to-noise ratio) of the data is evaluated.
[0179] Phase 2, Inflection Point Candidate Generation:
[0180] Potential inflection point locations are identified based on the zero-crossing of the first derivative and the extrema of the second derivative. Multi-scale consistency verification is employed, repeating the detection process at different window scales to retain candidate points that appear simultaneously at multiple scales, thus improving robustness. A symmetrical window is extracted as the sample region centered on each candidate inflection point, establishing a data foundation for feature computation and classification.
[0181] Zero-crossing detection: Find zero-crossing points of the first derivative and points with the largest second derivatives within the previous preset percentage, and extract inflection point candidates;
[0182] Multi-scale validation: The consistency of candidate points is validated using 22 scales ranging from 0.5 to 2.0 m to improve the robustness of the results;
[0183] Candidate window: Centered on the inflection point, extract a symmetrical window, define the sample region, and determine the analysis range of the candidate point.
[0184] Phase 3, Traditional Feature Engineering:
[0185] A comprehensive feature system incorporating time-domain, frequency-domain, morphological, and prior knowledge is constructed. Time-domain features include peak and valley amplitudes, statistical moments, energy, and symmetry indices; derivative features include maximum slope, left-right slope ratio, and curvature surrogate; morphological features extract full width at half maximum (FWHM), baseline crossing width, and three-angle similarity; frequency-domain and wavelet features capture spectral characteristics and multi-scale energy distribution. Combined with engineering prior knowledge, such as facility spacing and pitch constraints, 40-80 dimensional feature vectors are constructed.
[0186] Feature extraction from 5 dimensions:
[0187] Temporal and geometric characteristics: peak and trough amplitude, mean, variance, length of left and right monotonic segments, symmetry index, and relative height;
[0188] Derivatives, curvature characteristics: maximum slope, left-right slope ratio, curvature proxy, etc., angle between fitted line segments, area of curvature;
[0189] Morphological characteristics: full width at half maximum (FWHM), baseline crossing width three-angle similarity, V-shaped template correlation coefficient, periodic repeatability;
[0190] Frequency domain and wavelet features: bandpass energy ratio, spectral centroid and bandwidth, wavelet coefficient energy, L1 / L2 ratio, and multi-scale features;
[0191] Prior and contextual features: facility spacing, pitch prior range, local MAD, SNR, outlier ratio, and engineering constraint features.
[0192] Phase 4, Sample Construction and XGBoost Training:
[0193] Regarding sample labeling and balancing:
[0194] A positive sample set is constructed based on manual annotation or engineering records. Waveform segments far from the facility points are selected as negative samples, with special attention paid to the collection of difficult negative examples. To address the problem of severe imbalance between positive and negative samples, a sampling ratio of 1:10 to 1:30 is adopted. Through sample weight adjustment and difficult sample mining strategies, it is ensured that the model can effectively learn facility feature patterns.
[0195] XGBoost model training:
[0196] Gradient boosting decision trees are used for binary classification learning, and 5-fold hierarchical cross-validation and early stopping mechanisms are employed to prevent overfitting. Key hyperparameters include tree depth of 4-6, learning rate of 0.03-0.1, and feature-to-sample ratio of 0.7-0.9. The `scale_pos_weight` parameter is used to handle class imbalance, and PR-AUC and Recall@FP / km are used as the main evaluation metrics. Feature set is optimized through feature importance analysis.
[0197] Evaluation and optimization: Probabilistic calibration is performed using PR-AUC as the core indicator, combined with evaluations such as Recall@FP / km.
[0198] Phase 5, Reasoning and Post-processing:
[0199] Sliding Window Reasoning:
[0200] Candidate points are generated by sliding across the complete dataset, features are extracted, and facility probabilities are predicted using a trained XGBoost model. Probability thresholding is employed to filter candidate points, merging detections that are too close together into a single facility region. Region boundaries are precisely located using derivative zero-crossing and curvature peaks, and unreasonable isolated detection points are eliminated based on engineering prior constraints.
[0201] III. Algorithm for detecting images from line cameras:
[0202] To address the need for feature detection in linear scan camera images (long strip images), an improved RT-DETR algorithm is adopted as the core detection algorithm. This scheme achieves simultaneous detection and localization of multiple target types through an end-to-end Transformer architecture, avoiding the complexity of traditional NMS post-processing, and is particularly suitable for dense small target scenes.
[0203] Data preprocessing and augmentation strategies:
[0204] The original grayscale images from the linear scan camera are standardized. Image quality is improved through preprocessing techniques such as single-channel to three-channel expansion, histogram equalization, and reflection suppression. Data augmentation strategies encompass geometric transformations, illumination variations, small target copying and pasting, and dirt occlusion simulation, comprehensively enhancing the model's adaptability to complex environments.
[0205] Network architecture improvement design:
[0206] refer to Figure 8 This is a schematic diagram of the architecture of an improved RT-DETR model according to a specific embodiment of the present invention. Figure 8 As shown, targeted improvements are made to the RT-DETR architecture, employing a lightweight backbone feature extraction network (ResNet-50, ConvNeXt-T) combined with multi-scale feature fusion to enhance small target detection capabilities. The HybridEncoder, combined with a deformable attention mechanism, effectively aggregates feature information at different scales. The Transformer Decoder uses a 6-layer decoder structure, achieving joint optimization of classification and regression tasks through the Hungarian matching algorithm and a decoupled detection head. Considering the characteristics of railway infrastructure, the number of queries is adjusted to 300-400, and relative position encoding is optimized to accommodate slender targets. The detection results are output, including: bounding box [x1, y1, x2, y2], confidence score (conf), class_id, and mileage mapping.
[0207] Training strategy and loss function:
[0208] A multi-task joint training strategy is adopted. FocalLoss is used for classification loss to handle class imbalance, and L1 and GIoU are combined for regression loss to achieve accurate localization. The formula is shown below:
[0209]
[0210]
[0211]
[0212] in, Indicates the total loss; FocalLoss is the classification loss; This is a combination of L1 loss and GIoU loss; For L1 loss; For GIoU loss; For classification loss weights; For regression loss weights; y represents the probability of the facility location predicted by the model; y is the true label (0 or 1). For positive samples, the weighting factor is used. For focusing parameters; b, represents the coordinates of the ground truth bounding box and the predicted bounding box, respectively; C is the area of the minimum bounding rectangle. is the intersection-union ratio of the ground truth bounding box and the predicted bounding box; A and B are the areas of the ground truth bounding box and the predicted bounding box, respectively.
[0213] Training stability is ensured through AdamW optimizer and Cosine annealing learning rate scheduling, combined with mixed-precision training and gradient pruning. EMA parameter tracking and early stopping mechanisms are implemented to prevent overfitting and improve model generalization ability.
[0214] Reasoning and post-processing mechanisms:
[0215] An efficient sliding window inference process is designed to intelligently merge detection results in overlapping areas, avoiding duplicate detections. Through a class-adaptive threshold strategy and coordinate write-back mechanism, detection results are converted into cumulative pixel coordinates, and accurate pixel-to-mileage mapping is achieved based on imaging calibration. Combined with confidence filtering and geometric constraints, false detections are effectively suppressed.
[0216] IV. Facility Matching and Sequence Consistency Optimization:
[0217] A facility matching algorithm based on multi-constraint optimization is established to achieve precise correspondence between image coordinates and mileage coordinates. A complete technical route from coarse registration to fine optimization is constructed through classifier candidate construction, multi-dimensional scoring mechanism, robust mapping fitting, and sequence consistency optimization. This scheme specifically addresses engineering problems such as detection latency, local missed detections and false detections, and interference from dense facilities, ensuring the monotonicity, stability, and high coverage of the final mapping. (Reference) Figure 9 This is a schematic diagram illustrating the facility matching and sequence consistency optimization process according to a specific embodiment of the present invention. Figure 9 As shown, the specific process includes:
[0218] S901, Classification of Candidates Construction and Gating:
[0219] The facility detection results from both the image and one-dimensional data sides are grouped by category to establish a candidate matching pool for facilities of the same type. By employing mileage prior constraints and a neighborhood gating mechanism, obviously unreasonable candidate pairs are pre-emptively eliminated, effectively compressing the search space. Mileage priors utilize coarse mapping or ranking gating to limit the candidate range, while neighborhood gating ensures the reasonableness of the spacing between adjacent facilities, laying the foundation for subsequent accurate matching.
[0220] The input data includes image facility data, one-dimensional facility data, and coarse mapping prior data.
[0221] S902, multi-dimensional candidate scores:
[0222] A comprehensive scoring function is constructed, incorporating proximity, order relation, contextual consistency, and confidence weighting. Proximity scoring is based on a Gaussian kernel function of coordinate distance; order relation scoring considers the consistency of relative ranking; and contextual consistency establishes spatial constraints through the distance relationship between left and right neighbors. Combining a weighting mechanism for detection confidence and a penalty term for cross-risk, a comprehensive matching quality metric is formed.
[0223] S903, Intra-class Optimal Matching and Preliminary Screening:
[0224] For each facility category, the Hungarian algorithm is used to solve the maximum weighted matching problem, obtaining the optimal one-to-one correspondence within the category. High-confidence matches are retained through soft thresholding, while low-confidence candidates are left for subsequent global optimization. This step ensures optimal matching within the same facility category and provides a reliable foundation for global consistency optimization across categories.
[0225] S904, Robust Mapping Fit:
[0226] The RANSAC algorithm is used for robust mileage-to-pixel mapping fitting, supporting both global affine and piecewise affine models. Iterative sampling and interior point determination effectively suppress the influence of outliers.
[0227] S905, Sequence Consistency Optimization:
[0228] The Longest Increasing Subsequence (LIS) algorithm is used to eliminate cross-matches, ensuring strict monotonicity of the mapping. Dynamic Time Warping (DTW) is applied within adjacent anchor intervals for local sequence alignment, supplementing reliable matching points. Through iterative cross-match elimination and fine-tuning mechanisms, global sequence consistency optimization is achieved, ensuring the stability and continuity of the final mapping.
[0229] S906, Global Smoothing and Lay-out:
[0230] A secondary mapping fit is performed based on the optimized anchor point sequence. A piecewise linear unfolding strategy is implemented, performing linear interpolation between adjacent anchor points to ensure local linearity of the mapping. Let the matching pair be (m i ,x i ), m i Let x be the mileage of the i-th known point. i For the i-th facility image pixel, in the interval [m i ,m i+1 Linear interpolation is used within the interval to achieve "uniform spreading" and obtain the pixels corresponding to the data points within the interval.
[0231] When performing piecewise linear unfolding, the calculation formula used is:
[0232]
[0233] In the formula, f(m) is the output result at the input point, representing the image pixel coordinates of the facility point's mileage m; x i Let x represent the image pixel coordinates of the i-th known point. i With m i The coordinates of the i-th known point (m) i ,x i );m i The distance to the i-th known point is represented by m; the distance to the facility point is represented by x. i+1 m i+1 Let represent the image pixel coordinates and mileage of the (i+1)th known point, respectively.
[0234] The output includes the mileage-to-pixel mapping function f(m), the confidence anchor sequence, and the unified coordinate system registration basis.
[0235] The present invention provides a method for unified registration of multi-disciplinary data for high-speed railways based on visual features through a specific embodiment.
[0236] Sub-step 1: Data Acquisition and Preprocessing
[0237] One-dimensional track geometry data: equidistant resampling (Δm=0.05~0.20m) with mileage as the independent variable, using rolling median or high-pass detrending and zero-mean unit variance standardization; masks are created for saturated and missing segments, which are not included in subsequent statistics.
[0238] One-dimensional data of overhead contact system: amplitude normalization and zero-phase smoothing for noise reduction, output residual sequence and first or second derivative, used to construct support candidates and features.
[0239] Linear camera images: distortion correction, banner stitching and illumination equalization (CLAHE, reflection suppression) are completed, and the images are cropped to a fixed size along the direction of travel using a sliding window, while preserving the original image coordinate mapping.
[0240] Sub-step two, track geometry facility identification (1D ResNet):
[0241] Model and input: End-to-end 1D ResNet (multi-scale convolution + residual structure), with input being equidistant mileage sequence segments of length L (1024 or 2048).
[0242] Output and Decision: Dense prediction yields facility probability, subpixel offset, and peak width estimation; thresholding + 1DNMS is used to obtain anchor points for typical facilities such as turnouts. The output set M = {m, cls, conf}, where m is the mileage of the facility point, cls is the facility category label, and conf is the detection confidence.
[0243] Results Presentation: Reference Figure 10This is a schematic diagram of a turnout facility for one-dimensional data recognition of track geometry according to a specific embodiment of the present invention. Figure 10 As shown in the figure, the location of the turnout and the mileage markings are indicated.
[0244] Sub-step 3, Contact Line 1C Support Recognition (Traditional Features + XGBoost):
[0245] Candidates and Features: Candidates are nominated based on derivative zero crossover or curvature extrema. A window is constructed around the candidates to extract 40-80 dimensions of features, including time-domain statistics, left-right slope ratio, curvature energy, full width at half maximum (FWHM), bandpass energy ratio, and local SNR.
[0246] Classification and Regionalization: Train an XGBoost binary classifier to output pillar probabilities, then threshold them and merge adjacent candidates and refine the boundaries (based on derivative zero crossover and curvature peaks), outputting a set M={m,cls,conf}.
[0247] Results Presentation: Reference Figure 11 This is a schematic diagram of a support pillar for one-dimensional data recognition of overhead contact lines according to a specific embodiment of the present invention. Figure 11 As shown in the figure, the center of the support column and the range of the interval are displayed.
[0248] Sub-step four, image facility detection (improved RT-DETR):
[0249] Model and Categories: An improved RT-DETR (lightweight Backbone + Hybrid Encoder + Decoupled Detection Head) is used, covering categories such as "turnouts and supports"; end-to-end detection is performed on the images.
[0250] Inference and Writeback: The recognition results are used to calculate the cumulative coordinates based on the image number, resulting in the pixel anchor point set X. img ={x,cls,conf}; x is the accumulated pixels along the direction of travel, cls is the facility category label, and conf is the detection confidence.
[0251] Results show that the locations of turnout facilities and supports can be identified through images.
[0252] Sub-step five, facility matching and sequence optimization (cross-modal alignment):
[0253] Candidate construction by category: grouping by category; introducing a coarse mapping prior g(m) = a o m+b o Or, order gating limits the search range for candidate pairs. m represents mileage; g(m) represents the pixel coordinates of the corresponding mileage m in the image (or the registered target coordinates, output result); a o The slope of the mapping reflects the proportional relationship between the change in mileage and the change in pixels (e.g., a). o=2 indicates that the pixel coordinate increases by 2 units for every 1m increase in mileage); b o The intercept of the mapping is the pixel coordinate reference value corresponding to the mileage m=0.
[0254] Multidimensional scoring and intra-class matching: Candidate pairs are scored based on proximity, order relation (relative order), context consistency (left and right spacing ratio), and confidence weighting; the Hungarian algorithm is used to find the maximum weighted intra-class match to obtain the initial pair set.
[0255] Robust mapping fitting: RANSAC is used to fit the global or piecewise affine interior point weighted least squares refinement, and order-preserving regression or monotonic splines are used to refine the results to ensure monotonicity.
[0256] Sequence consistency optimization: Perform LIS to remove cross-matches for all pairs, then use DTW for local time alignment within adjacent anchor intervals to fill in the reliable points; perform LIS verification again to eliminate jump points and minor violations.
[0257] Output mapping: Obtain a monotonic and stable mapping function f(m) from mileage to pixel and its anchor point sequence; pixels within the interval are laid out linearly in segments.
[0258] The main improvements of this invention are as follows:
[0259] Using long strip images from a line scan camera as a unified coordinate system, the three data streams (track inspection, network inspection 1C, and image) are aligned together to achieve correlation analysis of track inspection and network inspection data.
[0260] Three-way facility identification links: 1D ResNet identifies typical track inspection facilities; "Feature Engineering + XGBoost" identifies network inspection 1C pillars; improved RT DETR detection image facilities and write-back accumulated pixels.
[0261] Candidate pairs are constructed by category, and the optimal match within each category is obtained by using a multi-dimensional score based on proximity, order relation, context consistency and confidence.
[0262] Robust mapping fitting: RANSAC mileage estimation → pixel affine / piecewise affine, interior point weighted least squares refinement; order-preserving regression / monotonic splines ensure monotonicity.
[0263] Sequence optimization: LIS decrossing and constraint DTW local alignment are combined to remove jump points and replenish reliable points.
[0264] Prior and mutual verification: Facility pairing / pitch / topology priors drive false alarm / missed alarm filtering and completion, and image and one-dimensional results mutual verification improves confidence.
[0265] The piecewise linear unfolding generates a monotonic mapping function f(m) and an interval pixel grid.
[0266] In practical applications, this invention can be used in high-speed rail, conventional rail and subway sections, and can also be extended to other facilities.
[0267] The method and apparatus for unified registration of multi-disciplinary data in high-speed railways based on visual features proposed in this invention have at least the following beneficial effects:
[0268] Provide a unified coordinate system: using linear array images as a common carrier, output a monotonic, segmentable "mileage → pixel" mapping function f(m) to achieve high-precision alignment and linked display of track inspection, network inspection and images.
[0269] Strong robust registration capability: Multidimensional scoring + Hungarian matching, RANSAC + order-preserving splines, LIS + DTW sequence optimization effectively suppress delay, velocity fluctuations and local non-monotonic distortion, avoiding mismatch and local optima.
[0270] Reduce reliance on sparse anchor points: By using visual facilities as dense anchor points, the reliance on GNSS / transponders / artificial landmarks is reduced, and stable alignment can still be achieved in environments such as tunnels and obstructions.
[0271] False alarm / missed alarm self-check: Integrating facility pairing / pitch / topology priors with multi-source mutual verification, automatically filtering out erroneous matches and filling in missing points, improving the confidence of anchor sequence.
[0272] Small target detection optimization: Improved RT-DETR for long strips of small and dense targets, with inference result coordinates written back, improving facility detection rate and positioning consistency.
[0273] Explainable and traceable: Anchor points, residuals, and scores are retained throughout the entire process, facilitating auditing, review, and operational decisions.
[0274] Easy to expand: The solution is open to data sources and categories, and can be seamlessly connected to new professional data and more facility types to continuously improve system capabilities.
[0275] After introducing the method of exemplary embodiments of the present invention, the following references are made. Figure 12 This paper introduces a visual feature-based unified registration device for multi-disciplinary data in high-speed railways, based on an exemplary embodiment of the present invention.
[0276] The implementation of the high-speed railway multi-disciplinary data unified registration device based on visual features can refer to the implementation of the above-described method, and repeated details will not be elaborated upon. The term "module" or "unit" used below can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0277] Based on the same inventive concept, this invention also proposes a unified registration device for multi-disciplinary data in high-speed railways based on visual features, such as... Figure 12 As shown, the device includes:
[0278] The data acquisition module 1210 is used to acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and line scan camera image data, and to preprocess the multi-source data.
[0279] The facility feature point advance module 1220 is used to extract track inspection facility feature points and catenary inspection facility feature points from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively, and output the facility category and mileage location after recognition processing; it detects the facility feature points in the image data from the preprocessed linear array camera image data, outputs the facility category and image pixel coordinates after detection, and constructs a unified coordinate system based on the image pixel coordinates;
[0280] The cross-modal matching module 1230 is used to group facilities according to facility type, match feature points of track inspection facilities with image facility points, and network inspection facility points with image facility points to obtain matching results; based on the matching results, mileage location and image pixels, the random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency.
[0281] The verification module 1240 uses a prior knowledge base to perform consistency verification on the matching results, and fills in missing points and removes erroneous points through a multi-source mutual verification mechanism to obtain verified matching pairs;
[0282] The registration module 1250 is used to perform piecewise linear expansion between adjacent matching pairs, generate a monotonic mileage-to-image pixel mapping function based on the mileage-to-image pixel mapping relationship, and map the multi-source data to the unified coordinate system to obtain the registration result.
[0283] It should be noted that although several modules of the high-speed railway multi-disciplinary data unified registration device based on visual features have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0284] Based on the aforementioned inventive concept, such as Figure 13As shown, the present invention also proposes a computer device 1300, including a memory 1310, a processor 1320, and a computer program 1330 stored in the memory 1310 and executable on the processor 1320. When the processor 1320 executes the computer program 1330, it implements the aforementioned unified registration method for multi-disciplinary data of high-speed railway based on visual features.
[0285] Based on the aforementioned inventive concept, this invention proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for unified registration of multi-disciplinary data for high-speed railways based on visual features.
[0286] Based on the aforementioned inventive concept, this invention proposes a computer program product, which includes a computer program that, when executed by a processor, implements a unified registration method for multi-disciplinary data of high-speed railway based on visual features.
[0287] This invention proposes a unified registration method and device for multi-disciplinary data in high-speed railways based on visual features. This constructs an automated and robust registration process using visual features as a common benchmark, fundamentally solving the problem of difficulty in unifying and aligning multi-source detection data in high-speed railways caused by system delays, dynamic distortion, and sparse anchor point dependence. This invention accurately extracts facility feature points from one-dimensional track inspection and network inspection data and linear array images, respectively, and achieves precise cross-modal association between waveform facilities and image facilities based on classification matching and random sample consistency algorithms. Furthermore, by introducing a multi-source mutual verification mechanism driven by sequence consistency optimization and prior knowledge base, it effectively suppresses mismatches and missed matches, ensuring the global monotonicity and stability of the mapping from mileage to image pixels. With a piecewise linear unfolding strategy, it generates a continuous and high-precision unified coordinate mapping function, seamlessly aligning previously isolated multi-disciplinary data to the same visual coordinate system. This provides a strong data foundation for joint defect location, root cause analysis, and trend judgment, significantly improving the accuracy and efficiency of railway infrastructure operation and maintenance analysis.
[0288] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.
[0289] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0290] This invention is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0291] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0292] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0293] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for unified registration of multi-disciplinary data in high-speed railways based on visual features, characterized in that, The method includes: Acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and preprocess the multi-source data; Feature points of track inspection facilities and feature points of catenary inspection facilities are extracted from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively. After identification and processing, the facility category and mileage location are output. Image facility feature points are detected from preprocessed linear scan camera image data. After detection, facility category and image pixel coordinates are output. A unified coordinate system is constructed based on the image pixel coordinates. Grouped by facility type, the feature points of track inspection facilities are matched with image facility points, and the network inspection facility points are matched with image facility points to obtain matching results; Based on the matching results, mileage location, and image pixels, a random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency. The consistency of the matching results is checked using a prior knowledge base. Missing points are filled and erroneous points are removed through a multi-source mutual verification mechanism to obtain the verified matching pairs. Piecewise linear tiling is performed between adjacent matching pairs. A monotonic mileage-to-image pixel mapping function is generated based on the mileage-to-image pixel mapping relationship. The multi-source data is then mapped to the unified coordinate system to obtain the registration result.
2. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Preprocessing the multi-source data includes: The one-dimensional track geometry data is denoised by digital filtering, bandpass filtering, or wavelet transform, and abnormal segments are verified in conjunction with odometer data. The one-dimensional data of the overhead contact system is subjected to amplitude normalization, baseline drift correction, and robust filtering. The image data from the line scan camera is subjected to distortion correction, illumination equalization, and image stitching processing.
3. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Feature points of track inspection facilities and catenary inspection facilities are extracted from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively. After identification processing, the facility category and mileage location are output, including: Feature points of track inspection facilities are extracted from the preprocessed one-dimensional track geometry data, and an end-to-end 1DResNet depth model is used for identification, outputting facility category, mileage location and confidence level; The 1DResNet deep model includes a Stem layer, multiple residual stages, and a classification head. The Stem layer is a one-dimensional convolutional layer equipped with a BatchNorm layer and the SiLU activation function; The residual stage includes narrow spike feature extraction, preset scale feature extraction, wide spike feature extraction, and contextual feature aggregation, which use parallel convolution kernels and attention mechanisms, respectively. The classification head outputs the probability of facility point existence, and FocalLoss is used as the loss function.
4. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Feature points of track inspection facilities and catenary inspection facilities are extracted from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively. After identification processing, the facility category and mileage location are output, including: Feature points of the overhead contact system are extracted from the preprocessed one-dimensional data of the overhead contact system. The identification is performed based on feature engineering and XGBoost classifier, and the facility category, mileage location and confidence score are output. The feature engineering process generates candidate inflection points based on the zero crossover of the first derivative and the extrema of the second derivative, and extracts time-domain, frequency-domain, morphological, and prior features. The XGBoost classifier is trained using the XGBoost model for binary classification, and the evaluation metrics include at least PR-AUC and recall.
5. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Image facility feature points are detected from preprocessed linear scan camera image data. After detection, the facility category and image pixel coordinates are output. A unified coordinate system is constructed based on the image pixel coordinates, including: Facility feature points are detected from preprocessed linear scan camera image data. An improved RT-DETR model is used for end-to-end detection, and the facility category, image pixel coordinates, and confidence score are output. The improved RT-DETR model includes: a Backbone feature extraction network, a hybrid editor, and a decoder; The improved RT-DETR model employs a Backbone feature extraction network combined with multi-scale feature fusion; a hybrid editor combined with a deformable attention mechanism aggregates feature information at different scales; the decoder uses a matching algorithm and a decoupled detection head to jointly optimize classification and regression tasks; the training strategy of the improved RT-DETR model uses a multi-task loss function and adopts the AdamW optimizer and Cosine annealing learning rate scheduling.
6. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Grouped by facility type, the feature points of track inspection facilities are matched with image facility points, and the feature points of network inspection facilities are matched with image facility points to obtain matching results, including: A cross-modal matching strategy is constructed, in which candidate pairs are grouped by facility category and selected based on mileage prior constraints and neighborhood gating mechanism. The candidate pairs are scored according to multidimensional scoring, including neighboring Gaussian kernel, order relation consistency, context distance ratio and confidence weighting. Candidate pairs that meet the preset confidence requirements are selected based on the score and soft threshold.
7. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, The matching results are checked for consistency using a prior knowledge base. A multi-source mutual verification mechanism is used to fill in missing points and remove erroneous points, resulting in verified matching pairs, including: The prior knowledge base includes facility pairing rules, typical spacing ranges, and relative topological order; The multi-source mutual verification mechanism includes cross-verification using the matching results of track inspection and network inspection.
8. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, Piecewise linear tiling is performed between adjacent matching pairs. A monotonic mileage-to-image-pixel mapping function is generated based on the mileage-to-image-pixel mapping relationship. The multi-source data is then mapped to the unified coordinate system to obtain the registration result, including: When performing piecewise linear unfolding, the calculation formula used is: In the formula, To output the result at the input point, the image pixel coordinates representing the mileage m of the facility point; x i Let x represent the image pixel coordinates of the i-th known point. i With m i The coordinates of the i-th known point are formed ;m i The distance to the i-th known point is represented by m; the distance to the facility point is represented by x. i+1 m i+1 Let represent the image pixel coordinates and mileage of the (i+1)th known point, respectively.
9. The method for unified registration of multi-disciplinary data in high-speed railways based on visual features according to claim 1, characterized in that, The method also includes: A fallback strategy is constructed, in which, when the quality index of the registration result does not meet the standard, the registration effect is optimized by relaxing the candidate threshold, increasing the prior weight, or switching to sparse anchor point alignment; when the registration quality index does not meet the standard, the optimization is carried out by adjusting the candidate threshold, prior weight, or switching to sparse anchor point alignment.
10. A unified registration device for multi-disciplinary data in high-speed railways based on visual features, characterized in that, The device includes: The data acquisition module is used to acquire multi-source data, including one-dimensional track geometry data, one-dimensional catenary data, and linear array camera image data, and to preprocess the multi-source data. The facility feature point advance module is used to extract track inspection facility feature points and catenary inspection facility feature points from the preprocessed one-dimensional track geometry data and one-dimensional catenary data, respectively, and output the facility category and mileage location after identification processing; it detects facility feature points from the preprocessed linear scan camera image data, outputs the facility category and image pixel coordinates after detection, and constructs a unified coordinate system based on the image pixel coordinates; The cross-modal matching module is used to group facilities according to facility type and match feature points of track inspection facilities with image facility points and network inspection facility points with image facility points to obtain matching results. Based on the matching results, mileage location and image pixels, the random sample consensus algorithm is used to fit the mapping relationship from mileage to image pixels, and the monotonicity of the mapping relationship is optimized by sequence consistency. The verification module uses a prior knowledge base to perform consistency checks on the matching results, and fills in missing points and removes erroneous points through a multi-source mutual verification mechanism to obtain verified matching pairs. The registration module is used to perform piecewise linear expansion between adjacent matching pairs, generate a monotonic mileage-to-image pixel mapping function based on the mileage-to-image pixel mapping relationship, and map the multi-source data to the unified coordinate system to obtain the registration result.
11. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 9.