Method and device for processing high-speed railway infrastructure detection data
By using dynamic feature modeling and a unified intermediate model (UDM), the problems of insufficient mileage calibration accuracy and difficulty in multi-source data fusion in high-speed railway infrastructure inspection are solved, achieving high-precision calibration and accurate analysis of data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA RAILWAY SIYUAN SURVEY & DESIGN GRP CO LTD
- Filing Date
- 2025-05-24
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies for high-speed railway infrastructure inspection suffer from insufficient mileage calibration accuracy, difficulty in multi-source data fusion, and limited data processing methods, resulting in inadequate data quality and analytical accuracy.
A nonlinear mileage calibration method based on dynamic feature modeling is adopted, which combines a hierarchical optimization strategy to dynamically calibrate mileage offsets. Data calibration is performed using a Gaussian mixture model and a hierarchical optimization algorithm. Data cleaning is performed using a joint weighted interpolation method based on spatiotemporal proximity and an adaptive kernel density estimation outlier detection method. A unified intermediate model (UDM) is constructed to achieve semantic alignment and format conversion of heterogeneous data.
It improves the accuracy of mileage calibration, enhances the adaptability of multi-source data fusion, ensures data quality and analysis accuracy, solves the adaptability limitations of traditional methods in complex orbital scenarios, and realizes the spatiotemporal correlation and noise robustness of data.
Smart Images

Figure CN120597198B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of high-speed railway data processing, and in particular relates to a method and apparatus for processing high-speed railway infrastructure detection data. Background Technology
[0002] With the continuous development of technology, the requirements for traffic efficiency have become increasingly higher, leading to the emergence of high-speed railways. High-speed railways, or HSR for short, refer to railway systems designed to high standards and capable of allowing trains to travel safely at high speeds.
[0003] In the service condition inspection of high-speed railway infrastructure, accurately grasping the condition of facilities such as tracks, bridges, and tunnels is crucial to ensuring train operation safety. With the advancement of intelligent inspection technology, integrated inspection equipment such as track inspection vehicles and inspection robots utilize multi-source devices such as sensors, images, and videos to collect massive amounts of data in real time, including geometric parameters, vibration signals, and images of defects.
[0004] In high-speed railway infrastructure inspection, mileage calibration, data governance, and data fusion are three key aspects to ensure data quality and analytical accuracy. Mileage calibration, through precise coordinate alignment, ensures that data collected by different inspection devices remains consistent both spatially and temporally. Data governance, by standardizing data collection, storage, and processing procedures, ensures data accuracy and consistency. Data standardization, through unified data models and semantic definitions, reduces data format and semantic differences between different inspection devices, facilitating the fusion and analysis of multi-source data. These three aspects are interdependent and work together to improve the usability of inspection data and the accuracy of analysis. Summary of the Invention
[0005] In view of this, this application provides a method and apparatus for processing high-speed railway infrastructure inspection data, which aims to correct the impact of mileage error on multi-source data fusion.
[0006] Firstly, this application provides a method for processing high-speed railway infrastructure inspection data, including:
[0007] Acquire detection data, which includes structured data and unstructured data;
[0008] A dynamic functional feature modeling alignment method is used to perform mileage calibration on structured data;
[0009] Perform data cleaning on the test data;
[0010] Construct an intermediate data model to store structured and unstructured data in a mixed manner, and perform mileage correction on the unstructured data through the intermediate data model;
[0011] A data quality scoring model is constructed to score the detection data; when the score of the detection data meets the requirements, the detection data is stored through an intermediate data model.
[0012] Optionally, the detection data includes structured data and unstructured data. The structured data includes one or more of curvature, track orientation, and track gauge, while the unstructured data includes one or more of vibration signals, images, and videos.
[0013] Optionally, the steps for mileage calibration of structured data using a dynamic functional feature modeling alignment method include:
[0014] Structured data is extracted from the detection data, and extreme points of the structured data are marked;
[0015] The line is divided into straight segments and curved segments based on the curvature in the structured data. Initial feature diffusion is set for the straight segments and curved segments respectively. Gaussian mixture model is constructed based on the extreme points and the initial feature diffusion.
[0016] Initialize the dynamic parameters and, based on the segment type of the line, initialize the weight coefficients in the Gaussian mixture model.
[0017] The initial feature diffusion in the Gaussian mixture model is determined based on the accuracy calibration of the detection equipment.
[0018] Based on the Gaussian mixture model and odometer calibration values, a multi-objective optimization model is constructed to minimize the difference between the theoretical feature function and the actual detection data.
[0019] A hierarchical optimization calibration method is used to solve for the mileage calibration values in the multi-objective optimization model;
[0020] The structured data is calibrated based on the mileage calibration values obtained from the solution.
[0021] Alternatively, the Gaussian mixture model is as follows:
[0022]
[0023] In the formula, Indicates the reference data at the mileage position The first The values of feature points in structured data, This indicates the first [number] determined based on design drawings or historical benchmark data. The theoretical mileage of feature points in structured data, Indicates the first The diffusion of feature points in structured data These are the weight coefficients in a Gaussian mixture model, used to represent the weights of the first Gaussian mixture model. The weights of feature points in structured data.
[0024] Alternatively, the multi-objective optimization model is as follows:
[0025]
[0026] In the formula, This represents the matching error term. Indicates the position at the offset mileage. The first The values of feature points in structured data, This indicates the measured data at the mileage location. The first The values of feature points in structured data, Represents the regularization term, used to constrain... The magnitude of the mutation, This represents the smoothness constraint coefficient for mileage offset. This represents a feature stabilization term used to suppress the influence of noise-sensitive features on mileage offset. This indicates the mileage offset.
[0027] Optionally, the steps for solving the mileage calibration values in the multi-objective optimization model using hierarchical optimization calibration include:
[0028] The mileage calibration value is solved iteratively using the Levenberg-Marquardt optimization algorithm, which is as follows:
[0029]
[0030] In the formula, Indicates the mileage calibration value. Represents the Jacobian matrix. Indicates the damping factor. This represents the set of feature points of structured data. Let represent the residual vector, where .
[0031] Optionally, the steps of mileage correction for unstructured data using an intermediate data model include:
[0032] Determine the error between theoretical and actual mileage for unstructured data. ;
[0033] According to the error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. ;
[0034] The relative displacement increment is determined based on the instantaneous velocity and instantaneous acceleration of the detection equipment. ;
[0035] Based on preset track feature points, determine the feedback correction mileage. ;
[0036] Based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data .
[0037] Optionally, determine the error between the theoretical mileage and the actual mileage of the unstructured data. The formula is as follows:
[0038]
[0039] In the formula, This indicates the theoretical mileage of the testing equipment. Indicates the actual mileage of the testing equipment.
[0040] Indicates the first Fixed deviations caused by the sliding or sampling errors of the detection equipment within each detection section. This represents the dynamic error that accumulates over time.
[0041] Optionally, based on the error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. The formula is as follows:
[0042]
[0043] In the formula, Indicates time The BeiDou absolute coordinates of the time detection equipment This indicates the geographical coordinates corresponding to the mileage shown on the design drawings.
[0044] Optionally, the relative displacement increment is determined based on the instantaneous velocity and instantaneous acceleration of the detection equipment. The formula is as follows:
[0045]
[0046] In the formula, Indicates the instantaneous speed of the detection equipment. This indicates the instantaneous acceleration of the detection equipment.
[0047] Optionally, the feedback correction mileage is determined based on preset track feature points. The formula is as follows:
[0048]
[0049] In the formula, Represents the true mileage of the orbital feature points. This indicates the recorded mileage when the detection equipment detects a feature point on the track.
[0050] Optionally, based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data The formula is as follows:
[0051]
[0052]
[0053]
[0054]
[0055] In the formula, Indicates the weighting coefficient. The standard deviation of satellite positioning error. The standard deviation of inertial navigation data error. This represents the standard deviation of the feature point matching error.
[0056] Optionally, the data cleaning steps for the detection data include:
[0057] Spatiotemporal weighted interpolation is used to clean the detection data and repair missing values in the data.
[0058] An adaptive probability density outlier detection method is used to filter outliers in the detection data.
[0059] Optionally, the steps of cleaning and repairing missing values in the detection data using spatiotemporal weighted interpolation include:
[0060] Constructing spatiotemporal nearest neighbor regions for missing points;
[0061] For each missing point, calculate the spatial weight of neighboring measurement points. and time weight ;
[0062] Interpolation is calculated using a spatiotemporal weighted formula to obtain estimated values for missing points, which are then filled into the missing locations.
[0063] Optionally, the spatiotemporal weighting formula is as follows:
[0064]
[0065] In the formula, To calculate the estimated value of the missing points, Indicates time Mileage location The multidimensional feature vector at the location, This indicates the number of nearby measurement points.
[0066] Optionally, the steps of filtering outliers in the detection data using an adaptive probability density outlier detection method include:
[0067] Standardize the raw data;
[0068] Use kernel density estimation to estimate the probability density function of the data;
[0069] Local probability density calculation: For each data point, calculate its probability density value;
[0070] The threshold for outliers is dynamically determined based on the probability density distribution of data points.
[0071] Data points with probability density values below a threshold are marked as outliers and filtered out.
[0072] Optionally, the data quality scoring model includes:
[0073]
[0074] In the formula, Used to indicate the type of detection data. This is the maximum permissible value for the noise ratio.
[0075] Secondly, a device for processing high-speed railway infrastructure inspection data is provided, comprising:
[0076] An acquisition module is used to acquire detection data, which includes structured data and unstructured data;
[0077] The mileage calibration module is used to perform mileage calibration on structured data using a dynamic functional feature modeling alignment method.
[0078] The cleaning module is used to clean the detection data.
[0079] The first building module is used to build an intermediate data model, which is used to mix and store structured and unstructured data, and to perform mileage correction on the unstructured data through the intermediate data model;
[0080] The second building module is used to construct a data quality scoring model to score the detection data; when the score of the detection data meets the requirements, the detection data is stored through an intermediate data model.
[0081] Thirdly, an electronic device is provided, including a processing device for high-speed railway infrastructure detection data as described above.
[0082] Fourthly, a computer-readable storage medium is provided, wherein at least one piece of program code is stored therein, the program code being executed by a processor to implement the method for processing high-speed railway infrastructure detection data as described in any of the preceding claims.
[0083] The beneficial effects of the technical solution provided in this application include at least the following:
[0084] This application provides a method for processing high-speed railway infrastructure inspection data. Firstly, it proposes a nonlinear mileage calibration method based on dynamic feature modeling. This method utilizes multi-dimensional features such as curvature extrema to construct a Gaussian mixture model, combined with a hierarchical optimization strategy to dynamically calibrate mileage offsets. This overcomes the limitations of traditional static methods in adapting to complex track scenarios, thus addressing the technical problem of insufficient mileage calibration accuracy. Secondly, it provides a unified intermediate model (UDM) with unique index coding. This model uses a dynamic adapter to achieve semantic alignment and format conversion of heterogeneous data, and includes a data quality quantification scoring mechanism to address the technical challenge of multi-source data fusion. Attached Figure Description
[0085] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0086] Figure 1 A flowchart illustrating a method for processing high-speed railway infrastructure inspection data according to an embodiment of this application;
[0087] Figure 2 A structural block diagram of a high-speed railway infrastructure inspection data processing device provided in an embodiment of this application;
[0088] Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0089] The attached figures are labeled as follows:
[0090] 11: Acquisition Module; 12: Mileage Calibration Module; 13: Cleaning Module; 14: First Construction Module; 15: Second Construction Module;
[0091] 21: Processor; 22: Memory. Detailed Implementation
[0092] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0093] Traditional methods have the following shortcomings:
[0094] First, the accuracy of mileage calibration is insufficient: traditional methods do not take into account the dynamic changes of line geometry (such as curvature extreme points, differences between curve segments and straight segments), and lack effective compensation for cumulative errors caused by equipment slippage and signal drift, resulting in feature misalignment when multi-source data is fused.
[0095] Second, the data processing methods are too simplistic: missing value repair relies on single-dimensional spatial or temporal information (such as linear interpolation based solely on mileage distance), which cannot take into account spatiotemporal correlation; outlier detection does not combine the data probability density distribution to dynamically adjust the threshold, which is prone to misjudging disease data or retaining abnormal noise.
[0096] Third, multi-source data fusion is difficult: the data formats and semantic definitions of different detection devices (such as sensors and cameras) differ significantly. There is a lack of a unified data model to support the mixed storage and semantic alignment of structured (such as numerical indicators) and unstructured (such as images) data, resulting in semantic ambiguity (such as the different meanings of "displacement" indicators) and lack of quality assessment in cross-source data.
[0097] To address the aforementioned issues, this application provides a method for processing high-speed railway infrastructure inspection data. Firstly, it proposes a nonlinear mileage calibration method based on dynamic feature modeling. This method utilizes multi-dimensional features such as curvature extrema to construct a Gaussian mixture model, combined with a hierarchical optimization strategy to dynamically calibrate mileage offsets. This overcomes the limitations of traditional static methods in adapting to complex track scenarios, thus addressing the technical problem of insufficient mileage calibration accuracy. Secondly, it proposes a spatiotemporal neighborhood joint weighted interpolation method, combined with adaptive kernel density estimation for outlier detection, balancing data spatiotemporal correlation and noise robustness to compensate for the limitations of single data processing methods. Thirdly, it provides a unified intermediate model (UDM) with unique index encoding. This model uses a dynamic adapter to achieve semantic alignment and format conversion of heterogeneous data, and includes a data quality quantification scoring mechanism to address the technical challenge of multi-source data fusion.
[0098] Figure 1A flowchart illustrating a method for processing high-speed railway infrastructure inspection data according to an embodiment of this application. See also... Figure 1 The processing method for the high-speed railway infrastructure inspection data includes the following steps:
[0099] S101. Obtain detection data, which includes structured data and unstructured data.
[0100] In one example, the detection data includes structured data and unstructured data, wherein the structured data includes one or more of curvature, track orientation, and track gauge, and the unstructured data includes one or more of vibration signals, images, and videos.
[0101] In this embodiment, mileage calibration is performed on structured data.
[0102] S102. The dynamic functional feature modeling alignment method is used to perform mileage calibration on structured data.
[0103] In one example, step S102 includes:
[0104] S1021. Extract structured data from the detection data and mark the extreme points of the structured data.
[0105] S1022. Based on the curvature in the structured data, the line is divided into straight segments and curved segments. Initial feature diffusion is set for the straight segments and curved segments respectively. Based on the extreme points and the initial feature diffusion, a Gaussian mixture model is constructed.
[0106] The Gaussian mixture model is as follows:
[0107]
[0108] In the formula, Indicates the reference data at the mileage position The first The values of feature points in structured data, This indicates the first [number] determined based on design drawings or historical benchmark data. The theoretical mileage of feature points in structured data, Indicates the first The diffusion of feature points in structured data These are the weight coefficients in a Gaussian mixture model, used to represent the weights of the first Gaussian mixture model. The weights of feature points in structured data.
[0109] S1023. Initialize the dynamic parameters. Based on the section type of the line, initialize the weight coefficients in the Gaussian mixture model.
[0110] S1024. Determine the initial feature diffusion (also weighting coefficient) in the Gaussian mixture model based on the accuracy calibration of the detection equipment.
[0111] S1025. Based on the Gaussian mixture model and mileage calibration values, a multi-objective optimization model is constructed to minimize the difference between the theoretical characteristic function and the actual detection data.
[0112] The multi-objective optimization model is as follows:
[0113]
[0114] In the formula, This represents the matching error term. Indicates the position at the offset mileage. The first The values of feature points in structured data, This indicates the measured data at the mileage location. The first The values of feature points in structured data, Represents the regularization term, used to constrain... The magnitude of the mutation, This represents the smoothness constraint coefficient for mileage offset. This represents a feature stabilization term used to suppress the influence of noise-sensitive features on mileage offset. This indicates the mileage offset.
[0115] S1026. Hierarchical optimization calibration is adopted to solve the mileage calibration value in the multi-objective optimization model.
[0116] In one example, step S1026 includes:
[0117] Hierarchical optimization calibration is employed to solve the global problem across the entire line. Adopting large Quickly locate the approximate offset range for coarse calibration, then optimize the local Δ in segments (e.g., every 2km), combined with small... Fine calibration is performed using curvature continuity constraints.
[0118] The specific steps of step S1026 are as follows:
[0119] The mileage calibration value is solved iteratively using the Levenberg-Marquardt optimization algorithm, which is as follows:
[0120]
[0121] In the formula, Indicates the mileage calibration value. Represents the Jacobian matrix. Indicates the damping factor. This represents the set of feature points of structured data. Let represent the residual vector, where .
[0122] S1027. Perform mileage calibration on the test data based on the mileage calibration value obtained from the solution.
[0123] S103. Perform data cleaning on the detection data.
[0124] In one example, step S103 includes:
[0125] S1031. Remove duplicate values from the detection data.
[0126] S1032. Use spatiotemporal weighted interpolation to clean the detection data and repair missing values in the detection data.
[0127] In one example, step S1032 includes:
[0128] The first step is to construct the spatiotemporal nearest neighbor for the missing points.
[0129] In the spatial dimension, for each missing measurement point Select the nearest coordinate by mileage. Valid measuring points (row)
[0130] (excluding missing points), denoted as In the time dimension, for each missing time point Choose before and after Each effective time step is denoted as . .
[0131] Step 2: For each missing point Calculate the spatial weights of neighboring measurement points and time weight .
[0132] in,
[0133]
[0134] In the formula, The point to be interpolated With the Distance between neighboring measuring points; The spatial kernel width is used to control the influence range of adjacent measuring points; It is the number of spatially adjacent measurement points; It is the first A timestamp at a specific point in time; It is the time decay factor; It is the half-length of the time window, and the total length of the window is .
[0135] The third step is to use the spatiotemporal weighted formula to calculate the interpolation, obtain the estimated value of the missing point, and fill it into the missing position.
[0136] The spatiotemporal weighting formula is as follows:
[0137]
[0138] In the formula, To calculate the estimated value of the missing points, Indicates time Mileage location The multidimensional feature vector at the location, This indicates the number of nearby measurement points.
[0139] S1033. An adaptive probability density outlier detection method is used to filter outliers in the detection data.
[0140] In one example, step S1033 includes:
[0141] The first step is to standardize the raw data.
[0142] The mean of the data after standardization is 0, and the standard deviation is 1.
[0143] The second step is to use kernel density estimation to estimate the probability density function of the data.
[0144] The probability density function is:
[0145]
[0146] In the formula, yes The probability density function at that location; It refers to the number of data points; It is the first One data point; It is a Gaussian kernel function; The bandwidth parameter is used to control the width of the kernel function.
[0147] The third step is to calculate the local probability density for each data point. Calculate its probability density value .
[0148] Step 4: Dynamically determine the threshold for outliers based on the probability density distribution of the data points. .
[0149]
[0150] In the formula, It is a coefficient that controls the size of the threshold; It is the median of the probability density values; yes
[0151] A coefficient that controls the strictness of the threshold; It is the standard deviation of the probability density value.
[0152] Step 5: Set the probability density value below the threshold. The data points were marked as outliers and filtered out.
[0153] S104. Construct an intermediate data model for mixed storage of structured and unstructured data, and perform mileage correction on the unstructured data through the intermediate data model.
[0154] The intermediate data model is as follows:
[0155]
[0156] Among them, ID: globally unique identifier.
[0157] SourceType: Data source type (e.g., "sensor data" or "image").
[0158] Timestamp: Data acquisition time, supports time sequence alignment.
[0159] Metadata: Standardized metadata (such as device ID, line, damage type, etc.).
[0160] RawData: Raw data.
[0161] Features: Structured features (e.g., {"crack length": 3mm}).
[0162] A standardized indicator system is constructed and deeply coupled with the UDM to ensure semantic consistency and eliminate semantic ambiguity in multi-source data. A globally unique indicator code is defined (e.g., indicator number "U0V0W2-X5" corresponds to the indicator name "Starting Mileage," representing the "starting mileage of the point or range where the track's abnormal service status occurs"). The adapter automatically matches the code and writes it into the UDM's metadata during parsing. For cases of conflicting names in multi-source data (e.g., "displacement" might refer to "bridge deflection" or "track offset"), the indicator system is used to convert them into unified standard terminology, ensuring semantic consistency of UDM fields.
[0163] For different types of monitoring and inspection data in high-speed railway operation and maintenance, a dedicated adapter is used to achieve standardized mapping from multi-source data to a unified intermediate model (UDM). Equipment time-series data is parsed through a time-series adapter, stored in the UDM's RawData, and key fields (equipment number, timestamp, numerical indicators, etc.) are extracted and mapped to the UDM's Features. Image / video data undergoes a visual adapter to extract metadata (resolution, timestamp) and key defect features (defect features can be determined from images and videos using pre-trained models), generating standardized descriptions which are then stored in the UDM's Metadata and Features.
[0164] The adapters support dynamic configuration and semantic self-recognition. The text / database adapter can automatically recognize CSV / JSON delimiters or table structures, align fields to the UDM's metadata using mapping rules, and directly map structured data to the UDM's Features fields. The image / video adapter can extract features from unstructured data based on a pre-trained model, generating standardized key-value pairs (such as images) stored in Features. {"Defect Type": "Steel rail fragment", "Area": 1cm} 2}).
[0165] Intermediate data models are also used for mileage correction of unstructured data. In one example, the mileage correction process for unstructured data is as follows:
[0166] S1041. Determine the error between the theoretical mileage and the actual mileage of the unstructured data. .
[0167] Among these, the error between the theoretical mileage and the actual mileage of the unstructured data is determined. The formula is as follows:
[0168]
[0169] In the formula, This indicates the theoretical mileage of the testing equipment. Indicates the actual mileage of the testing equipment.
[0170] Indicates the first Fixed deviations caused by the sliding or sampling errors of the detection equipment within each detection section. This represents the dynamic error that accumulates over time.
[0171] S1042, Based on the aforementioned error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. .
[0172] According to the error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. The formula is as follows:
[0173]
[0174] In the formula, Indicates time The BeiDou absolute coordinates of the time detection equipment This indicates the geographical coordinates corresponding to the mileage shown on the design drawings.
[0175] S1043. Determine the relative displacement increment based on the instantaneous velocity and instantaneous acceleration of the detection equipment. .
[0176] The relative displacement increment is determined based on the instantaneous velocity and instantaneous acceleration of the detection equipment. The formula is as follows:
[0177]
[0178] In the formula, Indicates the instantaneous speed of the detection equipment. This indicates the instantaneous acceleration of the detection equipment.
[0179] S1044. Determine the feedback correction mileage based on the preset track feature points. .
[0180] Based on preset track feature points, determine the feedback correction mileage. The formula is as follows:
[0181]
[0182] In the formula, Represents the true mileage of the orbital feature points. This indicates the recorded mileage when the detection equipment detects a feature point on the track.
[0183] S1045, Based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data .
[0184] Based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data The formula is as follows:
[0185]
[0186]
[0187]
[0188]
[0189] In the formula, Indicates the weighting coefficient. The standard deviation of satellite positioning error. The standard deviation of inertial navigation data error. This represents the standard deviation of the feature point matching error.
[0190] S105. Construct a data quality scoring model to score the test data; when the score of the test data meets the requirements, store the test data through an intermediate data model.
[0191] The data quality scoring models include:
[0192]
[0193] In the formula, Used to indicate the type of detection data. This is the maximum allowable value for the noise ratio. The missing rate is the proportion of missing information in the Features field of the UDM, while the noise ratio is the proportion of invalid information in the Raw Data.
[0194] Only when Data is stored in the warehouse only when it is ready; otherwise, it is deemed unqualified and needs to be re-collected. Seamless integration with the data warehouse is achieved based on UDM, ensuring data standardization. Multi-source data is stored in its original format in the data lake layer. The adapter parses the original data into the standardized UDM format and then stores it in the data warehouse layer, supporting efficient querying and analysis. Data lineage traceability in the data warehouse is achieved through the SourceType and ID fields of the UDM. Based on high-speed railway line maintenance rules and high-speed railway bridge and tunnel structure repair rules, exceedance thresholds are set for the treated data, automatically assessing the exceedance level of defects.
[0195] Figure 2 A structural block diagram of a high-speed railway infrastructure inspection data processing device provided in one embodiment of this application. See also... Figure 2 The data processing device for high-speed railway infrastructure inspection includes:
[0196] The acquisition module 11 is used to acquire detection data, which includes structured data and unstructured data;
[0197] Mileage calibration module 12 is used to perform mileage calibration on structured data using a dynamic functional feature modeling alignment method.
[0198] Cleaning module 13 is used to clean the detection data;
[0199] The first construction module 14 is used to build an intermediate data model, which is used to mix and store structured data and unstructured data, and to perform mileage correction on unstructured data through the intermediate data model;
[0200] The second construction module 15 is used to construct a data quality scoring model to score the detection data; when the score of the detection data meets the requirements, the detection data is stored through an intermediate data model.
[0201] It should be noted that since the above steps have already explained the entire process of processing high-speed railway infrastructure inspection data, the function of the high-speed railway infrastructure inspection data processing device can be referred to in the above steps and will not be repeated here.
[0202] Figure 3 This is a structural block diagram of an electronic device provided according to an embodiment of this application. See also... Figure 3 Electronic devices may include Figure 2 The aforementioned device for processing high-speed railway infrastructure inspection data. Typically, the electronic equipment includes a processor 21 and a memory 22.
[0203] Processor 21 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 21 may also include a main processor and a coprocessor. The main processor is used to process data in the wake-up state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. Memory 22 may include one or more computer-readable storage media, which may be non-transitory. Memory 22 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 22 is used to store at least one instruction, which is executed by processor 21 to implement the high-speed railway infrastructure detection data processing method performed by an electronic device provided in the method embodiments of this application.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing inspection data of high-speed railway infrastructure, characterized in that, include: Acquire detection data, which includes structured data and unstructured data; A dynamic functional feature modeling alignment method is used to perform mileage calibration on structured data, specifically including: Structured data is extracted from the detection data, and extreme points of the structured data are marked; Based on the curvature in the structured data, the line is divided into straight segments and curved segments. Initial feature diffusion is set for both straight and curved segments. Based on the extreme points and the initial feature diffusion, a Gaussian mixture model is constructed, as follows: In the formula, Indicates the reference data at the mileage position The first The values of feature points in structured data, This indicates the first [number] determined based on design drawings or historical benchmark data. The theoretical mileage of feature points in structured data, Indicates the first The diffusion of feature points in structured data These are the weight coefficients in a Gaussian mixture model, used to represent the weights of the first Gaussian mixture model. The weights of feature points in structured data; Initialize the dynamic parameters and, based on the segment type of the line, initialize the weight coefficients in the Gaussian mixture model. The initial feature diffusion in the Gaussian mixture model is determined based on the accuracy calibration of the detection equipment. Based on the Gaussian mixture model and odometer calibration values, a multi-objective optimization model is constructed to minimize the difference between the theoretical feature function and the actual detection data. The multi-objective optimization model is as follows: In the formula, This represents the matching error term. Indicates the position at the offset mileage. The first The values of feature points in structured data, This indicates the measured data at the mileage location. The first The values of feature points in structured data, Represents the regularization term, used to constrain... The magnitude of the mutation, This represents the smoothness constraint coefficient for mileage offset. This represents a feature stabilization term used to suppress the influence of noise-sensitive features on mileage offset. Indicates the mileage offset; A hierarchical optimization calibration method is used to solve for the mileage calibration values in the multi-objective optimization model, as detailed below: The mileage calibration value is solved iteratively using the Levenberg-Marquardt optimization algorithm, which is as follows: In the formula, Indicates the mileage calibration value. Represents the Jacobian matrix. Indicates the damping factor. This represents the set of feature points of structured data. Let represent the residual vector, where ; The structured data is calibrated based on the mileage calibration values obtained from the solution. Perform data cleaning on the test data; Construct an intermediate data model to store structured and unstructured data in a mixed manner, and perform mileage correction on the unstructured data through the intermediate data model; A data quality scoring model is constructed to score the detection data; when the score of the detection data meets the requirements, the detection data is stored through an intermediate data model.
2. The method for processing high-speed railway infrastructure inspection data according to claim 1, characterized in that, The detection data includes structured data and unstructured data. The structured data includes one or more of curvature, track orientation, and track gauge, while the unstructured data includes one or more of vibration signals, images, and videos.
3. The method for processing high-speed railway infrastructure inspection data according to claim 1, characterized in that, The steps for mileage correction of unstructured data using an intermediate data model include: Determine the error between theoretical and actual mileage for unstructured data. ; According to the error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. ; The relative displacement increment is determined based on the instantaneous velocity and instantaneous acceleration of the detection equipment. ; Based on preset track feature points, determine the feedback correction mileage. ; Based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data .
4. The method for processing high-speed railway infrastructure inspection data according to claim 3, characterized in that, Determine the error between theoretical and actual mileage for unstructured data. The formula is as follows: In the formula, This indicates the theoretical mileage of the testing equipment. Indicates the actual mileage of the testing equipment. Indicates the first Fixed deviations caused by the sliding or sampling errors of the detection equipment within each detection section. This represents the dynamic error that accumulates over time.
5. The method for processing high-speed railway infrastructure inspection data according to claim 3, characterized in that, According to the error Determine the geographic coordinates corresponding to the mileage on the design drawings, and determine the initial offset based on the absolute coordinates of the satellite positioning and the geographic coordinates corresponding to the mileage on the design drawings. The formula is as follows: In the formula, Indicates time The BeiDou absolute coordinates of the time detection equipment This indicates the geographical coordinates corresponding to the mileage shown on the design drawings.
6. The method for processing high-speed railway infrastructure inspection data according to claim 3, characterized in that, The relative displacement increment is determined based on the instantaneous velocity and instantaneous acceleration of the detection equipment. The formula is as follows: In the formula, Indicates the instantaneous speed of the detection equipment. This indicates the instantaneous acceleration of the detection equipment.
7. The method for processing high-speed railway infrastructure inspection data according to claim 3, characterized in that, Based on preset track feature points, determine the feedback correction mileage. The formula is as follows: In the formula, Represents the true mileage of the orbital feature points. This indicates the recorded mileage when the detection equipment detects a feature point on the track.
8. The method for processing high-speed railway infrastructure inspection data according to claim 3, characterized in that, Based on the initial offset Relative displacement increment and feedback correction mileage Determine the actual correction mileage of unstructured data The formula is as follows: In the formula, Indicates the weighting coefficient. The standard deviation of satellite positioning error. The standard deviation of inertial navigation data error. This represents the standard deviation of the feature point matching error.
9. The method for processing high-speed railway infrastructure inspection data according to any one of claims 1 to 8, characterized in that, The steps for cleaning the test data include: Spatiotemporal weighted interpolation is used to clean the detection data and repair missing values in the data. An adaptive probability density outlier detection method is used to filter outliers in the detection data.
10. The method for processing high-speed railway infrastructure inspection data according to claim 9, characterized in that, The steps for cleaning and repairing missing values in the detection data using spatiotemporal weighted interpolation include: Constructing spatiotemporal nearest neighbor regions for missing points; For each missing point, calculate the spatial weight of neighboring measurement points. and time weight ; Interpolation is calculated using a spatiotemporal weighted formula to obtain estimated values for missing points, which are then filled into the missing locations.
11. The method for processing high-speed railway infrastructure inspection data according to claim 10, characterized in that, The spatiotemporal weighting formula is as follows: In the formula, To calculate the estimated value of the missing points, Indicates time Mileage location The multidimensional feature vector at the location, This indicates the number of nearby measurement points.
12. The method for processing high-speed railway infrastructure inspection data according to claim 9, characterized in that, The steps for filtering outliers in the detection data using the adaptive probability density outlier detection method include: Standardize the raw data; Use kernel density estimation to estimate the probability density function of the data; Local probability density calculation: For each data point, calculate its probability density value; The threshold for outliers is dynamically determined based on the probability density distribution of data points. Data points with probability density values below a threshold are marked as outliers and filtered out.
13. The method for processing high-speed railway infrastructure inspection data according to any one of claims 1 to 8, characterized in that, Data quality scoring models include: In the formula, Used to indicate the type of detection data. This is the maximum permissible value for the noise ratio.
14. A device for processing high-speed railway infrastructure inspection data, characterized in that, include: An acquisition module is used to acquire detection data, which includes structured data and unstructured data; The mileage calibration module is used to perform mileage calibration on structured data using a dynamic functional feature modeling alignment method. Specifically, it includes: Structured data is extracted from the detection data, and extreme points of the structured data are marked; Based on the curvature in the structured data, the line is divided into straight segments and curved segments. Initial feature diffusion is set for both straight and curved segments. Based on the extreme points and the initial feature diffusion, a Gaussian mixture model is constructed, as follows: In the formula, Indicates the reference data at the mileage position The first The values of feature points in structured data, This indicates the first [number] determined based on design drawings or historical benchmark data. The theoretical mileage of feature points in structured data, Indicates the first The diffusion of feature points in structured data These are the weight coefficients in a Gaussian mixture model, used to represent the weights of the first Gaussian mixture model. The weights of feature points in structured data; Initialize the dynamic parameters and, based on the segment type of the line, initialize the weight coefficients in the Gaussian mixture model. The initial feature diffusion in the Gaussian mixture model is determined based on the accuracy calibration of the detection equipment. Based on the Gaussian mixture model and odometer calibration values, a multi-objective optimization model is constructed to minimize the difference between the theoretical feature function and the actual detection data. The multi-objective optimization model is as follows: In the formula, This represents the matching error term. Indicates the position at the offset mileage. The first The values of feature points in structured data, This indicates the measured data at the mileage location. The first The values of feature points in structured data, Represents the regularization term, used to constrain... The magnitude of the mutation, This represents the smoothness constraint coefficient for mileage offset. This represents a feature stabilization term used to suppress the influence of noise-sensitive features on mileage offset. Indicates the mileage offset; A hierarchical optimization calibration method is used to solve for the mileage calibration values in the multi-objective optimization model, as detailed below: The mileage calibration value is solved iteratively using the Levenberg-Marquardt optimization algorithm, which is as follows: In the formula, Indicates the mileage calibration value. Represents the Jacobian matrix. Indicates the damping factor. This represents the set of feature points of structured data. Let represent the residual vector, where ; The structured data is calibrated based on the mileage calibration values obtained from the solution. The cleaning module is used to clean the detection data. The first building module is used to build an intermediate data model, which is used to mix and store structured and unstructured data, and to perform mileage correction on the unstructured data through the intermediate data model; The second building module is used to construct a data quality scoring model to score the detection data; when the score of the detection data meets the requirements, the detection data is stored through an intermediate data model.
15. An electronic device, characterized in that, It includes the processing device for high-speed railway infrastructure inspection data as described in claim 14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is executed by a processor to implement the method for processing high-speed railway infrastructure detection data as described in any one of claims 1 to 13.