A steel rail crack detection method based on multi-modal fusion

By employing a multimodal fusion-based rail crack detection method, dynamically configuring detection equipment parameters and performing closed-loop verification, the problems of low detection efficiency and insufficient accuracy in existing technologies are solved, achieving efficient and accurate rail crack detection.

CN121324381BActive Publication Date: 2026-04-14CRRC HANGZHOU DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CRRC HANGZHOU DIGITAL TECH CO LTD
Filing Date
2025-12-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing multimodal rail crack detection technologies suffer from low detection efficiency, insufficient accuracy, and poor decision reliability, failing to meet the demands of high precision and engineering applications.

Method used

The rail crack detection method using multimodal fusion dynamically configures subsequent equipment parameters based on the results of preceding modal diagnostics, performs focused detection of suspicious areas, and generates a final detection result with joint confidence through hierarchical fusion and closed-loop verification.

Benefits of technology

It significantly improves detection efficiency and the measurement accuracy of quantitative parameters, ensures high confidence in the final diagnostic results, provides direct and reliable basis for maintenance decisions, and enhances the engineering application effect of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121324381B_ABST
    Figure CN121324381B_ABST
Patent Text Reader

Abstract

The application provides a rail crack detection method based on multi-modal fusion, and belongs to the technical field of rail detection. The method comprises the following steps: a first detection device is used to preliminarily detect a rail to generate a first diagnosis result containing a defect type and a confidence; the severity score of each defect area of the rail is calculated based on the first diagnosis result, and a suspicious area with a severity score reaching a preset threshold is locked; according to the suspicious area and the defect type, the parameters of a second detection device are configured, a control instruction is generated to drive the second detection device to perform focused detection and obtain second modal data; the deep features of the second modal data are extracted, a pre-trained model is used to generate a second diagnosis result containing a defect quantization parameter and a second confidence; the two diagnosis results are spatio-temporally aligned and fused to form fusion evidence; the parameters of a third detection device are configured based on the fusion evidence to complete verification detection, and the final detection conclusion is output after type verification and confidence arbitration, so that accurate recognition and scientific determination of rail cracks are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail inspection technology, and more specifically to a rail crack detection method based on multimodal fusion. Background Technology

[0002] As a core load-bearing component of the railway transportation system, the structural integrity of rails directly affects train operation safety and transportation efficiency. Under the combined effects of long-term loads, environmental corrosion, and material fatigue, rails are prone to surface cracks, internal inclusions, and other defects. If these defects are not detected promptly and accurately, they may lead to crack propagation or even rail breakage, causing serious safety accidents. Single-modal inspection technologies (such as visual inspection, which can only identify visible surface defects, and ultrasonic inspection, which has limited ability to identify micro-cracks) are no longer sufficient to meet the needs of comprehensive and high-precision inspection. Multimodal fusion inspection, due to its ability to integrate the complementary advantages of different modes, has become the mainstream development direction for rail crack detection.

[0003] Existing multimodal rail crack detection technologies still have significant limitations: On the one hand, the detection process often adopts a "uniform acquisition across the entire area + fixed parameters" approach, without specifically focusing on potential defect areas. This results in a large amount of redundant data in normal areas, reducing detection efficiency, and the lack of precise detection in key areas leads to insufficient accuracy in measuring quantitative parameters such as crack length, width, and depth, making it difficult to meet the needs of both large-scale coverage and precise detection. On the other hand, multimodal data fusion often uses simple superposition or weighted voting methods, failing to effectively resolve type conflicts and parameter deviations in the diagnostic results of each mode. Furthermore, the lack of a closed-loop verification process for the preliminary diagnostic results leads to low confidence in the final diagnostic results, making it impossible to directly provide accurate and reliable basis for subsequent maintenance decisions and limiting the engineering application of the technology. Summary of the Invention

[0004] This invention provides a rail crack detection method based on multimodal fusion, which solves the problems of imbalance between multimodal detection efficiency and accuracy, insufficient decision reliability and practicality in the existing technology.

[0005] To achieve the above objectives, this invention provides a rail crack detection method based on multimodal fusion, comprising the following steps: Step S1: Scanning the rail with a first detection device to obtain first modal data; performing a preliminary diagnosis on the rail based on the first modal data to generate a first diagnostic result; locating suspicious areas on the rail based on the first diagnostic result and extracting initial attribute information of the suspicious areas; Step S2: Dynamically configuring the detection parameters of a second detection device to focus on detecting the suspicious areas based on the initial attribute information to obtain second modal data; diagnosing the rail based on the second modal data to generate a second diagnostic result; Step S3: Performing consistency verification and fusion of the first and second diagnostic results to generate fused diagnostic evidence and joint confidence; Step S4: Configuring the detection parameters of a third detection device based on the fused diagnostic evidence and joint confidence; the third detection device performing verification detection on the suspicious areas based on the configured detection parameters to obtain third modal data; Step S5: Generating a third diagnostic result based on the third modal data, and verifying it in conjunction with the fused diagnostic evidence and joint confidence, generating a final detection result based on the verification result.

[0006] Optionally, the first diagnostic result includes a set of potential defect regions, a first defect type, and its corresponding first confidence level. The generation of the first diagnostic result includes: based on the first modal data, identifying all defect regions using an anomaly detection algorithm, and labeling the physical coordinates of each defect region to form a set of potential defect regions; extracting features from each defect region in the set of potential defect regions to obtain feature parameters related to cracks; matching the feature parameters of each defect region with a preset defect feature template library, calculating the matching degree of each template, and taking the defect type corresponding to the template with the highest matching degree as the first defect type of the defect region; calculating the signal-to-noise ratio of the first modal data, and fusing the matching degree of the template corresponding to the first defect type to obtain the first confidence level.

[0007] Optionally, the location of the suspicious area includes: calculating the severity score of each defect area based on the first defect type and its first confidence level of each defect area in the set of potential defect areas, comparing it with a preset threshold, and filtering out defect areas with scores higher than the threshold; and determining the spatial range defined by the outer rectangle of the filtered defect areas as the suspicious area.

[0008] Optionally, the initial attribute information includes the spatial size and geometry of the suspicious area and the defect type of the suspicious area. The dynamic configuration of the detection parameters of the second detection device includes: mapping the required scanning path density and resolution of the second detection device according to the spatial size and geometry of the suspicious area; matching the optimal detection mode from a variety of pre-stored detection modes of the second detection device according to the first defect type; binding the scanning path density, resolution and detection mode to the physical coordinates of the suspicious area, generating control commands and sending them to the second detection device, so that the second detection device performs focused detection on the suspicious area based on the control commands.

[0009] Optionally, the second diagnostic result includes a second defect type, a second confidence level, and defect quantification parameters. The generation of the second diagnostic result includes: extracting features from the second modality data to obtain deep features related to the precise quantification of the defect; inputting the deep features into a pre-trained second defect classification and regression model to obtain the second defect type and defect quantification parameters; and calculating the second confidence level of the second diagnostic result based on the classification probability of the deep features in the second defect classification and regression model and the signal-to-noise ratio of the second modality data during the focusing detection process.

[0010] Optionally, the consistency verification and fusion includes: spatiotemporally aligning the relevant diagnostic features in the first and second diagnostic results based on the physical coordinates of the suspicious area, and establishing the correlation between the features to form a preliminary fusion evidence set; performing consistency verification on the defect type judgments from the first and second diagnostic results in the preliminary fusion evidence set; generating a collaborative verification instruction when the verification results are consistent; generating a priority verification instruction pointing to the high-confidence diagnostic result based on the comparison result of the first confidence level and the second confidence level; and calculating the joint confidence level of the fusion diagnostic evidence based on the first confidence level and the second confidence level using a weighted average algorithm.

[0011] Optionally, the configuration process of the detection parameters of the third detection device includes: comparing the joint confidence level with the preset interval, and determining the verification mode of the third device based on the comparison result; extracting the estimated depth range and spatial orientation of the defect from the fused diagnostic evidence, and mapping them to the detection parameters of the third detection device; binding the determined verification mode with the detection parameters, generating control commands and sending them to the third detection device to configure the detection parameters of the third detection device.

[0012] Optionally, the verification process of the third diagnostic result includes: type consistency verification: comparing the defect type in the third diagnostic result with the dominant defect type in the fused diagnostic evidence to determine whether the types are consistent; confidence dominance verification: if the types are consistent, comparing the confidence of the third diagnostic result with the joint confidence to determine whether the former is higher than the latter and the difference exceeds a preset dominance threshold; confidence arbitration comparison: if the types are inconsistent or the confidence difference does not exceed the dominance threshold, comparing the joint confidence with the confidence of the third diagnostic result.

[0013] Optionally, generating the final test result based on the verification results includes: if the types are consistent and the confidence comparison meets the advantage condition, the third diagnostic result is taken as the final conclusion; if the confidence arbitration comparison is entered, the diagnostic result corresponding to the one with higher confidence is taken as the final conclusion; if the difference in confidence between the two in the confidence arbitration comparison is less than a preset tolerance threshold, an uncertain conclusion is generated and marked as requiring manual review.

[0014] The present invention provides a rail crack detection method based on multimodal fusion. Through a time-series multimodal fusion process of "pre-modal diagnosis guiding subsequent equipment dynamic parameter configuration, layered fusion to generate joint confidence, and third-modal targeted verification," it not only achieves focused detection of suspicious areas and significantly reduces redundant data to improve detection efficiency, but also effectively resolves conflicts between different modal diagnoses through progressive fusion and closed-loop verification. This significantly improves the accuracy of crack presence determination and the measurement of quantitative parameters such as length, width, and depth, while ensuring high confidence of the final diagnostic results. It takes into account both the need for wide coverage and precise detection, and can provide direct and reliable basis for rail maintenance decisions, significantly improving the engineering application effect of rail crack detection. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0016] Figure 1 This is a flowchart of the rail crack detection method provided in the embodiments of the present invention;

[0017] Figure 2 This is a flowchart of dynamic configuration and focusing detection provided in an embodiment of the present invention;

[0018] Figure 3 This is a flowchart of intelligent verification and decision-making provided in an embodiment of the present invention. Detailed Implementation

[0019] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0020] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0021] With the upgrading of railway transportation towards high speed and heavy load, the structural integrity of rails, as the core load-bearing component, directly determines transportation safety, making the demand for accurate and efficient crack detection increasingly urgent. Existing multimodal rail crack detection technologies have inherent limitations: the uniform acquisition mode across the entire area leads to a 30%-50% increase in redundant data, resulting in low detection efficiency; fixed detection parameters cannot adapt to different defect scenarios, causing excessive deviations in crack quantification parameter measurement; moreover, multimodal data is mostly simply superimposed and fused, making it difficult to resolve diagnostic conflicts, resulting in high false alarm and false negative rates; and the lack of closed-loop verification leads to insufficient confidence in diagnostic results, failing to meet the precise decision-making needs of engineering operation and maintenance.

[0022] To address this issue, this invention proposes a rail crack detection method based on multimodal fusion. This method deploys multimodal detection equipment sequentially, dynamically configuring the scanning density, resolution, and detection mode of subsequent equipment based on the diagnostic results of the preceding modalities, achieving focused acquisition of suspicious areas. A hierarchical fusion algorithm is used to perform spatiotemporal alignment and consistency verification of the multimodal diagnostic results, generating fused diagnostic evidence and joint confidence scores. A third modal-based targeted verification forms a closed loop, constructing a full-process intelligent detection system of "acquisition-diagnosis-configuration-verification." This significantly improves crack detection efficiency, quantitative parameter measurement accuracy, and diagnostic result reliability, providing direct and accurate decision-making basis for rail graded maintenance.

[0023] The following is combined Figures 1-3 This invention is described in detail.

[0024] like Figure 1As shown, this embodiment of the invention provides a rail crack detection method based on multimodal fusion, including the following steps: Step S1: Scanning the rail with a first detection device to obtain first modal data; performing a preliminary diagnosis on the rail based on the first modal data to generate a first diagnostic result; locating suspicious areas on the rail based on the first diagnostic result and extracting initial attribute information of the suspicious areas; Step S2: Dynamically configuring the detection parameters of a second detection device to focus on detecting the suspicious areas according to the initial attribute information to obtain second modal data; diagnosing the rail based on the second modal data to generate a second diagnostic result; Step S3: Performing consistency verification and fusion on the first diagnostic result and the second diagnostic result to generate fused diagnostic evidence and joint confidence; Step S4: Configuring the detection parameters of a third detection device according to the fused diagnostic evidence and joint confidence; the third detection device performing verification detection on the suspicious areas based on the configured detection parameters to obtain third modal data; Step S5: Generating a third diagnostic result based on the third modal data, and verifying it in conjunction with the fused diagnostic evidence and joint confidence, generating a final detection result based on the verification result.

[0025] Multimodal data refers to a heterogeneous dataset of rail cracks collected by devices using different detection principles, reflecting different physical characteristics. This includes, but is not limited to, ultrasonic modal data (e.g., acoustic signal data for detecting internal defects), three-dimensional contour modal data (e.g., point cloud data characterizing surface geometry), and visual modal data (e.g., image data capturing surface texture and edges). Suspicious areas refer to rail spatial regions identified in the preliminary diagnosis where the probability of a crack is higher than a preset threshold, defined by physical coordinate intervals. Dynamically configured detection parameters refer to the adaptive adjustment of the detection equipment's operating parameters based on initial attribute information, including but not limited to scan path density, data sampling frequency, detection resolution, and detection mode. Fusion diagnostic evidence refers to a structured dataset formed by integrating effective diagnostic information from various modalities after consistency verification. This dataset includes a unified defect location, dominant defect type, preliminary quantitative parameters, and supporting evidence for each modality.

[0026] The rail crack detection method provided in this invention comprises the following steps: A first detection device scans the rail to obtain first modal data and completes a preliminary diagnosis, locating suspicious areas and extracting initial attribute information. Then, based on this initial attribute information, the detection parameters of a second detection device are dynamically configured to focus on the suspicious areas, obtaining second modal data and generating a second diagnostic result. Subsequently, the consistency of the first two diagnostic results is verified and fused to form fused diagnostic evidence and joint confidence. Based on this, the detection parameters of a third detection device are configured to verify the suspicious areas and obtain third modal data. Finally, a third diagnostic result is generated based on the third modal data, and verification is completed by combining the fused diagnostic evidence and joint confidence, outputting the final detection result. This process achieves precise focusing on suspicious areas by dynamically guiding the parameters of subsequent detection devices through the preceding diagnostic results. Layered verification fusion and closed-loop verification eliminate diagnostic conflicts between modalities, reducing redundant data, improving detection efficiency, and enhancing the reliability and accuracy of diagnostic results, providing strong support for rail maintenance decisions.

[0027] Preferably, the first diagnostic result includes a set of potential defect regions, a first defect type, and its corresponding first confidence level. The generation of the first diagnostic result includes: based on the first modal data, identifying all defect regions using an anomaly detection algorithm, and labeling the physical coordinates of each defect region to form a set of potential defect regions; extracting features from each defect region in the set of potential defect regions to obtain feature parameters related to cracks; matching the feature parameters of each defect region with a preset defect feature template library, calculating the matching degree of each template, and taking the defect type corresponding to the template with the highest matching degree as the first defect type of the defect region; calculating the signal-to-noise ratio of the first modal data, and fusing the matching degree of the template corresponding to the first defect type to obtain the first confidence level.

[0028] Among them, the anomaly detection algorithm refers to a targeted algorithm applicable to rail crack detection. It extracts normal distribution features of the first modal data, such as signal amplitude and texture patterns, identifies anomalous signal segments deviating from this distribution, and then locates areas where cracks may exist. This includes, but is not limited to, improved algorithms based on thresholding and clustering analysis. The potential defect region set refers to the collection of all rail regions suspected of containing cracks identified by the anomaly detection algorithm, each region being identified by unique physical coordinates. Crack-related feature parameters refer to quantitative indicators extracted from the first modal data of the defect region that characterize the physical properties of the crack, including but not limited to signal amplitude change rate, texture complexity, defect region duty cycle, and signal mutation frequency. The defect feature template library refers to a pre-constructed structured database storing standard feature parameter templates for different types of rail cracks, such as surface transverse cracks, internal longitudinal cracks, and weld cracks. Each template corresponds one-to-one with a specific defect type and is used for matching and comparing with measured feature parameters. Template matching degree refers to a numerical index that quantifies the similarity between the feature parameters of the measured defect area and a standard template in the template library. The value ranges from 0 to 1. The higher the matching degree, the greater the probability that the defect area belongs to the defect type identified by the corresponding template.

[0029] Specifically, the template matching degree is calculated using the cosine similarity algorithm to quantify the similarity between the feature parameters and the standard template, as shown in the following formula:

[0030]

[0031] in, For the first The matching degree of a standard template (0≤ ≤1); This is the feature parameter vector of the measured defect area; For the first in the template library Standard feature parameter vectors for each defect type; Represents the vector dot product. , Let represent the magnitudes of the vectors. The first confidence score is calculated by combining the fusion template matching score and the signal-to-noise ratio of the first modality data, using a weighted product method, as shown in the following formula:

[0032]

[0033] in, First confidence level (0≤ ≤1); The matching value corresponding to the template with the highest matching degree; The weighting coefficient for matching degree is set according to the priority of defect type identification; SNR is the signal-to-noise ratio of the first modality data; This is a normalization constant used to map the signal-to-noise ratio to the 0-1 interval.

[0034] For example, after scanning the rail with an electromagnetic ultrasonic probe (the first detection device) to obtain first-mode data in the form of acoustic signals, the data is analyzed using an improved threshold method to identify three abnormal segments that deviate from the normal signal distribution. The rail mileage and lateral position corresponding to each segment are marked to form a set of potential defect areas. For a specific defect area, feature parameters related to cracks, such as the rate of change of signal amplitude, texture complexity, and number of signal mutations, are extracted. These feature parameters are compared with a preset defect feature template library. The matching degree of each template is calculated using a cosine similarity algorithm. The surface transverse crack template has the highest matching degree, and the first defect type of the defect area is determined to be a surface transverse crack. Subsequently, the signal-to-noise ratio of the first-mode data of the area is calculated. The signal-to-noise ratio is fused with the highest matching degree, and the first confidence level corresponding to the defect area is obtained through a weighted product method. Finally, a first diagnostic result containing the set of potential defect areas, the first defect type of each area, and the corresponding first confidence level is generated.

[0035] In a preferred embodiment of the present invention, potential defect areas of rails are accurately identified through anomaly detection, the defect type is determined by feature matching, and the diagnostic confidence is quantified by fusing matching degree and data signal-to-noise ratio. This generates a first diagnostic result that is complete and quantifiable in reliability, providing accurate and reliable preliminary basis for subsequent location of suspicious areas and dynamic configuration of detection parameters.

[0036] Preferably, the location of the suspicious area includes: calculating the severity score of each defect area based on the first defect type and its first confidence level of each defect area in the set of potential defect areas, comparing it with a preset threshold, and filtering out defect areas with scores higher than the threshold; and determining the spatial range defined by the outer rectangle of the filtered defect areas as the suspicious area.

[0037] The severity score refers to the comprehensive assessment of the risk level and confidence level of the first defect type in the defect area, which is used to quantitatively evaluate the degree of hazard posed by the crack in that area. The score ranges from 0 to 10. A higher score indicates a greater potential risk to the safe operation of the rail. The calculation formula is as follows:

[0038]

[0039] in, Score the severity of the defect area; The risk weight corresponding to the first defect type is set according to the degree of impact of the defect on the rail's load-bearing capacity. For example, the weight of a transverse surface crack is 10, the weight of an internal longitudinal crack is 8, and the weight of a micro-crack at the weld is 5. This represents the first confidence level corresponding to the defective region.

[0040] In a preferred embodiment of the present invention, a severity score is calculated by combining defect type and confidence level, and high-risk areas are screened. The spatial range of the suspicious area is precisely defined by an circumscribed rectangle, thereby achieving precise targeting of high-risk defects. This provides a clear and accurate range basis for subsequent detection equipment to focus on detection, improving the targeting and efficiency of the detection.

[0041] like Figure 2 As shown, preferably, the initial attribute information includes the spatial size and geometry of the suspicious area and the defect type of the suspicious area. The dynamic configuration of the detection parameters of the second detection device includes: mapping the required scanning path density and resolution of the second detection device according to the spatial size and geometry of the suspicious area; matching the optimal detection mode from a variety of pre-stored detection modes of the second detection device according to the first defect type; binding the scanning path density, resolution and detection mode to the physical coordinates of the suspicious area, generating control commands and sending them to the second detection device, so that the second detection device performs focused detection on the suspicious area based on the control commands.

[0042] The scanning path density refers to the number of scanning paths per unit area within the suspected region, directly affecting the ability to capture the details of defects within the area; higher density results in more thorough detail capture. Resolution refers to the detection resolution of the second detection device, i.e., the smallest defect size or smallest signal change that the device can distinguish, a core parameter ensuring the accuracy of defect quantification. Detection mode refers to the preset working modes of the second detection device for different types of rail cracks, such as a "high-frequency shallow detection mode" for surface cracks and a "low-frequency deep detection mode" for internal cracks. Different modes correspond to different combinations of signal transmission and reception parameters. Control commands include structured instructions for the scanning path density, resolution, detection mode, and physical coordinates of the suspected region, used to drive the device to accurately execute focused detection.

[0043] Specifically, the formula for calculating the scan path density mapping is:

[0044]

[0045] in, For scanning path density; This is the path density coefficient, set according to the accuracy level of the detection equipment, with a value range of 1.2-2.0; L represents the width of the suspicious region; L represents the length of the suspicious region. The smaller the width and the shorter the length, the higher the density. The resolution mapping calculation formula is:

[0046]

[0047] in, The detection resolution of the second detection device; The minimum identifiable crack width is set according to rail safety standards, such as 0.1 mm. This is a safety margin, used to reserve redundancy for detection accuracy.

[0048] For example, the initial attribute information of the suspicious area is as follows: the spatial dimensions are 50mm in length and 8mm in width, the geometric shape is elongated, and the defect type is a transverse surface crack. Based on the spatial dimensions and elongated geometric shape, the scanning path density of the second detection device is mapped to 1.8 lines / mm and the detection resolution is 0.08mm. Then, based on the defect type of transverse surface crack, the optimal "high-frequency shallow detection mode" is matched from the various detection modes pre-stored by the second detection device. Subsequently, the above scanning path density, resolution, and high-frequency shallow detection mode are bound with the rail mileage and transverse position coordinates corresponding to the suspicious area to generate a structured control command and send it to the second detection device. After receiving the command, the second detection device accurately aligns with the suspicious area to perform focused detection.

[0049] In a preferred embodiment of the present invention, the parameters of the second detection device are precisely configured based on the spatial size, geometry and defect type of the suspicious area, so as to achieve focused detection of the suspicious area, greatly reduce redundant data and improve the targeting, accuracy and efficiency of the detection.

[0050] Preferably, the second diagnostic result includes a second defect type, a second confidence level, and defect quantification parameters. The generation of the second diagnostic result includes: extracting features from the second modality data to obtain deep features related to the precise quantification of the defect; inputting the deep features into a pre-trained second defect classification and regression model to obtain the second defect type and defect quantification parameters; and calculating the second confidence level of the second diagnostic result based on the classification probability of the deep features in the second defect classification and regression model and the signal-to-noise ratio of the second modality data during the focusing detection process.

[0051] Among them, depth features refer to the deep abstract features extracted from the second modality data that can accurately characterize the three-dimensional physical properties of cracks. These directly support the accurate identification of defect types and the precise quantification of size parameters, including but not limited to the three-dimensional contour gradient of the crack, cross-sectional morphology parameters, and depth-related signal features. The second defect classification and regression model refers to an intelligent model pre-trained based on a large amount of rail crack sample data. It has dual functions of classification and regression, capable of both determining the defect type and simultaneously outputting the specific size parameters of the crack, without requiring additional separate modeling. Defect quantification parameters refer to the core indicators that accurately describe the physical size of the crack and are the key basis for assessing the severity of the crack. These include, but are not limited to, the actual length, maximum width, and penetration depth of the crack, which are quantitative data that can be directly used for engineering judgment.

[0052] Specifically, the formula for calculating the second confidence level is:

[0053]

[0054] in, For the second confidence level, 0 ≤ ≤1; The weighting coefficients for classification probabilities, 0 < <1, Prioritize defect type identification and data quality settings; The classification probability of the second defect type output by the regression model and the second defect classification; The signal-to-noise ratio of the second modality data during the focusing and detection process; It is a normalization constant used to... Mapping to the 0-1 range ensures a consistent computational dimension.

[0055] Specifically, after acquiring the second modality data, deep feature extraction is first performed to obtain deep abstract features that can accurately characterize the three-dimensional physical properties of the crack. Then, the deep features are input into the pre-trained second defect classification and regression model, and the model outputs the second defect type and the corresponding defect quantification parameters simultaneously. Subsequently, the classification probability of the second defect type given by the model is extracted, and combined with the signal-to-noise ratio of the second modality data in the focusing detection process, the second confidence level is obtained through weighted fusion calculation. Finally, a second diagnostic result containing the second defect type, defect quantification parameters and the second confidence level is formed.

[0056] In a preferred embodiment of the present invention, by extracting deep features and utilizing a pre-trained second defect classification and regression model, the defect type is accurately identified and the quantitative parameters are precisely obtained. The classification probability and data signal-to-noise ratio are fused to improve the diagnostic confidence, providing a high-quality second diagnostic result for subsequent consistency verification and fusion.

[0057] Preferably, the consistency verification and fusion includes: spatiotemporally aligning the relevant diagnostic features in the first and second diagnostic results based on the physical coordinates of the suspicious area, and establishing the correlation between the features to form a preliminary fusion evidence set; performing consistency verification on the defect type judgments from the first and second diagnostic results in the preliminary fusion evidence set; generating a collaborative verification instruction when the verification results are consistent; generating a priority verification instruction pointing to the high-confidence diagnostic result based on the comparison result of the first confidence level and the second confidence level; and calculating the joint confidence level of the fusion diagnostic evidence based on the first confidence level and the second confidence level using a weighted average algorithm.

[0058] Spatiotemporal alignment refers to matching and calibrating the core information describing the same defect in the first and second diagnostic results based on the physical coordinates and detection sequence of the suspected area. This eliminates fusion interference caused by spatial location deviations and temporal misalignments, ensuring that both are used to verify the same detection object. Relevant diagnostic features refer to the set of core information directly related to defect determination in the first and second diagnostic results, including defect type, location coordinates, key feature parameters, and confidence level. This is the core data object for consistency verification and fusion. The preliminary fusion evidence set is a temporary structured data set formed by integrating the effective relevant diagnostic features from the first and second diagnostic results after spatiotemporal alignment. Consistency verification refers to judging the matching degree of the core dimensions of the first and second diagnostic results in the preliminary fusion evidence set, verifying whether the conclusions of the two modalities are consistent. This is a key step in resolving diagnostic conflicts between modalities. The collaborative verification instruction refers to the identification instruction generated when the defect type judgments of the first and second diagnostic results are consistent. This is used to confirm that the two modalities have reached a consensus, and the weight of this consensus information is strengthened during subsequent fusion processes. The priority verification instruction is a guiding instruction generated when the defect type judgments of the first and second diagnostic results conflict. It clearly states that the diagnostic result with higher confidence is used as the core of fusion, and the effective diagnostic information of the high-confidence modality is retained first.

[0059] For example, using the rail mileage and lateral position of the suspected area as a benchmark, the relevant diagnostic features such as defect type, location, and confidence in the first diagnostic result (surface lateral crack, confidence level 0.92, signal-to-noise ratio 28dB) and the second diagnostic result (surface lateral crack, confidence level 0.95, signal-to-noise ratio 32dB) are spatiotemporally aligned to establish a correlation and form a preliminary fusion evidence set. The consistency of the defect type judgments of the two is checked, and a collaborative verification instruction is generated because the results are consistent. Then, the weights are calculated based on the signal-to-noise ratio. The weight of the first diagnostic result is 28 / (28+32)=0.47, and that of the second is 32 / (28+32)=0.53. The joint confidence is calculated using a weighted average algorithm as 0.47×0.92 + 0.53×0.95≈0.94. Finally, fusion diagnostic evidence containing surface lateral crack, unified location information, and a joint confidence of 0.94 is generated.

[0060] In a preferred embodiment of the present invention, the misalignment interference of two-modal diagnostic information is eliminated by spatiotemporal alignment, diagnostic conflicts are resolved by consistency verification, and highly reliable fusion diagnostic evidence and joint confidence are generated by combining data quality weighted fusion, providing accurate and credible basis for subsequent parameter configuration and verification testing of third detection equipment.

[0061] Preferably, the configuration process of the detection parameters of the third detection device includes: comparing the joint confidence level with the preset interval, and determining the verification mode of the third device based on the comparison result; extracting the estimated depth range and spatial orientation of the defect from the fused diagnostic evidence, and mapping them to the detection parameters of the third detection device; binding the determined verification mode with the detection parameters, generating control commands and sending them to the third detection device to configure the detection parameters of the third detection device.

[0062] The verification mode refers to the preset verification working mode of the third-party detection equipment for different confidence levels of fused diagnostic evidence. The core differences lie in the detection precision, time consumption, and parameter combination, such as "rapid review mode, high confidence, efficient verification", "fine verification mode, medium confidence, balancing accuracy and efficiency", and "comprehensive investigation mode, low confidence, in-depth verification". The estimated depth range refers to the possible depth range of the crack penetrating the thickness of the rail, extracted from the fused diagnostic evidence, such as 0.5-2mm. Spatial orientation refers to the extension direction of the crack on and inside the rail, such as longitudinal extension, transverse extension, and oblique extension, used to guide the direction planning of the scanning path of the third-party detection equipment. Detection parameters refer to the core working parameters adapted to the verification mode of the third-party detection equipment, including but not limited to detection frequency, signal transmission power, scanning step size, and data sampling density, which directly determine the accuracy and efficiency of the verification detection.

[0063] Specifically, the formula for calculating the detection frequency is:

[0064]

[0065] in, The detection frequency of the third detection device; This is the frequency coefficient, set according to the type of detection equipment, with a value range of 5-8MHz•mm; This is the average value for the estimated depth range.

[0066] Specifically, assuming the joint confidence level of the fusion diagnostic evidence is 0.93, the preset intervals are 0.9-1.0 (high confidence), 0.7-0.9 (medium confidence), and 0-0.7 (low confidence). After comparison, the verification mode of the third detection device is determined to be "rapid verification mode". The defect prediction depth range of 1-3mm and the spatial orientation of lateral extension along the rail are extracted from the fusion diagnostic evidence. The average predicted depth of 2mm is substituted into the formula to calculate the detection frequency. k=6MHz•mm, f=6 / 2=3MHz, and the scanning path direction is planned in combination with the lateral orientation to obtain the detection parameters. Finally, the "rapid verification mode" is bound to the detection parameters, control commands are generated and sent to the third detection device to complete the parameter configuration.

[0067] The preferred embodiment of the present invention provides a joint confidence matching verification mode that combines the defect prediction depth range with the spatial orientation to accurately map detection parameters, thereby achieving adaptive configuration of the parameters of the third detection device, improving the targeting, efficiency and accuracy of verification detection, and providing reliable parameter support for verification detection of suspicious areas.

[0068] like Figure 3 As shown, preferably, the verification process of the third diagnostic result includes: type consistency verification: comparing the defect type in the third diagnostic result with the dominant defect type in the fusion diagnostic evidence to determine whether the types are consistent; confidence dominance verification: if the types are consistent, comparing the confidence of the third diagnostic result with the joint confidence to determine whether the former is higher than the latter and the difference exceeds a preset dominance threshold; confidence arbitration comparison: if the types are inconsistent or the confidence difference does not exceed the dominance threshold, comparing the joint confidence with the confidence of the third diagnostic result.

[0069] Further preferably, the step of generating the final test result based on the verification result includes: if the types are consistent and the confidence comparison meets the advantage condition, the third diagnostic result is taken as the final conclusion; if the confidence arbitration comparison is entered, the diagnostic result corresponding to the one with higher confidence is taken as the final conclusion; if the difference in confidence between the two in the confidence arbitration comparison is less than the preset tolerance threshold, an uncertain conclusion is generated and marked as requiring manual review.

[0070] Among them, the dominant defect type refers to the defect type with the highest reliability determined after consistency verification and fusion of the fused diagnostic evidence, and serves as the benchmark for type consistency verification. The preset advantage threshold refers to the confidence difference benchmark set based on detection accuracy requirements, used to determine whether the confidence of the third diagnostic result has a significant advantage relative to the joint confidence (e.g., 0.05). The preset tolerance threshold refers to the critical value of the confidence difference set based on acceptable engineering error (e.g., 0.03), used to determine whether the two confidence levels are in a state of near indistinguishable superiority. An uncertain conclusion indicates a diagnostic result generated when the difference between the two confidence levels is less than the tolerance threshold, indicating that the existing data cannot accurately determine the reliability of the conclusion, and human intervention is required.

[0071] Specifically, if the dominant defect type of the fused diagnostic evidence is a transverse surface crack, with a joint confidence level of 0.94, a preset dominance threshold of 0.05, and a tolerance threshold of 0.03, after generating a third diagnostic result based on the third modality data, the defect type is first compared with the dominant type: if it is a transverse surface crack and the third confidence level is 0.99, and the confidence difference is 0.05 exceeding the dominance threshold, then the result is adopted; if it is an internal longitudinal crack and the third diagnostic result confidence level is 0.92, the types are inconsistent and arbitration is initiated, and the fused conclusion is adopted because 0.92 < 0.94; if it is a transverse surface crack and the third diagnostic result confidence level is 0.95, and the confidence difference is 0.01 less than the tolerance threshold, then an uncertain conclusion is generated and a manual review mark is added, finally outputting a reliable or manually reviewed test result.

[0072] In a preferred embodiment of the present invention, the consistency verification of the type of the third diagnostic result and the fusion diagnostic evidence, the judgment of confidence advantage, and the arbitration mechanism are used to efficiently adopt highly reliable diagnostic conclusions, identify difficult scenarios through tolerance thresholds, and accurately mark the areas for manual review. This effectively avoids misjudgment and reduces ineffective manual intervention, and finally outputs a scientific, credible, highly accurate rail crack detection conclusion that is adapted to the actual needs of engineering.

[0073] In summary, the rail crack detection method based on multimodal fusion provided by this invention generates a first diagnostic result through preliminary detection, calculates a severity score to identify suspicious areas, precisely configures the parameters of a second detection device based on regional attributes to achieve focused detection, extracts depth features and uses a pre-trained model to generate a second diagnostic result, performs spatiotemporal alignment and fusion of the two modal diagnostic information to form highly reliable fusion evidence, adaptively configures the parameters of a third detection device to complete verification detection, and finally outputs the final conclusion through type verification, confidence judgment, and arbitration mechanism. This entire process design significantly improves the targeting and efficiency of rail crack detection, effectively eliminates information bias between modalities, resolves diagnostic conflicts, efficiently adopts highly reliable conclusions, accurately identifies difficult scenarios and marks areas for manual review, avoids misjudgment and ineffective manual intervention, and ultimately achieves accurate identification, quantitative assessment, and scientific judgment of rail cracks, fully adapting to the actual safety inspection needs of engineering projects.

[0074] The above description is merely a preferred embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A rail crack detection method based on multimodal fusion, characterized in that, Includes the following steps: Step S1: Scan the rail using the first detection device to obtain the first modal data; A preliminary diagnosis of the rail is performed based on the first modal data to generate a first diagnostic result, which includes a set of potential defect areas, a first defect type and its corresponding first confidence level; based on the first diagnostic result, suspicious areas on the rail are located and the initial attribute information of the suspicious areas is extracted. Step S2: Based on the initial attribute information, dynamically configure the detection parameters of the second detection device to focus on the detection of suspicious areas and obtain second modal data; diagnose the rail based on the second modal data and generate a second diagnostic result; the second diagnostic result includes a second defect type, a second confidence level, and defect quantification parameters; Step S3: Perform consistency verification and fusion of the first diagnostic result and the second diagnostic result to generate fused diagnostic evidence and joint confidence level; The consistency verification and fusion includes: spatiotemporally aligning the relevant diagnostic features in the first and second diagnostic results based on the physical coordinates of the suspicious area, establishing the correlation between the features, and forming a preliminary fusion evidence set; performing consistency verification on the defect type judgments from the first and second diagnostic results in the preliminary fusion evidence set; generating a collaborative verification instruction when the verification results are consistent; generating a priority verification instruction pointing to the high-confidence diagnostic result based on the comparison result of the first confidence level and the second confidence level; calculating the joint confidence level of the fusion diagnostic evidence based on the first confidence level and the second confidence level using a weighted average algorithm; the fusion diagnostic evidence refers to a structured data set formed by integrating effective diagnostic information from various modalities after consistency verification. Step S4: Configure the detection parameters of the third detection device based on the fused diagnostic evidence and joint confidence level; the third detection device performs verification detection on the suspicious area based on the configured detection parameters to obtain third modality data; Step S5: Generate a third diagnostic result based on the third modality data, and verify it by combining the fused diagnostic evidence and joint confidence. Generate the final detection result based on the verification result.

2. The rail crack detection method according to claim 1, characterized in that, The generation of the first diagnostic result includes: Based on the first modality data, all defective regions are identified using an anomaly detection algorithm, and the physical coordinates of each defective region are labeled to form a set of potential defective regions. Feature extraction is performed on each defect region in the set of potential defect regions to obtain feature parameters related to the crack; The feature parameters of each defect region are matched with the preset defect feature template library, the matching degree of each template is calculated, and the defect type corresponding to the template with the highest matching degree is taken as the first defect type of the defect region. The signal-to-noise ratio of the first modal data is calculated, and the matching degree of the template corresponding to the first defect type is fused to obtain the first confidence level.

3. The rail crack detection method according to claim 2, characterized in that, The location of the suspicious area includes: calculating the severity score of each defect area based on the first defect type and its first confidence level of each defect area in the set of potential defect areas, comparing it with a preset threshold, and filtering out defect areas with scores higher than the threshold; and determining the spatial range defined by the outer rectangle of the filtered defect areas as the suspicious area.

4. The rail crack detection method according to claim 2, characterized in that, The initial attribute information includes the spatial dimensions and geometry of the suspicious area, and the defect type of the suspicious area. The dynamic configuration of the detection parameters of the second detection device includes: Based on the spatial dimensions and geometry of the suspected area, the required scanning path density and resolution of the second detection device are mapped. Based on the first defect type, the optimal detection mode is matched from a variety of pre-stored detection modes in the second detection device; The scanning path density, resolution, and detection mode are bound to the physical coordinates of the suspicious area, control commands are generated and sent to the second detection device, so that the second detection device performs focused detection on the suspicious area based on the control commands.

5. The rail crack detection method according to claim 1, characterized in that, The generation of the second diagnostic result includes: Feature extraction is performed on the second modality data to obtain deep features related to the accurate quantification of defects; The deep features are input into a pre-trained second defect classification and regression model to obtain the second defect type and defect quantification parameters. The second confidence level of the second diagnostic result is calculated based on the classification probability of the deep features in the second defect classification and regression model and the signal-to-noise ratio of the second modal data in the focusing detection process.

6. The rail crack detection method according to claim 1, characterized in that, The configuration process for the detection parameters of the third detection device includes: By comparing the joint confidence level with the preset interval, the verification mode of the third device is determined based on the comparison results. The estimated depth range and spatial orientation of the defects are extracted from the fused diagnostic evidence and mapped to the detection parameters of the third detection device. The determined verification mode is bound to the detection parameters, control commands are generated and sent to the third detection device to configure the detection parameters of the third detection device.

7. The rail crack detection method according to claim 1, characterized in that, The verification process for the third diagnostic result includes: Type consistency verification: Compare the defect type in the third diagnostic result with the dominant defect type in the fusion diagnostic evidence to determine whether the types are consistent; Confidence advantage verification: If the types are the same, the confidence of the third diagnostic result is compared with the joint confidence to determine whether the former is higher than the latter and the difference exceeds the preset advantage threshold. Confidence Arbitration Comparison: If the types are inconsistent or the confidence difference does not exceed the advantage threshold, the joint confidence is compared with the confidence of the third diagnostic result.

8. The rail crack detection method according to claim 7, characterized in that, The step of generating the final test result based on the verification results includes: if the types are consistent and the confidence comparison meets the advantage condition, the third diagnostic result is taken as the final conclusion; if the confidence arbitration comparison is entered, the diagnostic result corresponding to the one with higher confidence is taken as the final conclusion; if the difference in confidence between the two in the confidence arbitration comparison is less than the preset tolerance threshold, an uncertain conclusion is generated and marked as requiring manual review.

Citation Information

Patent Citations

  • Defect detection method and system

    CN120891006A

  • Inspection method

    US20120130666A1