Crack detection method and system for double-block sleeper
Through the fusion of multi-view image, ultrasonic and infrared data, combined with target filters and deep learning models, efficient and accurate crack detection of double-block sleepers is achieved, solving the problems of low detection efficiency and poor effectiveness in the prior art, and improving the accuracy of detection and the robustness of the system.
Patent Information
- Application Number
- CN202510530203.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art is inefficient and poor in detecting cracks in double-block sleepers. Especially in the spatial fusion of three-dimensional data and two-dimensional detection data, it is difficult to accurately establish the three-dimensional correspondence between surface deformation and internal cracks, resulting in insignificant cracks and internal cracks of the sleepers being difficult to detect.
By acquiring multi-view images, ultrasonic detection data and infrared thermal imaging data of double-block sleepers, the initial three-dimensional reconstruction model is generated using the pre-trained multi-view fusion model, combining ultrasonic and infrared data for multimodal feature fusion, the target filter is used to extract crack features of different scales, and the deep learning model is used to detect and evaluate real cracks.
It realizes efficient and accurate crack detection of double-block sleepers, can identify various cracks, improves detection efficiency and accuracy, and generates detailed inspection reports by evaluating the degree of damage, ensuring the safety and reliability of the track.
Smart Images

Figure CN120064299A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of rail transit disease detection, and in particular, to a method and system for crack detection of double-block sleepers. Background Art
[0002] The double-block sleeper is a new type of sleeper form, mainly composed of two independent concrete blocks, which are connected by an elastic cushion plate in the middle.
[0003] This design can not only effectively disperse the train load, improve the stability and comfort of the track, but also reduce the manufacturing and maintenance costs of the sleepers.
[0004] Therefore, double-block sleepers are widely used in high-speed railways and urban rail transit systems.
[0005] However, in order to ensure the quality and performance of double-block sleepers, it is particularly important to detect cracks in a timely manner.
[0006] At present, the more advanced detection scheme uses the method of fusing laser three-dimensional scanning and ultrasonic detection data to inspect the sleepers and identify the cracks in the sleepers.
[0007] In the current method, during the spatial fusion process of three-dimensional data and two-dimensional detection data, due to the lack of in-depth feature interaction of multi-modal data, it is difficult to accurately establish the three-dimensional correspondence between surface deformation and internal cracks, resulting in low detection efficiency, and it is difficult to detect non-obvious cracks and internal cracks in the sleepers. Therefore, there is an urgent need for a crack detection method for double-block sleepers with high efficiency and the ability to effectively identify various cracks. Summary of the Invention
[0008] The embodiments of the present application provide a method and system for crack detection of double-block sleepers to solve the problems of low crack detection efficiency and poor effectiveness existing in the prior art.
[0009] In a first aspect, the embodiments of the present application provide a method for crack detection of double-block sleepers, including: Obtaining multi-view images, ultrasonic detection data, and infrared thermal imaging data of the double-block sleeper; Inputting the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the double-block sleeper; Fusing the ultrasonic detection data, infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; Using a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identifying and marking suspected crack regions; Use a deep learning model to detect real cracks in the suspected crack area, and obtain the crack position, crack size and shape of the real cracks; In the case of multiple real cracks, analyze the crack distribution, and evaluate the damage degree of the double-block sleeper according to the crack position, crack size and shape of each real crack, as well as the crack distribution, and generate a crack detection report of the double-block sleeper according to the damage degree evaluation result.
[0010] Optionally, the fusion of the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model includes: Preprocess the ultrasonic detection data to obtain preprocessed ultrasonic detection data. Based on the preprocessed ultrasonic detection data, use wavelet transform combined with multi-resolution analysis to calculate the wavelet coefficients of each layer to obtain feature information at different scales, generate a feature extraction result, and based on the feature extraction result, apply short-time Fourier transform to extract the time-frequency domain features of the ultrasonic detection data to obtain a time-frequency domain feature map; Perform temperature correction on the infrared thermal imaging data to obtain corrected infrared thermal imaging data, and use clustering analysis technology to identify abnormal points in the infrared thermal imaging data to obtain an anomaly detection result; According to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model, perform multi-modal feature fusion through feature stitching technology and attention mechanism to generate an intermediate three-dimensional reconstruction model; According to the time-frequency domain feature map and the anomaly detection result, perform decision-level fusion by combining ensemble learning and Bayesian technology to obtain a decision-level fusion result, and optimize the intermediate three-dimensional reconstruction model according to the decision-level fusion result to obtain a target three-dimensional reconstruction model.
[0011] Optionally, the performing multi-modal feature fusion through feature stitching technology and attention mechanism according to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model to generate an intermediate three-dimensional reconstruction model includes: Adopt an attention mechanism, combine the historical attention weights of each modal feature, and calculate the initial attention weights of the modal features corresponding to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model respectively; Adjust the attention weights of each modal feature according to the mutual relationship between different modal features to obtain the target attention weights of each modal feature; Adopt the feature splicing technology, and perform weighted splicing processing on the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model according to the target attention weights of each modality feature to obtain a multi-modal weighted feature vector, and generate an intermediate three-dimensional reconstruction model according to the multi-modal weighted feature vector.
[0012] Optionally, the decision-level fusion is performed according to the time-frequency domain feature map and the anomaly detection result, combining the ensemble learning and Bayesian techniques to obtain a decision-level fusion result, and the intermediate three-dimensional reconstruction model is optimized according to the decision-level fusion result to obtain a target three-dimensional reconstruction model, including: Construct multiple base classifiers based on ensemble learning, and use different machine learning algorithms to train each base classifier respectively according to the time-frequency domain feature map and the anomaly detection result to generate the prediction results of each base classifier; Calculate the prior probability and likelihood function of each base classifier, use the Bayesian algorithm to calculate the posterior probability, and obtain the decision-level fusion result according to the posterior probability and the prediction results of each base classifier; Adjust the parameters of the intermediate three-dimensional reconstruction model according to the decision-level fusion result to obtain the target three-dimensional reconstruction model corresponding to the adjusted parameters.
[0013] Optionally, the deep learning model includes an object detection model and a semantic segmentation model. The use of the deep learning model to detect real cracks in the suspected crack area to obtain the crack position, crack size and shape of the real cracks includes: Use a pre-trained object detection model and combine an irregular recognition frame to determine the boundary data of the real crack; According to the boundary data of the real crack, combine a pre-trained semantic segmentation model to perform semantic segmentation on the image in the suspected crack area to obtain a pixel-level mask of the real crack; Based on the pixel-level mask of the real crack, determine the shape of the real crack, and calculate the pixel-level length and width of the real crack in the image in the suspected crack area; According to the pixel-level length and width of the real crack in the image in the suspected crack area, combine the conversion relationship and proportional relationship between the image three-dimensional coordinate system and the space coordinate system to determine the actual length and width of the real crack.
[0014] Optionally, the use of a pre-trained object detection model and combining an irregular recognition frame to determine the boundary data of the real crack includes: Use a pre-trained object detection model to detect the boundary of the real crack in the suspected crack area through an irregular recognition frame; Calculate the position offset value of the irregular recognition box using the regression bias function, determine the position perception degree of the irregular recognition box using the spatial attention mechanism, determine the weighted sum result of the position offset value and the position perception degree as the position difference, and use the position difference to correct the boundary of the real crack to obtain the boundary data of the real crack.
[0015] Optionally, according to the boundary data of the real crack, combined with a pre-trained semantic segmentation model, perform semantic segmentation on the image in the suspected crack area to obtain a pixel-level mask of the real crack, including: According to the boundary data of the real crack, use a pre-trained semantic segmentation model to classify each pixel in the image in the suspected crack area to achieve semantic segmentation of the image in the suspected crack area, and obtain a pixel-level label map. Each pixel-level mask in the pixel-level label map is used to indicate whether the corresponding pixel is marked as a crack or non-crack; Adjust the pixel-level masks in the pixel-level label map to obtain an adjusted pixel-level label map to make the real crack continuous and complete; Combine crack prior knowledge and the context information of the real crack to optimize the adjusted pixel-level label map to obtain a pixel-level mask of the real crack.
[0016] In a second aspect, an embodiment of the present application provides a crack detection system for a double-block sleeper, including: An acquisition module for acquiring multi-view images, ultrasonic detection data, and infrared thermal imaging data of the double-block sleeper; A generation module for inputting the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the double-block sleeper; A fusion module for fusing the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; An extraction and recognition module for using a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identifying and marking the suspected crack area; A detection module for using a deep learning model to detect real cracks in the suspected crack area to obtain the crack position, crack size, and shape of the real cracks; An evaluation and generation module for analyzing the crack distribution in the case of multiple real cracks, and evaluating the damage degree of the double-block sleeper according to the crack position, crack size, and shape of each real crack, and the crack distribution, and generating a crack detection report for the double-block sleeper according to the damage degree evaluation result.
[0017] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a method for detecting cracks in a double-block sleeper as described in any one of the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which when executed by a computer, implements a method for detecting cracks in a double-block sleeper as described in any one of the first aspect.
[0019] An embodiment of the present application provides a method for detecting cracks in a double-block sleeper, the method including: acquiring multi-view images, ultrasonic detection data, and infrared thermal imaging data of the double-block sleeper; inputting the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the double-block sleeper; fusing the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; using a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identifying and marking suspected crack regions; using a deep learning model to detect real cracks in the suspected crack regions to obtain the crack positions, crack sizes, and shapes of the real cracks; in the case of multiple real cracks, analyzing the crack distribution, and according to the crack positions, crack sizes, and shapes of each real crack, and the crack distribution, evaluating the damage degree of the double-block sleeper, and generating a crack detection report of the double-block sleeper according to the damage degree evaluation result.
[0020] The embodiments of the present application realize high-precision and all-round detection of sleeper cracks by integrating multi-modal data and combining deep learning and three-dimensional reconstruction technologies. Specifically, this method can not only generate a detailed initial three-dimensional reconstruction model, but also accurately identify suspected crack regions through multi-scale fusion and feature extraction, and further accurately locate the positions, sizes, and shapes of real cracks. In addition, the embodiments of the present application can comprehensively evaluate the damage degree of double-block sleepers by comprehensively analyzing the distribution of multiple cracks, thereby providing a scientific basis for maintenance decisions and ensuring the safety and reliability of railway infrastructure. The finally generated crack detection report not only improves the detection efficiency and accuracy, but also enhances the robustness and adaptability of the double-block sleeper crack detection system, and is applicable to complex and changeable actual application scenarios.
[0021] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of a method for detecting cracks in a double-block sleeper provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a system for detecting cracks in a double-block sleeper provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art of this technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application.
[0025] In some processes described in the specification, claims and the above-mentioned accompanying drawings of the present application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The serial numbers of the operations, such as 11, 12, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0027] A method for detecting cracks in a double-block sleeper provided in this embodiment can be carried out under exemplary environmental conditions, equipment conditions and detection conditions.
[0028] Among them, the environmental conditions include light conditions, temperature conditions, humidity conditions, and cleanliness conditions. Specifically, the setting of the light conditions in the embodiments of the present application can ensure sufficient and uniform light when collecting multi-view images, avoiding the influence of strong light direct irradiation or shadows on the image quality of the multi-view images. The setting of the temperature conditions in the embodiments of the present application can ensure the collection of infrared thermal imaging data under stable temperature conditions, avoiding the distortion of thermal imaging data caused by excessive temperature differences. The setting of the humidity conditions in the embodiments of the present application can keep the detection environment dry and avoid the influence of moisture on the ultrasonic detection data. The setting of the cleanliness conditions in the embodiments of the present application can ensure the cleanliness of the sleeper surface, without obvious dust, oil stains and other impurities, so as to ensure the accuracy of the multi-view images and the ultrasonic detection data.
[0029] The equipment conditions include a multi-view camera, an ultrasonic detection device, an infrared thermal imager, a computing device, and 3D reconstruction software. Specifically, the multi-view camera in the embodiments of the present application may refer to a high-resolution camera, which can capture images of the double-block sleeper from multiple perspectives. The ultrasonic detection device may refer to a high-precision ultrasonic flaw detector, which can accurately detect the defects inside the double-block sleeper. The infrared thermal imager may refer to a high-sensitivity infrared thermal imager, which can capture the temperature distribution on the sleeper surface. The computing device may refer to a high-performance computer, equipped with sufficient storage space and computing power for image processing and data analysis. The 3D reconstruction software may refer to 3D reconstruction software that supports multi-view fusion, and is used to generate an initial 3D reconstruction model of the double-block sleeper.
[0030] The detection conditions can be set for the detection of a single-block sleeper, or for the detection of four consecutive single-block sleepers to be deployed on one track. Therefore, when performing crack detection on single-block sleepers in the embodiments of the present application, the detection can be carried out under two different detection conditions. In this embodiment, crack detection can be performed only once under one of the detection conditions, or twice under both detection conditions. Among them, the first detection condition is to consider only one single-block sleeper. The second detection condition is to consider four consecutive single-block sleepers on each track. Therefore, in the second detection condition, an experimental track can be set up in this embodiment, and crack detection of four single-block sleepers can be carried out with four single-block sleepers as a unit, thereby improving the accuracy of single-block sleeper detection. The reasons are as follows: (1) Considering the combined action of multiple single-block sleepers in four consecutive single-block sleepers can provide a more comprehensive perspective to evaluate the stability of the track. (2) By inspecting four consecutive single-block sleepers, it is possible to better evaluate whether the connections between multiple single-block sleepers and between single-block sleepers and the rails are uniform, which is very important for ensuring the smooth operation of trains. (3) Four consecutive single-block sleepers as a unit, their condition directly affects the overall structural strength of the track. Damage to any single-block sleeper may affect the function of the entire unit and thus the safety of the track. (4) Detecting four consecutive single-block sleepers can help identify potential problems in advance, such as wear, cracks or other forms of damage, so as to take preventive measures to avoid greater failures.
[0031] It should be noted that in both the first detection condition and the second detection condition, the following crack detection method for single-block sleepers can be adopted in the embodiments of the present application.
[0032] Figure 1 The flowchart of a crack detection method for single-block sleepers provided by the embodiments of the present application is shown in Figure 1 As shown, the method includes: S11. Obtain multi-view images, ultrasonic detection data, and infrared thermal imaging data of the single-block sleeper.
[0033] Among them, the multi-view images include multiple different view images, and the view images include: visible light image feature data. In the first detection condition, various data of a single single-block sleeper are obtained. In the second detection condition, various data of 4 single-block sleepers are obtained each time.
[0034] S12. Input the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the single-block sleeper.
[0035] It should be understood that, in order to cope with different detection situations, the embodiments of the present application can provide a multi-view fusion model for the detection situation. Specifically, in the first detection situation, the multi-view fusion model adopted by the embodiments of the present application is for a single double-block sleeper track; in the second detection situation, the multi-view fusion model adopted by the embodiments of the present application is for 4 double-block sleeper tracks.
[0036] Exemplarily, the multi-view fusion model in all detection situations can include a deep learning framework and a feature pyramid network. The following is an analysis of the deep learning framework and the feature pyramid network respectively: (1) The deep learning framework includes a convolutional layer, a residual connection, and a pooling layer. Exemplarily, the expression of the convolutional layer is: ; where is the feature map of the th perspective image.
[0037] is the convolutional kernel weight matrix, with a size of , is the side length of the convolutional kernel, is the number of input channels, is the number of output channels. is the th perspective image. is the bias vector of the th perspective image, with a size of , is the activation function, such as the ReLU function. The expression of the residual connection is: ; where is the feature map after adding the residual connection, and ResBlock is the residual block, which is used to alleviate the gradient vanishing problem in the feature map of the th perspective image. The expression of the pooling layer is: ; where is the pooling feature map obtained based on the feature map after adding the residual connection, and pool is the pooling operation, such as max pooling or average pooling. In the embodiments of the present application, through the convolutional layer, the convolutional kernel can be used to extract the local features of the feature image, which helps to capture the subtle changes on the surface of the double-block sleeper. Through the residual connection, the gradient vanishing problem in the deep network can be effectively solved, and the training effect of the deep learning framework can be improved. Through the pooling layer, the spatial dimension of the feature map can be reduced, the amount of calculation can be reduced, and at the same time, important feature information can be retained.
[0038] (2) The feature pyramid network can be set with a top-down path and a lateral connection; specifically, the expression of the top-down path is: ; where is the feature map after fusion of the th layer, is the feature map of the previous layer of the pooled feature map; conv is the convolution operation. For example, the convolution operation can use a convolution kernel for feature fusion to ensure that feature maps of different levels can be aligned in the channel dimension. up is the upsampling operation. For example, the upsampling operation can be implemented using bilinear interpolation or transposed convolution, and the purpose is to restore the spatial resolution so that high-level features can be added to low-level features. The expression for the lateral connection is: ; where is the feature map after the lateral connection, and conv is the convolution operation. Exemplarily, a convolution kernel can be used for feature alignment. In the embodiment of the present application, in the top-down path, high-level feature maps are fused with low-level feature maps through the upsampling operation to enhance the model's perception ability of features at different scales. Moreover, feature maps of the same scale are fused through the lateral connection to retain more detailed information.
[0039] S13. Fuse the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain the target three-dimensional reconstruction model.
[0040] Based on the target three-dimensional reconstruction model, the embodiment of the present application can comprehensively reflect whether there are crack features in the double-block sleeper by combining visible light image feature data, ultrasonic detection data, and infrared thermal imaging data.
[0041] In an optional implementation manner, in the first detection case or the second detection case, the embodiment of the present application directly executes S13 after executing S12. In another optional implementation manner, in the second detection case, after executing S12, difference analysis can be performed on the ultrasonic detection data and the infrared thermal imaging data of 4 double-block sleepers. If there are no abnormal data in the ultrasonic detection data of the 4 double-block sleepers and there are no abnormal data in the infrared thermal imaging data of the 4 double-block sleepers, it is determined that there are no cracks in the 4 double-block sleepers, and no subsequent steps are performed. Otherwise, continue to execute S13 to generate the target three-dimensional reconstruction model.
[0042] S14. Use a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identify and mark the suspected crack areas.
[0043] Among them, in step S14, the target filter is an adaptive filter or a multi-scale Gaussian filter. The adaptive filter is used to dynamically adjust the filter parameters. Specifically, in the embodiments of the present application, the parameters of the adaptive filter can be dynamically adjusted according to the material and surface characteristics of the double-block sleeper, thereby improving the robustness of crack feature extraction. At the same time, this filter can also extract crack features at different scales to ensure that cracks of different sizes and shapes are captured. The multi-scale Gaussian filter is used for multi-scale analysis. Specifically, in the embodiments of the present application, crack features can be extracted at different scales through the multi-scale Gaussian filter to ensure that cracks of different sizes and shapes are captured. At the same time, the multi-scale Gaussian filter can also extract the edge and texture features of the cracks, thereby improving the accuracy of crack recognition.
[0044] S15. Use a deep learning model to detect real cracks in the suspected crack area to obtain the crack position, crack size, and shape of the real cracks.
[0045] S16. In the case of multiple real cracks, analyze the crack distribution, and evaluate the damage degree of the double-block sleeper according to the crack position, crack size, and shape of each real crack, as well as the crack distribution. Generate a crack detection report for the double-block sleeper according to the damage degree evaluation result.
[0046] In the embodiments of the present application, compliance judgment can be performed according to a preset detection standard, and then an inspection report can be automatically issued. Therefore, step S16 includes the following steps: Step 161. Obtain a detection standard file, and identify the detection standards corresponding to each damage degree based on the detection standard file. Step 162. Determine the corresponding damage degree based on the detection standards corresponding to each damage degree identified from the detection standard file, as well as the crack position, crack size, and shape, and the crack distribution. Exemplarily, mild damage: There are few cracks, and both the length and width are within the safe range. Moderate damage: The number of cracks is large, and the length or width of some cracks is close to the critical value. Severe damage: The cracks are dense, there are long cracks or wide cracks, which may affect the structural integrity of the sleeper. Step 163. Automatically generate a crack detection report. Exemplarily, the crack detection report can briefly introduce the detection purpose, method, and standard, and can also list information such as the number, position, length, and width of all detected real cracks; display the distribution of real cracks on the sleeper, and give the damage degree and treatment opinions of the double-block sleeper. Exemplarily, the generation of the crack detection report includes two processes: damage degree evaluation and automatic report generation. Among them, the following detection standards can be used for the damage degree evaluation: ; among them, is the damage degree of the double-block sleeper is the number of cracks, is the crack length, is the crack width, , are respectively the crack number threshold, crack length threshold and crack width threshold corresponding to mild damage; are respectively the crack number threshold, crack length threshold and crack width threshold corresponding to severe damage. The embodiments of the present application can evaluate the damage degree of the double-block sleeper by setting different thresholds, comprehensively considering the crack number, length and width, and can provide a reasonable decision-making basis.
[0047] The report content includes at least one of the following: (1) Detection purpose: Ensure the safety and reliability of the sleeper. (2) Methods and standards: Describe in detail the detection methods and standards used. (3) Crack information: List all detected crack numbers, crack positions, crack lengths and widths, and crack shapes. (4) Crack distribution: Display the distribution map of cracks on the sleeper. (5) Damage degree and treatment suggestions: Put forward specific treatment suggestions according to the evaluation results. The embodiments of the present application can provide comprehensive crack detection results and treatment suggestions through detailed report content, helping maintenance personnel take measures in time to ensure the safe operation of the track.
[0048] The above process can provide a crack detection method for double-block sleepers, ensuring its effectiveness and reliability in practical applications. It not only improves the accuracy of crack detection, but also enhances the robustness and generalization ability of the model, and is applicable to various complex application scenarios.
[0049] By executing steps S11 to S16, the embodiments of the present application integrate multi-view images, ultrasonic detection data and infrared thermal imaging data, and combine a deep learning framework and a feature pyramid network to achieve high-precision and all-round detection of double-block sleeper cracks. Among them, by acquiring and fusing visible light images, ultrasonic detection data and infrared thermal imaging data, the state information of the sleeper can be captured from multiple dimensions to ensure the comprehensiveness and accuracy of crack detection. For different detection situations (single or four sleepers), corresponding multi-view fusion models are adopted to enhance the adaptability and flexibility of the system. The top-down path and the lateral connection mechanism enable the model to better capture features at different scales and enhance the recognition ability of complex crack morphologies. The introduction of the top-down path and the lateral connection explains how to fuse features at different levels through specific mathematical expressions, improving the multi-scale perception ability of the model. The multi-scale Gaussian filter extracts crack features at different scales to ensure that cracks of various sizes and shapes are captured. Based on the preset detection standards, the damage degree is automatically evaluated and a detailed crack detection report is generated, simplifying the maintenance process and improving work efficiency. Based on the preset detection standards, the damage degree is automatically evaluated and a detailed crack detection report is generated, simplifying the maintenance process and improving work efficiency.
[0050] How to accurately fuse ultrasonic detection data, infrared thermal imaging data, and the initial 3D reconstruction model is the key to obtaining an accurate target 3D reconstruction model. Exemplarily, in one possible embodiment, S13. Fusing the ultrasonic detection data, infrared thermal imaging data, and the initial 3D reconstruction model to obtain the target 3D reconstruction model includes: Step 131. Preprocess the ultrasonic detection data to obtain the preprocessed ultrasonic detection data. Based on the preprocessed ultrasonic detection data, use wavelet transform combined with multi-resolution analysis to calculate the wavelet coefficients of each layer to obtain feature information at different scales, generate a feature extraction result. Based on the feature extraction result, apply the short-time Fourier transform to extract the time-frequency domain features of the ultrasonic detection data to obtain a time-frequency domain feature map. It should be understood that the time-frequency domain feature map can provide information about the crack's variation with time and frequency, which is very useful for understanding the crack development pattern. The time-frequency domain feature map can be used as additional features and input into deep learning or other machine learning models to help improve the accuracy of crack detection. In addition, the time-frequency domain feature map can also be used to assist in the qualitative and quantitative analysis of cracks, such as evaluating the length, width, and development speed of the cracks.
[0051] This step is the ultrasonic detection data processing flow, involving three steps: data preprocessing, feature extraction, and short-time Fourier transform.
[0052] (1) Data preprocessing includes processing methods such as denoising and smoothing. Specifically, the denoising formula is as follows: ; where is the ultrasonic detection data after denoising, DWT is the discrete wavelet transform, used to remove high-frequency noise, is the ultrasonic detection data before denoising; is the smoothing coefficient, used to control the intensity of median filtering; is the median filtering, used to remove impulse noise in the ultrasonic detection data; is the threshold coefficient, used to control the intensity of wavelet threshold denoising; is the wavelet threshold denoising, using the wavelet threshold to remove low-frequency noise in the ultrasonic detection data. Smoothing can be done using the following formula: ; where is the ultrasonic detection data after smoothing; GaussianFilter is the Gaussian filter, used to smooth the signal; is the standard deviation of the Gaussian filter; laplacian is the Laplacian operator, used to enhance the edge information of the ultrasonic detection data; is the sharpening coefficient, used to control the intensity of the Laplacian operator; is the bilateral filtering coefficient, used to control the influence of bilateral filtering; is the bilateral filtering, using the spatial standard deviation and the range standard deviation for smoothing. In the embodiments of the present application, discrete wavelet transform, median filtering, and wavelet threshold denoising are combined to effectively remove high-frequency noise and impulse noise and improve the signal quality. The ultrasonic detection data is smoothed and edge-enhanced by a Gaussian filter, a Laplace operator, and a bilateral filter, while retaining the detail features.
[0053] (2) Feature extraction includes wavelet transform and multi-resolution analysis. Specifically, the feature extraction can adopt the following formula: ; where is the wavelet coefficient of the th layer; is the preprocessed ultrasonic detection data at time ; is the wavelet basis function, is the complex conjugate of the wavelet basis function; is the enhancement coefficient, used to control the intensity of the additional wavelet transform; is the additional wavelet transform, using different wavelet basis functions to perform feature extraction; is the multi-resolution analysis coefficient, used to control the intensity of multi-resolution analysis; is used to represent multi-resolution analysis, which can use wavelet basis functions of different scales to perform feature extraction. In the embodiments of the present application, multiple wavelet basis functions and multi-resolution analysis are combined to extract features of different scales and improve the richness of features.
[0054] (3) The short-time Fourier transform can adopt the following formula: ; where is the result of the short-time Fourier transform, used to represent the time-frequency domain features of the ultrasonic detection data, is the frequency, is the time; is the integration variable, used to represent the time, is the value of the ultrasonic signal at time , that is, used to represent the ultrasonic detection data; is the window function; is the enhancement coefficient, used to control the intensity of the additional short-time Fourier transform; is the additional short-time Fourier transform, using another window function different from to perform feature extraction; is the short-time Fourier transform coefficient, which is used to control the intensity of the short-time Fourier transform. is another window function; short_time_fourier_transform is the short-time Fourier transform, which uses different window functions to extract features. In the embodiments of the present application, time-frequency domain features are extracted through different window functions, enhancing the model's perception ability of different frequency components.
[0055] Step 132: Perform temperature correction on the infrared thermal imaging data to obtain the corrected infrared thermal imaging data, and use the clustering analysis technique to identify the abnormal points in the infrared thermal imaging data to obtain the abnormal detection result. It should be understood that the abnormal detection result can help locate the specific positions where cracks or other structural defects may exist, and the abnormal detection result is very important for guiding subsequent detailed inspections or maintenance work. At the same time, the abnormal detection result can also be used to train or adjust the machine learning model to better adapt to specific types of abnormal situations, thereby improving the generalization ability and robustness of the model. In addition, when generating the final crack detection report, the abnormal detection result can be used as an important reference basis to help evaluate the damage degree of the double-block sleeper and propose corresponding treatment suggestions.
[0056] It should be understood that the infrared thermal imaging data analysis includes the temperature correction and abnormal detection processes. Among them, the temperature correction can adopt the following formula: ; where is the corrected temperature, is the original temperature, is the ambient temperature, is the reference temperature; is the temperature compensation coefficient, is the humidity influence coefficient, humidity is the humidity correction function, which is used to consider the influence of humidity on temperature measurement; is the radiometric correction coefficient, which is used to control the intensity of radiometric correction, radiometric_correction is the radiometric correction, which uses the emissivity and the transmittance to correct the original temperature. In the embodiments of the present application, the correction is combined with the ambient temperature and the reference temperature to eliminate the influence of environmental factors; the influence of humidity on temperature measurement is considered to improve the accuracy of temperature correction. The radiometric correction is performed through the emissivity and the transmittance to further improve the accuracy of temperature measurement.
[0057] To implement the abnormal detection, this embodiment can adopt the clustering analysis technique, and the clustering analysis technique can adopt the following formula: ; where is the comprehensive distance between the rd and the th samples, and are the th and th dimensional features of the th and is the feature dimension, is the similarity weight, used to control the influence of cosine similarity, cosine_similarity is the cosine similarity, used to measure the directional similarity between samples, is the Mahalanobis distance weight, used to control the influence of Mahalanobis distance. mahalanobis_distance is the Mahalanobis distance, calculated using the covariance matrix This application embodiment combines Euclidean distance, cosine similarity, and Mahalanobis distance, comprehensively considering the distance and directional similarity between samples, and improving the accuracy of clustering.
[0058] Step 133: According to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model, perform multi-modal feature fusion through feature stitching technology and attention mechanism to generate an intermediate three-dimensional reconstruction model.
[0059] As a possible implementation, Step 133: According to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model, perform multi-modal feature fusion through feature stitching technology and attention mechanism to generate an intermediate three-dimensional reconstruction model, including: Step a1: Adopt the attention mechanism, combine the historical attention weights of each modal feature, and calculate the initial attention weights of the modal features corresponding to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model respectively.
[0060] Step a2: Adjust the attention weights of each modal feature according to the mutual relationship between different modal features to obtain the target attention weights of each modal feature.
[0061] It should be understood that multi-modal feature fusion may include an attention mechanism and feature stitching technology; among them, the attention mechanism can adopt the following formula: ; where is the target attention weight of the th modal feature, where , when represents the first modal feature corresponding to the ultrasonic detection data, when represents the second modal feature corresponding to the corrected infrared thermal imaging data, when represents the third modal feature corresponding to the visible light image feature data in the initial 3D reconstruction model, is the th modal feature scoring function, and the function expression can be dot product or weighted sum, is the parameter in the modal feature scoring function, is the context attention coefficient, which is used to control the influence of context attention, is the function for calculating the mutual relationship between different modal features, is the parameter in the mutual relationship calculation function. In the embodiments of the present application, the importance of different modal features is dynamically adjusted through the attention mechanism, the sensitivity of the model to key features is improved, and the context attention mechanism is introduced to further enhance the model's understanding of context information.
[0062] Step a3: Adopt the feature splicing technology, and perform weighted splicing processing on the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial 3D reconstruction model according to the target attention weights of each modal feature to obtain a multi-modal weighted feature vector, and generate an intermediate 3D reconstruction model according to the multi-modal weighted feature vector.
[0063] Step a3 can adopt the following formula: ; where is the multi-modal weighted feature vector, , , are the target attention weights of the first modal feature, the second modal feature, and the third modal feature respectively, , , are the first modal feature, the second modal feature, and the third modal feature respectively. In the embodiments of the present application, weighted splicing processing is performed on different modal features, which can form a multi-modal feature vector and improve the richness of the input information of the model.
[0064] By performing steps a1 to a3, the embodiments of the present application introduce an attention mechanism and a context attention mechanism, which are used to dynamically adjust the importance of different modality features, and generate a multi-modal weighted feature vector through a weighted splicing technique, having the following advantages: By calculating the initial attention weights and adjusting the target attention weights according to the mutual relationship between the modality features, the model can dynamically focus on the most important features, improving the sensitivity to key information. The embodiments of the present application introduce a context attention mechanism, enabling the model to not only focus on individual modality features but also consider the mutual relationship between different modality features, enhancing the understanding ability of complex scenarios. The embodiments of the present application generate a multi-modal weighted feature vector through weighted splicing processing, integrating data from different modalities, providing richer input information, and helping to improve the detection accuracy of the model. The calculation method of the attention weights includes the parameters in the scoring function, the context attention coefficient, and the mutual relationship calculation function, making the entire process highly controllable and adjustable. The embodiments of the present application ensure the controllability and flexibility of the entire process through specific mathematical expressions and parameter settings, are applicable to more diverse application scenarios, enhance the technical depth of the crack detection method, and also enhance the robustness and accuracy of the system, providing more reliable technical support for practical applications.
[0065] Step 134: According to the time-frequency domain feature map and the anomaly detection result, perform decision-level fusion by combining ensemble learning and Bayesian techniques to obtain a decision-level fusion result, and optimize the intermediate three-dimensional reconstruction model according to the decision-level fusion result to obtain a target three-dimensional reconstruction model.
[0066] As a possible implementation, step 134: According to the time-frequency domain feature map and the anomaly detection result, perform decision-level fusion by combining ensemble learning and Bayesian techniques to obtain a decision-level fusion result, and optimize the intermediate three-dimensional reconstruction model according to the decision-level fusion result to obtain a target three-dimensional reconstruction model, includes: Step b1: Construct multiple base classifiers based on ensemble learning. According to the time-frequency domain feature map and the anomaly detection result, use different machine learning algorithms to train each base classifier respectively to generate the prediction results of each base classifier.
[0067] It should be understood that ensemble learning can be represented as ensemble , where are the prediction results of each base classifier in the m base classifiers respectively. Combining the detection results of multiple base classifiers, the robustness and accuracy of the detection results are improved through weighted voting or Bayesian fusion.
[0068] Step b2: Calculate the prior probability and likelihood function of each base classifier, calculate the posterior probability using the Bayesian algorithm, and obtain the decision-level fusion result according to the posterior probability and the prediction results of each base classifier.
[0069] It should be understood that decision-level fusion may include ensemble learning techniques and Bayesian fusion techniques. In the embodiments of the present application, by introducing the Bayesian algorithm, the ability of the model to handle the uncertainty of the decision-level fusion results is further enhanced.
[0070] Step b3: Adjust the parameters of the intermediate three-dimensional reconstruction model according to the decision-level fusion results to obtain a target three-dimensional reconstruction model corresponding to the adjusted parameters.
[0071] By executing Step b1 to Step b3, the embodiments of the present application provide a more specific implementation method of ensemble learning and Bayesian techniques. By constructing multiple base classifiers and training them using different machine learning algorithms, different types of features and patterns can be captured, thereby improving the diversity and generalization ability of the model. This part details the role of the Bayesian algorithm in calculating prior probabilities, likelihood functions, and posterior probabilities, enhancing the model's ability to handle uncertainty. This part clarifies how to adjust the parameters of the intermediate three-dimensional reconstruction model according to the decision-level fusion results, ensuring the accuracy and adaptability of the target three-dimensional reconstruction model. In summary, this process improves the technical depth of the crack detection method, enhances the robustness and accuracy of the system, and is applicable to more diverse application scenarios. In addition, by introducing specific Bayesian algorithms and parameter adjustment steps, the entire process becomes more transparent and controllable, providing a solid foundation for subsequent research and applications.
[0072] By executing Step 131 to Step 134, the embodiments of the present application perform more in-depth preprocessing and feature extraction on ultrasonic detection data and infrared thermography data, and enhance the technical depth and practical application effect of the crack detection method through high-level multi-modal feature fusion and decision-level fusion. This method not only improves the accuracy of crack detection, but also enhances the robustness and generalization ability of the system, and is applicable to more diverse application scenarios. The attention mechanism is introduced to highlight important features, making the multi-modal data fusion more efficient and targeted. Combining ensemble learning and Bayesian techniques for decision-level fusion can comprehensively consider multiple information sources and improve the reliability and accuracy of the final decision. Optimizing the intermediate three-dimensional reconstruction model according to the decision-level fusion results to obtain a more accurate target three-dimensional reconstruction model improves the accuracy and reliability of crack detection. In addition, by increasing the recognition of time-frequency domain features and outliers, the embodiments of the present application also provide important support for the research on the development mode of cracks and damage assessment.
[0073] In a possible embodiment, the deep learning model includes an object detection model and a speech segmentation model. S15: Use the deep learning model to detect real cracks in the suspected crack area to obtain the crack position, crack size, and shape of the real cracks, including: Step 151: Use a pre-trained object detection model and combine it with an irregular recognition frame to determine the boundary data of the real crack.
[0074] As a possible implementation, Step 151: Use a pre-trained object detection model and combine it with an irregular recognition frame to determine the boundary data of the real crack, including: Step c1: Use a pre-trained object detection model to detect the boundary of the real crack in the suspected crack area through an irregular recognition frame. It should be understood that the irregular recognition frame can be a deformed frame obtained by removing the non-crack area based on a rectangular frame.
[0075] Step c2: Calculate the position offset value of the irregular recognition frame using a regression bias function, determine the position perception degree of the irregular recognition frame using a spatial attention mechanism, determine the weighted sum result of the position offset value and the position perception degree as the position difference, and use the position difference to correct the boundary of the real crack to obtain the boundary data of the real crack.
[0076] Exemplarily, the object detection model can include the design of both bounding box regression and loss function. The bounding box regression can adopt the following formula: ; where, is the boundary data of the real crack before correction, is the boundary data of the real crack after correction, is the position offset value of the irregular recognition frame, is the position perception degree of the irregular recognition frame, is the bias coefficient, which is used to control the influence of regression bias. is the spatial attention coefficient, which is used to control the influence of spatial attention.
[0077] By executing Step c1 to Step c2, the embodiment of the present application uses a regression bias function and a spatial attention mechanism to correct the prediction deviation and improve the accuracy of the irregular recognition frame.
[0078] The loss function of the object detection model can adopt the following formula: ; where, is the total loss of the deep learning model. is the classification loss, such as cross-entropy loss; is the regression loss, such as smooth L1 loss, and are the classification loss weight and the regression loss weight; is the balance loss, which is used to balance the proportion of negative samples, is the balance coefficient, which is used to control the influence of the balance loss, focal_loss is the focal loss between the true detection result and the predicted detection result, which is used to solve the problem of class imbalance. is the focal loss coefficient, which is used to control the influence of the focal loss.
[0079] In the design of the loss function of the deep learning model in the embodiments of this application, considering the balanced loss and the focal loss can ensure that the deep learning model has good generalization ability on different types of samples.
[0080] Step 152: According to the boundary data of the true crack, combined with the pre-trained semantic segmentation model, perform semantic segmentation on the image in the suspected crack area to obtain the pixel-level mask of the true crack.
[0081] As a possible implementation manner, Step 152: According to the boundary data of the true crack, combined with the pre-trained semantic segmentation model, perform semantic segmentation on the image in the suspected crack area to obtain the pixel-level mask of the true crack, including: Step d1: According to the boundary data of the true crack, use the pre-trained semantic segmentation model to classify each pixel in the image in the suspected crack area to implement semantic segmentation on the image in the suspected crack area, and obtain a pixel-level label map. Each pixel-level mask in the pixel-level label map is used to represent marking the corresponding pixel as a crack or non-crack.
[0082] Step d2: Adjust the pixel-level mask in the pixel-level label map to obtain an adjusted pixel-level label map to make the true crack continuous and complete.
[0083] Step d3: Combine the crack prior knowledge and the context information of the true crack to optimize the adjusted pixel-level label map to obtain the pixel-level mask of the true crack.
[0084] Among them, the semantic segmentation model includes the design of two aspects: pixel-level classification and the loss function of the semantic segmentation model. Exemplarily, Steps d1 to d3 are used to implement pixel-level classification, and finally obtain the pixel-level mask of the true crack. The calculation formula of the pixel-level mask of the true crack is as follows: ; Among them, is the th pixel in the image x in the suspected crack area, is the pixel-level mask of the th pixel in the image x in the suspected crack area, which is used to represent the probability that the th pixel belongs to the crack category or the non-crack category, is the adjusted pixel-level label map, is the prior knowledge coefficient, which is used to control the influence of crack prior knowledge, prior_knowledge is the prior knowledge function, which is used to introduce the crack prior knowledge of the double-block sleeper is the coefficient of the contextual information, which is used to control the influence of the contextual information, contextual_information is the contextual information, which is used to enhance the semantic segmentation model's understanding of the local context. In the embodiments of the present application, by introducing the prior knowledge function and the contextual information, the semantic segmentation model's ability to identify various types of cracks is improved
[0085] The loss function of the semantic segmentation model can adopt the following formula ; where is the segmentation loss of the semantic segmentation model is the total number of pixels , is the label of the th pixel belonging to the th category represents a crack represents non-crack is the prediction probability that the th pixel belongs to the th category is the Dice loss is the one-hot encoded vector of the true label, representing the true category of the th pixel is the predicted probability distribution vector, which contains the prediction probabilities that the th pixel belongs to each category is the Lovasz-Hinge loss, which is used to optimize the segmentation boundary is the weight of the Dice loss is the weight of the Lovasz-Hinge loss, which is used to control the influence of the Lovasz-Hinge loss. When designing the loss function in the embodiments of the present application, the Dice loss and the Lovasz-Hinge loss are comprehensively considered, and the detection accuracy of the semantic segmentation model for fine cracks is improved
[0086] By performing steps d1 to d3, the embodiments of the present application introduce a more detailed semantic segmentation process and optimization mechanism, especially in aspects such as pixel-level classification, label map adjustment, combination of prior knowledge and context information. Specifically, each pixel within the suspected crack area is classified by a pre-trained semantic segmentation model to generate a pixel-level label map, ensuring a fine distinction between cracks and non-cracks. By adjusting the pixel-level masks in the pixel-level label map, the continuity and integrity of real cracks are ensured, avoiding breaks or discontinuities. The adjustment process can remove noise and small errors, making the crack boundaries smoother and more natural. By introducing a crack prior knowledge function and context information, the model's recognition ability for different types of cracks is improved, especially for thin or complex cracks. The loss function of the semantic segmentation model comprehensively considers cross-entropy loss, Dice loss, and Lovasz-Hinge loss, improving the detection accuracy for thin cracks.
[0087] Step 153: Based on the pixel-level mask of the real crack, determine the shape of the real crack, and calculate the pixel-level length and width of the real crack in the image within the suspected crack area.
[0088] Step 154: According to the pixel-level length and width of the real crack in the image within the suspected crack area, and in combination with the conversion relationship and proportional relationship between the image three-dimensional coordinate system and the spatial coordinate system, determine the actual length and width of the real crack.
[0089] By performing steps 151 to 154, the embodiments of the present application introduce a more specific and detailed processing method, especially in aspects such as crack boundary detection, semantic segmentation, shape determination, and actual length and width calculation. Specifically, by combining a pre-trained object detection model with an irregular recognition frame, the boundary data of the real crack can be determined more accurately, avoiding errors that may be brought by traditional rectangular frames. Combining a pre-trained semantic segmentation model, semantic segmentation is performed on the image within the suspected crack area to obtain the pixel-level mask of the real crack, ensuring a refined expression of the crack position and shape. Based on the pixel-level mask of the real crack, the shape of the crack can be determined more accurately, providing reliable geometric information for subsequent analysis. According to the conversion relationship and proportional relationship between the image three-dimensional coordinate system and the spatial coordinate system, the pixel-level dimensions are converted into actual length and width, ensuring the accuracy of the measurement results. Ultimately, the technical depth of the crack detection method is improved, and the robustness and accuracy of the system are enhanced, making it applicable to more diverse application scenarios. In addition, by introducing specific processing steps and technical details, the entire process becomes more transparent and controllable, providing solid technical support for practical applications.
[0090] Figure 2 This is a schematic structural diagram of a crack detection system for a double-block sleeper provided by the embodiments of the present application, as Figure 2As shown in the figure, the system includes: An acquisition module 21, configured to acquire multi-view images, ultrasonic detection data, and infrared thermal imaging data of a double-block sleeper.
[0091] A generation module 22, configured to input the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the double-block sleeper.
[0092] A fusion module 23, configured to fuse the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model.
[0093] An extraction and recognition module 24, configured to use a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identify and mark suspected crack areas.
[0094] A detection module 25, configured to use a deep learning model to detect real cracks in the suspected crack areas to obtain the crack positions, crack sizes, and shapes of the real cracks.
[0095] An evaluation and generation module 26, configured to analyze the crack distribution in the case of multiple real cracks, and based on the crack positions, crack sizes, and shapes of each real crack, and the crack distribution, evaluate the damage degree of the double-block sleeper, and generate a crack detection report of the double-block sleeper according to the evaluation result of the damage degree.
[0096] Figure 2 The crack detection system of the double-block sleeper described above can execute Figure 1 The crack detection method of the double-block sleeper described in the embodiments shown, and its implementation principle and technical effects will not be elaborated. For the crack detection system of the double-block sleeper in the above embodiments, the specific ways in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0097] In a possible design, Figure 2 The crack detection system of the double-block sleeper in the embodiments shown can be implemented as a computing device, such as Figure 3 As shown in the figure, the computing device may include a storage component 31 and a processing component 32.
[0098] The storage component 31 stores one or more computer instructions, where the one or more computer instructions are called and executed by the processing component 32.
[0099] The processing component 32 is configured to: obtain multi-view images, ultrasonic detection data, and infrared thermal imaging data of the double-block sleeper; input the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the double-block sleeper; fuse the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; use a target filter to extract crack features of different scales from the target three-dimensional reconstruction model, and based on all the extracted crack features of different scales, identify and mark suspected crack areas; use a deep learning model to detect real cracks in the suspected crack areas to obtain the crack positions, crack sizes, and crack shapes of the real cracks; in the case of multiple real cracks, analyze the crack distribution, and based on the crack positions, crack sizes, and crack shapes of each real crack, as well as the crack distribution, evaluate the damage degree of the double-block sleeper, and generate a crack detection report of the double-block sleeper according to the damage degree evaluation result.
[0100] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for executing the above method.
[0101] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0102] Of course, the computing device necessarily may also include other components, such as input / output interfaces, display components, communication components, etc.
[0103] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above-mentioned peripheral interface module may be an output device, an input device, etc.
[0104] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0105] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above-mentioned processing component, storage component, etc. may be basic server resources leased or purchased from a cloud computing platform.
[0106] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 crack detection method of the double-block sleeper shown in the embodiment.
[0107] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0108] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0109] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A crack detection method for a double-block sleeper, characterized in that: include: Acquire multi-view images, ultrasonic detection data and infrared thermal imaging data of dual-block sleepers; Inputting the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the dual-block sleeper; The ultrasonic detection data and the infrared thermal imaging data are fused with the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; Using a target filter, extracting crack features of different scales from the target three-dimensional reconstructed model, and identifying and marking suspected crack areas based on the extracted crack features of all scales; Using a deep learning model to detect real cracks in the suspected crack area to obtain the crack position, crack size and shape of the real cracks; In the presence of multiple real cracks, the crack distribution is analyzed, and the damage degree of the dual-block sleeper is evaluated based on the crack position, crack size and shape, and crack distribution of each of the real cracks, and a crack detection report for the dual-block sleeper is generated based on the damage degree evaluation results.
2. The method according to claim 1, characterized in that The step of fusing the ultrasonic detection data, the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model includes: Preprocessing the ultrasonic detection data to obtain preprocessed ultrasonic detection data, based on the preprocessed ultrasonic detection data, using wavelet transform combined with multi-resolution analysis to calculate the wavelet coefficients of each layer to obtain feature information of different scales, generate feature extraction results, based on the feature extraction results, apply short-time Fourier transform to extract time-frequency domain features of the ultrasonic detection data, and obtain a time-frequency domain feature map; Perform temperature correction on the infrared thermal imaging data to obtain corrected infrared thermal imaging data, use cluster analysis technology to identify abnormal points in the infrared thermal imaging data, and obtain abnormal detection results; According to the pre-processed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model, multi-modal feature fusion is performed through feature splicing technology and attention mechanism to generate an intermediate three-dimensional reconstruction model; According to the time-frequency domain feature graph and the anomaly detection result, the decision-level fusion is performed in combination with ensemble learning and Bayesian technology to obtain a decision-level fusion result, and the intermediate 3D reconstruction model is optimized according to the decision-level fusion result to obtain a target 3D reconstruction model.
3. The method according to claim 2, characterized in that The method generates an intermediate 3D reconstruction model by performing multimodal feature fusion through feature splicing technology and attention mechanism based on the pre-processed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial 3D reconstruction model, including: Using an attention mechanism and combining the historical attention weights of each modal feature, the initial attention weights of the modal features corresponding to the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model are calculated; According to the relationship between different modal features, the attention weight of each modal feature is adjusted to obtain the target attention weight of each modal feature; By adopting feature stitching technology, the preprocessed ultrasonic detection data, the corrected infrared thermal imaging data, and the visible light image feature data in the initial three-dimensional reconstruction model are weighted stitching processed according to the target attention weight of each modal feature to obtain a multi-modal weighted feature vector, and an intermediate three-dimensional reconstruction model is generated according to the multi-modal weighted feature vector.
4. The method according to claim 2, characterized in that: The method of performing decision-level fusion according to the time-frequency domain feature graph and the anomaly detection result in combination with ensemble learning and Bayesian technology to obtain a decision-level fusion result, optimizing the intermediate 3D reconstruction model according to the decision-level fusion result, and obtaining a target 3D reconstruction model includes: Based on ensemble learning, multiple base classifiers are constructed. According to the time-frequency domain feature graphs and anomaly detection results, different machine learning algorithms are used to train each base classifier separately to generate the prediction results of each base classifier. Calculate the prior probability and likelihood function of each base classifier, use the Bayesian algorithm to calculate the posterior probability, and obtain the decision-level fusion result based on the posterior probability and the prediction results of each base classifier; The parameters of the intermediate 3D reconstruction model are adjusted according to the decision-level fusion result to obtain a target 3D reconstruction model corresponding to the adjusted parameters.
5. The method according to claim 1, characterized in that The deep learning model includes a target detection model and a speech segmentation model. The deep learning model is used to detect the real cracks in the suspected crack area to obtain the crack position, crack size and shape of the real cracks, including: Use the pre-trained target detection model and the irregular recognition box to determine the boundary data of the real crack; According to the boundary data of the real crack, combined with a pre-trained semantic segmentation model, semantic segmentation is performed on the image in the suspected crack area to obtain a pixel-level mask of the real crack; Based on the pixel-level mask of the real crack, the shape of the real crack is determined, and the pixel-level length and width of the real crack in the image within the suspected crack area are calculated; According to the pixel-level length and width of the real crack in the image of the suspected crack area, combined with the conversion relationship and proportional relationship between the image three-dimensional coordinate system and the spatial coordinate system, the actual length and width of the real crack are determined.
6. The method according to claim 5, characterized in that The method of using a pre-trained target detection model and an irregular recognition frame to determine the boundary data of a real crack includes: Use the pre-trained target detection model to detect the boundaries of real cracks in the suspected crack area through irregular recognition boxes; The regression bias function is used to calculate the position bias value of the irregular recognition frame, and the spatial attention mechanism is used to determine the position perception of the irregular recognition frame. The weighted sum of the position bias value and the position perception is determined as the position difference. The position difference is used to correct the boundary of the real crack to obtain the boundary data of the real crack.
7. The method according to claim 5, characterized in that The step of performing semantic segmentation on the image in the suspected crack region based on the boundary data of the real crack and combining with a pre-trained semantic segmentation model to obtain a pixel-level mask of the real crack includes: According to the boundary data of the real crack, each pixel in the image in the suspected crack area is classified using a pre-trained semantic segmentation model to achieve semantic segmentation of the image in the suspected crack area and obtain a pixel-level label map, wherein each pixel-level mask in the pixel-level label map is used to indicate whether the corresponding pixel is a crack or a non-crack; The pixel-level mask in the pixel-level label map is adjusted to obtain an adjusted pixel-level label map so that the real cracks are continuous and complete; The adjusted pixel-level label map is optimized by combining the prior knowledge of cracks and the context information of the real cracks to obtain the pixel-level mask of the real cracks.
8. A crack detection system for a double-block sleeper, characterized in that: include: An acquisition module, used to acquire multi-view images, ultrasonic detection data and infrared thermal imaging data of the dual-block sleeper; A generation module, used for inputting the multi-view images into a pre-trained multi-view fusion model to generate an initial three-dimensional reconstruction model of the dual-block sleeper; A fusion module, used for fusing the ultrasonic detection data and the infrared thermal imaging data, and the initial three-dimensional reconstruction model to obtain a target three-dimensional reconstruction model; An extraction and recognition module is used to extract crack features of different scales from the target three-dimensional reconstruction model using a target filter, and identify and mark suspected crack areas based on the extracted crack features of all scales; A detection module, used to detect real cracks in the suspected crack area using a deep learning model to obtain the crack position, crack size and shape of the real cracks; The evaluation generation module is used to analyze the crack distribution in the presence of multiple real cracks, and evaluate the damage degree of the double-block sleeper according to the crack position, crack size and shape of each real crack, as well as the crack distribution, and generate a crack detection report for the double-block sleeper according to the damage degree evaluation result.
9. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a crack detection method for a dual-block sleeper as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a crack detection method for a double-block sleeper as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Crack detection method for doubled-block sleeper
CN110044905A
Deep learning-based offshore wind turbine blade defect detection system and method
CN118582351A
Intelligent detection system for house outer wall cracks
CN119394191A
Civil engineering structure defect detection method and system based on image processing technology
CN119810319A
Apparatus and methods for generating a three-dimensional (3D) model of an anatomical object via machine-learning
US20250117929A1
Cited By
Self-navigation tunnel crack detection method, system, equipment and medium
CN120259891A
Concrete crack leakage damage assessment method and system
CN120470452A
Method and system for evaluating concrete crack leakage damage
CN120470452B
Infrared thermal imaging and machine learning fused metal plate crack real-time detection system
CN121186128A
Real-time detection system for metal plate crack based on infrared thermal imaging and machine learning
CN121186128B