A high-speed rail part crack real-time detection method

By integrating a multimodal feature fusion network that combines hyperspectral imaging, structured light 3D reconstruction, and pulsed eddy current detection, the problem of low efficiency and complex structure identification in crack detection of high-speed rail components has been solved, achieving efficient and stable online crack detection and improving detection accuracy and robustness.

CN120525826BActive Publication Date: 2025-11-25QINGDAO NANYANG SANCHENG MASCH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510604469.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-11-25
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Traditional methods for detecting cracks in high-speed rail components suffer from low efficiency, easy to miss, difficulty in achieving online monitoring and identification of complex internal structures. Existing computer vision-based methods are prone to losing local detail features when processing high-resolution images, and their multi-sensor data fusion capabilities are insufficient.

Method used

A cross-modal feature alignment network and attention enhancement mechanism are constructed, integrating hyperspectral imaging, structured light 3D reconstruction and pulsed eddy current detection data. Through a multimodal feature fusion network, all-weather, non-contact intelligent diagnosis of sub-millimeter cracks is achieved. Crack segmentation and quantization are performed by combining an adaptive crack segmentation method.

Benefits of technology

Breaking through the bottleneck of single-mode detection, it improves the detection rate of cracks under coatings, reduces the quantification error of micro-cracks, enhances the detection stability under complex working conditions, and realizes the innovation of operation and maintenance mode from traditional downtime detection to online real-time monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525826B_ABST
    Figure CN120525826B_ABST
Patent Text Reader

Abstract

The application provides a high-speed rail part crack real-time detection method, and belongs to the image detection technical field based on computer vision. Hyperspectral images, visible light images, three-dimensional point cloud data and eddy current signal data of a high-speed rail part surface are acquired; the hyperspectral images and the visible light images of the high-speed rail part surface are spatially registered, and the three-dimensional point cloud data is time-registered based on the eddy current signal data; and feature enhancement and standardization processing are performed; the obtained standardized multi-modal data are subjected to feature extraction and fusion based on a multi-modal feature fusion network, so as to obtain a fusion feature map representing crack details of the high-speed rail part; and based on an adaptive crack segmentation method, a crack contour is extracted from the point cloud, and then whether the part has a crack, the length and the depth of the crack are calculated. The application innovatively integrates three kinds of modal data, breaks through the information dimension limitation of single-modal detection, and realizes all-weather, non-contact and intelligent efficient diagnosis of sub-millimeter cracks under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image detection technology based on computer vision, and particularly relates to a real-time crack detection method for high-speed rail components. Background Technology

[0002] Crack detection of high-speed rail components is a core technical aspect of ensuring the safe operation of rail transit systems, preventing major accidents, and extending equipment lifespan. As a major manifestation of metal fatigue damage, early and accurate identification of cracks can effectively prevent fracture failures in critical components such as bogies, wheelsets, and brake discs during high-speed operation. Traditional detection methods mainly rely on manual visual inspection and non-destructive testing (NDT) techniques, which have significant limitations: 1) Manual inspection requires downtime maintenance, resulting in a significant conflict between inspection efficiency and the high-frequency operation requirements of trains, and is also susceptible to missed detections due to personnel experience and fatigue; 2) Conventional NDT methods such as ultrasonic testing require contact coupling, making it difficult to achieve online monitoring of moving parts; 3) Industrial endoscope inspection is limited by its narrow field of view, limiting its ability to identify cracks within complex structures.

[0003] With the development of intelligent detection technology, crack detection methods based on computer vision are gradually being applied in the rail transit field, mainly divided into two categories: traditional image analysis and deep learning methods.

[0004] Traditional image analysis methods extract crack features through edge detection (such as the Canny operator) and texture analysis (such as gray-level co-occurrence matrix). However, due to factors such as strong vibration, uneven lighting, and oil pollution in the high-speed rail operating environment, they are prone to false defect identification, especially for micron-level cracks and complex background noise.

[0005] Deep learning approach: An end-to-end detection model is constructed using a convolutional neural network (CNN) to achieve crack localization and classification through feature self-learning. However, existing models have three bottlenecks: 1) Single visible light modality is sensitive to interference such as surface oxidation and corrosion, and cannot penetrate coatings to detect hidden cracks; 2) Local detail features are easily lost during pooling operations when processing high-resolution images, leading to missed detection of tiny cracks; 3) The fusion capability of multi-sensor data (such as infrared thermal imaging and laser 3D scanning) is insufficient, failing to fully leverage the complementary advantages of cross-modal data. Summary of the Invention

[0006] To address the aforementioned problems, this invention proposes a real-time crack detection method for high-speed railway components. By constructing a cross-modal feature alignment network and an attention enhancement mechanism, it innovatively integrates hyperspectral imaging, structured light 3D reconstruction, and pulsed eddy current detection data, breaking through the information dimension limitations of single-modal detection. This enables all-weather, non-contact intelligent diagnosis of sub-millimeter-level cracks under complex operating conditions, providing core technical support for the operational safety of high-speed railways.

[0007] This invention provides a method for real-time detection of cracks in high-speed rail components, comprising the following steps:

[0008] S1, acquire the collected hyperspectral images of the surface of high-speed rail components, visible light images of the surface of high-speed rail components, three-dimensional point cloud data of high-speed rail components, and eddy current signal data;

[0009] S2. Spatial registration is performed on the hyperspectral image and visible light image of the surface of high-speed rail components. Sensor perspective differences are eliminated through affine transformation. Temporal registration is performed on the 3D point cloud data of high-speed rail components based on eddy current signal data. Feature enhancement and standardization processing are performed on the above three types of data to obtain standardized multimodal data.

[0010] S3, based on the multimodal feature fusion network, performs feature extraction and fusion on the obtained standardized multimodal data to obtain a fused feature map characterizing the crack details of high-speed rail components;

[0011] S4, based on the adaptive crack segmentation method, uses the output fused feature map and dynamically adjusts the segmentation threshold by combining the curvature of the three-dimensional point cloud to extract the crack contour from the point cloud. Then, it combines the three-dimensional point cloud data of standardized high-speed rail components to calculate whether the component has cracks and the length and depth of the cracks.

[0012] Preferably, in step S2, the spatial registration of the hyperspectral image and the visible light image of the high-speed rail component surface is achieved by extracting key matching points between the SIFT feature points of the visible light image and the band correlation features of the hyperspectral image of the high-speed rail component surface using the Scale Invariant Feature Transform (SIFT) algorithm. The key matching points are then used to calculate the affine transformation matrix, which is used to map the pixels in the hyperspectral image to the pixel coordinate system of the visible light image, thus realizing cross-modal spatial registration of the hyperspectral image and the visible light image of the high-speed rail component surface.

[0013] The time registration of 3D point cloud data of high-speed rail components based on eddy current signal data involves using a linear interpolation algorithm to complete the 3D point cloud data of high-speed rail components using the timestamp of the eddy current signal data, and eliminating the timing deviation caused by sensor response delay through dynamic time warping to obtain the registered 3D point cloud data of high-speed rail components.

[0014] Preferably, the feature enhancement in S2 specifically involves:

[0015] For the registered hyperspectral images of high-speed rail component surfaces, principal component analysis (PCA) is used to retain the first three principal components, compress redundant bands, and enhance material difference features to obtain feature-enhanced hyperspectral images of high-speed rail component surfaces. For the registered 3D point cloud data of high-speed rail components, statistical outlier filtering and Poisson surface reconstruction are used to repair the missing scan areas to obtain feature-enhanced 3D point cloud data of high-speed rail components. Histogram equalization is used to perform histogram equalization on the registered visible light images of high-speed rail component surfaces to obtain feature-enhanced visible light images of high-speed rail component surfaces.

[0016] Preferably, the multimodal feature fusion network in S3 includes a multimodal feature extraction module, a cross-modal attention module, and a residual dense feature enhancement module;

[0017] The multimodal feature extraction module includes an improved ResNet-50 network, an improved 3D-UNet network, and a hierarchical point Transformer network, which respectively extract features from the obtained hyperspectral images of the surface of standardized high-speed rail components, visible light images of the surface of standardized high-speed rail components, and three-dimensional point cloud data of standardized high-speed rail components to obtain multimodal coding features with modality-specific characteristics.

[0018] The cross-modal attention module complements the advantages of multimodal encoded features through adaptive weight allocation, and fuses the encoded datasets with attention to obtain a multimodal feature map after attention fusion.

[0019] The residual dense feature enhancement module enhances the detail retention capability of the fused features through dense cross-layer connections and residual learning, thereby obtaining a small detail feature map of component cracks.

[0020] Preferably, the multimodal feature extraction module specifically includes:

[0021] The improved ResNet-50 network modifies the 3×3 convolutional layers in Stage 3 and Stage 4 of the traditional ResNet-50 network's residual blocks. Deformable convolutions replace the standard 3×3 convolutions in Stages 3 and 4 of the traditional ResNet-50 network. In the deformable convolutions of Stage 3, a dilation rate of 2 is set to expand the receptive field and capture the continuity features of large-scale cracks. In the deformable convolutions of Stage 4, the dilation rate is kept at 1 to retain the ability to model the local deformation of small cracks. Simultaneously, the global average pooling layer and fully connected layer at the end of the traditional ResNet-50 network are removed, and a 1×1 convolutional layer is added at the end of Stage 4, adjusting the number of output channels to 256 to adapt to the input dimension of the subsequent cross-modal attention module. Standardized visible light images of high-speed rail component surfaces are input into the improved ResNet-50 network, which outputs four levels of crack feature maps.

[0022] The improved 3D-UNet network comprises a 3D convolutional encoder and a nonlocal attention layer. The 3D convolutional encoder is responsible for extracting local spatiotemporal spectral features of hyperspectral data, while the nonlocal attention layer is embedded at the end of the 3D convolutional encoder to model global spectral dependencies and capture cross-band nonlocal correlations. The 3D convolutional encoder consists of four layers: the first input layer consists of one 3D convolutional layer, one batch normalization layer, and one ReLU activation function layer; the second, third, and fourth layers are downsampling layers, consisting of two 3D convolutional layers and one 3D max pooling layer; the nonlocal attention layer consists of nonlocal blocks and an attention mechanism to model long-range spectral dependencies. The network inputs standardized hyperspectral images of high-speed rail component surfaces into the 3D-UNet network and outputs crack feature maps containing both local spatiotemporal spectral features and global spectral features.

[0023] The hierarchical point Transformer network includes a farthest point sampling layer, local point Transformer modules, and hierarchical feature transfer. The farthest point sampling layer consists of four downsampling layers. The local point Transformer modules consist of a self-attention computation layer, residual connections, and layer normalization layers, and use two fully connected layers for further nonlinear mapping. The hierarchical feature transfer is performed after each point Transformer module processes the data, and then uses max pooling to transfer the features to adjacent layers to achieve multi-scale feature fusion. Standardized high-speed rail component 3D point cloud data is input into the hierarchical point Transformer network. First, the standardized high-speed rail component 3D point cloud data is sampled from the farthest point to downsample the point cloud to four layers to construct a multi-scale hierarchical structure. Second, the point Transformer modules are used to perform local geometric encoding on the multi-scale hierarchical structure, and local neighborhood features are aggregated through a self-attention mechanism to output 3D point cloud features for characterizing geometric deformation.

[0024] Preferably, the cross-modal attention module consists of four parts: a projection layer, a bidirectional cross-attention layer, a geometric enhancement fusion layer, and a multi-scale feature aggregation layer. The projection layer consists of three one-dimensional convolutions and two normalization layers. The bidirectional cross-attention layer consists of four cross-attention modules, with the dimension of each module reduced from 64 to 16, and the forward and reverse attention paths are structurally symmetrical. The geometric enhancement fusion layer consists of pointwise multiplication and weighted fusion operations. The multi-scale feature aggregation layer uses four max pooling layers to downsample different scales, then unifies the resolution through bilinear interpolation, and outputs the concatenated features through one-dimensional convolution.

[0025] Preferably, the residual dense feature enhancement module includes a 3×3 convolutional layer, a GRU gated unit layer, and a feature stitching layer. First, the multimodal feature map is input into the residual dense connection module, and after 3×3 convolution, a 64-channel feature map is generated. Second, the 64-channel feature map and the multimodal feature map are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to suppress noise propagation, generating a 32-channel feature map. Third, the 64-channel and 32-channel feature maps are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to generate a 16-channel feature map. Finally, the 64-channel, 32-channel, and 16-channel feature maps are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to generate an 8-channel component crack micro-detail feature fusion map.

[0026] Preferably, the adaptive crack segmentation method in S4 specifically includes:

[0027] S11, fused feature maps using 1×1 convolution. Mapped to a single-channel confidence heatmap This represents the probability that each pixel belongs to a crack;

[0028] S12, Extracting the local curvature of the standardized 3D point cloud data of high-speed rail components obtained in S2. A larger curvature value indicates a more significant surface deformation;

[0029] S13, based on the confidence heatmap M conf A pixel-level segmentation threshold T(p) is generated by combining the curvature feature K(p) with the threshold to suppress false detections in low-confidence regions. The specific calculation formula is as follows:

[0030] T(p)=α·M conf (p)+β·tanh(K(p) / γ)

[0031] Where α = 0.6 and β = 0.4 are weighting coefficients, γ = 0.2 controls curvature sensitivity, the tanh function normalizes the curvature to [-1, 1], and finally, M... conf Elements greater than T(p) are set to 1, and elements less than T(p) are set to 0, generating a segmentation mask M. mask .

[0032] Preferably, by combining standardized high-speed rail component 3D point cloud data, the presence, length, and depth of cracks in the components are calculated. The specific process is as follows:

[0033] S21, First, perform 3D contour extraction, and then perform binary mask M... maskMorphological closing operations are performed to connect fracture regions and extract the contours of connected regions to obtain the two-dimensional contours of cracks in high-speed rail components.

[0034] S22, the two-dimensional contour of the crack in the high-speed rail component is projected onto the standardized three-dimensional point cloud data of the high-speed rail component to obtain enhanced three-dimensional point cloud data. The improved α-shape algorithm is then used to reconstruct the three-dimensional geometry of the crack in the enhanced three-dimensional point cloud data to obtain the three-dimensional contour point set C. 3D ;

[0035] S23, Crack centerline fitting; for the three-dimensional contour point set C 3D Principal component analysis was used to determine the main crack propagation direction v. dir And along v dir The direction is to slice the point cloud, and within each slice, the local center point {c} is fitted using the RANSAC algorithm. k}, forming the crack centerline path

[0036] S24, Crack length and depth calculation; first, calculate the length along the centerline path. Integrating the Euclidean distance between adjacent points yields the crack length L. crack The formula is as follows:

[0037]

[0038] Among them, c k Indicates the local center point;

[0039] Next, depth calculation is performed on the enhanced 3D point cloud data P, along the centerline normal n(c) of each point. k The formula for extracting the depth profile is as follows:

[0040]

[0041] Where, n(c k ) is c k The crack depth is taken as the maximum value within the local neighborhood (radius 0.2 mm).

[0042] S25, Extended Direction Assessment: The rate of change of direction is calculated based on the centerline path L, using the following formula:

[0043]

[0044] Among them, v k =c k+1 -c k When θ>15°, the crack is determined to have branched cracks.

[0045] Preferably, the multimodal feature fusion network is based on the crack quantization label of S4, i.e., three types of output results. A multi-task loss function is designed, and the multimodal feature fusion network is compressed into a MobileNetV3 lightweight model through knowledge distillation to reduce the number of parameters and obtain a trained lightweight model.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] (1) Breakthrough in single-mode detection bottleneck: By combining hyperspectral imaging with multi-mode complementarity of pulsed eddy current, the problem of surface oxide layer covering hidden cracks is solved, and the detection rate of cracks under coating is significantly improved.

[0048] (2) Improve the accuracy of micro-crack identification: Based on three-dimensional point cloud curvature analysis and detail enhancement of the RDCM module, the quantization error of cracks in the 0.1-0.5mm range is greatly reduced;

[0049] (3) Enhanced robustness under complex working conditions: The dynamic threshold segmentation algorithm still maintains high detection stability under extreme conditions such as strong vibration and oil stains;

[0050] (4) Achieve operational and maintenance model innovation: Through the 5G edge computing architecture, the detection cycle of key components is transformed from traditional downtime detection to online real-time monitoring, reducing operation and maintenance costs and providing core technical support for the safe operation of high-speed rail. Attached Figure Description

[0051] Figure 1 This is a flowchart illustrating the overall technical route of the present invention.

[0052] Figure 2 This is a schematic diagram of the multimodal feature fusion network structure of the present invention.

[0053] Figure 3 This is a schematic diagram of the 3D-UNet network structure of the present invention.

[0054] Figure 4 This is a schematic diagram of the layered point Transformer network structure of the present invention.

[0055] Figure 5 This document outlines a real-time crack detection process for high-speed rail components after model deployment.

[0056] Figure 6 This is a comparison chart of the accuracy of three crack detection methods in the embodiments of the present invention.

[0057] Figure 7 This is a comparison chart of the robustness of three crack detection methods in the embodiments of the present invention. Detailed Implementation

[0058] This invention proposes a real-time crack detection method for high-speed rail components, the overall process of which is as follows: Figure 1 As shown: S1, hyperspectral images, visible light images, 3D point cloud data, and eddy current signal data of high-speed rail component surfaces are simultaneously acquired using a hyperspectral camera, a laser 3D scanner, and a pulsed eddy current sensor; S2, spatial registration, temporal synchronization, feature enhancement, and standardization are performed on the three modal data of high-speed rail component surfaces (hyperspectral images, visible light images, and 3D point cloud data) to eliminate feature misalignment caused by sensor perspective differences, obtaining consistent multimodal data in both spatial and temporal dimensions, including standardized hyperspectral images, standardized visible light images, and standardized 3D point cloud data of high-speed rail component surfaces; S3, a multimodal feature fusion network is then constructed to extract features from the three standardized images and fuse the features to obtain fused features of local details with minute cracks; S4, subsequently, an adaptive crack segmentation method is designed to achieve dynamic crack segmentation and quantization, and the presence, length, and depth of cracks in the component are calculated using the fused feature map and standardized high-speed rail components.

[0059] Meanwhile, a multi-task lightweight loss function is designed to train and compress the multimodal feature fusion network. Finally, the multimodal data alignment and preprocessing method, the adaptive crack segmentation method, and the trained lightweight model are deployed at the edge to develop an embedded real-time inference engine that can detect the crack status of components at 30 frames per second under train operation conditions.

[0060] The specific implementation process of the present invention will be described in detail below with reference to specific embodiments.

[0061] I. Multimodal Data Acquisition

[0062] A hyperspectral camera, a structured light 3D scanner, and a pulsed eddy current sensor are deployed to achieve simultaneous acquisition of multi-source data through hardware triggering. This results in the acquisition of hyperspectral images of the surface of high-speed rail components, visible light images of the surface of high-speed rail components, 3D point cloud data of high-speed rail components, and eddy current signal data. The specific steps are as follows:

[0063] S1-1, a hyperspectral camera, a structured light 3D scanner (scanning accuracy ±0.05mm, frame rate 30fps), and a pulsed eddy current sensor (frequency 1-100kHz, sensitivity 0.1mV / μm) are installed on the inspection platform, and a constant current light source and an environmental shield are configured to eliminate external light and electromagnetic interference; the hyperspectral camera is used to acquire hyperspectral images and visible light images of the surface of high-speed rail components, the structured light 3D scanner is used to acquire 3D point cloud data of high-speed rail components, and the pulsed eddy current sensor is used to acquire eddy current signal data;

[0064] S1-2, based on FPGA, develops a synchronous control module to generate hardware trigger signals (rising edge jitter <1μs) to drive a hyperspectral camera (exposure time 10ms), a structured light 3D scanner (fringe projection period 20ms), and an eddy current sensor (sampling rate 1MHz) to synchronously acquire data.

[0065] S1-3 performs integrity verification on the data collected in S1-2, verifies the spatial consistency between hyperspectral and visible light images by feature point matching, evaluates scanning accuracy using the three-dimensional point cloud back projection error (RMSE<0.1mm), detects the signal-to-noise ratio of eddy current signals (SNR>35dB), constructs the original multimodal database, and obtains the original data after integrity verification.

[0066] II. Multimodal Data Alignment and Preprocessing

[0067] Spatial registration was performed on the acquired hyperspectral and visible light images of the high-speed rail component surfaces. Affine transformation was used to eliminate sensor perspective differences. Temporal registration was performed on the acquired 3D point cloud data of the high-speed rail components. Feature enhancement and standardization were then applied to the three types of data to obtain standardized multimodal data. The details are as follows:

[0068] S2-1, Cross-modal spatial registration of hyperspectral and visible light images of high-speed rail component surfaces: Key matching points are extracted from the SIFT feature points of the visible light image of the high-speed rail component surface and the band correlation features of the hyperspectral image of the high-speed rail component surface using the Scale Invariant Feature Transform (SIFT) algorithm. The affine transformation matrix is ​​calculated using the key matching points. The affine transformation matrix is ​​then used to map the pixels in the hyperspectral image of the high-speed rail component surface to the pixel coordinate system of the visible light image of the high-speed rail component surface, thus achieving cross-modal spatial registration of the hyperspectral and visible light images of the high-speed rail component surface, and obtaining the registered hyperspectral and visible light images of the high-speed rail component surface.

[0069] S2-2, Time registration of 3D point cloud data of high-speed rail components: Based on the timestamp of the eddy current signal data, the 3D point cloud data of high-speed rail components is linearly interpolated and completed using a linear interpolation algorithm. Dynamic time warping is used to eliminate the timing deviation caused by sensor response delay (maximum time difference <2ms) to obtain the registered 3D point cloud data of high-speed rail components.

[0070] S2-3, Feature enhancement is performed on the registered multimodal data: For the registered hyperspectral image of the high-speed rail component surface, principal component analysis (PCA) is used to retain the first three principal components (cumulative variance contribution rate > 92%), compress redundant bands, and enhance material difference features to obtain a feature-enhanced hyperspectral image of the high-speed rail component surface; For the registered 3D point cloud data of the high-speed rail component, statistical outlier filtering (threshold 2σ) and Poisson surface reconstruction are used to repair the missing scanning areas to obtain a feature-enhanced 3D point cloud data of the high-speed rail component; Histogram equalization is used to perform histogram equalization on the registered visible light image of the high-speed rail component surface to obtain a feature-enhanced visible light image of the high-speed rail component surface.

[0071] S2-4, Multimodal data format standardization; normalize the pixels in the hyperspectral and visible light images of the high-speed rail component surfaces after feature enhancement, transforming the pixel coordinates to the [0, 1] interval; perform voxel mesh generation on the feature-enhanced 3D point cloud data of the high-speed rail components, setting the voxel mesh resolution to 0.1 mm. 3 The normalized hyperspectral images and visible light images of the surface of high-speed rail components, as well as the meshed 3D point cloud data of high-speed rail components, are uniformly converted into 256×256×6 multi-channel tensors to obtain standardized hyperspectral images, standardized visible light images, and standardized 3D point cloud data of the surface of high-speed rail components.

[0072] III. Constructing a Multimodal Feature Fusion Network

[0073] A multimodal feature fusion network is constructed to extract and fuse features from standardized multimodal data, resulting in a fused feature map characterizing the details of cracks in high-speed rail components. This invention designs a multimodal feature fusion network based on a multimodal fusion architecture, such as... Figure 2 As shown, it includes a multimodal feature extraction module, a cross-modal attention module (CMAF), and a residual dense feature enhancement module. The specific steps are as follows:

[0074] S3-1, Construct a multimodal feature extraction module to extract multimodal data features. This module includes an improved ResNet-50 network, an improved 3D-UNet network, and a hierarchical point Transformer network to extract features from the standardized high-speed rail component surface hyperspectral images, standardized high-speed rail component surface visible light images, and standardized high-speed rail component 3D point cloud data obtained in S2, respectively, achieving efficient encoding of modality-specific features. Specifically, this includes:

[0075] 1) An improved ResNet-50 network was constructed to extract features from the visible light images of standardized high-speed rail parts to obtain crack features. In order to enhance the model's ability to capture local features of micro-cracks, the traditional ResNet-50 network was improved by introducing deformable convolution. Deformable convolution enhances the ability to model the geometric deformation of irregular cracks by adding offset parameters to the standard convolution kernel. Specifically, in the residual blocks of the traditional ResNet-50 network, the 3×3 convolutional layers in Stage 3 and Stage 4 are modified. Deformable convolutions replace the standard 3×3 convolutions in Stages 3 and 4 of the traditional ResNet-50 network. In the deformable convolutions of Stage 3, the dilation rate is set to 2 to expand the receptive field and capture the continuity features of large-scale cracks. In the deformable convolutions of Stage 4, the dilation rate is kept to 1 to retain the ability to model the local deformation of small cracks. Furthermore, the global average pooling layer and fully connected layer at the end of the traditional ResNet-50 network are removed, and a 1×1 convolutional layer is added at the end of Stage 4, adjusting the number of output channels to 256 to adapt to the input dimension of the subsequent cross-modal attention module. Standardized visible light images of high-speed rail component surfaces are input into the improved ResNet-50 network, outputting four levels of crack feature maps, as follows:

[0076]

[0077] in, To standardize the visible light images (H=256, W=256) of high-speed rail component surfaces, This is the crack feature map of the l-th layer output by the improved ResNet-50 network. and The dimensions are I vis 1 / 4, 1 / 8, 1 / 16, 1 / 32; ResNet50 def (·) represents an improved ResNet-50 network;

[0078] 2) Construct a 3D-UNet network to extract features from hyperspectral images of standardized high-speed rail component surfaces, obtaining material-sensitive features; the 3D-UNet network structure is as follows: Figure 3As shown, it includes a 3D convolutional encoder and a nonlocal attention layer. The 3D convolutional encoder is responsible for extracting the local spatiotemporal spectral features of hyperspectral data, while the nonlocal attention layer is embedded at the end of the 3D convolutional encoder to model global spectral dependencies and capture cross-band nonlocal correlations. The organic combination of the 3D convolutional encoder and the nonlocal attention layer realizes the enhancement of local-global dual-flow features of materials, and solves the contradiction between local noise sensitivity and long-range material correlation in hyperspectral images of standardized high-speed rail component surfaces. The 3D convolutional encoder consists of four layers. The input layer (first layer) comprises a 3D convolutional layer (kernel size 3×3×3), a batch normalization layer, and a ReLU activation function layer. The second, third, and fourth layers are downsampling layers (sampling stride of 2), consisting of two 3D convolutional layers and a 3D max pooling layer. The nonlocal attention layer consists of nonlocal blocks and an attention mechanism, used to model long-range spectral dependencies. Standardized hyperspectral images of high-speed rail component surfaces are input into the 3D-UNet network, and the output is a crack feature map containing local spatiotemporal spectral features and global spectral features, as detailed below:

[0079]

[0080] Where θ(·), φ(·), and g(·) are the query, key, and value matrices generated by a 1×1×1 convolution, C is the number of channels, and I hyp To standardize the hyperspectral images of high-speed rail component surfaces, 3DUNet(·) is a 3D-UNet network; The output is a feature map containing local spatiotemporal spectral features and global spectral features.

[0081] 3) Construct a hierarchical point Transformer network to extract features from the 3D point cloud data of standardized high-speed rail components, obtaining the geometric features of the components. The model structure is as follows: Figure 4As shown. Specifically, the hierarchical point Transformer network structure includes a farthest point sampling layer, local Point Transformer modules, and hierarchical feature transfer. The farthest point sampling layer consists of four downsampling layers; the local Point Transformer module consists of a self-attention computation layer, residual connections, and layer normalization layers, and also uses two fully connected layers for further nonlinear mapping; hierarchical feature transfer refers to the process of transferring features to adjacent layers through max pooling after processing by each Point Transformer module, achieving multi-scale feature fusion. Standardized high-speed rail component 3D point cloud data is input into the hierarchical point Transformer network. First, the standardized high-speed rail component 3D point cloud data is sampled for the farthest point, downsampling the point cloud to 4 layers (1024, 256, 64, and 16 points respectively), constructing a multi-scale hierarchical structure. Second, the Point Transformer module is used to perform local geometric encoding on the multi-scale hierarchical structure, and local neighborhood features are aggregated through a self-attention mechanism to output 3D point cloud features for characterizing geometric deformation, as shown below:

[0082]

[0083] Wherein, Q, K, and V are query, key, and value vectors generated from standardized 3D point cloud data of high-speed rail components through linear transformation. Let K be the K nearest neighbors of point i (K = 32). It is a 3D point cloud feature used to characterize geometric deformation, where PT(·) is a layered point Transformer and Softmax(·) is the Softmax activation function.

[0084] In section S3-2, this invention designs a Cross-Modal Attention Fusion (CMAF) module. This module achieves complementary advantages of multimodal features through adaptive weight allocation, fusing the dataset encoded in S3-1. The CMAF module consists of four parts: a projection layer, a bidirectional cross-attention layer, a geometric enhancement fusion layer, and a multi-scale feature aggregation layer. The projection layer consists of three one-dimensional convolutions and two normalization layers. The bidirectional cross-attention layer consists of a four-head cross-attention module, with each head dimension reduced from 64 to 16, and the forward and reverse attention paths are structurally symmetrical. The geometric enhancement fusion layer consists of pointwise multiplication and weighted fusion operations. The multi-scale feature aggregation layer uses four max-pooling layers for downsampling at different scales, then unifies the resolution through bilinear interpolation, concatenates the features, and outputs them through a one-dimensional convolution. The features obtained in S4-1 are then processed... and The input to the cross-modal attention fusion module is calculated as follows:

[0085] 1) The input projection layer projects the 3D point cloud to obtain the geometric mask M. geo ;

[0086] 2) and The input is a bidirectional cross-attention layer, first with For query, Given a key and a value, calculate the association weight α between the two modalities as follows:

[0087]

[0088] Among them, W Q , The learnable projection matrix is ​​d = 64, which represents the attention dimension.

[0089] Secondly For Query, Generate a reverse attention weight β for the key and value;

[0090] 3) The attention weights α and β are used in the input geometry enhancement fusion layer for multimodal feature weighting and fusion, calculated as follows:

[0091]

[0092] Where γ∈[0,1] is the learnable modal importance coefficient, initialized to 0.5, ⊙ denotes element-wise multiplication, and M gco A geometric mask generated for 3D point cloud projection is used to enhance the fusion weights of geometrically sensitive regions;

[0093] 4) For and Repeat steps 2) and 3) to obtain different scales. Different scales The input multi-scale feature aggregation layer is used, and low-level details are combined with high-level semantic features through skip connections, as follows:

[0094]

[0095] in, This indicates channel concatenation. Upsample is a bilinear interpolation upsampling method to ensure that the feature map sizes of each level are aligned. Conv 1×1 (·) represents a 1×1 convolutional layer. It is the final multimodal feature map obtained after fusion attention;

[0096] S3-3, this invention designs a Residual Dense Connection Module (RDCM), which enhances fused features through dense cross-layer connections and residual learning. The module preserves details effectively. It includes 3×3 convolutional layers, GRU gated unit layers, and feature concatenation layers; firstly, it... The input residual dense connection module first performs a 3×3 convolution to generate a 64-channel feature map; then the 64-channel feature map and... After feature concatenation in the input feature stitching layer, the feature maps are then fed into a 3×3 convolutional layer for dimensionality reduction, followed by a GRU gated unit to suppress noise propagation, generating a 32-channel feature map. The 64-channel and 32-channel feature maps are then fed into the feature stitching layer again, followed by another 3×3 convolutional layer for dimensionality reduction, and finally into a GRU gated unit to generate a 16-channel feature map. Finally, the 64-channel, 32-channel, and 16-channel feature maps are fed into the feature stitching layer again, followed by another 3×3 convolutional layer for dimensionality reduction, and finally into a GRU gated unit to generate an 8-channel feature map of minute details of component cracks. The feature extraction mode of layer-by-layer dimensionality reduction in the residual dense connection module enables the model to better capture the minute details of cracks in high-speed rail components.

[0097] IV. Adaptive Crack Segmentation and Result Calculation

[0098] S4-1, based on the output 8-channel component crack micro-detail feature map This invention designs a dynamic threshold segmentation method that combines local geometric features with multimodal confidence to achieve robust crack segmentation under complex working conditions. The specific process is as follows:

[0099] 1) Fuse feature maps using 1×1 convolution. Mapped to a single-channel confidence heatmap This represents the probability that each pixel belongs to a crack, and the specific calculation formula is as follows:

[0100]

[0101] Where σ is the Sigmoid function, M conf ∈[0,1].

[0102] 2) Extracting the local curvature of the standardized high-speed rail component 3D point cloud data obtained in S2 A larger curvature value indicates a more significant surface deformation. The specific calculation formula is as follows:

[0103]

[0104] Where, λmin and λ max To standardize the minimum and maximum eigenvalues ​​of the neighborhood covariance matrix of point p in the 3D point cloud data of high-speed rail components, ∈=1e-5 to prevent division by zero.

[0105] 3) Dynamic threshold calculation, based on the confidence heatmap M conf A pixel-level segmentation threshold T(p) is generated by combining the curvature feature K(p) with the threshold to suppress false detections in low-confidence regions. The specific calculation formula is as follows:

[0106] T(p)=α·M conf (p)+β·tanh(K(p) / γ)

[0107] Where α = 0.6 and β = 0.4 are weighting coefficients, γ = 0.2 controls curvature sensitivity, and the tanh function normalizes the curvature to [-1, 1]. Finally, M... conf Elements greater than T(p) are set to 1, and elements less than T(p) are set to 0, generating a segmentation mask M. mask .

[0108] S4-2 Segmentation Mask M mask By combining the covered crack area with standardized 3D point cloud data of high-speed rail components, a hybrid geometric analysis algorithm is used to accurately quantify the crack length, depth, and propagation direction. The specific process is as follows:

[0109] 1) First, perform 3D contour extraction, using the binary mask M. mask Morphological closure operations (3×3 kernels) were performed to connect the fracture regions, and the contours of the connected regions were extracted to obtain the two-dimensional contours of the cracks in the high-speed rail components.

[0110] 2) Project the two-dimensional contour of the crack in the high-speed rail component onto the standardized three-dimensional point cloud data of the high-speed rail component to obtain enhanced three-dimensional point cloud data. Use the improved α-shape algorithm to reconstruct the three-dimensional geometry of the crack in the enhanced three-dimensional point cloud data to obtain the three-dimensional contour point set C. 3D The formula is as follows:

[0111]

[0112] Where α = 0.1 mm is an algorithm parameter that controls the smoothness of the contour, p i Within sphere S, sphere S does not contain any points other than point p in the enhanced 3D point cloud data P. i Other points besides those mentioned above.

[0113] 3) Crack centerline fitting; for the three-dimensional profile point set C 3D Principal component analysis (PCA) was used to determine the main crack propagation direction V. dir And along V dirThe direction is to slice the point cloud, and within each slice, the local center point {c} is fitted using the RANSAC algorithm. k}, forming the crack centerline path

[0114] 4) Crack length and depth calculation; First, calculate the length along the centerline path. Integrating the Euclidean distance between adjacent points yields the crack length L. crack The formula is as follows:

[0115]

[0116] Among them, c k Indicates the local center point.

[0117] Next, depth calculation is performed on the enhanced 3D point cloud data P, along the centerline normal n(c) of each point. k The formula for extracting the depth profile is as follows:

[0118]

[0119] Where, n(c k ) is c k The crack depth is taken as the maximum value within the local neighborhood (radius 0.2 mm).

[0120] 5) Expansion direction assessment, based on centerline path The formula for calculating the rate of change of direction is as follows:

[0121]

[0122] Among them, v k =c k+1 -c k When θ > 15°, the crack is determined to have branched cracks. Through the above steps, three types of results can be obtained: crack determination result, length, and depth.

[0123] V. Multi-task training and lightweight optimization of the model

[0124] Based on crack quantization labels, i.e., three types of output results, a multi-task loss function is designed. Through knowledge distillation, the multimodal feature fusion network is compressed into a lightweight MobileNetV3 model, reducing the number of parameters and obtaining a trained lightweight model. The specific design is as follows:

[0125] S5-1 This invention designs a multi-task loss function and a progressive training strategy to achieve efficient optimization of multimodal feature fusion networks, specifically including:

[0126] 1) First, the loss is divided. The Dice loss combined with FocalLoss is used to address the imbalance between positive and negative samples in the crack region. The process is as follows:

[0127]

[0128] Where γ = 2, λ dice =0.7, λ focal =0.3, θ i Let i be the predicted value of the rate of change of direction for the i-th sample. This represents the true value of the rate of change of direction for the i-th sample. A sample is considered a positive sample if it is not a positive sample, and vice versa.

[0129] 2) Secondly, regression loss The Huber loss function is applied to the crack length and depth parameters, and the process is as follows:

[0130]

[0131] Where δ=0.5, L crack This is the predicted crack length. This represents the actual crack length. If calculating the depth parameter, then L... crack Replace with D(c) k ).

[0132] 3) Finally, there is the cross-modal consistency loss, which constrains the consistency of the response of different modes to the same crack. The process is as follows:

[0133]

[0134] in A fixed indicator function indicates that the loss is calculated only in the crack region; It is the visible light feature of the i-th sample. It is the hyperspectral feature of the i-th sample.

[0135] This invention employs a progressive training strategy, with training divided into three stages. The first stage is single-modal pre-training, where the feature extraction networks for each modality are pre-trained using visible light, hyperspectral, and point cloud data, respectively. The second stage is cross-modal fine-tuning: the feature extraction network parameters are frozen, and only the CMAF and RDCM modules are trained, with the loss weights set to... The third stage is end-to-end joint training, which unfreezes all network parameters and uses the AdamW optimizer (learning rate 1e-4, weight decay 1e-5) for global fine-tuning.

[0136] The model is lightweighted and compressed, using the output crack segmentation mask and geometric quantization label as supervision signals. Knowledge distillation technology is used to compress the multimodal feature fusion network into a lightweight model, enabling efficient deployment on edge devices. First, a "teacher-student" distillation framework is constructed: the teacher model (Teacher) serves as the original fusion network, while the student model (Student) uses MobileNetV3 as the backbone network. A lightweight cross-modal feature fusion module is designed by combining depthwise separable convolution and channel pruning techniques. The distillation process consists of two stages:

[0137] First, knowledge transfer is achieved by minimizing the distribution difference of intermediate layer features between the student model and the teacher model. The feature imitation loss function is defined as follows:

[0138]

[0139] in, and Let represent the feature maps output by the i-th layer of the teacher and student models, respectively. The normalization operation (L2 norm) eliminates the influence of the difference in feature magnitude on the loss calculation.

[0140] Secondly, the segmentation and quantization task outputs of the student model are jointly optimized, and the final multi-task lightweight loss function is shown below:

[0141]

[0142] Wherein, the segmentation loss weight λ1 = 0.3, the regression loss weight λ2 = 0.3, and the cross-modal consistency loss weight λ2 = 0.3. The KL divergence loss is used with a weight λ4 = 0.1 to align the output probability distributions of the teacher and student models.

[0143] VI. Edge-Cloud Collaborative Detection Applications

[0144] The multimodal data alignment and preprocessing method, adaptive crack segmentation method, and trained lightweight multimodal feature fusion network model are deployed to edge devices. Real-time detection is achieved through 5G-MEC. The detection results are aggregated and the model is updated in the cloud to trigger graded early warnings. The real-time crack detection process for high-speed rail components is as follows: Figure 5 As shown, the details are as follows:

[0145] S6-1 deploys multimodal data alignment and preprocessing methods, adaptive crack segmentation methods, and lightweight models on edge devices to preprocess the data collected by sensors;

[0146] S6-2, the preprocessed data is input into the lightweight model to obtain fusion features, and the crack detection results and crack parameters are obtained by using the adaptive crack segmentation method;

[0147] S6-3 aggregates data detected by various edge nodes in the cloud;

[0148] S6-4 constructs a knowledge graph based on historical data, analyzes the spatiotemporal distribution patterns of crack data, and combines this with network prediction of expansion trends. When the prediction depth exceeds a threshold, a maintenance work order is triggered, and detection sensitivity is optimized to support preventative maintenance.

[0149] S6-5 monitors the false alarm rate and latency at the edge in the cloud. When an anomaly occurs, it backtracks the original data and semi-supervised annotation of new samples. It incrementally updates the model through an elastic weight algorithm to balance the old and new knowledge.

[0150] VII. Verification of Experimental Results

[0151] Experiments were conducted using the model proposed in this invention on a multimodal dataset of high-speed rail components. The experimental results, based on the verification method described above, are shown in Table 1. DeepCrack and TransFusion were used for comparison. This invention introduces hyperspectral and cross-modal attention mechanisms, achieving an F1 score of 0.94 and an average precision (mAP@0.5) of 93.6%, a 12.3% improvement over TransFusion. This verifies the advantages of multimodal complementarity. Figure 6 As shown. The crack length detection MAE (±0.03mm) and depth detection MAE (±0.01mm) of this invention are significantly better than TransFusion (±0.10mm / ±0.07mm), demonstrating superior geometric modeling capabilities for dynamic threshold segmentation and 3D curvature analysis. In strong vibration scenarios, the recall rate of this invention (89.5%) far exceeds that of DeepCrack (52.3%) and TransFusion (75.6%), thanks to hyperspectral anti-motion blurring and rigid point cloud registration; in oil-covered scenarios, this invention (87.2%) is 17% higher than TransFusion (70.2%), indicating that the ability of hyperspectral imaging to penetrate oil layers plays a key role, such as... Figure 7 As shown in the figure. Experimental results show that this method has high accuracy in detecting cracks in high-speed rail components.

[0152] Table 1 Experimental Errors of Multimodal Dataset for High-Speed ​​Railway Components

[0153]

[0154]

[0155] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0156] While the specific embodiments of the present invention have been described above, they are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for real-time detection of cracks in high-speed rail components, characterized in that, Includes the following steps: S1, acquire the collected hyperspectral images of the surface of high-speed rail components, visible light images of the surface of high-speed rail components, three-dimensional point cloud data of high-speed rail components, and eddy current signal data; S2. Spatial registration is performed on the hyperspectral image and visible light image of the surface of high-speed rail components. Sensor perspective differences are eliminated through affine transformation. Temporal registration is performed on the 3D point cloud data of high-speed rail components based on eddy current signal data. Feature enhancement and standardization processing are performed on the above three types of data to obtain standardized multimodal data. S3, based on the multimodal feature fusion network, performs feature extraction and fusion on the obtained standardized multimodal data to obtain a fused feature map characterizing the crack details of high-speed rail components; The multimodal feature fusion network includes a multimodal feature extraction module, a cross-modal attention module, and a residual dense feature enhancement module; The multimodal feature extraction module includes an improved ResNet-50 network, an improved 3D-UNet network, and a hierarchical point Transformer network, which respectively extract features from the obtained hyperspectral images of the surface of standardized high-speed rail components, visible light images of the surface of standardized high-speed rail components, and three-dimensional point cloud data of standardized high-speed rail components to obtain multimodal coding features with modality-specific characteristics. The cross-modal attention module complements the advantages of multimodal encoded features through adaptive weight allocation, and fuses the encoded datasets with attention to obtain a multimodal feature map after attention fusion. The residual dense feature enhancement module enhances the detail preservation capability of the fused features through dense cross-layer connections and residual learning, and obtains the minute detail feature map of the crack in the component. The improved ResNet-50 network modifies the 3×3 convolutional layers in Stage 3 and Stage 4 of the traditional ResNet-50 network's residual blocks. Deformable convolutions replace the standard 3×3 convolutions in Stage 3 and Stage 4 of the traditional ResNet-50 network. In the deformable convolutions of Stage 3, the dilation rate is set to 2 to expand the receptive field and capture the continuity features of large-scale cracks. In the deformable convolutions of Stage 4, the dilation rate is kept to 1 to retain the ability to model the local deformation of small cracks. Simultaneously, the global average pooling layer and fully connected layer at the end of the traditional ResNet-50 network are removed, and a 1×1 convolutional layer is added at the end of Stage 4, adjusting the number of output channels to 256 to adapt to the input dimension of the subsequent cross-modal attention module. The visible light images of the surface of standardized high-speed rail components are input into an improved ResNet-50 network, which outputs crack feature maps at four levels. The improved 3D-UNet network comprises a 3D convolutional encoder and a nonlocal attention layer. The 3D convolutional encoder is responsible for extracting local spatiotemporal spectral features of hyperspectral data, while the nonlocal attention layer is embedded at the end of the 3D convolutional encoder to model global spectral dependencies and capture cross-band nonlocal correlations. The 3D convolutional encoder consists of four layers: the first input layer consists of one 3D convolutional layer, one batch normalization layer, and one ReLU activation function layer; the second, third, and fourth layers are downsampling layers, consisting of two 3D convolutional layers and one 3D max pooling layer; the nonlocal attention layer consists of nonlocal blocks and an attention mechanism to model long-range spectral dependencies. The network inputs standardized hyperspectral images of high-speed rail component surfaces into the 3D-UNet network and outputs crack feature maps containing both local spatiotemporal spectral features and global spectral features. The hierarchical point Transformer network includes a farthest point sampling layer, a local point Transformer module, and hierarchical feature transfer; the farthest point sampling layer consists of four downsampling layers; the local point Transformer module consists of a self-attention computation layer, residual connections, and a layer normalization layer, and uses two fully connected layers for further nonlinear mapping; Hierarchical feature propagation involves passing features to adjacent layers via max pooling after processing by each Point Transformer module, achieving multi-scale feature fusion. Standardized 3D point cloud data of high-speed rail components is input into the hierarchical Point Transformer network. First, the farthest point of the standardized 3D point cloud data of high-speed rail components is sampled, and the point cloud is downsampled to 4 layers to construct a multi-scale hierarchical structure. Second, the Point Transformer module is used to perform local geometric encoding on the multi-scale hierarchical structure, and local neighborhood features are aggregated through a self-attention mechanism to output 3D point cloud features used to characterize geometric deformation. S4, based on the adaptive crack segmentation method, uses the output fused feature map and dynamically adjusts the segmentation threshold by combining the curvature of the three-dimensional point cloud to extract the crack contour from the point cloud. Then, it combines the three-dimensional point cloud data of standardized high-speed rail components to calculate whether the component has cracks and the length and depth of the cracks.

2. The method for real-time detection of cracks in high-speed rail components as described in claim 1, characterized in that: In step S2, spatial registration of the hyperspectral image and the visible light image of the high-speed rail component surface is achieved by extracting key matching points between the SIFT feature points of the visible light image and the band correlation features of the hyperspectral image using the Scale Invariant Feature Transform (SIFT) algorithm. The key matching points are then used to calculate the affine transformation matrix, which maps the pixels in the hyperspectral image to the pixel coordinate system of the visible light image, thus realizing cross-modal spatial registration of the hyperspectral image and the visible light image of the high-speed rail component surface. The time registration of 3D point cloud data of high-speed rail components based on eddy current signal data involves using a linear interpolation algorithm to complete the 3D point cloud data of high-speed rail components using the timestamp of the eddy current signal data, and eliminating the timing deviation caused by sensor response delay through dynamic time warping to obtain the registered 3D point cloud data of high-speed rail components.

3. The real-time crack detection method for high-speed rail components as described in claim 1, characterized in that: The feature enhancement in S2 specifically involves: For the registered hyperspectral images of high-speed rail component surfaces, principal component analysis (PCA) is used to retain the first three principal components, compress redundant bands, and enhance material difference features to obtain feature-enhanced hyperspectral images of high-speed rail component surfaces. For the registered 3D point cloud data of high-speed rail components, statistical outlier filtering and Poisson surface reconstruction are used to repair the missing scan areas to obtain feature-enhanced 3D point cloud data of high-speed rail components. Histogram equalization is used to perform histogram equalization on the registered visible light images of high-speed rail component surfaces to obtain feature-enhanced visible light images of high-speed rail component surfaces.

4. The real-time crack detection method for high-speed rail components as described in claim 1, characterized in that: The cross-modal attention module consists of four parts: a projection layer, a bidirectional cross-attention layer, a geometric enhancement fusion layer, and a multi-scale feature aggregation layer. The projection layer consists of three one-dimensional convolutions and two normalization layers. The bidirectional cross-attention layer consists of four cross-attention modules, with the dimension of each head reduced from 64 to 16, and the forward and reverse attention paths are structurally symmetrical. The geometric enhancement fusion layer consists of pointwise multiplication and weighted fusion operations. The multi-scale feature aggregation layer uses four max pooling layers to downsample for different scales, then unifies the resolution through bilinear interpolation, and outputs the concatenated features through one-dimensional convolution.

5. The real-time crack detection method for high-speed rail components as described in claim 1, characterized in that: The residual dense feature enhancement module includes a 3×3 convolutional layer, a GRU gated unit layer, and a feature stitching layer. First, the multimodal feature map is input into the residual dense connection module, and after 3×3 convolution, a 64-channel feature map is generated. Next, the 64-channel feature map and the multimodal feature map are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to suppress noise propagation, generating a 32-channel feature map. Then, the 64-channel and 32-channel feature maps are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to generate a 16-channel feature map. Finally, the 64-channel, 32-channel, and 16-channel feature maps are input into the feature stitching layer for feature stitching, then input into a 3×3 convolutional layer for dimensionality reduction, and finally input into a GRU gated unit to generate an 8-channel component crack micro-detail feature fusion map.

6. The real-time crack detection method for high-speed rail components as described in claim 1, characterized in that: The adaptive crack segmentation method in S4 specifically includes: S11, fused feature maps using 1×1 convolution. Mapped to a single-channel confidence heatmap , representing the probability that each pixel belongs to a crack; S12, Extracting the local curvature of the standardized 3D point cloud data of high-speed rail components obtained in S2. A larger curvature value indicates a more significant surface deformation; S13, based on the confidence heatmap With curvature characteristics Generate pixel-level segmentation threshold To suppress false detections in low-confidence regions, the specific calculation formula is as follows: ; in, These are the weighting coefficients. Controlling curvature sensitivity, The function normalizes the curvature to [-1, 1], and finally... Medium to large The element is set to 1, less than Set the element to 0 to generate a segmentation mask. .

7. The real-time crack detection method for high-speed rail components as described in claim 6, characterized in that: By combining standardized 3D point cloud data of high-speed rail components, the presence, length, and depth of cracks in the components are calculated. The specific process is as follows: S21, First, perform 3D contour extraction and then apply the binary mask. Morphological closing operations are performed to connect fracture regions and extract the contours of connected regions to obtain the two-dimensional contours of cracks in high-speed rail components. S22 projects the two-dimensional contour of cracks in high-speed rail components onto standardized three-dimensional point cloud data of high-speed rail components to obtain enhanced three-dimensional point cloud data, which is then used in an improved manner. The algorithm reconstructs and enhances the 3D geometry of cracks in 3D point cloud data, obtaining a 3D contour point set. ; S23, Crack centerline fitting; for the three-dimensional contour point set Principal component analysis was used to determine the main crack propagation direction. , and along The point cloud is sliced, and local center points are fitted within each slice using the RANSAC algorithm. , forming the crack centerline path ; S24, Crack length and depth calculation; first, calculate the length along the centerline path. Integrate the Euclidean distance between adjacent points to obtain the crack length. The formula is as follows: ; in, Indicates the local center point; Next, depth calculations are performed to enhance the 3D point cloud data. Normal to each point along the centerline The formula for extracting the depth profile is as follows: ; in, for The local neighborhood (radius 0.2 mm) has the maximum crack depth. ; S25, Extended Direction Assessment, Based on Centerline Path The formula for calculating the rate of change of direction is as follows: ; in, The presence of branched cracks was determined at that time.

8. The real-time crack detection method for high-speed rail components as described in claim 1, characterized in that: The multimodal feature fusion network, based on the crack quantization label of S4, i.e. three types of output results, designs a multi-task loss function, and compresses the multimodal feature fusion network into a MobileNetV3 lightweight model through knowledge distillation, reducing the number of parameters and obtaining a trained lightweight model.

Citation Information

Patent Citations

  • Aeroengine bearing fault diagnosis method based on STFT-IncepNext

    JP7628356B1

  • Object detection and classification using lidar range images for autonomous machine applications

    WO2021041854A1