An automatic detection and risk assessment method and device for track component defects based on multimodal data fusion and attention neural network.
By using multimodal data fusion and attention neural networks to collaboratively collect data from visible light, lidar, and acoustic emission sensors, the problems of perception blind spots and fusion instability in track component defect detection are solved, achieving high-precision defect detection and risk assessment, and supporting scientific decision-making for track maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU RUHUI INFORMATION TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-26
AI Technical Summary
Existing track component defect detection technologies suffer from problems such as blind spots caused by single-modal sensor data acquisition, simple multi-modal data fusion methods lacking dynamic weight adjustment, and lack of localization and risk assessment of defect detection results, making it difficult to meet the needs of track maintenance decision-making.
It employs visible light cameras, lidar, and acoustic emission sensors to collaboratively collect data. Through multimodal data fusion and attention neural networks, it achieves multi-dimensional defect perception, dynamic adaptive fusion, and multi-task joint discrimination, including defect segmentation, classification, spatial localization, and risk assessment.
It improves the accuracy and robustness of defect detection, enhances the practical value of detection results, supports scientific decision-making in track maintenance, and adapts to stability and real-time performance in complex environments.
Smart Images

Figure QLYQS_3 
Figure QLYQS_6 
Figure QLYQS_7
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, and in particular to a method and apparatus for automatic detection and risk assessment of defects in track components based on multimodal data fusion and attention neural networks. Background Technology
[0002] Currently, various technical solutions exist for defect detection in railway track components. Among these, more mature methods are mostly based on single-sensor data, such as detection systems using high-resolution visible light images alone, three-dimensional structural deformation detection systems using laser point cloud scanning, or using acoustic emission technology to monitor cracks and dynamic anomalies. These single-modal detection methods perform well in specific environments, but they suffer from drawbacks such as perception blind spots, poor environmental adaptability, and limited defect type recognition rates.
[0003] Some more advanced technical solutions attempt to integrate multimodal sensor data, such as multi-sensor information fusion detection combining images and point clouds, or defect identification methods based on convolutional neural networks (CNNs) combined with traditional feature extraction algorithms. However, these solutions mostly use simple feature stitching and lack deep fusion mechanisms that address the heterogeneity of multimodal data and dynamic weight adjustment. They also struggle to fully utilize the complementary information from each sensor, resulting in room for improvement in recognition accuracy and robustness.
[0004] Furthermore, most existing systems focus on single defect classification or detection, lacking the capabilities for three-dimensional spatial positioning of targets and comprehensive risk assessment, thus failing to provide more complete auxiliary information for track maintenance decisions. The existing technology suffers from at least the following technical problems: I. Single-modal sensor data acquisition limits defect perception.
[0005] Traditional technologies often rely on single sensors, such as visible light images or laser point clouds, for defect detection. These methods cannot fully capture the multidimensional physical characteristics of defects in track components, leading to missed or false detections in complex environments. This is because single sensors lack robustness to external interference and collect limited information, making it difficult to comprehensively reflect the diversity and complexity of defects.
[0006] Second, the multi-source data fusion method is simple and lacks a dynamic weight adjustment mechanism, which leads to unstable fusion results.
[0007] In existing technologies, multimodal data fusion often employs fixed-weight feature concatenation or weighting methods, failing to automatically adjust weights based on the spatiotemporal quality of data collected by each sensor. This results in the inability to fully leverage the complementary advantages between sensors. The root cause of this problem is the lack of reasonable consideration of the differences in sampling frequency, signal-to-noise ratio, and spatial distribution among the sensors, and the absence of effective adaptive fusion strategies.
[0008] Third, the defect detection results are mostly single-task outputs, lacking localization and risk level assessment, making it difficult to meet the needs of track maintenance decision-making.
[0009] Existing solutions typically perform defect detection or classification tasks independently, failing to integrate the spatial three-dimensional localization of defects with risk classification within a unified framework. This results in insufficient comprehensive judgment of defect information and an inability to provide a scientific basis for subsequent maintenance work. This deficiency is primarily due to the complex design of the multi-task joint model, insufficient training data, and fragmented system functional modules.
[0010] Therefore, there is an urgent need to propose an automatic detection and risk assessment method and device for track component defects based on multimodal data fusion and attention neural networks. Summary of the Invention
[0011] To address the shortcomings of existing track component defect detection technologies, this invention proposes an automatic defect detection and risk assessment method and apparatus for track components based on multimodal data fusion and attention neural networks, in order to solve the following technical problems: I. Enhancing the Multi-Dimensional Perception Capability of Track Component Defects. Existing technologies mostly rely on single sensors, making it difficult to comprehensively capture the multi-physical characteristics of defects, resulting in perception blind spots and high rates of missed and false detections. To address this, this invention employs a fixed configuration of a visible light camera, a lidar point cloud sensor, and an acoustic emission sensor to collaboratively collect track component information, achieving comprehensive perception of multi-modal data and improving the comprehensiveness and accuracy of defect identification.
[0012] II. Achieving Dynamic Adaptive Fusion of Multimodal Data. Addressing the limitations of existing fusion methods, which are often simplistic and have fixed weights, this invention designs a multimodal fusion neural network based on an attention mechanism. This network dynamically adjusts the fusion weights according to the spatiotemporal quality of data from different sensors, fully leveraging the complementary advantages of each sensor's information to enhance the expressive power of fused features and the robustness of defect identification.
[0013] III. Constructing a Multi-Task Joint Defect Judgment Framework. Existing solutions mostly only output single defect categories or segmentation results, lacking comprehensive functions for defect spatial localization and risk level assessment. This invention designs a multi-output network to achieve accurate defect segmentation, classification, spatial coordinate localization, and risk grading, effectively supporting the scientific decision-making needs of track maintenance.
[0014] According to one aspect of this disclosure, a method and apparatus for automatic detection and risk assessment of defects in track components based on multimodal data fusion and attention neural networks are provided, comprising the following steps: (1) Synchronous acquisition of multimodal data: (1.1) Data Acquisition Clock Synchronization: The visible light camera, lidar, and acoustic emission sensor are connected to the edge processing unit. A unified time base is generated based on the second pulse signal of the global satellite navigation system to synchronously trigger data acquisition from each sensor. (1.2) Data tagging: Assign a unique global time sequence number to each frame of data acquired by the sensor. , Establish a correspondence between the track kilometer markers and the location of the inspection vehicle; (2) Spatial alignment and data preprocessing: Spatial registration of laser point cloud data and visible light image data; spatial localization of sound source based on signals received by multiple acoustic emission sensors; and preprocessing of the image data, point cloud data and acoustic emission signals respectively. (3) Multimodal feature extraction: Image features, point cloud features and acoustic features are extracted from the preprocessed image data, point cloud data and acoustic emission signals respectively; (4) Multimodal attention fusion: Image features, point cloud features and acoustic features are spliced together, and channel attention mechanism is used to assign adaptive weights to the spliced multimodal features to generate fused features; (5) Defect identification and risk assessment: Based on the fusion features, defects are segmented and classified to determine the type of defect and its location in space; and risk level is predicted by combining historical defect data and component type metadata. (6) Result output and synchronization: Generate an inspection report containing the original detection image, point cloud, acoustic signal time-frequency diagram, defect detection mask, defect type, coordinates, risk level, and credibility score, and push the defect-related information of risk level II or III to the maintenance task sheet and defect evidence package to the railway integrated maintenance system.
[0015] In some embodiments of this disclosure, the timestamps of the data in step (1) satisfy: ; in and For any two modes, the sampling time is .
[0016] In some embodiments of this disclosure, step (2) spatial alignment and data preprocessing specifically includes the following steps: (2.1) Laser point cloud and image spatial transformation: The external parameter rotation matrix and translation matrix between the lidar and the visible light camera are obtained through calibration. Based on the following homogeneous coordinate transformation formula, the laser point cloud data and image data are spatially registered. ; in This is the intrinsic parameter matrix. Here are the extrinsic rotation and translation matrices, (u,v) are the pixel coordinates, and (X... LIDAR ,Y LIDAR Z LIDAR () represents the point cloud coordinates in the radar coordinate system; (2.2) Acoustic emission localization: Using the synchronous pulse signals received by multiple acoustic emission sensors, the spatial coordinates of the sound source are calculated based on the trilateration method. The spatial coordinates of each sensor are known. The speed of sound, v Measuring the arrival time t of acoustic emission pulses at various points i Using the earliest sensing point as a reference, we set up a system of equations: The spatial location of the sound source is calculated using the least squares method, i=1,2,3,… (2.3) Data preprocessing: Adaptive histogram equalization is performed on visible light image data, voxel filtering downsampling is performed on laser point cloud data, and bandpass filtering and continuous wavelet transform are performed on acoustic emission signals to extract time-frequency energy spectrum.
[0017] In some embodiments of this disclosure, (2.3) the point cloud is subjected to VoxelGrid voxel filtering with a sampling resolution of δ=2mm per side; The acoustic signal is bandpass filtered, with a filtering range of 10kHz to 80kHz, and then the time-frequency energy spectrum is extracted for each 512ms window using continuous wavelet transform.
[0018] In some embodiments of this disclosure, step (3) multimodal feature extraction specifically includes the following steps: (3.1) Image Feature Extraction: Image features are extracted from the image using the ResNet50 network; let the single-frame input be... Its output characteristics are: ; (3.2) Point cloud feature extraction: For each frame of preprocessed point cloud Point cloud features are obtained by processing with PointNet++ network: ; (3.3) Acoustic emission feature extraction: acoustic signal after continuous wavelet transform In the formula, L is the time window length, F is the number of frequency bands, and a one-dimensional convolutional network is used to process and extract acoustic emission features. .
[0019] In some embodiments of this disclosure, step (4) multimodal attention fusion specifically includes the following steps: stitching together image features, point cloud features and acoustic features, and using a channel attention mechanism to assign adaptive weights to the stitched multimodal features to generate fused features; After unfolding the image features, point cloud features, and acoustic features respectively, they are concatenated and combined to obtain the final concatenated feature vector: ; in The output of ResNet after global average pooling (GAP) is shown. A channel attention mechanism is employed to assign adaptive weights to the multimodal fusion features in the concatenated feature vector. A fully connected layer is defined and normalized to obtain the fusion weights α1, α2, and α3. ; Final fusion features: .
[0020] In some embodiments of this disclosure, step (5) defect identification and risk assessment specifically includes the following steps: based on fusion features, performing defect segmentation and classification to determine the type of defect and its location in space; and combining historical defect data and component type metadata to predict the risk level; (5.1) Defect segmentation mask output: The fused features from step (4) are input into a multi-layer fully convolutional network, and a defect segmentation mask of the same size as the original input image is output: ; In the formula, σ is the Sigmoid function. This represents the high-value area of the mask corresponding to the defect area; (5.2) Defect classification: Input the fused features from step (4) into the classification network: ; Output a probability distribution of defect types, wherein the defect types include at least cracks, corrosion, and loose connections; (5.3) Spatial positioning: Based on the defect segmentation mask Y seg Regions where all pixel values exceed the threshold θ are mapped to the laser point cloud space, and the coordinates of the corresponding defects are calculated through the spatial transformation established in step (2.1). ; (5.4) Risk level prediction: Combine the fused features obtained in step (4) with the historical defects and component type metadata of the detected road section. meta After concatenation, the data is input into a multilayer perceptron, and the Softmax function outputs the probability distribution of risk levels. ; The risk level with the highest probability will be used as the final risk assessment level, and the risk levels will include at least Level I, Level II, and Level III.
[0021] An automatic detection and risk assessment device for defects in track components based on multimodal data fusion and attention neural network; It includes a multimodal synchronous acquisition module, an edge processing and synchronization module, and a data fusion and discrimination module connected according to the signal flow direction; The multimodal synchronous acquisition module includes a visible light camera mounted above the detection platform, a lidar fixed next to the track component to be detected, and an acoustic emission sensor embedded in the track contact area. The edge processing and synchronization module integrates the clocks of the Global Navigation Satellite System and the Inertial Measurement Unit for synchronizing sensor sampling triggers; the processor architecture edge computing unit of the edge processing and synchronization module is used to run synchronization and preprocessing algorithms.
[0022] The data fusion and discrimination module includes a graphics processing unit and a field-programmable gate array (FPGA) computing board, which is used to integrate a multimodal neural network.
[0023] The beneficial effects of this invention are as follows: I. Comprehensive perception of multimodal information significantly improves defect detection accuracy. This invention employs fixed joint acquisition and deep fusion of visible light images, laser point clouds, and acoustic emission data, effectively overcoming the limitations of single sensors due to environmental interference and incomplete information. This enables comprehensive capture of multidimensional defect features of track components, significantly improving defect identification accuracy.
[0024] Second, the adaptive attention-based fusion network enhances system robustness. The dynamic weight adjustment mechanism enables the system to flexibly adjust the weights of each sensor based on the quality of real-time acquired data, significantly improving stability and detection reliability in complex field environments, and reducing false alarms and false negatives.
[0025] III. A multi-task joint judgment architecture enables precise defect location and risk assessment, enhancing maintenance decision support capabilities. Through the linkage of defect segmentation, classification, spatial coordinate location, and risk level output, the detection results are structured and refined, significantly improving the practical value of defect information and effectively supporting track maintenance priority scheduling and maintenance work arrangements.
[0026] IV. A robust spatiotemporal synchronization and preprocessing mechanism ensures the system's real-time performance and adaptability. Employing high-precision clock synchronization, multi-sensor data spatial alignment, and effective noise reduction techniques ensures the accuracy of data fusion and the efficiency of the processing, making this system widely applicable to track-based motion detection platforms and on-site automated inspection environments.
[0027] The collaborative acquisition system for multi-source heterogeneous sensor data is equipped with a multi-modal synchronous data acquisition device including a visible light camera, a lidar point cloud acquisition device, and an acoustic emission sensor, enabling real-time acquisition of multi-dimensional information of track components.
[0028] A multimodal adaptive fusion neural network based on the attention mechanism is designed with a dynamic weight adjustment module. The attention mechanism is used to achieve weighted fusion of features from different sensors, thereby enhancing the complementarity and robustness of heterogeneous data.
[0029] The defect multi-task joint discrimination model structure constructs a multi-head output network that includes defect segmentation, classification, spatial three-dimensional localization, and risk level assessment, realizing end-to-end linkage between defect detection and assessment.
[0030] Spatiotemporal synchronization calibration and multi-sensor data preprocessing methods enable high-precision time synchronization, spatial alignment and noise reduction of multi-source sensor data, ensuring the accuracy of fused data and the real-time performance of the system.
[0031] For robust design in the complex environment of the track site, a multi-level filtering and feature enhancement strategy is adopted to improve the stability and detection accuracy of the system under adverse environmental conditions such as vibration, occlusion and changes in lighting.
[0032] In summary, the technical solution of this invention not only improves the productivity and quality of defect detection for track components, but also enhances the accuracy and efficiency of the detection system, thereby strengthening the railway track safety assurance capability. It has good engineering application prospects and promotion value. Detailed Implementation
[0033] The preferred embodiments of the present invention are described below. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Example 1
[0034] This example discloses an automatic defect detection and risk assessment method for track components based on multimodal data fusion and attention neural networks, including the following steps: (1) Synchronous acquisition of multimodal data: (1.1) Data Acquisition Clock Synchronization: The visible light camera, lidar, and acoustic emission sensor are connected to the edge processing unit. A unified time base is generated based on the second pulse signal of the global satellite navigation system to synchronously trigger data acquisition from each sensor. (1.2) Data tagging: Assign a unique global time sequence number to each frame of data acquired by the sensor. , Establish a correspondence between the track kilometer markers and the location of the inspection vehicle; (2) Spatial alignment and data preprocessing: Spatial registration of laser point cloud data and visible light image data; spatial localization of sound source based on signals received by multiple acoustic emission sensors; and preprocessing of the image data, point cloud data and acoustic emission signals respectively. (3) Multimodal feature extraction: Image features, point cloud features and acoustic features are extracted from the preprocessed image data, point cloud data and acoustic emission signals respectively; (4) Multimodal attention fusion: Image features, point cloud features and acoustic features are spliced together, and channel attention mechanism is used to assign adaptive weights to the spliced multimodal features to generate fused features; (5) Defect identification and risk assessment: Based on the fusion features, defects are segmented and classified to determine the type of defect and its location in space; and risk level is predicted by combining historical defect data and component type metadata. (6) Result output and synchronization: Generate an inspection report containing the original detection image, point cloud, acoustic signal time-frequency diagram, defect detection mask, defect type, coordinates, risk level, and credibility score, and push the defect-related information of risk level II or III to the maintenance task sheet and defect evidence package to the railway integrated maintenance system.
[0035] The timestamps of the data in step (1) satisfy: ; in and For any two modes, the sampling time is .
[0036] Step (2), spatial alignment and data preprocessing, specifically includes the following steps: (2.1) Laser point cloud and image spatial transformation: The external parameter rotation matrix and translation matrix between the lidar and the visible light camera are obtained through calibration. Based on the following homogeneous coordinate transformation formula, the laser point cloud data and image data are spatially registered. ; in This is the intrinsic parameter matrix. Here are the extrinsic rotation and translation matrices, (u,v) are the pixel coordinates, and (X... LIDAR ,Y LIDAR Z LIDAR () represents the point cloud coordinates in the radar coordinate system; (2.2) Acoustic emission localization: Using the synchronous pulse signals received by multiple acoustic emission sensors, the spatial coordinates of the sound source are calculated based on the trilateration method. The spatial coordinates of each sensor are known. The speed of sound, v Measuring the arrival time t of acoustic emission pulses at various pointsi Using the earliest sensing point as a reference, we set up a system of equations: The spatial location of the sound source is calculated using the least squares method, i=1,2,3,… (2.3) Data preprocessing: Adaptive histogram equalization is performed on visible light image data, voxel filtering downsampling is performed on laser point cloud data, and bandpass filtering and continuous wavelet transform are performed on acoustic emission signals to extract time-frequency energy spectrum.
[0037] (2.3) VoxelGrid voxel filtering is performed on the point cloud, with a sampling resolution of δ=2mm per side; The acoustic signal is bandpass filtered, with a filtering range of 10kHz to 80kHz, and then the time-frequency energy spectrum is extracted for each 512ms window using continuous wavelet transform.
[0038] Step (3) multimodal feature extraction specifically includes the following steps: (3.1) Image Feature Extraction: Image features are extracted from the image using the ResNet50 network; let the single-frame input be... Its output characteristics are: ; (3.2) Point cloud feature extraction: For each frame of preprocessed point cloud Point cloud features are obtained by processing with PointNet++ network: ; (3.3) Acoustic emission feature extraction: acoustic signal after continuous wavelet transform In the formula, L is the time window length, F is the number of frequency bands, and a one-dimensional convolutional network is used to process and extract acoustic emission features. .
[0039] The step (4) multimodal attention fusion specifically includes the following steps: stitching together image features, point cloud features and acoustic features, and using a channel attention mechanism to assign adaptive weights to the stitched multimodal features to generate fused features; After unfolding the image features, point cloud features, and acoustic features respectively, they are concatenated and combined to obtain the final concatenated feature vector: ; in The output of ResNet after global average pooling (GAP) is shown. A channel attention mechanism is employed to assign adaptive weights to the multimodal fusion features in the concatenated feature vector. A fully connected layer is defined and normalized to obtain the fusion weights α1, α2, and α3. ; Final fusion features: .
[0040] The step (5) defect identification and risk assessment specifically includes the following steps: based on the fusion features, perform defect segmentation and classification, determine the type of defect and its location in space; and combine historical defect data and component type metadata to predict the risk level; (5.1) Defect segmentation mask output: The fused features from step (4) are input into a multi-layer fully convolutional network, and a defect segmentation mask of the same size as the original input image is output: ; In the formula, σ is the Sigmoid function. This represents the high-value area of the mask corresponding to the defect area; (5.2) Defect classification: Input the fused features from step (4) into the classification network: ; Output a probability distribution of defect types, wherein the defect types include at least cracks, corrosion, and loose connections; (5.3) Spatial positioning: Based on the defect segmentation mask Y seg Regions where all pixel values exceed the threshold θ are mapped to the laser point cloud space, and the coordinates of the corresponding defects are calculated through the spatial transformation established in step (2.1). ; (5.4) Risk level prediction: Combine the fused features obtained in step (4) with the historical defects and component type metadata of the detected road section. meta After concatenation, the data is input into a multilayer perceptron, and the Softmax function outputs the probability distribution of risk levels. ; The risk level with the highest probability will be used as the final risk assessment level, and the risk levels will include at least Level I, Level II, and Level III.
[0041] An automatic detection and risk assessment device for track component defects based on multimodal data fusion and attention neural networks includes a multimodal synchronous acquisition module, an edge processing and synchronization module, and a data fusion and discrimination module connected according to the signal flow direction. The multimodal synchronous acquisition module includes a visible light camera mounted above the detection platform, a lidar fixed next to the track component to be detected, and an acoustic emission sensor embedded in the track contact area. The edge processing and synchronization module integrates the clocks of the Global Navigation Satellite System and the Inertial Measurement Unit for synchronizing sensor sampling triggers; the processor architecture of the edge processing and synchronization module is an edge computing unit used to run synchronization and preprocessing algorithms.
[0042] The data fusion and discrimination module includes a graphics processing unit and a field-programmable gate array (FPGA) computing board, which is used to integrate multimodal neural networks. Example 2
[0043] The principle of this example is the same as that of Example 1, but the specific differences are: multimodal data fusion can be achieved using a graph neural network (GNN) or transformer architecture. A multi-stage cascaded model is used to complete defect detection and multi-task discrimination. Visible light images are replaced with infrared imaging, or other acoustic sensors are used to replace the acoustic emission device.
[0044] The shortcomings of the alternative solutions include: the high complexity of the fusion algorithm, leading to reduced efficiency in model training and inference; sensor replacement schemes are prone to decreased data complementarity, affecting overall detection performance; multi-stage models suffer from error accumulation, lack end-to-end optimization, and have poor stability; and the alternative solutions lack adaptability to complex field environments and real-time processing capabilities.
[0045] The advantages of this patented solution are: It introduces a dynamic adaptive weight fusion network based on an attention mechanism, effectively improving the complementary utilization rate of multi-source data; it enables joint defect discrimination across multiple tasks, enhancing the accuracy and application value of detection results; it possesses strong real-time performance and robustness, adapting to complex railway inspection environments; and its system architecture is simple and efficient, meeting the needs of practical engineering applications.
[0046] Although some preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0047] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this application and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for automatic detection and risk assessment of defects in track components based on multimodal data fusion and attention neural networks, characterized in that, Includes the following steps: (1) Synchronous acquisition of multimodal data: (1.1) Data Acquisition Clock Synchronization: The visible light camera, lidar, and acoustic emission sensor are connected to the edge processing unit. A unified time base is generated based on the second pulse signal of the global satellite navigation system to synchronously trigger data acquisition from each sensor. (1.2) Data tagging: Assign a unique global time sequence number to each frame of data acquired by the sensor. , Establish a correspondence between the track kilometer markers and the location of the inspection vehicle; (2) Spatial alignment and data preprocessing: Spatial registration of laser point cloud data and visible light image data; spatial localization of sound source based on signals received by multiple acoustic emission sensors; and preprocessing of the image data, point cloud data and acoustic emission signals respectively. (3) Multimodal feature extraction: Image features, point cloud features and acoustic features are extracted from the preprocessed image data, point cloud data and acoustic emission signals respectively; (4) Multimodal attention fusion: Image features, point cloud features and acoustic features are spliced together, and channel attention mechanism is used to assign adaptive weights to the spliced multimodal features to generate fused features; (5) Defect identification and risk assessment: Based on the fusion features, defects are segmented and classified to determine the type of defect and its location in space; Risk level prediction is made by combining historical defect data and component type metadata; (6) Result output and synchronization: Generate an inspection report containing the original detection image, point cloud, acoustic signal time-frequency diagram, defect detection mask, defect type, coordinates, risk level, and credibility score, and push the defect-related information of risk level II or III to the maintenance task sheet and defect evidence package to the railway integrated maintenance system.
2. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: The timestamps of the data in step (1) satisfy: ; in and For any two modes, the sampling time is .
3. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: Step (2), spatial alignment and data preprocessing, specifically includes the following steps: (2.1) Laser point cloud and image spatial transformation: The external parameter rotation matrix and translation matrix between the lidar and the visible light camera are obtained through calibration. Based on the following homogeneous coordinate transformation formula, the laser point cloud data and image data are spatially registered. ; in This is the intrinsic parameter matrix. Here are the extrinsic rotation and translation matrices, (u,v) are the pixel coordinates, and (X... LIDAR ,Y LIDAR Z LIDAR () represents the point cloud coordinates in the radar coordinate system; (2.2) Acoustic emission localization: Using the synchronous pulse signals received by multiple acoustic emission sensors, the spatial coordinates of the sound source are calculated based on the trilateration method. ; Known spatial coordinates of each sensor The speed of sound, v Measuring the arrival time t of acoustic emission pulses at various points i Using the earliest sensing point as a reference, we set up a system of equations: The spatial location of the sound source is calculated using the least squares method, i=1,2,3,… (2.3) Data preprocessing: Adaptive histogram equalization is performed on visible light image data, voxel filtering downsampling is performed on laser point cloud data, and bandpass filtering and continuous wavelet transform are performed on acoustic emission signals to extract time-frequency energy spectrum.
4. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 3, characterized in that: In step (2.3), the point cloud data is subjected to VoxelGrid voxel filtering with a sampling resolution of δ=2mm per side; the acoustic signal is bandpass filtered with a filtering range of 10kHz~80kHz, and then the time-frequency energy spectrum of each 512ms window is extracted using continuous wavelet transform.
5. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: Step (3) multimodal feature extraction specifically includes the following steps: (3.1) Image Feature Extraction: Image features are extracted from the image using the ResNet50 network; let the single-frame input be... Its output characteristics are: ; (3.2) Point cloud feature extraction: For each frame of preprocessed point cloud Point cloud features are obtained by processing with PointNet++ network: ; (3.3) Acoustic emission feature extraction: acoustic signal after continuous wavelet transform In the formula, L is the time window length, F is the number of frequency bands, and a one-dimensional convolutional network is used to process and extract acoustic emission features. 。 6. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: The step (4) multimodal attention fusion specifically includes the following steps: stitching together image features, point cloud features and acoustic features, and using a channel attention mechanism to assign adaptive weights to the stitched multimodal features to generate fused features; After unfolding the image features, point cloud features, and acoustic features respectively, they are concatenated and combined to obtain the final concatenated feature vector: ; in The output of ResNet after global average pooling (GAP) is shown. A channel attention mechanism is employed to assign adaptive weights to the multimodal fusion features in the concatenated feature vector. A fully connected layer is defined and normalized to obtain the fusion weights α1, α2, and α3. ; final Fusion characteristics: 。 7. The automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: The step (5) defect identification and risk assessment specifically includes the following steps: based on the fusion features, perform defect segmentation and classification, determine the type of defect and its location in space; and combine historical defect data and component type metadata to predict the risk level; (5.1) Defect segmentation mask output: The fused features from step (4) are input into a multi-layer fully convolutional network, and a defect segmentation mask of the same size as the original input image is output: ; In the formula, σ is the Sigmoid function. This represents the high-value area of the mask corresponding to the defect area; (5.2) Defect classification: Input the fused features from step (4) into the classification network: ; Output a probability distribution of defect types, wherein the defect types include at least cracks, corrosion, and loose connections; (5.3) Spatial positioning: Based on the defect segmentation mask Y seg Regions where all pixel values exceed the threshold θ are mapped to the laser point cloud space, and the coordinates of the corresponding defects are calculated through the spatial transformation established in step (2.1). ; (5.4) Risk level prediction: Combine the fused features obtained in step (4) with the historical defects and component type metadata of the detected road section. meta After concatenation, the data is input into a multilayer perceptron, and the Softmax function outputs the probability distribution of risk levels. ; The risk level with the highest probability will be used as the final risk assessment level, and the risk levels will include at least Level I, Level II, and Level III.
8. An automatic detection and risk assessment device for track component defects based on multimodal data fusion and attention neural network, applicable to the automatic detection and risk assessment method for track component defects based on multimodal data fusion and attention neural network as described in claim 1, characterized in that: It includes a multimodal synchronous acquisition module, an edge processing and synchronization module, and a data fusion and discrimination module connected according to the signal flow direction; The multimodal synchronous acquisition module includes a visible light camera mounted above the detection platform, a lidar fixed next to the track component to be detected, and an acoustic emission sensor embedded in the track contact area. The edge processing and synchronization module integrates the clocks of the Global Navigation Satellite System and the Inertial Measurement Unit for synchronizing sensor sampling triggers; the processor architecture edge computing unit of the edge processing and synchronization module is used to run synchronization and preprocessing algorithms. The data fusion and discrimination module includes a graphics processing unit and a field-programmable gate array (FPGA) computing board, which is used to integrate a multimodal neural network.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.