A Deep Learning-Based Intelligent Analysis Method and System for Multimodal Heterogeneous Data

By performing timestamp synchronization and dimension normalization on multimodal data of urban rail transit trains, and combining modality-specific feature extraction and cross-modal feature alignment networks, accurate identification and graded early warning of faults in key train components were achieved, solving the problems of interactive redundancy and insufficient identification accuracy in existing multimodal data analysis technologies.

CN122087593APending Publication Date: 2026-05-26XIAMEN SOFT CLOUD NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN SOFT CLOUD NETWORK TECH CO LTD
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for monitoring the condition of key components in urban rail transit trains suffer from problems such as asynchronous timestamps from multiple data sources, differences in dimensions, misalignment of feature dimensions between modes, and time-series phase shifts. These issues lead to high redundancy in cross-modal information interaction, accumulation of forward propagation errors in deep learning models, insufficient early detection of minor faults, low warning accuracy, and frequent false alarms and missed alarms.

Method used

Multimodal data is processed by timestamp synchronization alignment and dimensional normalization. Modal-specific feature extraction network and signal-to-noise ratio adaptive cross-modal feature alignment network are used to project the feature space onto the key physical structure of the train, construct a non-coplanar three-dimensional distortion correction datum, calculate cross-modal phase alignment compensation coefficient, dynamically calibrate modal interaction weights, perform cross-modal feature fusion, and input the data into the fault identification and early warning classification network through forward propagation deviation correction processing.

Benefits of technology

It has achieved accurate identification and graded early warning of bogie faults, wheelset wear and power supply anomalies, improved the efficiency of cross-modal information fusion and the accuracy of fault identification, and reduced the false alarm and missed alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087593A_ABST
    Figure CN122087593A_ABST
Patent Text Reader

Abstract

This invention provides a deep learning-based intelligent analysis method and system for multimodal heterogeneous data, relating to the field of multimodal data processing technology. The method includes: acquiring multimodal raw data during urban rail transit train operation; performing timestamp synchronization alignment and dimension normalization on the raw multimodal data to obtain a standardized multimodal data sequence; inputting the standardized multimodal data sequence into corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels; inputting the modality feature vector sets with quality assessment labels into a signal-to-noise ratio adaptive cross-modal feature alignment network; and spatially projecting and mapping the data onto the train's key physical structures to locate three multidimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the traction motor rotor air gap monitoring point. This invention enables accurate identification and graded early warning of bogie faults, wheelset wear, and power supply anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data processing technology, and in particular to a method and system for intelligent analysis of multimodal heterogeneous data based on deep learning. Background Technology

[0002] Currently, the status monitoring of key components of urban rail transit trains mostly adopts a single sensor data source analysis mode. For example, vibration sensors are used to determine the operating status of bogies, and temperature data is used to monitor the operating condition of traction motors. Such methods are difficult to fully capture the multi-dimensional coupling characteristics in the fault evolution process. Although some solutions have attempted to introduce multi-modal data such as video and audio for auxiliary analysis, they generally suffer from problems such as asynchronous timestamps of multi-source data, differences in dimensions, misalignment of feature dimensions between modes, and time sequence phase shifts. Moreover, they often directly use conventional feature splicing methods for simple fusion without combining the actual physical structure of the train to achieve accurate feature space mapping, resulting in high redundancy of cross-modal information interaction. At the same time, existing deep learning models are prone to hierarchical error accumulation during feature forward propagation, which is insufficient for identifying early weak fault features in key weak parts such as bogie side beams, wheelset treads, and traction motor air gaps. This can easily lead to problems such as delayed fault identification, low early warning accuracy, and frequent false alarms and missed alarms, making it difficult to meet the actual needs of high reliability and intelligent operation and maintenance of urban rail transit trains. Summary of the Invention

[0003] This invention provides a method and system for intelligent analysis of multimodal heterogeneous data based on deep learning, enabling accurate identification and graded early warning of bogie faults, wheelset wear, and power supply anomalies.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, a deep learning-based intelligent analysis method for multimodal heterogeneous data, the method comprising:

[0006] The process involves acquiring multimodal raw data during the operation of urban rail transit trains, performing timestamp synchronization and alignment and dimensional normalization on the multimodal raw data to obtain standardized multimodal data sequences. These standardized multimodal data sequences are then input into corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels.

[0007] The modal feature vector set with quality assessment labels is input into a cross-modal feature alignment network with adaptive signal-to-noise ratio, and the spatial projection is mapped to the key physical structure of the train to locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0008] Based on the spatiotemporal response trajectories of three multidimensional sensor anchors, a non-coplanar three-dimensional distortion correction datum is constructed. Orthogonal meshing and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient.

[0009] The modal interaction weights within the network are dynamically calibrated by using cross-modal phase alignment compensation coefficients, and then cross-modal alignment and feature fusion are performed on the modal feature vector set to obtain a multimodal fusion feature representation.

[0010] A forward propagation bias correction process is performed on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature;

[0011] The calibrated multimodal fusion features are input into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

[0012] Furthermore, multimodal raw data during urban rail transit train operation is acquired, and the multimodal raw data undergoes timestamp synchronization alignment and dimension normalization to obtain standardized multimodal data sequences. These standardized multimodal data sequences are then input into corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels, including:

[0013] Collect time-series data from vehicle vibration sensors, video image data from on-board cameras, numerical data from bearing temperature sensors, and audio monitoring data from traction motors to form multimodal raw data.

[0014] The original multimodal data is processed by timestamp synchronization alignment and dimension normalization to eliminate clock deviations and physical dimension differences from multi-source acquisition, resulting in a standardized multimodal data sequence.

[0015] Standardized multimodal data sequences are input into the corresponding modality-specific feature extraction networks, and multi-scale spatiotemporal feature convolution operations are performed to extract the initial feature representations of each modality.

[0016] The signal-to-noise ratio evaluation index and feature sparsity are calculated simultaneously for the initial feature representation, and quality evaluation labels are generated based on the preset reliability judgment threshold.

[0017] The initial feature representations and quality assessment labels are concatenated in the feature domain to obtain the modal feature vector set.

[0018] Furthermore, the modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network, and the spatial projection is mapped onto the key physical structure of the train to locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the air gap monitoring point of the traction motor rotor. This includes:

[0019] The modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network, and spatial projection mapping to the key physical structure of the train is performed through a high-dimensional manifold coordinate transformation algorithm;

[0020] Based on the geometric mapping relationship of spatial projection mapping and the spatial distribution weight of quality assessment labels, multi-source feature correlation optimization calculation is performed;

[0021] Based on the response extreme value coordinates obtained from the correlation optimization calculation, three multi-dimensional sensor anchors are located: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor. The multi-dimensional state response sequence of the three multi-dimensional sensor anchors in each sampling time step is extracted.

[0022] Furthermore, a non-coplanar three-dimensional distortion correction datum is constructed based on the spatiotemporal response trajectories of three multi-dimensional sensing anchors. Orthogonal mesh generation and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain cross-modal phase alignment compensation coefficients, including:

[0023] An orthogonal meshing operation is performed on the non-coplanar three-dimensional distortion correction datum plane to obtain several regularly distributed mesh elements;

[0024] Extract the local strain tensor features inside each subdivided mesh element, and extend them along the calculated local principal strain direction to generate a spatially oriented envelope;

[0025] Perform spatial intersection determination operation between the spatial orientation envelope and the corresponding mesh element, extract the geometrically overlapping region and calculate the spatial overlap volume ratio to obtain the topological mask matrix that characterizes the effective range.

[0026] The topological mask matrix is ​​used as a spatial domain constraint to map the gradient distribution field of the non-coplanar three-dimensional distortion correction datum. Gradient potential field integration is performed within the effective region defined by the mask, the phase offset gradient value is accumulated and normalized, and the cross-modal phase alignment compensation coefficient is obtained.

[0027] Furthermore, the modal interaction weights within the network are dynamically calibrated using cross-modal phase alignment compensation coefficients, thereby performing cross-modal alignment and feature fusion on the modal feature vector set to obtain a multimodal fused feature representation, including:

[0028] The cross-modal phase alignment compensation coefficients are injected into the feature interaction layer of the signal-to-noise ratio adaptive cross-modal feature alignment network. The initial modal interaction weights are dynamically calibrated through weight scaling and bias compensation operations to generate a phase-compensated calibration interaction weight matrix.

[0029] The modal feature vector set with quality assessment labels is input into the feature interaction layer. Spatiotemporal coordinate mapping and phase deviation correction are performed on each modal feature by calibrating the interaction weight matrix to eliminate dimensional misalignment and temporal delay between heterogeneous data, complete cross-modal alignment processing, and obtain the aligned multimodal feature subsequence.

[0030] Cross-channel feature splicing and fully connected dimensionality reduction aggregation operations are performed on the aligned multimodal feature subsequences to fuse complementary representation information of each modality. After nonlinear transformation, a multimodal fused feature representation is obtained.

[0031] Furthermore, forward propagation bias correction is performed on the multimodal fusion feature representation to obtain the calibrated multimodal fusion features, including:

[0032] The multimodal fusion feature representation is input into the forward propagation bias correction process, the intermediate state features of each hidden layer output of the deep network are extracted layer by layer, and the layer-by-layer differential comparison operation is performed with the shallow baseline features to construct the hierarchical feature residual sequence.

[0033] Gradient magnitude normalization and error distribution equalization operations are performed on the hierarchical feature residual sequence to eliminate the statistical offset caused by local gradient fluctuations and obtain the deviation compensation mapping parameters that characterize the degree of global feature distortion.

[0034] Based on the bias compensation mapping parameters, reverse residual injection and adaptive smoothing filtering operations are performed on the multimodal fusion feature representation to remove the feature offset noise that accumulates continuously during the forward propagation of the deep network.

[0035] The feature tensors after noise stripping are subjected to dimensional reconstruction and normalization mapping to obtain calibrated multimodal fusion features.

[0036] Furthermore, the calibrated multimodal fusion features are input into the fault identification and early warning classification network to obtain the identification results and early warning levels for bogie faults, wheelset wear, and power supply anomalies, including:

[0037] The calibrated multimodal fusion features are input into the fault identification and early warning classification network. High-dimensional feature space mapping and nonlinear activation operations are performed through a fully connected classification layer to generate multi-class probability distribution sequences for bogie faults, wheelset wear and power supply anomalies.

[0038] By performing confidence threshold truncation and extreme value optimization on the multi-class probability distribution sequence, the fault category label corresponding to the confidence peak is selected to obtain the preliminary fault identification result;

[0039] The confidence peak is dynamically compared and matched with the preset early weak fault evolution threshold range, and the fault severity level range is divided according to the numerical deviation distance.

[0040] The preliminary fault identification results are correlated and mapped with the fault severity level range to obtain the final identification results and warning levels for bogie faults, wheelset wear, and power supply anomalies.

[0041] Secondly, a deep learning-based multimodal heterogeneous data intelligent analysis system includes:

[0042] The acquisition module is used to acquire multimodal raw data during the operation of urban rail transit trains. It performs timestamp synchronization and alignment and dimension normalization on the multimodal raw data to obtain a standardized multimodal data sequence. The standardized multimodal data sequence is then input into the corresponding modality-specific feature extraction network to obtain a set of modality feature vectors with quality assessment labels.

[0043] The alignment module is used to input the modal feature vector set with quality assessment labels into the signal-to-noise ratio adaptive cross-modal feature alignment network, and to map the spatial projection onto the key physical structure of the train and locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0044] The calculation module is used to construct a non-coplanar three-dimensional distortion correction datum based on the spatiotemporal response trajectories of three multi-dimensional sensor anchors. Orthogonal meshing and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient.

[0045] The fusion module is used to dynamically calibrate the modal interaction weights within the network through cross-modal phase alignment compensation coefficients, and then perform cross-modal alignment and feature fusion on the modal feature vector set to obtain a multimodal fused feature representation;

[0046] The calibration module is used to perform forward propagation bias correction on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature.

[0047] The processing module is used to input the calibrated multimodal fusion features into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

[0048] Thirdly, a computing device, comprising:

[0049] One or more processors;

[0050] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0051] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0052] The above-described solution of the present invention has at least the following beneficial effects:

[0053] This method involves time-stamp synchronization and dimensional normalization of raw multimodal data from urban rail transit trains, extracting modal feature vector sets with quality assessment labels using a modal-specific feature extraction network, and projecting the feature space onto key physical structures of the train using a signal-to-noise ratio adaptive cross-modal feature alignment network. This includes locating three multi-dimensional sensing anchors: the lateral stiffness nodes of the bogie side beams, and constructing a non-coplanar three-dimensional distortion correction datum based on the spatiotemporal response trajectories of the anchor points. Cross-modal phase alignment compensation coefficients are then calculated, and modal interaction weights are dynamically calibrated using these coefficients to achieve cross-modal feature alignment and fusion. Finally, forward propagation bias correction is performed on the multimodal fused features. The technical approach of processing and then inputting the calibrated fused features into the fault identification and early warning classification network overcomes the technical problems in existing technologies, such as the difficulty of a single data source to fully capture the multidimensional coupled features of faults, the asynchrony and inconsistent dimensions of multi-source heterogeneous data, the misalignment and offset of features between modes, the information redundancy caused by the lack of integration with the train's physical structure, the accumulation of feature propagation errors in deep learning models, the low identification accuracy of early weak faults and the low accuracy of early warning, and the frequent occurrence of false alarms and missed alarms. As a result, it achieves accurate identification and graded early warning of bogie faults, wheelset wear and power supply anomalies, improves the efficiency of cross-modal information fusion and the accuracy of fault identification, and reduces the false alarm and missed alarm rates. Attached Figure Description

[0054] Figure 1 This is a flowchart illustrating a deep learning-based intelligent analysis method for multimodal heterogeneous data, provided by an embodiment of the present invention.

[0055] Figure 2 This is a schematic diagram of a deep learning-based intelligent analysis system for multimodal heterogeneous data, provided by an embodiment of the present invention.

[0056] Figure 3 This is a simulation diagram illustrating the effect of multi-source acquisition clock synchronization and dimensional normalization.

[0057] Figure 4 This is a schematic diagram illustrating the convergence trend of three-dimensional sensor anchor positioning and correlation optimization iteration.

[0058] Figure 5 This is a statistical diagram illustrating forward propagation bias correction and characteristic residual convergence. Detailed Implementation

[0059] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art.

[0060] like Figure 1 As shown, an embodiment of the present invention proposes a method for intelligent analysis of multimodal heterogeneous data based on deep learning, the method comprising the following steps:

[0061] Step 1: Obtain multimodal raw data during the operation of urban rail transit trains; perform timestamp synchronization alignment and dimension normalization on the multimodal raw data to obtain standardized multimodal data sequences; input the standardized multimodal data sequences into the corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels.

[0062] Step 2: Input the modal feature vector set with quality assessment labels into the signal-to-noise ratio adaptive cross-modal feature alignment network, and map the spatial projection onto the key physical structure of the train to locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0063] Step 3: Construct a non-coplanar three-dimensional distortion correction datum based on the spatiotemporal response trajectories of the three multi-dimensional sensor anchors. Perform orthogonal meshing and gradient potential field integration on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient.

[0064] Step 4: Dynamically calibrate the modal interaction weights within the network using cross-modal phase alignment compensation coefficients, and then perform cross-modal alignment and feature fusion on the modal feature vector set to obtain a multimodal fusion feature representation;

[0065] Step 5: Perform forward propagation bias correction on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature;

[0066] Step 6: Input the calibrated multimodal fusion features into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

[0067] In this embodiment of the invention, multimodal raw data of urban rail transit trains are acquired and processed with timestamp synchronization alignment and dimension normalization. A modal feature vector set with quality assessment labels is obtained through a modal-specific feature extraction network. This feature vector set is input into a signal-to-noise ratio adaptive cross-modal feature alignment network and projected onto the three multi-dimensional sensor anchors of the train's key physical structure and positioning. A non-coplanar three-dimensional distortion correction datum is constructed based on the spatiotemporal response trajectory of the anchor points, and cross-modal phase alignment compensation coefficients are obtained through orthogonal mesh partitioning and gradient potential field integration. These coefficients are used to dynamically calibrate modal interaction weights and complete cross-modal feature alignment and fusion. Forward propagation is performed on the fused features. Deviation correction processing inputs the calibrated fused features into the fault identification and early warning classification network to output the identification results and early warning levels. Therefore, it overcomes the technical problems in the existing technology, such as the difficulty of a single data source to fully capture the multidimensional coupled features of faults, the asynchronous and inconsistent dimensions of multi-source heterogeneous data, the misalignment and offset of features between modes, the information redundancy caused by the lack of integration with the train's physical structure, the accumulation of feature propagation errors in deep learning models, the low identification rate and low early warning accuracy of early weak faults, and the frequent occurrence of false alarms and missed alarms. As a result, it achieves accurate identification and graded early warning of bogie faults, wheelset wear and power supply anomalies, improves the efficiency of cross-modal information fusion and the accuracy of fault identification, and reduces the false alarm and missed alarm rates.

[0068] In a preferred embodiment of the present invention, step 1 above may include:

[0069] Step 1.1 involves collecting time-series data from vehicle vibration sensors, video image data from onboard cameras, numerical data from bearing temperature sensors, and audio monitoring data from traction motors to form multimodal raw data. Specifically, this includes collecting various types of sensor data from the onboard sensing terminal throughout the entire process of normal operation and condition switching of the urban rail transit train. This includes continuously collecting time-series vibration sensor data during train operation using vibration sensors located on the side beams of the bogie, wheelsets, and the car body frame; collecting real-time video image data of the wheel-rail contact area and bogie operating status using onboard cameras installed at the bottom of the carriage and the running gear; collecting bearing temperature sensor data using temperature sensors attached to the outer wall of the bearing; and collecting audio monitoring data of the traction motor during operation using audio acquisition equipment located on the outside of the traction motor. The above four types of collected data are then aggregated to form the multimodal raw data required by this solution.

[0070] Step 1.2 involves performing timestamp synchronization alignment and dimensional normalization on the multimodal raw data to eliminate clock deviations and physical dimension differences between the multi-source acquisitions, resulting in a standardized multimodal data sequence. Specifically, this includes sequentially performing timestamp synchronization alignment on the acquired multimodal raw data, using the train's onboard unified clock as the standard reference, adjusting the acquisition time nodes of various modal data, correcting clock deviations between different acquisition devices, and ensuring that vibration data, video image data, temperature sensing data, and audio monitoring data remain consistent in the time dimension, thus completing the time synchronization of multi-source data.

[0071] The multimodal raw data after time synchronization is subjected to dimensional normalization to eliminate the dimensional differences between various physical quantities and to unify the numerical range of different modal data to the same interval. After the dual processing of time synchronization and dimensional unification, a standardized multimodal data sequence is finally formed.

[0072] Step 1.3 involves inputting the standardized multimodal data sequences into the corresponding modality-specific feature extraction networks, performing multi-scale spatiotemporal feature convolution operations, and extracting the initial feature representations for each modality. Specifically, this includes inputting the standardized vibration time-series data, video image data, temperature sensing data, and audio monitoring data into the feature extraction networks that match their modality types, and performing multi-scale spatiotemporal feature convolution operations for each type of modality data. Taking vibration time-series data and audio monitoring data as examples, a one-dimensional multi-scale convolution kernel is used to sequentially perform convolution operations along the time dimension. The convolution operation follows the principle that the temporal feature map is the sum of the pointwise products of the temporal data sequence and the corresponding scale convolution kernel. For video image data, a two-dimensional multi-scale convolution kernel is used to perform convolution operations along the spatial dimension. The convolution process follows the principle that the spatial feature map is the sum of the pointwise products of the image pixel matrix and the corresponding scale convolution kernel. Temperature sensing data is extracted with one-dimensional temporal features through a fully connected layer. Through layer-by-layer convolution operations with convolution kernels of different scales, feature information of each modality data at different temporal granularities and spatial dimensions is extracted, and finally the initial feature representation corresponding to each modality is obtained.

[0073] Step 1.4: Simultaneously calculate the signal-to-noise ratio (SNR) evaluation index and feature sparsity for the initial feature representations, and generate quality evaluation labels based on the preset reliability judgment thresholds. Specifically, this includes: simultaneously performing calculation operations on the extracted initial feature representations of each modality, calculating the SNR evaluation index and feature sparsity for each group of initial feature representations, and comparing the calculation results with the preset reliability judgment thresholds after the index calculations are completed. The reliability judgment threshold for SNR is set to 0.85, and the reliability judgment threshold for feature sparsity is set to 0.30.

[0074] When the signal-to-noise ratio of the initial feature representation is higher than 0.85 and the feature sparsity is lower than 0.30, the feature is judged as a high-quality and reliable feature; otherwise, it is judged as a low-quality feature. Based on the comparison and judgment results, a corresponding quality assessment label is generated for each set of initial feature representations.

[0075] Step 1.5 involves concatenating the initial feature representations with the quality assessment labels in the feature domain to obtain a set of modal feature vectors. Specifically, this includes concatenating the initial feature representations of each modality with the corresponding quality assessment labels in the feature domain. The concatenation process follows the principle that the modal feature vector is the concatenation of the dimensions of the initial feature representation vector and the quality assessment label vector. That is, the vector dimension of the initial feature representation remains unchanged, and the numerical vector corresponding to the quality assessment label is directly concatenated to the end of the vector. The numerical and spatial distribution characteristics of the original features are not changed during the concatenation process.

[0076] After completing the above cascading operation on the initial feature representations and quality assessment labels of all modalities, all cascaded feature vectors are summarized in a unified manner to finally form a modal feature vector set with quality assessment labels.

[0077] In this embodiment of the invention, multimodal raw data is constructed by collecting time-series data from vehicle vibration sensors, video image data from vehicle-mounted cameras, numerical data from bearing temperature sensors, and audio monitoring data from traction motors. This raw data undergoes timestamp synchronization alignment and dimensional normalization. The standardized multimodal data sequence is then input into a corresponding modality-specific feature extraction network to perform multi-scale spatiotemporal feature convolution operations to extract initial feature representations. The signal-to-noise ratio (SNR) evaluation index and feature sparsity are simultaneously calculated for the initial feature representations, and quality assessment labels are generated based on a preset reliability threshold. Finally, the initial feature representations and quality assessment labels are concatenated in the feature domain to obtain a modal feature vector set. This approach overcomes the technical problems of existing technologies, such as a single source of multimodal raw data, clock skew and dimensional differences in multi-source data, inaccurate initial feature extraction without considering feature quality, and low-quality features easily interfering with subsequent analysis. This achieves the following: enriching multimodal data sources, eliminating heterogeneous interference from multi-source data, accurately extracting effective initial features for each modality, selecting high-quality features through quality assessment labels, and improving the reliability and effectiveness of the modal feature vector set.

[0078] In a preferred embodiment of the present invention, step 2 above may include:

[0079] Step 2.1 involves inputting the modal feature vector set with quality assessment labels into a signal-to-noise ratio (SNR) adaptive cross-modal feature alignment network. A high-dimensional manifold coordinate transformation algorithm is then used to perform spatial projection mapping to the train's key physical structures. Specifically, this includes: inputting the obtained modal feature vector set with quality assessment labels into the SNR adaptive cross-modal feature alignment network. This network automatically adjusts its parameters based on the SNR of the input features to ensure the adaptability of the feature processing. Subsequently, the high-dimensional manifold coordinate transformation algorithm is initiated to perform a projection mapping operation from the high-dimensional feature space to the train's key physical structure space on the input modal feature vector set.

[0080] A three-dimensional coordinate system for the key physical structure of the train is determined. A three-dimensional rectangular coordinate system is established with the geometric center of the train body as the origin, and the specific coordinate range of key components such as bogies, wheelsets, and traction motors in this coordinate system is clarified.

[0081] The high-dimensional manifold coordinate transformation process follows the following calculation procedure: , These are the mapped physical coordinates, i.e., the coordinates projected onto the space of the train's key physical structures. These are the high-dimensional feature coordinates, i.e., the coordinates of the modal feature vectors in the original high-dimensional feature space. The transformation matrix is ​​pre-set based on the key physical structural geometric parameters of the train. The offset is the deviation between the origin of the high-dimensional feature space and the origin of the train's physical coordinates. This calculation accurately maps the high-dimensional modal feature vectors to the corresponding key physical structure space of the train, thus achieving the binding between the feature space and the train's real physical structure.

[0082] Step 2.2: Based on the geometric mapping relationship of spatial projection mapping and the spatial distribution weight of quality assessment labels, perform multi-source feature correlation optimization calculation. Specifically, after completing the high-dimensional manifold coordinate transformation and spatial projection mapping, first extract the geometric mapping relationship formed during the projection mapping process. This geometric mapping relationship is constructed based on the correspondence between high-dimensional feature coordinates and the coordinates of key physical structures of the train. It accurately represents the one-to-one correspondence between modal feature vectors in the original high-dimensional feature space and the actual physical structure space of the train, and clarifies the physical location of key train components corresponding to each modal feature vector. For example, the vibration feature vector corresponds to the specific coordinates of the bogie side beam, the video image feature vector corresponds to the specific coordinates of the wheelset tread, the temperature feature vector corresponds to the specific coordinates of the bearing, and the audio feature vector corresponds to the specific coordinates of the traction motor. Through this geometric mapping relationship, the deep binding between modal features and the actual physical structure of the train is realized.

[0083] The spatial distribution weights of the quality assessment labels are calculated. The core function of these weights is to assign higher weights to high-quality features, reduce the interference of low-quality features on correlation optimization, ensure more reliable correlation optimization results, and thus solve the problem of high redundancy in cross-modal information interaction. The signal-to-noise ratio and feature sparsity required to calculate these weights are derived from the calculation results of the initial feature representations of each modality, without the need for additional data collection or calculation, ensuring data consistency and coherence. The specific calculation process is as follows: Let the spatial distribution weight of the quality assessment label be the quality weight, the signal-to-noise ratio (SNR) of the feature be the SNR, and the feature sparsity be the feature sparsity. The calculation formula is: Quality weight = SNR × 0.7 + (1 - Feature sparsity) × 0.3, where 0.7 and 0.3 are preset weight coefficients. The reason for setting the weight of SNR to 0.7 and the weight of feature sparsity to 0.3 is that SNR directly reflects the effective signal ratio of the feature and has a greater impact on feature quality, while feature sparsity reflects the redundancy of the feature and has a relatively smaller impact. Through this weight allocation, high-quality features with high SNR and low feature sparsity receive higher spatial distribution weights, while low-quality features with low SNR and high feature sparsity receive lower spatial distribution weights, ensuring that the weight allocation is highly matched with the feature quality and providing a guarantee for the accuracy of subsequent correlation optimization.For each type of modal feature, the corresponding spatial distribution weight of the quality assessment label is calculated according to the above formula, resulting in a weight set for all modal features. After extracting the geometric mapping relationship and calculating the spatial distribution weight of the quality assessment label, a multi-source feature correlation optimization operation is performed based on these two core data. The core purpose of correlation optimization is to find the feature combination with the highest correlation between different modal features, eliminate redundant feature correlations, improve the effectiveness of cross-modal feature fusion, and solve the problem of high redundancy in cross-modal information interaction in the background technology. The specific operation process is as follows: Determine the types of modal features participating in the correlation calculation. In this scheme, there are four types of modal features: vibration sensor time series features, vehicle camera video image features, bearing temperature sensing features, and traction motor audio monitoring features. Calculate the correlation degree between any two types of modal features one by one. Let the multi-source feature correlation degree be... For correlation, the sum of the covariances of the projected coordinates of each modality feature is called the covariance sum, and the sum of the weights of the spatial distribution of the corresponding quality assessment labels is called the weight sum. The calculation formula is: Correlation = Covariance sum × Weight sum. Here, the covariance sum refers to the sum of the covariances of the projected coordinates of any two types of modal features. The covariance is used to measure the linear correlation between two types of modal features. The larger the covariance value, the stronger the correlation between the two types of modal features, and vice versa. By calculating and summing the covariances of the projected coordinates of all two types of modal features, the overall correlation between multimodal features is comprehensively reflected. The weight sum refers to the sum of the weights of the spatial distribution of the quality assessment labels corresponding to the two types of modal features involved in the correlation calculation. Through the weighting effect of the weight sum, the correlation between high-quality features is strengthened, and the correlation between low-quality features is weakened, ensuring that the correlation calculation results are more targeted and reliable.

[0084] After calculating the correlation degree of all pairwise modal feature combinations, a comparative analysis of all correlation degree values ​​is performed, and an extreme value optimization operation is executed to find the feature combination with the largest correlation degree value. This feature combination is the combination with the strongest correlation and the lowest redundancy among the multimodal features, thus completing the multi-source feature correlation degree optimization operation.

[0085] Step 2.3: Based on the response extreme value coordinates obtained from the correlation optimization calculation, locate the three multi-dimensional sensor anchors: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the traction motor rotor air gap monitoring point. Extract the multi-dimensional state response sequence of the three multi-dimensional sensor anchors within each sampling time step. Specifically, after completing the multi-source feature correlation optimization calculation, extract the response extreme value coordinates obtained from the correlation optimization calculation. These coordinates are the physical coordinates with the highest correlation to the fault-sensitive points of key train components. Locate the three multi-dimensional sensor anchors based on these response extreme value coordinates. The positioning process involves several steps. First, the lateral stiffness node of the bogie side beam is located by identifying the coordinates of the corresponding area of ​​the bogie side beam in the extreme response coordinates. This node is the most sensitive location to changes in bogie stiffness. Second, the center of the wheelset tread contact patch is located by identifying the center coordinates of the contact area between the wheelset and the track in the extreme response coordinates. Third, the air gap monitoring point of the traction motor rotor is located by identifying the core coordinates of the gap between the traction motor rotor and stator in the extreme response coordinates.

[0086] After the positioning of the three multi-dimensional sensor anchors is completed, the multi-dimensional state response sequence of each multi-dimensional sensor anchor is extracted in each sampling time step according to the sampling frequency of data acquisition. Each sampling time step corresponds to a set of anchor positioning state data. The state data of all sampling time steps are arranged in chronological order to form a complete multi-dimensional state response sequence. This sequence fully reflects the state changes of the three anchors during train operation.

[0087] In this embodiment of the invention, a modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network. A high-dimensional manifold coordinate transformation algorithm is used to spatially project and map this vector onto the key physical structure of the train. Then, based on the geometric mapping relationship of the spatial projection and the spatial distribution weights of the quality assessment labels, a multi-source feature correlation optimization operation is performed. Finally, based on the response extreme coordinates obtained from the correlation optimization, three multi-dimensional sensor anchors are located: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the air gap monitoring point of the traction motor rotor. The multi-dimensional state response sequence of each anchor is extracted within each sampling time step. Therefore, this method overcomes the technical problems of existing technologies, such as cross-modal feature alignment not being combined with the actual physical structure of the train, feature mapping being detached from actual working conditions, low multi-source feature correlation, inability to accurately locate key fault-sensitive points of the train, and incomplete extraction of anchor point state responses. This achieves accurate mapping between the modal feature space and the key physical structure of the train, improves multi-source feature correlation, accurately locates the three core sensor anchors, and comprehensively captures the spatiotemporal state responses of the anchors.

[0088] In a preferred embodiment of the present invention, step 3 above may include:

[0089] Step 3.1: Perform orthogonal meshing operation on the non-coplanar three-dimensional distortion correction datum to obtain several regularly distributed mesh elements. Specifically, this includes: clarifying the specific composition of the non-coplanar three-dimensional distortion correction datum. This datum is constructed based on the spatiotemporal response trajectories of three multi-dimensional sensor anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor. The spatiotemporal trajectories of the three anchors are not coplanar, thus forming a non-coplanar three-dimensional distortion correction datum. This datum accurately covers the sensitive areas of key train faults.

[0090] An orthogonal meshing operation is performed on the non-coplanar 3D distortion correction datum. Before meshing, the 3D spatial dimensions of the datum are calculated to determine the length, width, and height ranges. A preset number of meshes is set, with 50 meshes in the length direction, 50 in the width direction, and 30 in the height direction. By determining the specific size of each orthogonal mesh unit, the regular distribution of mesh units is ensured. According to the calculated mesh unit size, starting from the initial coordinates of the datum, the datum is divided sequentially along the three orthogonal directions of the x-axis, y-axis, and z-axis, dividing the entire non-coplanar 3D distortion correction datum into several uniformly sized and regularly arranged mesh units. Each mesh unit corresponds to a unique 3D coordinate.

[0091] Step 3.2 involves extracting the local strain tensor features within each subdivided grid cell and expanding it along the calculated local principal strain direction to generate a spatially oriented envelope. Specifically, after completing the orthogonal grid subdivision, for each subdivided grid cell, the local strain tensor features are extracted. The local strain tensor reflects the degree of deformation in the corresponding region of the grid cell and is a core indicator characterizing the distortion of the feature space. During the extraction process, the three-dimensional coordinates of the four vertices of each subdivided grid cell are obtained. Combined with the spatiotemporal response sequences of three multi-dimensional sensor anchors, the displacement change of each grid cell vertex at different sampling time steps is calculated. The displacement change is the difference between the vertex coordinates at the current sampling time step and the vertex coordinates at the initial sampling time step.

[0092] The local strain tensor is calculated based on the displacement change. The calculation process involves finding the symmetric part of the partial derivative matrix of the displacement change. Specifically, by taking the partial derivatives of the displacement changes in the x, y, and z directions, a three-dimensional partial derivative matrix is ​​constructed. Then, half the sum of this matrix and its transpose is taken to obtain the symmetric form of the local strain tensor. The specific calculation formula is as follows: ,in For local strain tensor, The partial derivative matrix constructed for the displacement change, The tensor is the transpose of the partial derivative matrix. This tensor comprehensively reflects the stretching, compression, and shear deformation of the mesh elements in the x, y, and z directions.

[0093] After extracting the local strain tensor, eigenvalue decomposition is performed on the strain tensor to extract eigenvalues ​​and corresponding eigenvectors. The direction of the eigenvector corresponding to the largest eigenvalue is determined as the local principal strain direction, which is the spatial direction in which the deformation of the mesh element is most significant.

[0094] Starting from the geometric center of the mesh element, linear expansion is performed along both the positive and negative directions of the local principal strain. The expansion length is the local principal strain value multiplied by a preset expansion coefficient of 0.1. The resulting vertices are then connected to construct the original spatially oriented envelope. This original envelope is an irregular convex polyhedron structure, initially covering the deformation-affected area of ​​the mesh element. However, its irregular shape will affect the accuracy of subsequent spatial intersection determination. Therefore, the OBB bounding box computational geometry algorithm for polyhedra needs to be applied to the original spatially oriented envelope to perform regularization and optimization.

[0095] The process begins by extracting the vertex set of the original envelope. This involves reading the coordinates of all 3D vertices from the original spatially oriented envelope generated in the previous step, and then compiling and summarizing these coordinates to form a complete vertex set. This vertex set contains all the spatial positional information of the original envelope and serves as the sole input data for subsequent OBB bounding box calculations, ensuring reliable basic data support for the computation process. Next, the centroid of the vertex set is calculated using a polyhedral centroid calculation method. Specifically, the total number of vertices in the vertex set is first counted, and then the average values ​​of the x, y, and z coordinates of all vertices are calculated. The 3D coordinate point corresponding to these three average values ​​is the geometric centroid of the vertex set, which will serve as the central reference point for the subsequent OBB bounding box.

[0096] To eliminate the offset interference inherent in the vertex coordinates and ensure more accurate subsequent calculations, a centralized vertex matrix is ​​constructed. The x, y, and z coordinates of each vertex in the vertex set are subtracted from the x, y, and z coordinates of the geometric centroid obtained in the previous step, respectively, to obtain the centralized coordinates of each vertex. These centralized coordinates are then organized into a matrix to form the centralized vertex matrix. This method standardizes the data format and reduces the impact of coordinate offsets on subsequent calculations. Based on the constructed centralized vertex matrix, a three-dimensional covariance matrix is ​​further constructed. This matrix accurately reflects the distribution characteristics of the vertex set in three-dimensional space. Eigenvalue decomposition is performed on the constructed three-dimensional covariance matrix, yielding three mutually perpendicular eigenvectors. These three eigenvectors determine the local coordinate axis directions of the OBB bounding box, and these directions maintain spatial adaptation with the previously calculated local principal strain directions, ensuring that the bounding box accurately conforms to the deformation direction of the original envelope.

[0097] To determine the size and extent of the OBB bounding box, project all vertices of the original spatially oriented envelope onto the three local coordinate axes obtained in the previous step. For each local coordinate axis, find the maximum and minimum values ​​of all projected vertices. The difference between these two extreme values ​​is the size of the OBB bounding box along that coordinate axis. The dimensions along the three coordinate axes correspond to the length, width, and height of the bounding box, thus defining the specific spatial extent of the OBB bounding box. Finally, generate a regular OBB bounding box. Using the geometric centroid calculated in the previous step as the center of the OBB bounding box, and the three mutually orthogonal eigenvectors as the local coordinate axis directions of the bounding box, and using the determined length, width, and height as the dimensions of the bounding box, gradually construct a regular OBB bounding box. This regular OBB bounding box is the optimized spatially oriented envelope, which will replace the original irregularly shaped envelope.

[0098] Step 3.3 involves performing a spatial intersection determination operation between the spatially oriented envelope and its corresponding mesh element, extracting the geometrically overlapping region, and calculating the proportion of the spatially overlapping volume to obtain a topological mask matrix representing the effective range. Specifically, this includes: for each mesh element and its corresponding spatially oriented envelope, performing a spatial intersection determination operation to determine whether there is a geometrically overlapping region between the spatially oriented envelope and its corresponding mesh element. The core of the intersection determination is to compare the three-dimensional coordinate ranges of the two, that is, to determine whether there is an intersection between the x-axis, y-axis, and z-axis coordinate ranges of the envelope and the corresponding coordinate ranges of the mesh element. If there is an intersection in all three axes, it is determined that there is a geometrically overlapping region; otherwise, it is determined that there is no overlapping region. For mesh elements and envelopes with geometrically overlapping regions, the three-dimensional coordinate ranges of the geometrically overlapping region are further extracted, the spatial volume of the overlapping region is calculated, and the total volume of the corresponding spatially oriented envelope is calculated simultaneously.

[0099] The spatial overlap volume ratio is calculated, which represents the effective range of the envelope on the mesh cell. The higher the ratio, the greater the influence of the envelope on the mesh cell, and the more important it is to correct the feature distortion of the corresponding region. A topological mask matrix is ​​generated based on the spatial overlap volume ratio. The dimension of the topological mask matrix is ​​the same as the number of mesh cells. Each matrix element corresponds to a mesh cell. When the spatial overlap volume ratio of a mesh cell is greater than 0.5, the matrix element corresponding to that mesh cell is set to 1, indicating that the region is an effective region. When the spatial overlap volume ratio is less than or equal to 0.5, the corresponding matrix element is set to 0, indicating that the region is an ineffective region. Finally, the topological mask matrix representing the effective range is obtained.

[0100] Step 3.4: Map the topological mask matrix as a spatial domain constraint to the gradient distribution field of the non-coplanar 3D distortion correction datum. Perform gradient potential field integration within the effective region defined by the mask, accumulate the phase shift gradient values, and perform normalized mapping to obtain the cross-modal phase alignment compensation coefficients. Specifically, the obtained topological mask matrix is ​​used as a spatial domain constraint and mapped to the gradient distribution field of the non-coplanar 3D distortion correction datum. The gradient distribution field is constructed based on the phase shift of multimodal features, reflecting the phase shift gradient of each region of the datum. The larger the gradient value, the more obvious the modal phase shift in that region. During the mapping process, the region with an element of 1 in the topological mask matrix retains its corresponding gradient distribution field data as the effective region for subsequent integration operations; the region with an element of 0 masks the corresponding gradient distribution field data and does not participate in the integration operation, thereby achieving precise constraint on the integration range.

[0101] Within the effective region defined by the mask, gradient potential field integration is performed. This integration process involves integrating the gradient value for each subdivided mesh cell within the effective region. The integral value for each mesh cell is the product of its gradient value and volume. The integral values ​​of all mesh cells within the effective region are summed to obtain the cumulative phase shift gradient value. After calculating the cumulative phase shift gradient value, it undergoes normalization mapping. The minimum cumulative gradient value is the minimum of the integral values ​​of all mesh cells within the effective region, and the maximum cumulative gradient value is the maximum of the integral values ​​of all mesh cells within the effective region. This normalization mapping maps the cumulative phase shift gradient value to the interval between 0 and 1, ultimately yielding the cross-modal phase alignment compensation coefficient.

[0102] In this embodiment of the invention, the technical means of orthogonally meshing the non-coplanar three-dimensional distortion correction datum, extracting the local strain tensor features of each mesh element and generating a spatially oriented envelope along the principal strain direction, obtaining the topological mask matrix through spatial intersection judgment operation, and then using this matrix as the spatial domain constraint to perform gradient potential field integration, accumulate phase shift gradient and normalize mapping to obtain the cross-modal phase alignment compensation coefficient, thus overcoming the technical problems of difficult accurate correction of multimodal feature spatial distortion and phase shift, redundant and ineffective gradient calculation range, and lack of spatial constraints on phase compensation coefficient, thereby achieving accurate acquisition of cross-modal phase alignment compensation coefficient and improving the accuracy of phase compensation.

[0103] In a preferred embodiment of the present invention, step 4 above may include:

[0104] Step 4.1 involves injecting the cross-modal phase alignment compensation coefficients into the feature interaction layer of the signal-to-noise ratio adaptive cross-modal feature alignment network. The initial modal interaction weights are then dynamically calibrated through weight scaling and bias compensation operations to generate a phase-compensated calibration interaction weight matrix. Specifically, this includes obtaining the obtained cross-modal phase alignment compensation coefficients, denoted as... This coefficient accurately characterizes the phase shift of multimodal features in different regions and is the core basis for achieving dynamic calibration of modal interaction weights. Subsequently, this coefficient is fully injected into the feature interaction layer of the signal-to-noise ratio adaptive cross-modal feature alignment network. The feature interaction layer is the core module responsible for the fusion of features across different modalities, and it contains pre-set initial modal interaction weights, denoted as . The weights are fixed weights preset based on the initial importance of each modal feature, without considering the impact of phase shift on feature interaction. This is also the key problem in the existing technology where the modal interaction weights are fixed and cannot adapt to changes in phase shift.

[0105] To achieve dynamic calibration of the initial modal interaction weights, weight scaling and bias compensation operations need to be performed. The specific calculation process is as follows: Calculate the weight scaling factor, denoted as... The calculation formula is: ,in To assess the spatial distribution weights of the labels, the calculation results from step 2.2 are used. This calculation assigns higher scaling weights to modes with significant phase shifts and high feature quality, ensuring that weight calibration is adapted to both phase shift and feature quality. Subsequently, the calibrated modal interaction weights are calculated, denoted as... The calculation formula is: Among them, the bias compensation amount The calculation formula is: 0.05 is the preset bias coefficient. These are cross-modal phase alignment compensation coefficients used to compensate for possible deviations during weight scaling and to avoid feature distortion caused by over-scaling of weights.

[0106] The above weight scaling and bias compensation operations are performed on each element in the initial modal interaction weights. All calibrated modal interaction weights are arranged in matrix form to generate a phase-compensated calibrated interaction weight matrix. The dimension of this matrix is ​​the same as that of the initial modal interaction weight matrix. Each element corresponds to an interaction weight between modes and is dynamically adjusted according to phase shift and feature quality.

[0107] Step 4.2: Input the modal feature vector set with quality assessment labels into the feature interaction layer. Perform spatiotemporal coordinate mapping and phase deviation correction on each modal feature through the calibration interaction weight matrix to eliminate dimensional misalignment and temporal delay between heterogeneous data, complete the cross-modal alignment process, and obtain the aligned multimodal feature subsequence. Specifically, after generating the calibration interaction weight matrix, input the obtained modal feature vector set with quality assessment labels into the feature interaction layer of the signal-to-noise ratio adaptive cross-modal feature alignment network to ensure that the input feature vector set and the calibration interaction weight matrix are accurately matched, providing basic data for feature alignment operation.

[0108] By calibrating the interaction weight matrix, spatiotemporal coordinate mapping and phase deviation correction operations are performed on each modal feature to eliminate dimensional misalignment and temporal delay issues between heterogeneous data. The calculation process for spatiotemporal coordinate mapping is as follows: Let the original modal feature's spatiotemporal coordinates be... The mapped spacetime coordinates are The calculation formula is: This calculation maps the feature coordinates of different modes to the same spatial coordinate system, solving the problem of feature dimension misalignment between modes; the phase deviation correction calculation process is as follows: let the original mode feature phase be... The corrected characteristic phase is , To calibrate the corresponding elements of the interaction weight matrix, The cross-modal phase alignment compensation coefficient is calculated using the following formula: This calculation compensates for the phase shift between different modal features. Combined with the foundation of timestamp synchronization, it further eliminates timing delays and ensures that each modal feature remains consistent in both time and space dimensions.

[0109] During the spatiotemporal coordinate mapping and phase deviation correction process, the screening function of quality assessment labels is combined simultaneously. For features judged as low quality by the quality assessment labels, the corresponding weights in the calibration interaction weight matrix are appropriately reduced to reduce the interference of low-quality features on the alignment results. After completing the spatiotemporal coordinate mapping and phase deviation correction of all modal features, each modal feature is classified according to modal type. Each type of modal feature forms an independent feature sequence. These sequences have achieved precise alignment in the time and space dimensions, which are the aligned multimodal feature subsequences.

[0110] Step 4.3 involves performing cross-channel feature splicing and fully connected dimensionality reduction aggregation operations on the aligned multimodal feature subsequences to fuse complementary representation information from each mode. After nonlinear transformation, a multimodal fused feature representation is obtained. Specifically, this includes performing cross-channel feature splicing operations on the obtained aligned multimodal feature subsequences. The splicing process follows a preset order of modal types, specifically splicing the vibration time-series feature subsequence, video image feature subsequence, temperature sensing feature subsequence, and traction motor audio feature subsequence sequentially. During the splicing process, the feature dimensions of each subsequence remain unchanged; expansion is only performed on the feature channel dimensions. Let the dimension of each modal feature subsequence be... Feature dimensions after splicing The calculation formula is: By splicing across channels, features from different modalities are integrated into the same feature space, thus initially realizing the aggregation of multimodal features.

[0111] After concatenation, a fully connected dimensionality reduction and aggregation operation is performed on the concatenated features. The purpose is to compress the feature dimension, remove redundant information, and fuse complementary representation information from each modality. The calculation process of the fully connected dimensionality reduction and aggregation is as follows: Set the weight matrix of the fully connected layer. and bias vector The weight matrix of the fully connected layer has the dimension of the concatenated feature dimension and the target dimension reduction dimension. The target dimension reduction dimension is preset to 1024 based on subsequent fault identification requirements, and the dimension of the bias vector is consistent with the target dimension reduction dimension. Let the concatenated feature be... The features after dimensionality reduction are The calculation formula is: This calculation maps high-dimensional spliced ​​features to a preset low-dimensional feature space, preserving the core complementary information of each modality and eliminating useless redundant information during the dimensionality reduction process, thereby improving the effectiveness of the features.

[0112] After performing fully connected dimensionality reduction and aggregation, a nonlinear transformation is performed on the reduced features using a rectified linear unit transformation. Let the features after the nonlinear transformation be... The calculation formula is: By suppressing negative interference information in the features through this nonlinear transformation, the nonlinear representation ability of the features is enhanced, the deep correlation between the features of each modality is further explored, and the recognizability of the fused features is improved. After a series of processes such as cross-channel splicing, fully connected dimensionality reduction aggregation and nonlinear transformation, the multimodal fused feature representation is finally obtained. This feature representation integrates the complementary information of all modalities and eliminates redundant and interference information.

[0113] In this embodiment of the invention, cross-modal phase alignment compensation coefficients are injected into the feature interaction layer, and the initial modal interaction weights are dynamically calibrated through weight scaling and bias compensation operations to generate a calibration interaction weight matrix. This matrix is ​​then used to perform spatiotemporal coordinate mapping and phase deviation correction on the modal features. Finally, the aligned multimodal feature subsequences are subjected to cross-channel splicing, fully connected dimensionality reduction aggregation, and nonlinear transformation. This overcomes the technical problems of fixed and single modal interaction weights, dimensional misalignment and temporal delay in heterogeneous data, and the inability of multimodal features to be effectively complementary and fused. As a result, the technical effects of accurately completing cross-modal feature alignment, fully fusing complementary representation information of each modality, and improving the integrity and effectiveness of multimodal fusion feature representation are achieved.

[0114] In a preferred embodiment of the present invention, step 5 above may include:

[0115] Step 5.1 involves inputting the multimodal fusion feature representation into the forward propagation bias correction process, extracting intermediate state features from the outputs of each hidden layer of the deep network layer by layer, and performing layer-by-layer difference comparison operations with the shallow baseline features to construct a hierarchical feature residual sequence. Specifically, this includes obtaining the multimodal fusion feature representation, denoted as... This feature representation integrates complementary information from various modalities and removes redundant interference. It serves as the basic input data for subsequent bias correction. The multimodal fusion feature representation is fully input into the forward propagation bias correction process. This process is synchronized with the forward propagation process of the deep learning model and captures the accumulation of hierarchical errors in the feature propagation process in real time.

[0116] Extracting intermediate state features from the outputs of each hidden layer of a deep network layer by layer, assuming the deep network contains... There are 1 hidden layer, from the first hidden layer to the 2nd hidden layer. Hidden layers are constructed, and intermediate state features from the outputs of each layer are extracted sequentially, denoted as follows: ,in These are intermediate state features output by the shallow hidden layer. The intermediate state features output by the deepest hidden layer are used to represent the intermediate state features output by the shallowest hidden layer. As a shallow baseline feature, this feature is least affected by forward propagation errors.

[0117] Perform layer-by-layer difference comparison operations on the intermediate state features of each layer and the shallow baseline features to construct a hierarchical feature residual sequence. The specific calculation process is as follows: for the ... layer( Values ​​range from 1 to intermediate state characteristics The corresponding residual The calculation formula is: This calculation accurately captures the offset of each layer's intermediate state features relative to the shallow layer's baseline features. A larger offset indicates a more severe accumulation of forward propagation error in that layer. The residuals of all layers are then analyzed. Arranged sequentially according to the hidden layer order, forming a complete hierarchical feature residual sequence. This sequence visually reflects the distribution of error accumulation during the forward propagation of a deep network.

[0118] Step 5.2 involves performing gradient magnitude normalization and error distribution equalization operations on the hierarchical feature residual sequence to eliminate statistical bias caused by local gradient fluctuations and obtain bias compensation mapping parameters characterizing the degree of global feature distortion. Specifically, this includes completing the hierarchical feature residual sequence... After construction, gradient magnitude normalization is performed on the sequence to eliminate magnitude differences between residuals at different levels and avoid statistical offset problems caused by excessive local gradient fluctuations. The specific calculation process is as follows: Calculate the hierarchical feature residual sequence. gradient magnitude gradient magnitude The rate of change of the residual sequence is characterized by the following formula: in This is the gradient operator, used to calculate the rate of change of the residual sequence in each dimension, and to evaluate the gradient magnitude. Normalization is performed to map the gradient magnitude to the range of 0 to 1, effectively suppressing the impact of local gradient fluctuations and ensuring the consistency of gradient information.

[0119] An error distribution balancing operation is performed on the hierarchical feature residual sequence to eliminate the problem of uneven error distribution in the residual sequence, making the error distribution more stable and improving the accuracy of bias compensation. The specific calculation process is as follows: calculate the hierarchical feature residual sequence... Error distribution variance ,variance The degree of dispersion of the error distribution is characterized by the following formula: ,in Given the mean of the hierarchical feature residual sequence, an error distribution equalization operation is performed based on the variance. Let the equalized error be... The calculation formula is: , where 0.001 is a preset small constant used to avoid the denominator being 0 when the variance is 0, thus ensuring the stability of the operation.

[0120] After completing the gradient magnitude normalization and error distribution equalization operations, the normalized gradient magnitude is... Error after equalization Performing a product operation yields the bias compensation mapping parameters that characterize the degree of global feature distortion. The calculation formula is: This parameter comprehensively reflects the global distortion of multimodal fusion features during the forward propagation of deep networks. The higher the degree of distortion, the more significant the bias compensation mapping parameter. The larger the value, the better.

[0121] Step 5.3 involves performing inverse residual injection and adaptive smoothing filtering operations on the multimodal fusion feature representation based on the bias compensation mapping parameters to remove feature offset noise that accumulates continuously during the forward propagation of the deep network. Specifically, this includes: based on the obtained bias compensation mapping parameters... The obtained multimodal fusion feature representation The reverse residual injection operation is performed. The core of reverse residual injection is to inject the hierarchical feature residual sequence in reverse into the multimodal fusion features, thus offsetting the feature shift error accumulated during the forward propagation of the deep network. The specific calculation process is as follows: Let the hierarchical feature residual sequence be and the features after reverse residual injection be . The calculation formula is: Through this calculation, the deviation compensation mapping parameters The system dynamically adjusts the weight of residual injection based on the degree of global feature distortion. The higher the degree of distortion, the greater the weight of residual injection, which accurately offsets the corresponding feature shift and suppresses the impact of hierarchical error accumulation.

[0122] After completing the reverse residual injection, the injected features are... The purpose of performing adaptive smoothing filtering is to remove feature offset noise that accumulates during the forward propagation of deep networks, thereby further improving the purity of features. The size of the adaptive smoothing filter kernel is adjusted according to the bias compensation mapping parameters. Dynamic adjustment, specifically based on the filter kernel size. ,in The function is the floor function, and 5 is a preset adjustment coefficient to ensure that the filter kernel size is always an odd number, thus guaranteeing the symmetry of the filtering process. Let the characteristics after smoothing filtering be... The Gaussian filtering algorithm is used, and the calculation formula is as follows: ,in It is a Gaussian filter function. For features injected with inverse residuals, this filtering operation adaptively filters out feature offset noise under different distortion levels, retains the core effective information in the features, and removes redundant noise.

[0123] Step 5.4 involves performing dimensionality reconstruction and normalization mapping on the noise-removed feature tensor to obtain the calibrated multimodal fusion features. Specifically, this includes: after feature offset and noise removal, processing the smoothed feature tensor... Dimension reconstruction is performed. Since the feature dimensions may have slight deviations during forward propagation bias correction, the purpose of dimension reconstruction is to adjust the dimensions of the feature tensors to match the target dimensionality reduction dimension. The target dimensionality reduction dimension after fully connected dimensionality reduction aggregation is preset to 1024; therefore, the feature dimensions after dimension reconstruction are also adjusted to 1024. Let the reconstructed features be... The calculation formula is: in This is the dimension reconstruction function, which reconstructs the filtered feature tensor into a 1024-dimensional feature vector, ensuring the uniformity and standardization of feature dimensions.

[0124] Features after dimensional reconstruction Standardization mapping is performed to eliminate amplitude differences between feature values, making the feature distribution more uniform. The feature values ​​are mapped to the range of a standard normal distribution with a mean of 0 and a standard deviation of 1, further improving the stability and consistency of the features. After dual processing of dimension reconstruction and standardization mapping, the calibrated multimodal fusion features are finally obtained.

[0125] In this embodiment of the invention, a hierarchical feature residual sequence is constructed by performing layer-by-layer feature difference comparison on the multimodal fusion feature representation. Gradient magnitude normalization and error distribution equalization operations are performed on the residual sequence to obtain the deviation compensation mapping parameter. Based on this parameter, inverse residual injection and adaptive smoothing filtering are performed on the fusion feature to remove feature offset noise. Finally, the feature tensor is reconstructed in dimension and normalized. Therefore, this method overcomes the technical problems of feature error accumulation, statistical offset caused by local gradient fluctuations, and difficulty in removing feature offset noise during the forward propagation of deep networks. This achieves effective calibration of fusion features and improves feature purity and stability.

[0126] In a preferred embodiment of the present invention, step 6 above may include:

[0127] Step 6.1: Input the calibrated multimodal fusion features into the fault identification and early warning classification network. Perform high-dimensional feature space mapping and nonlinear activation operations through a fully connected classification layer to generate multi-category probability distribution sequences for bogie faults, wheelset wear, and power supply anomalies. Specifically, this includes: obtaining the calibrated multimodal fusion features, which have completely removed the errors and offset noise accumulated during the forward propagation of the deep network, significantly improving feature purity, stability, and consistency, accurately representing the operating status of key train components, and providing high-precision feature input for fault identification and early warning; and inputting the calibrated multimodal fusion features completely into the fault identification and early warning classification network. This network is a deep learning network specifically designed for fault identification and early warning level determination of key train components. Its core module is a fully connected classification layer, which is responsible for mapping the high-dimensional fusion features to the fault category probability space to achieve preliminary fault category discrimination.

[0128] A high-dimensional feature space mapping operation is performed through a fully connected classification layer, mapping the 1024-dimensional calibrated multimodal fusion features to a 3-dimensional fault category space, corresponding to three types of faults: bogie faults, wheelset wear, and power supply anomalies. The calibrated multimodal fusion features are denoted as... The weight matrix of the fully connected classification layer is denoted as... The bias vector is denoted as The mapped linear features are denoted as The calculation formula is: This calculation transforms high-dimensional fusion features into linear features corresponding to the three types of faults, thus achieving dimensionality transformation of the feature space.

[0129] A nonlinear activation operation is performed on the mapped linear features using the softmax activation function to map the linear features to the interval between 0 and 1, obtaining the probability values ​​of various faults. This ensures that the sum of the probabilities of all faults is 1. The probability of a fault class is denoted as The natural constant is denoted as , The linear features output by the fully connected classification layer The characteristic components corresponding to the type of fault, The temporary traversal indices used in the summation process are also 1, 2, and 3, used to traverse all three types of faults. The linear features output by the fully connected classification layer The characteristic components corresponding to a fault class are calculated using the following formula: ,in These correspond to three types of faults: bogie failure, wheelset wear, and power supply anomaly. Through this nonlinear activation operation, a multi-category probability distribution sequence is generated for bogie failure, wheelset wear, and power supply anomaly.

[0130] Step 6.2 involves performing confidence threshold truncation and extreme value optimization operations on the multi-class probability distribution sequence to filter out the fault category labels corresponding to the confidence peak, thus obtaining preliminary fault identification results. Specifically, after generating the multi-class probability distribution sequence, a confidence threshold truncation operation is performed on the sequence to eliminate fault categories with excessively low probabilities, reduce false alarms, and address the problem of frequent false alarms and missed alarms in existing technologies. Considering the actual needs of urban rail transit train operation and maintenance, a confidence threshold of 0.5 is preset; that is, when the probability value of a certain type of fault is lower than 0.5, it is determined that this type of fault has not occurred. The probability value of a fault is truncated to 0. When the probability value is higher than or equal to 0.5, the probability value is retained. The formula for calculating the confidence threshold truncation is: the truncated probability of the i-th type of fault = the original probability of the i-th type of fault (when the original probability of the i-th type of fault is ≥ 0.5); the truncated probability of the i-th type of fault = 0 (when the original probability of the i-th type of fault is < 0.5), where i corresponds to 1, 2, and 3, respectively, representing bogie fault, wheelset wear, and power supply abnormality. The truncated probability of the i-th type of fault is the truncated probability value of each type of fault. Through this calculation, invalid fault information with too low a probability is effectively filtered out, reducing the risk of false alarms.

[0131] The core of performing extreme value optimization on the truncated probability distribution sequence is to find the confidence peak, i.e., the maximum probability value, in the probability distribution sequence. The fault category corresponding to this peak is the most likely fault type. The confidence peak is set as the confidence peak value, and the corresponding fault category index is set as the fault category index. The calculation formula is: Confidence peak value = maximum value (first probability value after truncation, second probability value after truncation, third probability value after truncation); Fault category index = index corresponding to the maximum value (when the maximum value is the first probability value after truncation, the index is 1; when it is the second probability value after truncation, the index is 2; when it is the third probability value after truncation, the index is 3).

[0132] Based on the fault category index, the corresponding fault category labels are filtered out: when the fault category index is 1, the fault category label is bogie fault; when the fault category index is 2, the fault category label is wheelset wear; when the fault category index is 3, the fault category label is power supply abnormality. The filtered fault category labels are used as the preliminary fault identification results, which preliminarily determine the most likely fault type of the train at present.

[0133] Step 6.3 involves dynamically comparing and matching the confidence peak with the preset early weak fault evolution threshold range, and dividing the fault severity level range according to the numerical deviation distance. Specifically, after generating the preliminary fault identification result, if the preliminary identification result is no fault, there is no need to divide the fault severity level range; if the preliminary identification result is a certain type of fault, the obtained confidence peak is dynamically compared and matched with the preset early weak fault evolution threshold range, and the fault severity level range is divided according to the numerical deviation distance between the two, thus solving the problems of low early warning accuracy and coarse level division in the existing technology. Combining the fault evolution law of key components of urban rail transit trains, the preset early weak fault evolution threshold range is divided into four levels, corresponding to early weak fault, light fault, moderate fault, and severe fault, respectively. The specific threshold range is set as follows: early weak fault 0.5 ≤ confidence peak < 0.6, light fault 0.6 ≤ confidence peak < 0.75, moderate fault 0.75 ≤ confidence peak < 0.9, and severe fault confidence peak ≥ 0.9.

[0134] Calculate the numerical deviation distance between the peak confidence level and the upper limit of the corresponding threshold interval. Let the numerical deviation distance be denoted as and the upper limit of the corresponding threshold interval as the upper limit of the interval. Using the numerical deviation distance, further refine the fault severity level intervals: When the peak confidence level is in the early, weak fault interval, if the numerical deviation distance < 0.05, it is determined to be in the late stage of the early, weak fault; if the numerical deviation distance ≥ 0.05, it is determined to be in the early stage of the early, weak fault. When the peak confidence level is in the mild fault interval, if the numerical deviation distance < 0.075, it is determined to be in the late stage of the mild fault; if the numerical deviation distance ≥ 0.075, it is determined to be in the late stage of the mild fault. If the confidence peak value is within the moderate fault range, and the numerical deviation distance is <0.075, it is determined to be in the late moderate fault range; if the numerical deviation distance is ≥0.075, it is determined to be in the early moderate fault range. If the confidence peak value is within the severe fault range, and the numerical deviation distance is <0.05 (i.e., confidence peak value ≥0.95), it is determined to be in the late severe fault range; if the numerical deviation distance is ≥0.05 (i.e., 0.9 ≤ confidence peak value <0.95), it is determined to be in the early severe fault range. Through the above dynamic comparison and matching and numerical deviation distance calculation, the fault severity level range is accurately divided.

[0135] Step 6.4 involves performing feature association mapping between the preliminary fault identification results and the fault severity level intervals to obtain the final identification results and warning levels for bogie faults, wheelset wear, and power supply anomalies. Specifically, this includes: after completing the division of fault severity level intervals, performing feature association mapping between the obtained preliminary fault identification results and the divided fault severity level intervals to establish a one-to-one correspondence between fault categories and severity levels, and finally obtaining the identification results and warning levels for bogie faults, wheelset wear, and power supply anomalies, thus meeting the actual needs of high reliability and intelligent operation and maintenance of urban rail transit trains.

[0136] Extract fault category labels from the preliminary fault identification results to determine the fault type; extract the corresponding fault severity level range to determine the fault evolution stage; bind the fault type with the severity level range, and set a warning level according to the severity level range. The warning level is divided into four levels, which correspond one-to-one with the fault severity level: early and weak faults correspond to level four warning, mild faults correspond to level three warning, moderate faults correspond to level two warning, and severe faults correspond to level one warning, where level one warning is the highest level and level four warning is the lowest level.

[0137] If the initial fault identification result is a bogie fault, and the severity level range is in the early or late stage of an early minor fault, then the final identification result is an early minor bogie fault, and the warning level is level four; if the severity level range is in the early or late stage of a minor fault, then the final identification result is a minor bogie fault, and the warning level is level three; if the severity level range is in the early or late stage of a moderate fault, then the final identification result is a moderate bogie fault, and the warning level is level two; if the severity level range is in the early or late stage of a severe fault, then the final identification result is a severe bogie fault, and the warning level is level one.

[0138] The association mapping rules for wheelset wear and power supply anomalies are consistent with those for bogie faults. Specifically, wheelset wear corresponds to different severity levels, generating Level 4 warnings for early-stage weak wheelset wear, Level 3 warnings for mild wheelset wear, Level 2 warnings for moderate wheelset wear, and Level 1 warnings for severe wheelset wear. Similarly, power supply anomalies correspond to different severity levels, generating Level 4 warnings for early-stage weak power supply anomalies, Level 3 warnings for mild power supply anomalies, Level 2 warnings for moderate power supply anomalies, and Level 1 warnings for severe power supply anomalies. If the initial fault identification result is no fault, the final identification result is no anomaly in the key train components, and the warning level is no warning. Through this feature association mapping, the fault identification result is accurately bound to the warning level, intuitively presenting the fault status and risk level of the key train components.

[0139] In this embodiment of the invention, the calibrated multimodal fusion features are input into the fault identification and early warning classification network. A multi-category fault probability distribution sequence is generated through a fully connected classification layer. The probability sequence is truncated with a confidence threshold and optimized for extreme values ​​to obtain preliminary fault identification results. The confidence peak value is dynamically compared with the early weak fault evolution threshold interval to classify the fault severity level. The identification results are then correlated with the severity level. Therefore, this technical means overcomes the technical problems of inaccurate fault type identification, inability to distinguish fault severity, difficulty in identifying early weak faults, and coarse early warning level classification. As a result, it achieves accurate identification of bogie faults, wheelset wear, and power supply anomalies, realizes fault severity classification and refined early warning, and improves the reliability of fault diagnosis for rail transit trains.

[0140] like Figure 2 As shown, embodiments of the present invention also provide a multimodal heterogeneous data intelligent analysis system based on deep learning, comprising:

[0141] The acquisition module is used to acquire multimodal raw data during the operation of urban rail transit trains. It performs timestamp synchronization and alignment and dimension normalization on the multimodal raw data to obtain a standardized multimodal data sequence. The standardized multimodal data sequence is then input into the corresponding modality-specific feature extraction network to obtain a set of modality feature vectors with quality assessment labels.

[0142] The alignment module is used to input the modal feature vector set with quality assessment labels into the signal-to-noise ratio adaptive cross-modal feature alignment network, and to map the spatial projection onto the key physical structure of the train and locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0143] The calculation module is used to construct a non-coplanar three-dimensional distortion correction datum based on the spatiotemporal response trajectories of three multi-dimensional sensor anchors. Orthogonal meshing and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient.

[0144] The fusion module is used to dynamically calibrate the modal interaction weights within the network through cross-modal phase alignment compensation coefficients, and then perform cross-modal alignment and feature fusion on the modal feature vector set to obtain a multimodal fused feature representation;

[0145] The calibration module is used to perform forward propagation bias correction on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature.

[0146] The processing module is used to input the calibrated multimodal fusion features into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

[0147] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0148] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0149] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0150] Experimental example:

[0151] I. Experiment Overview

[0152] This experiment selected an operating train from Metro Line 3 in a certain city as the experimental subject. The line is 38.6 km long with 28 stations and an average daily passenger volume of approximately 860,000. The experimental train is a 6-car Type A train with a maximum operating speed of 80 km / h and a cumulative operating mileage of approximately 1.2 million km. The experiment focused on monitoring three key physical structures of the train: the lateral stiffness nodes of the bogie side beams, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0153] The experimentally collected multimodal raw data included four modes: time-series data from vehicle vibration sensors (IEPE accelerometer, sampling frequency 5120Hz, 16 channels), video image data from onboard cameras (industrial camera, resolution 1920x1080, frame rate 30fps), numerical data from bearing temperature sensors (PT100 platinum resistance thermometer, accuracy ±0.1°C, sampling rate 1Hz), and audio monitoring data from traction motors (MEMS microphone array, sampling frequency 44100Hz, 4 channels). The experiment lasted 60 days, accumulating 3462 valid data segments, including 186 manually labeled and confirmed fault samples (72 bogie faults, 64 wheelset wear faults, and 50 power supply abnormalities) and 3276 normal samples.

[0154] II. Experimental Conditions

[0155] The experimental platform is deployed in a two-tier architecture consisting of an onboard maintenance server and a ground data center. The onboard server is equipped with two industrial-grade embedded computing units (NVIDIA Jetson AGX Orin, 32GB memory) responsible for real-time acquisition and preprocessing of multimodal data; the ground data center is equipped with four GPU servers (NVIDIA A100 80GB×2, 512GB DDR4-3200 memory) for deep learning model training and offline analysis.

[0156] In terms of software environment, the operating system used was Ubuntu 22.04LTS, the deep learning framework was PyTorch 2.1.0, data preprocessing used Python 3.10 and NumPy 1.26, and the database used was MySQL 8.0.35 to store historical operation and maintenance records. The modality-specific feature extraction network used a ResNet-18 backbone network (vibration mode and temperature mode) and a MobileNetV3-Small backbone network (video mode and audio mode). The cross-modal feature alignment network was a custom attention mechanism network, and the fault identification and early warning classification network was a 4-layer fully connected classifier (output dimension 13 classes: 3 fault types × 4 severity + 1 normal). Regarding experimental parameters, the training iterations were 200 times, the batch size was 64, the initial learning rate was 0.001 (cosine annealing strategy), the ratio of fault samples to normal samples was balanced by oversampling at 1:4, the test set accounted for 20%, and 5-fold cross-validation was used.

[0157] III. Experimental Procedures and Results

[0158] Step 1: Synchronization, Alignment, and Dimensional Normalization of Multimodal Raw Data Timestamps

[0159] After collecting raw multimodal data (vibration, video, temperature, and audio) from the vehicle, timestamp synchronization and alignment were first performed. Due to inherent biases in the independent clocks of the four sensor types, the reference bias was 58.3 ms for the vibration mode, 47.6 ms for the video mode, 42.1 ms for the temperature mode, and 63.8 ms for the audio mode. Through a multi-source clock hard synchronization and software interpolation compensation algorithm based on GPS second pulses, the clock biases of the four modes were reduced to 2.1 ms, 3.5 ms, 1.8 ms, and 2.8 ms, respectively, improving the average alignment accuracy by 95.2%.

[0160] Dimensional normalization was performed using the Z-score standardization method. The mean and standard deviation of each modal data were calculated before standardization. Vibration acceleration data were normalized to the [-1, 1] interval, temperature data to the [0, 1] interval, video pixel values ​​to the [0, 1] interval, and audio amplitude to the [-1, 1] interval. After normalization, the standard deviation of each modal data converged to the range of 0.95–1.05, eliminating the influence of differences in physical dimensions on subsequent feature extraction. Detailed data are shown in Table 1 and [Table data would be inserted here]. Figure 3 .

[0161] Table 1. Synchronization and alignment effect of timestamps for each modality of data.

[0162] Step 3: 3D sensor anchor positioning and correlation optimization

[0163] The modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network, and spatial projection mapping to key physical structures of the train is performed through a high-dimensional manifold coordinate transformation algorithm. Based on the geometric mapping relationship of the spatial projection mapping and the spatial distribution weights of the quality assessment labels, a multi-source feature correlation optimization operation is performed. According to the response extreme value coordinates of the correlation optimization operation, three multi-dimensional sensor anchors are located: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the air gap monitoring point of the traction motor rotor.

[0164] The correlation scores of the three anchor points all showed a rapid convergence trend during the iteration process. The lateral stiffness node of the bogie side beam reached the convergence threshold of 0.92 in the 18th iteration, with a final convergence value of 0.953; the center of the wheelset tread contact patch reached the convergence threshold in the 21st iteration, with a final convergence value of 0.946; and the air gap monitoring point of the traction motor rotor reached the convergence threshold in the 16th iteration, with a final convergence value of 0.938. All three anchor points converged within 25 iterations, with an average of 18.3 iterations, indicating high efficiency in correlation optimization. See Table 2 for detailed data. Figure 4 .

[0165] Table 2 Optimization results of the correlation degree of three-dimensional sensor anchor positioning

[0166] Step 5: Cross-modal alignment and fusion and forward propagation bias correction

[0167] Cross-modal phase alignment compensation coefficients are injected into the feature interaction layer of a signal-to-noise ratio adaptive cross-modal feature alignment network. The initial modal interaction weights are dynamically calibrated through weight scaling and bias compensation operations to generate a phase-compensated calibrated interaction weight matrix. A set of modal feature vectors with quality assessment labels is input to the feature interaction layer. The calibrated interaction weight matrix is ​​used to perform spatiotemporal coordinate mapping and phase deviation correction on each modal feature, eliminating dimensional misalignment and temporal delay between heterogeneous data, thus completing the cross-modal alignment process. Cross-channel feature concatenation and fully connected dimensionality reduction aggregation operations are performed on the aligned multimodal feature subsequences to fuse complementary representation information from each modality, resulting in a multimodal fused feature representation.

[0168] Forward propagation bias correction was performed on the multimodal fusion feature representation. Intermediate state features from the outputs of each hidden layer of the deep network were extracted layer by layer and compared with the shallow baseline features through layer-by-layer differential comparison to construct a hierarchical feature residual sequence. Gradient magnitude normalization and error distribution equalization were performed on the hierarchical feature residual sequence to obtain the bias compensation mapping parameters. Experimental results show that the feature residual magnitudes of each network layer were significantly reduced after correction. The residual of the FC1 layer (fully connected first layer) decreased from 0.324 to 0.083, a reduction of 74.4%, which is the largest reduction for the FC1 layer; the residual of the Conv3 layer decreased from 0.267 to 0.064, a reduction of 76.0%. The average residual reduction across all layers was 73.8%. Detailed data are shown in Table 3 and [Table data would be inserted here]. Figure 5 .

[0169] Table 3 Comparison of characteristic residuals before and after forward propagation bias correction at each network level

[0170] Figure 5 The comparison of characteristic residual amplitudes before and after forward propagation bias correction is shown for each network layer.

[0171] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A deep learning-based intelligent analysis method for multimodal heterogeneous data, characterized in that, The method includes: The process involves acquiring multimodal raw data during the operation of urban rail transit trains, performing timestamp synchronization and alignment and dimensional normalization on the multimodal raw data to obtain standardized multimodal data sequences. These standardized multimodal data sequences are then input into corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels. The modal feature vector set with quality assessment labels is input into a cross-modal feature alignment network with adaptive signal-to-noise ratio, and the spatial projection is mapped to the key physical structure of the train to locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor. Based on the spatiotemporal response trajectories of three multidimensional sensor anchors, a non-coplanar three-dimensional distortion correction datum is constructed. Orthogonal meshing and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient. The modal interaction weights within the network are dynamically calibrated by using cross-modal phase alignment compensation coefficients, and then cross-modal alignment and feature fusion are performed on the modal feature vector set to obtain a multimodal fusion feature representation. A forward propagation bias correction process is performed on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature; The calibrated multimodal fusion features are input into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

2. The intelligent analysis method for multimodal heterogeneous data based on deep learning according to claim 1, characterized in that, The system acquires multimodal raw data during the operation of urban rail transit trains, performs timestamp synchronization and alignment and dimension normalization on the multimodal raw data, and obtains a standardized multimodal data sequence. The standardized multimodal data sequences are input into the corresponding modality-specific feature extraction networks to obtain modality feature vector sets with quality assessment labels, including: Collect time-series data from vehicle vibration sensors, video image data from on-board cameras, numerical data from bearing temperature sensors, and audio monitoring data from traction motors to form multimodal raw data. The original multimodal data is processed by timestamp synchronization alignment and dimension normalization to eliminate clock deviations and physical dimension differences from multi-source acquisition, resulting in a standardized multimodal data sequence. Standardized multimodal data sequences are input into the corresponding modality-specific feature extraction networks, and multi-scale spatiotemporal feature convolution operations are performed to extract the initial feature representations of each modality. The signal-to-noise ratio evaluation index and feature sparsity are calculated simultaneously for the initial feature representation, and quality evaluation labels are generated based on the preset reliability judgment threshold. The initial feature representations and quality assessment labels are concatenated in the feature domain to obtain the modal feature vector set.

3. The intelligent analysis method for multimodal heterogeneous data based on deep learning according to claim 2, characterized in that, The modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network, and the spatial projection is mapped onto the key physical structure of the train to locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheelset tread contact patch, and the air gap monitoring point of the traction motor rotor. The modal feature vector set with quality assessment labels is input into a signal-to-noise ratio adaptive cross-modal feature alignment network, and spatial projection mapping to the key physical structure of the train is performed through a high-dimensional manifold coordinate transformation algorithm; Based on the geometric mapping relationship of spatial projection mapping and the spatial distribution weight of quality assessment labels, multi-source feature correlation optimization calculation is performed; Based on the response extreme value coordinates obtained from the correlation optimization calculation, three multi-dimensional sensor anchors are located: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor. The multi-dimensional state response sequence of the three multi-dimensional sensor anchors in each sampling time step is extracted.

4. The intelligent analysis method for multimodal heterogeneous data based on deep learning according to claim 3, characterized in that, A non-coplanar three-dimensional distortion correction datum is constructed based on the spatiotemporal response trajectories of three multi-dimensional sensor anchors. Orthogonal mesh generation and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain cross-modal phase alignment compensation coefficients, including: An orthogonal meshing operation is performed on the non-coplanar three-dimensional distortion correction datum plane to obtain several regularly distributed mesh elements; Extract the local strain tensor features inside each subdivided mesh element, and extend them along the calculated local principal strain direction to generate a spatially oriented envelope; Perform spatial intersection determination operation between the spatial orientation envelope and the corresponding mesh element, extract the geometrically overlapping region and calculate the spatial overlap volume ratio to obtain the topological mask matrix that characterizes the effective range. The topological mask matrix is ​​used as a spatial domain constraint to map the gradient distribution field of the non-coplanar three-dimensional distortion correction datum. Gradient potential field integration is performed within the effective region defined by the mask, the phase offset gradient value is accumulated and normalized, and the cross-modal phase alignment compensation coefficient is obtained.

5. The method for intelligent analysis of multimodal heterogeneous data based on deep learning according to claim 4, characterized in that, The modal interaction weights within the network are dynamically calibrated using cross-modal phase alignment compensation coefficients. This leads to cross-modal alignment and feature fusion of the modal feature vector set, resulting in a multimodal fused feature representation, including: The cross-modal phase alignment compensation coefficients are injected into the feature interaction layer of the signal-to-noise ratio adaptive cross-modal feature alignment network. The initial modal interaction weights are dynamically calibrated through weight scaling and bias compensation operations to generate a phase-compensated calibration interaction weight matrix. The modal feature vector set with quality assessment labels is input into the feature interaction layer. Spatiotemporal coordinate mapping and phase deviation correction are performed on each modal feature by calibrating the interaction weight matrix to eliminate dimensional misalignment and temporal delay between heterogeneous data, complete cross-modal alignment processing, and obtain the aligned multimodal feature subsequence. Cross-channel feature splicing and fully connected dimensionality reduction aggregation operations are performed on the aligned multimodal feature subsequences to fuse complementary representation information of each modality. After nonlinear transformation, a multimodal fused feature representation is obtained.

6. The intelligent analysis method for multimodal heterogeneous data based on deep learning according to claim 5, characterized in that, The forward propagation bias correction process is performed on the multimodal fusion feature representation to obtain the calibrated multimodal fusion features, including: The multimodal fusion feature representation is input into the forward propagation bias correction process, the intermediate state features of each hidden layer output of the deep network are extracted layer by layer, and the layer-by-layer differential comparison operation is performed with the shallow baseline features to construct the hierarchical feature residual sequence. Gradient magnitude normalization and error distribution equalization operations are performed on the hierarchical feature residual sequence to eliminate the statistical offset caused by local gradient fluctuations and obtain the deviation compensation mapping parameters that characterize the degree of global feature distortion. Based on the bias compensation mapping parameters, reverse residual injection and adaptive smoothing filtering operations are performed on the multimodal fusion feature representation to remove the feature offset noise that accumulates continuously during the forward propagation of the deep network. The feature tensors after noise stripping are subjected to dimensional reconstruction and normalization mapping to obtain calibrated multimodal fusion features.

7. The method for intelligent analysis of multimodal heterogeneous data based on deep learning according to claim 6, characterized in that, The calibrated multimodal fusion features are input into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear, and power supply anomalies, including: The calibrated multimodal fusion features are input into the fault identification and early warning classification network. High-dimensional feature space mapping and nonlinear activation operations are performed through a fully connected classification layer to generate multi-class probability distribution sequences for bogie faults, wheelset wear and power supply anomalies. By performing confidence threshold truncation and extreme value optimization on the multi-class probability distribution sequence, the fault category label corresponding to the confidence peak is selected to obtain the preliminary fault identification result; The confidence peak is dynamically compared and matched with the preset early weak fault evolution threshold range, and the fault severity level range is divided according to the numerical deviation distance. The preliminary fault identification results are correlated and mapped with the fault severity level range to obtain the final identification results and warning levels for bogie faults, wheelset wear, and power supply anomalies.

8. A deep learning-based intelligent analysis system for multimodal heterogeneous data, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to acquire multimodal raw data during the operation of urban rail transit trains. It performs timestamp synchronization and alignment and dimension normalization on the multimodal raw data to obtain a standardized multimodal data sequence. The standardized multimodal data sequence is then input into the corresponding modality-specific feature extraction network to obtain a set of modality feature vectors with quality assessment labels. The alignment module is used to input the modal feature vector set with quality assessment labels into the signal-to-noise ratio adaptive cross-modal feature alignment network, and to map the spatial projection onto the key physical structure of the train and locate three multi-dimensional sensing anchors: the lateral stiffness node of the bogie side beam, the center of the wheel tread contact patch, and the air gap monitoring point of the traction motor rotor. The calculation module is used to construct a non-coplanar three-dimensional distortion correction datum based on the spatiotemporal response trajectories of three multi-dimensional sensor anchors. Orthogonal meshing and gradient potential field integration are performed on the non-coplanar three-dimensional distortion correction datum to obtain the cross-modal phase alignment compensation coefficient. The fusion module is used to dynamically calibrate the modal interaction weights within the network through cross-modal phase alignment compensation coefficients, and then perform cross-modal alignment and feature fusion on the modal feature vector set to obtain a multimodal fused feature representation; The calibration module is used to perform forward propagation bias correction on the multimodal fusion feature representation to obtain the calibrated multimodal fusion feature. The processing module is used to input the calibrated multimodal fusion features into the fault identification and early warning classification network to obtain the identification results and early warning levels of bogie faults, wheelset wear and power supply anomalies.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal positioning method and system based on intelligent wearable device

    CN121028142A

  • Intelligent elevator operation and maintenance management method and system based on multi-mode neural network

    CN121493741A

  • Railway anomaly detection method and system based on multi-modal data fusion

    WO2025092018A1