Intelligent information processing system based on multi-modal data fusion

By combining multi-camera and lidar with edge computing data processing technology, real-time and accurate monitoring of track status is achieved, solving the detection difficulties caused by the heterogeneity of multimodal data and network instability in high-speed train environments, and improving the real-time and accuracy of track deformation detection.

CN120705822AInactive Publication Date: 2025-09-26孙达
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511040213.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a high-speed train environment, the heterogeneity of multimodal data and unstable network conditions lead to a decrease in the real-time and accuracy of track deformation detection, which cannot meet the high-precision and high-real-time requirements of modern railway safety monitoring.

Method used

Multi-camera and lidar are used for track data collection, combined with edge computing nodes for data preprocessing and feature extraction. Through dynamic trust configuration and 3D point cloud fusion technology, fast, real-time data processing and unified standards are achieved, reducing the impact of network delays and packet loss.

Benefits of technology

It improves the real-time and accuracy of track deformation detection, solves the technical problems of multimodal data heterogeneity and unstable transmission, and realizes comprehensive monitoring of track status and efficient fault warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705822A_ABST
    Figure CN120705822A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent information processing system based on multi-modal data fusion, belongs to the technical field of data processing, and aims to solve the problems of multi-modal data isomerism, time consumption in preprocessing and unstable transmission in high-speed operation. The system comprises a data acquisition module which is used for acquiring track data by using a multi-view camera and a laser radar of a track detection vehicle and establishing a time sequence multi-mode data set; the data preprocessing and feature extraction module carries out real-time preprocessing and feature extraction through edge computing nodes; the dynamic trust configuration module configures the dynamic trust degree according to the preprocessed data; the three-dimensional point cloud fusion module executes three-dimensional point cloud fusion of the time series data according to the credibility to generate a fusion result; the driving data acquisition module acquires a bullet train driving data set in real time; the multi-modal three-dimensional target detection module is used for performing target attention detection on the fusion result and the driving data set; and the result output module generates a three-dimensional bounding box and classification confidence of the track deformation area according to the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an intelligent information processing system based on multimodal data fusion. Background Art

[0002] Track deformation detection, a crucial component of railway transportation safety monitoring, has garnered widespread attention in recent years. With the development of railway transportation toward higher speeds, safety, and intelligent capabilities, the demands placed on real-time monitoring of track conditions and fault warnings during high-speed train operation have increased. To this end, various sensors (such as laser ranging, image acquisition, and inertial measurement) are being applied to track detection, enabling comprehensive monitoring of track conditions through multimodal data acquisition. However, due to the varying sensor response times and diverse data formats of the collected data, the effective fusion and rapid processing of this heterogeneous data has become a key technical challenge in achieving real-time detection.

[0003] In high-speed train environments, the network conditions surrounding the train are often unstable, leading to significant communication delays and packet loss. This unstable network severely impacts real-time data transmission and processing. This is especially true when fusing multimodal data. Data inconsistencies and latency variations further reduce the overall system's response speed and detection accuracy, leading to a decrease in the real-time performance of track deformation detection, making it impossible to meet the high-precision and real-time detection requirements of modern railway safety monitoring.

[0004] The disclosure of the above background technology content is only used to assist in understanding the concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the Invention

[0005] The present application provides an intelligent information processing system based on multimodal data fusion, which is used to solve the technical problems caused by multimodal data heterogeneity, time-consuming preprocessing and unstable transmission in a high-speed operation environment.

[0006] To achieve the above objectives, the present application discloses the following technical solutions:

[0007] An intelligent information processing system based on multimodal data fusion, comprising:

[0008] The data acquisition module is used to collect track data using the multi-camera and lidar integrated on the track inspection vehicle and to establish a time-series multimodal dataset;

[0009] A data preprocessing and feature extraction module, configured to perform real-time data preprocessing and feature extraction on the time series multimodal dataset using an edge computing node to obtain a preprocessed time series multimodal dataset;

[0010] A dynamic trust configuration module, configured to configure dynamic trust according to the preprocessed time series multimodal data set;

[0011] A three-dimensional point cloud fusion module is used to perform time series data three-dimensional point cloud fusion on the pre-processed time series multimodal data set according to the dynamic trust degree to generate a three-dimensional point cloud fusion result;

[0012] Driving data acquisition module, used to obtain the driving data set of the motor vehicle in real time;

[0013] a multimodal 3D target detection module, configured to perform target attention detection on the 3D point cloud fusion result and the driving data set using a multimodal 3D target detection network, and output a target attention detection result;

[0014] The result output module is used to generate a three-dimensional bounding box and classification confidence of the track deformation area based on the detection results.

[0015] In the embodiment of the present application, by integrating multi-cameras and lidar to obtain track surface details and three-dimensional spatial structure information, a time-series multimodal data set is constructed, which fundamentally realizes comprehensive monitoring of the track status. Edge computing nodes are used to perform data preprocessing and feature extraction at the data collection site to eliminate the interference caused by sensor response time difference and data format diversity, so that the information has a unified standard and high processing efficiency. In addition, the introduction of edge computing also reduces the uncertainty caused by network delay or packet loss during data transmission, and can complete most computing tasks on site, realizing fast and real-time data preprocessing. Then, by real-time evaluation of preprocessing results and configuring dynamic trust, the sensor collected data is given different weights according to the transmission environment and data quality, so that the information is automatically corrected for inconsistencies during fusion, reducing the impact of network delay and data loss on subsequent processing. The time-series three-dimensional point cloud fusion technology is used to accurately match data information from different collection devices in time and space to form a data model that truly reflects the three-dimensional structure of the track, fundamentally improving the real-time and accuracy of track deformation detection, and effectively solving the technical problems caused by multi-modal data heterogeneity, time-consuming preprocessing and unstable transmission in high-speed operation environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of data interaction between modules of an intelligent information processing system based on multimodal data fusion provided in some embodiments of the present application;

[0017] Figure 2 for Figure 1Schematic diagram of data interaction between the modules of the data preprocessing and feature extraction module;

[0018] Figure 3 for Figure 1 Schematic diagram of data interaction between the units of the image feature extraction module;

[0019] Figure 4 for Figure 1 Schematic diagram of data interaction between the modules of the dynamic trust configuration module;

[0020] Figure 5 for Figure 1 Schematic diagram of data interaction between the modules of the 3D point cloud fusion module;

[0021] Figure 6 for Figure 1 Schematic diagram of data interaction between the various units of the driving data acquisition module;

[0022] Figure 7 for Figure 1 Schematic diagram of data interaction between the various units of the multimodal 3D object detection module;

[0023] Figure 8 for Figure 4 Schematic diagram of data interaction between the units of the weight dynamic adjustment module;

[0024] Figure 9 for Figure 7 Schematic diagram of data interaction between the subunits of the multimodal feature fusion unit;

[0025] Figure 10 for Figure 5 Schematic diagram of data interaction between the various units of the alignment module described in . DETAILED DESCRIPTION

[0026] Specific embodiments of the present invention will now be mentioned in detail. Although the present invention is described in conjunction with these specific embodiments, it should be appreciated that the present invention is not intended to be limited to these specific embodiments. On the contrary, these embodiments are intended to cover substitutions, changes or equivalent embodiments that may be included in the spirit and scope of the invention defined by the claims. In the following description, a large amount of specific details are set forth to provide a comprehensive understanding of the present invention. The present invention can be implemented without some or all of these specific details. In other cases, in order not to make the present invention unnecessarily obscure, well-known process operations are not described in detail.

[0027] When used in conjunction with "including," "methods comprising," or similar language in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0028] Application Overview: Track deformation detection, as a key component of railway transportation safety monitoring, has received widespread attention in recent years. With the development of railway transportation toward high-speed, safe, and intelligent transportation, higher requirements are being placed on real-time monitoring of track conditions and fault warnings during high-speed train operation. To this end, various sensors (such as laser ranging, image acquisition, inertial measurement, etc.) are being used in track detection, enabling comprehensive monitoring of track conditions through multimodal data acquisition. However, due to the varying sensor response times and diverse data formats of the collected data, the effective fusion and rapid processing of these heterogeneous data has become a key technical challenge in achieving real-time detection.

[0029] In high-speed train environments, the network conditions surrounding the train are often unstable, leading to significant communication delays and packet loss. This unstable network severely impacts real-time data transmission and processing. This is especially true when fusing multimodal data. Data inconsistencies and latency variations further reduce the overall system's response speed and detection accuracy, leading to a decrease in the real-time performance of track deformation detection, making it impossible to meet the high-precision and real-time detection requirements of modern railway safety monitoring.

[0030] For the above technical issues, please refer to Figure 1 , an embodiment of the present application provides an intelligent information processing system based on multimodal data fusion, comprising:

[0031] Data acquisition module 1, used to collect track data using a multi-camera and lidar integrated on the track inspection vehicle and to establish a time-series multimodal dataset;

[0032] Data preprocessing and feature extraction module 2, used to use edge computing nodes to perform real-time data preprocessing and feature extraction on the time series multimodal dataset to obtain a preprocessed time series multimodal dataset;

[0033] A dynamic trust configuration module 3, configured to configure dynamic trust according to the preprocessed time series multimodal data set;

[0034] A three-dimensional point cloud fusion module 4 is configured to perform time series data three-dimensional point cloud fusion on the pre-processed time series multimodal data set according to the dynamic trust degree to generate a three-dimensional point cloud fusion result;

[0035] Driving data acquisition module 5, used to obtain the driving data set of the motor vehicle in real time;

[0036] a multimodal 3D target detection module 6, configured to perform target attention detection on the 3D point cloud fusion result and the driving data set using a multimodal 3D target detection network, and output a target attention detection result;

[0037] The result output module 7 is used to generate a three-dimensional bounding box and classification confidence of the track deformation area according to the detection results.

[0038] This system integrates multiple cameras and lidar to capture track surface details and three-dimensional spatial structure, constructing a time-series multimodal dataset, fundamentally enabling comprehensive monitoring of track status. Edge computing nodes are used to perform data preprocessing and feature extraction at the data collection site, eliminating interference caused by sensor response time differences and data format diversity, thereby ensuring uniform information standards and high processing efficiency. Furthermore, the introduction of edge computing reduces uncertainty caused by network delays and packet loss during data transmission, enabling the majority of computational tasks to be completed on-site, enabling rapid, real-time data preprocessing. Subsequently, by real-time evaluation of preprocessing results and configuring dynamic trustworthiness, sensor data is weighted based on the transmission environment and data quality. This allows for automatic correction of inconsistencies during information fusion, minimizing the impact of network delays and data loss on subsequent processing. Time-series 3D point cloud fusion technology precisely matches data from different acquisition devices in time and space, forming a data model that truly reflects the track's three-dimensional structure. This fundamentally improves the real-time and accuracy of track deformation detection, effectively addressing the technical challenges of heterogeneous multimodal data, time-consuming preprocessing, and unstable transmission in high-speed operating environments.

[0039] See also Figure 2 In some embodiments, the data preprocessing and feature extraction module 2 includes:

[0040] An image and point cloud segmentation module 21 is used to separate image data from laser point cloud data in a temporal multimodal dataset in real time through an edge computing node;

[0041] An image preprocessing module 22 is used to perform illumination normalization processing on the segmented image data to generate normalized image data;

[0042] The point cloud preprocessing module 23 is used to perform outlier filtering on the segmented laser point cloud data to generate filtered point cloud data. Specifically, the following outlier judgment criteria are used:

[0043]

[0044] Where, Represents any data point in the point cloud; Representing point clouds The mean of all data points in ; Representing point clouds The standard deviation of all data points in ; It is a preset threshold coefficient used to control the sensitivity of outlier judgment; Represents the output of the decision function after outlier filtering, 1 means keep, 0 means discard;

[0045] An image feature extraction module 24 is configured to extract texture features from the normalized image data based on a convolutional neural network and generate an image feature vector;

[0046] The point cloud feature extraction module 25 is used to extract curvature features from the filtered point cloud data based on a geometric feature extraction algorithm to generate a point cloud feature vector;

[0047] The feature fusion module 26 is used to fuse the image feature vector and the point cloud feature vector to generate a pre-processed temporal multimodal data set. Specifically, the following linear weighted fusion formula is used:

[0048] ;

[0049] Where, Represents the feature vector extracted from the image data; Represents the feature vector extracted from the point cloud data; is the fusion weight of image features, and its value range is ; The feature vectors generated after fusion constitute the preprocessed time series multimodal dataset.

[0050] This solution performs preprocessing by separating image data from point cloud data, reduces the impact of lighting changes in unstable environments through lighting normalization, reduces background noise interference through outlier filtering, uses convolutional neural networks to obtain vectors describing texture features from normalized images, and uses geometric feature extraction algorithms to obtain vectors describing spatial structure information from filtered point clouds. The aforementioned image texture information and point cloud geometric information are then integrated to generate a unified data set describing temporal information, thereby improving the accuracy and comprehensiveness of information expression, thereby enhancing the detection accuracy of the subsequent multimodal three-dimensional target detection module 6 and the real-time response capability of the overall system.

[0051] See also Figure 3 In some embodiments, the image feature extraction module 24 specifically includes:

[0052] A data input unit 241 is used to input normalized image data into a pre-trained ResNet-18 network;

[0053] A feature acquisition unit 242 is used to extract the output feature map of the third convolutional layer of ResNet-18;

[0054] The pooling unit 243 is used to perform a global average pooling operation on the output feature map to generate an image feature vector; the global average pooling operation is calculated using the following formula:

[0055] ;

[0056] Where, Represents the output feature map of the third convolutional layer of ResNet-18 spatial position, The activation value of each channel; Represents the total number of spatial positions in the output feature map; Represents the image feature vector in The values ​​in the dimensions.

[0057] This solution uses a pre-trained ResNet-18 network to perform deep feature extraction on normalized image data. The pre-trained deep convolutional network captures the local details and global semantic information contained in the image through the output of the intermediate convolutional layer, and converts the high-dimensional feature map into a compact feature vector through the global average pooling operation, thereby achieving efficient expression and dimensionality reduction of image features. The visual knowledge formed after the pre-trained network has been trained on massive visual data empowers the deep feature extraction process, ensuring accurate description and robust expression of image content. The global average pooling operation reduces computational complexity while maintaining the integrity of image information, providing high-quality and efficient image feature support for subsequent multimodal data fusion, thereby improving the accuracy and response speed of the overall detection system.

[0058] See also Figure 4 In some embodiments, the dynamic trust configuration module 3 includes:

[0059] An image clarity evaluation module 31 is used to calculate a clarity evaluation value of image data in the preprocessed time series multimodal dataset;

[0060] The point cloud density evaluation module 32 is used to calculate the density evaluation value of the laser point cloud data in the pre-processed time series multimodal data set;

[0061] A weight dynamic adjustment module 33 is used to dynamically adjust the weight of each evaluation value based on the real-time network signal strength using the edge computing node;

[0062] The weighted calculation module 34 is used to calculate the weighted sum of the clarity evaluation value and the density evaluation value according to the adjusted weights to generate a dynamic trust level. The specific formula is as follows:

[0063] ;

[0064] Where, Indicates the clarity evaluation value calculated based on image data; Indicates the density evaluation value calculated based on point cloud data; represents the weight assigned to the image data; Represents the weight assigned to point cloud data, which is satisfy ; It represents the final calculated dynamic trust, which is used for weight configuration of subsequent data fusion.

[0065] The dynamic trust configuration module 3 obtains the quality of pre-processed data through the image clarity evaluation module 31 and the point cloud density evaluation module 32. It also uses the results of real-time monitoring of network signal strength on the edge computing node to dynamically adjust the evaluation weights and generate corresponding weighted results, which directly influence the quality of subsequent data fusion and target detection. The dynamic trust configuration module 3 quantifies the information clarity reflected in the image data extraction process and the sampling density reflected in the laser point cloud data, and selectively weights the network signal strength. The resulting virtuous cycle ensures that the highest quality data is always used for processing in the data fusion process, ensuring that data collected by different sensors can participate in the calculation in the multimodal fusion stage in an optimal manner, thereby making the detection results more stable and reliable, effectively alleviating the uncertainty caused by network environment fluctuations in data transmission and processing, and improving the entire system's ability to achieve real-time analysis and early warning under high-speed operation.

[0066] See also Figure 5 In some embodiments, the three-dimensional point cloud fusion module 4 includes:

[0067] A weight configuration module 41 is configured to configure image texture weights and point cloud geometry weights using the dynamic trust level;

[0068] The weighted fusion module 42 is used to fuse the image texture details in the preprocessed temporal multimodal dataset with the point cloud spatial geometric information using a weighted fusion algorithm. Specifically, the following weighted fusion formula is used:

[0069] ;

[0070] Where, Represents texture information extracted from image data; Represents the geometric information extracted from point cloud data; is the dynamic trust degree calculated by the dynamic trust configuration module 3; Represents the 3D point cloud data set generated after fusion, which is used for subsequent alignment processing;

[0071] The alignment module 43 is used to align the fused multimodal data using an iterative closest point algorithm to generate a final 3D point cloud fusion result.

[0072] 3D point cloud fusion module 4 uses dynamically configured weights to assign different weights to the image texture information and point cloud geometry information in the preprocessed time-series multimodal data. This weighted fusion achieves information complementarity, thereby ensuring that image details and spatial structure are fully preserved. An iterative closest point algorithm is used to finely align the fused multimodal data, effectively reducing errors caused by sensor response time differences and data format differences, ensuring accurate 3D spatial reconstruction. 3D point cloud fusion module 4 outputs high-quality 3D point cloud results, providing accurate spatial data for track status detection, improving the system's overall detection accuracy and response speed, and promoting the stable and efficient monitoring of track safety by intelligent information processing systems in high-speed operating environments.

[0073] See also Figure 6 In some embodiments, the driving data acquisition module 5 includes:

[0074] The acceleration data acquisition unit 51 is used to collect the acceleration sensor data of the motor vehicle in real time through the edge computing node;

[0075] A vibration data acquisition unit 52 is used to collect vibration sensor data of the motor vehicle in real time through the edge computing node;

[0076] GPS data acquisition unit 53, used to analyze and collect GPS trajectory data of the train in real time through edge computing nodes;

[0077] The data integration unit 54 is used to integrate the acceleration, vibration and GPS data to generate a driving data set. The generation of is based on the following integration formula:

[0078] ;

[0079] Where, Indicates the collected acceleration data; Indicates the collected vibration data; Indicates the collected GPS data; Indicates the weight coefficient corresponding to the acceleration data; Indicates the weight coefficient corresponding to the vibration data; Indicates the weight coefficient corresponding to GPS data, and satisfies ; is the integrated driving dataset.

[0080] The driving data acquisition module 5 receives acceleration sensor signals, vibration sensor information, and GPS trajectory data, and generates a comprehensive driving data set through the data integration unit 54. This comprehensive driving data set provides information on the train's operating status, ensuring a spatiotemporal correspondence between vehicle dynamic behavior and track status detection, and enabling comprehensive reflection of the train's operating status under real-time monitoring conditions. The driving data acquisition module 5 maintains data continuity and stability even when network transmission conditions are unstable, ensuring the integrity and accuracy of dynamic information collection. This provides reliable input for subsequent anomaly detection, improving the overall safety monitoring system's responsiveness and early warning effectiveness.

[0081] See also Figure 7 In some embodiments, the multimodal three-dimensional object detection module 6 includes:

[0082] A point cloud feature extraction unit 61 is used to input the 3D point cloud fusion result into the PointNet++ network to extract deformation features;

[0083] A time series feature extraction unit 62 is used to input the driving data set into the LSTM network to extract time series abnormal features;

[0084] The multimodal feature fusion unit 63 is used to fuse the deformation feature and the time series anomaly feature to generate a multimodal feature vector, using the following weighted fusion formula:

[0085] ;

[0086] Where, Represents the deformation features extracted by the PointNet++ network; Represents the time series anomaly features extracted by the LSTM network; is the weight coefficient of deformation feature fusion, and its value range should be Inside; Represents the multimodal feature vector formed after fusion, which is used for subsequent target detection operations;

[0087] Classification unit 64 is used to perform classification operations on the multimodal feature vectors through fully connected layers and generate target attention detection results. The multimodal 3D target detection module 6 uses a PointNet++ network to extract deformation features and an LSTM network to extract temporal anomaly features, ensuring full representation of both spatial geometry and motion state information. This intelligent integration of spatial deformation information and temporal anomaly information improves the accuracy of track deformation state determination during target attention detection. The fully connected layers perform classification operations to effectively distinguish between normal and abnormal states. By combining deep features, the sensitivity of target detection responses is improved, ensuring the stability and robustness of detection results in complex environments.

[0088] See also Figure 8 In some embodiments, the weight dynamic adjustment module 33 specifically includes:

[0089] The signal monitoring unit 331 is used to monitor the network signal strength in real time through the edge computing node and generate a signal strength index;

[0090] The weight configuration unit 332 is configured to: when the signal strength index is lower than a preset threshold, configure the weight of the image clarity to a first weight value and the weight of the point cloud density to a second weight value; when the signal strength index is higher than or equal to the preset threshold, configure the weight of the image clarity to a third weight value and the weight of the point cloud density to a fourth weight value;

[0091] The weight calculation unit 333 is used to calculate the weighted sum of the clarity evaluation value and the density evaluation value according to the configured weights to generate dynamic trust. This solution assigns dynamic weights to the image clarity evaluation value and the point cloud density evaluation value based on real-time monitoring of network signal strength and preset threshold judgment, and generates a credibility index to improve environmental adaptability during the data fusion process. The signal monitoring unit 331 continuously obtains network status information and adjusts the weight configuration according to the real-time network status to ensure that the evaluation index can still truly reflect the data information quality under network fluctuations and eliminate data errors caused by network communication interruptions and delays. The dynamic weight adjustment mechanism enables the reasonable distribution of the contribution of evaluation indicators during the data fusion process, thereby effectively improving the overall data processing robustness and detection accuracy, and ensuring that the track deformation detection system can still achieve high precision and high real-time requirements in an unstable network environment.

[0092] See also Figure 9 In some embodiments, the multimodal feature fusion unit 63 includes:

[0093] The deformation feature normalization subunit 631 is used to perform a first normalization process on the extracted deformation features to generate standardized deformation features;

[0094] The abnormal feature normalization subunit 632 is used to perform a second normalization process on the extracted time series abnormal features to generate standardized abnormal features;

[0095] The attention calculation subunit 633 is used to calculate the fusion weight between the normalized deformation feature and the normalized abnormal feature using the attention mechanism; the attention weight vector is calculated using the following formula;

[0096] The weighted splicing subunit 634 is used to perform weighted splicing on the normalized deformation features and the normalized abnormal features based on the fusion weights to generate a final multimodal feature vector.

[0097] The multimodal feature fusion unit 63 uses normalization to unify the numerical distribution of the extracted track deformation features and temporal anomaly features, ensuring consistent scaling of the different modal data during the fusion process. The unified features are objectively weighted in the attention mechanism, assigning a reasonable proportion and achieving weighted concatenation to generate a feature vector reflecting comprehensive information. The combined effects of normalization and weight calculation result in more comprehensive feature expression and more concentrated information, thereby improving the accuracy of abnormal state identification during detection and recognition, reducing the risk of false positives and missed positives, and enhancing the stability and reliability of the overall monitoring system.

[0098] See also Figure 10 In some embodiments, the alignment module 43 includes:

[0099] A key point extraction unit 431 is used to extract a key point set from the weighted fused multimodal data;

[0100] An initial registration unit 432 is configured to perform initial registration using an iterative closest point algorithm based on the extracted key point set to generate an initial transformation matrix;

[0101] The transformation optimization unit 433 is used to optimize the initial transformation matrix by using the distance error constraint to generate an optimized transformation matrix;

[0102] The data alignment unit 434 is used to align the fused multimodal data using the optimized transformation matrix, thereby generating the final 3D point cloud fusion result. Specifically, the alignment module 43 uses the key point extraction unit 431 to obtain a set of key points rich in information from the weighted fused multimodal data. The initial registration unit 432 uses the iterative closest point algorithm to generate an initial transformation matrix using the key point set. Position information between the data is obtained from the preliminary matching, providing a basis for optimization. The transformation optimization unit 433 adjusts the initial transformation matrix based on the distance error constraint to eliminate accumulated deviations, thereby achieving fine correction of the spatial position between the data. The data alignment unit 434 uniformly aligns the multimodal data based on the optimized transformation matrix, so that the 3D point cloud achieves a high degree of consistency in spatial expression, ensuring high-precision data output and providing accurate and reliable 3D reconstruction results for subsequent detection, fundamentally improving the accuracy and safety of the overall system detection.

[0103] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, persons skilled in the art will appreciate that modifications to the specific embodiments of the present invention or substitutions of some of the technical features may be made without departing from the spirit of the present invention and are intended to be encompassed within the scope of the technical solutions claimed herein.

Claims

1. An intelligent information processing system based on multimodal data fusion, characterized in that: include: The data acquisition module is used to collect track data using the multi-camera and lidar integrated on the track inspection vehicle and to establish a time-series multimodal dataset; A data preprocessing and feature extraction module, configured to perform real-time data preprocessing and feature extraction on the time series multimodal dataset using an edge computing node to obtain a preprocessed time series multimodal dataset; A dynamic trust configuration module, configured to configure dynamic trust according to the preprocessed time series multimodal data set; A three-dimensional point cloud fusion module is used to perform time series data three-dimensional point cloud fusion on the pre-processed time series multimodal data set according to the dynamic trust degree to generate a three-dimensional point cloud fusion result; Driving data acquisition module, used to obtain the driving data set of the motor vehicle in real time; a multimodal 3D target detection module, configured to perform target attention detection on the 3D point cloud fusion result and the driving data set using a multimodal 3D target detection network, and output a target attention detection result; The result output module is used to generate a three-dimensional bounding box and classification confidence of the track deformation area based on the detection results.

2. The intelligent information processing system based on multimodal data fusion according to claim 1, characterized in that: The data preprocessing and feature extraction module includes: Image and point cloud segmentation module, used to separate image data and laser point cloud data in time-series multimodal datasets in real time through edge computing nodes; An image preprocessing module, configured to perform illumination normalization processing on the segmented image data to generate normalized image data; The point cloud preprocessing module is used to perform outlier filtering on the segmented laser point cloud data to generate filtered point cloud data. The specific outlier judgment criteria are as follows: Where, Represents any data point in the point cloud; Representing point clouds The mean of all data points in ; Representing point clouds The standard deviation of all data points in ; It is a preset threshold coefficient used to control the sensitivity of outlier judgment; Represents the output of the decision function after outlier filtering, 1 means keep, 0 means discard; An image feature extraction module is used to extract texture features from normalized image data based on a convolutional neural network and generate an image feature vector; Point cloud feature extraction module, used to extract curvature features from filtered point cloud data based on geometric feature extraction algorithm and generate point cloud feature vectors; The feature fusion module is used to fuse the image feature vector and the point cloud feature vector to generate a preprocessed temporal multimodal dataset. The following linear weighted fusion formula is used: ; Where, Represents the feature vector extracted from the image data; Represents the feature vector extracted from the point cloud data; is the fusion weight of image features, and its value range is ; The feature vectors generated after fusion constitute the preprocessed time series multimodal dataset.

3. The intelligent information processing system based on multimodal data fusion according to claim 2, characterized in that: The image feature extraction module specifically includes: A data input unit, used to input normalized image data into the pre-trained ResNet-18 network; Feature acquisition unit, used to extract the output feature map of the third convolutional layer of ResNet-18; The pooling unit is used to perform a global average pooling operation on the output feature map to generate an image feature vector; the global average pooling operation is calculated using the following formula: ; Where, Represents the output feature map of the third convolutional layer of ResNet-18 spatial position, The activation value of each channel; Represents the total number of spatial positions in the output feature map; Represents the image feature vector in The values ​​in the dimensions.

4. The intelligent information processing system based on multimodal data fusion according to claim 1, characterized in that: The dynamic trust configuration module includes: An image clarity evaluation module is used to calculate the clarity evaluation value of image data in the preprocessed time series multimodal dataset; Point cloud density evaluation module, used to calculate the density evaluation value of laser point cloud data in the preprocessed time series multimodal dataset; The weight dynamic adjustment module is used to dynamically adjust the weight of each evaluation value based on the real-time network signal strength using the edge computing node; The weighted calculation module is used to calculate the weighted sum of the clarity evaluation value and the density evaluation value according to the adjusted weights to generate a dynamic trust level. The specific formula is as follows: ; Where, Indicates the clarity evaluation value calculated based on image data; Indicates the density evaluation value calculated based on point cloud data; represents the weight assigned to the image data; Represents the weight assigned to point cloud data, which is satisfy ; It represents the final calculated dynamic trust, which is used for weight configuration of subsequent data fusion.

5. The intelligent information processing system based on multimodal data fusion according to claim 1, characterized in that: The three-dimensional point cloud fusion module includes: A weight configuration module, configured to configure image texture weights and point cloud geometry weights using the dynamic trust level; The weighted fusion module is used to fuse the image texture details in the preprocessed temporal multimodal dataset with the point cloud spatial geometric information using a weighted fusion algorithm. Specifically, the following weighted fusion formula is used: ; Where, Represents texture information extracted from image data; Represents the geometric information extracted from point cloud data; is the dynamic trust calculated by the dynamic trust configuration module; Represents the 3D point cloud data set generated after fusion, which is used for subsequent alignment processing; The alignment module is used to align the fused multimodal data using an iterative closest point algorithm to generate the final 3D point cloud fusion result.

6. The intelligent information processing system based on multimodal data fusion according to claim 1, characterized in that: The driving data acquisition module includes: Acceleration data acquisition unit, used to collect acceleration sensor data of the train in real time through edge computing nodes; Vibration data acquisition unit, used to collect vibration sensor data of the train in real time through edge computing nodes; GPS data acquisition unit, used to analyze and collect GPS trajectory data of the train in real time through edge computing nodes; The data integration unit is used to integrate the above acceleration, vibration and GPS data to generate a driving data set. The generation of is based on the following integration formula: ; Where, Indicates the collected acceleration data; Indicates the collected vibration data; Indicates the collected GPS data; Indicates the weight coefficient corresponding to the acceleration data; Indicates the weight coefficient corresponding to the vibration data; Indicates the weight coefficient corresponding to GPS data, and satisfies ; is the integrated driving dataset.

7. The intelligent information processing system based on multimodal data fusion according to claim 1, characterized in that: The multimodal three-dimensional target detection module includes: Point cloud feature extraction unit, used to input the 3D point cloud fusion results into the PointNet++ network to extract deformation features; The time series feature extraction unit is used to input the driving data set into the LSTM network to extract time series anomaly features; The multimodal feature fusion unit is used to fuse deformation features and temporal anomaly features to generate a multimodal feature vector. The following weighted fusion formula is used: ; Where, Represents the deformation features extracted by the PointNet++ network; Represents the time series anomaly features extracted by the LSTM network; is the weight coefficient of deformation feature fusion, and its value range should be Inside; Represents the multimodal feature vector formed after fusion, which is used for subsequent target detection operations; The classification unit is used to perform classification operations on the multimodal feature vector through the fully connected layer and generate target attention detection results.

8. The intelligent information processing system based on multimodal data fusion according to claim 4, characterized in that: The weight dynamic adjustment module specifically includes: A signal monitoring unit is used to monitor the network signal strength in real time through edge computing nodes and generate signal strength indicators; a weight configuration unit, configured to: when the signal strength index is lower than a preset threshold, configure the weight of the image clarity to a first weight value and the weight of the point cloud density to a second weight value; when the signal strength index is higher than or equal to the preset threshold, configure the weight of the image clarity to a third weight value and the weight of the point cloud density to a fourth weight value; The weight calculation unit is used to calculate the weighted sum of the clarity evaluation value and the density evaluation value according to the configured weights to generate a dynamic trust level.

9. The intelligent information processing system based on multimodal data fusion according to claim 7, characterized in that: The multimodal feature fusion unit includes: a deformation feature normalization subunit, configured to perform a first normalization process on the extracted deformation features to generate standardized deformation features; The abnormal feature normalization subunit is used to perform a second normalization process on the extracted time series abnormal features to generate standardized abnormal features; The attention calculation subunit is used to calculate the fusion weight between the normalized deformation features and the normalized abnormal features using the attention mechanism; the attention weight vector is calculated using the following formula; The weighted splicing subunit is used to perform weighted splicing on the standardized deformation features and the standardized abnormal features based on the fusion weight to generate the final multimodal feature vector.

10. The intelligent information processing system based on multimodal data fusion according to claim 5, characterized in that: The alignment module includes: A key point extraction unit, used to extract a key point set from the weighted fused multimodal data; An initial registration unit, configured to perform an initial registration using an iterative closest point algorithm based on the extracted key point set to generate an initial transformation matrix; A transformation optimization unit is used to optimize the initial transformation matrix through the distance error constraint to generate an optimized transformation matrix; The data alignment unit is used to align the fused multimodal data by applying the optimized transformation matrix to generate the final 3D point cloud fusion result.