Traffic Rail Deformation Detection Method and System Based on Multimodal 3D Point Cloud Fusion
By integrating multi-eye cameras and lidar on the track detection vehicle for data acquisition and fusion, and combining with a multi-modal three-dimensional target detection network, the accuracy of track deformation detection is solved, and accurate identification and real-time monitoring of track deformation is achieved.
Patent Information
- Application Number
- CN202510535923.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the prior art, due to the failure of different data to be effectively integrated, the dynamic deformation detection capability of the track under different operating states is weak, affecting the accuracy of the track deformation detection.
By using multi-eye cameras and lidar integrated on the track detection vehicle for track data acquisition, a timing multi-modal data set is established, dynamic trust is configured for timing data three-dimensional point cloud fusion, and target attention detection is combined with a multi-modal three-dimensional target detection network to output the three-dimensional bounding box and classification confidence of the track deformation area.
Accurate monitoring of orbits under different dynamic states is achieved, the accuracy and recognition ability of orbit deformation detection is improved, and the track deformation and abnormality can be detected in a timely manner.
Smart Images

Figure CN120068001B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a traffic track deformation detection method and system based on multi-modal three-dimensional point cloud fusion. Background Art
[0002] Traffic track deformation detection is an important means to ensure the safety and stability of the track. Existing traffic track deformation detection methods combine various technical means, and the main goals are to improve detection accuracy and efficiency, reduce manual intervention, and timely detect track deformation, damage or abnormalities. Some existing track deformation detection methods will fuse data from multiple sensors, usually relying on technologies such as data synchronization, data quality assessment and denoising processing, and strive to achieve coordination among data from different sensors, so as to improve the accuracy of track deformation detection. The data of different sensors have different accuracies, resolutions and response characteristics, and there are also differences in the acquisition accuracy in time and space, resulting in the inability to achieve truly efficient fusion during the multi-source data fusion process, leading to increased data redundancy, information loss or inconsistent processing, thus affecting the overall accuracy of track deformation detection.
[0003] In summary, there is a technical problem in the prior art that due to the ineffective fusion of different data, the dynamic deformation detection ability of the track in different operating states is weak, further affecting the accuracy of track deformation detection. Summary of the Invention
[0004] The purpose of this application is to provide a traffic track deformation detection method and system based on multi-modal three-dimensional point cloud fusion, so as to solve the technical problem in the prior art that due to the ineffective fusion of different data, the dynamic deformation detection ability of the track in different operating states is weak, further affecting the accuracy of track deformation detection.
[0005] In view of the above problems, this application provides a traffic track deformation detection method and system based on multi-modal three-dimensional point cloud fusion.
[0006] In a first aspect, the present application provides a traffic track deformation detection method based on multimodal three-dimensional point cloud fusion. The traffic track deformation detection method based on multimodal three-dimensional point cloud fusion is implemented through a traffic track deformation detection system based on multimodal three-dimensional point cloud fusion. Among them, the traffic track deformation detection method based on multimodal three-dimensional point cloud fusion includes: using a multi-camera and a lidar integrated on a track detection vehicle to collect track data, and establishing a time-series multimodal data set, where the time-series multimodal data set includes an image data set and lidar point cloud data; configuring the dynamic trust degree of the time-series multimodal data set, and performing three-dimensional point cloud fusion on the time-series multimodal data set according to the dynamic trust degree to establish a three-dimensional point cloud fusion result; establishing a driving data set, where the driving data set is a data set of a bullet train in different states, and the driving data set includes acceleration data, vibration data, and trajectory data; using a multimodal three-dimensional object detection network to perform object attention detection on the three-dimensional point cloud fusion result and the driving data set, and outputting a three-dimensional bounding box and a classification confidence degree of the track deformation area according to the object attention detection result.
[0007] In a second aspect, the present application also provides a traffic track deformation detection system based on multimodal three-dimensional point cloud fusion for executing the traffic track deformation detection method based on multimodal three-dimensional point cloud fusion as described in the first aspect. Among them, the traffic track deformation detection system based on multimodal three-dimensional point cloud fusion includes: a data acquisition module for using a multi-camera and a lidar integrated on a track detection vehicle to collect track data and establish a time-series multimodal data set, where the time-series multimodal data set includes an image data set and lidar point cloud data; a data fusion module for configuring the dynamic trust degree of the time-series multimodal data set and performing three-dimensional point cloud fusion on the time-series multimodal data set according to the dynamic trust degree to establish a three-dimensional point cloud fusion result; a driving data set establishment module for establishing a driving data set, where the driving data set is a data set of a bullet train in different states, and the driving data set includes acceleration data, vibration data, and trajectory data; an object attention detection module for using a multimodal three-dimensional object detection network to perform object attention detection on the three-dimensional point cloud fusion result and the driving data set, and outputting a three-dimensional bounding box and a classification confidence degree of the track deformation area according to the object attention detection result.
[0008] One or more technical solutions provided in the present application have at least the following beneficial effects:
[0009] By using multi - camera and lidar integrated in the track inspection vehicle to collect track data, a time - series multi - modal dataset is established. The time - series multi - modal dataset includes an image dataset and lidar point cloud data. Configure the dynamic trust degree of the time - series multi - modal dataset, and perform three - dimensional point cloud fusion of the time - series data of the time - series multi - modal dataset according to the dynamic trust degree to establish a three - dimensional point cloud fusion result. Establish a driving dataset, where the driving dataset is a dataset of the train under different states, and the driving dataset includes acceleration data, vibration data, and trajectory data. Use a multi - modal three - dimensional object detection network to perform target - focused detection on the three - dimensional point cloud fusion result and the driving dataset, and output the three - dimensional bounding box and classification confidence of the track deformation area according to the target - focused detection result. That is to say, through multi - camera and lidar for multi - modal acquisition of track data, and fusing the time - series multi - modal dataset according to the dynamic trust degree of the data, the monitoring of the track under different dynamic states is realized. Using a multi - modal three - dimensional object detection network to perform real - time target - focused detection on the three - dimensional point cloud fusion result and the driving dataset, by fusing multi - modal data and using artificial intelligence algorithms, the track deformation area is accurately identified, and the three - dimensional bounding box and classification confidence are output, further improving the accuracy and recognition ability of track deformation detection.
[0010] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understood through the following description. Brief Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in this application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0012] Figure 1 It is a flowchart of the traffic track deformation detection method based on multi - modal three - dimensional point cloud fusion of this application.
[0013] Figure 2 It is a structural diagram of the traffic track deformation detection system based on multi - modal three - dimensional point cloud fusion of this application.
[0014] Description of the drawing reference numerals: data acquisition module 11, data fusion module 12, driving data set establishment module 13, target attention detection module 14. Detailed implementation manners
[0015] This application provides a traffic track deformation detection method and system based on multi-modal three-dimensional point cloud fusion, which solves the technical problem in the prior art that due to the failure to effectively fuse different data, the dynamic deformation detection ability of the track under different operating states is weak, further affecting the accuracy of track deformation detection. By using a multi-camera and a lidar to perform multi-modal acquisition of track data, and fusing the time-series multi-modal data set according to the dynamic trust degree of the data, the monitoring of the track under different dynamic states is realized. By using a multi-modal three-dimensional object detection network to perform real-time target attention detection on the three-dimensional point cloud fusion result and the driving data set, and by fusing multi-modal data and using artificial intelligence algorithms, the track deformation area is accurately identified, and a three-dimensional bounding box and classification confidence are output, further improving the accuracy and recognition ability of track deformation detection.
[0016] Next, the technical solutions in this application will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described here. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application. In addition, it should be noted that for the sake of description, only the parts related to this application are shown in the accompanying drawings rather than all of them.
[0017] Embodiment 1, please refer to the attached Figure 1 , this application provides a traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion. Among them, the traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion is executed by a traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion specifically includes the following steps:
[0018] S100: Use the multi-camera and lidar integrated on the track detection vehicle to collect track data and establish a time-series multi-modal data set, where the time-series multi-modal data set includes an image data set and lidar point cloud data.
[0019] Specifically, an orbital inspection vehicle is a vehicle specifically used to detect the status and performance of railway tracks. It is usually equipped with various sensors and detection devices and can travel along the tracks to collect track data in real time. A multi-camera and lidar are integrated on the orbital inspection vehicle for track data collection. Among them, the multi-camera refers to a device integrated with multiple cameras that captures image information from multiple perspectives, thereby improving the depth and accuracy of the data. It is usually used to obtain image information on the track surface, has good stereo perception ability, and can perform three-dimensional reconstruction and more accurate defect identification. Lidar is a technology that uses lasers to scan the surface of an object and obtains the distance of the object by measuring the laser reflection time. High-precision three-dimensional point cloud data is generated by lidar to depict the geometric shape and spatial structure of the track, and is used to detect problems such as track settlement, dislocation, and cracks.
[0020] The multi-camera captures images of the track surface from different perspectives through multiple lenses. The images captured by each lens are recorded and synchronized in real time to form a time-sequential image data set, which contains detailed features of the track surface, such as cracks and settlements. The lidar emits lasers to the track and receives the reflected signals to obtain three-dimensional spatial point cloud data of the track surface, recording information such as the height, position change, and surface unevenness of the track surface. A time-sequential multi-modal data set refers to a data set that contains different types of data (such as image data and laser point cloud data), and these data are collected in chronological order. The time-sequential nature means that the data has time stamps, so that the change trend of the data over time can be tracked.
[0021] By integrating the multi-camera and lidar, track data can be collected from multiple angles and dimensions simultaneously, forming a comprehensive and coherent time-sequential multi-modal data set, which can provide richer information for analyzing and detecting track deformation.
[0022] Furthermore, the present application S100 includes:
[0023] Using image annotation to perform image annotation segmentation on the image data in the collected data, establishing an annotation segmentation result; extracting the track image from the annotation segmentation result, and using the track image to construct the time-sequential multi-modal data set.
[0024] Specifically, image annotation and segmentation are performed on the image data in the collected track data through an image annotation tool. Image annotation is the process of identifying and classifying specific regions or objects in an image, usually by using an image annotation tool to identify and classify specific regions or objects in the image. During the image annotation process, the annotation tool helps the user automatically mark the parts related to the track in the image, usually in the form of bounding boxes or polygons, to distinguish the track area from other parts (such as the background, obstacles, etc.). For example, suppose there is an image containing a track, a station, and a green area. The segmentation algorithm will separately segment the track area and label it as the track category. The annotation segmentation result is the output result after image annotation and segmentation, usually a segmented image with the track area marked, and each segmented area may correspond to different categories (such as track, background, obstacle, etc.).
[0025] In image annotation and segmentation, artificial intelligence trains a deep learning model to automatically identify and segment target objects or regions from images. Prepare the collected track image data, including the track itself and the background of the surrounding environment (such as trees, buildings, etc.). The image data usually has a high resolution and can clearly show the track and its deformation features. The data set needs to have timestamps so that these images can be synchronized with other sensor data (such as laser point cloud data) according to the time sequence. Obtain a data set containing annotated regions, including images with the track area, background area, and other important features (such as cracks, settlements, etc.) marked. Preprocess the image data set, such as scaling, normalization, etc. Build a convolutional neural network. The convolutional neural network processes the image data through structures such as convolutional layers, pooling layers, and fully connected layers, and can effectively extract the feature information in the image. During the training process, the image is processed through the convolutional layer of the network to extract the features in the image. By calculating the difference (i.e., loss) between the output and the actual label, the network will perform backpropagation and continuously optimize the weights of the network until the network can efficiently perform the image segmentation task. Suppose 5000 annotated image data are used during model training. After multiple training epochs, the convolutional neural network model gradually improves the accuracy of image annotation and segmentation.
[0026] After training is completed, the image data in the collected data is input into the trained convolutional neural network model to automatically identify targets such as the track area, background area, crack or settlement area, etc., and perform pixel-level segmentation to obtain the annotation segmentation result. The annotation segmentation result is usually an image with color markings, where different target areas are distinguished by different colors or labels. For example, the track part may be marked red, the crack part marked blue, and the background part marked green.
[0027] Extract the track image according to the annotation segmentation result, which contains the structure and deformation characteristics of the track, usually including information such as the longitudinal and lateral offsets and settlement of the track. The extracted track image is fused with the laser point cloud data (including its corresponding timestamp) to construct a time-series multi-modal dataset, which contains the track image data and laser point cloud data at different time points and can provide comprehensive multi-dimensional information for subsequent track deformation detection.
[0028] Through image annotation segmentation, the part related to the track can be accurately extracted and combined with the laser point cloud data to provide more accurate track deformation information, realize the automatic detection of track deformation, and thus improve the efficiency and accuracy of track maintenance.
[0029] S200: Configure the dynamic trust of the time-series multi-modal dataset, and perform three-dimensional point cloud fusion of the time-series data of the time-series multi-modal dataset according to the dynamic trust to establish a three-dimensional point cloud fusion result.
[0030] Furthermore, S200 of the present application includes:
[0031] Perform time-series analysis of the detection scenarios of the track detection vehicle to generate a time-series detection scenario set; predict the acquisition effects of multi-cameras and lidar on the time-series detection scenario set, and generate a first dynamic factor according to the acquisition effect prediction result; evaluate the data acquisition effect of the time-series multi-modal dataset, and generate a second dynamic factor according to the data acquisition effect evaluation result; establish the dynamic trust using the first dynamic factor and the second dynamic factor.
[0032] Specifically, analyze the detection scenarios of the track detection vehicle, that is, the physical environment reflected by the data collected by the sensors when the track detection vehicle is performing the detection task, including the geometric characteristics of the track, the surrounding environment, and weather and other factors. Analyze the detection scenarios of the track detection vehicle over time, and evaluate the acquisition effects of multi-cameras and lidar under these different conditions according to the environmental changes at different time points. During the driving process of the track detection vehicle, environmental factors (such as lighting, weather, field of view, track status, etc.) are constantly changing, which will affect the effect of the sensors.
[0033] Perform time-series analysis on the environmental factors at different times (such as the lighting changes in time periods, different weather conditions, and changes in surrounding obstacles), not only considering the environmental changes, but also evaluating how these changes affect the acquisition effects of multi-cameras and lidar. According to the changes of environmental factors at different time points, form a time-series detection scenario set. Based on time-series analysis, analyze the influence of different environmental characteristics (such as lighting, weather, obstacles, etc.) on the sensors.
[0034] Build a convolutional neural network to predict the acquisition effect based on environmental features (such as weather, lighting) and sensor data (images, point cloud data). Obtain a historical dataset, including image data (from multi-cameras) and point cloud data (from lidar) collected by past track inspection vehicles at different times and under different environmental conditions, containing timestamps and environmental labels, such as weather conditions (sunny, cloudy, rainy, foggy, etc.), lighting intensity, geometric shape of the track, etc. Process the historical dataset for image data, adjust the size of the images, perform normalization, and scale the pixel values to between 0 and 1; point cloud data usually requires preprocessing such as sampling, denoising, and spatial regularization to ensure the quality and consistency of the data.
[0035] Use convolutional layers to extract spatial features (such as edges, textures, etc.) in the images; use pooling layers to reduce the size of the feature maps and extract the most significant features; use fully connected layers to transform the extracted features into prediction outputs, such as acquisition effect scores. Use the ReLU activation function after each convolutional layer and fully connected layer to enhance the non-linear modeling ability. Input the image data and environmental features into the convolutional neural network model; calculate the difference (loss) between the prediction result and the actual acquisition effect label; update the model weights through the backpropagation algorithm to minimize the loss function; perform multiple training epochs until the model converges to obtain a trained effect prediction model.
[0036] Input the time-series detection scenario set (such as weather, lighting conditions) of the track inspection vehicle into the effect prediction model to predict the acquisition effects of the multi-cameras and lidar, and calculate the first dynamic factor. The first dynamic factor is a quantitative value indicating the reliability of the sensor acquisition effect in the current environment. For example, assume that on a sunny day, the predicted score for the multi-camera is 0.9 and for the lidar is 0.95, then the first dynamic factor is calculated by the weighted average method as 0.925, indicating that the sensor acquisition effect is good in this scenario; on a rainy day, the predicted effect is poor, so the first dynamic factor is low (such as 0.4).
[0037] At the same time, conduct a quality assessment of the collected time-series multi-modal dataset, and the evaluation criteria include data integrity (whether there is missing data), accuracy (such as image clarity, point cloud density, etc.), noise level, etc. Conduct quality checks on the actually collected image and point cloud data. For example, check whether there are problems such as blurring, overexposure, and low contrast in the images; check whether the point cloud data is complete, dense, and accurate. Generate the second dynamic factor based on the evaluation of the data acquisition effect. The second dynamic factor is a quantitative index based on the quality of the actually collected data, reflecting the performance of the data under actual conditions.
[0038] By combining the first dynamic factor and the second dynamic factor, a comprehensive dynamic trust level is established, which respectively includes the dynamic trust levels of the multi-camera and the lidar. The dynamic trust level is a quantified value representing the credibility of the data under the current environment and actual acquisition conditions. The dynamic trust level can combine the first dynamic factor and the second dynamic factor through weighted averaging. For example, if the first dynamic factor is 0.9 and the second dynamic factor is 0.8, the dynamic trust level calculated using the weighted average method is 0.85. If the predicted acquisition effect is very good, but the quality of the actually acquired data is poor, then the dynamic trust level will also decrease accordingly.
[0039] By performing timing analysis, acquisition effect prediction, and data acquisition effect evaluation, a dynamic trust level can be established, which is a key indicator reflecting the reliability of the data set. The introduction of the dynamic trust level enables the track deformation detection to adjust its operations according to the actual data quality and environmental conditions, thereby improving the accuracy and robustness of the detection.
[0040] Furthermore, the present application further includes the following steps:
[0041] Perform timestamp alignment on the time series multi-modal data set to establish a time series multi-modal data set after timestamp alignment; obtain the time series running speed of the track detection vehicle, and synchronously obtain the acquisition positions of the multi-camera and the lidar; establish time compensation according to the time series running speed and the acquisition positions, and perform alignment and reconstruction on the time series modal data set after timestamp alignment according to the time compensation to establish an alignment and reconstruction result; use the alignment and reconstruction result and the dynamic trust level to complete the three-dimensional point cloud fusion of the time series data.
[0042] Extract the general key nodes of the track, perform feature extraction on the image data and lidar point cloud data in the alignment and reconstruction result based on the general key nodes to establish key matching features; obtain the adjacent distances of the general key points, perform demand matching for additional key points according to the adjacent distances and the recognition accuracy requirements to establish a demand matching result; perform feature recognition on the image data and lidar point cloud data between the key nodes according to the demand matching result to establish additional key matching features; use the key matching features to perform a primary registration search on the image data and lidar point cloud data in the alignment and reconstruction result to establish a primary registration search result; use the additional key matching features to perform a secondary registration search based on the primary registration search result to establish a secondary registration search result; use the secondary registration search result and the dynamic trust level to complete the three-dimensional point cloud fusion of the time series data.
[0043] Specifically, due to the different acquisition frequencies and time delays of the multi-camera and lidar, directly merging the data of the two may lead to temporal misalignment. Therefore, it is necessary to align the temporal multi-modal dataset according to the acquisition timestamps. By matching the timestamp of each data point with the data timestamps of other sensors, it can be ensured that the data of different sensors are synchronized at the same time point. Use a time synchronization algorithm (such as linear interpolation or nearest neighbor interpolation) to align the image and point cloud data from the multi-camera and lidar. For example, the timestamps of the image data are [1, 2, 3,..., 600], while the timestamps of the lidar point cloud data are [1.05, 2.05, 3.05,..., 600.05]. Calculate the time difference between the image data and the lidar point cloud data, which is 0.05 seconds, and adjust the timestamp of the lidar point cloud data forward by 0.05 seconds to make it exactly aligned with the timestamp of the image data. If there is missing data in the timestamp, it can be filled by interpolation.
[0044] Since the speed of the track inspection vehicle may vary at different time points, the running speed must be considered during the alignment process of the temporal data. Obtain the temporal running data of the track inspection vehicle, and synchronously obtain the acquisition positions of the multi-camera and lidar, that is, the acquisition position and running speed of the track inspection vehicle at each acquisition. According to the running speed and position of the inspection vehicle, adjust the timestamps of the data to ensure that the data of different sensors are synchronized in time and space, and eliminate the time deviation caused by different acquisition time delays, running speeds, etc. between sensors. Calculate the actual acquisition time at each time point based on the speed data and acquisition position of the track inspection vehicle, and adjust the timestamps of different sensors according to the time compensation algorithm.
[0045] Estimate the distance traveled by the track inspection vehicle within the time difference based on the speed and position of the track inspection vehicle, and use this as the basis for time compensation. For example, assume that the multi-camera captures an image, and the lidar captures a point cloud data at this moment. However, due to sensor delay or speed effects, the timestamp of the point cloud data may lag behind the image data. The multi-camera captures the image at timestamp t1, and the lidar captures the point cloud data at timestamp t2 (t2 > t1), and the speed of the track inspection vehicle during this period is 10 m / s. Assume the time difference △t = 0.5 seconds and the vehicle speed is 10 m / s, then the vehicle travels 5 meters during this period. Based on the distance traveled by the vehicle, the deviation that needs time compensation can be calculated, and the timestamp of the lidar can be adjusted to align it with the timestamp of the multi-camera.
[0046] Through the time compensation mechanism, the time-series data is aligned. According to the deviation calculated by the time compensation, the data timestamps of each sensor are corrected to ensure that the acquired data of the multi-view camera and the lidar are perfectly aligned in time. The aligned dataset will have the same timestamps and spatial positions, ensuring data synchronization. That is, the images and point cloud data for each timestamp are reconstructed according to the time compensation results to ensure that the image and point cloud data correspond at the same time point. Combining the acquisition positions of the sensors, the consistency of the image data and the point cloud data in space is ensured. Through spatial mapping with the positioning data, the data of different sensors are located at the correct geographical positions.
[0047] Through the time compensation and reconstruction of the time-series data, a time-series multi-modal dataset after alignment and reconstruction is finally generated, which contains synchronized image and point cloud data, and each data point has the same timestamp and precise spatial position.
[0048] The general key nodes refer to the spatial position points with important representativeness and stability in track detection, usually the points with significant marks on the track (such as track joints, sleeper centers, track centerlines, etc.). The general key nodes are determined through the track design drawings. According to the key nodes, the features of the general key nodes are extracted from the image data and the laser point cloud data in the alignment and reconstruction results, and the key matching features of the image data and the laser point cloud data are established.
[0049] The feature extraction of the image data and the laser point cloud data is to find the features in the images and point clouds that can represent these key nodes. The general key nodes are extracted from the images, such as corner points, edges, textures, etc. in the images. For example, if the camera takes a photo of a track joint, the feature points representing the track joint in this photo are found through existing artificial intelligence algorithms. The laser point cloud data is three-dimensional data collected through lidar technology, which records the precise positions of the surfaces of each object in the environment. The lidar emits laser beams, and the reflected laser signals are used to calculate the distances of the objects, so it can provide the positions of each object in space. The features extracted from the laser point cloud data include normal vectors (vectors describing the surface directions) and curvatures (describing the degree of surface curvature), and these features can understand the shapes and structures of areas such as track joints in the point cloud data.
[0050] By extracting these features from the image and point cloud data, a feature descriptor is created for each key node, just like attaching a unique label to each node to identify its position and features in the image and point cloud data. The key matching features are found by comparing the features in the image and the laser point cloud data to identify the similarities and correspondences between them. For example, the feature points of the track joint in the image should be spatially corresponding to the feature points of the track joint in the laser point cloud. By comparing the feature descriptors of the image data with those of the point cloud data, the matching relationship between the two is found.
[0051] After finding the key matching features between the image data and the laser point cloud data, the image data and the laser point cloud data are aligned to match them in the same coordinate system, which is registration. The purpose of a single registration search is to find a way to roughly align the image data and the point cloud data by comparing these matching features. For this purpose, the RANSAC (Random Sample Consensus) algorithm can be used. Some points are randomly selected from a large number of matching points to determine whether the required matching result conforms to the matching of most points. If some matching points are wrongly paired, the RANSAC algorithm will automatically remove them to ensure that the remaining matching points can provide the most accurate alignment result.
[0052] The adjacent distance of the general key points refers to the spatial distance between these key nodes, usually obtained by calculating the geometric distance between different key points. For example, track joints and sleepers are usually arranged at a certain interval, and this interval can be obtained by measuring the distance between adjacent nodes. By measuring the adjacent distance, it is judged which key points need to be more finely matched. For example, in some areas, the spacing of track joints is relatively large, and more additional points are needed for more accurate matching to ensure higher accuracy during alignment.
[0053] The recognition accuracy requirement refers to the accuracy requirements for the distance and data matching between different nodes during the data alignment and matching process. Based on the recognition accuracy requirement and the adjacent distance, it is judged whether additional key points are needed for matching. If the adjacent distance between some key points is relatively large, or the accuracy requirement for this area is relatively high, additional key points need to be introduced. The additional key points are located between the original key points to improve the matching accuracy. The required matching result is the matching result generated based on the adjacent distance and the accuracy requirement, including which additional key points need to be matched to meet the higher accuracy requirement.
[0054] Based on the demand matching results of additional key points, perform image data and laser point cloud data feature recognition again, and establish a corresponding relationship for the matching features in the image data and point cloud data. That is to say, similar to the previous steps, match the image data and point cloud data at the attachment key points to establish additional key matching features. The additional key matching features are the matching features of the additional key points obtained through feature recognition. The additional key points are those key points introduced based on the demand matching results and accuracy requirements after the initial matching. The additional key points are usually located between two original key points or in areas where higher-precision alignment is required.
[0055] Based on the results of the first registration search, use the additional key matching features for a second registration search to further improve the registration accuracy. For example, in the first registration search, the positions of the track joint and the sleeper are roughly aligned, but the matching accuracy in these areas may be insufficient. To further improve the accuracy, introduce additional key matching features (such as some additional points between the track joint and the sleeper) for refined matching to ensure that these areas can be aligned more precisely.
[0056] Use the additional key matching features to perform a second registration search based on the preliminary alignment results obtained from the first registration search. The purpose of the second registration search is to further optimize the alignment of the image data and the laser point cloud data based on the rough alignment results of the first registration. This process usually uses the ICP algorithm, which is a commonly used point cloud registration method that finely adjusts the registration results through iterative optimization. In the second registration search, the ICP algorithm will continuously match the closest points in the image data and the point cloud data, calculate the transformation (rotation and displacement) between them, and then further optimize the accuracy of the data alignment based on this transformation. The goal of the ICP algorithm is to precisely adjust the alignment of the image data and the point cloud data by minimizing the distance between the two sets of point clouds, so that they coincide as accurately as possible in space.
[0057] Extract general key nodes and calculate the adjacent distances. Then, through the demand matching and feature recognition of the additional key points, the registration accuracy is further improved. Next, through the first registration search and the second registration search, the precise alignment of the image and point cloud data is ensured. Finally, through the three-dimensional point cloud fusion of time-series data, not only the registration accuracy of the data is improved, but also the robustness and reliability in complex environments are enhanced, and more accurate results can be provided for the detection of details such as track deformation and cracks.
[0058] The ICP algorithm is an algorithm that optimizes the alignment of point cloud data by minimizing the distance between corresponding points in the point cloud. It finds the optimal transformation parameters (including rotation and translation) through an iterative approach to make the point cloud data and the target data as aligned as possible. The ICP algorithm performs secondary registration through the following steps: Based on the rough alignment result obtained from the primary registration search, the image data and the point cloud data already have a general spatial relationship. Next, the rough alignment result is used as the initial registration state. In the ICP algorithm, first, the closest point pairs between the image data and the point cloud data are found. These point pairs are the closest in space, and a corresponding relationship is established between them. By continuously finding the nearest neighbor of each point, the ICP algorithm can find the optimal alignment between the image data and the point cloud data.
[0059] Based on the matched point pairs, the ICP algorithm calculates a transformation matrix, usually a combined transformation including rotation and translation. The transformation matrix adjusts the image data and the point cloud data to make them as coincident as possible. The ICP algorithm optimizes the transformation matrix by minimizing the distance between the matched point pairs. In each iteration, the registration error under the current transformation matrix is calculated, and the transformation parameters are updated until the error reaches the minimum value or meets the convergence condition. Through iterative optimization, the ICP algorithm minimizes the error between the image data and the point cloud data, thus achieving high-precision alignment.
[0060] The secondary registration search result provides an accurate transformation matrix, representing the rotation and translation relationship between the image data and the point cloud data. The dynamic trust degree is a quantitative index used to evaluate the quality of different sensor data. The data of different sensors are weighted and fused using the dynamic trust degree. The data source with a higher trust degree will have a greater impact on the fusion result, while the data with a lower trust degree will have a smaller impact. Through weighted averaging or other fusion algorithms, the final fusion result, namely the three-dimensional point cloud fusion of sequential data, is obtained.
[0061] Furthermore, the present application further includes the following steps:
[0062] Configure the complementary attention mechanism for the spatial geometric information of the point cloud and the image texture detail information using the dynamic trust degree; perform the fusion of the registered sequential multi-modal data sets based on the complementary attention mechanism to establish the three-dimensional point cloud fusion result.
[0063] Specifically, according to the evaluation results of the dynamic trust degree, a complementary attention mechanism for configuring the geometric information of the point cloud space and the texture detail information of the image is configured. The calculation of the dynamic trust degree is based on the quality of each sensor data, environmental conditions, and other factors (such as noise level, data clarity, etc.). Data with a higher trust degree will be assigned a higher weight, while data with a lower trust degree will be assigned a smaller weight. The complementary attention mechanism is a deep learning method used to automatically focus on key information helpful for the final task when fusing data of different modalities. The complementary attention mechanism combines the spatial geometric information of the point cloud and the texture detail information of the image, and assigns different weights to different information.
[0064] For example, assume the dynamic trust degrees of the image data and the point cloud data are as follows: the trust degree of the image data is 0.85 (under sunny conditions, the image data is clear and the trust degree is high), and the trust degree of the point cloud data is 0.75 (under hazy weather, the point cloud data is relatively sparse and the trust degree is low). According to the dynamic trust degree, the complementary attention mechanism will assign a higher weight to the image data and give priority to focusing on the texture detail information in the image. In some areas (such as track joints), the texture information in the image may be more important for detecting track deformation or cracks, and the attention mechanism will automatically focus on these details.
[0065] Apply the configured complementary attention mechanism to the three-dimensional point cloud fusion of the time-series data after secondary registration, that is, the fusion of the registered time-series multi-modal data set. The complementary attention mechanism automatically adjusts the weights of the point cloud and the image data according to the dynamic trust degree, and preferentially retains the information with a higher trust degree during fusion to ensure that the fusion result can better reflect the actual state of the track. After the three-dimensional point cloud fusion of the time-series data, a three-dimensional point cloud fusion result is obtained, which combines the data of different sensors at different time points to generate a complete three-dimensional scene. Due to the weighted fusion of the attention mechanism, details such as track deformation and cracks can be detected more accurately.
[0066] By using the complementary attention mechanism, different sensor data can be dynamically weighted during the data fusion process to ensure that data with a higher trust degree has a greater impact on the final result, thereby improving the fusion accuracy.
[0067] S300: Establish a driving data set, where the driving data set is a data set of the bullet train in different states, and the driving data set includes acceleration data, vibration data, and trajectory data.
[0068] Specifically, a driving dataset is established to record various data of the bullet train in different states during driving, including multiple data such as the acceleration, vibration, and trajectory of the bullet train. The acceleration data is measured by an acceleration sensor (such as a triaxial accelerometer) to obtain the acceleration changes of the bullet train in various directions, which helps to judge the dynamic performance of the bullet train and detect possible abnormalities during the acceleration or deceleration process. The vibration data is collected by a vibration sensor (such as an accelerometer or a piezoelectric sensor) to obtain the vibration information generated by factors such as uneven ground and vibration of the vehicle body structure during the driving of the bullet train, which helps to evaluate the smoothness of the track and the running state of the bullet train. The trajectory data is recorded by a positioning system (such as GPS, inertial navigation system) to record the trajectory of the bullet train, including the position (latitude and longitude) and time, which helps to track the precise driving route of the bullet train and detect whether it deviates from the predetermined track. Exemplarily, the data shown in Table 1 was collected during the driving task of the bullet train:
[0069] Table 1 Driving Dataset
[0070]
[0071] According to the data in Table 1, it can be seen that the vibration of the train is abnormal at 15 minutes of operation. The sudden increase in vibration is usually a sign of an uneven track or vehicle body, indicating that there may be unevenness, looseness, or other structural problems in the track here.
[0072] By analyzing the acceleration data, the acceleration, deceleration conditions of the bullet train and possible braking failures can be understood. Observing the acceleration curve can detect whether there are abnormal acceleration changes or unstable running states. By analyzing the vibration data, it can be judged whether the track is smooth. For example, if the vibration data of some parts is significantly higher than that of other parts, there may be unevenness or damage to the track. The vibration data also helps to analyze the structural health status of the vehicle body. The trajectory data helps to track the driving path of the bullet train and analyze whether there is a phenomenon of deviating from the predetermined track. If the trajectory deviates far from the center line of the track, it may indicate track deformation or other obstacles. By establishing a driving dataset and conducting data analysis, the running state of the bullet train can be effectively monitored, and problems such as track deformation and cracks can be detected in a timely manner.
[0073] S400: Use a multi-modal three-dimensional object detection network to perform target attention detection on the three-dimensional point cloud fusion result and the driving dataset, and output the three-dimensional bounding box and classification confidence of the track deformation area according to the target attention detection result.
[0074] Furthermore, S400 of the present application includes:
[0075] Invoke the deformation recognition layer of the multi-modal 3D object detection network to perform deformation defect recognition on the 3D point cloud fusion result, and establish a deformation defect recognition result, where the deformation defect recognition result is marked with a position identifier; invoke the additional recognition layer of the multi-modal 3D object detection network to perform driving anomaly detection on the driving data set, and establish a driving anomaly detection result, where the driving anomaly detection result is marked with a position identifier; activate the authentication recognition layer of the multi-modal 3D object detection network, perform authentication recognition based on the deformation defect recognition result and the driving anomaly detection result, and complete target attention detection according to the authentication recognition result.
[0076] After performing vibration anomaly recognition on the vibration data, perform hysteresis analysis of vibration transmission and acceleration data effects to generate hysteresis time series compensation; perform trajectory anomaly recognition on the trajectory data to establish a trajectory anomaly; establish a driving anomaly detection result based on the trajectory anomaly, vibration anomaly, and hysteresis time series compensation.
[0077] Specifically, the multi-modal 3D object detection network is a deep learning network that integrates multiple data sources and is used to extract and identify information from different types of data. The multi-modal 3D object detection network combines image data, point cloud data (from lidar, etc.), and driving data (such as acceleration, vibration, and trajectory data) for comprehensive analysis to identify targets (such as track deformation, vehicle body faults, etc.). The deformation defect recognition layer in the multi-modal 3D object detection network processes the input 3D point cloud fusion result. By analyzing the point cloud data (such as track data collected by lidar), this layer identifies deformations and defects on the track.
[0078] Input the 3D point cloud fusion result after secondary registration and fusion into the deformation recognition layer for deformation defect recognition, extract the deformation features in the point cloud data, and detect deformations, cracks, or other defects on the track. The identified defects will be marked with specific position identifiers, such as cracks at the track joints. The output deformation defect recognition result will include the position of the defect (such as the coordinates of the track joint, the degree of deformation, etc.). For example, the deformation defect recognition result is shown in Table 2:
[0079] Table 2 Deformation Defect Recognition Result
[0080]
[0081] The additional recognition layer utilizes a driving dataset (including acceleration, vibration, and trajectory data) to perform driving anomaly detection and identify whether abnormal behaviors occur during the driving of high-speed trains or other vehicles. For example, a high-speed train may exhibit anomalies such as speeding, sudden braking, or track deviation, which can lead to unstable driving and may damage the track or the vehicle body. The acceleration, vibration, and trajectory data collected during driving are input into the additional recognition layer to analyze the driving data and identify abnormal behaviors. For example, if the acceleration suddenly increases, the vibration exceeds the normal range, or the trajectory deviates from the predetermined route, it may indicate a driving anomaly. The detected anomalies will be marked with specific location identifiers, indicating the time and location where the anomaly occurred.
[0082] Vibration anomaly recognition is performed on the collected vibration data to find abnormal points where the vibration exceeds the normal threshold. Abnormalities in vibration data may indicate uneven tracks, loose vehicle structures, or other operating anomalies. The real-time vibration data collected by vibration sensors is usually collected once per second or every few seconds. Based on the historical operating data of the train and expert experience, a reasonable vibration range (such as 0.1g to 0.6g) is preset in advance. When the vibration data exceeds the preset vibration range, it is marked as abnormal, and the abnormal location and time are recorded. For example, at 180s, the vibration value is 0.75, and here the track anomaly is caused by uneven tracks.
[0083] Hysteresis refers to the fact that a change in the current input will affect the output after a period of time, especially there is a certain time delay between vibration and acceleration data. In the track monitoring system, there may be a hysteresis effect between vibration data and acceleration data, especially when the vehicle is accelerating, braking, or passing over uneven tracks. The purpose of hysteresis time series compensation is to model the hysteresis effect between vibration data and acceleration data, correct the errors caused by time delay, and thus provide more accurate analysis results. Usually, a regression model is used to capture the time delay relationship between vibration and acceleration. Assuming that the relationship between acceleration and vibration is linear, the relationship between acceleration and vibration is described as: acceleration data = α · vibration data (t - △t) + β, where α and β are regression coefficients, and △t is the time delay (i.e., the degree of hysteresis).
[0084] The optimal matching time lag is found by calculating the correlation (similarity) between two signals (vibration and acceleration) at different time delays. For example, when the lag is 1 second, the dot product of the vibration signal (0.2, 0.3, 0.35, 0.4) and the acceleration signal (0.5, 0.8, 1.0, 1.2) is (0.2·0.5)+(0.3·0.8)+(0.35·1.0)+(0.4·1.2)=1.17; when the lag is 2 seconds, the dot product of the vibration signal (0.2, 0.3, 0.35) and the acceleration signal (0.0, 0.5, 0.8) is 0.43. The cross-correlation value at a lag of 1 second is greater than that at a lag of 2 seconds. Therefore, a lag of 1 second is the optimal matching lag time between vibration and acceleration.
[0085] After determining the optimal lag time, hysteresis time series compensation is performed on the acceleration data. The purpose of hysteresis compensation is to correct the error caused by hysteresis so that the acceleration data can more accurately reflect the current track dynamic response. For example, since the vibration data lags behind the acceleration data by 1 second, the acceleration data is shifted forward by 1 second to obtain the compensated acceleration data. By compensating for the hysteresis effect, the compensated acceleration data can more accurately reflect the current track dynamic response, helping to reduce the error caused by time delay.
[0086] Track anomaly identification is performed on the track data, the deviation of the track is analyzed, the distance between the actual track and the preset track is compared, and abnormal points deviating from a certain threshold are detected. Mark the time and location of the track deviation to establish a track anomaly. The track anomaly may be caused by factors such as track damage, equipment failure, or improper driver operation. Combine the results of vibration anomaly identification, track anomaly identification, and hysteresis time series compensation to determine whether there is an abnormal running of the bullet train and identify what abnormalities exist to obtain the running anomaly detection result. The running anomaly detection result has a location identifier, that is, what kind of anomaly occurred at which time point and which location. For example, there is a vibration anomaly at 45 meters of track when running for 15 minutes.
[0087] The authentication and recognition layer of the multi-modal 3D object detection network contains the deformation defect recognition result and the running anomaly detection result, further performs authentication and recognition, and performs target attention detection according to the authentication and recognition result. The authentication and recognition layer will judge whether there is an association between track defects (such as cracks or settlements) and abnormal vehicle behaviors (such as sudden acceleration, track deviation). For example, a sudden acceleration behavior of the vehicle appears in the track settlement area, which may mean that there are major problems with the track in this area. Combining the severity of the defect and the severity of the abnormal behavior, the authentication and recognition layer will comprehensively judge whether there is a potential safety hazard. Finally, an authentication and recognition result is generated to determine whether there are areas that need immediate attention, and mark the attention level for each area, such as high priority, medium priority, and low priority.
[0088] The authentication and recognition results are ultimately used for target concern detection, that is, to identify the areas that require key attention and handling due to track deformation or abnormal driving, or potential safety hazards caused by the interaction between the two. Target concern detection not only identifies the location of the problem area but also generates alerts or notifies relevant staff. Areas with high and medium priorities are marked as target areas that need to be inspected or repaired immediately.
[0089] Generally speaking, the multi-modal three-dimensional target detection network includes multiple recognition layers for performing deformation defect recognition, abnormal driving detection, and ultimately target concern detection. The deformation recognition layer is responsible for processing and analyzing three-dimensional point cloud data to identify deformation defects of the track, such as cracks, settlements, bends, etc.; the additional recognition layer processes the driving data set (acceleration, vibration, trajectory, etc.) to identify abnormal driving of the vehicle, such as sudden acceleration, sudden braking, track deviation, etc.; the authentication recognition layer, based on the deformation defect recognition results and abnormal driving detection results, comprehensively analyzes the relationship between the two, identifies potential safety hazards, and determines the areas that need attention. By comprehensively analyzing three-dimensional point cloud data and driving data, deformation defects on the track and abnormal driving of the vehicle can be effectively identified, potential safety hazards can be discovered in a timely manner, and areas of concern can be generated, effectively improving the accuracy and response speed of the track health monitoring system, providing data support for track maintenance, and ensuring the safety and stability of railway transportation.
[0090] Furthermore, the present application further includes the following steps:
[0091] Match the early warning level according to the classification confidence to establish a first early warning signal; identify abnormal coordinates based on the three-dimensional bounding box and establish a second early warning signal according to the abnormal coordinate recognition results; after fusing the first early warning signal and the second early warning signal, perform deformation early warning.
[0092] Specifically, according to the target concern detection results, output the three-dimensional bounding box and classification confidence of the track deformation area. The target concern detection results include the track areas that require key attention. The three-dimensional bounding box is a cuboid frame in three-dimensional space used to represent the spatial position of the track deformation, which can accurately mark the coordinate range of the deformation area, helping maintenance personnel locate the problem area. The three-dimensional bounding box includes the starting coordinates (the starting position of the bounding box), length, width, and height (the span of the bounding box on the x, y, and z axes, representing the size of the deformation area), and the rotation angle (if the deformation area is inclined, the bounding box may need to be rotated to accurately fit the track deformation area).
[0093] The classification confidence represents the confidence value regarding whether a certain track deformation area belongs to a specific problem (such as cracks, settlement, etc.). Through the deformation defect recognition layer, information such as the location and type of the track deformation area is obtained, and the probability of whether this area is a track crack, settlement, or other defect is calculated. The classification confidence is usually a numerical value between 0 and 1, indicating the confidence level of the model for a certain category. The higher the confidence, the more certain the model is that this area belongs to a specific category (such as track cracks, settlement, etc.).
[0094] The calculation of the classification confidence usually depends on a trained classification model, which analyzes the input feature data and calculates the probability of belonging to each category. The classification model is usually a softmax classifier, which will return a probability distribution for each category. The classification confidence is the maximum probability value output by softmax, indicating the confidence level of the model for this category. Suppose the model outputs 0.1, 0.7, 0.2, indicating three categories. The confidence in the first category is 0.1, the confidence in the second category is 0.7, and the confidence in the third category is 0.2. The output classification confidence is 0.7, indicating that the model believes that this target area is most likely to belong to the second category (possibly a track crack). Each target area will obtain a classification confidence, and this confidence can be output together with the location identification of this area. Based on the classification confidence, the severity of the track defect is evaluated.
[0095] Set a confidence threshold. According to the classification confidence, determine the corresponding warning level and generate a first warning signal. Compare the three-dimensional bounding box with the actual geometric data of the track to identify whether there are abnormal coordinates outside the normal range. According to the recognition result of the abnormal coordinates, generate a second warning signal. By fusing the first warning signal and the second warning signal, perform track warning, comprehensively consider the classification confidence and abnormal coordinate information, and judge whether further processing or emergency response is required. Perform weighted fusion according to the priorities of the two signals for deformation warning, and send a warning to the maintenance team through methods such as interface, text message, and email, including the warning level, abnormal coordinates, possible influence range, and recommended inspection time. For example, if the first warning signal is of high priority and the second warning signal is of high priority, then the comprehensive track deformation warning for this area is of high priority; if one of the signals is of low priority, then judge whether further monitoring is required according to the comprehensive analysis. By fusing the classification confidence and the three-dimensional bounding box, accurately identify the area of track deformation and quickly locate potential problems.
[0096] In summary, the traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion provided by this application has the following beneficial effects:
[0097] By using the multi - camera and lidar integrated on the track inspection vehicle to collect track data, a time - series multi - modal dataset is established. The time - series multi - modal dataset includes an image dataset and lidar point cloud data; configure the dynamic trust degree of the time - series multi - modal dataset, and perform 3D point cloud fusion of the time - series data of the time - series multi - modal dataset according to the dynamic trust degree to establish a 3D point cloud fusion result; establish a driving dataset, where the driving dataset is a dataset of the bullet train in different states, and the driving dataset includes acceleration data, vibration data, and trajectory data; use a multi - modal 3D object detection network to perform object - focused detection on the 3D point cloud fusion result and the driving dataset, and output the 3D bounding box and classification confidence of the track deformation area according to the object - focused detection result. That is to say, through multi - camera and lidar for multi - modal acquisition of track data, and fusing the time - series multi - modal dataset according to the dynamic trust degree of the data, the monitoring of the track in different dynamic states is realized. Using the multi - modal 3D object detection network to perform real - time object - focused detection on the 3D point cloud fusion result and the driving dataset, by fusing multi - modal data and using artificial intelligence algorithms, the track deformation area is accurately identified, and the 3D bounding box and classification confidence are output, further improving the accuracy and recognition ability of track deformation detection.
[0098] Embodiment 2. Based on the same inventive concept as the traffic track deformation detection method based on multi - modal 3D point cloud fusion in the foregoing Embodiment 1, the present application also provides a traffic track deformation detection system based on multi - modal 3D point cloud fusion. Please refer to the attached Figure 2 , the traffic track deformation detection system based on multi - modal 3D point cloud fusion includes:
[0099] A data acquisition module 11, which is used to use the multi - camera and lidar integrated on the track inspection vehicle to collect track data, establish a time - series multi - modal dataset, where the time - series multi - modal dataset includes an image dataset and lidar point cloud data; a data fusion module 12, which is used to configure the dynamic trust degree of the time - series multi - modal dataset, and perform 3D point cloud fusion of the time - series data of the time - series multi - modal dataset according to the dynamic trust degree to establish a 3D point cloud fusion result; a driving dataset establishment module 13, which is used to establish a driving dataset, where the driving dataset is a dataset of the bullet train in different states, and the driving dataset includes acceleration data, vibration data, and trajectory data; an object - focused detection module 14, which is used to use a multi - modal 3D object detection network to perform object - focused detection on the 3D point cloud fusion result and the driving dataset, and output the 3D bounding box and classification confidence of the track deformation area according to the object - focused detection result.
[0100] Further, the data acquisition module 11 in the traffic track deformation detection system based on multi - modal 3D point cloud fusion is further used for:
[0101] Perform image annotation segmentation on the image data in the collected data using image annotation, and establish an annotation segmentation result; extract the track image from the annotation segmentation result, and construct the time-series multi-modal data set using the track image.
[0102] Furthermore, the data fusion module 12 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0103] Align the timestamps of the collected time-series multi-modal data set to establish a time-series multi-modal data set after timestamp alignment; obtain the time-series running speed of the track detection vehicle, and synchronously obtain the acquisition positions of the multi-view camera and the lidar; establish time compensation according to the time-series running speed and the acquisition position, and perform alignment and reconstruction on the time-series modal data set after timestamp alignment according to the time compensation to establish an alignment and reconstruction result; perform three-dimensional point cloud fusion of the time-series data using the alignment and reconstruction result and the dynamic trust degree.
[0104] Furthermore, the data fusion module 12 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0105] Extract the general key nodes of the track, perform feature extraction on the image data and the lidar point cloud data in the alignment and reconstruction result based on the general key nodes to establish key matching features; obtain the adjacent distances of the general key points, perform demand matching of additional key points according to the adjacent distances and the recognition accuracy requirements to establish a demand matching result; perform feature recognition of the image data and the lidar point cloud data between the key nodes according to the demand matching result to establish additional key matching features; perform a primary registration search on the image data and the lidar point cloud data in the alignment and reconstruction result using the key matching features to establish a primary registration search result; perform a secondary registration search on the basis of the primary registration search result using the additional key matching features to establish a secondary registration search result; perform three-dimensional point cloud fusion of the time-series data according to the secondary registration search result and the dynamic trust degree.
[0106] Furthermore, the data fusion module 12 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0107] Perform time-series analysis of the detection scenarios of the track detection vehicle to generate a time-series detection scenario set; perform prediction of the acquisition effects of the multi-view camera and the lidar on the time-series detection scenario set, and generate a first dynamic factor according to the acquisition effect prediction result; perform evaluation of the data acquisition effects on the time-series multi-modal data set, and generate a second dynamic factor according to the data acquisition effect evaluation result; establish the dynamic trust degree using the first dynamic factor and the second dynamic factor.
[0108] Furthermore, the data fusion module 12 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0109] Utilize the dynamic trust degree to configure the complementary attention mechanism of the point cloud spatial geometric information and the image texture detail information; perform the fusion of the time-series multi-modal data sets after registration based on the complementary attention mechanism, and establish the three-dimensional point cloud fusion result.
[0110] Furthermore, the target attention detection module 14 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0111] Invoke the deformation recognition layer of the multi-modal three-dimensional target detection network to perform the deformation defect recognition of the three-dimensional point cloud fusion result, and establish the deformation defect recognition result, where the deformation defect recognition result is marked with a position identifier; invoke the additional recognition layer of the multi-modal three-dimensional target detection network to perform the driving anomaly detection of the driving data set, and establish the driving anomaly detection result, where the driving anomaly detection result is marked with a position identifier; activate the authentication recognition layer of the multi-modal three-dimensional target detection network, perform authentication recognition based on the deformation defect recognition result and the driving anomaly detection result, and complete the target attention detection according to the authentication recognition result.
[0112] Furthermore, the target attention detection module 14 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0113] After performing vibration anomaly recognition on the vibration data, perform the hysteresis analysis of the vibration transmission and the influence of the acceleration data, and generate the hysteresis time-series compensation; perform trajectory anomaly recognition on the trajectory data, and establish the trajectory anomaly; establish the driving anomaly detection result according to the trajectory anomaly, vibration anomaly and hysteresis time-series compensation.
[0114] Furthermore, the target attention detection module 14 in the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion is further configured to:
[0115] Perform early warning level matching according to the classification confidence degree, and establish the first early warning signal; perform abnormal coordinate recognition based on the three-dimensional bounding box, and establish the second early warning signal according to the abnormal coordinate recognition result; after fusing the first early warning signal and the second early warning signal, perform deformation early warning.
[0116] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The foregoing Figure 1The traffic track deformation detection method and specific examples in Embodiment 1 based on multi-modal three-dimensional point cloud fusion are equally applicable to the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion in this embodiment. Through the foregoing detailed description of the traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion, those skilled in the art can clearly know the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion in this embodiment. Therefore, for the sake of simplicity of the specification, it will not be elaborated here.
[0117] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein.
[0118] Obviously, for those skilled in the art, without departing from the principle of the present application, several improvements and modifications can also be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion, characterized in that Including: Using multi - vision cameras and lidars integrated on the track inspection vehicle to collect track data, and establishing a time - series multi - modal data set, where the time - series multi - modal data set includes an image data set and lidar point cloud data; Configuring the dynamic trust degree of the time - series multi - modal data set, and performing time - series data three - dimensional point cloud fusion on the time - series multi - modal data set according to the dynamic trust degree, and establishing a three - dimensional point cloud fusion result. Performing time - series data three - dimensional point cloud fusion on the time - series multi - modal data set according to the dynamic trust degree includes: Aligning the timestamps of the collected time - series multi - modal data set to establish a time - series multi - modal data set after timestamp alignment; Obtaining the time - series running speed of the track inspection vehicle, and synchronously obtaining the acquisition positions of the multi - vision camera and the lidar; Establishing time compensation according to the time - series running speed and the acquisition position, and performing alignment and reconstruction on the time - series modal data set after timestamp alignment according to the time compensation to establish an alignment and reconstruction result; Completing time - series data three - dimensional point cloud fusion by using the alignment and reconstruction result and the dynamic trust degree; Establishing a driving data set, where the driving data set is a data set of the EMU under different states, and the driving data set includes acceleration data, vibration data, and trajectory data; Using a multi - modal three - dimensional object detection network to perform target - focused detection on the three - dimensional point cloud fusion result and the driving data set, and outputting the three - dimensional bounding box and classification confidence of the track deformation area according to the target - focused detection result. Using a multi - modal three - dimensional object detection network to perform target - focused detection on the three - dimensional point cloud fusion result and the driving data set includes: Invoking the deformation recognition layer of the multi - modal three - dimensional object detection network to perform deformation defect recognition on the three - dimensional point cloud fusion result, and establishing a deformation defect recognition result, where the deformation defect recognition result has a position identifier; Invoking the additional recognition layer of the multi - modal three - dimensional object detection network to perform driving anomaly detection on the driving data set, and establishing a driving anomaly detection result, where the driving anomaly detection result has a position identifier; Activating the authentication recognition layer of the multi - modal three - dimensional object detection network, performing authentication recognition based on the deformation defect recognition result and the driving anomaly detection result, and completing target - focused detection according to the authentication recognition result.
2. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 1, wherein, The step of completing time - series data three - dimensional point cloud fusion by using the alignment and reconstruction result and the dynamic trust degree includes: Extracting the general key nodes of the track, and performing feature extraction on the image data and lidar point cloud data in the alignment and reconstruction result based on the general key nodes to establish key matching features; Obtaining the adjacent distance of the general key points, and performing demand matching of additional key points according to the adjacent distance and the recognition accuracy requirement to establish a demand matching result; Performing feature recognition of the image data and lidar point cloud data between the key nodes according to the demand matching result to establish additional key matching features; Performing a primary registration search on the image data and lidar point cloud data in the alignment and reconstruction result by using the key matching features to establish a primary registration search result; Perform a secondary registration search based on the results of the primary registration search using the additional key matching features, and establish the results of the secondary registration search. Complete the three-dimensional point cloud fusion of the time-series data according to the results of the secondary registration search and the dynamic trust level.
3. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 1, characterized in that, The configuration of the dynamic trust level of the time-series multi-modal data set includes: Perform the time-series analysis of the detection scenarios of the track detection vehicle to generate a set of time-series detection scenarios. Predict the acquisition effects of the multi-view cameras and lidar on the set of time-series detection scenarios, and generate a first dynamic factor according to the prediction results of the acquisition effects. Evaluate the data acquisition effects of the time-series multi-modal data set, and generate a second dynamic factor according to the evaluation results of the data acquisition effects. Establish the dynamic trust level using the first dynamic factor and the second dynamic factor.
4. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 3, wherein, The three-dimensional point cloud fusion of the time-series multi-modal data set according to the dynamic trust level to establish the three-dimensional point cloud fusion results includes: Configure the complementary attention mechanism for the spatial geometric information of the point cloud and the image texture detail information using the dynamic trust level. Fuse the registered time-series multi-modal data set based on the complementary attention mechanism to establish the three-dimensional point cloud fusion results.
5. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 1, characterized in that, The execution of the driving anomaly detection of the driving data set to establish the driving anomaly detection results includes: After identifying the vibration anomalies in the vibration data, perform the hysteresis analysis of the vibration transmission and the influence of the acceleration data to generate the hysteresis time-series compensation. Identify the trajectory anomalies in the trajectory data and establish the trajectory anomalies. Establish the driving anomaly detection results according to the trajectory anomalies, vibration anomalies, and hysteresis time-series compensation.
6. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 1, characterized in that The establishment of the time-series multi-modal data set includes: Use image annotation to perform image annotation segmentation on the image data in the collected data, and establish the annotation segmentation results. Extract the track images from the annotation segmentation results, and use the track images to construct the time-series multi-modal data set.
7. The traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to claim 1, characterized in that, The output of the three-dimensional bounding box and the classification confidence level of the track deformation area according to the target attention detection results includes: Perform the warning level matching according to the classification confidence level to establish the first warning signal. Perform the abnormal coordinate identification based on the three-dimensional bounding box, and establish the second warning signal according to the results of the abnormal coordinate identification. After fusing the first warning signal and the second warning signal, perform the deformation warning.
8. A traffic track deformation detection system based on multimodal three-dimensional point cloud fusion, characterized in that Steps for implementing the traffic track deformation detection method based on multi-modal three-dimensional point cloud fusion according to any one of claims 1 to 7, the traffic track deformation detection system based on multi-modal three-dimensional point cloud fusion includes: A data acquisition module, which is used to collect track data using the multi-view cameras and lidar integrated on the track detection vehicle, and establish a time-series multi-modal data set, and the time-series multi-modal data set includes an image data set and lidar point cloud data. A data fusion module, which is used to configure the dynamic trust level of the time-series multi-modal data set, and perform the three-dimensional point cloud fusion of the time-series data of the time-series multi-modal data set according to the dynamic trust level to establish the three-dimensional point cloud fusion results. A driving data set establishment module for establishing a driving data set, where the driving data set is a data set of the motor car in different states, and the driving data set includes acceleration data, vibration data, and trajectory data; A target attention detection module for using a multi-modal three-dimensional target detection network to perform target attention detection on the three-dimensional point cloud fusion result and the driving data set, and outputting a three-dimensional bounding box and a classification confidence of the track deformation area according to the target attention detection result.
Citation Information
Patent Citations
Three-dimensional target detection method and system based on multi-modal fusion
CN115937819A
Track fastener elastic strip deformation detection method, system and equipment and storage medium
CN117853496A