Multi-mode rainy day area road flatness detection method by means of Carla training
By training a multimodal fusion neural network using the Carla simulator and combining real and simulated data, the problems of low efficiency and accuracy in road roughness detection in rainy environments were solved, achieving efficient and low-cost road roughness detection.
Patent Information
- Application Number
- CN202510808296.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional road roughness detection methods are inefficient and costly in rainy environments, lidar performance is poor, multimodal methods have poor generalization in harsh environments, and deep learning models have limited accuracy.
With the help of the Carla simulator to construct a rainy day environment, the real data set and the simulated data set were combined to train a multimodal fusion neural network. Data was collected through lidar and vibration sensors, and multi-view convolutional neural networks and GRU networks were used to extract features. The attention mechanism was used for feature fusion to predict the variance and range of road height.
The data collection cost in real rainy environments is reduced, and the generalization performance of the model in real environments is improved. The lidar provides high-precision geometric information, and the vibration sensor captures small undulations in the road surface, which improves the robustness and accuracy of detection.
Smart Images

Figure CN120747897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road smoothness detection, and in particular to a multimodal rainy area road smoothness detection method using Carla training. Background Art
[0002] Traditional road roughness detection technology relies primarily on manual calibration (such as a roughness meter) or lidar. However, manual calibration is costly and inefficient, making it difficult to apply on a large scale. Lidar is also susceptible to interference from raindrops, which shortens detection range and reduces signal-to-noise ratio. Furthermore, point cloud data blurs in rainy environments, affecting target contour recognition. With the development of society and scientific progress, existing technologies use multimodal detection methods that combine mechanics and vision to detect road roughness. However, these methods rely on real, annotated data. In harsh environments such as rainy areas, data collection is expensive, deep learning models have poor generalization, and accuracy is limited when radar is distorted. The true value error is large, leaving room for improvement.
[0003] Based on this, a multimodal road roughness detection method for rainy areas trained with Carla is now provided, which can eliminate the drawbacks of existing technical solutions. Summary of the Invention
[0004] The purpose of the present invention is to provide a multimodal road roughness detection method in rainy areas with the help of Carla training, so as to solve the problems of low efficiency, high cost, poor performance of lidar in rainy environments and poor generalization of deep learning models of existing multimodal methods in harsh environments in the background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A multimodal rainy area road roughness detection method trained with Carla is applied to a multimodal rainy area road roughness detection system trained with Carla. The detection method specifically includes the following steps:
[0007] S1. Collect real road images and vehicle driving parameters under different rainy conditions, and construct a real dataset after annotation.
[0008] S2. Build a road scene in a rainy environment in the Carla simulator and introduce the ProxyParticleSpawn component, including but not limited to building road terrain, setting wet ground materials, and simulating rainy weather;
[0009] S3. Configure a laser radar and a vibration sensor, control the vehicle to drive stably on a simulated road, and collect multimodal data. The multimodal data includes laser radar point cloud data, vehicle three-axis acceleration data, vehicle parameters during driving, and parameter configurations of the laser radar and vibration sensor. After annotation, a simulation dataset is constructed.
[0010] S4, combining the data of the above real data set and the simulated data set into a training data set, and inputting the data into a multimodal fusion neural network for training;
[0011] S5. Use multimodal fusion neural network to extract and fuse features of point cloud and acceleration data;
[0012] S6. Based on the fused features, predict the road height variance and range, and detect the road smoothness based on the prediction results.
[0013] Preferably, constructing the rainy road scene in step S2 specifically includes the following steps:
[0014] S21. Use image editing software to create a grayscale image as a terrain height map, and import the terrain height map into the Carla simulator to generate diverse road terrains;
[0015] S22. Copy Carla's default road surface material and adjust physical parameters to create a slippery ground material. The physical parameters include but are not limited to road surface material parameters and static and dynamic friction parameters.
[0016] S23. Using the ProxyParticleSpawn component, simulate weather parameters of raindrop interference, where the weather parameters include but are not limited to raindrop density, raindrop size, and wind direction.
[0017] Preferably, configuring the laser radar and the vibration sensor in step S3 specifically includes:
[0018] S31. Create a Waymo-style lidar. Set the lidar's installation pitch angle to a negative value to tilt the beam downward and focus on the road ahead. The installation height and angle are dynamically adjusted based on vehicle speed and road type.
[0019] S32. Using a PID controller to stabilize the vehicle at a preset speed, the vehicle records the Z coordinate, the lidar signal, and the real-time acceleration data of the vehicle on the X, Y, and Z axes at a high frequency;
[0020] The parameter configuration of the laser radar includes but is not limited to installation position, pitch angle, scanning frequency, laser beam, wavelength and transmission power. The parameter configuration of the vibration sensor includes but is not limited to type, installation position, acquisition frequency, measurement range and signal processing parameters.
[0021] Preferably, the multimodal fusion neural network in step S4 includes:
[0022] Point cloud feature extraction module, which is used to extract point cloud geometric features from 3D point cloud data through multi-view convolutional neural network and self-attention mechanism;
[0023] The acceleration sequence processing module is used to extract three-axis acceleration dynamic features from vehicle vibration data through the GRU network and attention mechanism;
[0024] Feature interaction fusion module, which is used to achieve cross-modal fusion of point cloud and acceleration features through the scaled dot product attention mechanism;
[0025] A task-specific prediction module is used to output the road height variance and range based on the fused features.
[0026] Preferably, the point cloud feature extraction module specifically includes:
[0027] Multi-view feature extraction unit, used to project the point cloud to multiple viewpoints to generate a two-dimensional image, and extract local and global features through the convolution layer;
[0028] The point cloud self-attention encoding unit captures the spatial dependencies of point clouds and encodes point cloud features through a multi-head attention mechanism.
[0029] Preferably, the acceleration sequence processing module specifically includes:
[0030] The three-axis GRU unit is used to process the X-axis, Y-axis, and Z-axis acceleration sequences respectively and retain the Z-axis acceleration residual direct connection;
[0031] Acceleration attention fusion unit, used to weightedly fuse three-axis features through the attention mechanism.
[0032] Preferably, the feature interaction fusion module specifically includes:
[0033] Cross-modal interaction attention unit, which is used to input acceleration features and point cloud features, realizes modal interaction by scaling dot product attention, and outputs fused features;
[0034] The fused feature self-attention unit is used to perform self-attention encoding on the fused features.
[0035] Preferably, the task-specific prediction module specifically includes:
[0036] A shared feature extraction unit is used to extract shared features through dimensionality increase in a fully connected layer;
[0037] The task branch prediction unit is used to output the road height variance and road height range using an activation function.
[0038] Preferably, in step S4, the multimodal fusion neural network optimizes the training model parameters through a loss function, and performs back propagation and parameter update through the adaptive gradient optimization algorithm Adam, and the loss function adopts a combined loss function composed of a weighted variance loss and a range loss.
[0039] Preferably, the road smoothness detection system includes:
[0040] The scenario generation module is used to generate corresponding road scenarios and autonomous driving scenarios in the Carla simulator according to the scenario parameter data;
[0041] The sensor configuration module is used to obtain the parameter configuration and simulation data of the vibration sensor and lidar corresponding to the autonomous driving vehicle in the rainy scene of the scene parameter data;
[0042] Neural network training module, used to build a multimodal fusion neural network and train it using the training data set;
[0043] The smoothness detection module is used to detect the smoothness of the road using the trained neural network;
[0044] A memory for storing a program for implementing the multimodal rainy area road roughness detection method using Carla training;
[0045] The processor is used to execute a program for implementing the multimodal rainy area road roughness detection method trained with Carla, so as to implement the steps of the multimodal rainy area road roughness detection method trained with Carla.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The present invention uses the Carla simulator to generate synthetic data, combined with a small amount of real data, to train a multimodal neural network to perform road smoothness detection operations, which greatly reduces the cost and risk of collecting data in a real rainy environment, avoids the high cost of field testing and equipment loss, and the synthetic data can simulate rainfall and terrain of different intensities, covering edge scenes that are difficult to reproduce in the real world. After combining with a small amount of real data, the generalization performance of the model in the real environment is significantly improved, avoiding the overfitting problem. The lidar point cloud provides high-precision geometric information, and the vibration sensor captures tiny undulations on the road surface. The fusion of the two makes up for the limitations of a single sensor in rainy days, and dynamically weights the features of different modes through the attention mechanism. When the interference is severe in rainy days, the weight of the vibration data is automatically enhanced to improve robustness. The lightweight design of the neural network model can run in real time on the on-board edge device, meeting the low latency requirements of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic diagram of the step structure of the present invention.
[0049] Figure 2 Schematic diagram of the process of step S2 of the present invention.
[0050] Figure 3 Schematic diagram of the process of step S3 of the present invention.
[0051] Figure 4 Schematic diagram of the structure of the multimodal fusion neural network of the present invention.
[0052] Figure 5 It is a structural schematic diagram of the road flatness detection system of the present invention.
[0053] Notes on figure marks: point cloud feature extraction module 10, multi-view feature extraction unit 11, point cloud self-attention encoding unit 12, acceleration sequence processing module 20, three-axis GRU unit 21, acceleration attention fusion unit 22, feature interaction fusion module 30, cross-modal interaction attention unit 31, fused feature self-attention unit 32, task-specific prediction module 40, shared feature extraction unit 41, task branch prediction unit 42, scene generation module 50, sensor configuration module 60, neural network training module 70, flatness detection module 80, memory 90, processor 100. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.
[0055] In this embodiment, if Figure 1-Figure 5 As shown, the multimodal rainy area road roughness detection method trained with Carla is applied to the multimodal rainy area road roughness detection system trained with Carla. The detection method specifically includes the following steps:
[0056] S1. Collect real road images and vehicle driving parameters under different rainy conditions, and construct a real dataset after annotation.
[0057] Specifically, use acquisition equipment to measure the actual height and width of the road, the potholes on the road surface, mark the height variance and range of each section of road as the true value label, and collect data under different rainfall intensity conditions. The real road surface images include but are not limited to asphalt or concrete roads in highways, downtown areas, and urban suburbs, including flat road sections, potholes, speed bumps and other typical uneven areas. The rainy day environment includes but is not limited to light rain, moderate rain, heavy rain and rainstorm weather. The lighting conditions include but are not limited to dawn, morning, noon, dusk, night, etc. to ensure diversity. The real vehicle travels at a constant speed such as 40km / h, 60km / h, 80km / h, etc. Each acquisition time is ≥5min, and the lidar point cloud, three-axis acceleration raw voltage signal and synchronized video frame are recorded;
[0058] The acquisition equipment includes a sensor system and environmental monitoring equipment. The sensor system includes a lidar, a vibration sensor, and auxiliary sensors. The lidar type can be selected according to the actual environmental requirements. For example, a 32-line or 64-line mechanical / solid-state lidar can be used, with a wavelength set to 905nm or 1550nm, a horizontal field of view set to 360°, a vertical field of view set to ±15°, and a scanning frequency set to 10Hz. It is installed on the top of a real vehicle, and the pitch angle is adjusted to -5° to -10° to focus on the road surface. The vibration sensor can use a three-axis MEMS accelerometer with a range set to ±8g and a sampling frequency set to 100Hz. It is installed on the real vehicle chassis near the suspension system. The auxiliary sensor can record the vehicle's attitude angle and angular velocity through the inertial measurement unit. The positioning system provides positioning data and trajectory annotation. The camera collects synchronized road surface images for visual auxiliary annotation. The environmental monitoring equipment can use the weather station to record rainfall (mm / h), visibility (m), wind speed (m / s), and temperature and humidity in real time. The road surface moisture sensor quantifies the road surface water film thickness. The PTP protocol is used to synchronize multi-sensor hardware and reduce time synchronization errors.
[0059] S2. Build a road scene in a rainy environment in the Carla simulator and introduce the ProxyParticleSpawn component, including but not limited to building road terrain, setting wet ground materials, and simulating rainy weather;
[0060] Specifically, ProxyParticleSpawn is a Carla plug-in that simulates the physical effects of raindrops through a particle system. Its reflectivity model is based on Mie scattering theory. The code can be found in the Carla open source library. Use image editing software such as Global Mapper or Photoshop to convert the satellite image into a 2048×2048 grayscale image, which is used as the terrain height map. The terrain height range is set to 0 to 50 meters, with black (RGB0,0,0) representing the lowest point of 0 meters and white (RGB255,255,255) representing the highest point of 50 meters. A noise algorithm is used to generate diverse road terrain. The noise algorithm formula is: Among them A i is the amplitude, and the terrain roughness can be set to A i =20m, n=4, import the terrain height map into the Carla simulator, select the terrain actor in the Unreal Editor, click "Import Heightmap", set the Heightmap Max Height to 50 meters, check "Use Luminance" to map grayscale values to height, set the road slope range to -5° to 5°, and locally reduce the grayscale value in the terrain height map to simulate potholes with a depth of 5 to 15 cm;
[0061] Copy Carla's default road material, such as asphalt, and then adjust the physical parameters to create a slippery ground material. The physical parameters include but are not limited to road material parameters and static and dynamic friction parameters. Refer to "Rain-induced friction reduction in asphalt pavements". The static friction coefficient of dry road surface defaults to about 0.7, which is adjusted to 0.3-0.5. The dynamic friction of dry road surface defaults to about 0.6, which is adjusted to 0.2-0.4. The road surface roughness of dry road surface defaults to about 0.1, which is adjusted to 0.05 to meet rainy weather requirements. Materials can be assigned according to different rainfall intensities.
[0062] In the Carla Blueprint Editor, add a ProxyParticleSpawn component to simulate the weather parameters of raindrop interference. Weather parameters include but are not limited to raindrop density, raindrop size, and wind direction. Adjust the raindrop density to 0.5-2.0, the raindrop size diameter to 0.01m-0.05m, the wind direction to a random setting of 0°-360°, and the raindrop falling speed to 5-9m / s. This covers test requirements from mild wetness to extreme rainstorms. Key parameters can be adjusted according to different rainfall intensities as shown in the following table;
[0063]
[0064] Synchronously adjust the weather parameters in the Carla simulator, such as humidity to 80%-100%, cloud cover to 90%-100%, and visibility to 0.5-1.0, corresponding to visibility adjusted to 50-200m. In this process, raindrops will affect the effective detection range of the lidar and will be interfered by the reflectivity of raindrops. The attenuation formula of the effective detection range is: D rain =D clear ·e -βR , where R is the rainfall, β is the attenuation coefficient, and the raindrop reflectivity formula based on Mie scattering theory is: Where m is the refractive index of water, λ is the laser wavelength, and d is the raindrop diameter;
[0065] S3. Configure the LiDAR and vibration sensors, control the vehicle to drive stably on the simulated road, and collect multimodal data. The multimodal data includes LiDAR point cloud data, vehicle three-axis acceleration data, vehicle parameters during driving, and parameter configurations of the LiDAR and vibration sensors. After annotation, a simulation dataset is constructed.
[0066] Specifically, based on the working principles of lidar and vibration sensors, the CARLA simulator outputs real-world sensor and lidar data as control instructions. These instructions update the vehicle's motion state based on the vehicle dynamics model. This allows for mathematical models and formulas to describe the propagation characteristics and signal changes of lasers in different scenarios, regardless of whether the vehicle encounters an obstacle or not. This modeling combines physical principles to model reflection, scattering, and attenuation, providing a scientific basis for the high-precision generation of point cloud data.
[0067] A Waymo-style (64-line equivalent) lidar was created in the CARLA simulator. This is equivalent to a 64-line mechanical radar with a horizontal resolution of 0.1° and a vertical resolution of 0.2°, in line with Waymo's public sensor specifications. The lidar's installation pitch angle was set to a negative value to tilt the beam downward and focus on the road ahead, such as -10° to -15°, so that the beam is focused on the road 5 to 20 meters ahead. The installation height and angle were dynamically adjusted based on vehicle speed and road type. The installation height was set to 1.5 to 2 meters. The scanning frequency was automatically adjusted to 10 Hz to 20 Hz based on vehicle speed (e.g., 80 km / h) to ensure point cloud density. The transmit power was increased by 20% to 50% to compensate for energy attenuation caused by raindrops. The wavelength was set to 905 nm. A PID controller was used to stabilize the vehicle at a preset speed. The Z coordinate, lidar signal, and real-time acceleration data of the vehicle on the X, Y, and Z axes were recorded at a high frequency. The sampling frequency was increased from 100 Hz to 200 Hz. A 2 to 20 Hz bandpass filter was used to remove low-frequency noise from raindrop impacts.
[0068] The parameter configuration of the LiDAR includes but is not limited to the installation position, pitch angle, scanning frequency, laser beam, wavelength, and transmission power. The parameter configuration of the vibration sensor includes but is not limited to the type, installation position, acquisition frequency, measurement range, and signal processing parameters. The LiDAR and vibration sensor data are synchronized through hardware timestamps to control the error range.
[0069] S4, combining the data of the above real data set and the simulated data set into a training data set, and inputting the data into a multimodal fusion neural network for training;
[0070] Specifically, the real data set includes lidar point cloud, three-axis acceleration, vehicle status and annotated true value (σ 2 ,R), the simulation dataset includes synthetic data generated by Carla, covering light rain to heavy rain scenes, containing the same modal data and true values. The dataset consists of lidar point cloud data, three-axis acceleration data, vehicle driving parameters, lidar and vibration sensor parameter configurations, and road height variance and range. The point cloud frame and acceleration sequence are matched by timestamps, and the training set is divided into training set, validation set, and test set in a ratio of 70:15:15;
[0071] The above data is preprocessed and the point cloud is projected to 12 viewing angles (top view, front view, side view, etc.) to generate a 256×256 pixel two-dimensional image. Each pixel contains the average height z and density d, as shown in the following formula: Where N is the number of points in the pixel, N max To normalize the threshold, the global point cloud z coordinate is zero-meaned, the acceleration is filtered and normalized, and normalized to zero mean unit variance by axis: μ z is the mean, σ z is the standard deviation;
[0072] The multimodal fusion neural network optimizes the training model parameters through the loss function and performs backpropagation and parameter updates through the adaptive gradient optimization algorithm Adam. The loss function adopts a combined loss function consisting of a weighted variance loss and a range loss. The optimizer adopts the AdamW algorithm, with an initial learning rate of 0.001, a weight decay coefficient of 1e-5, a momentum parameter of 0.9, a training batch size of 32, and a maximum number of training rounds of 150. When the performance of the validation set does not improve after 15 rounds, the early stopping mechanism is triggered. The learning rate adjustment adopts a cosine decay strategy. In the data preprocessing stage, the input features are standardized to maintain a zero-mean unit variance distribution. Enhancement strategies such as rotation and scaling are applied to point cloud data. Time warping is used to enhance the acceleration sequence to increase the diversity of training data. The loss function consists of a weighted variance loss and a range loss, with weights of 0.55 and 0.45, respectively. The variance loss uses the logarithmic mean square error, and the range loss uses the Huber loss, with a threshold of 1.0.
[0073] S5. Use multimodal fusion neural network to extract and fuse features of point cloud and acceleration data;
[0074] Specifically, the lidar point cloud data is input into the multimodal fusion neural network, with the goal of extracting the geometric features of the road surface. The multi-view feature extraction unit 11 projects the three-dimensional point cloud to 12 fixed view angles (such as top view, front view, side view, etc.) to generate 12 two-dimensional images. Each image is subjected to a four-layer convolutional neural network (CNN) to extract features. The first layer uses 64 3×3 convolution kernels, ReLU activation, and outputs a 64-channel feature map. The second layer uses 128 3×3 convolution kernels, ReLU activation, and outputs a 128-channel feature map. The third layer uses 128 3×3 convolution kernels, ReLU activation, and outputs a 128-channel feature map. The first layer uses 256 3×3 convolution kernels, ReLU activation, and outputs a 256-channel feature map. The fourth layer uses 512 3×3 convolution kernels, ReLU activation, and outputs a 512-channel feature map. Global average pooling is performed on the multi-view features to obtain 12 512-dimensional vectors, which are concatenated to form a 1024-dimensional geometric feature. The point cloud self-attention encoding unit 12 uses an 8-head attention mechanism to encode the 1024-dimensional features, capturing the spatial dependency of the point cloud and outputting 1024-dimensional geometric features to represent the global and local geometric properties of the road.
[0075] Three-axis acceleration data (X, Y, and Z axes, 128 frames, 100Hz sampling rate) is input into the multimodal fusion neural network. The goal is to extract vehicle vibration characteristics and reflect road roughness. The three-axis GRU unit 21 is used to process the X, Y, and Z axis acceleration sequences respectively. 128 frames of data are input for each axis. A two-layer GRU network (hidden layer 128 dimensions) is used to extract temporal dynamic features. Each axis generates 128-dimensional features, and the output is 384 dimensions in total (X, Y, and Z axis splicing). The acceleration attention fusion unit 22 is used to perform four-head attention fusion on the three-axis features and calculate the weight distribution: Where Q, K, and V are query, key, and value matrices respectively, and d k is the head dimension, and the output is 128-dimensional dynamic features;
[0076] The feature interaction fusion module 30 inputs point cloud geometric features (1024 dimensions) and acceleration dynamic features (128 dimensions). The goal is to associate geometric and vibration features and enhance the representation of road roughness. The cross-modal interactive attention unit 31 calculates the scaled dot product attention and outputs a 128-dimensional interactive feature, with the acceleration feature as the query (Q) and the point cloud feature as the key (K) and value (V). The fused feature self-attention unit 32 performs self-attention encoding on the 128-dimensional features after interaction, further refining the cross-modal association and outputting a 64-dimensional fused feature.
[0077] S6. Based on the fused features, predict the road height variance and range, and detect the road smoothness according to the prediction results;
[0078] Specifically, the 64-dimensional fusion features output by the feature interaction fusion module are input to the task-specific prediction module 40, with the goal of predicting the road height variance σ 2 and range R, and uses the shared feature extraction unit 41 to extract high-order abstract features from the fusion features to serve the subsequent task branches. The shared feature extraction unit 41 has two fully connected layers. The first fully connected layer has a 64-dimensional input and a 512-dimensional output. The activation function uses the GELU function, and the formula is: GELU(x) = x·Φ(x), where Φ(x) is the standard normal distribution CDF. The second fully connected layer has a 512-dimensional input and a 256-dimensional output. The activation function uses the GELU function, and the Dropout ratio is set to 0.2. Some neurons are randomly shielded to prevent overfitting, and the final output is a 256-dimensional shared feature.
[0079] The task branch prediction unit 42 is used to predict the road height variance σ 2 And the range R, the road height variance formula is: z i The task branch prediction unit 42 has two fully connected layers. The first fully connected layer has a 256-dimensional input and a 128-dimensional output. The activation function uses the GELU function. The second fully connected layer has a 128-dimensional input and a 1-dimensional output. The activation function uses the Softplus function. The formula is: Softplus(x)=log(1+e x ), the loss function uses logarithmic mean square error, and its formula is: The range formula is: R = max (z i )-min(z i ), z i∈ road point cloud, the first fully connected layer has a 256-dimensional input and a 128-dimensional output, and the activation function uses the GELU function. The second fully connected layer has a 128-dimensional input and a 1-dimensional output, and the activation function uses the Softplus function, whose formula is: Softplus(x)=log(1+e x ), the loss function adopts Huber loss function, and its formula is: The total loss function formula is
[0080] Determine the flatness by predicting the results: variance σ 2 The typical range is set to 0~0.1m 2 (0: completely flat, 0.1: severely uneven), the typical range of the extreme difference R is set to 0~0.3m (0: no height difference, 0.3: large potholes), and the flatness classification standard can be shown in the following table:
[0081]
[0082]
[0083] Among them Figure 1 and Figure 4 As shown, the multimodal fusion neural network in step S4 includes:
[0084] A point cloud feature extraction module 10 is used to extract point cloud geometric features from three-dimensional point cloud data through a multi-view convolutional neural network and a self-attention mechanism;
[0085] Specifically, the point cloud feature extraction module 10 includes a multi-view feature extraction unit 11 and a point cloud self-attention encoding unit 12. The multi-view feature extraction unit 11 is used to project the point cloud to multiple views to generate a two-dimensional image, and extract local and global features through the convolution layer. The point cloud self-attention encoding unit 12 captures the spatial dependency of the point cloud through the multi-head attention mechanism and encodes the point cloud features.
[0086] The acceleration sequence processing module 20 is used to extract three-axis acceleration dynamic features from vehicle vibration data through a GRU network and an attention mechanism;
[0087] Specifically, the acceleration sequence processing module 20 includes a three-axis GRU unit 21 and an acceleration attention fusion unit 22. The three-axis GRU unit 21 is used to process the X-axis, Y-axis, and Z-axis acceleration sequences respectively and retain the Z-axis acceleration residual direct connection. The acceleration attention fusion unit 22 is used to weightedly fuse the three-axis features through the attention mechanism.
[0088] A feature interaction fusion module 30 is used to achieve cross-modal fusion of point cloud and acceleration features through a scaled dot product attention mechanism;
[0089] Specifically, the feature interaction fusion module 30 includes a cross-modal interaction attention unit 31 and a fusion feature self-attention unit 32. The cross-modal interaction attention unit 31 is used to input acceleration features and point cloud features, realize modal interaction by scaling dot product attention, and output fusion features. The fusion feature self-attention unit 32 is used to perform self-attention encoding on the fused features.
[0090] A task-specific prediction module 40 for outputting road height variance and range based on the fused features;
[0091] Specifically, the task-specific prediction module 40 includes a shared feature extraction unit 41 and a task branch prediction unit 42. The shared feature extraction unit 41 is used to extract shared features through dimensionality increase of the fully connected layer, and the task branch prediction unit 42 is used to output the road height variance and road height range using an activation function.
[0092] Among them Figure 5 As shown, the road roughness detection system includes:
[0093] The scenario generation module 50 is used to generate corresponding configured road scenarios and autonomous driving scenarios in the Carla simulator according to the scenario parameter data. Based on the input parameters (rainfall intensity, road friction coefficient, terrain height map), the module outputs a configured simulation environment, including a road grid, weather particle effects, and physical parameters. It also provides the synthetic point cloud and acceleration data required for training, with the data format aligned with the real dataset.
[0094] The sensor configuration module 60 is used to obtain parameter configurations and simulation data of the vibration sensor and lidar corresponding to the autonomous driving vehicle in the rainy scene of the scene parameter data, set the parameters of the lidar and vibration sensor, and synchronously collect data. The Carla simulator is connected to the scene generation module 50 and the sensor configuration module 60;
[0095] A neural network training module 70 is used to construct a multimodal fusion neural network and train it using a training data set, receive the synthetic data from the scene generation module 50 and the sensor configuration module 60, and train the multimodal fusion network in conjunction with real data;
[0096] The roughness detection module 80 is used to detect road roughness using a trained neural network, load the trained model, process sensor data in real time, and output a roughness index. The input is a real-time point cloud and acceleration sequence, and the output is a roughness result or grade, which is displayed in a visual graph.
[0097] Memory 90, for storing a program for implementing a multi-modal rainy area road roughness detection method using Carla training;
[0098] The processor 100 is configured to execute a program for implementing a multimodal rainy area road roughness detection method using Carla training, thereby implementing the steps of the multimodal rainy area road roughness detection method using Carla training;
[0099] Specifically, the system also includes a collection device for collecting real data, a background management device and a user terminal. The above modules and the multimodal fusion neural network are all arranged inside the background management device. The background management device has a built-in database, a media interface, a storage and an actuator. The background management device is equipped with a server, which can be an electronic device such as a desktop computer, a laptop computer, a server, etc. It can be an independent proxy server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This server is an electronic device used to provide background services. The Carla simulator is the underlying data engine, the multimodal fusion neural network is the core algorithm, and the system module is the scheduling center to achieve closed-loop detection;
[0100] Example 1
[0101] Assume that the rainfall intensity in a downtown area is moderate (1.5 mm / h), and the road contains frequent speed bumps, local potholes (5 cm deep), and asphalt repair areas. There is also water splash from other vehicles.
[0102] An urban road model was constructed in Carla, with the raindrop density set to 1.2 and the road friction coefficient adjusted to 0.35 to simulate a slippery asphalt material. The ProxyParticleSpawn component was used to dynamically generate water splash particle effects from vehicles to enhance the authenticity of interference. The lidar pitch angle was set to -12°, and the scanning frequency was increased to 15Hz to cope with dense vehicle scenes. The real-world rainy day data collected in downtown areas (including annotated speed bumps and potholes) was integrated with the synthetic data generated by Carla. The multimodal neural network automatically suppressed the point cloud noise caused by water splash interference through the attention mechanism and weighted the vibration sensor data. The above data was compared with the actual road values, and the predicted road height variance error (RMSE) was 0.015m. 2 The range error is 0.04m, which is significantly improved compared with the traditional pure lidar method;
[0103] Example 2
[0104] Assuming a complex urban traffic environment, the simulation involved multiple autonomous vehicles cooperating on a rainy road. The vehicles were 10 meters apart, traveling at a speed of 40 km / h (the urban speed limit), with a moderate rain density of 1.2 and visibility of 3 km. The road included slope variations (-3° to 5°) and local potholes (5 cm deep). Each vehicle was equipped with a 64-line lidar (scanning frequency of 15 Hz) with a pitch angle of -10° and a vibration sensor (sampling frequency of 150 Hz).
[0105] In Carla, a city road mesh was generated, and a height map was imported to simulate potholes and slopes. Moderate rain parameters were set using the ProxyParticleSpawn component, and the road friction coefficient was adjusted to 0.35. The OpenCDA framework was used to control vehicles to maintain platooning. Real-time point cloud and vibration data were shared through V2V communication, making the collected data more realistic for multi-vehicle interaction on real roads. This increased the complexity and authenticity of the data, thereby improving the generalization ability of subsequent model training. The lidar point cloud (including raindrop noise), three-axis acceleration (including multi-vehicle interaction vibration interference), and relative positions of each vehicle were recorded. The variance and range of road height in the multi-vehicle scenario were annotated. The results showed that the road height range prediction error was small and highly correlated with the International Roughness Index (IRI), which was a significant improvement over traditional lidar methods.
[0106] Example 3
[0107] Assume that the road is a mountain road with continuous curves and heavy rain (rainfall 3mm / h). There are many subsidence areas (height difference 20cm) and local water accumulation on the road surface.
[0108] A height-difference terrain map (grayscale range 0-50m) was imported into Carla. The raindrop density was set to 2.5, and the wind speed was 8m / s to simulate crosswinds. The lidar wavelength was adjusted to 1550nm to penetrate the rain curtain, and the transmission power was increased by 40%. A pre-trained model that had never seen this terrain was used for direct prediction. Real-time point cloud and vibration data (sampling frequency 200Hz) were input. The feature fusion module enhanced the vibration feature weights of the flooded areas through cross-modal interaction. The results showed that the road height range prediction error was small and highly correlated with the International Roughness Index (IRI), which was a significant improvement over traditional lidar methods.
[0109] Example 4: Migration verification using a real environment
[0110] Data from a 10-kilometer section of a highway was collected on a real rainy day (rainfall of 1.5 mm / h). A model trained solely on synthetic data was used to generate predictions. Compared to the results from a professional roughness meter, the correlation coefficient between the model output and the International Roughness Index (IRI) reached 0.92, with a false alarm rate of less than 5%. In areas of accumulated water (2 cm deep), vibration sensor data corrected for LiDAR misjudgments caused by water reflections.
[0111] Example 5: Performance comparison experiment under different rainfall intensities
[0112] In Carla, we constructed a similar road terrain, such as an asphalt pavement with 5cm-deep potholes. We used the ProxyParticleSpawn component to simulate four different rainfall intensities. We maintained the same sensor configuration, such as a lidar pitch angle of -12° and a vibration sensor sampling rate of 200Hz. We then compared this method's multimodal fusion approach (point cloud + vibration) with a lidar-only approach.
[0113]
[0114] The results show that the detection distance of the lidar decreases during heavy rain. However, this solution compensates for the vibration data, and the error increase is lower than the traditional method. The traditional method has an error of more than 0.1m in heavy rain. 2 ,cannot meet the road maintenance standards, while this method automatically increases the weight of vibration data in heavy rain,which can compensate for the performance degradation of the lidar;
[0115] Example 6: Comparative experiment with a single sensor method
[0116] Assuming the same moderate rain scenario (rainfall of 1.5 mm / h), the test results of multimodal fusion (lidar + vibration sensor), pure lidar (64 lines, pitch angle -12°), and pure vibration sensor (three-axis MEMS, 100 Hz sampling) were compared;
[0117]
[0118] The results show that lidar has limitations. Reflections from accumulated water can lead to misjudgments (identifying water surfaces as potholes). Vibration sensors, on the other hand, can distinguish high-frequency impact features, reducing the probability of misjudgment. Vibration sensors also have limitations and are insensitive to gentle slopes (slopes < 3°). Lidar's geometric measurement error is small. This solution, through feature fusion, can effectively reduce the probability of these defects.
[0119] In summary, the present invention uses the Carla simulator to generate synthetic data, combines it with a small amount of real data, and trains a multimodal neural network to perform road smoothness detection operations. This greatly reduces the cost and risk of collecting data in a real rainy environment, avoids the high cost of field testing and equipment loss. The synthetic data can simulate rainfall of different intensities and terrains, the lidar point cloud provides high-precision geometric information, and the vibration sensor captures tiny undulations in the road surface. The fusion of the two makes up for the limitations of a single sensor in rainy days and has good application prospects.
[0120] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multimodal rainy area road roughness detection method trained with Carla is applied to a multimodal rainy area road roughness detection system trained with Carla, characterized in that: The detection method specifically comprises the following steps: S1. Collect real road images and vehicle driving parameters under different rainy conditions, and construct a real dataset after annotation. S2. Build a road scene in a rainy environment in the Carla simulator and introduce the ProxyParticleSpawn component, including but not limited to building road terrain, setting wet ground materials, and simulating rainy weather; S3. Configure a laser radar and a vibration sensor, control the vehicle to drive stably on a simulated road, and collect multimodal data. The multimodal data includes laser radar point cloud data, vehicle three-axis acceleration data, vehicle parameters during driving, and parameter configurations of the laser radar and vibration sensor. After annotation, a simulation dataset is constructed. S4, combining the data of the above real data set and the simulated data set into a training data set, and inputting the data into a multimodal fusion neural network for training; S5. Use multimodal fusion neural network to extract and fuse features of point cloud and acceleration data; S6. Based on the fused features, predict the road height variance and range, and detect the road smoothness based on the prediction results.
2. The multimodal rainy area road roughness detection method based on Carla training according to claim 1 is characterized in that: Constructing the rainy road scene in step S2 specifically includes the following steps: S21. Use image editing software to create a grayscale image as a terrain height map, and import the terrain height map into the Carla simulator to generate diverse road terrains; S22. Copy Carla's default road surface material and adjust physical parameters to create a slippery ground material. The physical parameters include but are not limited to road surface material parameters and static and dynamic friction parameters. S23. Using the ProxyParticleSpawn component, simulate weather parameters of raindrop interference, where the weather parameters include but are not limited to raindrop density, raindrop size, and wind direction.
3. The multimodal rainy area road roughness detection method based on Carla training according to claim 1 is characterized in that: The configuration of the laser radar and the vibration sensor in step S3 specifically includes: S31. Create a Waymo-style lidar. Set the lidar's installation pitch angle to a negative value to tilt the beam downward and focus on the road ahead. The installation height and angle are dynamically adjusted based on vehicle speed and road type. S32. Using a PID controller to stabilize the vehicle at a preset speed, the vehicle records the Z coordinate, the lidar signal, and the real-time acceleration data of the vehicle on the X, Y, and Z axes at a high frequency; The parameter configuration of the laser radar includes but is not limited to installation position, pitch angle, scanning frequency, laser beam, wavelength and transmission power. The parameter configuration of the vibration sensor includes but is not limited to type, installation position, acquisition frequency, measurement range and signal processing parameters.
4. The multimodal rainy area road roughness detection method based on Carla training according to claim 1, characterized in that: The multimodal fusion neural network in step S4 includes: A point cloud feature extraction module (10), for extracting point cloud geometric features from three-dimensional point cloud data through a multi-view convolutional neural network and a self-attention mechanism; An acceleration sequence processing module (20) is used to extract three-axis acceleration dynamic features from vehicle vibration data through a GRU network and an attention mechanism; A feature interaction fusion module (30) is used to achieve cross-modal fusion of point cloud and acceleration features through a scaled dot product attention mechanism; A task-specific prediction module (40) is used to output road height variance and range based on the fused features.
5. The multimodal rainy area road roughness detection method based on Carla training according to claim 4 is characterized in that: The point cloud feature extraction module (10) specifically includes: A multi-view feature extraction unit (11) is used to project the point cloud to multiple views to generate a two-dimensional image, and extract local and global features through a convolution layer; The point cloud self-attention encoding unit (12) captures the spatial dependencies of point clouds through a multi-head attention mechanism and encodes point cloud features.
6. The multimodal rainy area road roughness detection method based on Carla training according to claim 4 is characterized in that: The acceleration sequence processing module (20) specifically includes: The three-axis GRU unit (21) is used to process the X-axis, Y-axis, and Z-axis acceleration sequences respectively and retain the Z-axis acceleration residual direct connection; The acceleration attention fusion unit (22) is used to weightedly fuse the three-axis features through the attention mechanism.
7. The multimodal rainy area road roughness detection method based on Carla training according to claim 4 is characterized in that: The feature interaction fusion module (30) specifically includes: Cross-modal interaction attention unit (31), which is used to input acceleration features and point cloud features, realize modal interaction by scaling dot product attention, and output fusion features; The fused feature self-attention unit (32) is used to perform self-attention encoding on the fused features.
8. The multimodal rainy area road roughness detection method based on Carla training according to claim 4 is characterized in that: The task-specific prediction module (40) specifically includes: A shared feature extraction unit (41), configured to extract shared features by performing dimension increase through a fully connected layer; The task branch prediction unit (42) is used for outputting the road height variance and the road height range using an activation function.
9. The multimodal rainy area road roughness detection method based on Carla training according to claim 1, characterized in that: In step S4, the multimodal fusion neural network optimizes the training model parameters through the loss function, and performs back propagation and parameter update through the adaptive gradient optimization algorithm Adam. The loss function adopts a combined loss function composed of weighted variance loss and range loss.
10. The multimodal rainy area road roughness detection method based on Carla training according to claim 1, characterized in that: The road smoothness detection system comprises: A scenario generation module (50) is used to generate corresponding configured road scenarios and autonomous driving scenarios in the Carla simulator according to scenario parameter data; A sensor configuration module (60) is used to obtain parameter configuration and simulation data of a vibration sensor and a laser radar corresponding to the autonomous driving vehicle in a rainy day scene of the scene parameter data; A neural network training module (70) is used to construct a multimodal fusion neural network and perform training using a training data set; A smoothness detection module (80) is used to detect the smoothness of the road using a trained neural network; A memory (90) for storing a program for implementing the multimodal rainy area road roughness detection method using Carla training; A processor (100) is used to execute a program for implementing the multimodal rainy area road roughness detection method trained with Carla, so as to implement the steps of the multimodal rainy area road roughness detection method trained with Carla as described in any one of claims 1 to 9.
Citation Information
Cited By
Hardware fitting size detection system and method based on visual detection
CN121230623A
Road surface flatness measuring method and system based on multi-angle weighted correction
CN121655432A
Automobile anti-flooding online intelligent early warning system and method
CN121929089A