Vehicle vibration response prediction method based on multi-modal feature deep fusion
Through the multimodal feature fusion method combined with EfficientNet and Transformer, the problems of accuracy and computing resources in traditional vehicle vibration response signal prediction are solved, and high-accuracy vehicle vibration response prediction is achieved, which is suitable for smart city infrastructure and unmanned system clusters and other fields.
Patent Information
- Application Number
- CN202510456188.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
Smart Images

Figure CN120387131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and signal processing, and mainly relates to a method for predicting vehicle vibration response signals. Background Art
[0002] Currently, for the problem of predicting vehicle vibration response signals, it is mainly achieved by using traditional numerical simulation and analysis methods. The general processing flow is to first construct a vehicle vibration response model according to vehicle system dynamics theory or finite element analysis theory, and finally use the collected road excitation signals for simulation to achieve signal prediction. For the problem of collecting road excitation signals, due to the complex and diverse road surface conditions, including different degrees of roughness, potholes and bumps, existing measurement technologies are difficult to accurately capture all the characteristics of these irregular conditions. For the problem of vibration signal prediction, the commonly used vehicle system dynamics analysis method and finite element analysis method both have certain limitations. The vehicle system dynamics model usually simplifies the complex structure of the vehicle and the interactions between components, ignoring some high-order effects and non-linear characteristics, thus affecting the accuracy of subsequent signal prediction. The finite element analysis model requires a large amount of computing resources and time, and it is difficult to adapt to the dynamically changing driving environment.
[0003] Regarding the application of vehicle vibration response signal prediction, the development in many fields is relatively mature. Currently, vehicle vibration response signal prediction has developed from single fault diagnosis to a core technology covering multiple fields such as structural optimization, intelligent driving, and energy management. In the future, with the breakthrough of multi-modal perception fusion and edge computing technologies, its application scope will be further expanded to emerging scenarios such as smart city infrastructure and unmanned system clusters, becoming a key technical support for intelligent transportation. Therefore, to accurately predict the vehicle vibration response under complex and changing road excitations in real time, it is necessary to establish an efficient and accurate vehicle vibration response signal prediction classification method, effectively utilize the characteristics of road excitation signals, improve the prediction accuracy of vehicle vibration response signals, and provide dynamic and accurate vibration response signal predictions for many application fields of vehicle intelligent perception and control decision-making technologies, so that the vehicle can achieve better performance under various complex road conditions. Summary of the Invention
[0004] Aiming at the problems existing in the above-mentioned prior art, the technical problem to be solved by the present invention is to provide a vehicle vibration response prediction method based on deep fusion of multi-modal features. High-level semantic features of road surface images are extracted by EfficientNet, the long-term dependence relationship of vibration signals is modeled by Transformer, and finally multi-modal features are fused through a bidirectional cross-attention mechanism that follows the Dirichlet distribution to achieve accurate prediction of vehicle vibration.
[0005] The implementation steps of the technical solution are as follows:
[0006] Step 1: Based on the vehicle driving road surface image data and vibration response time series signals, construct a vehicle vibration excitation and response triple dataset;
[0007] Step 2: Based on multi-modal feature deep fusion, construct a vehicle vibration response prediction model. Use EfficientNet as the image feature extractor and Transformer as the time series feature extractor to extract the features of road surface images and vibration signals respectively. Combine the bidirectional cross-attention mechanism that follows the Dirichlet distribution for multi-modal feature fusion, and realize signal prediction through a fully connected network;
[0008] Step 3: Divide the triple dataset into a training set, a validation set, and a test set. Use the training set to train the model, select the model and tune the hyperparameters through the validation set, and finally evaluate the performance of the model on the test set;
[0009] Step 4: Use the trained vehicle vibration response prediction model to predict the vehicle vibration response signals for the uploaded vehicle driving road surface image data and vibration response time series signals.
[0010] Furthermore, the specific steps for constructing the dataset based on multi-modal data in Step 1 are as follows:
[0011] ① Real-time collect the RGB images of the front road surface through an in-vehicle camera, and synchronously collect the vehicle vibration response time series signals through an acceleration sensor, including physical quantities such as vertical acceleration and speed;
[0012] ② Perform normalization and enhancement processing on the road surface images. Scale the images to 224×224 pixels, normalize the pixel values to the range of [0,1], and standardize them according to the mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225] of the ImageNet dataset. Perform image enhancement operations such as random cropping, horizontal flipping, and slight perturbations of brightness / contrast on the images;
[0013] ③ Perform noise addition or time warping operations on the vibration response time series signals to improve the generalization of the model, and then segment the signals with a fixed time window. Each segment is called a frame, which is aligned with the image frame rate;
[0014] ④ Construct a dataset with the current vibration response signal frame as the supervision label. Each sample i is a triple extracted from the preprocessed data, consisting of the current road surface image I(i), the historical vibration signal H(i), and the current vibration signal frame V(i). The historical vibration signal is a time series containing the vibration responses within a past period of time, expressed as H(i) = [V(i - n), V(i - n + 1),..., V(i - 1)].
[0015] Further, the specific steps for constructing the vehicle vibration response prediction model in the second step are as follows:
[0016] ① The image feature extractor adopts the EfficientNet architecture and inputs a three-channel image. The first layer of the image encoder uses a 3×3 convolutional kernel to downsample the original image and initially extract basic features such as edges and textures. By stacking multiple lightweight convolutional modules (MBConv), a five-stage feature deepening structure of "shallow - middle - high - fusion - compression" is constructed. In the shallow feature extraction stage, small-kernel depth convolution is used to capture edge textures. In the middle feature abstraction stage, local patterns are extracted by combining the expansion of the convolutional kernel and channel expansion with spatial downsampling. In the high-order semantic modeling stage, dilated convolution is introduced to expand the receptive field and model the global context. In the multi-scale fusion stage, different granularity features are integrated through stacked modules, and specific features are weighted and screened. In the global compression stage, the spatial resolution is gradually compressed through stacked modules. Finally, after global average pooling and fully connected layer processing, the high-level semantic feature vector of the image is output.
[0017] ② The time series feature extractor adopts the Transformer network architecture and inputs the historical vibration time series signal, which is mapped to the latent space through learnable position encoding and the embedding layer. Multiple stacked encoding modules are used to extract long-term dependence information. Each encoding module contains a multi-head self-attention mechanism, dilated causal convolution, and a GeLU gated feed-forward network. The vibration signal sequence input to the encoding module is split into multiple attention heads. Each head projects features through a linear transformation, calculates the attention weights, and sums them up after weighting. The outputs of each head are concatenated. One-dimensional dilated causal convolution with a specific convolutional kernel is performed on the output. The dilation rate increases exponentially in a hierarchical manner. After the convolutional result is screened by the gating mechanism, processed by residual connection and layer normalization, it is projected to a high-dimensional latent space. Nonlinearity is introduced using the GeLU activation function, and then it is compressed back to the original dimension and gated-fused with the original input. Finally, the time series feature vector is output through mean aggregation in the time dimension.
[0018] ③ The multi-modal feature fusion layer of the vehicle vibration response prediction model applies a bidirectional cross-attention mechanism that follows the Dirichlet distribution. It takes as input the semantic feature vector of the road surface image and the temporal feature vector of the vibration signal. The bidirectional cross-attention interaction divides into an image-temporal attention flow and a temporal-image attention flow. The former transforms the image feature vector into a query matrix through an independent linear transformation, transforms the temporal feature into key and value matrices, calculates the correlation score through scaled dot product, and uses the reparameterization trick to sample and generate an attention weight matrix that follows the Dirichlet distribution, and weighted aggregates the key segments. The latter process is symmetric to it, and an extended attention head is used to set different position biases, and a dynamic masking mechanism is used to mask invalid temporal segments; Subsequently, through a gated adaptive fusion module, the outputs of the bidirectional attentions are concatenated to form a joint representation, and two-modal dynamic weights are generated through a fully connected layer and a Sigmoid activation function. The two-modal features are weighted and summed according to the weights, and the final fusion vector is output.
[0019] Further, the specific steps for training, validating, and testing the vehicle vibration response prediction model in step three are as follows:
[0020] ① After randomly shuffling the triple dataset, it is divided proportionally into a 70% training set, a 15% validation set, and a 15% test set;
[0021] ② Use the training set to train the model. The model calculates the prediction result through forward propagation, calculates the loss value according to the loss function, and then uses the optimization algorithm to update the model's parameters through backpropagation. By iteratively adjusting the model's parameters, the loss value is gradually reduced to make the model's prediction result closer to the true value;
[0022] ③ Use the validation set to monitor the model tuning, calculate the loss value of the model on the validation set, judge the model fitting situation by observing the performance change of the model on the validation set, adopt the grid search algorithm, traverse all key hyperparameter combinations, screen out the hyperparameter configuration that makes the model performance reach the optimal, and introduce an early stopping mechanism to effectively avoid the overfitting phenomenon of the model and improve the training efficiency;
[0023] ④ Use the test set to evaluate the model performance. Input the current road surface image and historical vibration signal in the test set into the selected model, calculate the prediction result of the model, and compare it with the current vibration signal frame in the test set. Calculate a series of evaluation indicators such as root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R 2 ) and signal peak relative error (PEAK-ERR) to measure the model performance.
[0024] Advantages of the present invention over the prior art:
[0025] (1) The present invention overcomes the problems of poor measurement accuracy in traditional road excitation and difficulty in balancing computational complexity and structural fineness in vibration signal prediction. The collected signals are more accurate and the extracted features are more diverse, which can effectively improve the prediction accuracy of vehicle vibration response signals.
[0026] (2) The present invention applies the image feature extraction advantages of the EfficientNet architecture and the temporal feature extraction advantages of the Transformer architecture to the regression prediction of vehicle vibration response signals, and fuses multi-modal features through a bidirectional cross-attention mechanism that follows the Dirichlet distribution, achieving a high prediction accuracy. This shows that when the present invention is used to predict vehicle vibration response signals, a good prediction effect can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To better understand the present invention, the following further description is made with reference to the accompanying drawings.
[0028] Figure 1 is a flowchart of the steps for establishing a vehicle vibration response prediction method based on deep fusion of multi-modal features;
[0029] Figure 2 is an algorithm flowchart of a vehicle vibration response prediction method based on deep fusion of multi-modal features;
[0030] Figure 3 is a schematic structural diagram of the image feature extractor proposed by the present invention;
[0031] Figure 4 is a schematic structural diagram of the temporal feature extractor proposed by the present invention;
[0032] Figure 5 is the result of predicting four groups of vehicle vibration responses using the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following further clarifies the present invention with reference to the accompanying drawings. The vehicle vibration response prediction method based on deep fusion of multi-modal features provided by the present invention, the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0034] The overall process of the vehicle vibration response prediction method based on deep fusion of multi-modal features provided by the present invention is as Figure 1 shown, and the specific steps are as follows:
[0035] Step 1: Based on vehicle driving road surface image data and vibration response time series signals, construct a vehicle vibration excitation and response triple dataset;
[0036] Step 2: Build a vehicle vibration response prediction model based on the deep fusion of multi-modal features. Use EfficientNet as the image feature extractor and Transformer as the time series feature extractor to extract the features of the road surface image and vibration signal respectively. Combine the bidirectional cross-attention mechanism that follows the Dirichlet distribution for multi-modal feature fusion, and implement signal prediction through a fully connected network;
[0037] Step 3: Divide the triple dataset into a training set, a validation set, and a test set. Use the training set for model training, select the model and tune the hyperparameters through the validation set, and finally evaluate the performance of the model on the test set;
[0038] Step 4: Use the trained vehicle vibration response prediction model to predict the vehicle vibration response signal for the uploaded vehicle driving road surface image data and vibration response time series signal.
[0039] Furthermore, the specific steps of constructing the dataset based on multi-modal data in Step 1 are as follows:
[0040] ① Real-time collect the road surface RGB image within 5 meters ahead through an in-vehicle camera (resolution 1280×720, frame rate 30fps), and synchronously collect the vehicle vibration response time series signal through an acceleration sensor (range ±16g, sampling rate 1000Hz), including physical quantities such as vertical acceleration and speed;
[0041] ② Perform normalization and enhancement processing on the road surface image. Scale the image to 224×224 pixels, normalize the pixel values to the range of [0,1], and standardize according to the mean [0.485,0.456,0.406] and standard deviation [0.229,0.224,0.225] of the ImageNet dataset. Perform image enhancement operations such as random cropping, horizontal flipping, and slight perturbations of brightness / contrast on the image;
[0042] ③ Perform noise addition or time warping operations on the vibration response time series signal to improve the generalization of the model. Then segment the signal with a fixed time window. Each segment is called a frame, which is aligned with the image frame rate. Use a 50ms time window (corresponding to an image frame rate of 30fps). Each frame contains 50 vibration signal points (50 points in 50ms at a sampling rate of 1000Hz), and the length of the historical signal window is 1.5 seconds (30 frames);
[0043] ④Using the current vibration response signal frame as the supervision label, a triplet dataset containing 360,000 groups of samples is constructed. Each sample i is a triplet composed of the current road surface image I(i), the historical vibration signal H(i), and the current vibration signal frame V(i) extracted from the preprocessed data. The historical vibration signal contains the vibration response within the past 1.5 seconds, expressed as H(i) = [V(i - 30), V(i - 29),..., V(i - 1)].
[0044] Further, the specific steps for constructing the vehicle vibration response prediction model in step 2 are as follows:
[0045] ①The image feature extractor uses the EfficientNet network architecture. The input is a three-channel image of 224×224. The first layer of the image encoder uses 16 convolutional kernels of size 3×3 to perform preliminary feature extraction on the image. The stride of this layer is set to 2, directly downsampling the original image to quickly reduce the resolution of the feature map while capturing basic features such as edges and textures in the image. The output feature map size is 112×112×16. Then, through the successive stacking of 11 lightweight convolutional modules (MBConv), a five-stage feature deepening structure of "shallow - middle - high - fusion - compression" is constructed to extract the high-level semantic features of the image. In the shallow feature extraction stage, 1 time of MBConv1 (convolution kernel size = 3, expansion rate = 1) is used to perform 3×3 depthwise separable convolution and use residual connections to retain details. In the middle feature abstraction stage, 2 times of MBConv6 (convolution kernel size = 5, expansion rate = 6) are used to extract refined features of the image through 5×5 depth convolution and SE attention, and max pooling (pooling kernel size = 2, stride = 2) is used for downsampling. In the high-order semantic modeling stage, 2 times of MBConv6 (convolution kernel size = 5, expansion rate = 6) are used to expand the channels by 6 times and then compress them, and dilated convolution (dilation rate = 2) is introduced to expand the receptive field. In the multi-scale fusion stage, 3 times of MBConv6 (convolution kernel size = 5, expansion rate = 6) are used to stack the modules to fuse features of different granularities and use the SE attention module to weight and screen high-frequency vibration-related features. In the global compression stage, 3 times of MBConv6 (convolution kernel size = 5, expansion rate = 6) are used to gradually compress the spatial resolution and expand the channels to 1280. Finally, through global average pooling to aggregate features and a fully connected layer mapping, a 512-dimensional image feature vector is output to comprehensively represent the high-level semantic features of the road surface image;
[0046] ② The temporal feature extractor adopts a Transformer network architecture. The input is the historical vibration time series signal (length = 1500, including 2D acceleration data), which is mapped to a 512-dimensional hidden space through learnable position encoding and an embedding layer. Six stacked encoding modules are used to gradually extract long-term dependency information. Each encoding contains an 8-head self-attention mechanism, dilated causal convolution, and a GeLU gated feed-forward network. When the input vibration signal sequence enters the encoding module, it is simultaneously split into 8 different attention heads. Each head projects the input features into the c vector space through an independent linear transformation matrix, calculates the attention weights between elements at each position through dot product and the Softmax function, and after weighted summation with the value vector, the outputs of each head are concatenated into a 1500×512 feature matrix. A 5×5 convolutional kernel is used to perform one-dimensional dilated causal convolution on the output. The dilation rate increases exponentially by level. The convolution result is split into two parts, features are screened through a gating mechanism, and through residual connection and layer normalization, a feature matrix with a size of 1500×512 is output. The feature matrix is projected into a 2048-dimensional hidden space, and a non-linear factor is introduced through the GeLU activation function. The expanded features are compressed back to 512 dimensions and gate-fused with the original input as the output. Finally, a 512-dimensional temporal feature vector is aggregated through the mean in the time dimension.
[0047] ③ The multi-modal feature fusion layer of the vehicle vibration response prediction model applies a bidirectional cross-attention mechanism that follows the Dirichlet distribution. The input is a 512-dimensional road surface image semantic feature vector and a 512-dimensional vibration signal temporal feature vector. The bidirectional cross-attention interaction is divided into an image-temporal attention flow and a temporal-image attention flow. The former converts the image feature vector into a query matrix through an independent linear transformation, converts the temporal feature into key and value matrices, calculates the correlation score through scaled dot product, and uses the reparameterization trick to sample and generate an attention weight matrix that follows the Dirichlet distribution, and weighted aggregates the key segments. The latter process is symmetric, and an extended attention head is used to set different position biases, and a dynamic masking mechanism is used to mask invalid temporal segments. Subsequently, a gated adaptive fusion module is used to dynamically integrate the bimodal information. By concatenating the 1024-dimensional joint representation output by the bidirectional attention, two-modal dynamic weights are generated through a fully connected layer plus the Sigmoid function, and the two-modal features are weighted and summed according to the weights to output a fusion vector.
[0048] Further, the specific steps for training, validating, and testing the vehicle vibration response prediction model in step three are as follows:
[0049] ① After randomly shuffling the triple dataset, it is divided into a 70% training set, a 15% validation set, and a 15% test set according to a ratio. The training set contains 252,000 groups of samples, the validation set contains 54,000 groups of samples, and the test set contains 54,000 groups of samples.
[0050] ②Train the model using the training set. The model calculates the prediction result through forward propagation, calculates the loss value according to the loss function, and then uses the optimization algorithm to update the model parameters through backpropagation. By iteratively adjusting the model parameters, the loss value is gradually reduced to make the prediction result of the model closer to the true value;
[0051] ③Use the validation set to monitor the model tuning. Calculate the loss value of the model on the validation set. By observing the performance change of the model on the validation set, judge the model fitting situation, adopt the grid search algorithm, traverse all key hyperparameter combinations, screen out the hyperparameter configuration that makes the model performance reach the optimal, and introduce the early stopping mechanism to effectively avoid the model overfitting phenomenon and improve the training efficiency;
[0052] ④Use the test set to evaluate the model performance. Input the current road surface image and historical vibration signal in the test set into the selected model, calculate the prediction result of the model, and compare it with the current vibration signal frame in the test set. Calculate a series of evaluation indexes such as root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R 2 ) and signal peak relative error (PEAK-ERR) to measure the model performance.
[0053] To verify the accuracy of the vehicle vibration response signal prediction of the present invention, four groups of vibration signal prediction experiments were carried out on the present invention, and the experimental results are as Figure 5 shown. As Figure 5 can be seen, the accuracy rate of the vehicle vibration response prediction method established by the present invention for predicting the vehicle vibration response remains above 95%. It can achieve a high accuracy rate on the basis of ensuring stability, and the prediction effect is good. This shows that the vehicle vibration response prediction method established by the present invention is effective, provides a better method for establishing an accurate vibration signal prediction model, and has a certain practicality.
Claims
1. A vehicle vibration response prediction method based on deep fusion of multi-modal features, characterized in that, It includes the following steps: Step 1: Based on the vehicle driving road surface image data and vibration response time series signals, construct a vehicle vibration excitation and response triple dataset. The specific steps are as follows: ① Real-time collect the RGB images of the road surface ahead through an in-vehicle camera, and synchronously collect the vehicle vibration response time series signals through an acceleration sensor, including physical quantities such as vertical acceleration and speed; ② Normalize and enhance the road surface images, scale the images to 224×224 pixels, normalize the pixel values to the range of [0,1], and standardize them according to the mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225] of the ImageNet dataset. Perform image enhancement operations such as randomly cropping, horizontally flipping, and slightly perturbing the brightness / contrast of the images; ③ Perform noise addition or time warping operations on the vibration response time series signals to improve the generalization of the model, and then segment the signals with a fixed time window. Each segment is called a frame, which is aligned with the image frame rate; ④ Construct a dataset with the current vibration response signal frame as the supervision label. Each sample i is a triple composed of the current road surface image I(i), historical vibration signal H(i), and current vibration signal frame V(i) extracted from the preprocessed data. The historical vibration signal is a time series containing vibration responses within a past period of time, expressed as H(i) = [V(i - n), V(i - n + 1),..., V(i - 1)]; Step 2: Based on multi-modal feature deep fusion, construct a vehicle vibration response prediction model. Use EfficientNet as the image feature extractor and Transformer as the time series feature extractor to extract the features of the road surface image and vibration signal respectively. Combine the bidirectional cross-attention mechanism subject to the Dirichlet distribution for multi-modal feature fusion, and realize signal prediction through a fully connected network. The specific steps are as follows: ① The image feature extractor adopts the EfficientNet architecture. Input a three-channel image, and the first layer of the image encoder uses a 3×3 convolutional kernel to downsample the original image and initially extract basic features such as edges and textures; Adopt the stacking of multiple lightweight convolutional modules (MBConv) to construct a five-stage feature deepening structure of "shallow - middle - high - fusion - compression". In the shallow feature extraction stage, use small kernel depth convolution to capture edge textures. In the middle feature abstraction stage, combine the expansion of the convolutional kernel and channel expansion with spatial downsampling to extract local patterns. In the high-order semantic modeling stage, introduce dilated convolution to expand the receptive field and model the global context. In the multi-scale fusion stage, stack modules to integrate features of different granularities and weighted screen specific features. In the global compression stage, stack modules to gradually compress the spatial resolution layer by layer; Finally, after global average pooling and full connection layer processing, output the high-level semantic feature vector of the image; ② The described temporal feature extractor adopts a Transformer network architecture. It takes the historical vibration time series signal as input and maps it to the latent space through learnable position encoding and the embedding layer. It uses multiple stacked encoding modules to extract long-term dependency information. Each encoding module contains a multi-head self-attention mechanism, dilated causal convolution, and a GeLU gated feed-forward network. The vibration signal sequence input to the encoding module is split into multiple attention heads. Each head projects features through a linear transformation, calculates the attention weights, and sums them up after weighting. The outputs of each head are concatenated, and one-dimensional dilated causal convolution with a specific convolution kernel is performed on the output. The dilation rate grows exponentially at each level. After the convolution result is screened by the gating mechanism, residual connection, and layer normalization, it is projected to a high-dimensional latent space. The GeLU activation function is used to introduce non-linearity, and then it is compressed back to the original dimension and fused with the original input through gating. Finally, the temporal feature vector is aggregated by taking the mean along the time dimension. ③ The multi-modal feature fusion layer of the described vehicle vibration response prediction model applies a bidirectional cross-attention mechanism that follows the Dirichlet distribution. It takes the semantic feature vector of the road surface image and the temporal feature vector of the vibration signal as input. The bidirectional cross-attention interaction is divided into an image-temporal attention flow and a temporal-image attention flow. The former transforms the image feature vector into a query matrix through an independent linear transformation, transforms the temporal features into key and value matrices, calculates the correlation scores through scaled dot product, and uses the reparameterization trick to sample and generate an attention weight matrix that follows the Dirichlet distribution, and then weighted aggregates the key segments. The latter process is symmetric to it, and an extended attention head is used to set different position biases and a dynamic masking mechanism to mask invalid temporal segments. Subsequently, through the gating adaptive fusion module, the outputs of the bidirectional attention are concatenated to form a joint representation, and two-modal dynamic weights are generated through a fully connected layer and the Sigmoid activation function. The two-modal features are weighted and summed according to the weights, and the final fusion vector is output. Step 3: Divide the triple dataset into a training set, a validation set, and a test set. Use the training set to train the model, select the model and tune the hyperparameters through the validation set, and finally evaluate the performance of the model on the test set. The specific steps are as follows: ① After randomly shuffling the triple dataset, it is divided into a 70% training set, a 15% validation set, and a 15% test set according to a ratio. ② Use the training set to train the model. The model calculates the prediction result through forward propagation, calculates the loss value according to the loss function, and then uses the optimization algorithm to update the model parameters through backpropagation. By iteratively adjusting the model parameters, the loss value is gradually reduced to make the prediction result of the model closer to the true value. ③ Use the validation set to monitor the model tuning. Calculate the loss value of the model on the validation set. By observing the performance change of the model on the validation set, judge the model fitting situation. Adopt the grid search algorithm to traverse all key hyperparameter combinations, screen out the hyperparameter configuration that makes the model performance reach the optimal, and introduce an early stopping mechanism to effectively avoid the phenomenon of model overfitting and improve the training efficiency. ④Evaluate the model performance using the test set. Input the current road surface image and historical vibration signals in the test set into the selected model, calculate the prediction results of the model, and compare them with the current vibration signal frames in the test set. Calculate a series of evaluation metrics such as root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R 2 ) and peak relative error of the signal (PEAK-ERR) to measure the model performance; Step 4: Use the trained vehicle vibration response prediction model to predict the vehicle vibration response signal for the uploaded vehicle driving road surface image data and vibration response time series signal.
Citation Information
Cited By
Drill string vibration identification and regulation method based on multi-modal data fusion
CN120763676A
Road repair state monitoring method and system based on edge calculation
CN120804632A
A road repair state monitoring method and system based on edge computing
CN120804632B
Real-time vehicle collision prediction method based on multi-modal depth fusion and time sequence modeling
CN121564670A
Self-adaptive network security situation awareness method and device combined with online learning
CN121690666A