Adaptive Driving Behavior Evaluation and Training System Based on Multimodal Large Models
Through the adaptive driving behavior evaluation and training system of multimodal large model, driving behavior is evaluated in real time and personalized training tasks are generated, which solves the problem that existing systems cannot be dynamically adjusted, and improves real-time feedback and efficiency of driving training.
Patent Information
- Application Number
- CN202411542373.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-31
AI Technical Summary
The existing driving training and evaluation systems cannot make dynamic adjustments based on the driving environment and driver behavior in real time, lack real-time feedback and personalized training task generation, cannot optimize driving training in combination with cloud platform technology, and fail to effectively analyze the driver's psychological state.
Adaptive driving behavior evaluation and training system based on multimodal large models is adopted, including multimodal driving data collection and situational perception, data storage, real-time driving behavior evaluation and feedback, adaptive training strategy generation, cloud-based big model-driven self-learning and driver psychological state monitoring and adjustment modules. Through machine learning and deep learning technology, driving behavior can be evaluated in real time, personalized training tasks are generated, and driver psychological state is monitored and adjusted.
Real-time feedback and the generation of personalized training tasks are achieved, the efficiency and safety of driving training are improved, and the training content can be dynamically adjusted according to the driving situation and driver performance, and the driving training process is optimized.
Smart Images

Figure CN119513809B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence for driving training, and specifically refers to an adaptive driving behavior evaluation and training system based on a multimodal large model. Background Art
[0002] In recent years, significant progress has been made in multimodal large model technology. These models can process and understand data from different modalities, such as images, videos, audio, and text. Through multimodal alignment, they can perform various tasks more efficiently, including image classification, aligning text with corresponding videos, and speech detection. The development of these technologies has provided new perspectives and tools for the field of autonomous driving.
[0003] Traditional driving training and evaluation systems usually rely only on human coaches and pre - determined driving courses, and cannot be dynamically adjusted in real - time according to changes in the driving environment and the behavior of the driver.
[0004] The development of multimodal large model technology has brought new opportunities and challenges to the field of driving training, especially in the aspect of adaptive driving behavior evaluation and training. The progress of these technologies makes it possible to build a platform that can understand and predict driving behavior, thereby improving the safety and efficiency of the driving system.
[0005] Existing driving evaluation systems are difficult to provide real - time feedback, cannot quickly and accurately generate personalized training tasks according to complex driving situations, cannot optimize the multimodal large model relied on for driving training in combination with cloud platform technology, and lack the ability to adjust training tasks by analyzing the driver's mental state. Summary of the Invention
[0006] In order to solve the problems in the above - mentioned existing technologies that driving evaluation systems are difficult to provide real - time feedback, cannot quickly and accurately generate personalized training tasks according to complex driving situations, and cannot optimize the multimodal large model relied on for driving training in combination with cloud platform technology, etc., the present invention proposes an adaptive driving behavior evaluation and training system based on a multimodal large model to improve the above - mentioned problems.
[0007] Specifically, this application is as follows:
[0008] An adaptive driving behavior evaluation and training system based on a multimodal large model includes a multimodal driving data collection and situation awareness module, a data storage module, a driving behavior real - time evaluation and feedback module, an adaptive training strategy generation module, a cloud - based large model - driven self - learning module, and a driver mental state monitoring and adjustment module;
[0009] The multimodal driving data collection and situation awareness module is used to collect multimodal data of the vehicle and the driver and sense the driving situation;
[0010] The data storage module is used to store the data collected by the multi-modal driving data acquisition and situation awareness module;
[0011] The real-time driving behavior evaluation and feedback module is used to evaluate the driver's driving behavior in real time through a machine learning evaluation large model and provide instant feedback;
[0012] The adaptive training strategy generation module is used to generate personalized training tasks and learning strategies through a deep learning training large model;
[0013] The cloud large model-driven self-learning module conducts long-term tracking and self-learning for the machine learning evaluation large model and the deep learning training large model based on multi-modal data through the cloud platform;
[0014] The driver's mental state monitoring and adjustment module is used to monitor the driver's physiological and mental states and adjust the driving tasks and course content when necessary.
[0015] Furthermore, the multi-modal driving data acquisition and situation awareness module includes an in-vehicle multi-modal data acquisition sub-module and a driving situation awareness sub-module;
[0016] The in-vehicle multi-modal data acquisition sub-module collects the driver's behavior, the driver's physiological data, the vehicle environment, and the vehicle state in real time through cameras, LiDAR, GPS, seat sensors, and eye tracker sensors;
[0017] The driving situation awareness sub-module perceives the current driving situation based on the data of the in-vehicle multi-modal data acquisition sub-module, and analyzes the driving difficulty in combination with real-time traffic conditions and external information such as weather data. The driving situation at least includes urban roads, rural roads, rainy days, and night driving;
[0018] The data storage module includes a distributed storage unit, a data backup unit, and a data recovery unit;
[0019] The distributed storage unit is used to perform distributed storage on the database, store it on multiple servers, and ensure the security and reliability of the data;
[0020] The data backup unit is used to back up the data regularly to prevent data loss;
[0021] The data recovery unit is used to quickly recover the data when the data is lost or damaged.
[0022] Furthermore, the real-time driving behavior evaluation and feedback module includes a driving behavior evaluation sub-module, a situation adaptability evaluation sub-module, and a real-time feedback and suggestion sub-module;
[0023] The driving behavior evaluation sub-module evaluates the driver's performance in real time through a machine learning evaluation large model according to the driver's operating habits, and identifies potential errors. The operating habits at least include steering, shifting gears, accelerating, and braking;
[0024] The situation adaptability evaluation sub-module evaluates the driver's performance in different driving situations through the machine learning evaluation large model to form an evaluation result. The performance in different driving situations at least includes the reaction speed and operation accuracy in bad weather, night driving, and congested traffic;
[0025] The real-time feedback and suggestion sub-module provides real-time feedback and suggestions during driving according to the evaluation result of the situation adaptability evaluation sub-module. The system providing real-time feedback and suggestions during driving at least includes that when the driver speeds, makes a wrong operation, or reacts slowly, the system corrects and guides through voice or visual prompts.
[0026] Furthermore, the specific implementation process of the machine learning evaluation large model is as follows:
[0027] L1: Obtain the camera image, LiDAR data, map image, and voice data at the current moment collected by the multi-modal driving data collection and situation awareness module;
[0028] L2: Process the LiDAR data to obtain a 2D LiDAR image, process the map image to obtain a 2D grayscale image, and fuse the image features of the camera image, the features of the 2D LiDAR image, and the map features of the 2D grayscale image to obtain the fusion features at the current moment;
[0029] L3: Calculate the Euclidean similarity between the fusion features at the current moment and the fusion features of the previous frame of interest, and determine whether the fusion features at the current moment are the frame of interest according to the Euclidean similarity;
[0030] L4: If the fusion features at the current moment are the frame of interest, use the pre-trained main model to process the fusion features at the current moment to obtain the fine-grained features at the current moment, and then use multiple 3D convolutional kernels to fuse the fine-grained features of the cached frame of interest after time alignment and the fine-grained features at the current moment to obtain the evaluation features at the current moment. The main model uses a ResNet network;
[0031] L5: If the fused feature at the current moment is a non - interested frame, use the pre - trained auxiliary model to process the fused feature at the current moment to obtain the coarse - grained feature at the current moment, perform feature transformation on the coarse - grained feature to obtain the fine - grained feature, and then use multiple two - dimensional convolutional kernels to fuse the fine - grained feature of the cached interested frame after time alignment and the fine - grained feature at the current moment to obtain the evaluation feature at the current moment. The auxiliary model uses the MobileNetV2 network;
[0032] L6: Use natural language technology to extract the speech feature of the current speech data, and use the evaluation network model to fuse - process the evaluation feature and the speech feature at the current moment to obtain the target evaluation result at the current moment.
[0033] Further, the specific implementation of L2 includes:
[0034] Using the conversion relationship between the LiDAR radar coordinate system and the camera imaging coordinate system, project the LiDAR lidar data onto the pixel two - dimensional plane to obtain two - dimensional lidar pixels. The features of the two - dimensional lidar pixels include: x, y, z, r, g, and a; where x, y, z are the three - dimensional coordinates of the pixel center point, r is the reflection coefficient of the LiDAR lidar, g is the point frequency of the LiDAR lidar, and a is the angular resolution of the LiDAR lidar;
[0035] Extract the image features of the camera image. The image features include resolution HW, frame rate F, red component R, green component G, and blue component B;
[0036] Extract the image features of the map image. The map features include vehicle longitude m and vehicle latitude n;
[0037] Then the fused feature at the current moment includes: resolution HW, frame rate F, red component R, green component G, blue component B, x, y, z, the reflection coefficient r of the LiDAR radar, the LiDAR radar point frequency g, the LiDAR radar angular resolution a, vehicle longitude m, and vehicle latitude n;
[0038] In L3, calculating the Euclidean similarity between the fused feature at the current moment and the fused feature of the previous interested frame, and judging whether the fused feature at the current moment is an interested frame according to the Euclidean similarity includes:
[0039] Calculate the Euclidean similarity P between the fused feature at the current moment and the fused feature of the previous interested frame t :
[0040]
[0041] where, H t is the three - dimensional vector after caching the fused feature at the current moment, Hlast_i is a three-dimensional vector after caching the fusion features of the previous interesting frame, and d(H t , H last_i ) is the Euclidean distance between two vectors H t and H last_i ;
[0042] Judge whether the Euclidean similarity P t is greater than a preset threshold. If so, the fusion feature at the current moment is a non - interesting frame. If not, the fusion feature at the current moment is an interesting frame;
[0043] In the L4, multiple three - dimensional convolutional kernels are used to fuse the fine - grained features of the cached interesting frames after time alignment and the fine - grained features at the current moment to obtain the evaluation features at the current moment, including:
[0044] L41: Obtain the fine - grained feature F t of the current moment output by the main model;
[0045] L42: Calculate the correlation matrix K(d1, d2) between the fine - grained feature map at the current moment position and the fine - grained feature map M p2 of the cached interesting frame at the p2 position:
[0046]
[0047] Among them, transform the fine - grained feature map into a three - dimensional matrix, and β2(M d2 ) transform the fine - grained feature map M d2 into a three - dimensional matrix, is a three - dimensional matrix, and the fine - grained feature of the cached interesting frame is the previous interesting fine - grained feature;
[0048] L43: Calculate the cached feature map aligned to the d1 position
[0049]
[0050] Fuse the cached feature map with the fine - grained feature map to obtain the evaluation feature at the d1 position
[0051]
[0052] Among them, β v(.) is a 3*3 convolution operation, concat(.) represents a concatenation operation in the channel dimension, and Q(.) represents a convolution operation with three consecutive three-dimensional convolutional kernels, and the sizes of the three consecutive three-dimensional convolutional kernels are all 3×3;
[0053] L44: Combine all the evaluation features at the d1 position to form the evaluation feature at the current moment
[0054] The calculation method for obtaining the evaluation feature at the current moment in L5 is the same as that for obtaining the evaluation feature at the current moment in L4.
[0055] Furthermore, the adaptive training strategy generation module includes a personalized training task push sub-module and a dynamic training adjustment sub-module;
[0056] The personalized training task push sub-module evaluates the driver's portrait through deep learning and big data mining technologies and constructs a large deep learning training model;
[0057] The dynamic training adjustment sub-module dynamically adjusts the difficulty and content of the training tasks according to the driver's real-time performance, optimizes the large deep learning training model, and adjusts the driver's training content at different learning stages through the output of the large deep learning training model.
[0058] Furthermore, the implementation process of the large deep learning training model is as follows:
[0059] S1: Obtain the driver's historical behavior data and obtain a target training model based on the historical behavior data;
[0060] S2: Real-time collect the driver's target operation data, target physiological data, target vehicle state data, and target environment data through the multi-modal driving data collection and situation awareness module;
[0061] S3: Process the target operation data, the target physiological data, the target vehicle state data, and the target environment data to obtain target integrated data;
[0062] S4: According to the target integrated data, use multi-modal feature fusion and deep learning methods to judge the driver's intention, cognitive state, and visual attention to obtain a target judgment result;
[0063] S5: According to the target judgment result and the target training model, send a target voice reminder to the driver and display a target reminder message on the in-vehicle display;
[0064] S6: Receive the target feedback data of the driver on the target voice reminder and the target reminder information to evaluate the training effect, and obtain target training evaluation data;
[0065] S7: Adjust the target training model according to the target training evaluation data, and the output of the target training model is used to dynamically adjust the difficulty and content of the driver training task.
[0066] Further, the specific implementation process of the target training model in S1 is as follows:
[0067] S11: Construct a driver portrait through the driver's historical behavior data and basic information;
[0068] S12: Based on the driver portrait, construct a target training model through big data mining technology and deep learning technology;
[0069] The output of the target training model is real-time training driving evaluation, personalized course recommendation, and dynamic adjustment of the difficulty and content of the driver training task. The input of the target training model is the driver's historical behavior data and real-time target integration data. The real-time training driving evaluation includes accelerating driving, decelerating driving, steady driving, and lane-changing driving.
[0070] Further, the cloud large model-driven self-learning module includes a learning progress tracking and evaluation sub-module and a cloud self-learning model optimization sub-module;
[0071] The learning progress tracking and evaluation sub-module continuously tracks the driver's learning progress, evaluates their performance, analyzes their progress or bottlenecks, and provides phased summaries and feedback;
[0072] The cloud self-learning model optimization sub-module continuously learns a large amount of driving data through the cloud multi-modal large model to optimize the machine learning evaluation large model and the deep learning training large model;
[0073] The cloud multi-modal large model uses a variational autoencoder (VAE) to continuously learn and optimize the driver behavior model. The VAE realizes the reconstruction and generation of driver behavior through maximizing the likelihood estimation;
[0074] The cloud multi-modal large model uses a federated learning architecture to enhance the self-learning ability of the cloud platform. The local models of different vehicles share the gradient updates to the cloud platform without transmitting the original data.
[0075] Further, the driver mental state monitoring and adjustment module includes a physiological data collection sub-module, an emotion analysis and state adjustment sub-module, and a fatigue driving warning sub-module;
[0076] The physiological data acquisition sub-module collects the physiological state data of the driver in real time through seat sensors, heart rate monitoring sensors, and eye movement tracker sensors, and evaluates their psychological conditions through the physiological state data;
[0077] The emotion analysis and state adjustment sub-module is used to dynamically adjust the difficulty and rhythm of the driving training tasks through a deep emotion analysis model, and adaptively adjust the driver's training tasks through a reinforcement learning algorithm to ensure that the task difficulty matches the driver's emotional state;
[0078] The dynamic adjustment of the difficulty and rhythm of the driving training tasks at least includes that when the driver is in a tense or fatigued state, the system will reduce high-intensity training tasks and provide a more relaxing driving course. The deep emotion analysis model uses the BERT model. The BERT model analyzes the driver's psychological state by extracting physiological data features. The input data of the BERT model is the driver's voice, heart rate, and skin conductance response. The multi-modal analysis formula of the BERT model is:
[0079] Y = softmax(W t .BERT(x t ) + W s .s t + b),
[0080] where x t is the text data converted from the voice, s t is the physiological data corresponding to the heart rate and skin conductance response, W t and W s are the weights corresponding to the text data and physiological data respectively, b is the correction coefficient, and softmax is the activation function;
[0081] The formula for dynamically adjusting the difficulty of the driving training task is:
[0082] D ajusted = f(D current , S fatigue ),
[0083] where D ajusted is the adjusted driving task, and S fatigue is the driver's fatigue level;
[0084] The fatigue driving warning sub-module is used to detect the driver's fatigue and attention. If the fatigue and attention exceed the preset threshold, the system will ensure driving safety through voice prompts and active interventions. The active interventions include autonomous driving takeover.
[0085] The beneficial effects of the adaptive driving behavior evaluation and training system based on the multi-modal large model of the present invention are as follows:
[0086] The present invention collects multi-modal data of the vehicle and the driver through a multi-modal driving data collection and situation awareness module, and perceives the driving situation; a data storage module stores the data collected by the multi-modal driving data collection and situation awareness module; a driving behavior real-time evaluation and feedback module conducts real-time evaluation of the driver's driving behavior through a machine learning evaluation large model and provides immediate feedback; an adaptive training strategy generation module generates personalized training tasks and learning strategies through a deep learning training large model; a cloud large model-driven self-learning module conducts long-term tracking and self-learning of the machine learning evaluation large model and the deep learning training large model based on multi-modal data through a cloud platform; a driver mental state monitoring and adjustment module monitors the physiological and mental states of the driver and adjusts the driving tasks and course content when needed; the present invention combines multi-modal large model technology, collects multi-modal data of the driving vehicle and the driver in real time, generates personalized training tasks and learning strategies according to the changes in the driving situation and the real-time evaluation of the driver's performance, adjusts the training tasks by analyzing the driver's mental state, the system can provide detailed feedback on the driver's behavior, and optimize the driving training process through a cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 It is a system diagram of an adaptive driving behavior evaluation and training system based on a multi-modal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0088] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0089] Embodiment 1
[0090] Figure 1 It is a system diagram of an adaptive driving behavior evaluation and training system based on a multi-modal large model of the present invention.
[0091] The system includes a multi-modal driving data collection and situation awareness module, a data storage module, a driving behavior real-time evaluation and feedback module, an adaptive training strategy generation module, a cloud large model-driven self-learning module, and a driver mental state monitoring and adjustment module;
[0092] The multi-modal driving data collection and situation awareness module is used to collect multi-modal data of the vehicle and the driver and perceive the driving situation;
[0093] The data storage module is used to store the data collected by the multi-modal driving data collection and situation awareness module;
[0094] The real-time driving behavior evaluation and feedback module is used to evaluate the driving behavior of the driver in real time through a machine learning evaluation large model and provide instant feedback;
[0095] The adaptive training strategy generation module is used to generate personalized training tasks and learning strategies through a deep learning training large model;
[0096] The cloud large model-driven self-learning module, through the cloud platform, conducts long-term tracking and self-learning for the machine learning evaluation large model and the deep learning training large model based on multimodal data;
[0097] The driver mental state monitoring and adjustment module is used to monitor the physiological and mental states of the driver and adjust the driving tasks and course content when needed.
[0098] Furthermore, the multimodal driving data collection and situation awareness module includes an in-vehicle multimodal data collection sub-module and a driving situation awareness sub-module;
[0099] The in-vehicle multimodal data collection sub-module, through cameras, LiDAR, GPS, seat sensors, and eye tracker sensors, collects the driver's behavior, the driver's physiological data, the vehicle environment, and the vehicle state in real time;
[0100] The driving situation awareness sub-module, based on the data of the in-vehicle multimodal data collection sub-module, perceives the current driving situation, and analyzes the driving difficulty in combination with real-time traffic conditions, weather data, and other external information. The driving situation at least includes urban roads, rural roads, rainy days, and night driving;
[0101] Specifically, multimodal sensor fusion is a key technology that involves integrating data collected from multiple sensors, including cameras, LiDAR, radar, GPS, and seat sensors, to obtain a global understanding of the driving situation. The fusion methods used include Kalman filtering and Bayesian fusion algorithms;
[0102] The Kalman filter is used to fuse dynamic time series data, such as the data fusion of GPS and inertial sensors, to accurately estimate the trajectory of the vehicle. The formula is:
[0103] x t =x t-1 +K t (z t -Hx t-1 ),
[0104] where, x t is the estimation of the vehicle state at the current moment, K t is the Kalman gain, zt is the sensor observation value, and H is the observation matrix.
[0105] The Bayesian fusion algorithm estimates the prior and posterior probabilities of different sensor data to achieve the fusion of multi-modal data. Its fusion formula is:
[0106]
[0107] Among them, x is the fused driving situation state, Z is the data from each sensor, and P(x|Z) is the posterior probability.
[0108] Situation awareness is achieved through a deep learning model, specifically by processing based on the convolutional neural network CNN and long short-term memory network LSTM for time series data:
[0109] Convolutional neural network CNN: used to process visual data, including cameras and LiDAR, and automatically extract important features in the driving situation, including pedestrians, obstacles, and lane markings;
[0110] Long short-term memory network LSTM: used to process time series data, combine the current driving state with past driving information to judge driving behavior in complex scenarios;
[0111] External data processing of situation awareness:
[0112] Weather and traffic data integration: The system obtains real-time weather data, including temperature, humidity, rainfall, and traffic data, such as congestion and traffic signal status, through an API interface. These data are integrated into the situation analysis model through a decision tree algorithm to optimize the understanding of the driving situation.
[0113] The data storage module includes a distributed storage unit, a data backup unit, and a data recovery unit;
[0114] The distributed storage unit is used for distributed storage of the database, stored on multiple servers to ensure the security and reliability of the data;
[0115] The data backup unit is used to regularly back up the data to prevent data loss;
[0116] The data recovery unit is used to quickly recover the data when the data is lost or damaged.
[0117] Furthermore, the real-time driving behavior evaluation and feedback module includes a driving behavior evaluation sub-module, a situation adaptability evaluation sub-module, and a real-time feedback and suggestion sub-module;
[0118] The driving behavior evaluation sub-module evaluates the driver's performance in real time through a machine learning evaluation large model according to the driver's operating habits, and identifies potential errors. The operating habits at least include steering, shifting gears, accelerating, and braking;
[0119] The situation adaptability evaluation sub-module evaluates the driver's performance in different driving situations through the machine learning evaluation large model, and forms an evaluation result. The performance in different driving situations at least includes the reaction speed and operation accuracy in bad weather, night driving, and congested traffic;
[0120] The real-time feedback and suggestion sub-module provides real-time feedback and suggestions during the driving process according to the evaluation result of the situation adaptability evaluation sub-module. The system providing real-time feedback and suggestions during the driving process at least includes that when the driver exceeds the speed limit, makes a wrong operation, or reacts slowly, the system corrects and guides through voice or visual prompts.
[0121] Specifically, the system gives driving suggestions in real time through a voice assistant or visual prompts on the dashboard, and the feedback form is based on the score curve of driving behavior. The model can determine the direction of driver improvement through a linear regression formula.
[0122] Furthermore, the specific implementation process of the machine learning evaluation large model is as follows:
[0123] L1: Obtain the camera image, LiDAR data, map image, and voice data at the current moment collected by the multi-modal driving data collection and situation awareness module;
[0124] L2: Process the LiDAR data to obtain a two-dimensional LiDAR image, process the map image to obtain a two-dimensional grayscale image, and fuse the image features of the camera image, the features of the two-dimensional LiDAR image, and the map features of the two-dimensional grayscale image to obtain the fusion features at the current moment;
[0125] L3: Calculate the Euclidean similarity between the fusion features at the current moment and the fusion features of the previous interesting frame, and determine whether the fusion features at the current moment are an interesting frame according to the Euclidean similarity;
[0126] L4: If the fusion features at the current moment are an interesting frame, use the pre-trained main model to process the fusion features at the current moment to obtain the fine-grained features at the current moment, and then use multiple three-dimensional convolutional kernels to fuse the fine-grained features of the cached interesting frames after time alignment and the fine-grained features at the current moment to obtain the evaluation features at the current moment. The main model uses a ResNet network;
[0127] L5: If the fused feature at the current moment is a non - interesting frame, use the pre - trained auxiliary model to process the fused feature at the current moment to obtain the coarse - grained feature at the current moment. Perform feature transformation on the coarse - grained feature to obtain the fine - grained feature. Then, use multiple two - dimensional convolutional kernels to fuse the fine - grained feature of the cached interesting frame after time alignment and the fine - grained feature at the current moment to obtain the evaluation feature at the current moment. The auxiliary model uses the MobileNetV2 network;
[0128] L6: Use natural language technology to extract the speech feature of the current speech data, and use the evaluation network model to fuse - process the evaluation feature and the speech feature at the current moment to obtain the target evaluation result at the current moment.
[0129] Further, the specific implementation of L2 includes:
[0130] Using the conversion relationship between the LiDAR radar coordinate system and the camera imaging coordinate system, project the LiDAR radar data onto the pixel two - dimensional plane to obtain two - dimensional LiDAR pixels. The features of the two - dimensional LiDAR pixels include: x, y, z, r, g, and a; where x, y, z are the three - dimensional coordinates of the pixel center point, r is the reflection coefficient of the LiDAR radar, g is the point frequency of the LiDAR radar, and a is the angular resolution of the LiDAR radar;
[0131] Extract the image features of the camera image. The image features include resolution HW, frame rate F, red component R, green component G, and blue component B;
[0132] Extract the image features of the map image. The map features include vehicle longitude m and vehicle latitude n;
[0133] Then the fused feature at the current moment includes: resolution HW, frame rate F, red component R, green component G, blue component B, x, y, z, the reflection coefficient r of the LiDAR radar, the LiDAR radar point frequency g, the LiDAR radar angular resolution a, vehicle longitude m, and vehicle latitude n;
[0134] In L3, calculating the Euclidean similarity between the fused feature at the current moment and the fused feature of the previous interesting frame, and judging whether the fused feature at the current moment is an interesting frame according to the Euclidean similarity includes:
[0135] Calculate the Euclidean similarity P between the fused feature at the current moment and the fused feature of the previous interesting frame t :
[0136]
[0137] where, H t is the three - dimensional vector after caching the fused feature at the current moment, Hlast_i is a three-dimensional vector after caching the fusion features of the previous interesting frame, and d(H t , H last_i ) is the Euclidean distance between two vectors H t and H last_i ;
[0138] Judge whether the Euclidean similarity P t is greater than a preset threshold. If so, the fusion feature at the current moment is a non - interesting frame. If not, the fusion feature at the current moment is an interesting frame;
[0139] In the L4, multiple three - dimensional convolutional kernels are used to fuse the fine - grained features of the cached interesting frames after time alignment and the fine - grained features at the current moment to obtain the evaluation features at the current moment, including:
[0140] L41: Obtain the fine - grained feature F t of the current moment output by the main model;
[0141] L42: Calculate the correlation matrix K(d1, d2) between the fine - grained feature map at the current moment position and the fine - grained feature map M p2 of the cached interesting frame at the p2 position:
[0142]
[0143] Among them, transform the fine - grained feature map into a three - dimensional matrix, and β2(M d2 ) transform the fine - grained feature map M d2 into a three - dimensional matrix, is a three - dimensional matrix, and the fine - grained feature of the cached interesting frame is the previous interesting fine - grained feature;
[0144] L43: Calculate the cached feature map aligned to the d1 position
[0145]
[0146] Fuse the cached feature map with the fine - grained feature map to obtain the evaluation feature at the d1 position
[0147]
[0148] Among them, β v(.) represents a 3*3 convolution operation, concat(.) represents a concatenation operation in the channel dimension, and Q(.) represents a convolution operation with 3 consecutive three-dimensional convolution kernels, and the sizes of the 3 consecutive three-dimensional convolution kernels are all 3×3;
[0149] L44: All the evaluation features at the d1 positions are composed into the evaluation feature at the current moment
[0150] The method for obtaining the evaluation feature at the current moment in L5 is the same as the method for obtaining the evaluation feature at the current moment in L4.
[0151] Specifically, for the evaluation features for evaluating the driver's adaptability in different situations, a unified description is used, and a fuzzy logic control method is used. The fuzzy rule base is defined as follows:
[0152] If "road congestion" is "severe" and "response time" is "slow", then "driving risk" is "high";
[0153] If "weather condition" is "bad" and "operation accuracy" is "low", then "training requirement" is "high";
[0154] The fuzzy inference process finally outputs the driver adaptability evaluation score through the fuzzification, inference, and defuzzification of the input variables.
[0155] Furthermore, the adaptive training strategy generation module includes a personalized training task push sub-module and a dynamic training adjustment sub-module;
[0156] The personalized training task push sub-module evaluates the driver's portrait through deep learning and big data mining technologies and constructs a deep learning training large model;
[0157] The dynamic training adjustment sub-module dynamically adjusts the difficulty and content of the training tasks according to the driver's real-time performance, optimizes the deep learning training large model, and adjusts the training content of the driver at different learning stages through the output of the deep learning training large model.
[0158] Furthermore, the implementation process of the deep learning training large model is as follows:
[0159] S1: Obtain the driver's historical behavior data and obtain the target training model according to the historical behavior data;
[0160] S2: Real-time collect the driver's target operation data, target physiological data, target vehicle state data, and target environment data through the multi-modal driving data collection and situation awareness module;
[0161] S3: Process the target operation data, the target physiological data, the target vehicle state data, and the target environmental data to obtain target integrated data;
[0162] S4: Based on the target integrated data, use multi-modal feature fusion and deep learning methods to judge the driver's intention, cognitive state, and visual attention to obtain a target judgment result;
[0163] S5: According to the target judgment result and the target training model, send a target voice reminder to the driver and display target reminder information on the in-vehicle display;
[0164] S6: Receive the target feedback data of the driver on the target voice reminder and the target reminder information to evaluate the training effect and obtain target training evaluation data;
[0165] S7: Adjust the target training model according to the target training evaluation data, and the output of the target training model is used to dynamically adjust the difficulty and content of the driver training task.
[0166] Furthermore, the specific implementation process of the target training model in S1 is as follows:
[0167] S11: Construct a driver portrait through the driver's historical behavior data and basic information;
[0168] S12: Based on the driver portrait, construct a target training model through big data mining technology and deep learning technology;
[0169] The output of the target training model is real-time training driving evaluation, personalized course recommendation, and dynamic adjustment of the difficulty and content of the driver training task. The input of the target training model is the driver's historical behavior data and real-time target integrated data. The real-time training driving evaluation includes accelerating driving, decelerating driving, steady driving, and lane-changing driving.
[0170] Furthermore, the cloud large model-driven self-learning module includes a learning progress tracking and evaluation sub-module and a cloud self-learning model optimization sub-module;
[0171] The learning progress tracking and evaluation sub-module continuously tracks the driver's learning progress, evaluates their performance, analyzes their progress or bottlenecks, and provides phased summaries and feedback;
[0172] The cloud self-learning model optimization sub-module continuously learns a large amount of driving data through the cloud multi-modal large model to optimize the machine learning evaluation large model and the deep learning training large model;
[0173] The cloud-based multi-modal large model uses a variational autoencoder (VAE) to continuously learn and optimize the driver behavior model. The VAE reconstructs and generates driver behavior by maximizing the likelihood estimate.
[0174] The cloud-based multi-modal large model utilizes a federated learning architecture to enhance the self-learning ability of the cloud platform. Without transmitting raw data, the local models of different vehicles share gradient updates to the cloud platform.
[0175] Specifically, the specific process of gradient update is as follows:
[0176]
[0177] where θ is the global model parameter, and g i is the gradient of the i-th device.
[0178] Furthermore, the driver mental state monitoring and adjustment module includes a physiological data collection sub-module, an emotion analysis and state adjustment sub-module, and a fatigue driving warning sub-module;
[0179] The physiological data collection sub-module real-time collects the physiological state data of the driver through seat sensors, heart rate monitoring sensors, and eye movement tracker sensors, and evaluates the driver's mental condition through the physiological state data;
[0180] The emotion analysis and state adjustment sub-module is used to dynamically adjust the difficulty and rhythm of the driving training task through a deep emotion analysis model, and adaptively adjust the driver's training task through a reinforcement learning algorithm to ensure that the task difficulty matches the driver's emotional state;
[0181] The dynamic adjustment of the difficulty and rhythm of the driving training task at least includes that when the driver is in a tense or fatigued state, the system will reduce high-intensity training tasks and provide a more relaxing driving course. The deep emotion analysis model uses the BERT model. The BERT model analyzes the driver's mental state by extracting physiological data features. The input data of the BERT model is the driver's voice, heart rate, and skin conductance response. The multi-modal analysis formula of the BERT model is:
[0182] Y = softmax(W t .BERT(x t ) + W s .s t + b),
[0183] where x t is the text data converted from voice, s t is the physiological data corresponding to the heart rate and skin conductance response, and W t and W sare the weights corresponding to the text data and physiological data respectively, b is the correction coefficient, and softmax is the activation function;
[0184] The formula for dynamically adjusting the difficulty of the driving training task is:
[0185] D ajusted = f(D current , S fatigue ),
[0186] where D ajusted is the adjusted driving task, and S fatigue is the driver's fatigue level;
[0187] The fatigue driving warning sub-module is used to detect the driver's fatigue and attention. If the fatigue and attention exceed the preset threshold, the system will ensure driving safety through voice prompts and active interventions, and the active interventions include autonomous driving takeover.
[0188] Example 2, Driver Mental State Monitoring and Dynamic Adjustment:
[0189] 1. Background description: A trainee is undergoing advanced driving training, and the system needs to generate personalized training tasks based on the trainee's driving performance and dynamically adjust the course content;
[0190] 2. Operation process:
[0191] The system collects the trainee's driving behaviors (such as steering, braking, accelerating, etc.) through multi-modal sensors, and combines environmental data (such as road conditions, weather) to evaluate their performance in real time.
[0192] The system identifies that the trainee has deficiencies in lane-changing operations, so it pushes a series of lane-changing practice tasks to help the trainee improve lane-changing skills. After the trainee completes the training, the system generates personalized feedback and adjusts the difficulty of subsequent training tasks according to the trainee's progress.
[0193] Example 3, Adaptive Course Generation in Advanced Driving Training:
[0194] 1. Background description: The trainee feels nervous due to bad weather during driving, and the system needs to monitor the trainee's mental state and adjust the training content;
[0195] 2. Operation process:
[0196] The system monitors that the trainee's heart rate has increased through physiological sensors, and combines driving behavior analysis to determine that the trainee is in a tense state;
[0197] The system automatically reduces the intensity of the training task, provides a simpler driving scenario, and encourages the trainee to relax through voice prompts. After the trainee's state returns to normal, the system gradually increases the difficulty of the training task.
[0198] The present invention collects multi-modal data of the vehicle and the driver through a multi-modal driving data collection and situation awareness module, and perceives the driving situation; a data storage module stores the data collected by the multi-modal driving data collection and situation awareness module; a driving behavior real-time evaluation and feedback module evaluates the driving behavior of the driver in real time through a machine learning evaluation large model and provides immediate feedback; an adaptive training strategy generation module generates personalized training tasks and learning strategies through a deep learning training large model; a cloud large model-driven self-learning module performs long-term tracking and self-learning on the machine learning evaluation large model and the deep learning training large model based on multi-modal data through a cloud platform; a driver mental state monitoring and adjustment module monitors the physiological and mental states of the driver and adjusts the driving tasks and course content when necessary; the present invention combines multi-modal large model technology, collects multi-modal data of the driving vehicle and the driver in real time, and generates personalized training tasks and learning strategies according to the changes in the driving situation and the real-time evaluation of the driver's performance. By analyzing the mental state of the driver to adjust the training tasks, the system can provide detailed feedback on the driver's behavior and optimize the driving training process through a cloud platform.
[0199] The above describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual content is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. An adaptive driving behavior evaluation and training system based on a multimodal large model, characterized in that It includes a multi-modal driving data collection and situation awareness module, a data storage module, a real-time driving behavior evaluation and feedback module, an adaptive training strategy generation module, a cloud large model-driven self-learning module, and a driver mental state monitoring and adjustment module; The multi-modal driving data collection and situation awareness module is used to collect multi-modal data of the vehicle and the driver and perceive the driving situation; The data storage module is used to store the multi-modal data; The real-time driving behavior evaluation and feedback module is used to evaluate the driving behavior of the driver in real time through a machine learning evaluation large model and provide instant feedback; The specific implementation process of the machine learning evaluation large model is as follows: L1: Obtain the camera image, LiDAR data, map image, and voice data at the current moment collected by the multi-modal driving data collection and situation awareness module; L2: Process the LiDAR data to obtain a two-dimensional LiDAR image, process the map image to obtain a two-dimensional grayscale image, and fuse the image features of the camera image, the features of the two-dimensional LiDAR image, and the map features of the two-dimensional grayscale image to obtain the fusion features at the current moment; L3: Calculate the Euclidean similarity between the fusion features at the current moment and the fusion features of the previous interesting frame, and determine whether the fusion features at the current moment are an interesting frame according to the Euclidean similarity; L4: If the fusion features at the current moment are an interesting frame, use the pre-trained main model to process the fusion features at the current moment to obtain the fine-grained features at the current moment, and then use multiple three-dimensional convolutional kernels to fuse the fine-grained features of the cached interesting frame after time alignment and the fine-grained features at the current moment to obtain the evaluation features at the current moment. The main model uses a ResNet network; L5: If the fusion features at the current moment are a non-interesting frame, use the pre-trained auxiliary model to process the fusion features at the current moment to obtain the coarse-grained features at the current moment, perform feature transformation on the coarse-grained features to obtain fine-grained features, and then use multiple two-dimensional convolutional kernels to fuse the fine-grained features of the cached interesting frame after time alignment and the fine-grained features at the current moment to obtain the evaluation features at the current moment. The auxiliary model uses a MobileNetV2 network; L6: Use natural language technology to extract the voice features of the current voice data, and use the evaluation network model to fuse and process the evaluation features and voice features at the current moment to obtain the target evaluation result at the current moment; The adaptive training strategy generation module is used to generate personalized training tasks and learning strategies through a deep learning training large model; The cloud large model-driven self-learning module conducts long-term tracking and self-learning on the machine learning evaluation large model and the deep learning training large model based on multi-modal data through a cloud platform; The driver mental state monitoring and adjustment module is used to monitor the physiological and mental states of the driver and adjust the driving tasks and course content when necessary.
2. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 1, wherein, The multi-modal driving data collection and situation awareness module includes an in-vehicle multi-modal data collection sub-module and a driving situation awareness sub-module; The in-vehicle multi-modal data collection sub-module collects the driver's behavior, the driver's physiological data, the vehicle environment, and the vehicle state in real time through cameras, LiDAR radars, GPS, seat sensors, and eye tracker sensors; The driving situation awareness sub-module perceives the current driving situation based on the data of the in-vehicle multi-modal data collection sub-module, and analyzes the driving difficulty in combination with real-time traffic conditions, weather data and other external information. The driving situation at least includes urban roads, rural roads, rainy days, and night driving; The data storage module includes a distributed storage unit, a data backup unit, and a data recovery unit; The distributed storage unit is used for distributed storage of the database, stored on multiple servers to ensure the security and reliability of the data; The data backup unit is used for backing up data regularly to prevent data loss; The data recovery unit is used to quickly recover data when the data is lost or damaged.
3. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 2, wherein The real-time driving behavior evaluation and feedback module includes a driving behavior evaluation sub-module, a situation adaptability evaluation sub-module, and a real-time feedback and suggestion sub-module; The driving behavior evaluation sub-module evaluates the driver's performance in real time through a machine learning evaluation large model according to the driver's operating habits, and identifies potential errors. The operating habits at least include steering, shifting gears, accelerating, and braking; The situation adaptability evaluation sub-module evaluates the driver's performance in different driving situations through the machine learning evaluation large model to form an evaluation result. The performance in different driving situations at least includes the reaction speed and operation accuracy in bad weather, night driving, and congested traffic; The real-time feedback and suggestion sub-module provides real-time feedback and suggestions during the driving process according to the evaluation result of the situation adaptability evaluation sub-module. The system providing real-time feedback and suggestions during the driving process at least includes that when the driver exceeds the speed limit, makes a wrong operation, or reacts slowly, the system corrects and guides through voice or visual prompts.
4. The adaptive driving behavior evaluation and training system based on a multi-modal large model according to claim 3, characterized in that, The specific implementation of L2 includes: Using the conversion relationship between the LiDAR radar coordinate system and the camera imaging coordinate system, project the LiDAR lidar data onto the pixel two-dimensional plane to obtain two-dimensional lidar pixels. The features of the two-dimensional lidar pixels include: x, y, z, r, g, and a; where x, y, z are the three-dimensional coordinates of the pixel center point, r is the reflection coefficient of the LiDAR lidar, g is the point frequency of the LiDAR lidar, and a is the angular resolution of the LiDAR lidar; Extract the image features of the camera image. The image features include resolution HW, frame rate F, red component R, green component G, and blue component B; Extract the image features of the map image. The map features include vehicle longitude m and vehicle latitude n; The fused features at the current moment include: resolution HW, frame rate F, red component R, green component G, blue component B, x, y, z, reflection coefficient r of the LiDAR radar, point frequency g of the LiDAR radar, angular resolution a of the LiDAR radar, vehicle longitude m, and vehicle latitude n; In L3, calculating the Euclidean similarity between the fused features at the current moment and the fused features of the previous interesting frame, and judging whether the fused features at the current moment are an interesting frame according to the Euclidean similarity includes: Calculate the Euclidean similarity P between the fused feature at the current moment and the fused feature of the previous frame of interest t : Among them, H t is a three-dimensional vector after caching the fusion feature at the current moment, and H last_i is a three-dimensional vector after caching the fusion feature of the previous interesting frame. d(H t , H last_i ) is the Euclidean distance between the two vectors H t and H last_i ; Judge the Euclidean similarity P t to determine whether it is greater than a preset threshold. If so, the fused feature at the current moment is a non - interesting frame; if not, the fused feature at the current moment is an interesting frame. In L4, using multiple 3D convolutional kernels to fuse the fine-grained features of the cached interesting frame after time alignment and the fine-grained features at the current moment to obtain the evaluation features at the current moment, including: L41: Obtain the fine-grained feature F at the current moment output by the main model t ; L42: Calculate the current moment The fine-grained feature map at the position and the fine-grained feature map M of the cached frame of interest at the p2 position p2 The correlation matrix K(d1, d2) of: Among them, transform the fine-grained feature map into a three-dimensional matrix, and β2(M d2 ) transform the fine-grained feature map M d2 into a three-dimensional matrix, is a three-dimensional matrix, and cache the fine-grained features of the interested frame as the previous interested fine-grained features; L43: Calculate the cached feature map aligned to the d1 position For the cached feature map and the fine-grained feature map are fused to obtain the evaluation feature at position d1 where β v (.) is a 3*3 convolution operation, concat(.) represents a concatenation operation in the channel dimension, Q(.) represents a convolution operation of three consecutive three-dimensional convolutional kernels, and the sizes of the three consecutive three-dimensional convolutional kernels are all 3×3; L44: Combine the evaluation features at all d1 positions to form the evaluation features at the current moment The method for obtaining the evaluation features at the current moment in L5 is the same as that for obtaining the evaluation features at the current moment in L4.
5. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 4, wherein The adaptive training strategy generation module includes a personalized training task push sub-module and a dynamic training adjustment sub-module; The personalized training task push sub-module evaluates the driver's portrait through deep learning and big data mining technologies and constructs a deep learning training large model; The dynamic training adjustment sub-module dynamically adjusts the difficulty and content of the training task according to the driver's real-time performance, optimizes the deep learning training large model, and adjusts the driver's training content at different learning stages through the output of the deep learning training large model.
6. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 5, wherein The implementation process of the deep learning training large model is as follows: S1: Obtain the driver's historical behavior data and obtain a target training model according to the historical behavior data; S2: Real-time collect the driver's target operation data, target physiological data, target vehicle state data, and target environmental data through the multi-modal driving data collection and context awareness module; S3: Process the target operation data, the target physiological data, the target vehicle state data, and the target environmental data to obtain target integrated data; S4: According to the target integrated data, use multi-modal feature fusion and deep learning methods to judge the driver's intention, cognitive state, and visual attention to obtain a target judgment result; S5: According to the target judgment result and the target training model, send a target voice reminder to the driver and display target reminder information on the in-vehicle display; S6: Receive the target feedback data of the driver on the target voice reminder and the target reminder information to evaluate the training effect and obtain target training evaluation data; S7: Adjust the target training model according to the target training evaluation data, and the output of the target training model is used to dynamically adjust the difficulty and content of the driver training task.
7. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 6, wherein The specific implementation process of the target training model in S1 is as follows: S11: Construct a driver portrait through the driver's historical behavior data and the driver's basic information; S12: Based on the driver portrait, construct a target training model through big data mining technology and deep learning technology; S13: The output of the target training model is real-time training driving evaluation, personalized course recommendation, and dynamic adjustment of the difficulty and content of the driver training tasks. The input of the target training model is the historical behavior data of the driver and the real-time target integration data. The real-time training driving evaluation includes accelerating driving, decelerating driving, steady driving, and lane-changing driving.
8. The adaptive driving behavior evaluation and training system based on a multimodal large model according to claim 7, characterized in that The cloud large model-driven self-learning module includes a learning progress tracking and evaluation sub-module and a cloud self-learning model optimization sub-module; The learning progress tracking and evaluation sub-module continuously tracks the learning progress of the driver, evaluates their performance, analyzes their progress or bottlenecks, and provides phased summaries and feedback; The cloud self-learning model optimization sub-module continuously learns a large amount of driving data through the cloud multi-modal large model to optimize the machine learning evaluation large model and the deep learning training large model; The cloud multi-modal large model uses a variational autoencoder (VAE) to continuously learn and optimize the driver behavior model. The VAE realizes the reconstruction and generation of driver behavior by maximizing the likelihood estimation; The cloud multi-modal large model uses a federated learning architecture to enhance the self-learning ability of the cloud platform. Without transmitting the original data, the local models of different vehicles share the gradient updates to the cloud platform.
9. The adaptive driving behavior evaluation and training system based on a multi-modal large model according to claim 8, wherein, The driver mental state monitoring and adjustment module includes a physiological data collection sub-module, an emotion analysis and state adjustment sub-module, and a fatigue driving warning sub-module; The physiological data collection sub-module real-time collects the physiological state data of the driver through seat sensors, heart rate monitoring sensors, and eye movement tracker sensors, and evaluates their mental condition through the physiological state data; The emotion analysis and state adjustment sub-module is used to dynamically adjust the difficulty and rhythm of the driving training tasks through a deep emotion analysis model, and adaptively adjust the driver's training tasks through a reinforcement learning algorithm to ensure that the task difficulty matches the driver's emotional state; The dynamic adjustment of the difficulty and rhythm of the driving training tasks at least includes that when the driver is in a tense or fatigued state, the system will reduce high-intensity training tasks and provide a more relaxing driving course. The deep emotion analysis model uses a BERT model. The BERT model analyzes the driver's mental state by extracting physiological data features. The input data of the BERT model are the driver's voice, heart rate, and skin conductance response. The multi-modal analysis formula of the BERT model is: Y = softmax(W t .BERT(x t ) + W s .BERT(s t ) + b) where x t is the text data of speech conversion, s t is the physiological data corresponding to heart rate and galvanic skin response, W t and W s are the weights corresponding to the text data and physiological data respectively, b is the correction coefficient, and softmax is the activation function; The formula for realizing the dynamic adjustment of the difficulty of the driving training tasks is: D ajusted = f(D current , S fatigue ) Among them, D ajusted is the adjusted driving task, and S fatigue is the driver's fatigue level; The fatigue driving warning sub-module is used to detect the driver's fatigue and attention. If the fatigue and attention exceed the preset threshold, the system will ensure driving safety through voice prompts and active interventions. The active interventions include autonomous driving takeover.
Citation Information
Patent Citations
Cloud-side collaborative automatic driving method and system applied to auxiliary driving, and medium
CN117184131A
Driving behavior analysis and risk assessment system and method based on environmental information
CN117698741A