Driving simulation training method and system based on multi-mode human factor analysis and computing equipment

By using multimodal human factors analysis and deep learning algorithms, a driving simulation training system was constructed, which solved the problem of lack of real-time analysis and personalized feedback in traditional simulation training. It enabled a comprehensive understanding of the driver's state and personalized feedback, thereby improving driving skills and safety.

CN120910684APending Publication Date: 2025-11-07BEIJING HENGZHI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510983660.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional driving simulation training relies on visual and auditory feedback, lacks real-time analysis and personalized feedback, and cannot effectively improve drivers' driving skills and safety.

Method used

A multimodal human factors analysis method is adopted, which acquires driver physiological and vehicle data through multimodal sensors, and constructs a multimodal pre-trained neural network model by combining deep learning algorithms, and outputs driving suggestions and personalized training plans in real time.

Benefits of technology

It enables a comprehensive understanding of the driver's condition and personalized feedback, improving driving skills and safety, and providing real-time guidance in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910684A_ABST
    Figure CN120910684A_ABST
Patent Text Reader

Abstract

The invention provides a driving simulation training method and system based on multi-modal human factor analysis and computing equipment, and the method comprises the steps: setting different training scenes for driving simulation training, and obtaining multi-modal human factor data through a multi-modal sensor; performing data processing on the acquired multi-mode human factor data and the corresponding driving simulation training data; modeling a multi-modal human factor pre-training neural network classification model; and outputting a training result in real time according to the multi-mode human factor pre-training neural network classification model. According to the technical scheme, the neural network model is trained and predicted, and the real-time driving suggestion is output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of human factors engineering, and particularly relates to a method and system for driving simulation training based on multi-modal human factor analysis and a computing device. BACKGROUND

[0002] Traditional driving simulation training mainly relies on visual and auditory feedback to provide a safe and controllable environment for learning and training driving skills. Although this method can simulate real driving environments and create realistic traffic scenes and road conditions through computer graphics, virtual reality technology and other means, it has some limitations, especially in real-time analysis and personalized feedback.

[0003] Therefore, in order to overcome the above limitations, a technical solution is needed to introduce advanced technology based on multi-modal human factor analysis to improve the effectiveness of driving simulation training. SUMMARY

[0004] The present application aims to provide a method and system for driving simulation training based on multi-modal human factor analysis and a computing device to improve the effectiveness of driving simulation training and output real-time driving recommendations.

[0005] According to an aspect of the present application, a method for driving simulation training based on multi-modal human factor analysis is provided, the method comprising:

[0006] Setting different training scenarios for driving simulation training, and obtaining multi-modal human factor data through multi-modal sensors;

[0007] Data processing of the obtained multi-modal human factor data and corresponding driving simulation training data;

[0008] Constructing a multi-modal human factor pre-trained neural network classification model;

[0009] Real-time output of training results according to the multi-modal human factor pre-trained neural network classification model.

[0010] According to some embodiments, the different training scenarios include:

[0011] The first scenario is used to improve the driver's ability to identify and respond to potential dangers;

[0012] The second scenario is used to test the driver's decision-making ability and attention allocation under pressure;

[0013] The third scenario is used to study the driver's fatigue level;

[0014] The fourth scenario is used to test the driver's ability to maintain focus when disturbed by external factors;

[0015] The fifth scenario is used to evaluate the emotional management and emergency response ability of the driver in a stressful environment.

[0016] According to some embodiments, the acquired multi-modal human factor data and corresponding driving simulation training data are processed, including:

[0017] The multi-modal human factor data and corresponding driving simulation training data are processed using timestamps.

[0018] The acquired multi-modal human factor data and corresponding driving simulation training data are cleaned, standardized and feature extracted by a preprocessing module.

[0019] According to some embodiments, the multi-modal human factor pre-training neural network classification model includes an input layer, a multi-modal fusion layer, a long short-term memory network layer and an output layer, wherein:

[0020] The input layer includes timestamps, human factor physiological data, action data, vehicle motion data and scene data, and processes multi-modal time series data and organizes multi-dimensional vectors through the input layer.

[0021] The multi-modal fusion layer fuses data of different modalities before entering the long short-term memory network layer.

[0022] The long short-term memory network layer processes input sequences in chronological order and captures time-dependent relationships in the data.

[0023] The output layer is used for problem classification and output of result suggestions.

[0024] According to some embodiments, if the classification task is the first type, the output layer outputs driving suggestions.

[0025] If the classification task is the second type, the output layer outputs specific predicted values.

[0026] According to some embodiments, if the classification task is the first type, the output layer is a softmax layer with three nodes, each node representing a driving suggestion.

[0027] According to some embodiments, if the classification task is the second type, the output layer is a single-node linear layer.

[0028] According to some embodiments, it further includes dynamic difficulty adjustment using adaptive scenario adjustment, the system automatically adjusts the difficulty of the training scenario according to the current performance and state of the driver; and / or

[0029] The personalized learning path planning is based on long-term accumulated data to customize a personalized training plan for each driver.

[0030] According to another aspect of the present application, a system for driving simulation training based on multi-modal human factor analysis is provided, the system comprising a data acquisition module, a preprocessing module, a training inference module and a result generation module, wherein,

[0031] The data acquisition module is responsible for collecting multi-modal human factor data and raw data of corresponding driving simulation training data from various sensors and data sources;

[0032] The preprocessing module cleans, standardizes and extracts features from the raw data;

[0033] The training inference module trains the model using the preprocessed multi-modal human factor data and corresponding driving simulation training data, and performs real-time inference;

[0034] The result generation module generates a visual report or real-time feedback to help users understand and improve their driving behavior.

[0035] According to another aspect of the present application, a computing device is provided, comprising:

[0036] a processor; and

[0037] a memory storing a computer program which, when executed by the processor, causes the processor to perform the method of any one of the above.

[0038] According to an embodiment of the present application, a variety of driving training scenarios are constructed to simulate real or extreme driving environments, covering different road conditions and unexpected events, and improving the ability of drivers to cope with complex situations. Multi-modal human factor data is obtained through multi-modal sensors, which can achieve continuous data collection without interference and fully reflect the state of the driver. The multi-modal human factor data and corresponding driving simulation training data are processed to improve data quality, unify the time reference and adapt to the input requirements of deep learning. A multi-modal human factor pre-training neural network classification model is modeled, a neural network model based on long short-term memory network and multi-modal fusion architecture is designed and trained to capture temporal dependency, perform cross-modal information fusion and predict driver state or operation suggestions for advanced driving assistance system development.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows.

[0041] Figure 1A system architecture diagram illustrating a driving simulation based on multi-modal human factors analysis according to example embodiments is shown.

[0042] Figure 2 A method flow diagram illustrating a driving simulation training based on multi-modal human factors analysis according to example embodiments is shown.

[0043] Figure 3 A diagram illustrating pre-training of a neural network classification model according to example embodiments is shown.

[0044] Figure 4 A block diagram of a computing device according to example embodiments is shown. DETAILED DESCRIPTION

[0045] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the several views.

[0046] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the

[0047] The block diagrams in the drawings show only the functional entities and not necessarily the physical separate entities. That is, the functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0048] The flow diagrams in the drawings show only the example sequences of operations and not necessarily the complete sequence of operations and / or steps. For example, some operations / steps can be split into multiple operations / steps, and some operations / steps can be combined or partially combined, and the actual order of execution can be changed from that shown.

[0049] It should be understood that although the terms first, second, third, etc. can be used herein to describe various components, these components should not be limited by these terms. These terms are used only to distinguish one component from another. Thus, a first component discussed below could be termed a second component without departing from the teachings of the present inventive concept. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0050] The user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0051] Those skilled in the art can understand that the drawings are only schematic diagrams of example embodiments, and the modules or flows in the drawings are not necessarily necessary for implementing the present application, and therefore cannot be used to limit the protection scope of the present application.

[0052] Traditional driving simulation training generally refers to using a driving simulator to simulate a real driving environment to provide a safe and controllable environment for learning and training driving skills, which is a basic simulator technology. This training method mainly relies on visual and auditory feedback to simulate the real driving experience, and creates realistic traffic scenes and road conditions through computer graphics, virtual reality technology and other means.

[0053] In most cases, the feedback provided by traditional simulators is based on preset standards rather than specific conditions for each driver, which limits the effectiveness and relevance of the training. And traditional simulators usually do not have the ability to analyze the driver's behavior in real time, which means that the driver's state such as fatigue, distraction or emotional fluctuations cannot be identified immediately, and the training content cannot be adjusted or immediate feedback cannot be provided accordingly.

[0054] With the transformation of various fields to intelligence, simulated driving also needs to be more intelligent to assist drivers in better driving training.

[0055] To this end, the present application proposes a driving simulation training method based on multi-modal human factor analysis, which can make up for the shortcomings of traditional simulation training and open up new ways to improve driving safety and efficiency.

[0056] The method for driving simulation training based on multi-modal human factor analysis integrates multiple sensors, combines advanced data analysis techniques and machine learning algorithms, comprehensively understands the state of the driver, and accordingly provides more accurate, timely and personalized training experience. Compared with the feedback results provided by traditional simulators, there are improvements in concept, data dimension, analysis ability, feedback mechanism and even effect evaluation.

[0057] Before describing the embodiments of the present application, some terms or concepts related to the embodiments of the present application are explained and described.

[0058] Electroencephalography (EEG) is a non-invasive neuroscientific technology for measuring the electrical activity of the cerebral cortex. By placing multiple electrodes on the scalp, the electrical signals generated by the synchronous discharge of brain neuron groups are recorded, reflecting the activity of the brain in different states.

[0059] Perclos (PERCENT CLOSURE) is a physiological measurement index for evaluating the attention state and fatigue level of drivers. The state of the driver is determined by analyzing eye tracking data.

[0060] The example embodiments of the present application will be described below with reference to the accompanying drawings.

[0061] The present application obtains various data of human senses through multi-modal sensors, and analyzes the obtained multi-modal human factor data to better understand the state of the driver, such as attention, cognitive load, fatigue and emotion. Various driving data of the vehicle are also obtained through multi-modal sensors to provide a solid foundation for analyzing the behavior of the driver.

[0062] According to some embodiments, the eye tracking device is used to monitor the gaze position and movement pattern of the driver. By analyzing the fixation pattern and pupil changes, the attention and cognitive load can be inferred. For example, long time fixation on a certain point may be a sign of fatigue or concentration of attention, and frequent eye movement may indicate distraction or tension. In combination with the change frequency and distribution of the fixation point, the attention level of the driver can be evaluated. High-precision cameras are used to capture the changes in pupil diameter, and the changes in pupil size are related to psychological activity. Larger changes usually indicate higher psychological load or emotional excitement, and pupil dilation can be used to measure cognitive load and emotional response.

[0063] Wearing an EEG headset to record brain wave activity in real-time, monitor brain waves (Alpha / Theta) to analyze fatigue, cognitive load; Alpha waves (8-13 Hz) are typically enhanced in a relaxed state, while Theta waves (4-7 Hz) are associated with drowsiness. Calculating the power ratio of Alpha / Theta bands can be used as an indicator of cognitive load and fatigue level. A higher Alpha / Theta ratio may indicate lower cognitive load or better focus; conversely, it may indicate fatigue or higher cognitive load.

[0064] Capture facial expressions with high-definition cameras and identify subtle expression changes through computer vision algorithms. Certain specific micro-expressions (such as frowning, mouth corner drooping) may reveal negative emotions or stress. Recognized emotional signals help assess the driver's mental state, such as tension, anxiety, or anger. Frequent yawning is a clear sign of fatigue. Monitor yawning, and an increase in yawning frequency indicates the need for rest to avoid potential safety risks. Measure the proportion of eye closure through eye tracking technology. Get the percentage of eye closure (PERCLOS), when the PERCLOS value exceeds a certain threshold, indicating that the driver may be in a state of extreme fatigue, and a high PERCLOS value is directly related to the risk of fatigue driving.

[0065] By obtaining driving simulation training data such as steering angle, throttle / brake input, speed, lane position, etc., combined with the driver's performance indicators, important data support is provided for simulation training.

[0066] Figure 1 The architecture composition diagram of the driving simulation based on multi-modal human factor analysis according to an example embodiment is shown.

[0067] Referring to Figure 1 According to an example embodiment, the model of the driving simulation training based on multi-modal human factor analysis includes a data acquisition module, a preprocessing module, a training inference module, and a result generation module. The data acquisition module is responsible for collecting multi-modal human factor data and corresponding raw data of driving simulation training data from multiple sensors and data sources. Real-time capture of the driver's physiological signals, behavior data, and vehicle operation information. Ensure that all data have accurate time stamps and are transmitted and stored through a unified interface.

[0068] The preprocessing module cleans, standardizes, and extracts features from the raw data to make it suitable for subsequent model training. Missing values are filled using interpolation methods, and outliers are detected and processed. For data loss due to device failure or network delay, etc., linear interpolation, nearest neighbor interpolation or other more complex interpolation methods such as spline interpolation can be used to fill in missing values, ensuring the continuity of the time series.

[0069] Outliers can be identified using statistical methods (such as Z-score, IQR, or interquartile range) or machine learning algorithms (such as Isolation Forests, DBSCAN, etc.). Depending on the specific circumstances, outliers can be replaced, deleted, or corrected. For example, abnormal heart rate values ​​can be replaced with the average of several points before and after the abnormal value; for extreme speed mutations, if confirmed to be sensor errors, their removal or correction can be considered.

[0070] The training and inference module trains the model using preprocessed data and performs real-time inference on new data. A suitable deep learning framework (such as PyTorch or TensorFlow) is selected to construct the model structure. The pre-trained neural network classification model in the training and inference module includes an input layer, a multimodal fusion layer, a Long Short-Term Memory (LSTM) layer, and an output layer. See the architecture diagram of the training and inference module. Figure 3 .

[0071] The results generation module generates visual reports or real-time feedback to help users understand and improve their driving behavior. This module includes visualization tools, such as Matplotlib and Seaborn for plotting charts, and Plotly or Dash for creating interactive dashboards. It automatically generates comprehensive reports containing key indicators (such as attention level, fatigue level, and braking force recommendations) based on the model output. Real-time feedback is provided through voice prompts, HUD displays, and other methods to offer immediate guidance.

[0072] The results generation module uses a visualization library to create charts to display training results and trends. It develops scripts or programs to automatically summarize model prediction results and format them into easy-to-understand report documents. It also integrates a real-time feedback mechanism into the driving simulator interface to ensure that drivers can receive personalized improvement suggestions in a timely manner.

[0073] This example implementation demonstrates a driving simulation training model based on multimodal human factors analysis. Through effective collaboration between these modules, it can comprehensively analyze the driver's behavioral patterns and internal state from multiple dimensions, thereby providing scientific and reasonable training programs and personalized feedback, significantly improving driving skills and safety.

[0074] Figure 2 A flowchart illustrating a method for driving simulation training based on multimodal human factors analysis according to an example embodiment is shown.

[0075] In S101, different training scenarios are set up for driving simulation training, and multimodal human factors data are acquired through multimodal sensors.

[0076] According to some embodiments, in order to make the data obtained more comprehensive, targeted training scenarios are set. By setting targeted training scenarios for multi-modal human factor analysis in driving simulation training, the purpose is to induce different states of the driver, such as fatigue, distraction, stress, etc. by creating specific situations, so as to better evaluate and improve the ability of the driver.

[0077] The first scenario, "hazard perception scenario", is set to improve the driver's ability to quickly identify and respond to potential hazards.

[0078] Specific implementation: sudden appearance of pedestrians or animals crossing the road, sudden braking of the vehicle in front, road construction obstacles, etc. Use the technology of dynamically adjusting the difficulty to automatically increase or decrease the frequency and complexity of challenging events according to the driver's performance. Combined with virtual reality (VR) technology to provide realistic visual and auditory feedback, enhance the sense of immersion.

[0079] The second scenario, "complex navigation task", is set to increase the cognitive load and test the driver's decision-making ability and attention allocation under high pressure.

[0080] Specific implementation: create scenarios that include multiple traffic lights, complex intersections, merging and exit selection on highways, etc. Introduce time limits or require the driver to perform other tasks simultaneously (such as answering the phone) to increase task complexity and cognitive burden. Use simulation software to generate diverse urban or rural road layouts to ensure that each practice is different.

[0081] The third scenario, "fatigue induction scenario", is set to study the degree of driver fatigue and its impact under long-time monotonous driving conditions. Specific implementation: simulate long-time highway driving, reduce environmental changes, such as maintaining straight roads, stable weather conditions, etc. Monitor the driver's eye movement patterns (such as PERCLOS), head posture, and physiological indicators (such as heart rate variability HRV, galvanic skin response GSR) to determine the level of fatigue. Adjust the light intensity, play monotonous background music, etc. to further promote the feeling of fatigue.

[0082] The fourth scenario, "distraction scenario", is set to test the driver's ability to maintain focus when disturbed by external factors.

[0083] Specific implementation: insert tasks that require the driver to be distracted during driving, such as adjusting the car entertainment system, checking mobile phone messages, or talking with passengers. Observe how the driver balances the attention allocation between the main driving task and other secondary tasks. Use eye tracking technology and behavior analysis algorithms to detect the length and frequency of the driver's gaze away from the road.

[0084] The fifth scenario, "stress scenario", is set to evaluate the driver's emotional management and emergency response ability under high pressure environment.

[0085] Specific implementation: Set up emergency situations such as approaching accident sites, encountering severe weather conditions (heavy rain, heavy fog), and facing aggressive other road users. Record changes in the driver's heart rate, respiratory rate, skin conductance, and other physiological parameters as indicators of emotional tension. Provide immediate feedback mechanism to help drivers understand their performance under stress and guide them to learn effective coping strategies. Throughout the process, it is important to ensure that all data can be accurately collected, synchronized and analyzed to facilitate subsequent human state modeling and personalized feedback.

[0086] This embodiment helps to exercise the psychological endurance and emotional management ability of the driver through the setting of specific situations. For example, training in high-pressure environment can help drivers learn to remain calm and make correct decisions. By observing the driver's operating habits and response patterns in specific situations, it is possible to improve the design of the vehicle's human-machine interface to be more ergonomic, thereby reducing the driver's operating burden and improving driving comfort and safety. Thus, personalized training programs are generated, based on the analysis of different drivers' behavior patterns, physiological responses, etc., personalized training programs can be tailored for each driver. This not only effectively addresses individual differences, but also targets weaknesses in the driver.

[0087] Specialized data acquisition software and machine learning algorithms are used to process multi-modal data streams from different sensors. Deep learning frameworks such as TensorFlow / Keras and PyTorch are used for time series analysis to capture dynamic patterns in sequential data.

[0088] In S103, the acquired multi-modal human factor data and corresponding driving simulation training data are processed.

[0089] According to some embodiments, the acquired multi-modal human factor data and corresponding driving simulation training data are cleaned, standardized and feature extracted to adapt to subsequent model training. Interpolation methods are used to fill in missing values, detect and process outliers. In the event of data loss due to equipment failure or network delay, linear interpolation, nearest neighbor interpolation or other more complex interpolation methods such as spline interpolation can be used to fill in missing values to ensure the continuity of the time series.

[0090] Anomalies are identified using statistical methods (e.g. Z-score, IQR - Interquartile Range) or machine learning algorithms (e.g. Isolation Forests, DBSCAN, etc.). Depending on the specific case, the outliers are replaced, removed or corrected. For example, an abnormal heart rate value can be replaced by the average of several points before and after it; for extreme speed jumps, if confirmed as a sensor error, it can be considered to be removed or corrected.

[0091] In S105, modeling of the multi-modal human-in-the-loop pre-trained neural network classification model is performed.

[0092] According to some embodiments, referring to Figure 3 , the input layer of the multi-modal human-in-the-loop pre-trained neural network classification model is organized as a multi-dimensional vector, organizing data from different sensors such as heart rate, braking force, speed, acceleration, etc. into a multi-dimensional vector according to a fixed time window. For categorical data (such as scene data), an embedding layer is used to convert it to a numerical representation.

[0093] The design of the input layer aims to handle multi-modal time series data from different sensors and information sources, and each type of data has its specific role and purpose. The input layer includes timestamps, human physiological data, motion data, vehicle motion data, and scene data.

[0094] The timestamp (T) represents the current time step and can be used as an index for the sequence. The data type of the timestamp is numerical (integer or floating point), representing a specific time point or relative time interval. The timestamp is used to align data from different sources, ensuring that all input data is on the same time reference.

[0095] Human physiological data (N) reflects the physical state and mental load of the driver, helping to assess attention level, fatigue level, etc. It includes eye tracking data, heart rate, brain waves, etc. The data type of human physiological data is divided into numerical type (such as heart rate, eye movement coordinates) and categorical type (such as classification of different frequency bands of brain waves). For non-numerical data (such as certain brain wave features), an embedding layer is needed to convert it to a numerical representation.

[0096] Motion data (M) is used to describe the driver's operation behavior, which is used to analyze driving habits and decision-making processes.

[0097] Motion data includes steering wheel angle, accelerator pedal position, brake pedal position, etc. It is mainly numerical, reflecting the specific size or proportion of physical quantities. The motion data needs to be properly normalized to ensure that data of different scales can be effectively utilized in the same model.

[0098] Vehicle motion data (P) provides information about the dynamic characteristics of the vehicle, which is crucial for understanding the driving state of the vehicle. Vehicle motion data includes vehicle speed, acceleration, position, attitude, etc. The data type is numerical, including scalars (such as speed) and vectors (such as acceleration, position). The vehicle motion data needs to be synchronized with other types of data in order to accurately reflect the interaction between the vehicle and the driver.

[0099] Scene data (Q) is used to describe external environmental conditions, helping to understand how the driver responds to different road conditions. For example, road type, traffic signal status, lane line information, etc. The data type of scene data is categorical (such as road type), Boolean (such as traffic light status), and numerical (such as lane line distance). For categorical data, One-Hot encoding or embedding layers need to be applied to convert to numerical form.

[0100] According to some embodiments, the multi-modal fusion selects a fully connected layer, which preliminarily fuses the data of different modalities before entering the LSTM layer. Multi-modal fusion can also use an attention mechanism to dynamically weight different input features and emphasize more important information.

[0101] Temporal modeling is performed, and the LSTM layer processes the input sequence in time order, capturing the temporal dependence in the data. Multiple LSTM layers can be stacked as needed to increase the expressive power of the model.

[0102] The output layer performs question classification and outputs the result suggestion. If the classification task is to predict driving suggestions, such as increasing / decreasing / unchanging brake force, the output layer can be a softmax layer with three nodes, each representing a possible driving suggestion. If the task is to predict the specific brake force value, the output layer can be a single-node linear layer that directly outputs the predicted value.

[0103] If the output is a classification problem (such as increasing / decreasing / unchanging brake force), cross-entropy loss is used. Cross-entropy loss (Cross-Entropy Loss) with a softmax activation function is suitable for multi-class classification tasks.

[0104] If the output is a regression problem (such as a specific brake force value), mean squared error (MSE) or smooth L1 loss can be selected, where smooth L1 loss is more robust to outliers. According to some embodiments, Adam or AdamW optimizers can also be selected for adaptive learning rate optimization, and AdamW improves the weight decay mechanism. Or use Dropout to prevent overfitting, especially suitable for deep networks, and perform data augmentation on the data, such as adding noise (such as Gaussian noise) in the time series to improve the robustness of the model.

[0105] According to some embodiments, hyperparameter tuning is performed, the size of the time window affects the ability of the model to capture temporal dependencies, and the trade-off between long-term dependencies and computational cost needs to be balanced. The number of attention heads is adjusted to balance computational complexity and information processing capacity. The number of LSTM hidden units determines the memory capacity, and too many can lead to overfitting. Grid search or random search is used to optimize the hyperparameters using the validation set. Methods such as grid search, random search, or Bayesian optimization can be used to explore the best combination of hyperparameters.

[0106] Checkpoint saving is set to save the model state periodically to facilitate recovery to the best performance moment. Early stopping mechanism is set: monitor the performance of the validation set, and stop training once it no longer improves to prevent overfitting. An independent test set is used to evaluate the performance of the model to ensure its good generalization ability. Training with half-precision floating-point numbers (FP16) can significantly reduce memory usage and speed up the training process while maintaining model performance.

[0107] For large-scale data sets or complex models, distributed training can be used to speed up the training process. Combining the prediction results of multiple models can further improve the stability and accuracy of the model. Applying explainability tools and techniques such as LIME, SHAP to understand the model decision-making process is very important for trust building and model improvement.

[0108] By comprehensively using the above strategies, an efficient and stable driving simulation training system can be built, which not only provides accurate behavior analysis and feedback, but also customizes personalized training plans according to the specific situation of the driver, effectively improving driving skills and safety.

[0109] Model training and optimization are performed on the pre-trained neural network classification model to achieve adaptive scenario adjustment and personalized learning path planning.

[0110] According to some embodiments, dynamic difficulty adjustment is performed for adaptive scenario adjustment, and the system automatically adjusts the difficulty of the training scenario according to the current performance and state of the driver. When the driver shows low cognitive load and good performance, the traffic density is increased or more complex road challenges are introduced. If the driver is found to be in a highly stressed or fatigued state, the task is simplified or external interference factors are reduced to reduce the burden.

[0111] Personalized learning path planning is based on long-term accumulated data to customize personalized training plans for each driver. For example, if a driver frequently has problems with left turn operations, the system can schedule more exercises on this skill and gradually increase the complexity until mastery.

[0112] In S107, a training result is output in real time according to the multi-modal human factor pre-trained neural network classification model.

[0113] According to some embodiments, based on the pre-trained neural network classification model output real-time driving suggestions, if the classification task is the first type, the output result suggestion is a driving suggestion. For the classification task, the purpose is to predict the brake force adjustment direction (increase, decrease, or remain unchanged) that the driver should take. Based on the analysis of the multi-dimensional vector of the input layer, such as the input heart rate: 80 beats / min, brake force: 50%, speed: 60 km / h, acceleration: -2 m / s 2 , scene: urban road, distance to target: 10 meters. The processing method uses the Softmax function to generate the probability of each class (increase, decrease, and unchanged) at the output layer. For example, the model may output a probability of "brake force increase" of 70%, so according to the current data, it is most likely that the brake force needs to be increased to ensure safety. The applicable application scenario is when the system detects potential danger (such as sudden obstacles in front), it can give timely suggestions to help the driver respond.

[0114] If the classification task is the second type, the output result suggestion is a specific predicted value. For the regression task, the purpose is to directly predict the specific brake force value, rather than just providing a directional suggestion. Based on the analysis of the multi-dimensional vector of the input layer, the model directly outputs a continuous value representing the suggested brake force percentage. For example, the model may suggest adjusting the brake force to 60%. This method provides more precise operation guidance and is suitable for scenarios that require fine control. The applicable application scenario is that in complex traffic environments, providing more accurate brake force suggestions helps to optimize safety and smoothness during driving.

[0115] To ensure that the model inference speed is fast enough to guarantee real-time regulation, strategies such as model pruning, quantization, hardware acceleration, optimizing model architecture, and caching mechanisms can be used. Model pruning simplifies the model structure by removing weight connections that contribute less to the final output, thereby speeding up inference. Quantization converts model parameters from floating-point numbers to low-precision values (such as integers), which not only reduces computational complexity but also reduces storage requirements. Hardware acceleration uses specialized hardware such as GPUs or TPUs for acceleration, which are optimized for matrix operations and are very suitable for deep learning model inference. Optimizing model architecture selects lightweight model architectures that are more suitable for real-time applications, or designs specific network structures to reduce latency, such as reducing the number of LSTM layers or nodes. The caching mechanism can quickly respond to similar inputs that repeatedly appear by caching previous calculation results, avoiding repeated calculations.

[0116] Through the above method, the response speed of the model can be significantly improved under the premise of ensuring the accuracy of the model, so that it better adapts to the needs of real-time driving recommendations. Such a system not only improves the safety and operational efficiency of drivers, but also provides strong support for the further development of intelligent transportation systems.

[0117] Based on the driving simulation training result output of the multi-modal human factor analysis, all the data collected during the training process are processed and analyzed in detail, and a detailed training report is automatically generated by using the data analysis results.

[0118] The training report includes but is not limited to behavior data such as steering angle, speed change, lane deviation, etc.; physiological data such as heart rate variability (HRV), galvanic skin response (GSR) reflecting emotional and stress levels; eye movement data such as gaze position, pupil size change reflecting attention concentration; facial expression data, recognized expression changes for evaluating emotional state; environmental data, weather conditions, traffic density and other external factors.

[0119] The training report also supports the generation of visual playback. By developing a special software tool or using existing advanced visualization platforms (such as Unity, Unreal Engine), an interactive visual playback system can be created. This system allows users to replay the entire driving process, reproducing the driving path and events in first-person perspective or bird's eye view. Overlay key indicators on the video to display important physiological indicators (such as heart rate), eye movement trajectories, vehicle operation information, etc. in real time, helping to understand the behavior background at a specific moment. Use heat maps to represent the areas that the driver pays most attention to, and assist in analyzing visual search strategies.

[0120] Using the data analysis results to automatically generate a detailed training report, showing the task completion status, including whether the predetermined goal (such as reaching the destination without accidents) is successfully completed, and the number of any violations. Perform safety assessment based on near-miss frequency, emergency braking frequency and other indicators to evaluate overall safety. Analyze eye movement data and other related parameters for attention management, and give scores on attention distribution and transfer efficiency. According to the facial expression recognition and the change of physiological signals, describe the emotional fluctuations and their influence on driving performance.

[0121] Based on the above analysis, specific personalized suggestions are made for improvement, such as increasing practice intensity in problem areas that frequently occur, or learning how to better manage emotions.

[0122] The present application adopts data synchronization technology to ensure that all data from different sensors can be accurately synchronized and recorded, and a unified timestamp mechanism is adopted. Advanced machine learning algorithms are used for data mining and pattern recognition to extract valuable information from massive data.

[0123] By combining the above tools and techniques, a high-efficiency and flexible system is constructed to realize driving simulation training based on multi-modal human factor analysis, and real-time driving suggestions are output. The key is to select a suitable training neural network classification model to ensure that it can quickly respond and give accurate suggestions; at the same time, appropriate regularization techniques and optimization strategies are used to ensure the real-time performance of the model. In addition, the use of visualization tools can help better understand the data and the performance of the output model, thereby further optimizing the overall performance of the system.

[0124] Figure 4 A block diagram of a computing device according to an example embodiment of the application is shown.

[0125] As shown in Figure 4 , the computing device 30 includes a processor 12 and a memory 14. The computing device 30 can also include a bus 22, a network interface 16, and an I / O interface 18. The processor 12, the memory 14, the network interface 16, and the I / O interface 18 can communicate with each other through the bus 22.

[0126] The processor 12 can include one or more general-purpose CPUs (Central Processing Units), microprocessors, or application-specific integrated circuits, etc., for executing related program instructions.

[0127] The memory 14 can include a machine system readable medium in the form of volatile memory, such as random access memory (RAM), read-only memory (ROM), and / or cache memory. The memory 14 is used to store one or more programs containing instructions and data. The processor 12 can read the instructions stored in the memory 14 to execute the above-mentioned method according to the embodiments of the application.

[0128] The computing device 30 can also communicate with one or more networks through the network interface 16. The network interface 16 can be a wireless network interface.

[0129] The bus 22 can include an address bus, a data bus, a control bus, etc. The bus 22 provides a channel for exchanging information between the components.

[0130] It should be noted that in the specific implementation process, the computing device 30 can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above-mentioned device can also only contain the components necessary to implement the embodiments of the present application, and does not necessarily contain all the components shown in the figure.

[0131] The present application also provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the above method. The computer readable storage medium can include, but is not limited to, any type of disk including floppy disks, optical disks, DVDs, CD-ROMs, micro-drives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), network attached storage devices, cloud storage devices, or any type of medium or device suitable for storing instructions and / or data.

[0132] The embodiments of the present application also provide a computer program product, which comprises a non-transitory computer readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments.

[0133] Those skilled in the art can clearly understand that the technical solutions of the present application can be implemented by means of software and / or hardware. The "units" and "modules" in the specification refer to software and / or hardware that can independently complete or cooperate with other components to complete a specific function, and the hardware can be, for example, a field programmable gate array, an integrated circuit, etc.

[0134] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0135] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0136] In the several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of units is only a logical function division. There can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some services interfaces, devices or units, and can be electrical or other forms.

[0137] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0138] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0139] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable memory. Based on this understanding, the technical scheme of the present application or the part of the present application which contributes to the prior art in essence or the whole or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application.

[0140] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0141] The exemplary embodiments of the present application are specifically shown and described above. It should be understood that the present application is not limited to the detailed structure, arrangement or implementation method described herein; on the contrary, the present application is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the appended clauses.

Claims

1. A method of driving simulation training based on multi-modal human factors analysis, characterized in that, The method comprises the following steps: different training scenarios are set for driving simulation training, and multi-modal human factor data is obtained through multi-modal sensors; the obtained multi-modal human factor data and corresponding driving simulation training data are processed; a multi-modal human factor pre-training neural network classification model is constructed; and real-time training results are output according to the multi-modal human factor pre-training neural network classification model.

2. The method of claim 1, wherein, The different training scenarios include: a first scenario for improving the driver's ability to identify and respond to potential dangers; a second scenario for testing the driver's decision-making ability and attention allocation under pressure; a third scenario for studying the driver's fatigue level; a fourth scenario for testing the driver's ability to maintain focus when disturbed by external factors; a fifth scenario for evaluating the driver's emotional management and emergency response ability under pressure.

3. The method of claim 1, wherein, The multi-modal human factor data and corresponding driving simulation training data are processed, including: the multi-modal human factor data and corresponding driving simulation training data are processed using timestamps; the obtained multi-modal human factor data and corresponding driving simulation training data are cleaned, standardized and feature extracted through a preprocessing module.

4. The method of claim 1, wherein, The multi-modal human factor pre-training neural network classification model includes an input layer, a multi-modal fusion layer, a long short-term memory network layer and an output layer, wherein: the input layer includes timestamps, human physiological data, action data, vehicle motion data and scene data, and processes multi-modal time series data through the input layer for multi-dimensional vector organization; the multi-modal fusion layer fuses data of different modalities before entering the long short-term memory network layer; the long short-term memory network layer processes input sequences in chronological order and captures temporal dependencies in the data; the output layer is used for problem classification and output of result suggestions.

5. The method of claim 4, wherein: if the classification task is of a first type, the output layer outputs driving suggestions; if the classification task is of a second type, the output layer outputs specific prediction values.

6. The method of claim 5, wherein, If the classification task is of the first type, the output layer is a softmax layer with three nodes, each node representing a driving suggestion.

7. The method of claim 5, wherein, If the classification task is of the second type, the output layer is a linear layer with a single node.

8. The method of claim 1, wherein, Further comprising: dynamic difficulty adjustment using adaptive scenario adjustment, the system automatically adjusts the difficulty of the training scenario based on the current performance and state of the driver; and / or personalized learning path planning based on long-term accumulated data to customize personalized training plans for each driver.

9. A system for driving simulation training based on multi-modal human factors analysis, characterized by, The system includes a data acquisition module, a preprocessing module, a training inference module and a result generation module, wherein: the data acquisition module is responsible for collecting raw data of multi-modal human factor data and corresponding driving simulation training data from various sensors and data sources; the preprocessing module cleans, standardizes and extracts features from the raw data; the training inference module trains the model using the preprocessed multi-modal human factor data and corresponding driving simulation training data, and performs real-time inference; and the result generation module generates the final training results. The result generation module model generates a visual report or real-time feedback to help the user understand and improve driving behavior.

10. A computing device, comprising: Comprise: a processor; and a memory storing a computer program which, when executed by the processor, causes the processor to perform the method of any one of claims 1-8.

Citation Information

Patent Citations

  • High-speed rail driving simulation system and method

    CN118038735A

  • Driving skill level classification method based on multi-dimensional data fusion

    CN119293596A

  • Driving skill training method based on diversified virtual examination room construction

    CN119580561A

  • Enhanced closed scene capability evaluation method and system applied to intelligent driving training

    CN120125077A

  • Driving state monitoring and feedback method and system based on multi-modal human factors intelligent data analysis, and edge computing terminal device

    WO2025118936A1