Simulation training difficulty dynamic constraint method and system based on emotion recognition
Through the dynamic constraint method of simulation training difficulty based on emotion recognition, combined with physiological signals, voice signals and expression image information, the difficulty of driving simulation training is dynamically adjusted, which solves the problem that existing systems are difficult to adjust training difficulty according to user needs, and achieves a more personalized and efficient training effect.
Patent Information
- Application Number
- CN202510037504.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-09
AI Technical Summary
The existing driving simulation training system cannot flexibly adjust the training difficulty according to the user's skill level, learning progress and actual needs, resulting in poor training results and experience.
The dynamic constraint method of simulation training difficulty based on emotion recognition is adopted, and physiological signals, speech signals and expression image information are collected by wearing sensors, combined with the training difficulty level calibration database and emotion recognition model, and dynamically adjust the training difficulty mode.
It realizes dynamic adjustment of training difficulty according to the user's real-time emotions and physiological state, improves the personalization and adaptability of simulation training, and improves the training effect and experience.
Smart Images

Figure CN120014907A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a method and system for dynamically constraining the difficulty of simulation training based on emotion recognition. Background Art
[0002] Existing driving simulation training systems have been widely used in driver training, traffic safety education, military training and other fields. Most current simulation training systems use fixed training difficulty and preset task scenarios, and lack consideration for the driver's personalized needs and skill level. The training difficulty of existing systems is usually adjusted based on a unified standard, rather than flexibly adapting to the actual needs, progress and learning situation of each user. The fixed mode may cause some drivers to encounter tasks that are too simple or too complex during training, affecting the learning effect and training experience. Therefore, how to flexibly adjust the training difficulty according to the driver's skill level, learning progress and actual needs has become a key issue in improving the personalization and effectiveness of simulation training.
[0003] In the current relevant technologies, there is a technical problem that the difficulty of driving simulation training cannot be adaptively adjusted according to user needs, resulting in a poor degree of individualization of simulation training. Summary of the invention
[0004] The present application provides a driving skills training method based on a diversified virtual examination room, which is used to solve the technical problem that the difficulty of existing driving simulation training cannot be adaptively adjusted according to user needs, resulting in a poor degree of individualization of simulation training.
[0005] This application provides a dynamic constraint method for simulation training difficulty based on emotion recognition, including:
[0006] When the target user enters the driving simulation training area and starts training, the target user's physiological signal timing information and the target user's voice signal timing information are collected through wearable sensors, and the target user's facial expression image timing information is collected through the simulation training area camera; a training difficulty level calibration database is obtained; based on the training difficulty level calibration database, the training difficulty is calibrated for the target user's physiological signal timing information, the target user's voice signal timing information and the target user's facial expression image timing information to obtain a first matching training difficulty; when the first matching training difficulty is empty or not unique, the target user's physiological signal timing information, the target user's voice signal timing information and the target user's facial expression image timing information are processed through an emotion recognition model to obtain a second matching training difficulty; and the simulation training difficulty mode is adjusted according to the second matching training difficulty.
[0007] This application provides a dynamic constraint system for simulation training difficulty based on emotion recognition, including:
[0008] An expression image timing information acquisition module, which is used to collect the target user's physiological signal timing information and the target user's voice signal timing information through wearable sensors when the target user enters the driving simulation training area and starts training, and collect the target user's expression image timing information through the simulation training area camera; a calibration database acquisition module, which is used to obtain a training difficulty level calibration database; a training difficulty calibration module, which is used to calibrate the training difficulty of the target user's physiological signal timing information, the target user's voice signal timing information and the target user's expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty; a second matching training difficulty acquisition module, which is used to process the target user's physiological signal timing information, the target user's voice signal timing information and the target user's expression image timing information through an emotion recognition model when the first matching training difficulty is empty or not unique, to obtain a second matching training difficulty; a difficulty mode adjustment module, which is used to adjust the simulation training difficulty mode according to the second matching training difficulty.
[0009] The proposed method and system for dynamic constraint of simulation training difficulty based on emotion recognition in this application is to first collect physiological signals, voice signals and facial expression image information of the target user through wearable sensors when the target user enters the driving simulation training area, and calibrate the training difficulty of these signals based on the training difficulty level calibration database to obtain the first matching training difficulty. If the first matching training difficulty is empty or not unique, these signals are processed through the emotion recognition model to obtain the second matching training difficulty, and the difficulty mode of the simulation training is adjusted accordingly, achieving the technical effect of improving the individualization and adaptive adjustment ability of driving simulation training. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 A schematic diagram of a flow chart of a method for dynamically constraining the difficulty of simulation training based on emotion recognition provided in an embodiment of the present application;
[0012] Figure 2 A schematic diagram of the structure of a dynamic constraint system for simulation training difficulty based on emotion recognition provided in an embodiment of the present application.
[0013] Explanation of the accompanying reference numerals: expression image time sequence information acquisition module 10, calibration database acquisition module 20, training difficulty calibration module 30, second matching training difficulty acquisition module 40, difficulty mode adjustment module 50. DETAILED DESCRIPTION
[0014] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0015] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0016] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict, and the terms "first\second" involved are merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "including" and "having" and any variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by technicians in the technical field of this application. The terms used herein are for the purpose of describing the embodiments of the present application only.
[0017] The present application embodiment provides a method for dynamically constraining the difficulty of simulation training based on emotion recognition, such as Figure 1 As shown, the method includes:
[0018] Step S100, when the target user enters the driving simulation training area and starts training, the target user's physiological signal timing information and the target user's voice signal timing information are collected through the wearable sensor, and the target user's facial expression image timing information is collected through the simulated training area camera. Specifically, when the target user enters the driving simulation training area and starts training, the data acquisition system enters the working state. The wearable sensor pre-placed in a specific part uses biosensor technology to fit the skin tightly and collect physiological signal timing information including heart rate, blood pressure, skin conductivity, etc. The signal will fluctuate due to driving scenes and emotional changes. For example, when nervous, the heart rate rises, blood pressure rises, and skin conductivity increases, and the whole process is recorded in the form of a curve. At the same time, the built-in microphone of the sensor converts the voice signal generated by the user due to driving conditions and emotions into a digital signal and records it in time sequence. Its acoustic characteristics such as tone and speech speed can reflect emotions and psychological activities. The camera in the simulated training area captures the user's facial expressions at a certain frame rate. The changes in facial expressions, from nervousness at the beginning of training, to fear when encountering unexpected situations, to joy when completing the task, all form facial expression image timing information in chronological order, providing a comprehensive and solid data foundation for evaluating training status, calibrating training difficulty, etc.
[0019] Step S200, obtain the training difficulty level calibration database. Specifically, when obtaining the training difficulty level calibration database, first determine its data source, integrate multi-channel information, such as past driving training history records covering various road conditions and weather simulations, and record the basic information of the target user, the physiological and voice signal timing information collected by the wearable sensor, the expression image timing information collected by the camera, and the training difficulty and results. After collection, pre-process the data, clean the abnormal points of the physiological signal and standardize it, identify the speech semantics and quantify the acoustic features, identify the expression image category and count the frequency and duration, and then organize it according to the logic, and classify it by hierarchical indexing with task number and user number. Design the database architecture, determine the structure and fields of each table, such as the training task table, user information table, etc., and finally accurately enter the data according to the architecture to ensure completeness and consistency, and build a database for subsequent training difficulty calibration.
[0020] In a possible implementation, a training difficulty level calibration database is obtained, and step S200 further includes step S210, obtaining a simulation training task number and a simulation training difficulty mode number. Specifically, according to the currently ongoing simulation driving training task, the corresponding simulation training task number is obtained. The number is an identifier of a specific simulation driving training scene and process, and can uniquely identify the task among many different training tasks. For example, different driving route planning, traffic scene settings or training target settings will correspond to different task numbers. At the same time, the simulation training difficulty mode number is obtained, which indicates the difficulty level mode used in the current training, such as the primary difficulty mode number may correspond to a relatively simple road condition, fewer traffic rules restrictions and lower driving operation requirements; the intermediate difficulty mode number will involve a more complex road condition combination, more traffic rules constraints and a moderate degree of driving operation complexity; the advanced difficulty mode number means close to real and extremely challenging road conditions, strict traffic rules and difficult driving operations, such as highway driving simulation under bad weather conditions.
[0021] Step S220, the second matching training difficulty, the simulation training task number, the simulation training difficulty mode number, the target user physiological signal timing information, the target user voice signal timing information and the target user facial expression image timing information are associated and stored in the training difficulty level calibration database. Specifically, after obtaining the simulation training task number and the simulation training difficulty mode number, the second matching training difficulty and the number, as well as the target user physiological signal timing information, the target user voice signal timing information and the target user facial expression image timing information are associated and stored in the training difficulty level calibration database. During the storage process, the simulation training task number and the simulation training difficulty mode number are used as key indexes to integrate all relevant data under the same task number and difficulty mode number. For example, for a group of data with a specific task number of "T001" and a difficulty mode number of "D02", the physiological signal time series information generated by the target user during the training process under this task, such as the curve data of heart rate changes over time, blood pressure fluctuation data, skin conductivity change sequence, etc., voice signal time series information, such as voice intonation changes, speech speed and specific words said at different driving stages, facial expression image time series information, such as the panic expression when encountering an emergency, the relaxed expression when driving smoothly, etc., and the second matching training difficulty determined by the emotion recognition model are all stored in the database under the corresponding "T001~D02" data group. The associative storage method enables the database to easily query all relevant data under a specific training situation according to the task number and difficulty mode number, providing a rich and organized data foundation for subsequent training difficulty calibration, training effect evaluation and training strategy optimization, which helps to continuously improve the content and accuracy of the training difficulty level calibration database, thereby improving the performance and adaptability of the entire driving simulation training system.
[0022] Step S300, based on the training difficulty level calibration database, the physiological signal timing information of the target user, the voice signal timing information of the target user and the facial expression image timing information of the target user are calibrated for training difficulty, and the first matching training difficulty is obtained. Specifically, based on the training difficulty level calibration database, the current simulation training task number and the difficulty mode number of the target user are first used as search conditions to extract the corresponding first recorded physiological, voice, and facial expression image timing information and the first recorded training difficulty. The first similarity coefficient between the physiological signal timing information of the target user and the first recorded physiological signal timing information is calculated respectively, and the similarity of the change trend of multi-dimensional indicators such as heart rate, blood pressure, and skin conductivity is comprehensively analyzed; the second similarity coefficient between the voice signal timing information of the target user and the first recorded voice signal timing information is calculated, and the acoustic features such as intonation, speech speed, and emotional vocabulary are analyzed; the third similarity coefficient between the first recorded facial expression image timing information and the target user facial expression image timing information is calculated, and the facial muscle movement, expression duration, etc. are compared with the image recognition and expression analysis algorithm. When the first similarity coefficient, the second similarity coefficient, and the third similarity coefficient are respectively greater than or equal to their respective similarity thresholds, the first recorded training difficulty is determined as the first matching training difficulty for the target user. This can utilize past data experience to provide a reasonable initial difficulty setting basis for training, thereby improving the targetedness and effectiveness of training and facilitating personalized arrangements.
[0023] In a possible implementation, the training difficulty is calibrated for the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty, and step S300 further includes step S310, inputting the simulation training task number and the simulation training difficulty mode number into the training difficulty level calibration database, extracting the first recorded physiological signal timing information, the first recorded voice signal timing information, the first recorded facial expression image timing information, and the first recorded training difficulty. Specifically, the simulation training task number and the simulation training difficulty mode number corresponding to the current simulation training of the target user are input into the training difficulty level calibration database. The database performs retrieval and extraction operations in the data storage based on two key numbers. Specifically, the first recorded physiological signal timing information matching the task number and difficulty mode number will be extracted, and the first recorded physiological signal timing information contains detailed data on the changes over time of various physiological signals such as heart rate, blood pressure, skin conductivity, etc. generated by users who have participated in the same training in the past; at the same time, the first recorded voice signal timing information is extracted, which covers the voice intonation change sequence, speech speed change curve, voice content and other information reflecting the user's language expression characteristics; the first recorded facial expression image timing information will also be extracted, that is, the changes in the user's facial expressions at different training stages recorded by the image acquisition device, such as the time distribution and duration of expressions such as smiles, frowns, and surprise; and the corresponding first recorded training difficulty will be extracted at the same time, and the difficulty value is determined based on the previous comprehensive evaluation of the same training task and difficulty mode.
[0024] Step S320, calculate the first similarity coefficient between the first recorded physiological signal timing information and the target user physiological signal timing information. Specifically, after successfully extracting the above-mentioned first recorded information, start calculating the first similarity coefficient. Since the physiological signal contains multiple different indicators, corresponding distribution weights will be set for different indicators. For example, for some indicators that are sensitive to emotional reactions and have an important indicative role in driving simulation training, such as heart rate and skin conductivity, they will be given relatively high weights; while indicators that relatively indirectly reflect emotions, such as body temperature, have lower weights. After determining the weights, similarity calculations are performed on each physiological signal indicator according to the weights. Taking heart rate as an example, the heart rate change curve in the first recorded physiological signal timing information and the heart rate change curve in the target user physiological signal timing information will be compared to analyze the similarity between the two in terms of the average value of the heart rate, fluctuation amplitude, peak time point, and change trend; for blood pressure signals, the similarity of their numerical values and the consistency of fluctuation periods are also examined; skin conductivity focuses on the similarity of its change response when different training scenarios are converted. By comprehensively calculating each indicator according to its weight, a first similarity coefficient is finally obtained, which can accurately reflect the similarity between the first recorded physiological signal timing information and the target user's physiological signal timing information.
[0025] Step S330, calculate the second similarity coefficient between the target user voice signal timing information and the first recorded voice signal timing information. Specifically, the second similarity coefficient is then calculated, that is, the similarity between the target user voice signal timing information and the first recorded voice signal timing information. The calculation process analyzes multiple key features of the voice signal. In terms of intonation, compare the rising and falling changes in the intonation of the two during the entire training process, for example, whether the intonation rises when encountering complex driving scenes, and whether the amplitude and frequency of the increase are similar; for speech rate, examine its fast and slow change rhythm in different training stages, such as whether the speech rate is accelerated when the driving operation is more intense, and whether the degree of acceleration is similar; at the same time, the speech content will be semantically analyzed to count whether the frequency of occurrence of specific emotional words or words related to driving operations is consistent. By comprehensively comparing and quantitatively analyzing the intonation, speech rate, semantics and other features of the voice signal, a second similarity coefficient that can accurately represent the degree of similarity between the two is obtained.
[0026] Step S340, calculate the third similarity coefficient between the timing information of the first recorded expression image and the timing information of the expression image of the target user. Specifically, calculate the third similarity coefficient, that is, the similarity between the timing information of the first recorded expression image and the timing information of the expression image of the target user. With the help of image recognition technology and expression analysis algorithm, firstly, the facial muscle movement characteristics in the expression image are compared in detail. For example, when facing sudden traffic conditions, observe whether both have similar frowning muscle movements, and whether the depth, duration and position distribution of frowning on the face are similar; for smiling expressions, analyze the similarity in terms of the angle of the corner of the mouth, the degree of squinting of the eyes, and the degree of stretching of the entire facial expression; at the same time, the frequency and timing of expression conversion will be examined, such as the time point of conversion from tense expression to relaxed expression and whether the intermediate expression state in the conversion process is consistent. By comprehensively analyzing and quantitatively evaluating the multi-dimensional features of these expression images, the third similarity coefficient that can accurately reflect the similarity between the two is finally determined.
[0027] Step S350, when the first similarity coefficient is greater than or equal to the first similarity threshold, and the second similarity coefficient is greater than or equal to the second similarity threshold, and the third similarity coefficient is greater than or equal to the third similarity threshold, the first recorded training difficulty is set as the first matching training difficulty. Specifically, when the first similarity coefficient obtained through the above calculation process is greater than or equal to the first similarity threshold, and the second similarity coefficient is greater than or equal to the second similarity threshold, and the third similarity coefficient is greater than or equal to the third similarity threshold, this indicates that the physiological state, language expression state and facial expression state of the target user in the simulation training are highly similar to the user state in the same training task and difficulty mode recorded in the database. In this case, based on the reliability of the empirical data, the first recorded training difficulty is directly set as the first matching training difficulty of the target user. For example, if the first recorded training difficulty is "advanced difficulty", then the first matching training difficulty of the target user will also be determined as "advanced difficulty", thereby providing a reasonable and well-founded initial training difficulty setting for the subsequent simulated driving training, which helps to improve the pertinence and effectiveness of the entire training process.
[0028] Step S400, when the first matching training difficulty is empty or not unique, the target user's physiological signal timing information, the target user's voice signal timing information and the target user's facial expression image timing information are processed through the emotion recognition model to obtain the second matching training difficulty. Specifically, when the first matching training difficulty is empty or not unique, the emotion recognition model composed of the physiological signal pre-processing network, the voice signal pre-processing network, the facial expression image pre-processing network and the post-emotion recognition fitting network is used for processing. The target user's physiological signal timing information (covering data such as heart rate, blood pressure, skin conductivity, etc. that change with driving scenes and emotions), voice signal timing information (including emotion-related features such as intonation, speech rate, loudness and voice content) and facial expression image timing information (such as frowning, smiling, etc.) are respectively input into the corresponding pre-processing network. The physiological signal pre-processing network determines the physiological signal identification training difficulty and trains according to the historical driving training data (including the first monitoring physiological signal time series information and qualified training time record data) under the specific simulation training task and difficulty mode number through complex steps such as similarity calculation, clustering, grouping, and mode calculation. After inputting the target user data, the physiological signal matching training difficulty is output; the speech signal pre-processing network is similarly trained based on historical speech data, and the speech signal matching training difficulty is obtained by analyzing the target user's speech characteristics; the expression image pre-processing network uses a large amount of expression image historical data to learn and establish mapping relationships, and obtains the expression matching training difficulty based on the input target user's expression image. Finally, the three difficulty values are input into the post-emotion recognition fitting network, and the integrated optimization such as comprehensive analysis of each modal information and weight assignment is performed to output the second matching training difficulty that comprehensively reflects the target user's emotional state, so as to improve the effectiveness and adaptability of simulated driving training.
[0029] Step S500, adjusting the simulation training difficulty mode according to the second matching training difficulty. Specifically, obtain the operation step sequence corresponding to the simulation training task number, which covers a series of actions and decision-making processes from vehicle startup to coping with various traffic conditions, such as startup inspection, intersection start, speed limit adjustment, turning operation, etc. in the urban driving simulation task. The operation step sequence is divided into different numbers of steps to generate a simulation training difficulty mode, from a single-step segmentation to obtain the first simulation training difficulty mode, such as considering the basic difficulty of the vehicle startup step alone, to two-step segmentation to form the second simulation training difficulty mode, focusing on the difficulty increase of the coherence and synergy between steps, until the Kth simulation training difficulty mode is obtained by segmentation by K steps, K is the total number of operation steps, and the difficulty gradually becomes more complex and comprehensive as the number of segmentation steps increases. Then add these difficulty modes to the set, select the adaptation mode in the set according to the obtained second matching training difficulty, and if the target user's emotions cause the second matching training difficulty to change during training, the system will dynamically switch to the corresponding new mode to accurately adapt the difficulty mode to the user's emotions and matching training difficulty, so as to achieve the best training effect and avoid the inappropriate difficulty affecting the training effect.
[0030] In a possible implementation, the emotion recognition model includes a physiological signal pre-processing network, a speech signal pre-processing network, an expression image pre-processing network, and a post-emotion recognition fitting network. Specifically, the emotion recognition model is a comprehensive multi-module architecture, which is composed of a physiological signal pre-processing network, a speech signal pre-processing network, an expression image pre-processing network, and a post-emotion recognition fitting network. The four networks work together to accurately identify the emotional state of the target user from information sources of different dimensions, and then determine the training difficulty that is adapted to it.
[0031] According to the second matching training difficulty, the simulation training difficulty mode is adjusted. Step S500 further includes step S510, processing the target user's physiological signal timing information through the physiological signal pre-processing network to obtain the physiological signal matching training difficulty. Specifically, the physiological signal pre-processing network focuses on processing the target user's physiological signal timing information. The physiological signal timing information contains a variety of indicators that can reflect the user's internal emotional changes, such as heart rate data, which will fluctuate with the user's tension, excitement level or fatigue state in driving simulation training. When facing complex and dangerous driving scenes, the heart rate tends to increase significantly; the same is true for blood pressure data. In high-pressure situations, blood pressure may increase; skin conductivity is also a key indicator. When emotions fluctuate, changes in sweat gland secretion will cause changes in skin conductivity. The network is constructed based on a large amount of historical driving training data. During the construction process, historical driving training data under a specific simulation training task number and simulation training difficulty mode number are first collected, which includes the first monitoring physiological signal timing information and qualified training time record data. The qualified training time record data refers to the training time to qualified after triggering the first monitoring physiological signal timing information. Then, a training difficulty mapping table is set, and the qualified training duration record data is mapped by calculation, so as to obtain the training difficulty of physiological signal identification. Specifically, multiple groups of one-to-one corresponding first monitoring physiological signal timing information and qualified training duration record data are first obtained; then, the first monitoring physiological signal timing information is calculated for similarity between the two to form a first monitoring physiological signal similarity coefficient set; according to the set similarity coefficient threshold, the first monitoring physiological signal timing information is clustered in combination with the similarity coefficient set to obtain multiple clusters of first monitoring physiological signal timing information; then, according to the clustering result, the qualified training duration record data is grouped to obtain multiple groups of qualified training duration record data; the combination is traversed to calculate the mode value to obtain multiple qualified training duration identification data; finally, single data is randomly extracted from multiple clusters of first monitoring physiological signal timing information, and multiple qualified training duration identification data are processed in combination with the training difficulty mapping table to determine the training difficulty of physiological signal identification. Afterwards, the physiological signal identification training difficulty is supervised, and the first monitoring physiological signal timing information is used as input to train the physiological signal pre-processing network. When the target user's physiological signal timing information is input, the network performs in-depth analysis and feature extraction based on the trained model parameters and algorithms, and finally outputs the physiological signal matching training difficulty.
[0032] Step S520, the target user's voice signal timing information is processed by the voice signal pre-processing network to obtain the voice signal matching training difficulty. Specifically, the voice signal pre-processing network mainly processes the target user's voice signal timing information. In the driving simulation training process, the user's voice signal contains many features closely related to emotions. For example, the high and low changes in intonation can intuitively reflect the user's emotional state. When in excitement, tension or excitement, the intonation usually rises; the speed of speech is also an important emotional indicator. Fast speech is often associated with emotions such as anxiety and excitement, while slow speech may imply the user's relaxation, meditation or fatigue; the vocabulary selection and expression in the voice content are also not to be ignored. Some emotional words or specific driving-related words can further reveal the user's psychological state. The construction steps of this network are similar to those of the physiological signal pre-processing network, and are also trained based on a large amount of historical data related to voice signals. During the training process, the features such as intonation, speech speed, and voice content in the historical voice signal data are analyzed and extracted to establish a correlation model between voice signal features and training difficulty. When the target user's voice signal timing information is input, the network will quickly capture various feature information in the voice, such as identifying keywords and emotional words in the voice content through speech recognition technology, analyzing the changing trend of intonation and the changing rhythm of speech speed, etc., and calculating the difficulty of voice signal matching training based on the established association model.
[0033] Step S530, the target user's expression image timing information is processed by the expression image pre-processing network to obtain the expression matching training difficulty. Specifically, the expression image pre-processing network processes the target user's expression image timing information. Expression images are intuitive external manifestations of user emotions. For example, frowning is often associated with emotions such as confusion, anxiety, and dissatisfaction; smiling is likely to reflect that the user is in a good mood, is relatively satisfied with the driving situation, or is easy to deal with; wide eyes may mean that the user is surprised or alert; the shape and opening degree of the mouth can also reflect different emotions, such as a closed mouth may indicate tension or concentration, and an open mouth may be surprised or shouting. When constructing the network, a large amount of expression image historical data is used. In-depth study and analysis are performed on the facial muscle movement features in the expression image, the types of expressions (such as happiness, sadness, anger, surprise, etc.), the duration of expressions, and the frequency of expression conversion. Through machine learning algorithms, a mapping relationship between expression image features and training difficulty is established. When the target user's facial expression image time series information is input, the network can accurately identify the type and characteristics of the expression, such as determining whether the expression is a brief surprise or a long-term tension, and derive the difficulty of expression matching training based on the changes in the expression and the established mapping relationship.
[0034] Step S540, input the physiological signal matching training difficulty, the voice signal matching training difficulty and the expression matching training difficulty into the post-emotion recognition fitting network, and output the second matching training difficulty. Specifically, after the physiological signal pre-processing network, the voice signal pre-processing network and the expression image pre-processing network output the physiological signal matching training difficulty, the voice signal matching training difficulty and the expression matching training difficulty respectively, the three difficulty values are input into the post-emotion recognition fitting network. The post-emotion recognition fitting network will comprehensively consider the three training difficulty values from different modal information, and use the fitting algorithm to integrate and optimize. For example, different weights will be assigned to different modal information according to their differences in accuracy and importance of emotional expression, and then weighted summation and other operations will be performed. After being processed by the post-emotion recognition fitting network, the second matching training difficulty is finally output. The difficulty value integrates the emotional information reflected by the physiological state, voice expression state and expression state of the target user, and can more accurately determine the training difficulty suitable for the current emotional state of the target user, thereby improving the effectiveness and adaptability of simulated driving training.
[0035] In a possible implementation, the physiological signal pre-processing network processes the target user's physiological signal timing information to obtain the physiological signal matching training difficulty, and step S510 further includes step S511, collecting historical driving training data in the simulation training task number and the simulation training difficulty mode number, wherein the historical driving training data includes the first monitoring physiological signal timing information and qualified training duration record data, and the qualified training duration record data refers to the training duration to qualified after triggering the first monitoring physiological signal timing information. Specifically, for a specific simulation training task number and simulation training difficulty mode number, collect the historical driving training data related thereto. In the data, the focus is on the first monitoring physiological signal timing information, which covers the detailed records of various physiological indicators such as the heart rate change curve, blood pressure fluctuation data and dynamic changes of skin conductivity during the training process over time. At the same time, it also includes qualified training duration record data, which clearly records the time from the triggering of the first monitoring physiological signal timing information until the driver reaches the qualified standard in the training task and difficulty mode. For example, if the simulation training task is driving training under complex urban road conditions and the simulation training difficulty mode is medium difficulty, then the collected historical data will reflect the changes in physiological signals of different drivers in specific situations and the differences in time required for each of them to reach the qualification.
[0036] Step S512, set a training difficulty mapping table, map the qualified training duration record data, and obtain the physiological signal identification training difficulty. Specifically, set a special training difficulty mapping table. The mapping table establishes a corresponding relationship between qualified training duration record data and training difficulty based on the analysis and summary of a large amount of historical data. By inputting the collected qualified training duration record data into this training difficulty mapping table, performing corresponding mapping operations, the physiological signal identification training difficulty is determined. For example, if it is found that in a certain set of historical data, after triggering a specific physiological signal, the driver reaches the qualified standard in a short time, then according to the mapping table, the corresponding physiological signal identification training difficulty will be set to a lower level; conversely, if the qualified training duration is longer, it may correspond to a higher physiological signal identification training difficulty. The key is to accurately construct a reasonable mapping logic between qualified training duration and training difficulty, so as to provide a reliable identification basis for subsequent network training.
[0037] Step S513, taking the physiological signal identification training difficulty as supervision and the first monitoring physiological signal timing information as input, the physiological signal pre-processing network is trained. Specifically, taking the obtained physiological signal identification training difficulty as supervision information and the first monitoring physiological signal timing information as input data of the network, the physiological signal pre-processing network is trained. During the training process, the network continuously learns the intrinsic relationship between the various characteristic patterns in the first monitoring physiological signal timing information and the corresponding physiological signal identification training difficulty. For example, the network will sort out what kind of training difficulty prediction value should be output under specific heart rate change trends, blood pressure fluctuation patterns, and skin conductivity change laws. Through repeated training of a large amount of historical data, the physiological signal pre-processing network continuously optimizes its own model parameters and algorithm structure, so that when faced with new target users' physiological signal timing information, it can accurately predict the physiological signal matching training difficulty that matches it.
[0038] Step S514, wherein the speech signal pre-processing network and the expression image pre-processing network are constructed in the same steps as the physiological signal pre-processing network. Specifically, the construction steps of the speech signal pre-processing network and the expression image pre-processing network are the same as those of the physiological signal pre-processing network. For the speech signal pre-processing network, historical speech signal data under a specific simulation training task number and a simulation training difficulty mode number are collected, including information such as the voice intonation change sequence, the speech speed change curve, and the speech content, as well as the corresponding qualified training time record data (qualified training time refers to the training time associated with the speech signal feature to pass). A training difficulty mapping table for the speech signal is set, and the qualified training time record data is mapped to the speech signal identification training difficulty. The speech signal identification training difficulty is supervised, and the historical speech signal data is used as input to train the speech signal pre-processing network. For the expression image pre-processing network, we first collect historical expression image data under specific tasks and difficulty modes, such as facial muscle movement characteristics of different expressions, expression duration and other information, as well as qualified training time record data, set the expression image training difficulty mapping table, map the expression image identification training difficulty, and then use this as supervision and historical expression image data as input to train the expression image pre-processing network. Through the same construction steps, the three pre-processing networks can effectively convert the user's multimodal information into corresponding matching training difficulty information in their respective information processing fields, laying a solid foundation for the entire emotion recognition model to accurately evaluate the target user's emotional state and determine the appropriate training difficulty.
[0039] In a possible implementation, a training difficulty mapping table is set, the qualified training duration record data is mapped, and the physiological signal identification training difficulty is obtained. Step S512 further includes step S5121, obtaining a number of first monitoring physiological signal timing information and a number of qualified training duration record data corresponding to each other. Specifically, from the historical driving training data resources, a number of first monitoring physiological signal timing information and a number of qualified training duration record data corresponding to each other are obtained. The first monitoring physiological signal timing information records in detail the physiological signal change process of different drivers under a specific simulation training task number and a simulation training difficulty mode number, such as how the heart rate fluctuates with the switching of driving scenes, the fluctuation of blood pressure when encountering complex road conditions, and the subtle changes of skin conductivity at tense moments. The corresponding qualified training duration record data indicates the length of time it takes for the driver to finally reach the qualified standard in the training task from the beginning of a specific change in the physiological signal (i.e., triggering the first monitoring physiological signal timing information). For example, in a city road driving simulation training scenario with a medium level of difficulty, the timing information of a driver's first monitored physiological signal showed that his heart rate showed a specific fluctuation curve when passing multiple intersections and traffic conditions changed, while his qualified training time record data showed that it took a total of 30 minutes from the first obvious change in heart rate to qualified training.
[0040] Step S5122, perform pairwise similarity calculation on the plurality of first monitoring physiological signal time series information to obtain a first monitoring physiological signal similarity coefficient set. Specifically, perform pairwise similarity calculation on the plurality of first monitoring physiological signal time series information obtained. The calculation process is not a simple numerical comparison, but a comprehensive analysis of the change trend, fluctuation amplitude and similarity of signal characteristics of multiple physiological signal indicators. For example, for the heart rate signal, it is necessary not only to compare the closeness of the average heart rate value, but also to analyze whether the acceleration and deceleration change rules of the heart rate during the training process are consistent; for the blood pressure signal, analyze the similarity of the time point and numerical value of the peak and valley of the blood pressure; for the skin conductivity signal, pay attention to the change slope and the degree of fit of the fluctuation cycle in different training stages. Integrate the multi-dimensional comparison results into numerical values that can quantify the degree of similarity, and then construct a first monitoring physiological signal similarity coefficient set. Each of the similarity coefficients represents the degree of similarity between a pair of first monitoring physiological signal time series information.
[0041] Step S5123, according to the similarity coefficient threshold, combined with the first monitoring physiological signal similarity coefficient set, the plurality of first monitoring physiological signal timing information are clustered to obtain multiple clusters of first monitoring physiological signal timing information. Specifically, according to the pre-set similarity coefficient threshold, combined with the constructed first monitoring physiological signal similarity coefficient set, a plurality of first monitoring physiological signal timing information are clustered. When the similarity coefficient between two first monitoring physiological signal timing information is greater than or equal to this threshold, they are classified into the same category. Through the clustering process, a plurality of first monitoring physiological signal timing information are orderly divided into multiple clusters of first monitoring physiological signal timing information. The physiological signal timing information in each cluster has a high similarity, representing a group of drivers with similar physiological response characteristics in a specific driving scenario and difficulty mode. For example, a cluster of first monitoring physiological signal timing information corresponds to a group of drivers whose heart rate fluctuations are relatively gentle, blood pressure changes are relatively stable, and skin conductivity changes are small when facing frequent start-stop and complex traffic lights on urban roads.
[0042] Step S5124, grouping the several qualified training duration record data according to the multiple clusters of first monitoring physiological signal timing information, and obtaining multiple groups of qualified training duration record data. Specifically, after completing the clustering of the first monitoring physiological signal timing information, grouping the several qualified training duration record data according to the obtained multiple clusters of first monitoring physiological signal timing information. Since the qualified training duration record data and the first monitoring physiological signal timing information are in a one-to-one correspondence, when the first monitoring physiological signal timing information is clustered into different clusters, the corresponding qualified training duration record data is naturally divided into different groups, thereby obtaining multiple groups of qualified training duration record data. Grouping makes each group of qualified training duration record data associated with the first monitoring physiological signal timing information of a specific cluster, laying the foundation for the subsequent in-depth analysis of the training duration differences of different physiological signal feature groups.
[0043] Step S5125, traverse the multiple groups of qualified training duration record data to calculate the mode value, and obtain multiple qualified training duration identification data. Specifically, traverse each group of qualified training duration record data and calculate its mode value. The mode value represents the qualified training duration with the highest frequency in the driver group corresponding to the first monitoring physiological signal timing information of a specific cluster. By calculating the mode value, multiple qualified training duration identification data are obtained. The identification data can reflect the typical characteristics of the driver group corresponding to the group of data in terms of training duration to a certain extent. For example, if a group of qualified training duration record data is [25, 28, 25, 30, 25], then its mode value is 25, and 25 minutes is determined as the qualified training duration identification data of the group.
[0044] Step S5126, traverse the multiple clusters of first monitoring physiological signal timing information and randomly extract single data respectively, to obtain multiple first monitoring physiological signal timing information. Specifically, traverse the multiple clusters of first monitoring physiological signal timing information, randomly extract single data from each cluster respectively, to obtain multiple first monitoring physiological signal timing information. The randomly extracted data will be used as the basic data for subsequent processing based on the training difficulty mapping table. They not only retain the characteristic representativeness of each cluster of the first monitoring physiological signal timing information, but also have a certain degree of randomness, and can provide diverse input samples for the application of the training difficulty mapping table.
[0045] Step S5127, based on the training difficulty mapping table, the plurality of qualified training duration identification data are processed separately to obtain the physiological signal identification training difficulty. Specifically, based on the pre-set training difficulty mapping table, the plurality of qualified training duration identification data are processed separately. The training difficulty mapping table is a data conversion rule system constructed based on a large amount of historical data and professional knowledge, which can map the qualified training duration identification data to the corresponding physiological signal identification training difficulty according to the size and distribution of the qualified training duration identification data. For example, if a qualified training duration identification data is short, it indicates that in the physiological signal characteristic group, the driver can generally reach the training qualification standard quickly, then according to the training difficulty mapping table, the corresponding physiological signal identification training difficulty will be set to a lower level; conversely, if the qualified training duration identification data is long, it corresponds to a higher physiological signal identification training difficulty. Through the mapping process, the physiological signal identification training difficulty is finally obtained, which provides key supervision information and difficulty calibration basis for the subsequent training of the physiological signal pre-processing network.
[0046] In a possible implementation, the simulation training difficulty mode is adjusted according to the second matching training difficulty, and step S500 further includes step S560, obtaining an operation step sequence of the simulation training task number. Specifically, the operation step sequence corresponding to the simulation training task number is accurately obtained from the management system or data repository of the simulation driving training task. Taking an ordinary urban road driving simulation training task as an example, its operation step sequence covers the whole process from getting ready to finally parking and leaving the vehicle. Preparation for getting on the vehicle includes adjusting the seat to a suitable position, fastening the seat belt, checking whether the rearview mirror and various indicator lights on the instrument panel are normal, etc.; the steps for starting the vehicle include inserting the key or pressing the start button, observing the vehicle's self-check status, stepping on the clutch (manual transmission) or shifting into a suitable gear (automatic transmission) and releasing the handbrake; the operating steps during driving include starting according to the instructions of the traffic lights, maintaining a suitable speed under different road conditions, turning on the turn signal in advance when turning and observing the surrounding traffic conditions, confirming the safe distance and using the turn signal to indicate when changing lanes, correctly judging the right of way at the intersection and giving way to pedestrians and other vehicles, etc.; the parking steps involve turning on the turn signal in advance to indicate pulling over, slowly decelerating and stopping the vehicle steadily at the specified position, pulling up the handbrake, shifting into neutral (manual transmission) or P gear (automatic transmission), turning off the engine and unbuckling the seat belt, etc.
[0047] Step S570, the sequence of operation steps is divided according to a single operation step to obtain a first simulation training difficulty mode. Specifically, the obtained sequence of operation steps is divided according to a single operation step to construct a first simulation training difficulty mode. For example, for the single step of "adjusting the seat to a suitable position", in this difficulty mode, the trainee's familiarity with the seat adjustment function, such as the front and rear position of the seat, the height adjustment, and the adjustment method of the backrest inclination, may be set to some seat adjustment devices that are slightly stuck or not very sensitive, to see whether the trainee can operate correctly to achieve a comfortable and safe driving posture. For the step of "starting according to the traffic light indication", the trainee's reaction speed to the color change of the signal light, the coordination of the clutch and the throttle when starting (manual gear) or the switching of the brake and the throttle (automatic gear), and whether the vehicle starts smoothly, whether there is a slipping or rushing phenomenon, etc. This single-step segmentation method allows the difficulty of each operation step to be considered and set separately, which helps novice trainees to initially master the essentials and specifications of each driving action.
[0048] Step S580, divide the operation step sequence into two operation steps to obtain a second simulation training difficulty mode. Specifically, after the single-step segmentation is completed, the operation step sequence is further divided into two operation steps to obtain a second simulation training difficulty mode. Take the two operation step combinations of "turn on the turn signal in advance and observe the surrounding traffic conditions when turning" and "confirm the safe distance and use the turn signal to signal when changing lanes" as an example. The difficulty setting at this time is no longer limited to the basic operation of a single step, but must take into account the continuity and coordination between the two steps. Trainees must not only complete the turning operation accurately, but also change lanes at the right time, which requires trainees to have a certain comprehensive judgment ability and operational coordination ability.
[0049] Step S590, until the operation step sequence is divided into K operation steps, a Kth simulation training difficulty mode is obtained, wherein K is equal to the total number of operation steps. Specifically, according to the segmentation logic, the number of segmented operation steps is continuously increased, until the operation step sequence is divided into K operation steps, a Kth simulation training difficulty mode is obtained, wherein K is equal to the total number of operation steps. As the number of segmented steps increases, the difficulty mode will gradually become more complex and comprehensive.
[0050] Step S5100, adding the first simulation training difficulty mode, the second simulation training difficulty mode, and the Kth simulation training difficulty mode into the simulation training difficulty mode. Specifically, the generated first simulation training difficulty mode, the second simulation training difficulty mode, and the Kth simulation training difficulty mode are added into the simulation training difficulty mode set. During the simulation driving training process, the appropriate difficulty mode can be selected from the difficulty mode set to carry out training based on the actual situation of the trainees, such as learning progress, driving skill level, current emotional state, etc.
[0051] In the above, refer to Figure 1 The present invention describes in detail a method for dynamically constraining the difficulty of simulation training based on emotion recognition according to an embodiment of the present invention. Figure 2 A dynamic constraint system for simulation training difficulty based on emotion recognition according to an embodiment of the present invention is described.
[0052] The dynamic constraint system for simulation training difficulty based on emotion recognition according to the embodiment of the present invention is used to solve the technical problem that the difficulty of existing driving simulation training cannot be adaptively adjusted according to user needs, resulting in a poor degree of individualization of simulation training, and achieves the technical effect of improving the individualization and adaptive adjustment capabilities of driving simulation training. The dynamic constraint system for simulation training difficulty based on emotion recognition includes: an expression image time series information acquisition module 10, a calibration database acquisition module 20, a training difficulty calibration module 30, a second matching training difficulty acquisition module 40, and a difficulty mode adjustment module 50.
[0053] The expression image timing information acquisition module 10 is used to collect the target user's physiological signal timing information and the target user's voice signal timing information through wearable sensors when the target user enters the driving simulation training area and starts training, and to collect the target user's expression image timing information through the simulation training area camera.
[0054] The calibration database acquisition module 20 is used to obtain a training difficulty level calibration database.
[0055] The training difficulty calibration module 30 is used to calibrate the training difficulty of the target user's physiological signal timing information, the target user's voice signal timing information and the target user's facial expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty.
[0056] The second matching training difficulty acquisition module 40 is used to process the target user's physiological signal timing information, the target user's voice signal timing information and the target user's expression image timing information through an emotion recognition model to obtain a second matching training difficulty when the first matching training difficulty is empty or not unique.
[0057] The difficulty mode adjustment module 50 is used to adjust the simulation training difficulty mode according to the second matching training difficulty.
[0058] The specific configuration of the calibration database acquisition module 20 will be described in detail below. As described above, the training difficulty level calibration database is obtained, and the calibration database acquisition module 20 further includes: a number acquisition unit, the number acquisition unit is used to obtain the simulation training task number and the simulation training difficulty mode number; an information storage unit, the information storage unit is used to associate the second matching training difficulty, the simulation training task number, the simulation training difficulty mode number with the target user physiological signal timing information, the target user voice signal timing information and the target user expression image timing information and store them in the training difficulty level calibration database.
[0059] The specific configuration of the training difficulty calibration module 30 will be described in detail below. As described above, the training difficulty is calibrated for the target user's physiological signal timing information, the target user's voice signal timing information and the target user's facial expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty. The training difficulty calibration module 30 further includes: a timing information extraction unit, the timing information extraction unit is used to input the simulation training task number and the simulation training difficulty mode number into the training difficulty level calibration database, extract the first recorded physiological signal timing information, the first recorded voice signal timing information, the first recorded facial expression image timing information and the first recorded training difficulty; a first similarity coefficient calculation unit, the first similarity coefficient calculation unit is used to calculate the similarity between the first recorded physiological signal timing information and the target user a first similarity coefficient of the target user's physiological signal timing information; a second similarity coefficient calculation unit, the second similarity coefficient calculation unit is used to calculate the second similarity coefficient between the target user's voice signal timing information and the first recorded voice signal timing information; a third similarity coefficient calculation unit, the third similarity coefficient calculation unit is used to calculate the third similarity coefficient between the first recorded expression image timing information and the target user's expression image timing information; a first record training difficulty setting unit, the first record training difficulty setting unit is used to set the first record training difficulty to the first matching training difficulty when the first similarity coefficient is greater than or equal to a first similarity threshold, the second similarity coefficient is greater than or equal to a second similarity threshold, and the third similarity coefficient is greater than or equal to a third similarity threshold.
[0060] Next, the specific configuration of the difficulty mode adjustment module 50 will be described in detail. As described above, according to the second matching training difficulty adjustment simulation training difficulty mode, the difficulty mode adjustment module 50 further includes: an emotion recognition model composition unit, the emotion recognition model composition unit is used for the emotion recognition model to include a physiological signal pre-processing network, a voice signal pre-processing network, an expression image pre-processing network and a post-emotion recognition fitting network; a physiological signal matching training difficulty acquisition unit, the physiological signal matching training difficulty acquisition unit is used to process the target user's physiological signal timing information through the physiological signal pre-processing network to obtain the physiological signal matching training difficulty; a voice signal matching training difficulty acquisition unit, the voice signal matching training difficulty acquisition unit is used to process the target user's voice signal timing information through the voice signal pre-processing network to obtain the voice signal matching training difficulty; an expression matching training difficulty acquisition unit, the expression matching training difficulty acquisition unit is used to process the target user's expression image timing information through the expression image pre-processing network to obtain the expression matching training difficulty; a second matching training difficulty output unit, the second matching training difficulty output unit is used to input the physiological signal matching training difficulty, the voice signal matching training difficulty and the expression matching training difficulty into the post-emotion recognition fitting network to output the second matching training difficulty.
[0061] The target user's physiological signal timing information is processed by the physiological signal pre-processing network to obtain the physiological signal matching training difficulty, and the physiological signal matching training difficulty acquisition unit further includes: a driving training data acquisition subunit, the driving training data acquisition subunit is used to collect historical driving training data in the simulation training task number and the simulation training difficulty mode number, wherein the historical driving training data includes the first monitoring physiological signal timing information and qualified training time record data, and the qualified training time record data refers to the training time to pass after triggering the first monitoring physiological signal timing information; a record data mapping subunit, the record data mapping subunit is used to set a training difficulty mapping table, map the qualified training time record data, and obtain the physiological signal identification training difficulty; a pre-processing network training subunit, the pre-processing network training subunit is used to train the physiological signal pre-processing network with the physiological signal identification training difficulty as supervision and the first monitoring physiological signal timing information as input; a processing network construction subunit, the processing network construction subunit is used to wherein the speech signal pre-processing network and the expression image pre-processing network have the same construction steps as the physiological signal pre-processing network.
[0062] Among them, a training difficulty mapping table is set, and the qualified training duration recording data is mapped to obtain the physiological signal identification training difficulty. The recording data mapping subunit further includes: a training duration recording data acquisition microunit, and the training duration recording data acquisition microunit is used to obtain a number of first monitoring physiological signal timing information and a number of qualified training duration recording data corresponding to each other; a similarity coefficient set acquisition microunit, and the similarity coefficient set acquisition microunit is used to perform pairwise similarity calculations on the number of first monitoring physiological signal timing information to obtain a first monitoring physiological signal similarity coefficient set; a timing information clustering microunit, and the timing information clustering microunit is used to cluster the number of first monitoring physiological signal timing information according to a similarity coefficient threshold and in combination with the first monitoring physiological signal similarity coefficient set to obtain multiple clusters of first monitoring physiological signal timing information; a recording data grouping microunit Element, the recorded data grouping micro-unit is used to group the several qualified training duration recording data according to the multiple clusters of first monitoring physiological signal timing information, so as to obtain multiple groups of qualified training duration recording data; the majority value calculation micro-unit is used to traverse the multiple groups of qualified training duration recording data to perform majority value calculation, so as to obtain multiple qualified training duration identification data; the first monitoring physiological signal timing information acquisition micro-unit is used to traverse the multiple clusters of first monitoring physiological signal timing information and randomly extract single data respectively, so as to obtain multiple first monitoring physiological signal timing information; the physiological signal identification training difficulty acquisition micro-unit is used to process the multiple qualified training duration identification data respectively based on the training difficulty mapping table, so as to obtain the physiological signal identification training difficulty.
[0063] Among them, the difficulty mode adjustment module 50 further includes: an operation step sequence acquisition unit, which is used to obtain the operation step sequence of the simulation training task number; an operation step segmentation unit, which is used to segment the operation step sequence according to a single operation step to obtain a first simulation training difficulty mode; a second simulation training difficulty mode acquisition unit, which is used to segment the operation step sequence according to two operation steps to obtain a second simulation training difficulty mode; a traversal step segmentation unit, which is used until the operation step sequence is segmented according to K operation steps to obtain the Kth simulation training difficulty mode, where K is equal to the total number of operation steps; a training difficulty mode adding unit, which is used to add the first simulation training difficulty mode, the second simulation training difficulty mode until the Kth simulation training difficulty mode into the simulation training difficulty mode.
[0064] The dynamic constraint system for simulation training difficulty based on emotion recognition provided by the embodiment of the present invention can execute the dynamic constraint method for simulation training difficulty based on emotion recognition provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0065] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0066] The above specific implementation manner does not constitute a limitation to the protection scope of the present application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application. In some cases, the actions or steps recorded in the present application can be performed in an order different from that in the embodiment and can still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A dynamic constraint method for simulation training difficulty based on emotion recognition, characterized in that: The dynamic constraint system for simulation training difficulty based on emotion recognition is applied, the system is connected to a wearable sensor in communication, and the wearable sensor is deployed at a preset position of the target user, including: When the target user enters the driving simulation training area and starts training, the target user's physiological signal timing information and the target user's voice signal timing information are collected through wearable sensors, and the target user's facial expression image timing information is collected through the simulation training area camera; Obtain a training difficulty level calibration database; Based on the training difficulty level calibration database, the training difficulty of the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information is calibrated to obtain a first matching training difficulty; When the first matching training difficulty is empty or not unique, the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information are processed by an emotion recognition model to obtain a second matching training difficulty; The simulation training difficulty mode is adjusted according to the second matching training difficulty.
2. The method according to claim 1, characterized in that Also includes: Get the simulation training task number and simulation training difficulty mode number; The second matching training difficulty, the simulation training task number, the simulation training difficulty mode number, the target user physiological signal timing information, the target user voice signal timing information and the target user expression image timing information are associated and stored in the training difficulty level calibration database.
3. The method according to claim 2, characterized in that Based on the training difficulty level calibration database, the training difficulty of the target user physiological signal timing information, the target user voice signal timing information, and the target user expression image timing information is calibrated to obtain a first matching training difficulty, including: Input the simulation training task number and the simulation training difficulty mode number into the training difficulty level calibration database, and extract the first recorded physiological signal timing information, the first recorded voice signal timing information, the first recorded expression image timing information and the first recorded training difficulty; Calculating a first similarity coefficient between the first recorded physiological signal timing information and the target user physiological signal timing information; Calculating a second similarity coefficient between the target user voice signal timing information and the first recorded voice signal timing information; Calculating a third similarity coefficient between the first recorded expression image timing information and the target user expression image timing information; When the first similarity coefficient is greater than or equal to a first similarity threshold, the second similarity coefficient is greater than or equal to a second similarity threshold, and the third similarity coefficient is greater than or equal to a third similarity threshold, the first record training difficulty is set as the first matching training difficulty.
4. The method according to claim 2, characterized in that The emotion recognition model includes a physiological signal pre-processing network, a voice signal pre-processing network, an expression image pre-processing network and a post-emotion recognition fitting network. When the first matching training difficulty is empty, the target user's physiological signal timing information, the target user's voice signal timing information and the target user's expression image timing information are processed through the emotion recognition model to obtain a second matching training difficulty, including: Processing the target user's physiological signal timing information through the physiological signal pre-processing network to obtain the physiological signal matching training difficulty; Processing the target user's voice signal timing information through the voice signal pre-processing network to obtain the voice signal matching training difficulty; Processing the target user's expression image time sequence information through the expression image pre-processing network to obtain expression matching training difficulty; The physiological signal matching training difficulty, the speech signal matching training difficulty and the expression matching training difficulty are input into the post-emotion recognition fitting network, and the second matching training difficulty is output.
5. The method according to claim 4, characterized in that The steps of constructing the physiological signal pre-processing network include: Collecting historical driving training data in the simulation training task number and the simulation training difficulty mode number, wherein the historical driving training data includes the first monitoring physiological signal timing information and qualified training duration record data, and the qualified training duration record data refers to the training duration to pass after triggering the first monitoring physiological signal timing information; Setting a training difficulty mapping table, mapping the qualified training duration record data, and obtaining a physiological signal identification training difficulty; Taking the physiological signal identification training difficulty as supervision and taking the first monitored physiological signal time series information as input, training the physiological signal pre-processing network; The speech signal pre-processing network and the expression image pre-processing network are constructed in the same steps as the physiological signal pre-processing network.
6. The method according to claim 5, characterized in that Setting a training difficulty mapping table, mapping the qualified training duration record data, and obtaining the physiological signal identification training difficulty, including: Obtaining a plurality of first monitoring physiological signal time series information and a plurality of qualified training duration record data corresponding to each other; Performing pairwise similarity calculation on the plurality of first monitored physiological signal time series information to obtain a first monitored physiological signal similarity coefficient set; Clustering the plurality of first monitoring physiological signal time series information according to a similarity coefficient threshold value and in combination with a first monitoring physiological signal similarity coefficient set to obtain a plurality of clusters of first monitoring physiological signal time series information; According to the multiple clusters of first monitored physiological signal timing information, the plurality of qualified training duration record data are grouped to obtain multiple groups of qualified training duration record data; Traversing the plurality of groups of qualified training duration record data to perform mode value calculation to obtain a plurality of qualified training duration identification data; Traversing the multiple clusters of first monitoring physiological signal time series information and randomly extracting single data respectively, to obtain multiple first monitoring physiological signal time series information; Based on the training difficulty mapping table, the plurality of qualified training duration identification data are processed respectively to obtain the physiological signal identification training difficulty.
7. The method according to claim 1, characterized in that Adjusting the simulation training difficulty mode according to the second matching training difficulty includes: Obtain the sequence of operation steps for the simulation training task number; Dividing the operation step sequence into individual operation steps to obtain a first simulation training difficulty mode; Dividing the operation step sequence into two operation steps to obtain a second simulation training difficulty mode; until the operation step sequence is divided into K operation steps to obtain a Kth simulation training difficulty mode, wherein K is equal to the total number of operation steps; The first simulation training difficulty mode, the second simulation training difficulty mode, up to the Kth simulation training difficulty mode are added into the simulation training difficulty mode.
8. A dynamic constraint system for simulation training difficulty based on emotion recognition, characterized in that: The system comprises: An expression image timing information acquisition module, which is used to collect the target user's physiological signal timing information and the target user's voice signal timing information through wearable sensors when the target user enters the driving simulation training area to start training, and collect the target user's expression image timing information through the simulation training area camera; A calibration database acquisition module, wherein the calibration database acquisition module is used to obtain a training difficulty level calibration database; A training difficulty calibration module, the training difficulty calibration module is used to calibrate the training difficulty of the target user's physiological signal timing information, the target user's voice signal timing information and the target user's expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty; A second matching training difficulty acquisition module, wherein the second matching training difficulty acquisition module is used to process the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's expression image timing information through an emotion recognition model to obtain a second matching training difficulty when the first matching training difficulty is empty or not unique; A difficulty mode adjustment module is used to adjust the simulation training difficulty mode according to the second matching training difficulty.
Citation Information
Patent Citations
Game regulation and control system and method based on multi-physiological-parameter emotion estimation
CN106730812A
Relationship enhancement method and system based on EEG (electroencephalogram) multi-person cooperative regulation and control
CN116301309A
Electronic game difficulty dynamic adjustment method based on electroencephalogram signals
CN116603229A
Game difficulty adjusting method, system and equipment based on emotion recognition and medium
CN116943226A
Cognitive emotion interaction method and system for ADHD co-affected emotional disorder
CN118072953A