Emotion recognition-based simulation training difficulty dynamic constraint method and system

By collecting and analyzing the driver's physiological, voice, and facial expression information, and utilizing a training difficulty database and emotion recognition model, the simulation training difficulty is dynamically adjusted. This solves the problem that the training difficulty in existing systems cannot meet user needs, and achieves personalized and efficient simulation training results.

CN120014907BActive Publication Date: 2025-11-18WUHAN FUTURE MIRAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510037504.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-11-18
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing driving simulation training systems cannot flexibly adjust the training difficulty according to user needs, resulting in poor individualization and affecting learning outcomes and training experience.

Method used

By collecting physiological signals, voice signals, and temporal information of facial expressions from target users, and utilizing a training difficulty level calibration database and emotion recognition model, the simulation training difficulty is dynamically adjusted to adapt to the individual's skill level and emotional state.

Benefits of technology

It enables personalized and adaptive adjustments to driving simulation training, improving the relevance and effectiveness of training and enhancing the user's learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014907B_ABST
    Figure CN120014907B_ABST
Patent Text Reader

Abstract

The application discloses a simulation training difficulty dynamic constraint method and system based on emotion recognition, and relates to the technical field of human-computer interaction. The method comprises the following steps: when a target user enters a driving simulation training area, physiological signals, voice signals and expression image information of the target user are collected through a wearable sensor, and the signals are trained for difficulty calibration based on a training difficulty level calibration database to obtain a first matched training difficulty. If the first matched training difficulty is empty or not unique, the signals are processed through an emotion recognition model to obtain a second matched training difficulty, and the difficulty mode of the simulation training is adjusted accordingly. The technical problem that the driving simulation training difficulty cannot be adaptively adjusted according to user demand, resulting in poor individualization of the simulation training, is solved, and the technical effect of improving the individualization and adaptive adjustment capability of the driving simulation training is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technology, and in particular to a method and system for dynamic constraint of simulation training difficulty based on emotion recognition. Background Technology

[0002] Existing driving simulation training systems are widely used in driver training, traffic safety education, and military training. Most current simulation training systems employ fixed training difficulty and preset task scenarios, lacking consideration for the individual needs and skill levels of drivers. The training difficulty of existing systems is usually adjusted based on a uniform standard, rather than flexibly adapting to each user's actual needs, progress, and learning situation. A fixed pattern may result in some drivers encountering tasks that are too simple or too complex during training, affecting learning effectiveness and the training experience. Therefore, how to flexibly adjust the training difficulty according to the driver's skill level, learning progress, and actual needs has become a key issue in improving the personalization and effectiveness of simulation training.

[0003] Currently, there is a technical problem with related technologies: the difficulty of driving simulation training cannot be adaptively adjusted according to user needs, resulting in a poor degree of individualization in simulation training. Summary of the Invention

[0004] This application provides a driving skills training method based on the construction of diversified virtual test sites, which is used to solve the technical problem that the difficulty of existing driving simulation training cannot be adaptively adjusted according to user needs, resulting in poor individualization of simulation training.

[0005] This application provides a method for dynamically constraining the training difficulty of simulations based on emotion recognition, including:

[0006] When a target user enters the driving simulation training area to begin training, wearable sensors collect the target user's physiological signal timing information and voice signal timing information, while a camera in the simulation training area collects the target user's facial expression image timing information; a training difficulty level calibration database is obtained; based on the training difficulty level calibration database, the target user's physiological signal timing information, voice signal timing information, and facial expression image timing information are calibrated to obtain a first matching training difficulty; when the first matching training difficulty is empty or not unique, an emotion recognition model is used to process the target user's physiological signal timing information, voice signal timing information, and facial expression image timing information to obtain a second matching training difficulty; the simulation training difficulty mode is adjusted according to the second matching training difficulty.

[0007] This application provides a dynamic constraint system for simulation training difficulty based on emotion recognition, including:

[0008] The system includes: a facial expression image timing information acquisition module, used to acquire timing information of the target user's physiological signals and speech signals through wearable sensors, and timing information of the target user's facial expression images through a camera in the simulation training area when the target user enters the driving simulation training area to begin training; a calibration database acquisition module, used to obtain a training difficulty level calibration database; a training difficulty calibration module, used to calibrate the training difficulty of the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information based on the training difficulty level calibration database to obtain a first matching training difficulty; a second matching training difficulty acquisition module, used to process the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information through an emotion recognition model when the first matching training difficulty is empty or not unique, to obtain a second matching training difficulty; and a difficulty mode adjustment module, used to adjust the simulation training difficulty mode according to the second matching training difficulty.

[0009] The proposed method and system for dynamically constraining simulation training difficulty based on emotion recognition, as described in this application, firstly collects physiological signals, speech signals, and facial expression image information of the target user when they enter the driving simulation training area using wearable sensors. These signals are then calibrated for training difficulty based on a training difficulty level calibration database to obtain a first matching training difficulty. If the first matching training difficulty is empty or not unique, these signals are processed by an emotion recognition model to obtain a second matching training difficulty. The simulation training difficulty mode is then adjusted accordingly, achieving the technical effect of improving the individualization and adaptive adjustment capabilities of driving simulation training. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating the dynamic constraint method for simulation training difficulty based on emotion recognition provided in this application embodiment;

[0012] Figure 2 This is a schematic diagram of the structure of a dynamic constraint system for simulation training difficulty based on emotion recognition provided in an embodiment of this application.

[0013] Figure labeling: 10 for facial expression image time sequence information acquisition module, 20 for calibration database acquisition module, 30 for training difficulty calibration module, 40 for second matching training difficulty acquisition module, and 50 for difficulty mode adjustment module. Detailed Implementation

[0014] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below.

[0015] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.

[0017] This application provides a method for dynamically constraining the training difficulty of simulations based on emotion recognition, such as... Figure 1 As shown, the method includes:

[0018] In step S100, when the target user enters the driving simulation training area to begin training, wearable sensors collect the target user's physiological signal timing information and voice signal timing information. A camera in the simulation training area also collects the target user's facial expression image timing information. Specifically, the data acquisition system enters working mode when the target user steps into the driving simulation training area and starts training. Wearable sensors, pre-placed on specific body parts, utilize biosensor technology to closely adhere to the skin, collecting physiological signal timing information including heart rate, blood pressure, and skin conductivity. These signals fluctuate depending on the driving scenario and emotional changes; for example, heart rate increases, blood pressure rises, and skin conductivity increases during periods of tension. The data is recorded dynamically throughout the process in a curve format. Simultaneously, the sensor's built-in microphone converts the user's voice signals generated by driving conditions and emotions into digital signals and records them sequentially. The user's tone, speech rate, and other acoustic characteristics reflect emotions and psychological activities. The camera in the simulated training area captures the user's facial expressions at a certain frame rate. From the initial tension during training to the fear of encountering unexpected situations, and then to the joy of completing the task, the changes in facial expressions are formed in chronological order, providing a comprehensive and solid data foundation for assessing training status and calibrating training difficulty.

[0019] Step S200: Obtain the training difficulty level calibration database. Specifically, when obtaining the training difficulty level calibration database, first determine its data source and integrate information from multiple channels, such as past driving training history records covering various road conditions and weather simulations, recording basic information of the target user, time-series information of physiological and speech signals collected by wearable sensors, time-series information of facial expression images collected by cameras, and training difficulty and results. After collection, the data is preprocessed, physiological signal anomalies are cleaned and standardized, speech semantics are identified and acoustic features are quantified, facial expression image categories are identified and frequency and duration are statistically analyzed, and then logically organized and categorized by task number and user number hierarchical index. The database architecture is designed, the structure and fields of each table are determined, such as the training task table, user information table, etc. Finally, the data is accurately entered according to the architecture to ensure completeness and consistency, constructing a database for subsequent training difficulty calibration.

[0020] In one possible implementation, a training difficulty level calibration database is obtained. Step S200 further includes step S210, obtaining the simulation training task number and the simulation training difficulty mode number. Specifically, based on the currently ongoing simulated driving training task, its corresponding simulation training task number is obtained. The number is an identifier for a specific simulated driving training scenario and process, uniquely identifying the task among many different training tasks. For example, different driving route planning, traffic scenario settings, or training objective settings will correspond to different task numbers. Simultaneously, the simulation training difficulty mode number is obtained. This number indicates the difficulty level mode used in the current training. For example, the beginner difficulty mode number may correspond to relatively simple road conditions, fewer traffic rule restrictions, and lower driving operation requirements; the intermediate difficulty mode number involves more complex road condition combinations, more traffic rule constraints, and a moderate level of driving operation complexity; the advanced difficulty mode number signifies near-realistic and highly challenging road conditions, strict traffic rules, and high-difficulty driving operations, such as highway driving simulation under adverse weather conditions.

[0021] Step S220 involves associating and storing the second matching training difficulty, the simulated training task number, the simulated training difficulty mode number, and the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information in the training difficulty level calibration database. Specifically, after obtaining the simulated training task number and the simulated training difficulty mode number, the second matching training difficulty is associated with the number, as well as the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information, and stored in the training difficulty level calibration database. During storage, the simulated training task number and the simulated training difficulty mode number are used as key indexes to integrate all relevant data under the same task number and difficulty mode number. For example, for a set of data with a specific task number "T001" and a difficulty mode number "D02", the following information is stored in the database: the time-series information of physiological signals generated by the target user during training under this task, such as heart rate curve data, blood pressure fluctuation data, and skin conductivity change sequences; the time-series information of speech signals, such as changes in tone of voice, speech speed, and specific content of speech at different driving stages; the time-series information of facial expressions, such as expressions of fear when encountering emergencies and expressions of ease when driving smoothly; and the second matching training difficulty determined by the emotion recognition model. This associative storage method allows the database to easily query all relevant data for a specific training situation based on the task number and difficulty mode number. This provides a rich and organized data foundation for subsequent training difficulty calibration, training effect evaluation, and training strategy optimization, helping to continuously improve the content and accuracy of the training difficulty level calibration database, thereby enhancing the performance and adaptability of the entire driving simulation training system.

[0022] Step S300: Based on the training difficulty level calibration database, the training difficulty of the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information is calibrated to obtain a first matching training difficulty. Specifically, based on the training difficulty level calibration database, the target user's current simulated training task number and difficulty mode number are used as search conditions to extract the corresponding first recorded physiological, speech, and facial expression image timing information and the first recorded training difficulty. The first similarity coefficient between the target user's physiological signal timing information and the first recorded physiological signal timing information is calculated, and the similarity of the changing trends of multi-dimensional indicators such as heart rate, blood pressure, and skin conductivity is comprehensively analyzed. The second similarity coefficient between the target user's speech signal timing information and the first recorded speech signal timing information is calculated, and acoustic features such as tone, speech rate, and emotional vocabulary are analyzed. The third similarity coefficient between the first recorded facial expression image timing information and the target user's facial expression image timing information is calculated, and facial muscle movements and expression duration are compared using image recognition and expression analysis algorithms. When the first similarity coefficient, the second similarity coefficient, and the third similarity coefficient are greater than or equal to their respective similarity thresholds, the training difficulty of the first record is determined as the first matching training difficulty for the target user. This can leverage past data experience to provide a reasonable basis for setting the initial difficulty for training, improve the targeting and effectiveness of training, and facilitate personalized arrangements.

[0023] In one possible implementation, the training difficulty is calibrated based on the training difficulty level calibration database for the target user's physiological signal timing information, the target user's speech signal timing information, and the target user's facial expression image timing information to obtain a first matching training difficulty. Step S300 further includes step S310, where the simulated training task number and the simulated training difficulty mode number are input into the training difficulty level calibration database, and the first recorded physiological signal timing information, the first recorded speech signal timing information, the first recorded facial expression image timing information, and the first recorded training difficulty are extracted. Specifically, the simulated training task number and the simulated training difficulty mode number corresponding to the target user's current simulated training are input into the training difficulty level calibration database. The database performs retrieval and extraction operations in the data storage based on the two key numbers. Specifically, the system will extract the first recorded physiological signal time-series information that matches the task number and difficulty mode number. This information includes detailed data on the changes over time in various physiological signals such as heart rate, blood pressure, and skin conductivity generated by users who have previously participated in the same training. Simultaneously, it will extract the first recorded speech signal time-series information, which includes information reflecting the user's language expression characteristics, such as the speech intonation sequence, speech rate curve, and speech content. It will also extract the first recorded facial expression image time-series information, i.e., data on the time distribution and duration of facial expression changes, such as smiles, frowns, and surprise, recorded by image acquisition devices at different training stages. Finally, it will extract the corresponding first recorded training difficulty, the value of which is determined based on a comprehensive evaluation of previous training tasks and difficulty modes.

[0024] Step S320: Calculate the first similarity coefficient between the time-series information of the first recorded physiological signal and the time-series information of the target user's physiological signal. Specifically, after successfully extracting the aforementioned types of first recorded information, the first similarity coefficient is calculated. Since physiological signals contain multiple different indicators, corresponding distribution weights are assigned to different indicators. For example, indicators that are more sensitive to emotional responses and play an important indicative role in driving simulation training, such as heart rate and skin conductivity, are assigned relatively high weights; while indicators that indirectly reflect emotions, such as body temperature, are assigned lower weights. After determining the weights, the similarity of each physiological signal indicator is calculated based on the weights. Taking heart rate as an example, the heart rate change curve in the time-series information of the first recorded physiological signal is compared with the heart rate change curve in the time-series information of the target user's physiological signal, analyzing the similarity between the two in terms of average heart rate, fluctuation amplitude, peak occurrence time, and change trend; for blood pressure signals, the similarity of their numerical magnitude and the consistency of their fluctuation period are also examined; for skin conductivity, the similarity of its change response when switching between different training scenarios is considered. By comprehensively calculating each indicator according to its weight, a first similarity coefficient is finally obtained, which can accurately reflect the similarity between the timing information of the first recorded physiological signal and the timing information of the target user's physiological signal.

[0025] Step S330: Calculate the second similarity coefficient between the timing information of the target user's speech signal and the timing information of the first recorded speech signal. Specifically, the second similarity coefficient is then calculated, representing the similarity between the timing information of the target user's speech signal and the timing information of the first recorded speech signal. The calculation process analyzes several key features of the speech signals. Regarding intonation, the rise and fall patterns of intonation are compared throughout the training process. For example, whether intonation increases in complex driving scenarios, and whether the magnitude and frequency of the increase are similar. Regarding speech rate, the rhythm of changes in speed at different training stages is examined, such as whether speech rate increases during periods of high driving tension, and whether the degree of increase is similar. Semantic analysis of the speech content is also performed, statistically analyzing the frequency of specific emotional words or words related to driving operations. Through comprehensive comparison and quantitative analysis of multiple features of the speech signals, including intonation, speech rate, and semantics, a second similarity coefficient that accurately represents the degree of similarity between the two is obtained.

[0026] Step S340: Calculate the third similarity coefficient between the temporal information of the first recorded facial expression image and the temporal information of the target user's facial expression image. Specifically, the third similarity coefficient is calculated to represent the similarity between the temporal information of the first recorded facial expression image and the temporal information of the target user's facial expression image. Using image recognition technology and facial expression analysis algorithms, a detailed comparison of facial muscle movement features in the facial expression images is first performed. For example, when facing sudden traffic situations, it is observed whether similar frowning muscle movements are observed in both images, and whether the depth, duration, and positional distribution of the frown are similar. For smiling expressions, the similarity is analyzed in terms of the angle of the upturned corners of the mouth, the degree of squinting of the eyes, and the overall relaxation of the facial expression. The frequency and timing of expression transitions are also examined, such as the time point from a tense expression to a relaxed expression and whether the intermediate expression states during the transition are consistent. Through comprehensive analysis and quantitative evaluation of the multi-dimensional features of these facial expression images, a third similarity coefficient that accurately reflects the degree of similarity between the two images is finally determined.

[0027] Step S350: When the first similarity coefficient is greater than or equal to the first similarity threshold, the second similarity coefficient is greater than or equal to the second similarity threshold, and the third similarity coefficient is greater than or equal to the third similarity threshold, the training difficulty of the first record is set to the first matching training difficulty. Specifically, when the first similarity coefficient obtained through the above calculation process is greater than or equal to the first similarity threshold, the second similarity coefficient is greater than or equal to the second similarity threshold, and the third similarity coefficient is greater than or equal to the third similarity threshold, this indicates that the target user's current physiological state, language expression state, and facial expression state in the simulated training are highly similar to the user states recorded in the database under the same training tasks and difficulty modes in the past. In this case, based on the reliability of empirical data, the training difficulty of the first record is directly set to the target user's first matching training difficulty. For example, if the training difficulty of the first record is "advanced difficulty," then the target user's first matching training difficulty will also be determined as "advanced difficulty," thereby providing a reasonable and evidence-based initial training difficulty setting for subsequent simulated driving training, which helps to improve the pertinence and effectiveness of the entire training process.

[0028] Step S400: When the first matching training difficulty is empty or not unique, the target user's physiological signal timing information, speech signal timing information, and facial expression image timing information are processed through the emotion recognition model to obtain the second matching training difficulty. Specifically, when the first matching training difficulty is empty or not unique, the emotion recognition model, composed of a physiological signal preprocessing network, a speech signal preprocessing network, a facial expression image preprocessing network, and a post-emotion recognition fitting network, is used for processing. The target user's physiological signal timing information (including data such as heart rate, blood pressure, and skin conductivity that change with driving scenarios and emotions), speech signal timing information (including emotion-related features such as tone, speech rate, loudness, and speech content), and facial expression image timing information (such as facial expression changes reflecting emotions, such as frowning and smiling) are respectively input into the corresponding preprocessing networks. The physiological signal preprocessing network determines the physiological signal identification training difficulty and trains based on historical driving training data (including the first monitoring physiological signal time sequence information and qualified training duration records) under specific simulation training tasks and difficulty mode numbers. This is achieved through complex steps such as similarity calculation, clustering, grouping, and mode calculation. After inputting target user data, it outputs the physiological signal matching training difficulty. Similarly, the speech signal preprocessing network, trained on historical speech data, analyzes the target user's speech features to derive the speech signal matching training difficulty. The facial expression image preprocessing network learns and establishes mapping relationships using a large amount of historical facial expression image data, deriving the facial expression matching training difficulty based on the input target user facial expression image. Finally, the three difficulty values ​​are input into the post-emotion recognition fitting network. This network comprehensively analyzes the information of each modality and allocates weights for optimization, outputting a second matching training difficulty that comprehensively reflects the target user's emotional state, thereby improving the effectiveness and adaptability of simulated driving training.

[0029] Step S500: Adjust the simulation training difficulty mode according to the second matching training difficulty. Specifically, obtain the operation step sequence corresponding to the simulation training task number, which covers a series of actions and decision-making processes from vehicle start-up to dealing with various traffic conditions, such as start-up checks, intersection starts, speed limit adjustments, and turning operations in urban driving simulation tasks. Divide the operation step sequence into different numbers of steps to generate simulation training difficulty modes. From single-step division to obtain the first simulation training difficulty mode, such as considering the basic difficulty of starting the vehicle alone, to two-step division to form the second simulation training difficulty mode, focusing on the difficulty increase of the coherence and coordination between steps, until dividing into K steps to obtain the Kth simulation training difficulty mode, where K is the total number of operation steps. As the number of division steps increases, the difficulty gradually becomes more complex and comprehensive. Then, add these difficulty modes into a set, and select the appropriate mode from the set according to the obtained second matching training difficulty. If the target user's emotions cause the second matching training difficulty to change during training, the system will dynamically switch to the corresponding new mode to accurately adapt the difficulty mode to the user's emotions and match the training difficulty, achieve the best training effect, and avoid inappropriate difficulty affecting the training results.

[0030] In one possible implementation, the emotion recognition model includes a physiological signal preprocessing network, a speech signal preprocessing network, an facial expression image preprocessing network, and a post-emotion recognition fitting network. Specifically, the emotion recognition model is a comprehensive multi-module architecture, composed of the physiological signal preprocessing network, the speech signal preprocessing network, the facial expression image preprocessing network, and the post-emotion recognition fitting network. These four networks collaborate to accurately identify the target user's emotional state from information sources across different dimensions, thereby determining the appropriate training difficulty.

[0031] According to the second matching training difficulty, the simulation training difficulty mode is adjusted. Step S500 further includes step S510, which processes the target user's physiological signal timing information through the physiological signal preprocessing network to obtain the physiological signal matching training difficulty. Specifically, the physiological signal preprocessing network focuses on processing the target user's physiological signal timing information. The physiological signal timing information includes a variety of indicators that reflect the user's internal emotional changes, such as heart rate data, which fluctuates with the user's tension, excitement level, or fatigue state during driving simulation training. When facing complex and dangerous driving scenarios, the heart rate often increases significantly; the same is true for blood pressure data, which may rise under high pressure; skin conductivity is also a key indicator, as changes in sweat gland secretion during emotional fluctuations lead to changes in skin conductivity. This network is built based on a large amount of historical driving training data. During the construction process, historical driving training data under specific simulation training task numbers and simulation training difficulty mode numbers are first collected, which includes the first monitoring physiological signal timing information and qualified training duration record data. The qualified training duration record data refers to the training time to pass after triggering the first monitoring physiological signal timing information. Next, a training difficulty mapping table is established. This table maps qualified training duration records to determine the training difficulty of physiological signal identifiers. Specifically, multiple sets of one-to-one corresponding time-series information of the first monitored physiological signal and qualified training duration records are first acquired. Then, pairwise similarity calculations are performed on the time-series information of the first monitored physiological signal to form a set of similarity coefficients. Based on a set similarity coefficient threshold, the time-series information of the first monitored physiological signal is clustered using this set of similarity coefficients to obtain multiple clusters of time-series information of the first monitored physiological signal. The qualified training duration records are then grouped according to the clustering results to obtain multiple sets of qualified training duration records. The mode value is calculated by iterating through the combinations to obtain multiple qualified training duration identifier data. Finally, individual data are randomly extracted from each cluster of time-series information of the first monitored physiological signal, and the multiple qualified training duration identifier data are processed using the training difficulty mapping table to determine the training difficulty of the physiological signal identifiers. Then, using the training difficulty of the physiological signal identifiers as supervision and the time-series information of the first monitored physiological signal as input, a physiological signal preprocessing network is trained. When the target user's physiological signal time series information is input, the network performs in-depth analysis and feature extraction based on the pre-trained model parameters and algorithms, and finally outputs the physiological signal matching training difficulty.

[0032] Step S520 involves processing the temporal information of the target user's speech signal through the speech signal preprocessing network to obtain the speech signal matching training difficulty. Specifically, the speech signal preprocessing network primarily processes the temporal information of the target user's speech signal. During driving simulation training, the user's speech signal contains many features closely related to emotions. For example, changes in pitch can intuitively reflect the user's emotional state; when excited, tense, or thrilled, the pitch usually rises. Speech speed is also an important emotional indicator; a fast speech speed is often associated with anxiety or excitement, while a slow speech speed may suggest relaxation, contemplation, or fatigue. The choice of words and expression in the speech content are also crucial; some emotional words or specific driving-related words can further reveal the user's psychological state. The construction steps of this network are similar to those of the physiological signal preprocessing network, also based on training with a large amount of historical speech signal data. During training, features such as pitch, speech speed, and speech content in the historical speech signal data are analyzed and extracted to establish a correlation model between speech signal features and training difficulty. When the target user's speech signal timing information is input, the network will quickly capture various feature information in the speech, such as identifying keywords and emotional words in the speech content through speech recognition technology, analyzing the trend of tone change and the rhythm of speech rate change, and calculating the training difficulty of speech signal matching based on the established association model.

[0033] Step S530: The facial expression image preprocessing network processes the temporal information of the target user's facial expression images to obtain the training difficulty of facial expression matching. Specifically, the facial expression image preprocessing network processes the temporal information of the target user's facial expression images. Facial expressions are the direct external manifestation of a user's emotions. For example, a frown is often associated with confusion, anxiety, or dissatisfaction; a smile likely reflects a user's pleasant mood, satisfaction with the driving situation, or ease in handling it; wide eyes may indicate surprise or alertness; the shape and degree of opening of the mouth can also reflect different emotions, such as a tightly closed mouth indicating tension or focus, while an open mouth may indicate surprise or shouting. This network utilizes a large amount of historical facial expression image data during its construction. It deeply learns and analyzes features such as facial muscle movement characteristics, types of expressions (e.g., happiness, sadness, anger, surprise), duration of expressions, and frequency of expression transitions in facial expression images. Through machine learning algorithms, a mapping relationship between facial expression image features and training difficulty is established. When the target user's facial expression image time sequence information is input, the network can accurately identify the type and features of the expression, such as determining whether the expression is a brief surprise or a long-term tension. Based on the changes in the expression and the established mapping relationship, the training difficulty of expression matching is determined.

[0034] Step S540: The physiological signal matching training difficulty, the speech signal matching training difficulty, and the facial expression matching training difficulty are input into the post-emotion recognition fitting network, and the second matching training difficulty is output. Specifically, after the physiological signal preprocessing network, the speech signal preprocessing network, and the facial expression image preprocessing network output the physiological signal matching training difficulty, the speech signal matching training difficulty, and the facial expression matching training difficulty, respectively, the three difficulty values ​​are input into the post-emotion recognition fitting network. The post-emotion recognition fitting network comprehensively considers these three training difficulty values ​​from different modal information and uses a fitting algorithm to integrate and optimize them. For example, it assigns different weights to different modal information based on their differences in the accuracy and importance of emotional expression, and then performs weighted summation and other operations. After processing by the post-emotion recognition fitting network, the second matching training difficulty is finally output. This difficulty value integrates the emotional information reflected by the target user's physiological state, speech expression state, and facial expression state, and can more accurately determine the training difficulty suitable for the target user's current emotional state, thereby improving the effectiveness and adaptability of simulated driving training.

[0035] In one possible implementation, the physiological signal timing information of the target user is processed by the physiological signal preprocessing network to obtain physiological signal matching training difficulty. Step S510 further includes step S511, collecting historical driving training data at the simulated training task number and the simulated training difficulty mode number. The historical driving training data includes first monitoring physiological signal timing information and qualified training duration record data. The qualified training duration record data refers to the training duration from triggering the first monitoring physiological signal timing information to achieving qualification. Specifically, historical driving training data related to a specific simulated training task number and simulated training difficulty mode number is collected. The data focuses on the first monitoring physiological signal timing information, which covers detailed records of various physiological indicators such as the driver's heart rate change curve, blood pressure fluctuation data, and dynamic changes in skin conductivity over time. It also includes qualified training duration record data, which clearly records the time elapsed from triggering the first monitoring physiological signal timing information until the driver reaches the qualification standard in that training task and difficulty mode. For example, if the simulated training task is driving training in complex urban road conditions, and the simulation training difficulty mode is medium, then the collected historical data will reflect the changes in the physiological signals of different drivers in a specific situation and the differences in the time required for each of them to reach the passing grade.

[0036] Step S512: A training difficulty mapping table is set up to map the qualified training duration records to obtain the physiological signal identifier training difficulty. Specifically, a dedicated training difficulty mapping table is established. Based on the analysis and summary of a large amount of historical data, the mapping table establishes a correspondence between qualified training duration records and training difficulty. By inputting the collected qualified training duration records into this training difficulty mapping table and performing the corresponding mapping operation, the physiological signal identifier training difficulty is determined. For example, if it is found that in a certain set of historical data, the driver reaches the qualified standard in a short time after triggering a specific physiological signal, then according to the mapping table, the corresponding physiological signal identifier training difficulty will be set to a lower level; conversely, if the qualified training duration is longer, it may correspond to a higher physiological signal identifier training difficulty. The key is to accurately construct a reasonable mapping logic between qualified training duration and training difficulty to provide a reliable identifier basis for subsequent network training.

[0037] Step S513: Using the physiological signal identifier training difficulty as supervision and the first monitored physiological signal time series information as input, train the physiological signal preprocessing network. Specifically, use the obtained physiological signal identifier training difficulty as supervision information and the first monitored physiological signal time series information as input data to start training the physiological signal preprocessing network. During training, the network continuously learns the intrinsic correlation between various feature patterns in the first monitored physiological signal time series information and the corresponding physiological signal identifier training difficulty. For example, the network will figure out what training difficulty prediction value should be output under specific heart rate change trends, blood pressure fluctuation patterns, and skin conductivity change patterns. Through repeated training with a large amount of historical data, the physiological signal preprocessing network continuously optimizes its model parameters and algorithm structure, thereby being able to accurately predict the matching physiological signal training difficulty when facing new target user physiological signal time series information.

[0038] Step S514, wherein the construction steps of the speech signal preprocessing network and the facial expression image preprocessing network are the same as those of the physiological signal preprocessing network. Specifically, the construction steps of the speech signal preprocessing network and the facial expression image preprocessing network are the same as those of the physiological signal preprocessing network. For the speech signal preprocessing network, historical speech signal data under specific simulated training task numbers and simulated training difficulty mode numbers are collected, including information such as speech intonation change sequences, speech rate change curves, and speech content, as well as corresponding qualified training time record data (qualified training time refers to the training time associated with the speech signal feature until it is qualified). A training difficulty mapping table for speech signals is set up, mapping the qualified training time record data to the speech signal identifier training difficulty. The speech signal identifier training difficulty is used as supervision, and historical speech signal data is used as input to train the speech signal preprocessing network. For the facial expression image preprocessing network, historical facial expression image data under specific task and difficulty modes is first collected, such as facial muscle movement features of different expressions, expression duration, and records of qualified training time. An facial expression image training difficulty mapping table is then established to map the training difficulty of facial expression images. This table serves as supervision, and the network is trained using historical facial expression image data as input. Through the same construction steps, these three preprocessing networks can effectively transform the user's multimodal information into corresponding matching training difficulty information within their respective information processing domains. This lays a solid foundation for the entire emotion recognition model to accurately assess the target user's emotional state and determine the appropriate training difficulty.

[0039] In one possible implementation, a training difficulty mapping table is set to map the qualified training duration record data to obtain physiological signal indicators of training difficulty. Step S512 further includes step S5121, obtaining a one-to-one correspondence of several first monitoring physiological signal time sequence information and several qualified training duration record data. Specifically, a one-to-one correspondence of several first monitoring physiological signal time sequence information and several qualified training duration record data is obtained from historical driving training data resources. The first monitoring physiological signal time sequence information records in detail the physiological signal change process of different drivers under specific simulated training task numbers and simulated training difficulty mode numbers, such as how heart rate fluctuates with the switching of driving scenarios, blood pressure fluctuations when encountering complex road conditions, and subtle changes in skin conductivity during tense moments. The corresponding qualified training duration record data indicates the length of time from when the physiological signal begins to show a specific change (i.e., triggering the first monitoring physiological signal time sequence information) to when the driver finally reaches the qualified standard in the training task. For example, in a city road driving simulation training scenario with a medium difficulty level, the timing information of a driver's first monitored physiological signal showed that his heart rate exhibited a specific fluctuation curve when passing through multiple intersections and traffic conditions. The data recorded on his qualified training time showed that it took 30 minutes from the first obvious change in heart rate to the qualified training.

[0040] Step S5122 involves performing pairwise similarity calculations on the time-series information of the several first monitored physiological signals to obtain a set of similarity coefficients for the first monitored physiological signals. Specifically, pairwise similarity calculations are performed on the acquired time-series information of the several first monitored physiological signals. The calculation process is not a simple numerical comparison, but rather a comprehensive analysis of the changing trends, fluctuation amplitudes, and signal characteristics of various physiological signal indicators. For example, for heart rate signals, it is necessary not only to compare the closeness of the average heart rate values, but also to analyze whether the acceleration and deceleration patterns of the heart rate during training are consistent; for blood pressure signals, the similarity of the timing and magnitude of the peak and trough values ​​of blood pressure is analyzed; for skin conductivity signals, the compatibility of the slope of change and the fluctuation cycle at different training stages is considered. The multi-dimensional comparison results are integrated into numerical values ​​that can quantify the degree of similarity, thereby constructing a set of similarity coefficients for the first monitored physiological signals. Each similarity coefficient represents the degree of similarity between a pair of time-series information of the first monitored physiological signals.

[0041] Step S5123: Based on a similarity coefficient threshold and combined with a set of similarity coefficients for the first monitored physiological signals, cluster the time-series information of the several first monitored physiological signals to obtain multi-cluster time-series information of the first monitored physiological signals. Specifically, based on a pre-set similarity coefficient threshold and combined with a constructed set of similarity coefficients for the first monitored physiological signals, cluster the time-series information of the several first monitored physiological signals. When the similarity coefficient between two time-series information of the first monitored physiological signals is greater than or equal to this threshold, they are classified into the same category. Through the clustering process, the time-series information of the several first monitored physiological signals is orderly divided into multi-cluster time-series information of the first monitored physiological signals. The time-series information of the physiological signals within each cluster has high similarity, representing a group of drivers with similar physiological response characteristics in specific driving scenarios and difficulty modes. For example, a certain cluster of time-series information of the first monitored physiological signals corresponds to a group of drivers who experience relatively smooth heart rate fluctuations, relatively stable blood pressure changes, and small changes in skin conductivity when facing frequent starts and stops and complex traffic lights on urban roads.

[0042] Step S5124: Based on the multi-cluster first monitoring physiological signal time series information, the several qualified training duration record data are grouped to obtain multiple sets of qualified training duration record data. Specifically, after completing the clustering of the first monitoring physiological signal time series information, the several qualified training duration record data are grouped according to the obtained multi-cluster first monitoring physiological signal time series information. Since there is a one-to-one correspondence between the qualified training duration record data and the first monitoring physiological signal time series information, when the first monitoring physiological signal time series information is clustered into different clusters, its corresponding qualified training duration record data is naturally divided into different groups, thereby obtaining multiple sets of qualified training duration record data. Grouping ensures that each set of qualified training duration record data is associated with the first monitoring physiological signal time series information of a specific cluster, laying the foundation for subsequent in-depth analysis of the training duration differences of different physiological signal characteristic groups.

[0043] Step S5125: Traverse the multiple sets of grid training duration records and calculate the mode value to obtain multiple qualified training duration identifier data. Specifically, traverse each set of grid training duration records and calculate its mode value. The mode value represents the qualified training duration with the highest frequency in the driver group corresponding to the first monitoring physiological signal timing information of a specific cluster. By calculating the mode value, multiple qualified training duration identifier data are obtained. The identifier data can reflect, to a certain extent, the typical characteristics of the driver group corresponding to the data set in terms of training duration. For example, if a set of grid training duration records is [25, 28, 25, 30, 25], then its mode value is 25, and 25 minutes is determined to be the qualified training duration identifier data for that set.

[0044] Step S5126: Traverse the multiple clusters of first monitoring physiological signal time-series information and randomly extract individual data points to obtain multiple first monitoring physiological signal time-series information points. Specifically, traverse the multiple clusters of first monitoring physiological signal time-series information and randomly extract individual data points from each cluster to obtain multiple first monitoring physiological signal time-series information points. The randomly extracted data will serve as the basis for subsequent processing based on the training difficulty mapping table. They retain the characteristic representativeness of the first monitoring physiological signal time-series information of each cluster while also possessing a certain degree of randomness, providing diverse input samples for the application of the training difficulty mapping table.

[0045] Step S5127: Based on the training difficulty mapping table, the multiple qualified training duration identifier data are processed respectively to obtain the physiological signal identifier training difficulty. Specifically, based on the pre-set training difficulty mapping table, the multiple qualified training duration identifier data are processed respectively. The training difficulty mapping table is a data transformation rule system built on a large amount of historical data and professional knowledge. It can map qualified training duration identifier data to corresponding physiological signal identifier training difficulties according to the size and distribution of the data. For example, if a qualified training duration identifier data is short, it indicates that drivers in this physiological signal characteristic group can generally reach the training qualification standard relatively quickly. According to the training difficulty mapping table, the corresponding physiological signal identifier training difficulty will be set to a low level; conversely, if the qualified training duration identifier data is long, it corresponds to a higher physiological signal identifier training difficulty. Through mapping processing, the physiological signal identifier training difficulty is finally obtained, providing key supervision information and difficulty calibration basis for the subsequent training of the physiological signal preprocessing network.

[0046] In one possible implementation, the simulation training difficulty mode is adjusted according to the second matching training difficulty. Step S500 further includes step S560, obtaining the operation step sequence for the simulation training task number. Specifically, the operation step sequence corresponding to the simulation training task number is accurately obtained from the management system or data repository of the simulation driving training task. Taking a typical urban road driving simulation training task as an example, its operation step sequence covers the entire process from getting into the vehicle to finally parking and leaving the vehicle. Getting into the vehicle involves adjusting the seat to a comfortable position, fastening the seatbelt, and checking that the rearview mirrors and dashboard indicator lights are functioning properly. Starting the vehicle involves inserting the key or pressing the start button, observing the vehicle's self-check, depressing the clutch (manual transmission) or engaging the appropriate gear (automatic transmission) and releasing the handbrake. Driving procedures include starting according to traffic light signals, maintaining an appropriate speed under different road conditions, signaling in advance and observing surrounding traffic conditions when turning, confirming a safe distance and signaling with the turn signal when changing lanes, correctly judging right-of-way at intersections and yielding to pedestrians and other vehicles. Parking procedures involve signaling to pull over in advance, slowly decelerating and smoothly parking the vehicle in the designated location, engaging the handbrake, shifting to neutral (manual transmission) or P (automatic transmission), turning off the engine, and unfastening the seatbelt.

[0047] Step S570: The sequence of operation steps is divided into individual operation steps to obtain a first simulation training difficulty mode. Specifically, the obtained sequence of operation steps is divided into individual operation steps to construct the first simulation training difficulty mode. For example, for the single step of "adjusting the seat to a suitable position," this difficulty mode analyzes the student's familiarity with seat adjustment functions, such as the methods for adjusting the seat's fore-aft position, height, and backrest tilt. Slightly stiff or unresponsive seat adjustment devices may be introduced to see if the student can operate correctly to achieve a comfortable and safe driving posture. For the step of "starting according to traffic light instructions," the focus is on the student's reaction speed to changes in traffic light color, the coordination of the clutch and accelerator (manual transmission) or the switching of the brake and accelerator (automatic transmission) during starting, and whether the vehicle starts smoothly without any rolling back or sudden acceleration. This single-step segmentation method allows the difficulty of each operation step to be considered and set individually, helping novice learners to initially grasp the essentials and standards of each driving action.

[0048] Step S580: The sequence of operation steps is divided into two steps to obtain a second simulation training difficulty mode. Specifically, after the single-step division, the sequence of operation steps is further divided into two steps to obtain the second simulation training difficulty mode. Take the combination of two operation steps, "signaling the turn signal in advance and observing the surrounding traffic conditions when turning" and "confirming a safe distance and signaling with the turn signal when changing lanes," as an example. The difficulty setting here is no longer limited to the basic operation of a single step, but also takes into account the continuity and coordination between the two steps. Trainees must not only complete the turning operation accurately, but also change lanes at the appropriate time, requiring them to have a certain level of comprehensive judgment and operational coordination ability.

[0049] Step S590 continues until the operation step sequence is divided into K operation steps to obtain the Kth simulated training difficulty pattern, where K equals the total number of operation steps. Specifically, according to the segmentation logic, the number of operation steps in the segmentation is continuously increased until the operation step sequence is divided into K operation steps to obtain the Kth simulated training difficulty pattern, where K equals the total number of operation steps. As the number of segmented steps increases, the difficulty pattern gradually becomes more complex and comprehensive.

[0050] Step S5100: Add the first simulation training difficulty mode, the second simulation training difficulty mode, up to the Kth simulation training difficulty mode, into the simulation training difficulty mode set. Specifically, add the generated first simulation training difficulty mode, second simulation training difficulty mode, up to the Kth simulation training difficulty mode, into the simulation training difficulty mode set. During simulated driving training, a suitable difficulty mode can be selected from this set of difficulty modes based on the student's actual situation, such as learning progress, driving skill level, and current emotional state.

[0051] In the above text, refer to Figure 1 This paper describes in detail a dynamic constraint method for simulation training difficulty based on emotion recognition according to embodiments of the present invention. Next, reference will be made to... Figure 2 This invention describes a dynamic constraint system for simulation training difficulty based on emotion recognition, according to an embodiment of the present invention.

[0052] The emotion recognition-based dynamic constraint system for simulation training difficulty according to embodiments of the present invention addresses the technical problem that existing driving simulation training difficulty cannot be adaptively adjusted according to user needs, resulting in poor individualization of simulation training. It achieves the technical effect of improving the individualization and adaptive adjustment capabilities of driving simulation training. The emotion recognition-based dynamic constraint system for simulation training difficulty includes: an expression image temporal information acquisition module 10, a calibration database acquisition module 20, a training difficulty calibration module 30, a second matching training difficulty acquisition module 40, and a difficulty mode adjustment module 50.

[0053] The facial expression image timing information acquisition module 10 is used to acquire the timing information of the target user's physiological signals and voice signals through wearable sensors when the target user enters the driving simulation training area to start training, and to acquire the timing information of the target user's facial expression images through the camera in the simulation training area.

[0054] The calibration database acquisition module 20 is used to obtain the training difficulty level calibration database.

[0055] The training difficulty calibration module 30 is used to calibrate the training difficulty based on the training difficulty level calibration database for the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information, and to obtain the first matching training difficulty.

[0056] The second matching training difficulty acquisition module 40 is used to obtain the second matching training difficulty by processing the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information through an emotion recognition model when the first matching training difficulty is empty or not unique.

[0057] The difficulty mode adjustment module 50 is used to adjust the simulation training difficulty mode according to the second matching training difficulty.

[0058] The specific configuration of the calibration database acquisition module 20 will be described in detail below. As mentioned above, to obtain the training difficulty level calibration database, the calibration database acquisition module 20 further includes: a number acquisition unit, which is used to obtain the simulation training task number and the simulation training difficulty mode number; and an information storage unit, which is used to associate and store the second matching training difficulty, the simulation training task number, the simulation training difficulty mode number, and the target user physiological signal timing information, the target user voice signal timing information, and the target user facial expression image timing information in the training difficulty level calibration database.

[0059] The specific configuration of the training difficulty calibration module 30 will be described in detail below. As mentioned above, the training difficulty is calibrated based on the training difficulty level calibration database for the timing information of the target user's physiological signals, the timing information of the target user's speech signals, and the timing information of the target user's facial expression images to obtain a first matching training difficulty. The training difficulty calibration module 30 further includes: a timing information extraction unit, which is used to input the simulated training task number and the simulated training difficulty mode number into the training difficulty level calibration database to extract the timing information of the first recorded physiological signals, the timing information of the first recorded speech signals, the timing information of the first recorded facial expression images, and the first recorded training difficulty; and a first similarity coefficient calculation unit, which is used to calculate the similarity between the timing information of the first recorded physiological signals and the timing information of the target user's facial expression images. The system includes: a first similarity coefficient for the timing information of the user's physiological signals; a second similarity coefficient calculation unit, which calculates a second similarity coefficient between the timing information of the target user's speech signals and the timing information of the first recorded speech signals; a third similarity coefficient calculation unit, which calculates a third similarity coefficient between the timing information of the first recorded facial expression images and the timing information of the target user's facial expression images; and a first recording training difficulty setting unit, which sets the first recording training difficulty to the first matching training difficulty when the first similarity coefficient is greater than or equal to a first similarity threshold, the second similarity coefficient is greater than or equal to a second similarity threshold, and the third similarity coefficient is greater than or equal to a third similarity threshold.

[0060] The following will describe in detail the specific configuration of the difficulty mode adjustment module 50. As described above, the simulation training difficulty mode is adjusted according to the second matching training difficulty. The difficulty mode adjustment module 50 further includes: an emotion recognition model component unit, wherein the emotion recognition model includes a physiological signal preprocessing network, a speech signal preprocessing network, an facial expression image preprocessing network, and a post-emotion recognition fitting network; a physiological signal matching training difficulty acquisition unit, wherein the physiological signal matching training difficulty acquisition unit processes the target user's physiological signal temporal information through the physiological signal preprocessing network to obtain the physiological signal matching training difficulty; a speech signal matching training difficulty acquisition unit, wherein the speech signal matching training difficulty acquisition unit processes the target user's speech signal temporal information through the speech signal preprocessing network to obtain the speech signal matching training difficulty; a facial expression matching training difficulty acquisition unit, wherein the facial expression matching training difficulty acquisition unit processes the target user's facial expression image temporal information through the facial expression image preprocessing network to obtain the facial expression matching training difficulty; and a second matching training difficulty output unit, wherein the second matching training difficulty output unit inputs the physiological signal matching training difficulty, the speech signal matching training difficulty, and the facial expression matching training difficulty into the post-emotion recognition fitting network and outputs the second matching training difficulty.

[0061] The physiological signal preprocessing network processes the target user's physiological signal timing information to obtain physiological signal matching training difficulty. The physiological signal matching training difficulty acquisition unit further includes: a driving training data acquisition subunit, used to collect historical driving training data at the simulated training task number and the simulated training difficulty mode number, wherein the historical driving training data includes first monitored physiological signal timing information and qualified training duration record data, the qualified training duration record data referring to the training duration to qualified after triggering the first monitored physiological signal timing information; a recording data mapping subunit, used to set a training difficulty mapping table and map the qualified training duration record data to obtain physiological signal identifier training difficulty; a preprocessing network training subunit, used to train the physiological signal preprocessing network with the physiological signal identifier training difficulty as supervision and the first monitored physiological signal timing information as input; and a processing network construction subunit, wherein the voice signal preprocessing network and the facial expression image preprocessing network are constructed using the same steps as the physiological signal preprocessing network.

[0062] The system includes a training difficulty mapping table to map the qualified training duration recording data, obtaining physiological signals to indicate training difficulty. The data mapping subunit further comprises: a training duration recording data acquisition microunit, used to obtain a one-to-one correspondence between several first monitoring physiological signal time-series information and several qualified training duration recording data; a similarity coefficient set acquisition microunit, used to perform pairwise similarity calculations on the several first monitoring physiological signal time-series information to obtain a first monitoring physiological signal similarity coefficient set; a time-series information clustering microunit, used to cluster the several first monitoring physiological signal time-series information according to a similarity coefficient threshold and the first monitoring physiological signal similarity coefficient set to obtain multi-cluster first monitoring physiological signal time-series information; and a data recording grouping microunit. The system comprises: a data grouping micro-unit, which groups the plurality of qualified training duration record data according to the multi-cluster first monitoring physiological signal time series information to obtain multiple sets of qualified training duration record data; a mode value calculation micro-unit, which iterates through the multiple sets of qualified training duration record data to perform mode value calculation to obtain multiple qualified training duration identifier data; a first monitoring physiological signal time series information acquisition micro-unit, which iterates through the multi-cluster first monitoring physiological signal time series information and randomly extracts individual data to obtain multiple first monitoring physiological signal time series information; and a physiological signal identifier training difficulty acquisition micro-unit, which processes the multiple qualified training duration identifier data according to the training difficulty mapping table to obtain the physiological signal identifier training difficulty.

[0063] The difficulty mode adjustment module 50 further includes: an operation step sequence acquisition unit, which is used to obtain an operation step sequence of the simulated training task number; an operation step segmentation unit, which is used to segment the operation step sequence according to a single operation step to obtain a first simulated training difficulty mode; a second simulated training difficulty mode acquisition unit, which is used to segment the operation step sequence according to two operation steps to obtain a second simulated training difficulty mode; a traversal step segmentation unit, which is used to segment the operation step sequence according to K operation steps to obtain a Kth simulated training difficulty mode, where K is equal to the total number of operation steps; and a training difficulty mode addition unit, which is used to add the first simulated training difficulty mode, the second simulated training difficulty mode up to the Kth simulated training difficulty mode into the simulated training difficulty mode.

[0064] The emotion recognition-based simulation training difficulty dynamic constraint system provided in this embodiment of the invention can execute the emotion recognition-based simulation training difficulty dynamic constraint method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0065] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0066] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method for dynamically constraining the training difficulty of simulation based on emotion recognition, characterized in that, A dynamic constraint system for simulation training difficulty based on emotion recognition is applied. The system is communicatively connected to a wearable sensor, which is deployed at a preset location on the target user. The system includes: When the target user enters the driving simulation training area to start training, the system collects the target user's physiological signal timing information and voice signal timing information through wearable sensors, and collects the target user's facial expression image timing information through the camera in the simulation training area. Obtain a training difficulty level calibration database; Based on the training difficulty level calibration database, the training difficulty is calibrated for the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information to obtain the first matching training difficulty; When the first matching training difficulty is empty or not unique, the second matching training difficulty is obtained by processing the target user's physiological signal time sequence information, the target user's voice signal time sequence information and the target user's facial expression image time sequence information through the emotion recognition model. Adjust the simulation training difficulty mode according to the second matching training difficulty.

2. The method as described in claim 1, characterized in that, Also includes: Obtain the simulation training task number and the simulation training difficulty mode number; The second matching training difficulty, the simulated training task number, the simulated training difficulty mode number, the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information are associated and stored in the training difficulty level calibration database.

3. The method as described in claim 2, characterized in that, Based on the training difficulty level calibration database, the training difficulty is calibrated for the target user's physiological signal timing information, the target user's speech signal timing information, and the target user's facial expression image timing information to obtain a first matching training difficulty, including: Input the simulated training task number and the simulated training difficulty mode number into the training difficulty level calibration database, and extract the first recorded physiological signal timing information, the first recorded speech signal timing information, the first recorded facial expression image timing information, and the first recorded training difficulty. Calculate the first similarity coefficient between the first recorded physiological signal timing information and the target user's physiological signal timing information; Calculate the second similarity coefficient between the timing information of the target user's speech signal and the timing information of the first recorded speech signal; Calculate the third similarity coefficient between the temporal information of the first recorded facial expression image and the temporal information of the target user's facial expression image; When the first similarity coefficient is greater than or equal to the first similarity threshold, the second similarity coefficient is greater than or equal to the second similarity threshold, and the third similarity coefficient is greater than or equal to the third similarity threshold, the training difficulty of the first record is set to the first matching training difficulty.

4. The method as described in claim 2, characterized in that, The emotion recognition model includes a physiological signal preprocessing network, a speech signal preprocessing network, an facial expression image preprocessing network, and a post-emotion recognition fitting network. When the first matching training difficulty is empty or not unique, the emotion recognition model processes the target user's physiological signal temporal information, the target user's speech signal temporal information, and the target user's facial expression image temporal information to obtain a second matching training difficulty, including: The physiological signal preprocessing network processes the time sequence information of the target user's physiological signals to obtain the physiological signal matching training difficulty. The speech signal preprocessing network processes the temporal information of the target user's speech signal to obtain the speech signal matching training difficulty. The facial expression image preprocessing network is used to process the temporal information of the target user's facial expression image to obtain the training difficulty of facial expression matching. The training difficulty of matching the physiological signal, the training difficulty of matching the speech signal, and the training difficulty of matching the facial expression are input into the post-emotion recognition fitting network, and the second matching training difficulty is output.

5. The method as described in claim 4, characterized in that, The construction steps of the physiological signal preprocessing network include: Collect historical driving training data at the simulation training task number and the simulation training difficulty mode number, wherein the historical driving training data includes first monitoring physiological signal timing information and qualified training duration record data, wherein the qualified training duration record data refers to the training duration to qualified after triggering the first monitoring physiological signal timing information; A training difficulty mapping table is set up to map the qualified training duration records to obtain physiological signals indicating the training difficulty. The physiological signal preprocessing network is trained using the training difficulty of the physiological signal identifier as supervision and the timing information of the first monitored physiological signal as input. The construction steps of the speech signal preprocessing network and the facial expression image preprocessing network are the same as those of the physiological signal preprocessing network.

6. The method as described in claim 5, characterized in that, A training difficulty mapping table is established to map the qualified training duration records to obtain physiological signals indicating training difficulty, including: Obtain a one-to-one correspondence of several first-monitored physiological signal timing information and several qualified training duration record data; The pairwise similarity calculation is performed on the time-series information of the plurality of first monitoring physiological signals to obtain a set of similarity coefficients for the first monitoring physiological signals; Based on the similarity coefficient threshold and combined with the first monitoring physiological signal similarity coefficient set, the time series information of the several first monitoring physiological signals is clustered to obtain multi-cluster first monitoring physiological signal time series information. Based on the timing information of the first monitoring physiological signal of the multiple clusters, the several qualified training duration record data are grouped to obtain multiple sets of qualified training duration record data. The mode value is calculated by traversing the multiple sets of training duration records to obtain multiple qualified training duration identifier data; By traversing the multiple clusters of first monitoring physiological signal time-series information, a single data point is randomly extracted to obtain multiple first monitoring physiological signal time-series information; Based on the training difficulty mapping table, the multiple qualified training duration identifier data are processed respectively to obtain the physiological signal identifier training difficulty.

7. The method as described in claim 1, characterized in that, Adjust the simulation training difficulty mode according to the second matching training difficulty, including: The sequence of steps to obtain the simulation training task number; The sequence of operation steps is divided into individual operation steps to obtain the first simulation training difficulty mode; The sequence of operation steps is divided into two operation steps to obtain a second simulation training difficulty mode; The process continues until the sequence of operation steps is divided into K operation steps to obtain the Kth simulated training difficulty mode, where K equals the total number of operation steps. The first simulation training difficulty mode, the second simulation training difficulty mode, up to the Kth simulation training difficulty mode are added to the simulation training difficulty mode.

8. A dynamic constraint system for simulation training difficulty based on emotion recognition, characterized in that, The system includes: The facial expression image timing information acquisition module is used to acquire the timing information of the target user's physiological signals and voice signals through wearable sensors when the target user enters the driving simulation training area to start training, and to acquire the timing information of the target user's facial expression images through the camera in the simulation training area. A calibration database acquisition module is used to obtain a training difficulty level calibration database; The training difficulty calibration module is used to calibrate the training difficulty based on the training difficulty level calibration database for the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information, to obtain a first matching training difficulty. The second matching training difficulty acquisition module is used to obtain the second matching training difficulty by processing the target user's physiological signal timing information, the target user's voice signal timing information, and the target user's facial expression image timing information through an emotion recognition model when the first matching training difficulty is empty or not unique. The difficulty mode adjustment module is used to adjust the simulation training difficulty mode according to the second matching training difficulty.

Citation Information

Patent Citations

  • Game regulation and control system and method based on multi-physiological-parameter emotion estimation

    CN106730812A

  • Method and apparatus for dynamically adjusting game or other simulation difficulty

    US20080266250A1