Dynamic calibration method and system of vehicle-mounted emotion recognition system
Through multi-source data fusion and dynamic weight allocation, the individual and environmental adaptability problems of the on-board emotion recognition system are solved, and the stability and robustness of the on-board emotion recognition are improved, and it is suitable for on-board infotainment systems and advanced driving assistance systems.
Patent Information
- Application Number
- CN202510566039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
The existing on-board emotion recognition system lacks individual adaptability and poor environmental adaptability, the calibration process is static and lacks real-time feedback mechanism, resulting in low recognition accuracy and inability to continuously optimize.
By acquiring multi-source data, including facial images, speech signals and physiological signals, preprocessing and feature extraction, dynamically allocating weights with environmental credibility evaluation functions, multimodal fusion vectors are constructed for emotion recognition, and intelligent arbitration and closed-loop feedback optimization are used to use improved D-S evidence theory.
It has achieved the stability and robustness of emotional recognition in complex driving environments, adapted to individual differences and environmental changes, and continuously optimized the emotion recognition model, which is suitable for in-vehicle infotainment systems and advanced driving assistance systems.
Smart Images

Figure CN120448917A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intersection of smart cars and artificial intelligence, and specifically, to a dynamic calibration method and system for an in-vehicle emotion recognition system, and more specifically, to a dynamic calibration method and system for an in-vehicle emotion recognition system based on multimodal data fusion. Background Art
[0002] In the existing technology, in-vehicle emotion recognition systems mostly adopt emotion classification methods based on static models, that is, model training and calibration are completed through pre-collected data before system deployment, and the model is solidified on the vehicle side for subsequent recognition of user emotional states.
[0003] Although the above methods can achieve basic recognition of emotional states in some scenarios, they still have the following significant shortcomings in practical applications:
[0004] 1. Lack of individual adaptability: Existing systems generally rely on general models, ignoring individual differences in emotional expression among different users, resulting in low recognition accuracy;
[0005] 2. Poor environmental adaptability: Factors such as lighting, noise, and driving conditions in the vehicle environment are highly dynamic, and traditional static calibration models cannot adapt to these changes, making recognition failures or misjudgments more likely to occur.
[0006] 3. The calibration process is static and one-time: Traditional emotion recognition model calibration is usually completed only during the initial deployment phase. There is a lack of dynamic correction mechanisms during operation, making it impossible to continuously optimize model performance.
[0007] 4. Lack of real-time feedback mechanism: Existing systems are mostly closed-loop processes of "input-recognition-output". They lack a mechanism to obtain feedback from user interactions and use it for model correction, which limits the long-term availability of the system.
[0008] Patent document CN109017797B (application number: 201810942449.8) discloses a method for identifying driver emotions, which includes the following steps: collecting attribute data that can reflect the driver's emotions during a vehicle journey, and dividing the collected attribute data into multiple data segments at a first time interval; for all data segments, obtaining the data class to which each data segment belongs through a clustering algorithm based only on its own data; assigning a corresponding emotion label to each data segment in the obtained one or more data classes, so that the data segments in each data class correspond to an emotion category; performing machine learning on all data segments that have been assigned emotion labels to obtain an emotion recognition model; and collecting attribute data in real time during vehicle driving, and using the emotion recognition model to identify the driver's emotions at the time corresponding to the real-time attribute data. Summary of the Invention
[0009] In view of the defects in the prior art, the purpose of the present invention is to provide a dynamic calibration method and system for an in-vehicle emotion recognition system.
[0010] A dynamic calibration method for an in-vehicle emotion recognition system provided by the present invention includes:
[0011] Step S1: acquiring multi-source data, including facial images, voice signals, and physiological signals; and preprocessing the acquired multi-source data to obtain preprocessed multi-source data;
[0012] Step S2: performing feature extraction based on the preprocessed facial image, speech signal, and physiological signal to obtain a facial expression feature vector, an audio feature vector, and a physiological state feature vector;
[0013] Step S3: Evaluate the current environment credibility based on the environment credibility evaluation function;
[0014] Step S4: Dynamically assign weights to multi-source data based on the current environmental credibility and real-time scenario;
[0015] Step S5: Based on the dynamically assigned weights of the multi-source data, a multimodal fusion vector is constructed according to the facial expression feature vector, the audio feature vector, and the physiological state feature vector, and emotion recognition is performed using the constructed emotion recognition model based on the multimodal fusion vector.
[0016] Preferably, the step S1 includes:
[0017] Step S1.1: Capturing a facial image of a target subject through an in-vehicle camera, and performing preprocessing on the captured facial image including image denoising, alignment, and cropping to obtain a preprocessed facial image;
[0018] Step S1.2: collecting the target subject's voice information through the microphone array, and performing pre-processing including noise reduction and frame segmentation on the collected voice information to obtain pre-processed voice information;
[0019] Step S1.3: collecting physiological information of the target subject through the somatosensory monitoring device, and performing preprocessing including denoising and normalization on the collected physiological information to obtain preprocessed physiological information;
[0020] The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin electrical response and body surface temperature.
[0021] Preferably, step S2 includes:
[0022] Step S2.1: performing facial key point recognition based on the preprocessed facial image, and extracting facial expression feature vectors based on the recognized facial key points;
[0023] Step S2.2: Extracting audio feature vectors that can reflect emotional state based on the preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators;
[0024] Step S2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
[0025] Preferably, the step S3 includes: constructing an environment credibility evaluation function, and evaluating the credibility of the current environment using the constructed environment credibility evaluation function;
[0026]
[0027] Where L represents the light intensity; R valid Indicates the effective frame rate.
[0028] Preferably, step S4 includes:
[0029] Step S4.1: Adjust the weights of multi-source data based on the assessed credibility of the current environment and the real-time scenario;
[0030] Step S4.2: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights respectively; when the confidence of a multi-source data weight is less than the preset value, the driving scenario correction coefficient is introduced to correct the multi-source data weight to obtain the corrected multi-source data weight.
[0031] Preferably, the weights of the multi-source data in step S4.1 include: face weight W f , speech weight W s and physiological information weight W p ;
[0032] The facial weight W f include:
[0033]
[0034] Among them, W f (0) Assign weight to the base face; k1 is the vehicle speed influence coefficient; Q v,th is the visual quality critical value; h represents the vehicle speed threshold; Z is the normalization factor;
[0035] The speech weight W s include:
[0036]
[0037] Among them, W s (0) Assign weight to basic voice; SNR0 is the reference signal-to-noise ratio; SNR is the signal-to-noise ratio; η∈[0,1] represents the degree of window closing;
[0038] The physiological information weight W p include:
[0039]
[0040] Among them, W p (0) Assign weights to basic physiological information; R C is the road complexity index; γ(R C )={0.05R C ,R C ≤4; 0.2+0.1(R C -4), R C >4.
[0041] Preferably, the method further comprises: establishing an emotion-behavior association rule base, and when the target object is identified as having an emotion of "anger" and a rapid acceleration behavior occurs within a preset time, the corresponding emotion confidence is increased by ΔC;
[0042]
[0043] Among them, γ is the experience adjustment coefficient; A indicates that rapid acceleration occurs within the preset time, and E indicates that the target object’s emotion is recognized as “anger”.
[0044] A dynamic calibration system for an in-vehicle emotion recognition system provided by the present invention includes:
[0045] Module M1: Acquire multi-source data, including facial images, voice signals, and physiological signals; and pre-process the acquired multi-source data to obtain pre-processed multi-source data;
[0046] Module M2: Extract features based on the preprocessed facial image, speech signal, and physiological signal to obtain facial expression feature vectors, audio feature vectors, and physiological state feature vectors;
[0047] Module M3: Evaluate the current environment credibility based on the environment credibility evaluation function;
[0048] Module M4: Dynamically assign weights to multi-source data based on the current environmental credibility and real-time scenarios;
[0049] Module M5: Based on the dynamically assigned weights of multi-source data, a multimodal fusion vector is constructed according to the facial expression feature vector, audio feature vector, and physiological state feature vector. Emotion recognition is performed using the constructed emotion recognition model based on the multimodal fusion vector.
[0050] Preferably, the module M1 includes:
[0051] Module M1.1: Capture the target object's facial image through the in-vehicle camera and perform preprocessing on the captured facial image, including image denoising, alignment, and cropping, to obtain the preprocessed facial image;
[0052] Module M1.2: Collects the target subject's voice information through a microphone array and performs pre-processing on the collected voice information, including noise reduction and frame segmentation, to obtain pre-processed voice information;
[0053] Module M1.3: Collect physiological information of the target subject through the somatosensory monitoring device, and perform preprocessing including denoising and normalization on the collected physiological information to obtain preprocessed physiological information;
[0054] The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin galvanic response and body surface temperature;
[0055] The module M2 includes:
[0056] Module M2.1: Recognize facial key points based on preprocessed facial images and extract facial expression feature vectors based on the recognized facial key points;
[0057] Module M2.2: Extract audio feature vectors that reflect emotional state based on preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators;
[0058] Module M2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
[0059] Preferably, the module M3 includes: constructing an environment credibility evaluation function, and evaluating the credibility of the current environment using the constructed environment credibility evaluation function;
[0060]
[0061] Where L represents the light intensity; R valid Indicates the effective frame rate;
[0062] The module M4 includes:
[0063] Module M4.1: Adjust the weights of multi-source data based on the assessed credibility of the current environment and the real-time scenario;
[0064] Module M4.2: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights. When the confidence of a multi-source data weight is less than the preset value, the driving scenario correction coefficient is introduced to correct the multi-source data weight to obtain the corrected multi-source data weight.
[0065] The weights of the multi-source data in the module M4.1 include: face weight W f , speech weight W s and physiological information weight W p ;
[0066] The facial weight W f include:
[0067]
[0068] Among them, W f (0) Assign weight to the base face; k1 is the vehicle speed influence coefficient; Q v,th is the visual quality critical value; h represents the vehicle speed threshold; Z is the normalization factor;
[0069] The speech weight W s include:
[0070]
[0071] Among them, W s (0) Assign weight to basic voice; SNR0 is the reference signal-to-noise ratio; SNR is the signal-to-noise ratio; η∈[0,1] represents the degree of window closing;
[0072] The physiological information weight W p include:
[0073]
[0074] Among them, W p (0) Assign weights to basic physiological information; R C is the road complexity index; γ(R C )={0.05R C ,R C ≤4; 0.2+0.1(R C -4), R C >4.
[0075] Compared with the prior art, the present invention has the following beneficial effects:
[0076] 1. This invention addresses the challenges of static calibration models' adaptability in dynamic driving environments, conflicts and misjudgments in multimodal sensor data, and error accumulation due to a lack of closed-loop optimization through three key mechanisms: dynamic calibration of environmental perception, intelligent arbitration of conflicting evidence, and feedback-driven continuous evolution.
[0077] 2. This invention continuously optimizes the emotion recognition model by collecting real-time driving environment data, driver biometrics, and vehicle status information, combined with a dynamic weight allocation algorithm and a closed-loop feedback mechanism.
[0078] 3. This invention is particularly suitable for personalized emotion monitoring in complex driving scenarios and can be integrated into in-vehicle infotainment systems (IVI) or advanced driver assistance systems (ADAS);
[0079] 4. The present invention detects emotional states through multimodal data fusion, thereby improving the stability and robustness of emotion recognition, and can maintain a high recognition rate even under complex cabin environment changes such as lighting changes and noise interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0081] Figure 1 Flowchart of the dynamic calibration method for the in-vehicle emotion recognition system.
[0082] Figure 2 Flowchart of the human factors engineering testing method for the emotional cockpit.
[0083] Figure 3 Flowchart of the automatic adaptation method for emotional response strategies in multiple climate zones.
[0084] Figure 4 Flowchart of the subscription service approach for the Emotion Cockpit feature. DETAILED DESCRIPTION
[0085] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0086] Example 1
[0087] According to the present invention, a dynamic calibration method for a vehicle-mounted emotion recognition system is provided. Figure 1 Shown, including:
[0088] Step S1: acquiring multi-source data, including facial images, voice signals, and physiological signals; and preprocessing the acquired multi-source data to obtain preprocessed multi-source data;
[0089] Step S2: performing feature extraction based on the preprocessed facial image, speech signal, and physiological signal to obtain a facial expression feature vector, an audio feature vector, and a physiological state feature vector;
[0090] Step S3: Evaluate the current environment credibility based on the environment credibility evaluation function;
[0091] Step S4: Dynamically assign weights to multi-source data based on the current environmental credibility and real-time scenario;
[0092] Step S5: Based on the dynamically assigned weights of the multi-source data, a multimodal fusion vector is constructed according to the facial expression feature vector, the audio feature vector, and the physiological state feature vector, and emotion recognition is performed using the constructed emotion recognition model based on the multimodal fusion vector.
[0093] Specifically, step S1 includes:
[0094] Step S1.1: Capturing a facial image of a target subject through an in-vehicle camera, and performing preprocessing on the captured facial image including image denoising, alignment, and cropping to obtain a preprocessed facial image;
[0095] Step S1.2: collecting the target subject's voice information through the microphone array, and performing pre-processing including noise reduction and frame segmentation on the collected voice information to obtain pre-processed voice information;
[0096] Step S1.3: collecting physiological information of the target subject through the somatosensory monitoring device, and performing preprocessing including denoising and normalization on the collected physiological information to obtain preprocessed physiological information;
[0097] The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin electrical response and body surface temperature.
[0098] Specifically, step S2 includes:
[0099] Step S2.1: performing facial key point recognition based on the preprocessed facial image, and extracting facial expression feature vectors based on the recognized facial key points;
[0100] Step S2.2: Extracting audio feature vectors that can reflect emotional state based on the preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators;
[0101] Step S2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
[0102] Specifically, the step S3 includes: constructing an environment credibility evaluation function, and using the constructed environment credibility evaluation function to evaluate the credibility of the current environment;
[0103]
[0104] Where L represents the light intensity; R valid Indicates the effective frame rate.
[0105] Specifically, step S4 includes:
[0106] Step S4.1: Adjust the weights of multi-source data based on the assessed credibility of the current environment and the real-time scenario;
[0107] Step S4.2: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights respectively; when the confidence of a multi-source data weight is less than the preset value, the driving scenario correction coefficient is introduced to correct the multi-source data weight to obtain the corrected multi-source data weight.
[0108] Specifically, the weights of the multi-source data in step S4.1 include: face weight W f , speech weight W s and physiological information weight W p ;W f +W s +W p =1(0≤W f ,W s ,W p ≤1);
[0109] The facial weight W f include:
[0110]
[0111] Among them, W f (0) Assign weight to the base face; k1 is the vehicle speed influence coefficient; Q v,th is the visual quality critical value; h represents the vehicle speed threshold; Z is the normalization factor; in this embodiment, k1=0.1; Q v,th =0.6;
[0112] The speech weight W s include:
[0113]
[0114] Among them, W s (0) Assign weight to basic voice; SNR0 is the reference signal-to-noise ratio; SNR is the signal-to-noise ratio; η∈[0,1] represents the degree of window closing, where 0 = fully open; 1 = fully closed; in this embodiment, SNR0 = 15dB;
[0115] The physiological information weight W p include:
[0116]
[0117] Among them, W p (0) Assign weights to basic physiological information; when deceleration a>0.4g a>0.4g is detected, the weight increment of the physiological signal is: R C is the road complexity index; γ(R C )={0.05R C ,R C ≤4; 0.2+0.1(R C -4), R C >4.
[0118] In this embodiment, the basic weight distribution W f =0.5,W s =0.3,W p =0.2;
[0119] Specifically, the method further includes: establishing an emotion-behavior association rule base, and when the target object's emotion is identified as "anger" and rapid acceleration occurs within a preset time, the corresponding emotion confidence is increased by ΔC;
[0120]
[0121] Among them, γ is the experience adjustment coefficient; A indicates that rapid acceleration occurs within the preset time, and E indicates that the target object’s emotion is recognized as “anger”.
[0122] The present invention also provides a dynamic calibration system for a vehicle-mounted emotion recognition system. The dynamic calibration system for the vehicle-mounted emotion recognition system can be implemented by executing the process steps of the dynamic calibration method for the vehicle-mounted emotion recognition system. That is, those skilled in the art can understand the dynamic calibration method for the vehicle-mounted emotion recognition system as a preferred implementation of the dynamic calibration system for the vehicle-mounted emotion recognition system.
[0123] Example 2
[0124] Example 2 is a preferred example of Example 1
[0125] A dynamic calibration method for an in-vehicle emotion recognition system provided by the present invention includes:
[0126] Step 1: Acquire multi-source data. Collect multimodal data required for driver emotion recognition using relevant hardware, including environmental sensors, driver biosensors, and vehicle status interfaces.
[0127] Step 2: Dynamic weight allocation, which is responsible for dynamically allocating sentiment weight contributions to different data sources;
[0128] Step 3: Perform weight fusion based on the dynamically assigned emotion weights from different data sources, and perform emotion recognition using the constructed model based on multimodal fusion;
[0129] Step 4: Closed-loop feedback optimization: the emotion category automatically adjusts the cabin environment parameters according to the environment adjustment strategy.
[0130] Specifically, the step 1 includes: deploying data collection equipment to collect multimodal data related to the driver's emotions;
[0131] In this embodiment, the driver's facial image is collected by the in-car camera; the voice signal is collected by the microphone array; the physiological signal is obtained by the steering wheel grip sensor and the seat heart rate sensor; and the vehicle status data is read through the OBD-II interface;
[0132] The driver's facial image captured by the in-car camera has a resolution of ≥1280×720 and a frame rate of 30fps. The voice signal collected by the microphone array has a sampling rate of 16kHz and a signal-to-noise ratio of ≥60dB. The steering wheel grip force sensor has a range of 0-100N. The seat heart rate sensor has an accuracy of ±2bpm. Vehicle status data, including speed, acceleration, and steering angle, is read through the OBD-II interface.
[0133] The step 2 comprises the following steps:
[0134] Step 2.1: Construct an environment credibility evaluation function;
[0135]
[0136] Where L represents the light intensity / lux, R valid Indicates the effective frame rate.
[0137] Step 2.2: Adjust the fusion weights according to the real-time scenario.
[0138] In this embodiment, when Qv<0.6Qv<0.6 and the vehicle speed is >80km / h, the facial weight Wf is reduced from 0.5 to 0.3, and the speech weight Ws is increased to 0.6;
[0139] When sudden braking (deceleration > 0.4g) is detected, the physiological signal weight Wp is increased by Δ = 0.2;
[0140] Step 2.3: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights respectively; when the confidence of a multi-source data weight is less than the preset value, introduce the driving scenario correction coefficient to correct the multi-source data weight to obtain the corrected multi-source data weight;
[0141] The improved DS evidence theory is used to calculate the confidence of the multi-source data weights respectively, including:
[0142] (K: conflict factor);
[0143] Among them, m new (A) represents the synthesized multi-source confidence, A: the proposition to be verified (e.g., "fatigue state"), B, C: the focal elements of evidence (e.g., "abnormal blinking frequency," "reduced heart rate variability"), m1(B): the support of evidence source 1 for proposition B (value range [0,1]), m2(C): the support of evidence source 2 for proposition C;
[0144] In this embodiment, the fatigue confidence threshold dynamic adjustment formula is:
[0145] In this embodiment, in the night scene, the basic confidence threshold of the "fatigue" emotion is adjusted from 0.7 to 0.65.
[0146] The method further includes: when it is determined that the emotion meets the preset requirements, asking the driver for confirmation through voice interaction, such as "Is the current emotion accurate?", and triggering model parameter correction if the answer is negative.
[0147] The method further includes: establishing an emotion-behavior association rule base, for example, if a rapid acceleration behavior occurs within 5 seconds after the determination of "anger", the corresponding emotion confidence is increased by 0.15.
[0148] Example 3
[0149] Example 3 is a preferred example of Example 1
[0150] According to the human factors engineering testing method of an emotional cockpit provided by the present invention, Figure 2 Shown, including:
[0151] Step A1: Acquire multi-source data, including facial images, voice signals, and physiological signals; and pre-process the acquired multi-source data to obtain pre-processed multi-source data;
[0152] Step A2: performing feature extraction based on the preprocessed facial image, speech signal, and physiological signal to obtain a facial expression feature vector, an audio feature vector, and a physiological state feature vector;
[0153] Step A3: constructing a multimodal fusion vector based on the facial expression feature vector, the audio feature vector, and the physiological state feature vector;
[0154] Step A4: Construct a joint emotion recognition model, and use the multimodal fusion vector to perform emotion classification using the constructed joint emotion recognition model;
[0155] Step A5: The AI system automatically adjusts the cabin environment parameters based on the emotion category and the environment adjustment strategy;
[0156] The joint emotion recognition model uses a method of multimodal feature fusion and deep learning classification to identify the emotional state of the user in the car and achieve the purpose of intelligent dynamic adjustment of the cabin environment accordingly.
[0157] Specifically, the step A1 includes:
[0158] Step A1.1: Capture a facial image of the target subject using an in-vehicle camera, and perform preprocessing on the captured facial image, including image denoising, alignment, and cropping, to obtain a preprocessed facial image;
[0159] Step A1.2: collecting the target subject's voice information through the microphone array, and performing pre-processing including noise reduction and frame segmentation on the collected voice information to obtain pre-processed voice information;
[0160] Step A1.3: collecting physiological information of the target subject through the somatosensory monitoring device, and performing preprocessing including denoising and normalization on the collected physiological information to obtain preprocessed physiological information;
[0161] The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin electrical response and body surface temperature.
[0162] Specifically, step A2 includes:
[0163] Step A2.1: Recognize facial key points based on the preprocessed facial image, and extract facial expression feature vectors based on the recognized facial key points;
[0164] Step A2.2: Extract audio feature vectors that can reflect emotional state based on the preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators;
[0165] Step A2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
[0166] Specifically, the method further includes: collecting the target subject's feedback on cabin environment comfort in real time, and automatically adjusting the environment adjustment strategy according to the feedback;
[0167] At the same time, the joint emotion recognition model is optimized based on the collected target subject’s self-reported emotion evaluation.
[0168] Specifically, the method further includes: encrypting the collected multi-source data, and locally storing the encrypted multi-source data.
[0169] This embodiment deeply integrates emotion recognition with in-car environment adjustment, analyzes the driver's emotional state in real time and adjusts the in-car environment to ensure that the driver is always in the most comfortable state.
[0170] Example 4
[0171] Example 4 is a preferred example of Example 1
[0172] According to the present invention, a method for automatically adapting a multi-climate zone emotional response strategy is provided. Figure 3 Shown, including:
[0173] Step B1: Identify the climate zone type in which the vehicle is currently located;
[0174] Step B2: Collect multi-dimensional environmental data reflecting the environment in which the vehicle is located, and construct a structured environmental state vector based on the multi-dimensional environmental data;
[0175] Step B3: Collecting multimodal data of the target object, including facial images, voice information, and physiological information; extracting feature vectors reflecting the emotional state based on the collected multimodal data; and identifying the current emotion category based on the extracted feature vectors reflecting the emotional state;
[0176] Step B4: Based on the identified climate zone type, environmental state vector, and emotion category of the current vehicle, the optimal matching strategy template is retrieved from the emotion response strategy library to generate device control parameters, thereby automatically adjusting the cabin environment parameters.
[0177] This embodiment achieves context-aware adaptation of occupant emotional response strategies by jointly modeling multimodal emotion recognition results with climate zone types and current environmental parameters, addressing the problem of traditional policy libraries' single adjustment mode and lack of personalized responses. This method combines the occupant's emotional state identified through image, voice, and physiological modalities to match the most appropriate response plan in the policy library, achieving precise responses tailored to individual needs and local conditions.
[0178] Specifically, the step B1 includes: obtaining the vehicle's geographic location information in real time through the vehicle positioning module; matching the climate zone classification database based on the vehicle's geographic location information to output a climate zone label;
[0179] The climate zone label includes: climate zone code, climate zone name, regional distribution and climate characteristics.
[0180] Specifically, step B2 includes:
[0181] Step B2.1: Acquire multi-dimensional environmental perception data through vehicle-mounted sensors, wherein the multi-dimensional environmental perception data includes: vehicle interior temperature, vehicle interior humidity, CO2 concentration, light intensity, vehicle exterior temperature, and vehicle exterior humidity;
[0182] Step B2.2: Accessing the cloud meteorological platform through the communication module to obtain remote meteorological data, wherein the remote meteorological data includes: real-time weather conditions, external temperature and relative humidity, wind speed risk, ultraviolet intensity, and precipitation probability;
[0183] Step B2.3: Construct an environmental state vector based on the multi-dimensional environmental perception data and remote meteorological data using a weighted average method;
[0184] Env_Vector = [temperature, humidity, light, CO2, PM2.5, UV index, wind speed, weather type code];
[0185] Step B2.4: Correct the constructed environment state vector using the credibility correction strategy to obtain a corrected environment state vector.
[0186] This embodiment builds a complete environmental perception subsystem by integrating multi-parameter sensor data from the vehicle's internal and external environments, including temperature, humidity, sunlight intensity, PM2.5 concentration, and in-vehicle carbon dioxide concentration. This solves the problem of existing systems with insufficient parameter collection dimensions and inability to reflect the actual riding environment. This method enables real-time monitoring and linked invocation of environmental factors under extreme climate conditions, providing more accurate and time-sensitive environmental context support for policy selection.
[0187] Specifically, step B3 includes:
[0188] Step B3.1: Acquire a facial image of the target subject, locate the facial region using a face detection model based on the acquired facial image, and extract the facial expression vector using a lightweight neural network model based on the facial region location.
[0189] Step B3.2: Acquire the target subject's sound signal during natural interaction or subjective speech input, perform noise reduction and frame processing on the acquired sound signal to obtain a processed sound signal; extract audio feature parameters reflecting the emotional state based on the processed sound signal, including: Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation index; encode the extracted audio feature parameters reflecting the emotional state into a sound feature vector of a length that meets preset requirements;
[0190] Step B3.3: Acquire physiological signals of the target subject, including heart rate, galvanic skin response, and body surface temperature; perform preprocessing on the acquired physiological signals, including filtering, normalization, and detrending, to obtain preprocessed physiological signals; and construct a physiological state feature vector having a length of several dimensions based on the preprocessed physiological signals;
[0191] Step B3.4: Construct a multimodal fusion vector based on the facial expression vector, the sound feature vector, and the physiological state feature vector. Input the constructed multimodal fusion vector into the multimodal emotion recognition model for emotion classification to obtain the emotional state label of the target object.
[0192] Specifically, when using a multimodal emotion recognition model to perform emotion classification to obtain the emotional state label of the target object, its corresponding confidence score is obtained. When the confidence score is lower than the set threshold, the current emotion result is not used as the basis for strategy triggering; when the confidence score is not lower than the set threshold, it triggers the retrieval of the optimal matching strategy template in the emotion response strategy library and generates device control parameters to enable intelligent control of the device.
[0193] This embodiment introduces a climate recognition mechanism based on GPS positioning and matching with a climate zone geographic information database, thereby solving the technical problem that existing cockpit emotional response systems are unable to distinguish climate characteristics in different geographical regions and their response strategies are not adaptable. This mechanism uses the vehicle's current location to automatically determine the climate zone category and dynamically adjusts the parameter benchmarks of the emotional regulation strategy, providing basic situational criteria for subsequent response strategies, thus realizing the transformation of strategy regulation from "fixed logic" to "regional adaptation."
[0194] Example 5
[0195] Example 5 is a preferred example of Example 1
[0196] According to the present invention, a subscription service method for the emotional cockpit function is provided, such as Figure 4Shown, including:
[0197] Step M1: Collect multimodal data of the target object, including facial images, voice information, and physiological information; identify the target object's emotion category based on the collected multimodal data;
[0198] Step M2: Based on the target object's emotion category and the user subscription policy, call the preset adjustment scheme corresponding to the current emotion category;
[0199] Step M3: Based on the preset adjustment scheme corresponding to the current emotion category, the corresponding cabin experience service package is activated based on the user's selection.
[0200] Specifically, the step M1 includes:
[0201] Step M1.1: Acquire a facial image of a target subject and perform preprocessing on the acquired facial image of the target subject, including denoising, alignment, and cropping, to obtain a preprocessed facial image; perform facial key point recognition based on the preprocessed facial image, and extract a facial expression feature vector based on the recognized facial key points;
[0202] Step M1.2: Acquire the target subject's voice signal and perform noise reduction and frame preprocessing on the acquired voice signal to obtain a preprocessed voice signal; extract audio feature vectors that can reflect the emotional state based on the preprocessed voice information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators;
[0203] Step M1.3: Acquire physiological signals of the target subject, perform preprocessing including denoising and normalization on the acquired physiological information to obtain preprocessed physiological information; construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
[0204] Step M1.4: Construct a multimodal fusion vector based on the facial expression feature vector, audio feature vector, and physiological state feature vector, input the constructed multimodal fusion vector into the multimodal emotion recognition model for emotion classification, and obtain the emotional state label of the target object.
[0205] Specifically, the method also includes: triggering an emotion continuous warning mechanism when the emotion classification is negative for a continuous preset time; and pushing an emotion depth adjustment package that meets the preset requirements through the vehicle interface.
[0206] Specifically, the method further includes: obtaining adjustment feedback of the target object, and adjusting the user subscription strategy according to the adjustment feedback to optimize the intelligent matching.
[0207] This embodiment achieves precise, personalized, and continuous cabin emotion management through the deep integration of multimodal emotion perception and dynamic closed-loop regulation technologies, multi-dimensional data collaboration, intelligent subscription services, and a continuous feedback optimization mechanism. This subscription model also significantly differs from traditional full-featured services by breaking down and packaging different functions into distinct packages, guiding users to experience the entire service system from a single point to a comprehensive level.
[0208] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.
[0209] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A dynamic calibration method for an in-vehicle emotion recognition system, characterized in that: include: Step S1: Acquire multi-source data, including facial images, voice signals, and physiological signals; and preprocessing the acquired multi-source data respectively to obtain preprocessed multi-source data; Step S2: performing feature extraction based on the preprocessed facial image, speech signal, and physiological signal to obtain a facial expression feature vector, an audio feature vector, and a physiological state feature vector; Step S3: Evaluate the current environment credibility based on the environment credibility evaluation function; Step S4: Dynamically assign weights to multi-source data based on the current environmental credibility and real-time scenario; Step S5: Based on the dynamically assigned weights of the multi-source data, a multimodal fusion vector is constructed according to the facial expression feature vector, the audio feature vector, and the physiological state feature vector, and emotion recognition is performed using the constructed emotion recognition model based on the multimodal fusion vector.
2. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 1, characterized in that: The step S1 comprises: Step S1.1: Capturing a facial image of a target subject through an in-vehicle camera, and performing preprocessing on the captured facial image including image denoising, alignment, and cropping to obtain a preprocessed facial image; Step S1.2: collecting the target subject's voice information through the microphone array, and performing pre-processing including noise reduction and frame segmentation on the collected voice information to obtain pre-processed voice information; Step S1.3: collecting physiological information of the target subject through the somatosensory monitoring device, and performing preprocessing including denoising and normalization on the collected physiological information to obtain preprocessed physiological information; The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin electrical response and body surface temperature.
3. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 2, characterized in that: The step S2 comprises: Step S2.1: performing facial key point recognition based on the preprocessed facial image, and extracting facial expression feature vectors based on the recognized facial key points; Step S2.2: Extracting audio feature vectors that can reflect emotional state based on the preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators; Step S2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
4. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 1, characterized in that: The step S3 includes: constructing an environment credibility evaluation function, and using the constructed environment credibility evaluation function to evaluate the credibility of the current environment; Where L represents the light intensity; R valid Indicates the effective frame rate.
5. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 1, characterized in that: The step S4 comprises: Step S4.1: Adjust the weights of multi-source data based on the assessed credibility of the current environment and the real-time scenario; Step S4.2: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights respectively; when the confidence of a multi-source data weight is less than the preset value, the driving scenario correction coefficient is introduced to correct the multi-source data weight to obtain the corrected multi-source data weight.
6. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 5, characterized in that: The weights of the multi-source data in step S4.1 include: face weight W f , speech weight W s and physiological information weight W p ; The facial weight W f include: Among them, W f (0) Assign weight to the base face; k1 is the vehicle speed influence coefficient; Q v,th is the visual quality critical value; h represents the vehicle speed threshold; Z is the normalization factor; The speech weight W s include: Among them, W s (0) Assign weight to basic voice; SNR0 is the reference signal-to-noise ratio; SNR is the signal-to-noise ratio; η∈[0,1] represents the degree of window closing; The physiological information weight W p include: Among them, W p (0) Assign weights to basic physiological information; R C is the road complexity index; γ(R C )={0.05R C ,R C ≤4; 0.2+0.1(R C -4), R C >4.
7. The dynamic calibration method of the vehicle-mounted emotion recognition system according to claim 1, characterized in that: The method further includes: establishing an emotion-behavior association rule base, and when the target object's emotion is identified as "anger" and rapid acceleration occurs within a preset time, the corresponding emotion confidence is increased by ΔC; Among them, γ is the experience adjustment coefficient; a indicates the occurrence of rapid acceleration behavior within the preset time; E indicates that the target object's emotion is recognized as "anger".
8. A dynamic calibration system for an in-vehicle emotion recognition system, characterized in that: include: Module M1: Acquire multi-source data, including facial images, voice signals, and physiological signals; and pre-process the acquired multi-source data to obtain pre-processed multi-source data; Module M2: Extract features based on the preprocessed facial image, speech signal, and physiological signal to obtain facial expression feature vectors, audio feature vectors, and physiological state feature vectors; Module M3: Evaluate the current environment credibility based on the environment credibility evaluation function; Module M4: Dynamically assign weights to multi-source data based on the current environmental credibility and real-time scenarios; Module M5: Based on the dynamically assigned weights of multi-source data, a multimodal fusion vector is constructed according to the facial expression feature vector, audio feature vector, and physiological state feature vector. Emotion recognition is performed using the constructed emotion recognition model based on the multimodal fusion vector.
9. The dynamic calibration system for an in-vehicle emotion recognition system according to claim 8, characterized in that: The module M1 includes: Module M1.1: Capture the target object's facial image through the in-vehicle camera and perform preprocessing on the captured facial image, including image denoising, alignment, and cropping, to obtain the preprocessed facial image; Module M1.2: Collects the target subject's voice information through a microphone array and performs pre-processing on the collected voice information, including noise reduction and frame segmentation, to obtain pre-processed voice information; Module M1.3: Collect physiological information of the target subject through the somatosensory monitoring device, and perform preprocessing on the collected physiological information, including denoising and normalization, to obtain preprocessed physiological information; The voice information includes the sound signal of the target object during natural interaction or subjective voice input; the physiological information includes: heart rate, skin galvanic response and body surface temperature; The module M2 includes: Module M2.1: Recognize facial key points based on preprocessed facial images and extract facial expression feature vectors based on the recognized facial key points; Module M2.2: Extract audio feature vectors that reflect emotional state based on preprocessed speech information, including Mel-frequency cepstral coefficients, intonation contour, fundamental frequency variation, speech rate, and energy variation indicators; Module M2.3: Construct a physiological state feature vector with a length of several dimensions based on the preprocessed physiological information.
10. The dynamic calibration system for an in-vehicle emotion recognition system according to claim 8, characterized in that: The module M3 includes: constructing an environment credibility evaluation function, and using the constructed environment credibility evaluation function to evaluate the credibility of the current environment; Where L represents the light intensity; R valid Indicates the effective frame rate; The module M4 includes: Module M4.1: Adjust the weights of multi-source data based on the assessed credibility of the current environment and the real-time scenario; Module M4.2: Use the improved DS evidence theory to calculate the confidence of the multi-source data weights. When the confidence of a multi-source data weight is less than the preset value, the driving scenario correction coefficient is introduced to correct the multi-source data weight to obtain the corrected multi-source data weight. The weights of the multi-source data in the module M4.1 include: face weight W f , speech weight W s and physiological information weight W p ; The facial weight W f include: Among them, W f (0) Assign weight to the base face; k1 is the vehicle speed influence coefficient; Q v,th is the visual quality critical value; h represents the vehicle speed threshold; Z is the normalization factor; The speech weight W s include: Among them, W s (0) Assign weight to basic voice; SNR0 is the reference signal-to-noise ratio; SNR is the signal-to-noise ratio; η∈[0,1] represents the degree of window closing; The physiological information weight W p include: Among them, W p (0) Assign weights to basic physiological information; R C is the road complexity index; γ(R C )={0.05R C ,R C ≤4; 0.2+0.1(R C -4), R C >4.
Citation Information
Patent Citations
Driver emotion recognition method and onboard control unit implementing the method
CN109017797B
Cited By
Multi-mode driver emotion recognition method and system in real vehicle environment
CN120950889A
Employee emotion index evaluation method and device based on multi-modal data
CN120950901A
Employee sentiment index evaluation method and device based on multi-modal data
CN120950901B
Automatic monitoring method for infrastructure construction operation violation based on multi-modal data fusion
CN120974243A
Equipment operation and maintenance method and system based on multi-modal large model
CN121052807A