A cockpit occupant state monitoring method and device for operating in a high interference environment

By acquiring occupant physiological and behavioral data in a high-interference environment, and combining the environmental context to perform multimodal feature fusion and latent variable coupling modeling, the low reliability of emotion and cognitive state monitoring in existing technologies is solved. This enables collaborative quantification and forward-looking early warning of emotion and cognitive states, thereby improving occupant safety and human-machine collaboration efficiency.

CN121598236BActive Publication Date: 2026-04-10NORTHWESTERN POLYTECHNICAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify occupants' emotional and cognitive states in highly disruptive environments, especially in enclosed cabins such as armored vehicles, engineering machinery, and aircraft. Sensor signals are severely affected by noise, vibration, and light, making it impossible to effectively quantify the coupling relationship between emotions and cognition. This results in low monitoring reliability and makes it difficult to support proactive interventions to mitigate risks.

Method used

By acquiring physiological, behavioral, and environmental context data of passengers, multimodal feature extraction and fusion are performed. Latent variable coupling modeling is used for state prediction, generating human-computer interaction intervention instructions, and achieving coordinated quantification and proactive early warning of emotional and cognitive states.

Benefits of technology

It enables the coordinated quantification of emotional and cognitive states in highly disruptive environments, possesses strong anti-interference capabilities and forward-looking early warning capabilities, improves human-machine collaboration efficiency and occupant safety, and reduces the risk of operational errors caused by deteriorating states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598236B_ABST
    Figure CN121598236B_ABST
Patent Text Reader

Abstract

The application discloses a cockpit occupant state monitoring method and device for high-interference environment operation. The method comprises the following steps: acquiring physiological data, behavior data and environment context data of an occupant; performing feature extraction on the physiological data and the behavior data to obtain a multi-modal feature vector; fusing the multi-modal feature vector and the environment context data to generate an initial state vector; correcting the initial state vector through hidden variable coupling modeling to obtain a current state vector; performing time series prediction based on the current state vector and a historical state vector to obtain a predicted state vector; performing risk determination on the predicted state vector to distinguish between a normal state and non-normal states such as emotional state deterioration and cognitive overload; and generating a human-computer interaction intervention instruction and controlling a vehicle-mounted system to execute under the non-normal state. The application realizes collaborative quantification of emotional and cognitive states, has strong anti-interference and forward-looking early warning capabilities, and improves human-computer collaborative combat effectiveness and occupant task safety.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of man-machine ergonomics and intelligent auxiliary decision-making of special vehicles, and particularly relates to a cabin occupant state monitoring method and device for high-interference environment operation. BACKGROUND

[0002] With the improvement of automobile intelligence and driving safety needs, the mood and cognitive state of the occupant have a more significant impact on driving safety and driving experience. In daily long-distance driving, urban traffic congestion, driving in bad weather and other scenarios, the occupant is prone to cognitive overload due to long-time concentration, or negative emotions such as anxiety and irritability due to complex road conditions and travel delays. If these states cannot be identified and intervened in time, it may lead to operational judgment errors and increase the driving risk.

[0003] However, the current ordinary emotion and cognitive state monitoring technology has obvious limitations: first, the state perception means is one-sided, for example, only through a camera to identify facial expressions to judge fatigue, or only through a heart rate sensor to detect stress. In actual scenarios such as vehicle jolting, frequent changes in light, and the occupant wearing sunglasses / hats, a single modality signal is easily disturbed, resulting in a significant decline in monitoring reliability; second, the existing technology often evaluates cognitive fatigue or emotional state alone, ignoring the interactive influence of the two. For example, anxiety can significantly weaken attention allocation and decision-making ability; third, there is a lack of dynamic prediction capability, which cannot predict in advance the possible cognitive overload (such as entering a series of complex intersections) and emotional state deterioration (such as the escalation of irritability caused by congestion) in the near future, making it difficult to support proactive intervention to avoid risks.

[0004] In particular, in the closed cabin of armored vehicles, engineering machinery, aircraft and other special operation platforms, the occupant faces an extreme composite environment of continuous strong vibration, extremely high noise, sudden change of light and high-intensity tactical task pressure. The sensor signals relied on by existing civilian vehicle monitoring technology are severely deteriorated in signal-to-noise ratio in such an environment, and the algorithm model does not consider the strong interference of environmental context on physiological characteristics, nor does it model the nonlinear coupling relationship between emotional and cognitive states under high stress, resulting in complete system failure or extremely low reliability. Therefore, there is an urgent need for an emotional-cognitive collaborative monitoring and early warning technology that can adapt to extreme physical environments and high psychological load, with strong anti-interference capability. SUMMARY

[0005] Therefore, the cabin occupant state monitoring method and device for high-interference environment operation provided by the embodiments of the present application can realize the collaborative quantification of emotional and cognitive states, have strong anti-interference and proactive warning capabilities, and improve the man-machine collaborative efficiency and occupant safety. The cabin occupant state monitoring method and device for high-interference environment operation provided by the embodiments of the present application are implemented as follows:

[0006] The embodiment of the application provides a cockpit occupant state monitoring method for operation in a high-interference environment, comprising:

[0007] physiological data, behavior data and environment context data representing physical disturbance and task load of the cockpit of an occupant in a vehicle are acquired;

[0008] feature extraction processing is performed on the physiological data and the behavior data, so as to obtain a multi-modal feature vector;

[0009] fusion processing is performed on the multi-modal feature vector and the environment context data, so as to obtain an initial state vector;

[0010] hidden variable coupling modeling processing is performed on the initial state vector, so as to obtain a current state vector;

[0011] time series prediction is performed according to the current state vector and a historical state vector, so as to obtain a predicted state vector;

[0012] state risk determination processing is performed on the predicted state vector, so as to obtain an emotional and cognitive state, wherein the emotional and cognitive state comprises a normal state and an abnormal state, and the abnormal state comprises at least one of cognitive overload and emotional state deterioration;

[0013] in the case that the emotional and cognitive state is the abnormal state, a human-computer interaction intervention instruction is generated, and a vehicle-mounted system is controlled to execute the human-computer interaction intervention instruction.

[0014] In some embodiments, the fusion processing on the multi-modal feature vector and the environment context data to obtain the initial state vector comprises:

[0015] full connection network coding processing is performed on the environment context data, so as to obtain an environment modulation vector, and multi-modal initial embedding vector is obtained by performing multi-modal linear embedding processing on the multi-modal feature vector;

[0016] element-by-element addition processing is performed on the multi-modal initial embedding vector and the environment modulation vector, so as to obtain an environment-modulated modal embedding vector;

[0017] in-modal self-attention calculation processing is performed on the environment-modulated modal embedding vector, so as to obtain an in-modal optimized feature vector;

[0018] weighting processing is performed on the in-modal optimized feature vector, so as to obtain a multi-modal collaborative fusion feature vector;

[0019] fitting processing is performed on the multi-modal collaborative fusion feature vector, so as to obtain the initial state vector.

[0020] In some embodiments, the physiological data includes electroencephalogram-like data and peripheral physiological data, the feature extraction processing on the physiological data and the behavior data obtains a multi-modal feature vector, including:

[0021] The electroencephalogram-like data is subjected to band-pass filtering, power frequency notch filtering and artifact rejection processing to obtain electroencephalogram features;

[0022] The behavior data is subjected to statistical extraction processing to obtain eye movement features;

[0023] The peripheral physiological data is subjected to signal filtering separation and feature statistical processing to obtain processed peripheral physiological data;

[0024] The electroencephalogram features, eye movement features and processed peripheral physiological data are subjected to splicing processing to obtain the multi-modal feature vector.

[0025] In some embodiments, the hidden variable coupling modeling processing on the initial state vector obtains a current state vector, including:

[0026] The initial state vector is subjected to feature extraction processing to obtain an observation variable;

[0027] A hidden variable coupling model is constructed, and the observation variable is input into the hidden variable coupling model to obtain a hidden variable;

[0028] The hidden variable and the observation variable are subjected to partial least squares path analysis processing to obtain path coefficients;

[0029] The initial state vector is subjected to correction processing based on the path coefficients to obtain the current state vector.

[0030] In some embodiments, the time series prediction according to the current state vector and a historical state vector obtains a predicted state vector, including:

[0031] The state vectors on a historical time series are subjected to caching processing to construct a time series data set including the current state vector and historical state vectors within a preset time length;

[0032] The time series data set is subjected to preprocessing to obtain preprocessed time series input data;

[0033] The preprocessed time series input data is input into a preset time series prediction model to obtain an initial prediction vector;

[0034] The initial prediction vector is subjected to standardization processing to obtain the predicted state vector.

[0035] In some embodiments, the processing of the predicted state vector to determine a state risk judgment result comprises:

[0036] The predicted state vector is decomposed to obtain a cognitive load and an emotional state, the emotional state comprising an emotional arousal and an emotional valence;

[0037] The cognitive load and a preset cognitive safety threshold, and the emotional valence and a preset emotional normal threshold are compared respectively, and if the cognitive load does not exceed the cognitive safety threshold and the emotional valence is within the emotional normal threshold range, it is determined as a normal state;

[0038] Or, in the case that the cognitive load exceeds the cognitive safety threshold, it is determined as cognitive overload; in the case that the emotional valence is not within the emotional normal threshold range and the emotional arousal deviates from a normal fluctuation range, it is determined as emotional state deterioration, and is determined as an abnormal state.

[0039] In some embodiments, the preset time series prediction model is a time series prediction model based on a long short-term memory network.

[0040] The embodiments of the present application provide a cabin occupant state monitoring device for operation in a high-interference environment, comprising:

[0041] An acquisition module is configured to acquire physiological data, behavior data of an occupant in a vehicle, and environmental context data representing physical disturbance of a cabin and task load;

[0042] A processing module is configured to perform feature extraction processing on the physiological data and the behavior data to obtain a multi-modal feature vector;

[0043] The processing module is further configured to perform fusion processing on the multi-modal feature vector and the environmental context data to obtain an initial state vector;

[0044] The processing module is further configured to perform hidden variable coupling modeling processing on the initial state vector to obtain a current state vector;

[0045] A prediction module is configured to perform time series prediction according to the current state vector and a historical state vector to obtain a predicted state vector;

[0046] The processing module is further configured to perform state risk judgment processing on the predicted state vector to obtain an emotional and cognitive state, the emotional and cognitive state comprising a normal state and an abnormal state, the abnormal state comprising at least one of cognitive overload and emotional state deterioration;

[0047] The processing module is further configured to generate a human-machine interaction intervention instruction and control a vehicle-mounted system to execute the human-machine interaction intervention instruction when the emotional and cognitive state is an abnormal state.

[0048] The computer device provided in the embodiments of the present application comprises a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements the method provided in the embodiments of the present application when executing the program.

[0049] The computer readable storage medium provided in the embodiments of the present application stores a computer program, and the computer program is executed by a processor to implement the method provided in the embodiments of the present application.

[0050] The method and device for monitoring the state of a cabin occupant working in a high-interference environment provided in the embodiments of the present application obtain physiological data, behavior data and environmental context data representing physical disturbance of the cabin and task load of the occupant; feature extraction is performed on the physiological data and the behavior data to obtain a multi-modal feature vector; the multi-modal feature vector and the environmental context data are fused to generate an initial state vector; the initial state vector is corrected through hidden variable coupling modeling to obtain a current state vector; time series prediction is performed based on the current state vector and a historical state vector to obtain a predicted state vector; risk determination is performed on the predicted state vector to distinguish between a normal state and non-normal states such as emotional state deterioration and cognitive overload; a human-machine interaction intervention instruction is generated in the non-normal state and a vehicle-mounted system is controlled to execute the instruction. The embodiments of the present application realize the collaborative quantification of emotional and cognitive states in a high-interference environment, have strong anti-interference and forward-looking warning capabilities, improve the human-machine collaborative combat effectiveness and the task safety of the occupant, and solve the technical problems proposed in the background art. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without any creative labor.

[0052] Figure 1 An implementation flowchart of the method for monitoring the state of a cabin occupant working in a high-interference environment provided in the embodiments of the present application is shown in the figure.

[0053] Figure 2 An implementation flowchart of the method for monitoring the state of a cabin occupant working in a high-interference environment provided in the embodiments of the present application is shown in the figure.

[0054] Figure 3 An implementation flowchart of the method for monitoring the state of a cabin occupant working in a high-interference environment provided in the embodiments of the present application is shown in the figure. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0056] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0057] Traditional intelligent cockpit emotional interaction technologies generally lack a crew state-combat effectiveness coupling optimization mechanism for high-stress, high-interference mission scenarios. When crew members are in extreme environments such as severe vibration, high noise, and electromagnetic interference on platforms such as armored vehicles, existing technologies cannot quantify the coupling relationship between emotion and cognition and dynamically adjust system strategies, which can easily lead to a decrease in crew combat effectiveness, delay in mission response, and affect the accuracy and safety of tactical execution.

[0058] For example, the invention patent with publication number CN119577557A achieves digital human emotional interaction in various driving environments by integrating multimodal data, providing users with personalized feedback. Although it involves multimodal data application and emotional interaction, its core is to use a large model to achieve a more human-like interactive experience to optimize the comfort of ordinary driving, without addressing anti-interference design and combat effectiveness optimization for high-stress mission scenarios. It does not design anti-interference processing solutions for harsh physical environments, and its output is limited to comfort feedback such as audio and mode adjustments. In the extreme mission environment of armored vehicles, it cannot resist signal interference caused by strong vibrations and high noise, nor can it quantify the interaction between emotion and cognition, let alone support tactical-level dynamic allocation of human-machine functions.

[0059] For example, the invention patent with publication number CN120439964A identifies occupant emotions by integrating multimodal data such as facial images, voice, and physiological signals, and then adjusts cabin environmental parameters such as air conditioning and lighting. Although it involves the relationship between occupant emotions and the cabin system, its core is to improve the comfort of ordinary driving and riding through environmental adjustment, without addressing anti-interference design or combat effectiveness optimization for high-stress mission scenarios. It uses a general data preprocessing method, does not introduce environmental context data for signal decoupling, and the output is only the emotion category for basic environmental adjustment. In the extreme mission environment of armored vehicles, it cannot resist interference to obtain reliable status signals, nor can it quantify the interaction between emotion and cognition, let alone support tactical-level dynamic allocation of human-machine functions.

[0060] For example, the invention patent with publication number CN116101311A uses multimodal perception information from cameras, microphones, etc., for fusion analysis to provide voice or image interaction and reminder functions for ordinary smart cockpits. Although it focuses on the fusion application of multimodal information, its core is to generate basic human-machine interaction content to improve the convenience of ordinary driving and riding, without addressing anti-interference design and combat effectiveness optimization for high-stress mission scenarios. It does not emphasize the targeted overcoming of interference in extreme environments, and the output is only conventional interactive reminder information. In the extreme mission environment of armored vehicles, it cannot cope with the distortion of perception signals caused by strong electromagnetic fields and severe impacts, nor can it quantify the coupling relationship between emotion and cognition to predict effectiveness decay, and it is even more difficult to meet the high-precision requirements of human-machine collaborative decision-making in tactical missions.

[0061] Figure 1 This is a flowchart illustrating the implementation of a cockpit occupant status monitoring method for operations in highly disruptive environments, as provided in this application embodiment, including steps 101 to 107. Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for a cockpit occupant status monitoring method for operations in highly disruptive environments. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0062] Step 101: Obtain physiological data, behavioral data, and environmental context data characterizing cabin physical disturbances and task load of the occupants.

[0063] In this embodiment, physiological data is collected by using a vehicle-mounted, non-invasive sensing device, including multi-channel electroencephalogram (EEG) signals, binocular eye movement data, skin conductance signals, and pulse wave signals.

[0064] Behavioral data is acquired synchronously by relying on eye-tracking sensors to obtain visual behavioral data such as occupant gaze, saccades, and blinks.

[0065] The environmental context data is read through a vehicle bus interface to obtain cabin physical disturbance and task load information, including three-axis impact and vibration frequency spectrum data from a vehicle inertial navigation system, cabin noise and illumination data, and current task phase identification (such as maneuvering, reconnaissance, and engaging the enemy) and threat level obtained from a tactical task management system.

[0066] In step 102, feature extraction is performed on the physiological data and the behavioral data to obtain a multi-modal feature vector.

[0067] In the embodiments of the present application, the collected physiological data and behavioral data are preprocessed and feature-extracted to screen out features that are strongly correlated with emotional state and cognitive load.

[0068] The electroencephalogram feature extraction is performed by band-pass filtering the electroencephalogram signal in a preset frequency range, removing power frequency interference, and removing artifacts. A preset time window and a step are used for sliding analysis to extract electroencephalogram band power and power ratio reflecting cognitive load, and electroencephalogram band power and asymmetry index reflecting emotional state.

[0069] The eye movement feature extraction is performed by event detection based on a preset speed threshold on the eye movement data to identify behaviors such as fixation, saccade, and blinking. In a sliding time window, pupil change features reflecting cognitive load, saccade speed features reflecting attention state, and fixation point features reflecting visual attention distribution are counted.

[0070] The peripheral physiological feature extraction is performed by filtering the skin conductance signal to separate the skin conductance level representing the degree of tension and count the features in the time window. The peak value of the pulse wave signal is detected to generate a heartbeat interval sequence, and then time-domain and frequency-domain features reflecting autonomic nervous activity are extracted.

[0071] After the feature extraction is completed, different types of features are spliced according to the modal, and standardized processing is performed to eliminate the dimensional difference, and finally a multi-modal feature vector with unified dimensions and directly inputtable into a model is obtained.

[0072] In step 103, the multi-modal feature vector and the environmental context data are fused to obtain an initial state vector.

[0073] In the embodiments of the present application, a multi-modal fusion model based on an attention mechanism is used to introduce environmental context data to modulate feature weights and weaken the influence of environmental interference on state evaluation.

[0074] The feature subsets of different modalities in the multi-modal feature vector are respectively mapped to a unified model embedding space through an independent feature mapping layer to obtain initial embedding vectors of each modality; meanwhile, the environmental context data is encoded into an environmental modulation vector matching the dimension of the embedding vector, which is combined with the initial embedding vectors of each modality, so that the model can perceive and adapt to environmental interference. For example, when the environmental context data indicates that the current working condition is extremely high vibration, the model automatically reduces the weight of high-frequency tremor in eye movement features through modulation, thereby avoiding misjudgment of pupil tremor caused by mechanical vibration as cognitive effort.

[0075] First, intra-modal attention calculation is performed on the modality embedding vectors after environmental modulation to extract effective features within a single modality that are strongly related to the emotional-cognitive state; then, cross-modal attention weighting is performed on the optimized features of all modalities, taking the core physiological modality features as the reference to dynamically adjust the contribution weights of the features of each modality, and to strengthen effective information and suppress interference information.

[0076] The fused cross-modal features are input into a regression layer to generate a preliminary evaluation vector containing three core dimensions of cognitive load level, emotional arousal, and emotional valence, i.e., an initial state vector.

[0077] In step 104, the initial state vector is subjected to hidden variable coupling modeling processing to obtain a current state vector.

[0078] In the embodiments of the present application, the hidden variable coupling relationship is constructed through a structural equation model, the bias of the initial state vector is corrected, and the emotional and cognitive states are quantified cooperatively.

[0079] The cognitive load, emotional arousal, and emotional valence in the initial state vector are defined as observed variables. The three are driven by a high-order hidden variable representing the overall psychological and physiological state of the occupant, and the hidden variable reflects the interaction between emotion and cognition.

[0080] The structural equation model is fitted using sample data in a vehicle scene by using a partial least squares path analysis method, the path coefficients between the hidden variable and each observed variable are calculated, and the coupling strength and correlation law of emotion and cognition are verified.

[0081] The initial state vector is theoretically corrected and standardized by combining the hidden variable score obtained by fitting and the observed variable error after path coefficient correction, and finally a current state vector that can truly reflect the emotional-cognitive cooperative state is obtained.

[0082] In step 105, a time series prediction is performed according to the current state vector and the historical state vector to obtain a predicted state vector.

[0083] In the embodiments of the present application, a time series prediction model is used to capture the state evolution law to realize forward-looking prediction of the occupant's state in the future short time.

[0084] Maintain a fixed-length historical state data buffer to store continuous state vectors within the most recent preset time period. Combine the current state vector with the historical state vector in chronological order to form a time-series input dataset.

[0085] Preprocessing time series datasets can be done by using smoothing to suppress short-term random fluctuations or by using differencing to enhance the trend of state evolution and improve the stability of model predictions.

[0086] The preprocessed time series data is input into the time series prediction model (such as a long short-term memory network), and real-time environmental context data can be selectively introduced as auxiliary input to modulate the prediction process. By learning the evolution of historical states, the model outputs the state sequence within a preset time period in the future, i.e., the predicted state vector, and can also include prediction uncertainty information.

[0087] Step 106: Perform state risk assessment on the predicted state vector to obtain the emotional and cognitive states.

[0088] In this embodiment of the application, the normal state and abnormal state of the occupant are distinguished based on preset judgment rules and thresholds.

[0089] Based on a large amount of in-vehicle scenario sample data, a safe threshold for cognitive load, a normal range for emotional valence, and a reasonable fluctuation range for emotional arousal are preset.

[0090] The cognitive load level in the predicted state vector is compared with the safety threshold to determine whether there is a risk of cognitive overload; the emotional valence and emotional arousal are compared with the normal range and the reasonable fluctuation range, respectively, to determine whether there is a risk of emotional state deterioration.

[0091] If there is one or both of the risks of cognitive overload and deterioration of emotional state, it is judged as an abnormal state; if all dimensions are within the normal range, it is judged as a normal state.

[0092] Step 107: When the emotional and cognitive states are abnormal, generate human-computer interaction intervention commands and control the vehicle system to execute the human-computer interaction intervention commands.

[0093] In this embodiment of the application, an appropriate human-machine interaction intervention command is generated based on the specific risk type of the abnormal state, and the vehicle system is driven to execute it.

[0094] To address the risk of cognitive overload, instructions are generated to simplify the main combat situation display interface, suspend secondary information broadcasts, or suggest temporary transfer of mission permissions; to address the risk of worsening negative emotions, instructions are generated to trigger physiological regulation guidance (such as tactical breathing prompts), provide mission certainty feedback, or adjust the response thresholds of auxiliary systems.

[0095] If there are multiple risks, the above intervention instructions can be combined to generate.

[0096] All instructions are transmitted to corresponding functional modules (such as display module, audio module, task management module) through the vehicle-mounted system standard interface, and the driving module executes the intervention operation within the preset time to form a closed loop of monitoring-prediction-intervention, and improve the task efficiency and passenger safety.

[0097] The embodiment of the present application solves the fragmentation problem of the prior art by implementing a full-process closed loop of emotional and cognitive state monitoring. The complete link of data acquisition-feature extraction-fusion modeling-time series prediction-risk determination-intervention execution is covered, avoiding the fragmentation of the existing monitoring method which only stays in the data collection or current state evaluation stage. A complete solution from data to intervention is formed to adapt to the actual needs of ordinary vehicle scenes (such as long-distance driving and urban commuting). By synchronously acquiring physiological, behavioral and environmental data, combined with multi-modal feature fusion, the one-sided problem of state perception caused by the dependence of the prior art on single data is solved, especially for disturbances such as bumps and changes in light in the vehicle scene, making the emotional and cognitive state evaluation more comprehensive. Through hidden variable coupling modeling, the emotional state (such as anxiety and irritability) and the cognitive state (such as cognitive overload and attention distraction) are quantified together, avoiding the defects of the prior art of separately evaluating emotion or cognition, and better fitting the relevance of human psychological and physiological states (such as negative emotions weakening cognitive ability), and the evaluation result is more realistic. Based on time series prediction, future cognitive overload and emotional state deterioration can be predicted in advance to generate intervention instructions (such as simplifying the screen display and voice reminding), and the passenger is given adjustment time, effectively reducing the risk of operation failure caused by state deterioration, and ensuring driving safety.

[0098] On the basis of the above Figure 1 The embodiment of the present application also provides an implementation process diagram for obtaining an initial state vector, as shown in Figure 2 The embodiment of the present application also provides an implementation process diagram for obtaining an initial state vector, as shown in

[0099] Step 201, the environmental context data is processed by a full connection network to obtain an environmental modulation vector, and the multi-modal feature vector is processed by a multi-modal linear embedding to obtain an initial embedding vector.

[0100] In the embodiment of the present application, for the collected vehicle environmental context data (including vehicle vibration data, cabin noise data, environmental illumination data, and driving scene coding, road condition level coding provided by the vehicle information system, etc.), it is input into a small full connection network with adaptive vector dimension. Through the limited hidden layer, the environmental context data is mapped and dimensioned, and finally the environmental modulation vector consistent with the subsequent multi-modal feature embedding vector dimension is output. The purpose of this encoding process is to convert the unstructured or multi-dimensional environmental information into a numerical vector that can be directly combined with the feature vector, so that the model can perceive and adapt to environmental interference in advance (such as physiological signal artifacts caused by strong vibration of the vehicle).

[0101] For the standardized multi-modal feature vector, it is split into EEG feature subset, eye movement feature subset, and peripheral physiological feature subset according to the modal category. In order to make different modal features in a unified data space, three independent linear layers are used to perform linear embedding processing on the three feature subsets respectively. Each linear layer maps the corresponding modal feature dimension to the same model embedding space as the environmental modulation vector, and finally obtains the EEG modal initial embedding vector, the eye movement modal initial embedding vector, and the peripheral physiological modal initial embedding vector, which together constitute the multi-modal initial embedding vector set.

[0102] Step 202, element-wise addition processing is performed on the initial embedding vector of the multi-modal and the environmental modulation vector to obtain the environmental modulation modal embedding vector.

[0103] In the embodiment of the present application, the obtained multi-modal initial embedding vector and the environmental modulation vector are processed by element-wise addition. Specifically, the EEG modal initial embedding vector and the environmental modulation vector are added element by element to obtain the environmental modulation EEG embedding vector. Similarly, the eye movement modal initial embedding vector and the environmental modulation vector are added element by element to obtain the environmental modulation eye movement embedding vector, and the peripheral physiological modal initial embedding vector and the environmental modulation vector are added element by element to obtain the environmental modulation peripheral physiological embedding vector. The core role is to inject environmental information as prior knowledge into each modal feature. For example, when the environmental context data shows that the vehicle is in a severe jolt state, the environmental modulation vector will pre-correct the modal features such as eye movement and peripheral physiology that are easily disturbed by vibration through element-wise addition, so that the model can actively distinguish between feature fluctuations caused by environmental disturbance and feature fluctuations caused by changes in the real state of the occupant in subsequent feature processing.

[0104] Step 203, intra-modal attention calculation processing is performed on the environmental modulation modal embedding vector to obtain the intra-modal optimized feature vector.

[0105] In the embodiment of the present application, the modality-in self-attention calculation processing is performed on each modality embedding vector modulated by the environment. Taking the electroencephalogram embedding vector modulated by the environment as an example, the electroencephalogram embedding vector is input into a multi-head self-attention layer, and the correlation between different features in the electroencephalogram modality is calculated to automatically strengthen the feature weight of the feature more valuable for the emotional-cognitive state evaluation and weaken the feature weight of the redundant or disturbed feature. Similarly, the same modality-in self-attention calculation is performed on the eye movement embedding vector and the peripheral physiological embedding vector modulated by the environment. The effective features in each modality are refined and enhanced, for example, the correlation feature between the pupil diameter change and the cognitive load in the eye movement modality is strengthened, and the abnormal signal caused by accidental blinking is weakened, and finally the intra-modality optimized feature vector of the electroencephalogram modality, the intra-modality optimized feature vector of the eye movement modality, and the intra-modality optimized feature vector of the peripheral physiological modality are output, which together constitute the intra-modality optimized feature vector set.

[0106] In step 204, the intra-modality optimized feature vector is weighted to obtain a multi-modal collaborative fusion feature vector.

[0107] In the embodiment of the present application, the intra-modality optimized feature vector of the electroencephalogram modality, the intra-modality optimized feature vector of the eye movement modality, and the intra-modality optimized feature vector of the peripheral physiological modality are spliced to form a cross-modality feature set.

[0108] The intra-modality optimized feature vector of the electroencephalogram modality is taken as a query reference, and all feature vectors in the cross-modality feature set are taken as key-value pairs.

[0109] Through cross-attention calculation, the contribution weight of each modality feature in the fusion process is dynamically adjusted. For example, when the environmental context data shows that the sudden change of cabin illumination leads to a decrease in the reliability of eye movement features, the model will automatically reduce the weight of the eye movement modality feature, while increasing the weight of the modality feature with stronger anti-interference ability such as electroencephalogram and peripheral physiology. Conversely, when the environment is stable, the model will evenly distribute the weight of each modality.

[0110] Through weighting processing, the effective features of different modalities are collaboratively integrated, and finally a multi-modal collaborative fusion feature vector is output.

[0111] In step 205, the multi-modal collaborative fusion feature vector is fitted to obtain an initial state vector.

[0112] In the embodiment of the present application, the multi-modal collaborative fusion feature vector is input into a full connection layer, and regression fitting processing is performed through the full connection layer. The full connection layer maps the high-dimensional multi-modal collaborative fusion feature vector into a low-dimensional numerical vector according to the requirements of the emotional-cognitive state evaluation. In the fitting process, the parameters of the full connection layer are optimized through a preset training target (such as minimizing the error between the fitting result and the true state label) to ensure that the output vector can accurately reflect the current preliminary emotional-cognitive state of the occupant. Finally, the full connection layer outputs a vector containing the quantized values of the three dimensions of cognitive load level, emotional arousal, and emotional valence, i.e., the initial state vector.

[0113] The embodiment of the present application generates a modulation vector by coding the environmental context data through a full connection network, injects environmental interference as prior knowledge into multi-modal features, solves the feature deviation problem caused by ignoring environmental influence in existing fusion methods, and makes the fusion result more resistant to dynamic interference in the vehicle scene. First, effective features of a single modality are extracted through intra-modal self-attention, and then multi-modal information is integrated through cross-modal weighting, avoiding invalid information interference caused by equal weighting in existing fusion methods, and making the multi-modal collaborative fusion features more accurate.

[0114] In some embodiments, the physiological data includes electroencephalogram data and peripheral physiological data, and the physiological data and the behavior data are subjected to feature extraction processing to obtain a multi-modal feature vector, including: performing band pass filtering, power frequency notch filtering and artifact rejection processing on the electroencephalogram data to obtain electroencephalogram features.

[0115] Specifically, the original electroencephalogram data is input into a band pass filtering module, and the filtering frequency range is set to 1-45 Hz. Through filtering operation, low-frequency baseline drift (such as signal deviation caused by slight shaking of the occupant's body) below 1 Hz and high-frequency noise (such as high-frequency interference generated by vehicle electronic devices) above 45 Hz in the electroencephalogram signal are filtered out, and the electroencephalogram band signal strongly related to cognitive and emotional state is retained.

[0116] Considering that in-vehicle electrical equipment (such as air conditioner, central control system) is prone to 50 Hz power frequency interference, 50 Hz power frequency notch filtering is performed on the electroencephalogram signal after band pass filtering to suppress the interference signal of this specific frequency through narrowband filtering.

[0117] Independent component analysis method is adopted to separate and reject the artifacts of the above-mentioned pre-processed electroencephalogram signal. The electrooculogram artifacts (such as signal fluctuation caused by blinking) and electromyogram artifacts (such as signal interference caused by facial muscle activity) mixed in the electroencephalogram signal can be identified and separated, and the pure electroencephalogram signal component is retained.

[0118] After removing the artifacts from the electroencephalogram signal, a sliding analysis method with a 4-second time window and a 2-second step is used to extract two types of core features: one is the feature reflecting cognitive load (such as the average power of the theta band in the frontal lobe region, the theta / β power ratio), and the other is the feature reflecting emotional state (such as the alpha band power in the frontal lobe region, the frontal lobe alpha asymmetry index), which together constitute the electroencephalogram features.

[0119] Further, the behavior data is statistically extracted to obtain eye movement features.

[0120] Specifically, the original eye movement data is first de-noised (such as filtering out abnormal coordinate values caused by temporary device obstruction), and then event detection is performed based on a speed threshold. The speed determination criteria (such as determining saccadic behavior when the eye movement speed exceeds 30° / s, and determining fixation behavior when the speed is lower than 5° / s and the duration exceeds 100 milliseconds) are set to identify the three core behaviors of fixation, saccade and blink in the eye movement data.

[0121] The detected eye movement behaviors are statistically analyzed in units of 4-second sliding time windows, which are consistent with the electroencephalogram feature extraction. The sliding average and standard deviation of the pupil diameter (reflecting cognitive load changes, such as increased pupil diameter and increased fluctuations when cognitive overload occurs), the average peak speed of saccadic movement (reflecting attention transfer efficiency), and the information entropy of fixation point distribution (reflecting visual attention dispersion, with higher entropy indicating more dispersed attention) are calculated. The above statistical results together constitute the eye movement features.

[0122] Further, the peripheral physiological data is filtered and separated, and the processed peripheral physiological data is obtained.

[0123] Specifically, the original skin conductance signal is subjected to low-pass filtering to filter out high-frequency noise and separate the skin conductance level that can stably reflect the passenger's tension level. The mean value of the skin conductance level within a 4-second sliding time window is calculated as a skin conductance-related feature.

[0124] The peak value of the original photoplethysmogram signal is detected. By identifying the peak position of each heartbeat corresponding to the pulse wave, a sequence of successive heartbeat intervals is generated. Based on the sequence, two types of features are further extracted: one is the time domain feature (such as the average heartbeat interval and the standard deviation of the heartbeat interval, reflecting the stability of the heartbeat), and the other is the frequency domain feature (such as the high-frequency band power and the low-frequency band power ratio of the sequence, reflecting the balance between the sympathetic and parasympathetic nerves of the autonomic nervous system).

[0125] The above skin conductance-related features and photoplethysmogram-derived features are integrated to obtain the processed peripheral physiological data.

[0126] Further, the electroencephalogram features, eye movement features, and processed peripheral physiological data are spliced to obtain a multi-modal feature vector.

[0127] Specifically, the obtained electroencephalogram features, eye movement features, and processed peripheral physiological data are spliced in the order of electroencephalogram features, eye movement features, and peripheral physiological data to form an initial feature set. Considering the dimensional differences of the three types of features, Z-score standardization processing is performed on the initial feature set. By calculating the mean and standard deviation of each feature, all feature values are converted into standardized values with a mean of 0 and a standard deviation of 1, eliminating the influence of dimensional differences on subsequent model inputs. Finally, the feature set after splicing and standardization is the multi-modal feature vector.

[0128] The application embodiments can filter out vehicle-mounted electrical appliance interference and human artifacts through band-pass filtering, power frequency notch filtering, and artifact removal of electroencephalogram data, solving the problem of large feature noise caused by direct use of raw data. The filtering and separation of peripheral physiological data can extract features that stably reflect the state of tension and autonomic nervous activity, ensuring that the features of each modality have strong correlation with emotion-cognition properties. Through uniform sliding time window statistics of eye movement and peripheral physiological features, combined with Z-score standardization, the problem of large dimensional differences of different modal features is solved, making the multi-modal feature vector format uniform and the numerical values comparable, providing input data with strong adaptability for multi-modal fusion and avoiding fusion bias caused by feature format disorder.

[0129] In some embodiments, the initial state vector is subjected to hidden variable coupling modeling processing to obtain a current state vector, including: performing feature extraction processing on the initial state vector to obtain observation variables.

[0130] Specifically, the initial state vector is a preliminary emotion-cognition evaluation vector after multi-modal fusion, including three key dimensions: cognitive load level (quantifying the information processing pressure and attention allocation state of the occupant), emotional arousal (quantifying the emotional excitement level of the occupant, such as calm, tension, and irritability), and emotional valence (quantifying the positive or negative tendency of the occupant's emotion, such as pleasure, neutrality, and anxiety). Based on this, when performing feature extraction on the initial state vector, the quantification values of the above three dimensions are directly taken as independent features, which are respectively defined as cognitive load observation variables, emotional arousal observation variables, and emotional valence observation variables, and the three together constitute the observation variable set required for hidden variable coupling modeling. The core purpose of the extraction process is to select core indicators that can directly reflect the cognitive and emotional state.

[0131] Further, a hidden variable coupling model is constructed, and the observation variables are input into the hidden variable coupling model to obtain a hidden variable.

[0132] Specifically, a structural equation model is used to build a latent variable coupling model, and the core logic of the model is that the observed variables are driven by high-order latent variables. The cognitive load observed variable, the emotional arousal observed variable, and the emotional valence observed variable do not exist independently, but are jointly affected by an indirectly measurable passenger overall psychophysiological state latent variable. The latent variable comprehensively reflects the interaction of emotion and cognition (for example, when the overall psychological tension is high, it will not only cause the cognitive load to increase, but also make the emotional arousal rise and the emotional valence tend to be negative). In terms of model structure, the passenger overall psychophysiological state latent variable is taken as an exogenous latent variable, and the three observed variables are taken as endogenous observed variables, and a mapping relationship from the latent variable to each observed variable is established (that is, the change of the latent variable will affect the value of each observed variable through a specific correlation rule).

[0133] The three extracted observed variables are input into the constructed latent variable coupling model, and the measurement equation of the model is used to calculate the quantitative score of the passenger overall psychophysiological state latent variable (that is, the specific numerical representation of the latent variable) based on the sample data in the vehicle scene (covering observed variable data of different driving time, road conditions, and passenger states).

[0134] Further, a partial least squares path analysis is performed on the latent variable and the observed variable to obtain the path coefficient.

[0135] Specifically, the partial least squares path analysis method is used to quantitatively verify and calculate the correlation between the latent variable and the observed variable.

[0136] A large amount of sample data in the vehicle scene (such as the latent variable scores and corresponding observed variable values of different passengers in long-distance driving, urban congestion, and adverse weather conditions) are collected, and the data are input into the latent variable coupling model. The partial least squares algorithm is used to fit the model, so that the error between the predicted values of the observed variables output by the model and the actual sample data is minimized.

[0137] During the fitting process, the standardized path coefficient of the passenger overall psychophysiological state latent variable to each observed variable is calculated simultaneously. For example, the path coefficient of the latent variable to the cognitive load observed variable quantifies the influence strength of the overall psychophysiological state change on the cognitive load; and the path coefficient of the latent variable to the emotional valence observed variable quantifies the influence rule of the overall state change on the emotional positive or negative tendency. The value range of the path coefficient is usually [-1, 1], and the greater the absolute value, the higher the correlation strength (for example, when the path coefficient is 0.8, it means that the latent variable changes by 1 unit, and the corresponding observed variable changes by about 0.8 units).

[0138] Further, the initial state vector is corrected based on the path coefficient to obtain the current state vector.

[0139] Specifically, first, the residual of each observation variable in the initial state vector (i.e., the difference between the actual value of the observation variable and the value of the observation variable predicted by the model through the hidden variable) is calculated, which reflects the random error that may exist in the initial state vector due to single data fusion (such as the observation variable deviation caused by the temporary environmental interference of the eye movement feature at a certain moment). The residual is corrected by weighting in combination with the path coefficient obtained above. The observation variable with a larger absolute value of the path coefficient has a greater impact on the overall state evaluation, and a higher correction weight is required to reduce the interference of random error on the result.

[0140] The corrected observation variable residual and the passenger overall psychophysiological state hidden variable score are integrated to make the final state vector not only retain the specific quantitative information of each dimension, but also integrate the influence of the overall coordinated state. For example, if the hidden variable score shows that the overall psychological state is tense, and the path coefficient from the hidden variable to the emotional arousal degree is high, the quantitative value of the emotional arousal degree needs to be adjusted appropriately during correction to make it more consistent with the overall state logic.

[0141] The integrated vector is subjected to standardization processing to eliminate the dimensional difference that may be generated in the correction process, and finally the current state vector that can truly reflect the emotional-cognitive coordinated state is obtained. Compared with the initial state vector, the current state vector can better avoid single data bias and be more consistent with the actual psychophysiological state of the passenger.

[0142] The embodiments of the present application solve the random bias that may exist in the initial vector generated by relying on data fusion only by constructing the correlation between the hidden variable and the observation variable through the structural equation model, upgrade the state evaluation from data-driven to data plus theory-driven, and the result is more reliable. The correlation strength between the hidden variable and the observation variable is calculated through the path coefficient, the interaction rule of emotion and cognition is clear, the subjectivity of the existing technology in determining the relationship between emotion and cognition is avoided, the coupling modeling process is quantifiable and verifiable, and it is ensured that the current state vector can truly reflect the coordinated state of emotion and cognition. From the observation variable extraction, model construction to path coefficient correction, all are based on clear statistical methods, avoiding the subjectivity of the correction process, and ensuring the stability of the initial vector correction results under different samples and different scenes.

[0143] In some embodiments, a time series prediction is performed according to the current state vector and the historical state vector to obtain a predicted state vector, including: performing cache processing on the state vector in the historical time series to construct a time series data set including the current state vector and the historical state vector in the recent preset time length.

[0144] Specifically, the vehicle-mounted system maintains a fixed-length first-in-first-out (FIFO) data buffer, and the buffer capacity is determined according to a preset time length and a state vector generation interval. For example, if the history state vectors of the last 3 minutes are stored, and the state vectors are updated at a frequency of 1 every 2 seconds, the buffer can store 90 history state vectors (3 minutes x 60 seconds / 2 seconds = 90), ensuring that a sufficient history evolution period is covered.

[0145] Whenever a new current state vector is generated, it is stored at the tail of the buffer, and the earliest history state vector at the head of the buffer is removed (if the buffer is full), so that the buffer always stores continuous state vectors of the last preset time length; if the buffer is not full, the current state vector is directly stored until the preset capacity is reached.

[0146] All history state vectors stored in the buffer are arranged in order of generation time from early to late, and the latest current state vector is spliced to the tail of the sequence to form a time series data set with dimensions of (history vector number + 1) x state vector dimensions. The time series data set completely records the evolution process of the occupant state from the starting time of the last preset time length to the current time.

[0147] Further, the time series data set is preprocessed to obtain preprocessed time series input data.

[0148] Specifically, if there is a short-term state fluctuation in the time series data set due to accidental factors (such as the occupant briefly lowering his head, sudden external noise), a moving average method is used for processing. Set the size of the sliding window (such as 3 consecutive state vectors as 1 window), calculate the average value of each dimension feature (cognitive load, emotional arousal, emotional valence) in the window, replace the value of the state vector at the center of the window with the average value, and slide the window to complete the processing of the entire data set, thereby weakening the random fluctuations that mask the state trend.

[0149] If the time series data set shows a clear state evolution trend (such as the cognitive load gradually increasing and the emotional valence gradually decreasing during long-distance driving), a first-order difference method is used for processing. Calculate the difference between adjacent two state vectors (the value of the latter vector minus the value of the former vector) to obtain a difference sequence reflecting the state change rate, and convert the absolute state value to a relative change value, so that the model is more likely to capture the rising / falling trend of the state.

[0150] According to the actual characteristics of the time series data set, one or a combination of the above processing methods (such as smoothing first and then differentiating) is selected, and finally the preprocessed time series input data that eliminates noise and highlights trends is obtained.

[0151] Further, the preprocessed time series input data is input into a preset time series prediction model to obtain an initial prediction vector.

[0152] Specifically, a long short-term memory network (LSTM) is adopted as a time series prediction model, and a future state sequence is output by learning a historical state evolution rule through the model.

[0153] The preset LSTM model includes an LSTM layer and a fully connected output layer. The LSTM layer adopts a one-way LSTM structure with 128 hidden units, and is used to capture long-term dependencies of states in time series input data. The fully connected output layer maps the output of the LSTM layer to a state vector sequence corresponding to a preset prediction time length (such as 1-5 minutes in the future) according to the future preset prediction time length.

[0154] The preprocessed time series input data is directly input into the LSTM model. Meanwhile, real-time environmental context data (such as scene information such as heavy rain in the next 5 minutes on a congested road segment obtained by a vehicle-mounted navigation device) can be selectively encoded into an environmental feature vector matching the dimension of the time series input data, and injected into the cell state of the LSTM model, to modulate the prediction process (such as predicting that a congested road segment will increase cognitive load) and improve the degree of prediction fitting to the actual scene.

[0155] The parameters of the model formed through pre-training (using historical time series data and subsequent real states as training samples, and minimizing the error between the predicted state and the real state as the target) are used to calculate the input time series data, and a continuous state vector sequence (such as one state vector every 2 seconds in the next 5 minutes) in a preset time length in the future is output. The sequence is the initial prediction vector, which includes the preliminary quantitative values of cognitive load, emotional arousal, and emotional valence at each time in the future.

[0156] Further, the initial prediction vector is standardized to obtain the predicted state vector.

[0157] Specifically, the Z-score standardization method consistent with the current state vector of the multi-modal feature vector is adopted. The mean and standard deviation of each dimension feature (cognitive load, emotional arousal, and emotional valence) in the initial prediction vector in the historical training sample are calculated, and each feature value in the initial prediction vector is converted into a standardized value of (feature value-mean) / standard deviation, so as to eliminate the dimension difference of different features.

[0158] According to the output capability of the model, the prediction interval or uncertainty index (such as the probability of cognitive load exceeding the threshold in the next 2 minutes being 85%) can be attached to the standardized vector. These information is derived from the statistical analysis of the prediction error in the training process of the LSTM model.

[0159] After standardization, the predicted state vector with uniform format, consistent dimension, and directly usable for risk judgment is obtained, which fully reflects the evolution trend of the occupant's emotional-cognitive state in the preset time length in the future.

[0160] The embodiment of the application stores the history state vector of the preset time length by the fixed length FIFO buffer area, solves the trend capture deviation problem caused by the incomplete history data of the existing prediction method, makes the time series data set fully reflect the state evolution process from the past to the present, and provides continuous trend basis for the prediction model. The short-term fluctuation (such as temporary abnormality of cognitive load caused by the occasional lowering of the passenger) is inhibited by the smoothing processing, and the long-term trend (such as the gradual increase of cognitive load in long-distance driving) is strengthened by the difference processing, the prediction noise caused by the direct use of the original time series data by the existing prediction method is solved, the trend is not clear, and the preprocessed time series input data is more suitable for the demand of the prediction model. The initial prediction vector is standardized, the prediction result is uniform in format and consistent in dimension with the current state vector, and the problem that the prediction result format is chaotic, which leads to the difficulty in comparing with the current state and risk judgment, is solved.

[0161] In some embodiments, the state risk judgment processing is performed on the predicted state vector to obtain the emotional and cognitive state, including: performing disassembly processing on the predicted state vector to obtain the cognitive load and the emotional state, and the emotional state including the emotional arousal and the emotional valence.

[0162] Specifically, the predicted state vector is dimensionally disassembled in a fixed order from the cognitive load to the emotional arousal and then to the emotional valence, and the quantization values of the three dimensions are extracted respectively.

[0163] The cognitive load quantization value reflects the information processing pressure of the passenger at the future moment (such as 5.2 indicating moderate information processing pressure, and the higher the value, the greater the pressure).

[0164] The emotional arousal quantization value reflects the emotional excitement degree of the passenger at the future moment (such as 3.1 indicating mild excitement, and the value being too low is depression and too high is tension).

[0165] The emotional valence quantization value reflects the positive and negative tendency of the emotion of the passenger at the future moment (such as 0.6 indicating a positive tendency, and the value being lower than 0 is negative and higher than 0 is positive).

[0166] The extracted cognitive load quantization value is classified as a cognitive state indicator, and the emotional arousal quantization value and the emotional valence quantization value are jointly classified as an emotional state indicator, forming two kinds of judgment basis of cognitive load and emotional state.

[0167] Further, the cognitive load and the preset cognitive safety threshold and the emotional valence and the preset emotional normal threshold are compared respectively, and if the cognitive load does not exceed the cognitive safety threshold and the emotional valence is within the range of the emotional normal threshold, it is determined as a normal state.

[0168] Specifically, the cognitive load data of the passengers in the normal driving state under different vehicle-mounted scenes (such as long-distance uniform-speed driving, urban congestion following, driving in heavy rain, and driving in insufficient night lighting) is collected (a total of not less than 1000 valid samples are collected), the maximum value of the cognitive load in the samples is counted (for example, the statistical result is 6.0), and the maximum value is set as the cognitive safety threshold. That is, when the cognitive load quantitative value is less than or equal to 6.0, it is determined that the cognitive state is not overloaded.

[0169] Similarly, based on the sample data of the normal driving state, the mean (for example, 0.2) and the standard deviation (for example, 0.3) of the emotional valence are counted, and the range of the mean ± 1 times the standard deviation is set as the emotional normal threshold (that is, 0.2-0.3=-0.1 to 0.2+0.3=0.5). That is, when the emotional valence quantitative value is in the interval [-0.1, 0.5], it is determined that the emotional tendency is normal.

[0170] The minimum value (for example, 2.0) and the maximum value (for example, 4.0) of the emotional arousal in the normal driving state sample are counted, and the interval [2.0, 4.0] is set as the normal fluctuation range of the emotional arousal. That is, when the emotional arousal quantitative value is in the interval, it is determined that the emotional excitement degree is reasonable, and there is no excessive depression or excessive tension.

[0171] Further, or, in the case where the cognitive load exceeds the cognitive safety threshold, it is determined that the cognition is overloaded; in the case where the emotional valence is not in the emotional normal threshold range and the emotional arousal deviates from the normal fluctuation range, it is determined that the emotional state is deteriorated, and it is determined that the state is not normal.

[0172] Specifically, the normal state determination: when the two conditions are met at the same time, it is determined that the state is normal: the cognitive load quantitative value is less than or equal to the cognitive safety threshold (for example, less than or equal to 6.0): it indicates that the information processing pressure of the passenger at the future time is within the safe range, and the operation failure caused by cognitive overload will not occur; the emotional valence quantitative value is in the emotional normal threshold interval (for example, [-0.1, 0.5]): it indicates that the emotional tendency of the passenger at the future time is neutral or slightly positive, and there is no obvious negative emotion.

[0173] Example: If the prediction vector is decomposed into cognitive load 5.5, emotional arousal 3.2, and emotional valence 0.3, the cognitive load is not over the threshold, and the emotional valence is in the normal interval, it is determined that the state is normal.

[0174] Abnormal state determination: when any of the following conditions is met, it is determined to be an abnormal state, and further distinguish the specific risk type. Cognitive overload determination: only need to meet the cognitive load quantitative value > cognitive safety threshold (such as > 6.0), without considering the emotional state. Common in the future time passengers need to handle a large amount of information scene (such as about to enter the continuous complex intersection, long distance driving fatigue at the end of attention), example: predicted cognitive load 6.8, emotional valence 0.4, determine cognitive overload. Emotional state deterioration determination: need to meet two conditions: emotional valence quantitative value < emotional normal threshold lower limit (such as < -0.1) or > emotional normal threshold upper limit (such as > 0.5, the upper limit is out of limit, mostly due to excessive excitement leading to emotional out of control); emotional arousal quantitative value < normal fluctuation range lower limit (such as < 2.0, showing emotional depression) or > normal fluctuation range upper limit (such as > 4.0, showing emotional tension / frustration). This situation is common in the future time scene deterioration (such as about to enter the long time congestion road section, sudden severe weather), example: predicted emotional valence -0.3, emotional arousal 4.5, determine emotional state deterioration.

[0175] The final determination result (normal state abnormal state-cognitive overload abnormal state-emotional state deterioration) is used as the emotional and cognitive state output, which is directly used for the generation of subsequent human-machine interaction intervention instructions.

[0176] The embodiment of the application solves the problem of fuzzy emotional and cognitive state determination standard by decomposing the prediction vector into cognitive load, emotional arousal, and emotional valence, clearly determines the cognitive state and emotional state determination dimension, and makes the determination basis clearer. The cognitive safety threshold, emotional normal threshold, and arousal normal fluctuation range are all based on a large number of sample data statistics in the vehicle scene (such as the maximum value of cognitive load, the mean value ± standard deviation of emotional valence in the normal driving state), which avoids the determination deviation caused by subjective threshold setting, and ensures the uniformity and objectivity of the determination standard for different passengers and different scenes.

[0177] Although the application provides method operation steps such as embodiments or flowcharts, more or fewer operation steps can be included based on conventional or non-inventive labor. The order of steps listed in the embodiment is only one of the many step execution orders, and does not represent the only execution order. When the device or client product is executed in practice, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-thread processing environment) according to the method order shown in the embodiment or the drawing.

[0178] As shown in Figure 3 The embodiment of the application also provides a cabin passenger state monitoring device 300 for high interference environment operation. The device comprises:

[0179] The acquisition module 301 is used to acquire physiological data, behavioral data, and environmental context data characterizing cabin physical disturbances and task load of the occupants.

[0180] The processing module 302 is used to perform feature extraction processing on physiological and behavioral data to obtain multimodal feature vectors.

[0181] The processing module 302 is also used to fuse the multimodal feature vector with the environmental context data to obtain the initial state vector.

[0182] The processing module 302 is also used to perform latent variable coupling modeling on the initial state vector to obtain the current state vector.

[0183] The prediction module 303 is used to perform time-series prediction based on the current state vector and the historical state vector to obtain the predicted state vector.

[0184] The processing module 302 is also used to perform state risk judgment processing on the predicted state vector to obtain emotional and cognitive states, which include normal states and abnormal states. Abnormal states include at least one of cognitive overload and emotional state deterioration.

[0185] The processing module 302 is also used to generate human-computer interaction intervention commands when the emotional and cognitive states are abnormal, and to control the vehicle system to execute the human-computer interaction intervention commands.

[0186] In some embodiments, the processing module 302 is further configured to perform fully connected network encoding on the environmental context data to obtain an environmental modulation vector and perform modal linear embedding on the multimodal feature vector to obtain an initial embedding vector for the multimodality.

[0187] The processing module 302 is also used to perform element-wise addition of the initial embedding vector of the multimodal mode and the environmental modulation vector to obtain the modal embedding vector after environmental modulation.

[0188] The processing module 302 is also used to perform intramodal self-attention calculation on the modal embedding vector after environmental modulation to obtain the intramodal optimized feature vector.

[0189] The processing module 302 is also used to perform weighted processing on the intramodal optimized feature vector to obtain a multimodal collaborative fusion feature vector.

[0190] The processing module 302 is also used to fit the multimodal collaborative fusion feature vector to obtain the initial state vector.

[0191] In some embodiments, the processing module 302 is further configured to perform bandpass filtering, power frequency notch filtering, and artifact removal processing on the EEG data to obtain EEG features.

[0192] The processing module 302 is further configured to perform statistical extraction processing on the behavior data to obtain eye movement features.

[0193] The processing module 302 is further configured to perform signal filtering and separation and feature statistical processing on the peripheral physiological data to obtain processed peripheral physiological data.

[0194] The processing module 302 is further configured to perform splicing processing on the electroencephalogram features, the eye movement features, and the processed peripheral physiological data to obtain a multi-modal feature vector.

[0195] In some embodiments, the processing module 302 is further configured to perform feature extraction processing on the initial state vector to obtain an observation variable.

[0196] The processing module 302 is further configured to construct a latent variable coupling model and input the observation variable into the latent variable coupling model to obtain a latent variable.

[0197] The processing module 302 is further configured to perform partial least squares path analysis processing on the latent variable and the observation variable to obtain a path coefficient.

[0198] The processing module 302 is further configured to perform correction processing on the initial state vector based on the path coefficient to obtain a current state vector.

[0199] In some embodiments, the processing module 302 is further configured to perform cache processing on the state vectors on a historical time sequence to construct a time series data set including the current state vector and historical state vectors within a preset time length.

[0200] The processing module 302 is further configured to perform preprocessing on the time series data set to obtain preprocessed time series input data.

[0201] The prediction module 303 is further configured to input the preprocessed time series input data into a preset time series prediction model to obtain an initial prediction vector.

[0202] The prediction module 303 is further configured to perform standardization processing on the initial prediction vector to obtain a predicted state vector.

[0203] In some embodiments, the processing module 302 is further configured to perform disassembly processing on the predicted state vector to obtain a cognitive load and an emotional state, the emotional state including emotional arousal and emotional valence.

[0204] The prediction module 303 is further configured to compare the cognitive load with a preset cognitive safety threshold and the emotional valence with a preset emotional normal threshold, respectively, and determine that the state is normal when the cognitive load does not exceed the cognitive safety threshold and the emotional valence is within the range of the emotional normal threshold.

[0205] The prediction module 303 is further configured to determine that the cognitive load is overloaded when the cognitive load exceeds the cognitive safety threshold. When the emotional valence is not within the emotional normal threshold range and the emotional arousal deviates from the normal fluctuation range, it is determined that the emotional state is deteriorated, and it is determined that the state is abnormal.

[0206] Some of the modules in the apparatus described in the present application can be described in the general context of computer-executable instructions, such as program modules, which are executed by computers. Generally, program modules include routines, programs, objects, components, data structures, classes, and the like, which perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including storage devices.

[0207] The apparatus or modules described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above apparatus is described as various modules with functions. In the implementation of the embodiments of the present application, the functions of the modules can be implemented in the same or multiple software and / or hardware. Of course, the modules implementing certain functions can also be implemented by multiple sub-modules or sub-units.

[0208] The methods, apparatuses or modules described in the present application can be implemented in a computer readable program code in any appropriate manner, for example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code (for example, software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASIC), programmable logic controllers and embedded microcontrollers, examples of the controller include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in pure computer readable program code, the same function can be achieved by logically programming the method steps in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers. Therefore, such a controller can be considered as a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both a software module for implementing the method and a structure within the hardware component.

[0209] The embodiments of the present application also provide a device, which comprises: a processor; a memory for storing processor executable instructions; and the processor implements the method as described in the embodiments of the present application when executing the executable instructions.

[0210] The embodiments of the present application also provide a non-volatile computer readable storage medium, which stores a computer program or instructions, and when the computer program or instructions are executed, the method as described in the embodiments of the present application is implemented.

[0211] In addition, each functional module in each embodiment of the present application can be integrated in one processing module, or each module can exist independently, or two or more modules can be integrated in one module.

[0212] The storage medium described above includes but is not limited to random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD) or memory card. The memory can be used to store computer program instructions.

[0213] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary hardware. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product or in the form of data migration. The computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0214] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. The whole or part of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc.

[0215] The above embodiments are only used to illustrate the technical solutions of the present application, and not to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present application.

Claims

1. A method for monitoring the status of cabin occupants during operations in highly disruptive environments, characterized in that, include: Acquire physiological and behavioral data of occupants, as well as environmental context data characterizing cabin physical disturbances and task load; The physiological and behavioral data are processed for feature extraction to obtain multimodal feature vectors; The multimodal feature vector and the environmental context data are fused to obtain an initial state vector; The initial state vector is subjected to latent variable coupling modeling to obtain the current state vector; Based on the current state vector and the historical state vector, a time series prediction is performed to obtain the predicted state vector; The predicted state vector is subjected to state risk assessment to obtain emotional and cognitive states, which include normal and abnormal states. The abnormal states include at least one of cognitive overload and emotional state deterioration. When the emotional and cognitive states are abnormal, a human-computer interaction intervention command is generated, and the vehicle system is controlled to execute the human-computer interaction intervention command.

2. The method according to claim 1, characterized in that, The process of fusing the multimodal feature vector with the environmental context data to obtain the initial state vector includes: The environmental context data is encoded using a fully connected network to obtain an environmental modulation vector, and the multimodal feature vector is subjected to modal linear embedding to obtain an initial embedding vector for the multimodality. The initial embedding vector of the multimodal mode is added element by element to the environment modulation vector to obtain the environment-modulated mode embedding vector; Intramodal self-attention calculation is performed on the modal embedding vector after environmental modulation to obtain the intramodal optimized feature vector; The intra-modal optimized feature vectors are weighted to obtain multi-modal collaborative fusion feature vectors; The initial state vector is obtained by fitting the multimodal collaborative fusion feature vector.

3. The method according to claim 1, characterized in that, The physiological data includes electroencephalogram (EEG) data and peripheral physiological data. The feature extraction process performed on the physiological and behavioral data to obtain a multimodal feature vector includes: The EEG data is processed by bandpass filtering, power frequency notch filtering, and artifact removal to obtain EEG features; The behavioral data is statistically extracted to obtain eye movement features; The peripheral physiological data are subjected to signal filtering and separation and feature statistical processing to obtain the processed peripheral physiological data. The EEG features, eye movement features, and processed peripheral physiological data are spliced ​​together to obtain the multimodal feature vector.

4. The method according to claim 1, characterized in that, The process of performing latent variable coupling modeling on the initial state vector to obtain the current state vector includes: The initial state vector is subjected to feature extraction processing to obtain the observed variables; Construct a latent variable coupling model by inputting the observed variables into the latent variable coupling model to obtain the latent variables; Partial least squares path analysis is performed on the latent variables and the observed variables to obtain the path coefficients; The initial state vector is corrected based on the path coefficients to obtain the current state vector.

5. The method according to claim 1, characterized in that, The step of performing time-series prediction based on the current state vector and historical state vectors to obtain the predicted state vector includes: The state vectors in the historical time series are cached to construct a time series dataset that includes the current state vector and the historical state vectors within the most recent preset time period; The time series dataset is preprocessed to obtain preprocessed time series input data; The preprocessed time series input data is input into the preset time series prediction model to obtain the initial prediction vector; The initial prediction vector is standardized to obtain the predicted state vector.

6. The method according to claim 1, characterized in that, The process of performing state risk determination on the predicted state vector to obtain emotional and cognitive states includes: The predicted state vector is decomposed to obtain cognitive load and emotional state, which includes emotional arousal and emotional valence. The cognitive load is compared with a preset cognitive safety threshold, and the emotional valence is compared with a preset emotional normality threshold. If the cognitive load does not exceed the cognitive safety threshold and the emotional valence is within the range of the emotional normality threshold, it is determined to be a normal state. Alternatively, if the cognitive load exceeds the cognitive safety threshold, it is determined to be cognitive overload; if the emotional valence is not within the normal emotional threshold range and the emotional arousal deviates from the normal fluctuation range, it is determined to be a deterioration of the emotional state and judged as an abnormal state.

7. The method according to claim 5, characterized in that, The preset time series prediction model is a time series prediction model based on long short-term memory networks.

8. A cockpit occupant status monitoring device for operation in highly disruptive environments, characterized in that, include: The acquisition module is used to acquire physiological data, behavioral data, and environmental context data characterizing cabin physical disturbances and task load of the occupants. The processing module is used to perform feature extraction processing on the physiological data and behavioral data to obtain multimodal feature vectors; The processing module is further configured to fuse the multimodal feature vector with the environmental context data to obtain an initial state vector; The processing module is also used to perform latent variable coupling modeling on the initial state vector to obtain the current state vector; The prediction module is used to perform time-series prediction based on the current state vector and historical state vectors to obtain the predicted state vector. The processing module is also used to perform state risk judgment processing on the predicted state vector to obtain emotional and cognitive states, the emotional and cognitive states include normal states and abnormal states, the abnormal states include at least one of cognitive overload and emotional state deterioration. The processing module is also used to generate human-computer interaction intervention commands when the emotional and cognitive states are abnormal, and to control the vehicle system to execute the human-computer interaction intervention commands.

9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interaction method and device for intelligent cabin

    CN116101311A

  • Digital human emotion recognition and feedback system based on large model

    CN119577557A

  • Human factor engineering test system and method for emotion cockpit

    CN120439964A

  • Vehicle cabin sound effect control method and device, storage medium and equipment

    CN119239628A

  • Automobile cabin user experience evaluation method and system based on digital interconnection

    CN119721483A