Driving state evaluation method and device, electronic equipment and vehicle

By acquiring head and eye movement, facial video, and physiological electrical signal data, and utilizing multi-dimensional perception features and a driving fatigue status analysis model, the problem of low accuracy in driver fatigue detection is solved, and a highly accurate assessment of the driver's fatigue status is achieved.

CN120678435APending Publication Date: 2025-09-23BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410338964.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing technology has low accuracy in driver fatigue detection, which makes it difficult to truly reflect the driver's status and has low detection accuracy.

Method used

By acquiring head and eye movement data, facial video data, and physiological electrical signal data, and utilizing multi-dimensional perception features and driving fatigue status analysis models, including long-short-term memory networks and deep neural networks, the driver's multi-dimensional perception features are analyzed to determine the driving status.

Benefits of technology

It improves the accuracy of driver fatigue status detection, provides a reliable basis for judgment, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120678435A_ABST
    Figure CN120678435A_ABST
Patent Text Reader

Abstract

The invention provides a driving state evaluation method and device, electronic equipment and a vehicle, and relates to the technical field of mode recognition. The driving state evaluation method comprises the steps that head-eye movement data, face video data and physiological electric signal data are obtained, and the physiological electric signal data comprise electroencephalogram signals, electrocardiosignals and electromyographic signals; determining multi-dimensional perception features corresponding to the head-eye movement data, the face video data and the physiological electric signal data; and analyzing the multi-dimensional perception characteristics through a driving fatigue state analysis model, and determining a driving state. Through the technical scheme provided by the invention, the problem of low driving state detection accuracy is solved, and the driving state detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of pattern recognition, and in particular to a driving state assessment method, device, electronic equipment, and vehicle. Background Art

[0002] Driver fatigue is a significant safety issue in modern vehicles. Data indicates that up to 20% of traffic accidents are caused by driver fatigue, with the proportion of truck accidents attributable to driver fatigue reaching 54%. Unpredictable driving times, disruptions to circadian rhythms, and visual fatigue during driving all inevitably contribute to driver fatigue. Driving places high demands on the driver's ability to resist fatigue, spatial awareness, motion sickness resistance, attention span, attention distribution, and muscle strength. The driver's skills, training, and driving state directly impact driving effectiveness and safety. Real-time status monitoring technology, which assesses the driver's state during driving, can effectively improve driver skills and personal qualities, reduce and avoid unsafe factors caused by human error, and enhance driving safety.

[0003] In related technologies, the driver's fatigue level is identified through visual recognition technology, which can only be used to roughly estimate fatigue and non-fatigue. It is difficult to truly reflect the driver's status and the detection accuracy is low. Summary of the Invention

[0004] This disclosure provides a driving state assessment method, device, electronic device, chip, and medium to address the issue of low driving state detection accuracy. By analyzing multi-dimensional perception data such as gaze, facial expression, and physiological electrical signals, along with driver behavior data, and using a driving fatigue state analysis model, the method determines the driving state, improving the accuracy of driving state detection.

[0005] A first embodiment of the present disclosure provides a driving state assessment method, the method comprising:

[0006] Acquire head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals;

[0007] Determine multi-dimensional perceptual features corresponding to head and eye movement data, facial video data, and physiological electrical signal data;

[0008] The driving fatigue state analysis model is used to analyze multi-dimensional perception features and determine the driving state.

[0009] In one embodiment of the present disclosure, determining multi-dimensional perceptual features corresponding to head and eye movement data, facial video data, and physiological electrical signal data includes:

[0010] Extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data;

[0011] Sort the multiple visual gaze behavior features, multiple expression features, and multiple physiological electrical change features in chronological order to form visual gaze behavior data, expression data, and physiological electrical change data;

[0012] The visual gaze behavior data, facial expression data, and physiological electrical change data are integrated into multi-dimensional perception features.

[0013] In one embodiment of the present disclosure, extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data includes:

[0014] Preprocessing the head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively. The visual gaze data includes at least one of the first gaze duration, the second gaze duration, the number of gazes, the gaze duration, the pupil diameter, and the saccade distance. The first gaze duration is the duration of the driver's first stay on the gaze point, the second gaze duration is the duration of the second stay on the gaze point, the number of gazes is the number of times the gaze point has been stayed, and the gaze duration is the total duration of the gaze point stay.

[0015] Acquire driving control data, and extract a plurality of visual gaze behavior features through a first feature extraction network based on the visual gaze data and the driving control data;

[0016] Performing expression recognition on the facial emotion data using an expression recognition model to extract multiple expression features, wherein the multiple expression features include at least one of energetic, neutral, nervous, and sleepy;

[0017] Multiple physiological electrical change features of physiological electrical signals are extracted through deep neural networks. The multiple physiological electrical change features include multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands. The multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands correspond to different fatigue states respectively.

[0018] In one embodiment of the present disclosure, a driving fatigue state analysis model is used to analyze multi-dimensional perception features to determine the driving state, including:

[0019] Using a long short-term memory network, visual gaze behavior data, facial expression data, and driving control data are used to identify driver behavior patterns and obtain the first fatigue state. The long short-term memory network belongs to the driving fatigue state analysis model;

[0020] Determine the second fatigue state corresponding to the multi-dimensional perception characteristics through the state assessment model, which belongs to the driving fatigue state analysis model;

[0021] A weighted sum is performed on the first fatigue state and the second fatigue state to obtain a driving state.

[0022] In one embodiment of the present disclosure, the head and eye movement data, facial video data, and physiological electrical signal data are preprocessed into visual gaze data, facial emotion data, and physiological electrical data, respectively, including:

[0023] Filtering head and eye movement data and physiological electrical signal data;

[0024] Screen out abnormal data in head and eye movement data and physiological electrical signal data;

[0025] Correlating and collating the head-eye movement data and / or physiological electrical signal data with the vehicle cabin scene where the driver is located, so as to retain the corresponding head-eye movement data or physiological electrical signal data obtained in the vehicle cabin scene;

[0026] Recognize facial emotion data corresponding to facial video data through convolutional neural networks.

[0027] In one embodiment of the present disclosure, a long short-term memory network is used to perform driver behavior pattern recognition on visual gaze behavior data, expression data, and driving control data to obtain a first fatigue state, including:

[0028] The encoder of the attention mechanism and global average pooling are used to extract the gaze features of the visual gaze landing rate time series data of the visual gaze behavior data;

[0029] The encoder of the attention mechanism and global average pooling are used to extract the gaze landing point features of the region of interest gaze sequence data of the visual gaze behavior data. The region of interest gaze sequence data refers to the visual gaze behavior data recorded when the expression data is a preset expression.

[0030] Extract visual gaze category features of visual gaze behavior data through a deep neural network classification model;

[0031] Extracting control features of driving control data through attention mechanism encoder, long short-term memory network, and global average pooling;

[0032] Through the multi-classification model, the landing point sight features, sight gaze landing point features, visual gaze category and manipulation features are classified to obtain the first fatigue state.

[0033] A second embodiment of the present disclosure provides a driving state evaluation device, the device comprising:

[0034] An acquisition module is used to acquire head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals;

[0035] A fusion module is used to determine the multi-dimensional perceptual features corresponding to the head and eye movement data, facial video data, and physiological electrical signal data;

[0036] The determination module is used to analyze the multi-dimensional perception features through the driving fatigue state analysis model to determine the driving state.

[0037] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods in the first aspect embodiment of the present disclosure.

[0038] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to enable a computer to execute the method in the first aspect embodiment of the present disclosure.

[0039] The fifth aspect embodiment of the present disclosure proposes a chip, comprising at least one processor and a communication interface; the communication interface is used to receive signals input into the chip or signals output from the chip, the processor communicates with the communication interface and implements any one of the methods in the first aspect embodiment of the present disclosure through logic circuits or execution code instructions.

[0040] A sixth aspect embodiment of the present disclosure proposes a vehicle, comprising the driving state evaluation device in the second aspect embodiment of the present disclosure or the electronic device in the third aspect embodiment of the present disclosure.

[0041] In summary, the driving state assessment method proposed in this disclosure obtains head and eye movement data, facial video data, and physiological electrical signal data. The physiological electrical signal data includes EEG signals, ECG signals, and EMG signals, providing a highly practical data source for driving state assessment. Multidimensional perceptual features corresponding to the head and eye movement data, facial video data, and physiological electrical signal data are determined, providing a reliable basis for judging driving state assessment. The multidimensional perceptual features are analyzed using a driving fatigue state analysis model to determine the driving state and obtain an assessment of the driver's fatigue state. This improves the accuracy of driver fatigue state detection.

[0042] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0044] Figure 1 This is a structural diagram of the LSTM module;

[0045] Figure 2 This is a flowchart of a driving state assessment method according to an embodiment of the present disclosure;

[0046] Figure 3 This is a flow chart of an embodiment of the present disclosure for determining multi-dimensional perception features corresponding to head and eye movement data, facial video data, and physiological electrical signal data;

[0047] Figure 4 This is a flowchart for collecting multi-dimensional perception feature data;

[0048] Figure 5 This is a flowchart of an embodiment of the present disclosure for extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data;

[0049] Figure 6 Schematic diagram of the structure of the first feature extraction network;

[0050] Figure 7 Schematic diagram of the brain activity analysis module of the deep neural network;

[0051] Figure 8 This is a flow chart of an embodiment of the present disclosure for analyzing multi-dimensional perception features through a driving fatigue state analysis model to determine a driving state;

[0052] Figure 9 Schematic diagram of the accuracy analysis of single module and multi-dimensional fusion model;

[0053] Figure 10 Schematic diagram of multi-dimensional fusion model training;

[0054] Figure 11 A flowchart of an embodiment of the present disclosure for preprocessing head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively;

[0055] Figure 12 This is a comparison diagram of EEG signal data before and after filtering;

[0056] Figure 13This is a flowchart of an embodiment of the present disclosure that uses a long short-term memory network to perform driver behavior pattern recognition on visual gaze behavior data, expression data, and driving control data to obtain a first fatigue state;

[0057] Figure 14 Schematic diagram of the structure of a driving state evaluation device according to an embodiment of the present disclosure;

[0058] Figure 15 is a block diagram of an electronic device for implementing the driving state evaluation method disclosed herein according to an exemplary embodiment;

[0059] Figure 16 It is a schematic structural diagram of a chip according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout identify the same or similar components or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0061] First, a brief introduction to the relevant terms in this disclosure is given:

[0062] An electroencephalogram (EEG) is a graph obtained by amplifying and recording the brain's spontaneous biopotentials from the scalp using an electroencephalogram (EEG). It records the spontaneous, rhythmic electrical activity of groups of brain cells through electrodes. EEG includes conventional EEG, ambulatory EEG, and video EEG. EEGs contain signals related to fatigue. For example, 32-channel EEG data represents 32 sets of electrode frequency data. This frequency data is transformed to generate a power spectrum for easier analysis. Different states of fatigue exhibit distinct frequency variations. Fatigue-related EEG signals are primarily concentrated in the delta, theta, alpha, and beta frequencies, with frequency bands of 0-4 Hz, 4-7 Hz, 8-13 Hz, and 14-30 Hz, respectively. Delta waves reflect excessive fatigue and deep sleep, theta waves reflect frustration or depression, alpha waves reflect a state of calm and concentration, and beta waves reflect tension, emotionality, or excitement.

[0063] An electrocardiogram (ECG) is a graph that uses an electrocardiogram (ECG) to record the changes in electrical activity produced by the heart during each cardiac cycle. Studies have found that when a person is fatigued, the heart rate may increase or decrease, and the rhythm may become irregular. Certain waveforms on the ECG may change. For example, the P wave may decrease in amplitude, the QRS complex may lengthen, and the ST segment and T wave may be altered.

[0064] An electromyogram (EMG) is a measurement of muscle bioelectricity recorded using an electromyograph. While primarily used to detect muscle disorders, EMG can also reflect a number of physiological and psychological conditions, including fatigue. When a person is fatigued, changes in muscle electrical activity may occur. For example, the speed of muscle contraction and relaxation may slow, and the force of muscle contraction may decrease. These changes can be detected and recorded using EMG.

[0065] Long Short-Term Memory (LSTM): In the present disclosure, an improvement to the recurrent neural network (RNN). RNN can make predictions based on current information and historical information. However, as the complexity of neural networks gradually increases, RNNs often have problems with information overload and local over-optimization. As a variant of RNN, LSTM can use gate control units to make the network's information extraction more selective, thereby effectively improving information utilization and the accuracy of time series prediction. The LSTM module consists of three parts: a forget gate, an input gate, and an output gate. The forget gate determines which information from the previous memory cell state should be deleted. The input gate selects information from candidate storage cells to update the cell state. The output gate filters information from the storage cell and only outputs information about part of the cell state.

[0066] Figure 1 Figure 1 is a schematic diagram of the LSTM module. As shown in the figure, the LSTM module consists of a forget gate, an input gate, and an output gate.

[0067] For the forget gate, its function is to calculate the output value of the forget gate using the following formula.

[0068] f t =σ(W f *H t-1 +U f *X t +b f )

[0069] For the input gate, its function is to calculate the output value of the input gate using the following formula:

[0070] i t =σ(W i *H t-1 +U i *X t +b i )

[0071] For the output gate, its function is to calculate the output value of the output gate using the following formula:

[0072] O t =σ(W o *H t-1 +U o *X t +b o )

[0073] The following formula is used to calculate the storage unit value C t and the output H of the LSTM block t :

[0074]

[0075]

[0076] In the above formula, W f 、W i 、W o 、W c is the weight between the current hidden layer and the previous hidden layer, U f 、U i 、U o 、U c is the weight between the current input layer and the current hidden layer, b f 、b i 、b o 、b c is the bias vector, is the element-wise multiplication operator, σ is the sigmoid function, and tanh is the activation function.

[0077] Regarding driver fatigue detection, research both domestically and internationally focuses on two types of driver status: one, based on the driver's glance behavior, to obtain information about their attention allocation; the other, based on the driver's biological and physiological information, to detect the probability of fatigue. With the development of various wearable devices and charge-coupled devices, driver status assessment has evolved from the initial subjective analysis of questionnaires to various technical means for detecting various types of status, such as the measurement and analysis of various physiological data. However, most detection systems can only detect fatigue or non-fatigue, making it difficult to obtain accurate and reliable test results.

[0078] The method proposed in this disclosure is applied to driving state assessment tasks and has a wide range of applications. It can be used in long-distance freight transport, where drivers often drive continuously for long periods of time and are prone to fatigue. Fatigue recognition technology can promptly remind drivers to take a break or switch drivers, improving driving safety. It can also be used in the taxi industry, where drivers often work for hours on end and are prone to traffic accidents due to fatigue. Fatigue recognition technology can help taxi companies monitor driver fatigue and take timely measures to ensure the safety of both passengers and drivers. It can also be used in the bus industry, where drivers drive for long periods of time and fatigue is a major cause of traffic accidents. Fatigue recognition technology can monitor driver fatigue in real time, prompting them to take breaks or take other safety measures. It can also be used in the logistics industry, where drivers spend long hours delivering goods, where fatigue can lead to accidents and cargo loss. Fatigue recognition technology can help logistics companies monitor driver fatigue and ensure the safety and efficiency of the transportation process. It can also be used in the field of autonomous driving, where the development of autonomous driving technology is inseparable from the monitoring and analysis of driver behavior. Fatigue recognition technology can provide autonomous driving systems with information on driver fatigue, helping them make more accurate decisions and improving driving safety. The application scenarios in this disclosure are not limited.

[0079] The driving state assessment method provided by the present disclosure is described in detail below with reference to the accompanying drawings.

[0080] Figure 2 FIG. 1 is a flow chart of a driving state evaluation method according to an embodiment of the present disclosure. Figure 2 In the embodiment shown, the driving state assessment method includes:

[0081] Step 201 : Acquire head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals.

[0082] In this embodiment, head-eye movement data refers to data on head and eye movements. Facial video data refers to video data of the driver's face. Physiological signal data refers to various physiological signal data of the driver, including EEG signals, ECG signals, and EMG signals. In order to evaluate the driving status, head-eye movement data, facial video data, and physiological signal data are first obtained. For example, in a simulated driving platform, a camera is used to collect image data of the driver's head and eyes, a camera is used to record video data of the driver's face, an EEG monitor is used to collect the driver's EEG signals, an ECG meter is used to collect the driver's ECG signals, and an electromyograph is used to collect the driver's EMG signals.

[0083] Step 202: Determine multi-dimensional perceptual features corresponding to the head and eye movement data, facial video data, and physiological electrical signal data.

[0084] In this embodiment, multi-dimensional perception features refer to the characteristic representation of the human body's perception state in multiple dimensions. In the present disclosure, it refers to data that is pre-processed, feature-extracted, and aligned on the time axis for head and eye movement data, facial video data, and physiological electrical signal data. For example, at the first moment, the head is tilted down, the eyes are closed, the mouth is wide open, the facial expression is sleepy, and delta waves begin to appear in the brain waves. At the second moment, the head is facing down, the eyes are closed, the mouth is closed, the facial expression is sleepy, and delta waves are mostly in the brain waves. At the third moment, the head is raised, the eyes are open, the mouth is open, the facial expression is normal, and the alpha waves in the brain waves gradually increase, while the delta waves decrease. By arranging the first moment, the second moment, and the third moment in sequence on the time axis, a detection of a complete driving fatigue process can be formed, and all the features recorded in these three moments constitute the multi-dimensional perception features.

[0085] Step 203: Analyze the multi-dimensional perception features using a driving fatigue state analysis model to determine the driving state.

[0086] In this embodiment, the driving fatigue state analysis model refers to a detection model for analyzing a driver's fatigue state during simulated driving or driving, and includes an LSTM and a driver state assessment model. Driving state refers to the driver's physical and mental state during simulated driving or driving. Preferably, driving state refers to the driver's level of fatigue during fatigue driving.

[0087] In summary, the driving state assessment method proposed in this disclosure obtains head and eye movement data, facial video data, and physiological electrical signal data. The physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals, providing a highly practical data source for driving state assessment. Multidimensional perceptual features corresponding to the head and eye movement data, facial video data, and physiological electrical signal data are determined. The fused multidimensional perceptual features provide a reliable basis for driving state assessment. The multidimensional perceptual features are analyzed using a driving fatigue state analysis model to determine the driving state and obtain an assessment of the driver's fatigue state. This improves the accuracy of driver fatigue state detection.

[0088] Figure 3 This is a flowchart of an embodiment of the present disclosure for determining multi-dimensional perception features corresponding to head and eye movement data, facial video data, and physiological electrical signal data. Figure 3 Yes Figure 2 Further explanation of step 202 is based on Figure 3 The embodiment shown includes the following steps:

[0089] Step 301 : extract multiple visual gaze behavior features corresponding to the head and eye movement data, multiple expression features corresponding to the facial video data, and multiple physiological electrical change features corresponding to the physiological electrical signal data.

[0090] In this embodiment, visual gaze behavior features refer to the characteristic representation of the driver's gaze movements within the feature extraction neural network. For example, the feature vector representations of the driver's gaze shifting to the steering wheel, lingering on the steering wheel, and moving away from the steering wheel are represented. Facial expression features refer to the characteristics of the driver's facial expression. For example, when smiling, the corners of the mouth are raised and the eyes are slightly closed; when tired, the eyes are closed and the mouth is slightly open or closed. Physiological change features refer to the waveform characteristics of the driver's physiological electrical signal recordings.

[0091] After collecting head and eye movement data, facial video data, and physiological electrical signal data, in order to obtain multi-dimensional perceptual features, it is necessary to first extract the visual gaze behavior features corresponding to the head and eye movement data, the expression features corresponding to the facial video data, and the physiological electrical change features corresponding to the physiological electrical signal data. For example, for head and eye movement data, gaze tracking is performed using an eye tracking device or corneal reflection method. A face detection algorithm based on a depth camera is used to detect the driver's face and track the speed of facial movement, thereby obtaining the motion trajectory of the gaze point on the gaze plane. Features of the gaze trajectory are extracted in a feature extraction neural network. Optionally, features such as the maximum / minimum value points and rate of change of the gaze trajectory curve are extracted using a recurrent convolutional neural network. For facial video data, an expression recognition algorithm based on deep learning methods is used to identify the driver's facial expression data and obtain a feature vector representation of the corresponding expression. Optionally, the expression recognition algorithm based on deep learning methods can use a three-dimensional convolutional neural network (3DCNN). For physiological electrical signal data, the signal features of the electroencephalogram, electrocardiogram or electromyogram are learned through a deep learning model for physiological electrical signals. For example, the electroencephalogram signal is classified through the electroencephalogram signal channel network (EEG-ChannelNet), and the features corresponding to the category are represented by feature vectors.

[0092] In one implementation of this embodiment,

[0093] Figure 4This is a flowchart for multidimensional perceptual feature data acquisition. As shown in the figure, after the sensor data acquisition system collects head and eye movement data, EEG data, and physiological signal data, it uses a head and eye movement data analyzer to analyze the data and extract the first eigenvector; an EEG data analyzer to analyze the EEG data and extract the second eigenvector; and a physiological signal data analyzer to analyze the physiological signal data and extract the third eigenvector. The first, second, and third eigenvectors are then subjected to comprehensive data adaptation and aligned according to chronological order. The data is then stored in a database and merged to train a feature extraction model appropriate for each data type.

[0094] In step 302, the plurality of visual gaze behavior features, the plurality of expression features, and the plurality of physiological electrical change features are sorted in chronological order to form visual gaze behavior data, expression data, and physiological electrical change data.

[0095] In this embodiment, visual gaze behavior data refers to the action behavior data formed by multiple visual gaze behavior features in a time sequence. Expression data refers to the expression change data formed by multiple expression features in a time sequence. Physiological change data refers to the signal change data formed by multiple physiological change features in a time sequence.

[0096] After obtaining multiple visual gaze behavior features, facial expression features, and physiological electrical change features, they are organized into visual gaze behavior data, facial expression data, and physiological electrical change data in chronological order. For example, at the first moment, feature vector a1 represents the data gaze behavior features at the first moment, feature vector b1 represents the facial expression features at the first moment, and feature vector c1 represents the physiological electrical change features at the first moment. At the second moment, feature vector a2 represents the data gaze behavior features at the second moment, feature vector b2 represents the facial expression features at the second moment, and feature vector c2 represents the physiological electrical change features at the second moment. If the first moment is earlier than the second moment, and the chronological order is from the first moment to the second moment, feature vector (a1, a2) represents the visual gaze behavior data, feature vector (b1, b2) represents the facial expression data, and feature vector (c1, c2) represents the physiological electrical change data.

[0097] Step 303: Fusing the visual gaze behavior data, facial expression data, and physiological electrical change data into multi-dimensional perception features.

[0098] In this embodiment, after obtaining visual gaze behavior data, facial expression data, and physiological electrical change data, these three types of data are combined based on the time sequence to form a multi-dimensional perception feature. For example, based on the time sequence, that is, from the first moment to the second moment, the multi-dimensional perception feature generated is [(a1, a2), (b1, b2), (c1, c2)].

[0099] In this embodiment, feature extraction and fusion are performed based on head and eye movement data, facial video data, and physiological electrical signal data to obtain multi-dimensional perception features. The fused multi-dimensional perception features provide a reliable judgment basis for driving status assessment.

[0100] Figure 5 This is a flowchart of an embodiment of the present disclosure for extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data. Figure 5 Yes Figure 3 Further explanation of step 301 is based on Figure 5 The embodiment shown includes the following steps:

[0101] Step 501, preprocessing the head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively. The visual gaze data includes at least one of the first gaze point duration, the second gaze point duration, the number of gazes, the gaze duration, the pupil diameter, and the saccade distance. The first gaze point duration is the length of time the driver's gaze point stays for the first time, the second gaze point duration is the length of time the gaze point stays for the second time, the number of gazes is the number of times the gaze point stays, and the gaze duration is the total length of time the gaze point stays.

[0102] In this embodiment, visual gaze data refers to head and eye movement data that has been filtered, cleaned, screened, and correlated. It includes at least one of the following: first fixation duration, second fixation duration, number of fixations, fixation duration, pupil diameter, and saccade distance. Facial emotion data refers to facial expression data correctly identified from facial video data. Physiological data refers to data obtained after filtering out interference signals and addressing abnormal fluctuations in sensor sampling rates. Optionally, filtering or interpolation algorithms can be used to address issues such as sensor data sampling rate fluctuations and interference filtering of physiological signals. Data cleaning can be used to filter out erroneous and useless data. Data correlation and collation involves linking sensor signal data with the driving simulation cockpit scene. For example, visual gaze data is calibrated to align with the cockpit physical environment data, and visual gaze is associated with vehicle instrument information using regions of interest. For example, when the gaze falls on an object, that object is considered the visual fixation target. Six key characteristic indicators of visual fixation exist: first fixation duration, second fixation duration, number of fixations, fixation duration, pupil diameter, and saccade distance. The first fixation duration is the duration of the driver's first fixation, the second fixation duration is the duration of the second fixation, the number of fixations is the number of times the fixation was made, and the fixation duration is the total duration of the fixation. The focus is on the driver's gaze on the instrument. The driving simulation platform consists of three visual screens and one instrument screen as the display output. The head-up display (HUD) is displayed on the visual screen, and the remaining instruments are displayed on the instrument screen. The displayed instruments are divided into regions and modeled to obtain instrument data. When a person fixates on an instrument, the gaze falls on the corresponding instrument, and the instrument object is directly obtained. The gaze object detection data includes: gaze object attributes and instrument data, gaze start and end time, and gaze trajectory within the instrument. The gaze point motion trajectory, gaze object, and accompanying visual gaze feature data are arranged in time series to form visual gaze behavior data.

[0103] Step 502 : Acquire driving control data, and extract multiple visual gaze behavior features through a first feature extraction network based on the visual gaze data and the driving control data.

[0104] In this embodiment, driving control data refers to the driver's control data of the vehicle during simulated driving or driving, which is obtained by the vehicle computer from the vehicle's electronic control unit. The first feature extraction network is a deep neural network that extracts the driver's visual gaze behavior characteristics. It includes an attention mechanism encoder module, a classification network module, an LSTM module, a global average pooling module, and a feature merging module.

[0105] Figure 6is a schematic diagram of the structure of the first feature extraction network. In this embodiment, Figure 6 As shown, the visual gaze data includes visual gaze point rate time series data, area of ​​interest (AOI) gaze sequence data, and visual gaze features. The visual gaze data and driving control data are used as the input of the first feature extraction network. The visual gaze point rate time series data and the AOI gaze sequence data are sequentially passed through the attention mechanism encoder module and global average pooling to obtain line of sight features and gaze point features. The visual gaze features are classified through the deep neural network (DNN) classification network module to obtain visual gaze category features. The driving control data is passed through the attention mechanism encoder module, the LSTM module and the global average pooling to obtain the driving control features. The line of sight features, gaze point features, visual gaze category features and driving control features are merged and output as visual gaze behavior features associated with the driving operation behavior.

[0106] Step 503: Perform expression recognition on the facial emotion data through an expression recognition model to extract multiple expression features, where the multiple expression features include at least one of energetic, neutral, nervous, and sleepy.

[0107] In this embodiment, the expression recognition model refers to a network model that recognizes facial expressions in videos or images containing faces. Optionally, the expression recognition model may utilize a 3D CNN. Expression features include at least one of energetic, neutral, nervous, and sleepy.

[0108] Step 504: extract multiple physiological electrical change features of the physiological electrical signal through a deep neural network. The multiple physiological electrical change features include multiple brain wave frequency bands, electrocardiogram frequency bands, and electromyogram frequency bands. The multiple brain wave frequency bands, electrocardiogram frequency bands, and electromyogram frequency bands correspond to different fatigue states, respectively.

[0109] In this embodiment, a physiological electrical signal feature extraction network based on a deep neural network is used to extract multiple physiological electrical change features. For example, through EEG-ChannelNet, four frequency waves of δ, θ, α, and β are extracted from the brain wave signal, and their corresponding frequency bands are: 0-4Hz, 4-7Hz, 8-13Hz, and 14-30Hz, respectively. Among them, δ waves reflect excessive fatigue and deep sleep, θ waves reflect frustration or mental depression, α waves reflect a state of quietness and concentration, and β waves reflect a state of tension, emotion, or excitement. In this way, the driver's fatigue state can be determined by brain waves. Similarly, the electrocardiogram frequency band and the electromyogram frequency band correspond to different fatigue states in different segments.

[0110] In one implementation of this embodiment,

[0111] Figure 7 This is a schematic diagram of the brain activity analysis module of a deep neural network. As shown in the figure, the EEG motor sequence and feature data include EEG electrode timing data and EEG electrode feature data. The EEG electrode timing data is processed through the attention mechanism encoder module, LSTM module, and global average pooling to obtain EEG electrode timing features. The EEG electrode feature data is processed through the attention mechanism encoder module and global average pooling to obtain EEG electrode features. The EEG electrode timing features and EEG electrode features are combined and output as brain activity features, which are used to reflect the driver's mental state or fatigue from an EEG perspective.

[0112] In this embodiment, by extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data, scientific and reliable feature data is provided for driving status analysis.

[0113] Figure 8 This is a flowchart of an embodiment of the present disclosure for analyzing multi-dimensional perception features through a driving fatigue state analysis model to determine the driving state. Figure 8 Yes Figure 5 and Figure 2 The specific description of step 203 is based on Figure 8 The embodiment shown includes the following steps:

[0114] Step 801: Use a long short-term memory network to identify the driver's behavior pattern based on the visual gaze behavior data, the expression data, and the driving control data to obtain a first fatigue state. The long short-term memory network belongs to a driving fatigue state analysis model.

[0115] In this embodiment, the first fatigue state refers to the driver's fatigue state reflected in their gaze, facial expression, and driving control over a short period of time. The LSTM of the driving fatigue state analysis model identifies the driver's behavior pattern based on visual gaze behavior data, facial expression data, and driving control data to determine the first fatigue state. For example, when turning left, the driver looks forward instead of looking at the left rearview mirror, has a sleepy expression, and stops steering after turning the steering wheel slightly. The LSTM of the driving fatigue state analysis model identifies the driver's current behavior pattern as fatigued driving, and assigns the first fatigue state as severe fatigue.

[0116] Step 802: Determine the second fatigue state corresponding to the multi-dimensional perception feature through a state evaluation model. The state evaluation model belongs to a driving fatigue state analysis model.

[0117] In this embodiment, the state assessment model refers to a classification model, such as a model constructed by support vector machine (SVM), linear regression, K-nearest neighbor, etc. It is used to perform multi-classification of multi-dimensional perception features. It is a sub-model of the driving fatigue state analysis model. The second fatigue state refers to the fatigue state of the driver determined comprehensively from multi-dimensional data. For example, when the vehicle turns left, the driver's brain waves are extracted to show that the δ wave continues to appear, and the eyes are slightly closed, and the line of sight does not shift from the steering wheel for a long time. The second fatigue state is identified as severe fatigue by the state assessment model.

[0118] Step 803: Perform weighted summation on the first fatigue state and the second fatigue state to obtain a driving state.

[0119] In this embodiment, according to the first fatigue state and the second fatigue state determined in the above two steps, the first fatigue state and the second fatigue state are weighted and summed by preset weight values ​​to obtain the final driving state evaluation result for the driver. For example, the first fatigue state is severe fatigue (8, in the range of 1 to 10), and the second fatigue state is severe fatigue (7, in the range of 1 to 10). The first weight value is 70%, and the second weight value is 90%. Then the driving state s=8*70%+7*90%=0.56+0.63=1.19, and the driving state is 1 to represent the severe fatigue threshold, then s=1.19 exceeds the severe fatigue threshold, and its driving state is severe fatigue. At this time, it is necessary to warn or remind the driver through sound, vibration, light, etc., so that the driver maintains a clear driving state.

[0120] In this embodiment, the driving fatigue state analysis model is used to analyze multi-dimensional perception characteristics to determine the driving state, and the combination of the driver's multi-dimensional perception characteristics and operating behavior is used to achieve accurate detection of the fatigue state of multiple drivers.

[0121] As shown in the table below, in the comparison of results from various detection schemes, the multi-dimensional fusion model (driving fatigue state analysis model) proposed in this disclosure achieves the highest detection accuracy at the same level of driving mental workload, whether using visual fusion control behavior analysis alone or EEG data analysis alone.

[0122] Comparison table of results of multiple detection schemes

[0123] Controlling mental workload levels 1 2 3 4 5 average Visual fusion manipulation behavior analysis 87.10% 88.12% 85.54% 88.63% 87.56% 87.39% EEG data analysis 88.25% 91.31% 90.50% 89.41% 88.13% 89.82% Multidimensional fusion model 94.92% 95.13% 95.24% 94.58% 94.06% 94.73%

[0124] In one implementation of this embodiment,

[0125] Figure 9Figure 2 shows the analysis accuracy of single-module and multi-dimensional fusion models. Corresponding to the table above, across all five experimental data sets, the multi-dimensional data fusion model provided by this disclosure achieved the highest driving state detection accuracy compared to both the visual fusion operational behavior analysis model and the EEG data analysis model.

[0126] In one implementation of this embodiment,

[0127] Figure 10 This figure shows a diagram of the multi-dimensional fusion model training process. After 200 epochs of training, the driving state detection accuracy reached 95.18%, and the loss dropped to 0.14, demonstrating excellent state detection capabilities.

[0128] Figure 11 This is a flowchart of an embodiment of the present disclosure for preprocessing head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively. Figure 11 Yes Figure 5 The specific description of step 501 is based on Figure 11 The embodiment shown includes the following steps:

[0129] Step 1101: Filter the head and eye movement data and the physiological electrical signal data.

[0130] In this embodiment, to address the problem of noise interference in continuous data in multi-dimensional data, such as gaze tracking data or physiological electrical signal data, an extended Kalman filter (EKF) is used to filter the data. The EKF is an optimized autoregressive data processing algorithm that uses the kth frame data to predict the state of the k+1 frame data. If the data eigenvalue of a certain frame does not conform to the conventional rules, it is filtered. If the eigenvalue of a certain frame is weak, several Kalman predictions and corrections can be performed to enhance the eigenvalue.

[0131] In one implementation of this embodiment,

[0132] Figure 12 This is a comparison chart of EEG signal data before and after filtering. Figure 12 As shown in (a) the EEG spectrum before filtering and (b) the EEG spectrum after filtering, the collected EEG signal is subjected to spectrum analysis. It can be seen that before filtering, there is a stable interference signal at 50Hz and 80Hz throughout the entire experimental process. This interference signal is considered to be the interference of urban AC power and another electromagnetic interference source in the environment. Then, through filtering, these two interferences are filtered out, and the breathing and heartbeat interference signals are also filtered out in the process. Figure 12As shown in (c) the EEG time-frequency diagram before filtering and (d) the EEG time-frequency diagram after filtering, the interference signal is filtered.

[0133] Step 1102: filter out abnormal data in the head and eye movement data and the physiological electrical signal data.

[0134] In this embodiment, the raw head and eye movement data and physiological electrical signal data may contain a large amount of abnormal data, which needs to be cleaned and filtered out. This abnormal data can be filtered out based on prior knowledge of head and eye movement data. For example, head rotation angles within the blind spot of the visual field can be filtered out. Eye gaze movement trajectories that are too fast or too slow, or that fluctuate greatly, can also be filtered out. For example, if two gaze points on the HUD are more than 50 cm apart, the data will be filtered out.

[0135] Step 1103 , correlating and collating the head-eye movement data and / or physiological electrical signal data with the cabin scene where the driver is located, so as to retain the corresponding head-eye movement data or physiological electrical signal data obtained in the cabin scene.

[0136] In this embodiment, correlation verification involves linking sensor signal data with the simulated driving cockpit scene. For example, head and eye movement data is calibrated to align with the physical cabin environment data, and visual gaze is mapped to vehicle instrument information using regions of interest. This ensures that at least one of the head and eye movement data and physiological signal data is correlated with actual objects in the driver's cabin scene, resulting in critically coupled visual gaze data and physiological electrical data. For example, when a user turns left, the head and eye movement data indicating a leftward gaze shift and physiological electrical data indicating no significant fluctuations are detected and correlated.

[0137] Step 1104: Identify facial emotion data corresponding to the facial video data through a convolutional neural network.

[0138] In this embodiment, a convolutional neural network is used to identify the driver's facial emotion data from the facial video data. For example, from the driver's facial video data, the driver's facial expression is identified as calm.

[0139] In this embodiment, head and eye movement data, facial video data, and physiological electrical signal data are preprocessed into visual gaze data, facial emotion data, and physiological electrical data, respectively, providing a reliable and stable data source for driving state assessment. This is conducive to training a driving fatigue state analysis model.

[0140] Figure 13 This is a flowchart of an embodiment of the present disclosure that uses a long short-term memory network to identify driver behavior patterns based on visual gaze behavior data, expression data, and driving control data to obtain a first fatigue state. Figure 13 Yes Figure 8 The specific description of step 801 is based on Figure 13 The embodiment shown includes the following steps:

[0141] Step 1301: Use the encoder of the attention mechanism and global average pooling to extract the landing point sight features of the visual gaze landing point rate time series data of the visual gaze behavior data.

[0142] In this embodiment, the gaze features are extracted based on the time series data of the visual gaze rate. The region of interest gaze sequence data refers to the gaze sequence data formed when the driver's gaze point is within the region of interest when the driver's expression is serious or focused. That is, the region of interest gaze sequence data corresponds to the expression data.

[0143] Step 1302, using the encoder of the attention mechanism and global average pooling to extract the gaze landing point features of the region of interest gaze sequence data of the visual gaze behavior data, where the region of interest gaze sequence data refers to the visual gaze behavior data recorded when the expression data is a preset expression.

[0144] In this embodiment, the gaze point feature refers to the driver's gaze point feature extracted based on the ROI gaze sequence data. ROI gaze sequence data refers to the visual gaze behavior data recorded when the expression data is a preset expression. The gaze point features of the ROI gaze sequence data are extracted using an attention mechanism encoder and global average pooling. For example, when the driver's expression is focused or serious, the gaze behavior data sequence recorded can be used as the ROI gaze sequence data.

[0145] Step 1303: extract visual gaze category features of the visual gaze behavior data through a deep neural network classification model.

[0146] In this embodiment, visual gaze categories refer to visual gaze category features, such as features corresponding to categories such as distracted gaze, general gaze, and staring. Deep neural network classification models, such as LSTM and recurrent neural networks (RNN), are used to extract visual gaze categories from visual gaze behavior data. For example, the driver's gaze can be used to determine whether they are paying close attention to the dashboard.

[0147] Step 1304 , extracting control features of the driving control data through an encoder of an attention mechanism, a long short-term memory network, and global average pooling.

[0148] In this embodiment, the control feature of the driving control data refers to the feature vector obtained after the attention mechanism encoder, LSTM, and global average pooling process of the command data generated by the driver when controlling the vehicle. This feature vector is associated with the landing point gaze feature, the gaze landing point feature, and the visual gaze category feature.

[0149] Step 1305 , classifying the sight line features, the sight line features, the visual gaze category features, and the manipulation features through a multi-classification model to obtain a first fatigue state.

[0150] In this embodiment, a multi-classification model refers to a classification model used to map input data into multiple categories, such as softmax, support vector machines (SVMs), and linear regression. The landing point features, gaze landing point features, visual gaze category features, and manipulation features are merged and classified using the multi-classification model to obtain a first fatigue state.

[0151] The driving state assessment method proposed in the embodiments of the present disclosure obtains head and eye movement data, facial video data, and physiological electrical signal data. The physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals, providing a highly practical data source for driving state assessment. Multidimensional perceptual features corresponding to the head and eye movement data, facial video data, and physiological electrical signal data are determined. The fused multidimensional perceptual features provide a reliable basis for driving state assessment. The multidimensional perceptual features are analyzed using a driving fatigue state analysis model to determine the driving state and obtain an assessment of the driver's fatigue state. This improves the accuracy of driver fatigue state detection.

[0152] Corresponding to the methods provided in the above-mentioned embodiments, the present disclosure also provides a driving state evaluation device. Since the device provided in the embodiment of the present disclosure corresponds to the methods provided in the above-mentioned embodiments, the implementation method is also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.

[0153] Figure 14 FIG. 1 is a schematic diagram of the structure of a driving state evaluation device 1400 according to an embodiment of the present disclosure. Figure 14 As shown, the driving state evaluation device includes:

[0154] An acquisition module 1410 is configured to acquire head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals;

[0155] Fusion module 1420, for determining multi-dimensional perceptual features corresponding to head and eye movement data, facial video data, and physiological electrical signal data;

[0156] The determination module 1430 is configured to analyze the multi-dimensional perception features using a driving fatigue state analysis model to determine the driving state.

[0157] In some embodiments, the fusion module 1420 is configured to:

[0158] Extracting multiple visual gaze behavior features corresponding to head and eye movement data, multiple expression features corresponding to facial video data, and multiple physiological electrical change features corresponding to physiological electrical signal data;

[0159] sorting the multiple visual gaze behavior features, the multiple facial expression features, and the multiple physiological electrical change features in chronological order to form visual gaze behavior data, facial expression data, and physiological electrical change data;

[0160] The visual gaze behavior data, facial expression data, and physiological electrical change data are integrated into multi-dimensional perception features.

[0161] In some embodiments, the fusion module 1420 extracts multiple visual gaze behavior features corresponding to the head and eye movement data, multiple expression features corresponding to the facial video data, and multiple physiological electrical change features corresponding to the physiological electrical signal data in the following manner:

[0162] Preprocessing the head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively. The visual gaze data includes at least one of the first gaze duration, the second gaze duration, the number of gazes, the gaze duration, the pupil diameter, and the saccade distance. The first gaze duration is the duration of the driver's first stay on the gaze point, the second gaze duration is the duration of the second stay on the gaze point, the number of gazes is the number of times the gaze point has been stayed, and the gaze duration is the total duration of the gaze point stay.

[0163] Acquire driving control data, and extract a plurality of visual gaze behavior features through a first feature extraction network based on the visual gaze data and the driving control data;

[0164] Performing expression recognition on the facial emotion data using an expression recognition model to extract multiple expression features, wherein the multiple expression features include at least one of energetic, neutral, nervous, and sleepy;

[0165] Multiple physiological electrical change features of physiological electrical signals are extracted through deep neural networks. The multiple physiological electrical change features include multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands. The multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands correspond to different fatigue states respectively.

[0166] In some embodiments, the determination module 1430 is configured to:

[0167] Using a long short-term memory network, visual gaze behavior data, facial expression data, and driving control data are used to identify driver behavior patterns and obtain the first fatigue state. The long short-term memory network belongs to the driving fatigue state analysis model;

[0168] Determine the second fatigue state corresponding to the multi-dimensional perception characteristics through the state assessment model, which belongs to the driving fatigue state analysis model;

[0169] A weighted sum is performed on the first fatigue state and the second fatigue state to obtain a driving state.

[0170] In some embodiments, the fusion module 1420 pre-processes the head and eye movement data, facial video data, and physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively, in the following manner:

[0171] Filtering head and eye movement data and physiological electrical signal data;

[0172] Screen out abnormal data in head and eye movement data and physiological electrical signal data;

[0173] Correlating and collating the head-eye movement data and / or physiological electrical signal data with the vehicle cabin scene where the driver is located, so as to retain the corresponding head-eye movement data or physiological electrical signal data obtained in the vehicle cabin scene;

[0174] Recognize facial emotion data corresponding to facial video data through convolutional neural networks.

[0175] In some embodiments, the determination module 1430 uses a long short-term memory network to perform driver behavior pattern recognition on the visual gaze behavior data, the facial expression data, and the driving control data to obtain the first fatigue state in the following manner:

[0176] The encoder of the attention mechanism and global average pooling are used to extract the gaze features of the visual gaze landing rate time series data of the visual gaze behavior data;

[0177] The encoder of the attention mechanism and global average pooling are used to extract the gaze landing point features of the region of interest gaze sequence data of the visual gaze behavior data. The region of interest gaze sequence data refers to the visual gaze behavior data recorded when the expression data is a preset expression.

[0178] Extract visual gaze category features of visual gaze behavior data through a deep neural network classification model;

[0179] Extracting control features of driving control data through attention mechanism encoder, long short-term memory network, and global average pooling;

[0180] Through the multi-classification model, the landing point sight features, sight gaze landing point features, visual gaze category features and manipulation features are classified to obtain the first fatigue state.

[0181] In summary, the driving state assessment device acquires head and eye movement data, facial video data, and physiological electrical signal data, including EEG, ECG, and EMG signals. Multidimensional perceptual features corresponding to these data are determined. The driving fatigue state analysis model analyzes these multidimensional perceptual features to determine the driving state. This device addresses the issue of low driving state detection accuracy and improves it.

[0182] The embodiments provided above in this disclosure describe the methods and devices provided in these embodiments. To implement the various functions in the methods provided in these embodiments, electronic devices may include hardware structures and software modules, and implement these functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Certain of these functions may be implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules.

[0183] Figure 15 is a block diagram of an electronic device 1500 for implementing the above-mentioned driving state evaluation method according to an exemplary embodiment.

[0184] For example, the electronic device 1500 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0185] Reference Figure 15 , the electronic device 1500 may include one or more of the following components: a processing component 1502 , a memory 1504 , a power component 1506 , a multimedia component 1508 , an audio component 1510 , an input / output (I / O) interface 1512 , a sensor component 1514 , and a communication component 1516 .

[0186] The processing component 1502 generally controls the overall operation of the electronic device 1500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 1502 may include one or more processors 1520 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1502 may include one or more modules to facilitate interaction between the processing component 1502 and other components. For example, the processing component 1502 may include a multimedia module to facilitate interaction between the multimedia component 1508 and the processing component 1502.

[0187] The memory 1504 is configured to store various types of data to support operations on the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 1500, contact data, phone book data, messages, pictures, videos, etc. The memory 1504 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0188] The power supply component 1506 provides power to the various components of the electronic device 1500. The power supply component 1506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1500.

[0189] The multimedia component 1508 includes a screen that provides an output interface between the electronic device 1500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1508 includes a front camera and / or a rear camera. When the electronic device 1500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0190] The audio component 1510 is configured to output and / or input audio signals. For example, the audio component 1510 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1504 or transmitted via the communication component 1516. In some embodiments, the audio component 1510 also includes a speaker for outputting audio signals.

[0191] I / O interface 1512 provides an interface between processing component 1502 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.

[0192] The sensor assembly 1514 includes one or more sensors for providing various aspects of the status assessment of the electronic device 1500. For example, the sensor assembly 1514 can detect the open / closed state of the electronic device 1500, the relative positioning of components, such as the display and keypad of the electronic device 1500. The sensor assembly 1514 can also detect changes in the position of the electronic device 1500 or a component of the electronic device 1500, the presence or absence of user contact with the electronic device 1500, the orientation or acceleration / deceleration of the electronic device 1500, and changes in the temperature of the electronic device 1500. The sensor assembly 1514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1514 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1514 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0193] The communication component 1516 is configured to facilitate wired or wireless communication between the electronic device 1500 and other devices. The electronic device 1500 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 1516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0194] In an exemplary embodiment, the electronic device 1500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0195] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1504 including instructions, which can be executed by the processor 1520 of the electronic device 1500 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0196] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the driving state assessment method described in the above embodiment of the present disclosure.

[0197] An embodiment of the present disclosure further provides a computer program product, including a computer program, which executes the driving state assessment method described in the above embodiment of the present disclosure when a processor executes the computer program.

[0198] The embodiment of the present disclosure also provides a vehicle, comprising Figure 14 The driving state evaluation device shown or Figure 15 Electronic devices shown.

[0199] Figure 16 FIG. 1 is a schematic structural diagram of a chip 1600 for implementing the above-mentioned driving state evaluation method according to an exemplary embodiment.

[0200] Reference Figure 16 The chip 1600 includes at least one communication interface 1601 and a processor 1602; the communication interface 1601 is used to receive signals input to the chip 1600 or signals output from the chip 1600, and the processor 1602 communicates with the communication interface 1601 and implements the driving state evaluation method described in the above embodiment through logic circuits or execution code instructions.

[0201] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0202] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with an embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, the illustrative use of the above terms does not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0203] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0204] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0205] It should be understood that the various parts of the embodiments of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0206] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0207] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.

[0208] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A driving state evaluation method, characterized in that: The method comprises: Acquiring head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals; Determining multi-dimensional perceptual features corresponding to the head-eye movement data, the facial video data, and the physiological electrical signal data; The multi-dimensional perception characteristics are analyzed by a driving fatigue state analysis model to determine the driving state.

2. The method according to claim 1, characterized in that The determining of the multi-dimensional perceptual features corresponding to the head-eye movement data, the facial video data, and the physiological electrical signal data includes: Extracting a plurality of visual gaze behavior features corresponding to the head and eye movement data, a plurality of expression features corresponding to the facial video data, and a plurality of physiological electrical change features corresponding to the physiological electrical signal data; Sort the multiple visual gaze behavior features, the multiple expression features, and the multiple physiological electrical change features in chronological order to form visual gaze behavior data, expression data, and physiological electrical change data; The visual gaze behavior data, the facial expression data, and the physiological electrical change data are fused into the multi-dimensional perception feature.

3. The method according to claim 2, characterized in that The extracting of multiple visual gaze behavior features corresponding to the head and eye movement data, multiple expression features corresponding to the facial video data, and multiple physiological electrical change features corresponding to the physiological electrical signal data includes: Preprocessing the head and eye movement data, the facial video data, and the physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively, wherein the visual gaze data includes at least one of a first gaze point duration, a second gaze point duration, a number of gazes, a gaze duration, a pupil diameter, and a saccade distance, wherein the first gaze point duration is the duration of the driver's first stay at the gaze point, the second gaze point duration is the duration of the second stay at the gaze point, the number of gazes is the number of times the gaze point has stayed, and the gaze duration is the total duration of the stay at the gaze point; Acquire driving control data, and extract the plurality of visual gaze behavior features through a first feature extraction network based on the visual gaze data and the driving control data; Performing expression recognition on the facial emotion data using an expression recognition model to extract the plurality of expression features, wherein the plurality of expression features include at least one of energetic, neutral, nervous, and sleepy; The multiple physiological electrical change features of the physiological electrical signal are extracted through a deep neural network, and the multiple physiological electrical change features include multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands, and the multiple brain wave frequency bands, electrocardiogram frequency bands and electromyogram frequency bands correspond to different fatigue states respectively.

4. The method according to claim 3, characterized in that Analyzing the multi-dimensional perception features using a driving fatigue state analysis model to determine the driving state includes: performing driver behavior pattern recognition on the visual gaze behavior data, the facial expression data, and the driving control data using a long short-term memory network to obtain a first fatigue state, the long short-term memory network belonging to the driving fatigue state analysis model; determining a second fatigue state corresponding to the multidimensional perception feature through a state assessment model, wherein the state assessment model belongs to the driving fatigue state analysis model; A weighted sum is performed on the first fatigue state and the second fatigue state to obtain the driving state.

5. The method according to claim 3, characterized in that The preprocessing of the head-eye movement data, the facial video data, and the physiological electrical signal data into visual gaze data, facial emotion data, and physiological electrical data, respectively, includes: filtering the head-eye movement data and the physiological electrical signal data; Screening out abnormal data in the head-eye movement data and the physiological electrical signal data; Correlating and collating the head-eye movement data and / or the physiological electrical signal data with the vehicle cabin scene where the driver is located, so as to retain the corresponding head-eye movement data or the physiological electrical signal data obtained in the vehicle cabin scene; The facial emotion data corresponding to the facial video data is identified through a convolutional neural network.

6. The method according to claim 4, characterized in that The method of using a long short-term memory network to perform driver behavior pattern recognition on the visual gaze behavior data, the facial expression data, and the driving control data to obtain a first fatigue state includes: Extracting the gaze features of the visual gaze landing rate time series data of the visual gaze behavior data by using an encoder of the attention mechanism and global average pooling; Extracting gaze landing point features of region-of-interest gaze sequence data of the visual gaze behavior data using the encoder of the attention mechanism and the global average pooling, wherein the region-of-interest gaze sequence data refers to the visual gaze behavior data recorded when the expression data is a preset expression; extracting visual gaze category features of the visual gaze behavior data through a deep neural network classification model; Extracting control features of the driving control data through an encoder of an attention mechanism, a long short-term memory network, and global average pooling; The landing point sight feature, the sight gaze landing point feature, the visual gaze category feature and the manipulation feature are classified through a multi-classification model to obtain the first fatigue state.

7. A driving state evaluation device, characterized in that: The device comprises: An acquisition module, configured to acquire head and eye movement data, facial video data, and physiological electrical signal data, wherein the physiological electrical signal data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electromyography (EMG) signals; a fusion module, configured to determine multi-dimensional perceptual features corresponding to the head-eye movement data, the facial video data, and the physiological electrical signal data; The determination module is used to analyze the multi-dimensional perception characteristics through a driving fatigue state analysis model to determine the driving state.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

10. A vehicle, characterized in that: Includes the driving state evaluation device according to claim 7 or the electronic device according to claim 8.