Fatigue driving detection method and device and electronic equipment

By comparing and learning the characteristics of the fusion of historical biological information, historical eye movement information and historical visual images, a target fatigue driving detection model is formed, which solves the problem of low accuracy of detection results in the prior art and achieves more accurate driver fatigue driving detection.

CN120182950APending Publication Date: 2025-06-20ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510165235.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art fails to effectively learn data similarity and differences between different modes when detecting whether the driver is driving tired, resulting in low accuracy of the detection results.

Method used

A target fatigue driving detection method is adopted, by obtaining the driver's biological information, eye movement information and visual images, using the historical target fusion characteristics after fusing historical biological information and historical eye movement information, and comparing and learning with the historical image cross-frame characteristics of the extracted historical visual image to form a target fatigue driving detection model.

Benefits of technology

By learning the commonality and differences of different modal features, the performance of the detection model is improved, making the driver fatigue driving detection results based on this model more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182950A_ABST
    Figure CN120182950A_ABST
Patent Text Reader

Abstract

The invention discloses a fatigue driving detection method and device and electronic equipment. The method comprises the steps that biological information, eye movement information and a visual image of a to-be-detected driver in a T frame are acquired; the biological information, the eye movement information and the visual image are processed through a target fatigue driving detection model to obtain a detection result whether fatigue driving exists or not, and the target fatigue driving detection model is obtained by fusing historical biological information and historical eye movement information through the fatigue driving detection model. And comparing and learning with the extracted historical image cross-frame features of the historical visual image. Through the technical scheme provided by the embodiment of the invention, multi-modal information is fused, a comparative learning mode is adopted, and multi-modal commonality and difference characteristics are learned, so that the performance of the target fatigue driving detection model is optimal, and a fatigue driving detection result based on the target fatigue driving detection model is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation technology, and particularly to a method, device, and electronic device for detecting fatigue driving. Background Art

[0002] With the rapid development of the national economy, people's living standards have been increasing day by day. There are more and more vehicles on the road, and traffic accidents have become a social problem that cannot be ignored. Among the causes of traffic accidents, driver fatigue driving is an important inducement. Therefore, detecting whether a driver is fatigued is an important problem to be solved.

[0003] When the prior art detects whether a driver is fatigued, it can detect based on information such as the driver's electroencephalogram signal. Specifically, the physiological information such as the driver's electroencephalogram signal and the visual image are fused, and then the driver's state is analyzed through the fused information to determine whether the driver is fatigued.

[0004] However, the above method only fuses biological information and visual images, and does not learn the similarities and differences between data in different modalities (i.e., biological information and visual images), resulting in low accuracy of the detection result of driver fatigue driving. Summary of the Invention

[0005] This application provides a method, device, and electronic device for detecting fatigue driving, which is used to solve the problem that the accuracy of the detection result of whether a driver is fatigued in the prior art is low because the detection method cannot learn the similarities and differences between data in different modalities. The specific implementation solutions are as follows:

[0006] In a first aspect, this application provides a method for detecting fatigue driving, the method comprising:

[0007] Obtain the biological information, eye movement information, and visual image of the driver to be detected within T frames;

[0008] Process the biological information, the eye movement information, and the visual image through a target fatigue driving detection model to obtain a detection result for indicating that the driver to be detected is fatigued or not fatigued; wherein, the target fatigue driving detection model is obtained by contrastive learning of the historical target fusion feature obtained by fusing historical biological information and historical eye movement information and the historical image cross-frame feature of the extracted historical visual image.

[0009] Through the above application embodiments, the biological information, eye movement information, and visual images of the driver to be detected are processed using the target fatigue driving detection model to implement the detection of whether the driver to be detected is fatigued. Since the target fatigue driving detection model is obtained by the fatigue driving detection model through contrastive learning of the historical target fusion features that fuse multi-modal information and the extracted historical image cross-frame features, during the contrastive learning process, the fatigue driving detection model can learn the commonalities and differences of different modal features, making the performance of the obtained target fatigue driving detection model better. Therefore, when detecting whether a driver is fatigued based on the target fatigue driving detection model with the best performance, the detection result is more accurate.

[0010] In a possible implementation manner, the target fatigue driving detection model includes a first network, a second network, a third network, and a classification network;

[0011] Processing the biological information, the eye movement information, and the visual image through the target fatigue driving detection model to obtain a detection result for indicating that the driver to be detected is fatigued or not fatigued includes:

[0012] Fusing the biological information and the eye movement information through the first network to obtain target fusion features; and extracting the image cross-frame features of the visual image through the second network;

[0013] Processing the target fusion features and the image cross-frame features through the third network to obtain target features;

[0014] Processing the target features through the classification network to obtain the detection result for indicating that the driver to be detected is fatigued or not fatigued.

[0015] Through the above application embodiments, multi-modal information fusion is performed based on the first network, the second network, the third network, and the classification network in the target fatigue driving detection model, thereby achieving accurate detection of whether the driver to be detected is fatigued. Compared with using a single piece of information to detect whether a driver is fatigued, the information used is more comprehensive, making the obtained detection result more accurate.

[0016] In a possible implementation manner, the first network includes a first multi-layer perceptron MLP;

[0017] Fusing the biological information and the eye movement information through the first network to obtain target fusion features includes:

[0018] Obtaining the global fusion features between the biological information and the eye movement information, and the local fusion features between the biological information and the eye movement information;

[0019] Concatenate the global fusion feature and the local fusion feature to obtain a concatenated fusion feature;

[0020] Through the first MLP, perform feature extraction on the concatenated fusion feature to obtain the target fusion feature.

[0021] Through the above application embodiments, first concatenate the global fusion feature and the local fusion feature between the biological information and the eye movement information, and then process the concatenated fusion feature after concatenation through the first MLP to obtain the finally fused target fusion feature between the biological information and the eye movement information, thereby realizing the fusion of the biological information and the eye movement information, and making the fused target fusion feature consider both the global fusion feature between the biological information and the eye movement information and the local fusion feature between the biological information and the eye movement information, so that the information contained in the target fusion feature is more comprehensive and more accurate, which helps to make the contrast learning based on the target fusion feature more accurate.

[0022] In a possible implementation manner, the first network further includes a temporal encoder and a second MLP;

[0023] The obtaining of the global fusion feature between the biological information and the eye movement information includes:

[0024] Through the temporal encoder, perform global feature extraction on the biological information and the eye movement information respectively to obtain a global biological feature and a global eye movement feature;

[0025] Concatenate the global biological feature and the global eye movement feature to obtain a global concatenated feature;

[0026] Through the second MLP, perform feature extraction on the global concatenated feature and perform pooling dimensionality reduction on the extracted feature to obtain the global fusion feature.

[0027] Through the above application embodiments, based on the temporal encoder, concatenation processing, and the second MLP, the global features of the biological information and the eye movement information at all times are fused, so that the obtained global fusion feature can better reflect the global features of the biological information and the eye movement information, and thus the obtained global fusion feature is more comprehensive and more accurate.

[0028] In a possible implementation manner, the first network further includes a first cross-attention CA and a third MLP;

[0029] The obtaining of the local fusion feature between the biological information and the eye movement information includes:

[0030] Through the first CA, the biological information and the eye movement information are interactively processed to obtain critical moment features;

[0031] Through the third MLP, feature extraction is performed on the critical moment features to obtain the local fusion features.

[0032] Through the above application embodiments, based on the first CA and the third MLP, the local features of biological information and eye movement information at critical moments are fused, so that the obtained local fusion features can more accurately reflect the critical features of biological information and eye movement information at critical moments.

[0033] In a possible implementation manner, the second network includes H sub-networks, an average pooling layer AP, and a multi-head attention MHA, where H is a positive integer, and the image cross-frame features include first cross-frame features corresponding to image patches and second cross-frame features corresponding to class tokens;

[0034] The extraction of the image cross-frame features of the visual image through the second network includes:

[0035] Through block embedding, multiple frames of the visual image are respectively segmented into a patches to obtain patch features corresponding to each of the multiple frames of the visual image;

[0036] Based on the patch features, the initial class token, and the position encoding, the first initial input of the initial layer sub-network in the second network is determined, and the current layer number is determined as the first preset value;

[0037] Based on the first initial input, the patch output and the token output of the corresponding layer sub-network are obtained, and the layer number is added with the second preset value to obtain a new layer number, and it is determined whether the new layer number is less than or equal to H;

[0038] If so, the patch output is concatenated with the token fusion feature of the next layer sub-network to obtain the second initial input of the next layer sub-network; the first initial input is replaced with the second initial input, and based on the replaced first initial input, the patch output and the token output of the corresponding layer sub-network are obtained, the layer number is replaced with the new layer number, and based on the replaced layer number, a new layer number is determined, and it is determined whether the new layer number is less than or equal to H;

[0039] If not, the patch output of the obtained Hth layer sub-network is used as the first cross-frame feature, and based on the AP and the MHA, the processing of the sum of the token output of the Hth layer sub-network and the time position encoding is performed to obtain the second cross-frame feature.

[0040] Through the above application embodiments, based on the H-layer sub-network in the second network, the cross-frame features of visual images are extracted, the information before different frames of visual images is fused, and the connection before and after actions is considered, making the extracted cross-frame features of images (i.e., the first cross-frame feature and the second cross-frame feature) more accurate, which helps to make the detection of whether a driver is fatigued based on the cross-frame features of images more accurate.

[0041] In a possible implementation manner, the second network further includes a second CA and a first feed-forward neural network FNN;

[0042] The obtaining of the patch output and the token output of the corresponding layer sub-network based on the first initial input includes:

[0043] Normalize the first initial input of the corresponding frame to obtain a normalized input, and add the first initial input to the feature obtained by the second CA processing the normalized input to obtain the initial output of the corresponding frame in the initial layer sub-network; wherein, the initial output includes an initial patch output and the token output;

[0044] Add the initial patch output to the feature obtained by the first FFN processing the initial patch output to obtain the patch output of the corresponding frame in the initial layer sub-network.

[0045] Through the above application embodiments, first normalize the first initial input, then process the normalized input after normalization by the second CA, and then add the first initial input, making the initial output of the obtained initial layer sub-network more accurate. Moreover, add the feature obtained by the first FFN processing the initial patch output to the initial patch output in the initial output to obtain the patch output of the initial layer sub-network, thereby making the patch output corresponding to the partial information of the patch more accurate.

[0046] In a possible implementation manner, the second network further includes a linear unit and a first self-attention SA;

[0047] Before splicing the patch output and the token fusion feature of the next layer sub-network, it further includes:

[0048] Process the class tokens of the multi-frame visual images in the previous layer sub-network of the next layer sub-network through the linear unit to obtain the token information of the next layer sub-network;

[0049] Normalize the token information to obtain token normalized information;

[0050] Process the token normalization information through the first SA to obtain the token fusion feature of the next-layer sub-network.

[0051] Through the above application embodiments, based on the processing of the class token of the previous-layer sub-network of the next-layer sub-network, the token fusion feature of the next-layer sub-network is obtained, so that the global information from all patches is continuously aggregated through the class token, facilitating the formation of a feature vector containing rich context information. Furthermore, it helps the second network to better extract the cross-frame features of the visual image based on the token fusion feature.

[0052] In a possible implementation manner, the third network includes a feature interaction module, a cross-modal decoder, and a fourth MLP. The feature interaction module is a module composed of a second SA and a bidirectional attention BiA, and the cross-modal decoder is a module composed of a third CA, a fourth CA, and a second FFN.

[0053] The processing of the target fusion feature and the cross-frame feature of the visual image through the third network to obtain the target feature includes:

[0054] In the time dimension, perform average processing on the first cross-frame feature in the cross-frame feature of the visual image to obtain the cross-frame average feature.

[0055] Through the feature interaction module, perform interaction processing on the cross-frame average feature and the target fusion feature to obtain a first interaction feature and a second interaction feature.

[0056] Through the cross-modal decoder, process the first interaction feature, the second interaction feature, and the second cross-frame feature in the cross-frame feature of the visual image to obtain an intermediate feature.

[0057] Through the fourth MLP, process the intermediate feature to obtain the target feature.

[0058] Through the above application embodiments, the target fusion feature and the image cross-frame feature are processed based on time-dimensional average processing, a feature interaction module, a cross-modal decoder, and a fourth MLP, thereby fusing different information, making the obtained target feature better reflect the information of the driver to be detected, so as to further improve the accuracy of the detection result obtained based on the target feature. Moreover, through the feature interaction module composed of the second SA and BiA, the cross-frame average feature and the target fusion feature are interactively processed, so that the information between the cross-frame average feature and the target fusion feature can be more fully interacted, which helps the target fatigue driving detection model to better understand and learn the correlation and mutual influence between the cross-frame average feature and the target fusion feature, and further helps to improve the accuracy of the detection result. In addition, through the cross-modal decoder composed of the third CA, the fourth CA, and the second FFN, the first interaction feature, the second interaction feature, and the second cross-frame feature are processed, so that the first interaction feature, the second interaction feature, and the second cross-frame feature can be more fully fused, making the fused intermediate feature richer and more accurate, and further helping to improve the accuracy of the detection result.

[0059] In a possible implementation manner, before the target fatigue driving detection model processes the biological information, the eye movement information, and the visual image to obtain a detection result indicating that the driver to be detected is fatigued or non-fatigued, it further includes:

[0060] Obtain the historical biological information, the historical eye movement information, and the historical visual image;

[0061] Through the first network in the fatigue driving detection model, fuse the historical biological information and the eye movement information to obtain the historical target fusion feature; and, through the second network in the fatigue driving detection model, extract the historical image cross-frame feature of the historical visual image;

[0062] Through the third network in the fatigue driving detection model, process the historical target fusion feature and the historical image cross-frame feature to obtain a historical target feature, a first historical interaction feature, and a second historical interaction feature;

[0063] According to the historical target feature and the classification loss function, calculate the classification loss value; and, according to the first historical interaction feature, the second historical interaction feature, and the contrast loss function, calculate the contrast loss value;

[0064] Take the sum of the classification loss value and the contrast loss value as the target loss value, and train the fatigue driving detection model according to the minimization of the target loss value to obtain the target fatigue driving detection model.

[0065] Through the above application embodiments, the first network, the second network, and the third network of the fatigue driving detection model process historical biological information, historical eye movement information, and historical visual images, integrating different information, so as to facilitate the fatigue driving detection model to learn the information in the semantic space. Then, according to the obtained historical target features, the first historical interaction feature, and the second historical interaction feature, the target loss value is calculated, making the calculated target loss value more accurate, so as to train the fatigue driving detection model according to the minimization of the target loss value, and making the performance of the obtained target fatigue driving detection model optimal.

[0066] Moreover, the target loss value for training the fatigue driving detection model is calculated based on the classification loss value and the contrast loss value, making the obtained target loss value take into account both the classification loss value and the contrast loss value, so that the fatigue driving detection model can better learn the common features of drivers in the fatigued driving state. In addition, the contrast loss function is used to perform contrast learning on the first historical interaction feature and the second historical interaction feature, so that the obtained target fatigue driving detection model is more accurate, and further making the detection result of whether the to-be-detected driving is fatigued driving based on the optimal target fatigue driving detection model more accurate.

[0067] In a second aspect, the present application also provides a fatigue driving detection device, and the device includes:

[0068] An acquisition module, configured to acquire the biological information, eye movement information, and visual images of the to-be-detected driver within T frames;

[0069] A processing module, configured to process the biological information, the eye movement information, and the visual images through the target fatigue driving detection model to obtain a detection result for indicating whether the to-be-detected driver is fatigued driving or not; wherein, the target fatigue driving detection model is obtained by performing contrast learning on the historical target fusion feature obtained by fusing the historical biological information and the historical eye movement information by the fatigue driving detection model and the historical image cross-frame feature of the extracted historical visual image.

[0070] In a possible implementation manner, the target fatigue driving detection model includes a first network, a second network, a third network, and a classification network; the processing module is specifically configured to fuse the biological information and the eye movement information through the first network to obtain a target fusion feature; and extract the image cross-frame feature of the visual image through the second network; process the target fusion feature and the image cross-frame feature through the third network to obtain a target feature; and process the target feature through the classification network to obtain the detection result for indicating whether the to-be-detected driver is fatigued driving or not.

[0071] In a possible implementation, the first network includes a first multi-layer perceptron (MLP); the processing module is specifically configured to obtain the global fusion feature between the biological information and the eye movement information, and the local fusion feature between the biological information and the eye movement information; splice the global fusion feature and the local fusion feature to obtain a spliced fusion feature; and perform feature extraction on the spliced fusion feature through the MLP to obtain the target fusion feature.

[0072] In a possible implementation, the first network further includes a temporal encoder and a second MLP; the processing module is further configured to perform global feature extraction on the biological information and the eye movement information respectively through the temporal encoder to obtain a global biological feature and a global eye movement feature; splice the global biological feature and the global eye movement feature to obtain a global spliced feature; and perform feature extraction on the global spliced feature through the second MLP and perform pooling dimensionality reduction on the extracted feature to obtain the global fusion feature.

[0073] In a possible implementation, the first network further includes a first cross-attention (CA) and a third MLP; the processing module is further configured to perform interaction processing on the biological information and the eye movement information through the first CA to obtain a critical moment feature; and perform feature extraction on the critical moment feature through the third MLP to obtain the local fusion feature.

[0074] In a possible implementation, the second network includes an H-layer sub-network, an average pooling layer AP, and a multi-head attention MHA, where H is a positive integer. The cross-frame image features include first cross-frame features corresponding to image patches (patches) and second cross-frame features corresponding to class tokens. The processing module is further configured to, through patch embedding, respectively divide multiple frames of the visual images into a patches to obtain patch features corresponding to each of the multiple frames of visual images; based on the patch features, an initial class token, and positional encoding, determine a first initial input of an initial layer sub-network in the second network, and determine that the current layer number is a first preset value; based on the first initial input, obtain a patch output and a token output of the corresponding layer sub-network, add a second preset value to the layer number to obtain a new layer number, and determine whether the new layer number is less than or equal to H; if so, concatenate the patch output and a token fusion feature of the next layer sub-network to obtain a second initial input of the next layer sub-network; replace the first initial input with the second initial input, and based on the replaced first initial input, obtain a patch output and a token output of the corresponding layer sub-network, replace the layer number with the new layer number, and based on the replaced layer number, determine a new layer number and determine whether the new layer number is less than or equal to H; if not, use the patch output of the obtained H-th layer sub-network as the first cross-frame feature, and based on the AP and the MHA, process the sum of the token output of the H-th layer sub-network and the temporal positional encoding to obtain the second cross-frame feature.

[0075] In a possible implementation, the second network further includes a second CA and a first feed-forward neural network FNN; the processing module is further configured to normalize the first initial input of the corresponding frame to obtain a normalized input, and add the feature obtained by the second CA processing the normalized input to the first initial input to obtain an initial output of the corresponding frame in the initial layer sub-network; where the initial output includes an initial patch output and the token output; add the feature obtained by the first FFN processing the initial patch output to the initial patch output to obtain the patch output of the corresponding frame in the initial layer sub-network.

[0076] In a possible implementation manner, the second network further includes a linear unit and a first self-attention SA; the processing module is further configured to process the class tokens of multiple frames of the visual images in the previous sub-network of the next sub-network through the linear unit to obtain the token information of the next sub-network; normalize the token information to obtain token normalization information; and process the token normalization information through the first SA to obtain the token fusion feature of the next sub-network.

[0077] In a possible implementation manner, the third network includes a feature interaction module, a cross-modal decoder, and a fourth MLP. The feature interaction module is a module composed of a second SA and a bidirectional attention BiA, and the cross-modal decoder is a module composed of a third CA, a fourth CA, and a second FFN; the processing module is further configured to perform an average process on the first cross-frame feature in the image cross-frame feature in the time dimension to obtain a cross-frame average feature; perform an interaction process on the cross-frame average feature and the target fusion feature through the feature interaction module to obtain a first interaction feature and a second interaction feature; process the first interaction feature, the second interaction feature, and the second cross-frame feature in the image cross-frame feature through the cross-modal decoder to obtain an intermediate feature; and process the intermediate feature through the fourth MLP to obtain the target feature.

[0078] In a possible implementation manner, the device further includes a training module, which is specifically configured to obtain the historical biological information, the historical eye movement information, and the historical visual images; fuse the historical biological information and the eye movement information through the first network in the fatigue driving detection model to obtain the historical target fusion feature; and extract the historical image cross-frame feature of the historical visual images through the second network in the fatigue driving detection model; process the historical target fusion feature and the historical image cross-frame feature through the third network in the fatigue driving detection model to obtain a historical target feature, a first historical interaction feature, and a second historical interaction feature; calculate a classification loss value according to the historical target feature and a classification loss function; and calculate a contrast loss value according to the first historical interaction feature, the second historical interaction feature, and a contrast loss function; use the sum of the classification loss value and the contrast loss value as the target loss value, and train the fatigue driving detection model according to the minimization of the target loss value to obtain the target fatigue driving detection model.

[0079] In a third aspect, the present application provides an electronic device, including:

[0080] A memory for storing a computer program;

[0081] A processor, when executing the computer program stored on the memory, implements the steps of the above-mentioned method for detecting fatigue driving.

[0082] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for detecting fatigue driving are implemented.

[0083] For the various aspects in the second to fourth aspects above and the possible technical effects that each aspect may achieve, please refer to the description of the possible technical effects that can be achieved for the first aspect or various possible solutions in the first aspect above, and details will not be repeated here. Description of the Drawings

[0084] Figure 1a Schematic diagram of the first network in the fatigue driving detection model provided by the embodiment of the present application;

[0085] Figure 1b Schematic diagram of the second network in the fatigue driving detection model provided by the embodiment of the present application;

[0086] Figure 1c Schematic diagram of the third network in the fatigue driving detection model provided by the embodiment of the present application;

[0087] Figure 2 Schematic flowchart of a method for detecting fatigue driving provided by the embodiment of the present application;

[0088] Figure 3 Schematic flowchart of the second network extracting cross-frame features of an image provided by the embodiment of the present application;

[0089] Figure 4 Schematic diagram of a device for detecting fatigue driving provided by the embodiment of the present application;

[0090] Figure 5 Schematic diagram of an electronic device provided by the embodiment of the present application. Detailed Embodiments

[0091] In order to make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail in conjunction with the accompanying drawings. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. It should be noted that in the description of this application, "a plurality of" is understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The connection between A and B can represent: A is directly connected to B and A is connected to B through C. In addition, in the description of this application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.

[0092] The following will describe the embodiments of this application in detail in conjunction with the accompanying drawings.

[0093] When the prior art detects whether a driver is fatigued based on information such as the driver's electroencephalogram signal, it does not learn the similarities and differences between data in different modalities, so it cannot learn a more efficient data representation, resulting in a low accuracy of the detection result of the driver's fatigue driving.

[0094] Therefore, this application proposes a method for detecting fatigue driving. By using the historical target fusion features obtained after fusing the historical biological information and historical eye movement information in the training data, and comparing and learning with the historical image cross-frame features of the historical visual images in the training data, a target fatigue driving detection model is obtained. The biological information, eye movement information, and visual images of the driver to be detected within T frames are processed to obtain a detection result indicating whether the driver to be detected is fatigued or not. In this way, by using the historical target fusion features that fuse multi-modal information (such as fusing historical biological information and historical eye movement information) and the extracted historical image cross-frame features for comparative learning, the fatigue driving detection model can learn the commonalities and differences of different modal features, making the performance of the obtained target fatigue driving detection model better. Furthermore, when using the target fatigue driving detection model with the best performance to process the biological information, eye movement information, and visual images of the driver to be detected within T frames to detect whether the driver is fatigued, the detection result is more accurate.

[0095] The above-mentioned target fatigue driving detection model may include a first network, a second network, a third network, and a classification network. The target fatigue driving detection model is a model obtained after training the fatigue driving detection model. That is to say, the structure of the fatigue driving detection model is the same as that of the target fatigue driving detection model, and it also includes a first network, a second network, a third network, and a classification network. Among them, the parameters of the first network, the second network, the third network, and the classification network in the target driving detection model can be obtained through the training of the fatigue driving detection model.

[0096] The first network in the above-mentioned target fatigue driving detection model (or fatigue driving detection model) may include a global fusion module, a local fusion module, and a global and local fusion module. The global fusion module may include a temporal encoder and a second multi-layer perceptron (English: Multilayer Perceptron, abbreviated as MLP). The temporal encoder may be a long short-term memory network (English: Long Short-Term Memory, abbreviated as LSTM), or a Transformer model, but is not limited thereto, and can be adjusted according to specific application scenarios. The local fusion module may include a first cross-attention (English: Cross-Attention, abbreviated as CA) and a third MLP. The global and local fusion module may include a first MLP.

[0097] Specifically, as Figure 1a shown, in the global fusion module of the first network, first, through the temporal encoder, global feature extraction is respectively performed on the biological information (i.e., Figure 1a X1 in Figure 1a ) and the eye movement information (i.e., Figure 1a X2 in Figure 1a ) to obtain the global biological feature of the biological information and the eye movement information (i.e., Figure 1a X1' in Figure 1a ) and the global eye movement feature (i.e., Figure 1a X2' in

[0098] ). The global biological feature and the global eye movement feature may be the global features of the biological information and the eye movement information at multiple moments; they may also be the global features of the biological information and the eye movement information at all moments, so as to better reflect the global information of the biological information and the eye movement information.

[0098] Among them, the biological information is the information obtained by normalizing the collected original biological information; the eye movement information is the information obtained by normalizing the collected original eye movement information; X1 ∈ R 2×M×T , X2 ∈ R 8×N×T . T represents time or the number of frames, M represents the number of biological information channels, and N represents the number of eye movement information channels.

[0099] In the embodiments of the present application, the above-mentioned original biological information (or normalized biological information) may include information such as the heart rate (abbreviated as HR in English: Heart Rate) and breath (abbreviated as BR in English: Breath) of a driver (such as a driver to be detected). The collection of the original biological information can be carried out by a millimeter-wave radar, but is not limited thereto.

[0100] The above-mentioned original eye movement information (or normalized eye movement information) may include information such as the eye movement trajectory, line-of-sight change, eye movement state, three-dimensional (abbreviated as 3D in English: Three Dimensions) spatial position of the eyes, fixation time, pupil size, blink count, eyelid closure degree, etc. of a driver (such as a driver to be detected). The collection of the original eye movement information can be carried out using a telemetric eye tracking system, but is not limited thereto.

[0101] Through the above method, based on the temporal encoder, the time dependence and dynamic changes in biological information and eye movement information can be captured, and the internal laws of biological information and eye movement information can be better understood. Thus, global feature extraction can be performed on biological information and eye movement information respectively based on the temporal encoder, making the extracted global biological features and global eye movement features more comprehensive and accurate. Moreover, through the normalization of the original biological information and the normalization of the original eye movement information, the scale differences between different features are eliminated, enabling the temporal encoder to treat each feature more fairly when extracting the global biological features and global eye movement features of biological information and eye movement information, thereby further improving the accuracy of the global biological features and global eye movement features.

[0102] Then, the global biological features and global eye movement features are concatenated (such as concat) to obtain the global concatenated features (i.e., Figure 1a R in). This concatenation can be to concatenate the global biological features and global eye movement features along a certain dimension.

[0103] Then, the second MLP (i.e., Figure 1a MLP2 in) performs feature extraction on the global concatenated features, and the extracted features are pooled and dimension-reduced to obtain the fused features, i.e., the global fused features (i.e., Figure 1a Y1 in). By using this second MLP to process the global concatenated features, the complex relationships between different features in the global concatenated features can be learned and fused, making the obtained global fused features more comprehensive and accurate.

[0104] As Figure 1a shown, in the local fusion module of the first network, first, through the first CA (i.e., Figure 1a CA1 in), the normalized biological information and eye movement information are interactively processed to obtain the critical moment features (i.e.,Figure 1a Z1) in it. Thus, through the first CA, the biological information and the eye movement information are deeply interacted and fused at the feature level, and then the interaction features (i.e., critical moment features) of the biological information and the eye movement information at critical moments are better captured, making the extracted critical moment features more accurate.

[0105] For example, the features in the biological information can be used as Query, and the features in the eye movement information can be used as Key and Value; or, the features in the eye movement information can be used as Query, and the features in the biological information can be used as Key and Value, which can be adjusted according to specific application scenarios and requirements. Figure 1a Taking the features in the biological information as Key (i.e., Figure 1a K1) and Value (i.e., Figure 1a V1) in it, and the features in the eye movement information as Query (i.e., Figure 1a Q1) in it as an example. Thus, through the processing of K1, V1, and Q1 by the first CA, the critical moment features are obtained.

[0106] The process of obtaining the critical moment features through the first CA can be shown by the following formula:

[0107]

[0108] where d k represents the dimension of K1.

[0109] The larger the value corresponding to the above Z1, the more important the features corresponding to the corresponding moment are.

[0110] Since the biological information and the eye movement information input to the first CA are the biological information and the eye movement information at all moments, the importance of the features at each moment can be determined according to the weight (i.e., Z1, critical moment features) of each moment, and then the local fusion features determined based on the critical moment features are more accurate.

[0111] Furthermore, the critical moment features are input into the third MLP (i.e., Figure 1a MLP3) in it, and through the processing of the third MLP on the critical moment features, local fusion features (i.e., Figure 1a Y2) in it are obtained.

[0112] As Figure 1a shown, in the global and local fusion module in the first network, first, the global fusion features and the local fusion features are concatenated to obtain concatenated fusion features (i.e., Figure 1a [Y1, Y2]) in it. Then, the concatenated fusion features are sent into the first MLP (i.e., Figure 1ain the MLP1), the first MLP extracts features from the spliced and fused features to obtain the target fused features after the final fusion of biometric information and eye movement information (i.e., Figure 1a F in).

[0113] The second network in the aforementioned target fatigue driving detection model (or fatigue driving detection model) may include H sub-networks, where H is a positive integer, and may also include linear units, a first feed-forward neural network (English: Feed-Forward Network, abbreviated as FNN), a first self-attention (English: Self-Attention, abbreviated as SA), a multi-head attention (English: Multi-Head Attention, abbreviated as MHA), and an average pooling layer (English: AvgPool, abbreviated as AP).

[0114] In the embodiments of the present application, the second network may be a deep neural network. The deep neural network may be a convolutional neural network (English: Convolutional Neural Networks, abbreviated as CNN), or a vision transformer (English: Vision Transformer, abbreviated as vit), but is not limited thereto, and can be adjusted according to specific application scenarios. Figure 1b The second network shown is the cross-frame features of the image extracted taking vit as an example.

[0115] Specifically, as Figure 1b shown, first through patch embedding, the multiple frames of visual images in T frames of visual images (i.e., Figure 1b X3 in, X3 ∈ R W×S×C×T ; W represents the width of the visual image, S represents the height of the visual image, and C represents the number of channels of the visual image) are respectively segmented into a image patches (patches), obtaining the patch features corresponding to each of the multiple frames of visual images (i.e., Figure 1b y in 1 ,..., y t ,..., y T , t ∈ [1, T]; where, y 1 , y t , y T respectively represent the patch features corresponding to the first frame of visual image, the t-th frame of visual image, and the T-th frame of visual image). And take this patch feature as the first patch feature.

[0116] In the embodiments of the present application, splitting each of the multiple visual images in the T-frame visual image into a patches can be to split each visual image in the T-frame visual image into a patches respectively, so as to obtain the patch features corresponding to each visual image, which is convenient for subsequent processing of each visual image in the T-frame visual image, making the obtained cross-frame features of the image more accurate.

[0117] The above visual image may include facial feature information of a driver (such as a driver to be detected). The visual image can be collected by a camera, but is not limited thereto.

[0118] In the embodiments of the present application, the number a of patches in one frame of visual image can be determined in the following way:

[0119] First, calculate the first product between the width and height of the visual image, and calculate the second product between the patch window sizes. Then, take the ratio of the first product to the second product as the number a of patches in one frame of visual image, making the number of patches in one frame of visual image more reasonable. This process can be shown by the following formula:

[0120]

[0121] where P represents the window size of the patch.

[0122] Taking the t-th frame of visual image as an example, the first patch feature (i.e., y t ) corresponding to the t-th frame of visual image is:

[0123]

[0124] where represents the feature corresponding to the j-th patch in the t-th frame of visual image, and j ∈ [1, a]. Therefore, the first patch feature corresponding to the t-th frame of visual image includes the features corresponding to each of the a patches in the t-th frame of visual image.

[0125] Then, based on the first patch feature, the initial class token (i.e., Figure 1b in where respectively represent the initial class tokens corresponding to the first frame of visual image, the t-th frame of visual image, and the T-th frame of visual image) and the position encoding (i.e., Figure 1b the E in P_1 ,..., E P_t ,..., EP_T ; among them, E P_1 , E P_t , E P_T respectively represent the position encodings corresponding to the first-frame visual image, the t-frame visual image, and the T-frame visual image, and in the position encoding corresponding to one frame of visual image, the position encodings of different patches are different), determine the first initial input of the initial layer sub-network (i.e., the first layer sub-network in the second network) (i.e., Figure 1b 's i 1,1 , …, i t,1 , …, i T,1 ; among them, i 1,1 , i t,1 , i T,1 respectively represent the first initial inputs of the first-frame visual image, the t-frame visual image, and the T-frame visual image in the initial layer sub-network). This initial class token is a learnable vector, which can be obtained by initializing random values or in a specific initialization manner. And, the class tokens after this initial class token can be updated and optimized according to the training process of the fatigue driving detection model, so that it can better represent the category information of the entire visual image. Therefore, in the second network of the target fatigue driving detection model, the initial class token and the class tokens after the initial class token can be obtained through the training of the fatigue driving detection model.

[0126] Specifically, first, through a linear unit (regard this linear unit as the first linear unit, i.e., Figure 1b 's L1), process the first patch features of multiple frames of visual images respectively to obtain the second patch features corresponding to each frame of visual image in the multiple frames of visual images, so as to convert the first patch features into vectors, which is convenient for the second network to better learn the cross-frame features of the images in the visual image. Then, splice the second patch features with the initial class tokens of the corresponding frames, and add the position encoding of the corresponding frames to the spliced features to obtain the first initial input of the initial layer sub-network. Thus, through the introduction of the position encoding, the information of the relative position of the patch in the visual image can be provided, so that the second network can understand and utilize the spatial relationship between different patches, and thus more accurately understand the overall structure of the visual image.

[0127] Through the above method, the first initial input includes both the features of the patch part information, the information of the class token part, and the position encoding, making the information contained in this first initial input more comprehensive, so that the second network can better learn the cross-frame features of the visual image based on the first initial input.

[0128] Taking the t-th frame of visual image as an example, according to the above process, the first initial input of the initial layer sub-network determined based on the first patch feature and the initial class token corresponding to the t-th frame of visual image can be shown as follows:

[0129]

[0130] Then, based on the first initial input of the corresponding frame, the patch output of the corresponding frame in the initial layer sub-network (i.e., o in Figure 1b ... o 1,1 ... o t,1 ... o T,1 ; where o 1,1 o t,1 o T,1 respectively represent the patch outputs of the first frame of visual image, the t-th frame of visual image, and the T-th frame of visual image in the initial layer sub-network) and the token output (i.e., in Figure 1b ... where respectively represent the token outputs of the first frame of visual image, the t-th frame of visual image, and the T-th frame of visual image in the initial layer sub-network). The patch output is the output of the corresponding patch part, and the token output is the output of the corresponding token part.

[0131] Specifically, first, normalize the first initial input of the corresponding frame (i.e., Nor in Figure 1b ) to obtain a normalized input to eliminate the scale difference between different features in the first initial input. Then, process the normalized input through the second CA (i.e., CA2 in Figure 1b ). Then, add the first initial input to the feature obtained by processing the normalized input by the second CA to obtain the initial output of the corresponding frame in the initial layer sub-network. The initial output includes the initial patch output (i.e., (o Figure 1b )'... (o 1,1 )'... (o t,1 )'; where (o T,1 )'... (o 1,1 )'... (o t,1 )'... (o T,1 )' respectively represent the initial patch outputs of the first frame of visual image, the t-th frame of visual image, and the T-th frame of visual image in the initial layer sub-network) and the token output. The initial patch output is the initial output of the corresponding patch part.

[0132] Taking the t-th frame of visual image as an example, the above process can be shown by the following formula:

[0133]

[0134] Then output the initial patch, and add the features obtained by processing the output of the initial patch with the first FNN (i.e., the FFN1 in Figure 1b ) to obtain the patch output of the corresponding frame in the initial layer sub-network. Taking the t-th frame of visual image as an example, this process can be shown by the following formula:

[0135] o t,1 =(o t,1 )'+FFN1((o t,1 )')

[0136] Next, concatenate the patch output of the initial layer sub-network with the token fusion features of the second layer sub-network, and use the concatenated features as the initial input of the second layer sub-network (i.e., the second initial input). Then, based on the initial input of the second layer sub-network, obtain the patch output and token output of the second layer sub-network. The specific process of obtaining the patch output and token output of the second layer sub-network based on the initial input of the second layer sub-network is the same as the specific process of obtaining the patch output and token output of the initial layer sub-network based on the first initial input of the initial layer sub-network.

[0137] Continue to repeat the above process based on the next layer sub-network until obtaining the patch output (i.e., the o Figure 1b in 1,H , …, o t,H , …, o T,H ; where, o 1,H , o t,H , o T,H respectively represent the patch outputs of the 1st frame of visual image, the t-th frame of visual image, and the T-th frame of visual image in the H-th layer sub-network) and token output (i.e., the Figure 1b in where respectively represent the token outputs of the 1st frame of visual image, the t-th frame of visual image, and the T-th frame of visual image in the H-th layer sub-network).

[0138] Among them, the patch output of the H-th layer sub-network is the first cross-frame feature in the image cross-frame features, and this first cross-frame feature is the cross-frame feature corresponding to the patch. This first cross-frame feature F I can be shown as follows:

[0139] F I =[o 1,H , …, o t,H , …, o T,H

[0140] The above token output of the H-th layer sub-network can adopt F​C is represented as follows:

[0141]

[0142] Taking the h-th layer sub-network as an example, where h is an integer greater than 1 and less than or equal to H. Concatenate the patch outputs of the (h - 1)-th layer sub-network (i.e., Figure 1b the o in 1,h-1 ,..., o t,h-1 ,..., o T,h-1 ; where, o 1,h-1 , o t,h-1 , o T,h-1 respectively represent the patch outputs of the first frame visual image, the t-th frame visual image, and the T-th frame visual image in the (h - 1)-th layer sub-network) with the token fusion features of the h-th layer sub-network (i.e., Figure 1b the in where Figure 1b the i in 1,h ,..., i t,h ,..., i T,h ; where, i 1,h , i t,h , i T,h respectively represent the initial inputs of the first frame visual image, the t-th frame visual image, and the T-th frame visual image in the h-th layer sub-network, that is, the second initial input). Thus, the initial input of the h-th layer sub-network includes both the patch output of the previous layer sub-network and the token fusion features, so that the h-th layer sub-network can better learn the features in the visual image to obtain a more accurate output.

[0143] Taking the t-th frame visual image as an example, the initial input i t,h of the t-th frame visual image in the h-th layer sub-network is:

[0144] The determination method of the token fusion features of the above-mentioned h-th layer sub-network is as follows:

[0145] First, through a linear unit (regarding this linear unit as the second linear unit, that is, Figure 1b the L2 in Figure 1b the where respectively represent the class tokens of the first-frame visual image, the t-frame visual image, and the T-frame visual image in the sub-network of the (h - 1)th layer, and process them to obtain the token information of the sub-network of the hth layer (i.e., Figure 1b in where respectively represent the token information of the first-frame visual image, the t-frame visual image, and the T-frame visual image in the sub-network of the hth layer). This token information can be the token information of all frames of the sub-network of the hth layer. Thus, through the second linear unit, the feature information in the class token of the sub-network of the (h - 1)th layer can be better extracted and transformed, so as to be better processed in the sub-network of the hth layer.

[0146] This process can be shown by the following formula:

[0147]

[0148] where E h represents the token information of all frames of the sub-network of the hth layer.

[0149] Then, normalize the token information of the sub-network of the hth layer (i.e., Nor in Figure 1b ) to obtain the token normalization information, so that the numerical range of the token information is within a reasonable range to avoid numerical instability problems in subsequent processing. Then, through the first SA (i.e., SA1 in Figure 1b ), process the token normalization information, and then add the obtained feature to the token information of the sub-network of the hth layer to obtain the token fusion feature of the sub-network of the hth layer.

[0150] Taking the token information of all frames of the sub-network of the hth layer as an example, when normalizing the token information of all frames of the sub-network of the hth layer, the feature obtained by processing the token normalization information by the first SA is added to the token information of all frames of the sub-network of the hth layer to obtain the token fusion feature E′ h of all frames of the sub-network of the hth layer. This process can be shown by the following formula:

[0151]

[0152] Since SA allows the model to consider the relationships between each element in a sequence when processing the sequence, it can help the model better understand the context information in the sequence, so as to process the sequence data more accurately. Therefore, by processing the token normalization information through the first SA, the global dependencies in the token normalization information can be captured, so as to better understand the context information in the token normalization information. Moreover, through the first SA, the similarity or correlation between each feature in the token normalization information and all other features can be calculated, so that the importance of each feature can be adjusted. Then, based on the processed features, the token information is summed up to generate richer and more accurate new token information (i.e., token fusion features), so as to facilitate the second network to better understand the visual image, and further make the subsequent extracted cross-frame features of the image more accurate.

[0153] After obtaining the initial input of the h-th sub-network, based on this initial input, the patch output of the h-th sub-network (i.e., Figure 1b the o in 1,h ,..., o t,h ,..., o T,h ; where, o 1,h , o t,h , o T,h respectively represent the patch outputs of the first-frame visual image, the t-th frame visual image, and the T-th frame visual image in the h-th sub-network) and the token output (i.e., Figure 1b the where, respectively represent the token outputs of the first-frame visual image, the t-th frame visual image, and the T-th frame visual image in the h-th sub-network). This process is the same as the specific process of obtaining the patch output and token output of the initial sub-network based on the first initial input of the initial layer sub-network. Taking the t-th frame visual image as an example, it can be shown as the following formula:

[0154]

[0155] where, (o t,h )' represents the initial output of the corresponding patch part of the t-th frame visual image in the h-th sub-network, that is, the initial patch output corresponding to the t-th frame visual image in the h-th sub-network.

[0156] Furthermore, after obtaining the patch output and token output of the H-th sub-network, add the time position encoding to the token output (i.e., Figure 1b the E in time_1 ,..., E tiem_t ,..., E time_T ; where, E time_1 , Etiem_t , E time_T respectively represent the time position encodings corresponding to the first-frame visual image, the t-frame visual image, and the T-frame visual image), and then are fed into the MHA, and the obtained result is fed into the AP to obtain the second cross-frame feature in the image cross-frame feature (i.e., Figure 1b F in v ). Thus, the second cross-frame feature is made more accurate. Moreover, the feature obtained by processing the token output based on the time position encoding, MHA, and AP is used as the second cross-frame feature, making the second cross-frame feature more accurate and enabling better fusion of the second cross-frame feature with the target fusion feature in the subsequent process. And, by introducing the time position encoding, when the MHA processes the token output, it can consider the time position in the token output, which helps to make the obtained second cross-frame feature more accurate.

[0157] The above-mentioned second cross-frame feature is the cross-frame feature corresponding to the class token. The above process can be shown by the following formula:

[0158] F v = AP(MHA(F C + E time ))

[0159] where E time represents the time position encoding.

[0160] The third network in the aforementioned target fatigue driving detection model (or fatigue driving detection model) may include a feature interaction module, a cross-modal decoder, and a fourth MLP. The feature interaction module may be composed of a second SA and bi-attention (English: Bi-Attention, abbreviated as BiA); the cross-modal decoder may be composed of a third CA, a fourth CA, and a second FFN.

[0161] Specifically, in the third network shown in Figure 1c , first in the time dimension, the first cross-frame feature in the image cross-frame feature (i.e., Figure 1c F in I ) is averaged to obtain a cross-frame average feature (i.e., Figure 1c in ) to reduce the influence of fluctuations and noise, make the data more stable, and help capture the data trend in the first cross-frame feature.

[0162] Then, in the feature interaction module, the cross-frame average feature and the target fusion feature (i.e., Figure 1c F in Figure 1c ), and the second interaction feature corresponding to the target fusion feature (i.e.,Figure 1c in ). Specifically, the cross-frame average feature and the target fusion feature are respectively input into the second SA, and the obtained results are then input into the BiA. Through the processing of the BiA, the first interaction feature and the second interaction feature are obtained. This process can be shown by the following formula:

[0163]

[0164] The above BiA can be the BiA in the prior art and will not be elaborated here. When the above second SA processes the target fusion feature and the cross-frame average feature respectively, the parameters used can be the same or different, and can be determined according to the training of the fatigue driving detection model.

[0165] Through the above feature interaction module, first, the second SA processes the cross-frame average feature and the target fusion feature respectively, which can capture the complex relationships within the cross-frame average feature and within the target fusion feature, thereby generating a richer and more accurate feature representation, which helps the subsequent BiA to better understand and learn the correlation and mutual influence between the cross-frame average feature and the target fusion feature. Then, the BiA is used for processing, which can capture the bidirectional dependence relationship between the cross-frame average feature and the target fusion feature, so that both the influence of the cross-frame average feature on the target fusion feature and the influence of the target fusion feature on the cross-frame average feature can be considered, and thus a more comprehensive feature representation (i.e., the first interaction feature, the second interaction feature) can be generated.

[0166] Next, in the cross-modal decoder, the first interaction feature, the second interaction feature, and the second cross-frame feature in the image cross-frame feature are processed to obtain intermediate features (i.e., Figure 1c Z in

[0167] Specifically, first, through the third CA, the second interaction feature and the second cross-frame feature (i.e., Figure 1c F in vProcess it to obtain the first information, so as to better interact and fuse the second interaction feature and the second cross-frame feature to generate a richer and more comprehensive feature representation (i.e., the first information). Then process the first information and the first interaction feature through the fourth CA to obtain the second information, so as to better interact and fuse the first information and the first interaction feature to generate a richer and more comprehensive feature representation (i.e., the second information). Then process the second information through the second FFN to obtain the intermediate feature, so as to better capture the features in the second information, facilitate learning more advanced feature representations from the second information, and make the generated intermediate feature richer. Thus, through the cross-modal decoder composed of the third CA, the fourth CA, and the second FFN, different modal information and information across time dimensions can be fused, which is beneficial to correctly judge the driver's state to determine whether the driver is fatigued or not fatigued.

[0168] This process can be shown as follows:

[0169]

[0170] In the above cross-modal decoder, the first interaction feature and the second interaction feature can be used as the input of Key and Value, and the second cross-frame feature can be used as the input of Query. Specifically, in the process of the third CA, the second interaction feature (i.e., ) is used as Key (i.e., K2 in Figure 1c ), Value (i.e., V2 in Figure 1c ), and the second cross-frame feature (i.e., F v ) is used as Query (i.e., Q2 in Figure 1c ); in the process of the fourth CA, the first interaction feature (i.e., F') is used as Key (i.e., K3 in Figure 1c ), Value (i.e., V3 in Figure 1c ), and the processing result of the third CA (i.e., the first information) is used as Query.

[0171] Then process the intermediate feature through the fourth MLP (i.e., MLP4 in Figure 1c ) to obtain the target feature for recognition (i.e., Z Figure 1c in C ), so as to learn the information in the intermediate feature and make the obtained target feature more accurate.

[0172] In the target fatigue driving detection model, when the target feature is obtained based on the foregoing third network, the classification network can be used to further process the target feature, so as to obtain the detection result for indicating whether the driver to be detected is fatigued or not fatigued.

[0173] In addition, the third network may also include a learning module. Through this third module, the fatigue driving detection model can calculate the target loss value by using the historical target features, the first historical interaction feature, and the second historical interaction feature obtained based on the training data (such as historical biological information, historical eye movement information, and historical visual images), and thus perform training based on the target loss value to obtain the target fatigue driving detection model.

[0174] Specifically, in the learning module, first, the classification loss value is calculated according to the historical target features and the classification loss function. The historical target features are obtained by first fusing the historical biological information and historical eye movement information in the training data based on the first network in the fatigue driving detection model to obtain the historical target fusion features, and extracting the historical image cross-frame features of the historical visual images in the training data based on the second network in the fatigue driving detection model, and then processing the historical target fusion features and the historical image cross-frame features based on the third network in the fatigue driving detection model.

[0175] Among them, the specific process of fusing the historical biological information and historical eye movement information in the training data based on the first network in the fatigue driving detection model to obtain the historical target fusion features is the same as the specific process of fusing the biological information and eye movement information to obtain the target fusion features in the aforementioned first network; the specific process of extracting the historical image cross-frame features of the historical visual images in the training data based on the second network in the fatigue driving detection model is the same as the specific process of extracting the image cross-frame features of the visual images in the aforementioned second network; the specific process of processing the historical target fusion features and the historical image cross-frame features based on the third network in the fatigue driving detection model to obtain the historical target features is the same as the specific process of processing the target fusion features and the image cross-frame features to obtain the target features in the aforementioned third network, and will not be elaborated here.

[0176] The process of calculating the classification loss value according to the historical target features and the classification loss function can be shown by the following formula:

[0177] L C1 = f(softmax(Z C ), g)

[0178] Among them, L C1 represents the classification loss value; f(·) represents the classification loss function; g represents the true category (such as the driver is fatigued or the driver is not fatigued), and this true category can be obtained by pre-labeling the categories corresponding to the driver's historical biological information, historical eye movement information, and historical visual images (such as, fatigued driving, non-fatigued driving); softmax(Z C) represents the predicted category, which can be the category obtained after the classification network (i.e., softmax) in the fatigue driving detection model processes the historical target features (e.g., fatigue driving, non-fatigue driving).

[0179] The above classification loss function can be loss functions such as Cross-Entropy Loss, Focal Loss, Log Loss, etc., but is not limited to this, and the specific classification loss function can be selected according to the specific application scenario.

[0180] In addition, according to the first historical interaction feature and the second historical interaction feature, a contrast loss value is calculated. The first historical interaction feature and the second historical interaction feature are obtained after the third network in the fatigue driving detection model processes the historical target fusion feature and the historical image cross-frame feature. Among them, the specific process of obtaining the first historical interaction feature and the second historical interaction feature after the third network in the fatigue driving detection model processes the historical target fusion feature and the historical image cross-frame feature is the same as the specific process of obtaining the first interaction feature and the second interaction feature by processing the target fusion feature and the image cross-frame feature in the aforementioned third network, and will not be elaborated here.

[0181] The process of calculating the contrast loss value according to the first historical interaction feature and the second historical interaction feature can be shown by the following formula:

[0182]

[0183] where, L C2 represents the contrast loss value; L C (·) represents the contrast loss function. The contrast loss function can be loss functions such as Cosine Similarity Loss, Contrastive Loss, etc., but is not limited to this, and the specific contrast loss function can be selected according to the specific application scenario.

[0184] The above contrast loss function can be used to limit the first historical interaction feature and the second historical interaction feature to be close in a similar semantic space and far in a dissimilar semantic space, so that the fatigue driving detection model can learn the representations of similar semantics in different modalities, enhance the learning of multi-modalities in the semantic space, and then train the fatigue driving detection model in the way of using the contrast loss function, making the obtained target fatigue driving detection model more accurate, and further making the detection result of whether the driving to be detected is fatigued more accurate based on the optimal target fatigue driving detection model.

[0185] Then, the sum of the classification loss value and the contrast loss value is used as the target loss value, as shown in the following formula:

[0186] L all = L C1 + L C2

[0187] Among them, L all represents the target loss value.

[0188] Furthermore, according to the minimization of the target loss value, the fatigue driving detection model is trained to obtain the target fatigue driving detection model, so that the performance of the obtained target fatigue driving detection model is optimal, so that when detecting whether the driver is fatigued based on the target fatigue driving detection model with the optimal model performance, the detection result is more accurate.

[0189] In addition, in the embodiments of the present application, the parameters of the above-mentioned first MLP, second MLP, third MLP, and fourth MLP are different from each other, and the parameters corresponding to each of the first MLP, second MLP, third MLP, and fourth MLP can all be obtained through the training of the fatigue driving detection model.

[0190] The parameters of the above-mentioned first CA, second CA, third CA, and fourth CA are different from each other, and the parameters corresponding to each of the first CA, second CA, third CA, and fourth CA can all be obtained through the training of the fatigue driving detection model.

[0191] The parameters of the first SA and the second SA are different from each other, and the parameters corresponding to each of the first SA and the second SA can all be obtained through the training of the fatigue driving detection model.

[0192] The parameters of the above-mentioned first FFN and second FFN are different from each other, and the parameters corresponding to each of the first FFN and second FFN can all be obtained through the training of the fatigue driving detection model.

[0193] Refer to Figure 2 The flowchart of a method for detecting fatigue driving provided by the embodiments of the present application is shown below. The method includes:

[0194] S201, obtaining the biological information, eye movement information, and visual image of the driver to be detected within T frames.

[0195] Among them, the biological information, eye movement information, and visual image are all non-contact information.

[0196] Since in the prior art, when detecting whether the driver is fatigued according to information such as the electroencephalogram signal of the driver, the electroencephalogram signal needs to be collected by the driver wearing a professional brain-computer interface device, and the accuracy of the brain-computer interface device will be affected when the driver's head and body move frequently, thereby affecting the accuracy of the detection result of whether the driver is fatigued.

[0197] Therefore, the embodiments of the present application perform fatigue driving detection based on non-contact information (i.e., the biological information, eye movement information, and visual images of the driver to be detected), thereby avoiding the impact on the accuracy of the acquisition device (such as a brain-computer interface device) during the process of collecting information when the head and body of the driver to be detected move frequently. This makes the collected data more accurate, helps to further improve the accuracy of the detection result of whether the driver to be detected is fatigued while driving, and at the same time avoids the driver to be detected wearing the acquisition device for a long time, thus improving the driving comfort.

[0198] S202. Process the biological information, eye movement information, and visual images through the target fatigue driving detection model to obtain a detection result for indicating whether the driver to be detected is fatigued while driving or not.

[0199] Specifically, first, through the first network in the target fatigue driving detection model, fuse the biological information and eye movement information to obtain a target fusion feature, and through the second network in the target fatigue driving detection model, extract the image cross-frame feature of the visual image. Then, through the third network in the target fatigue driving detection model, process the target fusion feature and the image cross-frame feature to obtain a target feature. Further, through the classification network in the target fatigue driving detection model, process the target feature to obtain a detection result for indicating whether the driver to be detected is fatigued while driving or not. Thus, through the target fatigue driving detection model, the biological information, eye movement information, and visual images of the driver to be detected within T frames are used to accurately detect whether the driver to be detected is fatigued while driving.

[0200] Furthermore, if the detection result indicates that the driver to be detected is fatigued while driving, the device can also use voice prompts to remind the driver to take a break and make adjustments to avoid the harm caused by the driver's fatigue while driving.

[0201] Optionally, the specific process of extracting the image cross-frame feature of the visual image through the second network (such as the second network shown in Figure 1b ) in the target fatigue driving detection model can be as shown in steps S301 - S308 in Figure 3 .

[0202] S301. Through patch embedding, each frame of the visual image is respectively segmented into a patches to obtain the first patch feature corresponding to each frame of the visual image.

[0203] S302. Based on the first patch feature and the initial class token, determine the first initial input of the initial layer sub-network in the second network.

[0204] S303. Determine that the current number of layers is the first preset value.

[0205] The first preset value can be 1. That is, it is determined that the current layer number is 1.

[0206] S304. Based on the first initial input, obtain the patch output and token output of all frames of the sub-network of the corresponding layer.

[0207] When the current layer number is 1, based on the first initial input of the corresponding frame in the initial layer sub-network, obtain the patch output and token output of the corresponding frame in the initial layer sub-network.

[0208] S305. Add the second preset value to the current layer number to obtain a new layer number.

[0209] The second preset value can be 1. That is to say, every time the patch output and token output of the sub-network of the corresponding layer are obtained, the layer number is incremented by 1.

[0210] S306. Determine whether the new layer number is less than or equal to H.

[0211] If it is determined that the new layer number is less than or equal to H, it is determined that the patch output and token output of the H-th layer sub-network have not been obtained yet, that is, the sub-network in the second network has not been executed up to the H-th layer sub-network. At this time, execute step S307.

[0212] If it is determined that the new layer number is greater than H, it is determined that the patch output and token output of the H-th layer sub-network have been obtained, that is, at this time, the H layer sub-networks in the second network have been executed, and continue to execute step S308.

[0213] S307. Concatenate the patch output and the token fusion feature of the next layer sub-network to obtain the second initial input of the next layer sub-network. Replace the first initial input with the second initial input, and replace the current layer number with the new layer number.

[0214] After replacing the first initial input with the second initial input, based on the replaced first initial input, continue to execute step S304, that is, based on the replaced first initial input, obtain the patch output and token output of the sub-network of the corresponding layer. And replace the current layer number with the new layer number, so that when step S305 is executed again after step S304, the current new layer number is added with the second preset value to obtain a new new layer number, so that the layer number can be updated every time step S304 is executed.

[0215] S308. Use the patch output of all frames of the obtained H-th layer sub-network as the first cross-frame feature, and based on AP and MSA, process the sum of the token output and the time position encoding of all frames of the H-th layer sub-network to obtain the second cross-frame feature.

[0216] After determining that the new number of layers is greater than H, the patch output of the obtained H-th layer sub-network is used as the first cross-frame feature, and after processing the sum of the token outputs and the temporal position encodings of all the obtained frames based on AP and MSA, the second cross-frame feature is obtained.

[0217] Through the above steps S301 - S308, the second network in the target fatigue driving detection model extracts the image cross-frame features (such as the first cross-frame feature and the second cross-frame feature) of the T-frame visual images.

[0218] In summary, the fatigue driving detection method proposed in this application uses the fusion of multi-modal information (such as the historical biological information, historical eye movement information, and historical visual images of the driver within T frames) to obtain the target fatigue driving detection model. Thus, when detecting whether the driver to be detected is fatigued through the target fatigue driving detection model, it is also based on the fusion of the multi-modal information of the driver to be detected (such as the biological information, eye movement information, and visual images of the driver to be detected within T frames). Compared with detecting whether the driver is fatigued using a single piece of information, the information used is more comprehensive, making the detection result of fatigue driving more accurate.

[0219] Moreover, since the fatigue state cannot be fully determined by one or two actions, through the fusion of multi-modal information in the target driving detection model in the above application embodiment, and at the same time extracting the image cross-frame features of the T-frame visual images of the driver, feature fusion is performed in the time dimension, correlating the features of the driver's different states at different times. Thus, it can more accurately analyze whether the driver is in a fatigued state (i.e., whether fatigued driving occurs), further making the detection result of whether the driver is fatigued driving more accurate.

[0220] In addition, the embodiments of this application use the historical biological information, historical eye movement information, and historical visual images of the driver for multi-modal fusion, and at the same time perform feature fusion in the time dimension, correlating the features of the driver's different states at different times. At the same time, contrastive learning is used to align the features of different modalities semantically, enabling the features to better learn the information in the semantic space. Thus, the obtained target fatigue driving detection model can more accurately analyze whether the driver is in a fatigued state, that is, whether the driver is fatigued driving.

[0221] Meanwhile, the embodiments of this application use the fusion of multiple features (such as the target fusion feature and the image cross-frame feature) to determine whether the driver is fatigued driving. Compared with using a single detection method to determine whether the driver is fatigued driving, the detection result is more accurate.

[0222] Based on the same inventive concept, an embodiment of this application also provides a fatigue driving detection device, such asFigure 4 The following is a schematic structural diagram of a fatigue driving detection device provided by the present application. The device includes:

[0223] An acquisition module 401, configured to acquire biometric information, eye movement information, and visual images of a driver to be detected within T frames;

[0224] A processing module 402, configured to process the biometric information, the eye movement information, and the visual images through a target fatigue driving detection model to obtain a detection result for indicating whether the driver to be detected is fatigued or not. Among them, the target fatigue driving detection model is obtained by comparing and learning the historical target fusion features obtained by fusing historical biometric information and historical eye movement information with the historical image cross-frame features of the extracted historical visual images.

[0225] In a possible implementation manner, the target fatigue driving detection model includes a first network, a second network, a third network, and a classification network. The processing module 402 is specifically configured to fuse the biometric information and the eye movement information through the first network to obtain a target fusion feature; and extract the image cross-frame features of the visual images through the second network; process the target fusion feature and the image cross-frame features through the third network to obtain a target feature; and process the target feature through the classification network to obtain the detection result for indicating whether the driver to be detected is fatigued or not.

[0226] In a possible implementation manner, the first network includes a first multi-layer perceptron MLP. The processing module 402 is specifically configured to obtain the global fusion feature between the biometric information and the eye movement information, and the local fusion feature between the biometric information and the eye movement information; splice the global fusion feature and the local fusion feature to obtain a spliced fusion feature; and perform feature extraction on the spliced fusion feature through the first MLP to obtain the target fusion feature.

[0227] In a possible implementation manner, the first network further includes a time series encoder and a second MLP. The processing module 402 is further configured to respectively perform global feature extraction on the biometric information and the eye movement information through the time series encoder to obtain a global biometric feature and a global eye movement feature; splice the global biometric feature and the global eye movement feature to obtain a global spliced feature; perform feature extraction on the global spliced feature through the second MLP, and perform pooling dimensionality reduction on the extracted features to obtain the global fusion feature.

[0228] In a possible implementation, the first network further includes a first cross-attention CA and a third MLP; the processing module 402 is further configured to perform interactive processing on the biological information and the eye movement information through the first CA to obtain critical moment features; and perform feature extraction on the critical moment features through the third MLP to obtain the local fusion features.

[0229] In a possible implementation, the second network includes an H-layer sub-network, an average pooling layer AP, and a multi-head attention MHA, where H is a positive integer. The image cross-frame features include a first cross-frame feature corresponding to an image patch and a second cross-frame feature corresponding to a class token. The processing module 402 is further configured to respectively segment multiple frames of the visual images into a patches through patch embedding to obtain patch features corresponding to each of the multiple frames of visual images; determine a first initial input of the initial layer sub-network in the second network based on the patch features, the initial class token, and the position encoding, and determine that the current layer number is a first preset value; obtain a patch output and a token output of the corresponding layer sub-network based on the first initial input, add the second preset value to the layer number to obtain a new layer number, and determine whether the new layer number is less than or equal to H; if so, splice the patch output and the token fusion feature of the next layer sub-network to obtain a second initial input of the next layer sub-network; replace the first initial input with the second initial input, and based on the replaced first initial input, obtain a patch output and a token output of the corresponding layer sub-network, replace the layer number with the new layer number, and based on the replaced layer number, determine the new layer number and determine whether the new layer number is less than or equal to H; if not, use the patch output of the obtained H-layer sub-network as the first cross-frame feature, and based on the AP and the MHA, process the sum of the token output of the H-layer sub-network and the time position encoding to obtain the second cross-frame feature.

[0230] In a possible implementation, the second network further includes a second CA and a first feed-forward neural network FNN; the processing module 402 is further configured to normalize the first initial input of the corresponding frame to obtain a normalized input, and add the feature obtained by the second CA processing the normalized input to the first initial input to obtain an initial output of the corresponding frame in the initial layer sub-network; where the initial output includes an initial patch output and the token output; add the feature obtained by the first FFN processing the initial patch output to the initial patch output to obtain the patch output of the corresponding frame in the initial layer sub-network.

[0231] In a possible implementation, the second network further includes a linear unit and a first self-attention (SA); the processing module 402 is further configured to process the class tokens of multiple frames of the visual images in the previous sub-network of the next sub-network through the linear unit respectively to obtain the token information of the next sub-network; normalize the token information to obtain token normalization information; and process the token normalization information through the first SA to obtain the token fusion feature of the next sub-network.

[0232] In a possible implementation, the third network includes a feature interaction module, a cross-modal decoder, and a fourth MLP. The feature interaction module is a module composed of a second SA and a bidirectional attention (BiA), and the cross-modal decoder is a module composed of a third CA, a fourth CA, and a second FFN; the processing module 403 is further configured to perform an average process on the first cross-frame feature in the image cross-frame features in the time dimension to obtain a cross-frame average feature; perform an interaction process on the cross-frame average feature and the target fusion feature through the feature interaction module to obtain a first interaction feature and a second interaction feature; process the first interaction feature, the second interaction feature, and the second cross-frame feature in the image cross-frame features through the cross-modal decoder to obtain an intermediate feature; and process the intermediate feature through the fourth MLP to obtain the target feature.

[0233] In a possible implementation, the device further includes a training module, which is specifically configured to obtain the historical biological information, the historical eye movement information, and the historical visual images; fuse the historical biological information and the eye movement information through the first network in the fatigue driving detection model to obtain the historical target fusion feature; and extract the historical image cross-frame features of the historical visual images through the second network in the fatigue driving detection model; process the historical target fusion feature and the historical image cross-frame features through the third network in the fatigue driving detection model to obtain a historical target feature, a first historical interaction feature, and a second historical interaction feature; calculate a classification loss value according to the historical target feature and a classification loss function; and calculate a contrast loss value according to the first historical interaction feature, the second historical interaction feature, and a contrast loss function; use the sum of the classification loss value and the contrast loss value as the target loss value, and train the fatigue driving detection model according to the minimization of the target loss value to obtain the target fatigue driving detection model.

[0234] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present application. The above electronic device can implement the functions of the foregoing fatigue driving detection device. Refer toFigure 5 , the above electronic device includes:

[0235] at least one processor 501 and a memory 502 connected to the at least one processor 501. In the embodiments of the present application, the specific connection medium between the processor 501 and the memory 502 is not limited. Figure 5 In the example, the processor 501 and the memory 502 are connected through a bus 500. The bus 500 is Figure 5 represented by a thick line in the figure. The connection manners between other components are only for illustrative purposes and are not limiting. The bus 500 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 501 can also be called a controller, and the name is not limited.

[0236] In the embodiments of the present application, the memory 502 stores instructions executable by the at least one processor 501. By executing the instructions stored in the memory 502, the at least one processor 501 can execute the fatigue driving detection method described above. The processor 501 can implement Figure 4 the functions of each module in the device shown in the figure.

[0237] Among them, the processor 501 is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 502 and calling the data stored in the memory 502, various functions of the device and process data, so as to monitor the device as a whole.

[0238] In a possible design, the processor 501 may include one or more processing units. The processor 501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 501. In some embodiments, the processor 501 and the memory 502 can be implemented on the same chip. In some embodiments, they can also be implemented on separate chips independently.

[0239] The processor 501 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the output method of the landing area disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0240] The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 502 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory 502 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0241] By designing and programming the processor 501, the code corresponding to the fatigue driving detection method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 2 the steps of the fatigue driving detection method of the embodiments shown. How to design and program the processor 501 is a well-known technology to those skilled in the art and will not be elaborated here.

[0242] Based on the same inventive concept, the embodiments of the present application also provide a storage medium that stores computer instructions. When the computer instructions run on a computer, the computer is made to execute the fatigue driving detection method discussed above.

[0243] In some possible embodiments, each aspect of the fatigue driving detection method provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the control device to execute the steps in the fatigue driving detection method according to various exemplary embodiments of the present application described above in this specification.

[0244] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0245] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0246] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0247] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0248] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.

Claims

1. A method for detecting fatigue driving, characterized in that: include: Obtaining biological information, eye movement information and visual images of the driver to be detected in the T frame; The biological information, the eye movement information and the visual image are processed by a target fatigue driving detection model to obtain a detection result indicating whether the driver to be detected is fatigue driving or not; wherein the target fatigue driving detection model is a fatigue driving detection model that uses historical biological information and historical eye movement information to obtain historical target fusion features, and compares and learns the historical image cross-frame features extracted from historical visual images.

2. The method according to claim 1, characterized in that The target fatigue driving detection model includes a first network, a second network, a third network and a classification network; The target fatigue driving detection model is used to process the biological information, the eye movement information, and the visual image to obtain a detection result indicating fatigue driving or non-fatigue driving of the driver to be detected, including: By means of the first network, the biological information and the eye movement information are fused to obtain target fusion features; and by means of the second network, image cross-frame features of the visual image are extracted; Processing the target fusion feature and the image cross-frame feature through the third network to obtain the target feature; The target features are processed through the classification network to obtain the detection result indicating whether the driver to be detected is driving fatigued or not.

3. The method according to claim 2, characterized in that The first network includes a first multi-layer perceptron MLP; The step of fusing the biological information and the eye movement information through the first network to obtain a target fusion feature includes: Acquire a global fusion feature between the biological information and the eye movement information, and a local fusion feature between the biological information and the eye movement information; Splicing the global fusion feature and the local fusion feature to obtain a spliced ​​fusion feature; The first MLP is used to extract the spliced ​​fusion features to obtain the target fusion features.

4. The method according to claim 3, characterized in that The first network also includes a temporal encoder and a second MLP; The acquiring of the global fusion feature between the biological information and the eye movement information comprises: By using the temporal encoder, respectively extracting global features from the biological information and the eye movement information to obtain global biological features and global eye movement features; Splicing the global biological feature and the global eye movement feature to obtain a global splicing feature; The global concatenated features are extracted through the second MLP, and the extracted features are pooled and dimensionally reduced to obtain the global fusion features.

5. The method according to claim 3, characterized in that The first network also includes a first cross attention CA and a third MLP; The acquiring of local fusion features between the biological information and the eye movement information includes: By using the first CA, the biological information and the eye movement information are interactively processed to obtain key moment features; The key moment features are extracted through the third MLP to obtain the local fusion features.

6. The method according to claim 2, characterized in that The second network includes an H-layer subnetwork, an average pooling layer AP, and a multi-head attention MHA, H is a positive integer, and the image cross-frame feature includes a first cross-frame feature corresponding to the image block patch and a second cross-frame feature corresponding to the class token class token; The step of extracting the image cross-frame features of the visual image through the second network includes: By block embedding, the multiple frames of visual images are respectively divided into a patches, and the patch features corresponding to the multiple frames of visual images are obtained; Based on the patch feature, the initial class token, and the position code, determine a first initial input of an initial layer subnetwork in the second network, and determine that the current number of layers is a first preset value; Based on the first initial input, obtain the patch output and token output of the corresponding layer subnetwork, add the number of layers to a second preset value to obtain a new number of layers, and determine whether the new number of layers is less than or equal to H; If so, concatenate the patch output and the token fusion feature of the next layer of sub-network to obtain the second initial input of the next layer of sub-network; replace the first initial input with the second initial input, and based on the replaced first initial input, obtain the patch output and token output of the corresponding layer of sub-network, replace the number of layers with the new number of layers, and determine the new number of layers based on the replaced number of layers, and judge whether the new number of layers is less than or equal to H; If not, the obtained patch output of the H-th layer sub-network is used as the first cross-frame feature, and based on the AP and the MHA, the sum of the token output and the time position coding of the H-th layer sub-network is processed to obtain the second cross-frame feature.

7. The method according to claim 6, characterized in that The second network also includes a second CA and a first feedforward neural network FNN; The obtaining of the patch output and token output of the corresponding layer sub-network based on the first initial input includes: Normalizing the first initial input of the corresponding frame to obtain a normalized input, and adding the first initial input to the features of the normalized input processed by the second CA to obtain the initial output of the corresponding frame in the initial layer subnetwork; wherein the initial output includes an initial patch output and the token output; The initial patch output is added with the feature after the first FFN processes the initial patch output, so as to obtain the patch output of the corresponding frame in the initial layer sub-network.

8. The method according to claim 6, characterized in that The second network also includes a linear unit and a first self-attention SA; Before the patch output is concatenated with the token fusion feature of the next layer of sub-network, the method further includes: Through the linear unit, the class tokens of the previous sub-network of the next sub-network are processed for the multiple frames of the visual image to obtain the token information of the next sub-network; Normalizing the token information to obtain token normalized information; The token normalization information is processed through the first SA to obtain the token fusion features of the next layer of sub-network.

9. The method according to claim 2, characterized in that The third network includes a feature interaction module, a cross-modal decoder, and a fourth MLP, wherein the feature interaction module is a module composed of a second SA and a bidirectional attention BiA, and the cross-modal decoder is a module composed of a third CA, a fourth CA, and a second FFN; The process of processing the target fusion feature and the image cross-frame feature through the third network to obtain the target feature includes: In the time dimension, averaging the first cross-frame feature of the cross-frame features of the image to obtain a cross-frame average feature; By means of the feature interaction module, the cross-frame average feature and the target fusion feature are interactively processed to obtain a first interactive feature and a second interactive feature; Processing the first interaction feature, the second interaction feature, and the second cross-frame feature in the image cross-frame feature through the cross-modal decoder to obtain an intermediate feature; The intermediate features are processed by the fourth MLP to obtain the target features.

10. The method according to any one of claims 1 to 9, characterized in that: Before the biological information, the eye movement information, and the visual image are processed by the target fatigue driving detection model to obtain a detection result indicating fatigue driving or non-fatigue driving of the driver to be detected, the method further includes: Acquiring the historical biological information, the historical eye movement information, and the historical visual images; By means of a first network in the fatigue driving detection model, the historical biological information and the eye movement information are fused to obtain the historical target fusion feature; and by means of a second network in the fatigue driving detection model, the historical image cross-frame feature of the historical visual image is extracted; The historical target fusion feature and the historical image cross-frame feature are processed by the third network in the fatigue driving detection model to obtain the historical target feature, the first historical interaction feature and the second historical interaction feature; Calculating a classification loss value according to the historical target feature and the classification loss function; and calculating a contrast loss value according to the first historical interaction feature, the second historical interaction feature and the contrast loss function; The sum of the classification loss value and the contrast loss value is taken as the target loss value, and the fatigue driving detection model is trained according to minimization of the target loss value to obtain the target fatigue driving detection model.

11. A fatigue driving detection device, characterized in that: include: An acquisition module, used to acquire biological information, eye movement information and visual images of the driver to be detected in the T frame; A processing module is used to process the biological information, the eye movement information, and the visual image through a target fatigue driving detection model to obtain a detection result indicating whether the driver to be detected is fatigue driving or not; wherein the target fatigue driving detection model is obtained by comparing and learning the historical target fusion features obtained by fusing historical biological information with historical eye movement information and the historical image cross-frame features extracted from the historical visual image.

12. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to implement the method steps of any one of claims 1 to 10 when executing the computer program stored in the memory.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • Driver fatigue driving state detection method and system

    CN121614965A

  • Method and system for detecting a driver's fatigue driving state

    CN121614965B