Method and apparatus for estimating fatigue state of object

Through the method of single-channel EEG signal acquisition and self-attention capsule network processing, the problems of complexity and low accuracy of driver alertness detection equipment in the prior art are solved, and more easy to wear and high-precision fatigue detection is achieved.

CN120203587APending Publication Date: 2025-06-27ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311833895.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing driver alertness detection methods require multiple acquisition points and high computing capabilities, making the equipment complex and difficult to wear, and it is difficult to achieve high-precision fatigue detection in practical applications.

Method used

A single-channel EEG signal acquisition device is used to collect EEG signals from the driver's forehead, and a method and device for estimating the fatigue state of an object is formed through a combination of feature extraction, self-attention processing and capsule network modules.

Benefits of technology

The design of signal acquisition equipment is simplified, the size and complexity of the equipment is reduced, the demand for computing resources is reduced, and on the basis of ensuring detection accuracy, it achieves higher real-time and application convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120203587A_ABST
    Figure CN120203587A_ABST
Patent Text Reader

Abstract

The present application relates to a method for estimating a fatigue state of a subject, comprising: obtaining a physiological signal of the subject; extracting features from the physiological signal and forming a first feature map using the features; performing self-attention processing on the first feature map by a self-attention module in the fatigue detection module to obtain a second feature map; and a capsule network module in the fatigue detection module processes the second feature map to obtain a result representing a fatigue state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more particularly, to a method and apparatus for estimating the fatigue state of an object. Background Art

[0002] With the rapid increase in vehicles, driving safety has become a crucial issue. It is reported that fatigue driving is one of the most prominent causes of traffic accidents. Generally, due to the lack of vigilance caused by fatigue, the driver cannot drive safely. Therefore, it is very important to evaluate the driver's vigilance in real time. Electroencephalogram (EEG), as a signal that directly reflects brain activity, has been proven to be a reliable indicator of the human mental state.

[0003] Currently, there have been some studies on the detection of driver vigilance. Among them, the collected EEG signals come from the international 10–20 system (the number of channels: between 17 and 61), and accordingly, the accuracy of vigilance detection is improved by adopting a method of fusing multiple deep neural networks. This method of collecting EEG signals requires a large number of collection points to be arranged on the scalp of the object, so it is required to wear the collection device on the head of the object and cover the entire cerebral cortex, which brings difficulties to the daily wearing of the device. In addition, the data processing amount of EEG data based on multi-channel multi-feature fusion is large, so the processing of EEG data in the vigilance detection model requires greater computing power and power supply. The above reasons make the existing vigilance detection methods difficult to apply in actual products. Correspondingly, it would be beneficial if the data collection device could be simplified to make it more easily applicable to actual products, and at the same time, the accuracy of vigilance detection could be not reduced or even improved. Summary of the Invention

[0004] The following introduction is provided to introduce some selected concepts in a simple form, which will be further described in the detailed description below. This introduction is not intended to highlight the key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

[0005] According to one aspect of the present application, there is provided a method for estimating the fatigue state of an object, including: obtaining a physiological signal of the object; extracting features from the physiological signal and forming a first feature map using the features; performing self-attention processing on the first feature map by a self-attention module in a fatigue detection module to obtain a second feature map; and processing the second feature map by a capsule network module in the fatigue detection module to obtain a result representing the fatigue state.

[0006] According to one embodiment, the physiological signal includes an EEG signal, and obtaining the physiological signal of the subject includes: obtaining the electroencephalogram signal of the subject collected by a single-channel acquisition device. According to one embodiment, the single-channel acquisition device collects the electroencephalogram signal through the forehead of the subject. According to one embodiment, the feature includes the power spectral density (PSD) and differential quotient (DE) extracted from the physiological signal, and the first feature map includes a single-channel feature map composed of multiple PSDs and multiple DEs.

[0007] According to one aspect of the present application, a method for training a fatigue detection module for estimating the fatigue state of an object is provided, including: obtaining a first training data set, the first training data set including data representing physiological signals and labels representing fatigue states; training the fatigue detection module based on the first training data set, including: performing self-attention processing on the data representing physiological signals in the first training data set by a self-attention module in the fatigue detection module to obtain a second feature map; processing the second feature map by a capsule network module in the fatigue detection module to predict a result representing the fatigue state; and updating the fatigue detection module based on the predicted result representing the fatigue state and the corresponding label representing the fatigue state.

[0008] According to one aspect of the present application, a device for estimating the fatigue state of an object is provided, including: a feature extraction module for obtaining the physiological signal of the object and extracting features from the physiological signal; a feature map construction module for forming a first feature map using the features; and a fatigue detection module including a self-attention module and a capsule network module, the self-attention module performing self-attention processing on the first feature map to obtain a second feature map, and the capsule network module processing the second feature map to obtain a result representing the fatigue state.

[0009] According to one aspect of the present application, a device for estimating the fatigue state of an object is provided, including: one or more processors; one or more memories storing computer-executable instructions, the instructions when executed causing the one or more processors to execute the methods described in various embodiments of the present disclosure.

[0010] According to one aspect of the present application, a machine-readable storage medium is provided, which stores executable instructions, the instructions when executed causing one or more processors to execute the methods described in various embodiments of the present disclosure.

[0011] According to one aspect of the present application, a computer program product is provided, which includes computer-executable instructions, the instructions when executed causing one or more processors to execute the methods described in various embodiments of the present disclosure.

[0012] Using the method for fatigue detection according to an embodiment of the present application, a single-channel electroencephalogram (EEG) signal can be collected from the forehead of an object using a signal acquisition device for fatigue detection of the object, thereby simplifying the acquisition of EEG signals. Correspondingly, the size and complexity of the signal acquisition device can be reduced, making it possible to use the signal acquisition device in practical applications such as driver fatigue detection. Since the amount of data of the single-channel EEG signal is less, the corresponding computational amount is also reduced, making the computing device more power-efficient. At the same time, in order to ensure the accuracy of detection when the amount of EEG signal data is reduced, an improved method for fatigue detection according to the present application is adopted, ensuring the accuracy of fatigue detection while reducing the amount of EEG signal data. Other advantages of the technical solution of the present application will be described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] By referring to the following drawings, a further understanding of the essence and advantages of the present disclosure can be achieved. In the drawings, similar components or features may have the same reference numerals.

[0014] Figure 1 FIG. shows a schematic diagram of collecting physiological signals of an object according to an embodiment.

[0015] Figure 2A FIG. shows a schematic diagram of a device for fatigue detection according to an embodiment.

[0016] Figure 2B shows Figure 2A a schematic diagram of a partial processing procedure shown.

[0017] Figure 2C FIG. shows a schematic diagram of a feature map of an input signal for fatigue detection according to an embodiment.

[0018] Figure 3A FIG. shows a schematic diagram of a device for fatigue detection according to an embodiment.

[0019] Figure 3B FIG. shows a schematic diagram of a device for fatigue detection according to an embodiment.

[0020] Figure 4 FIG. shows a schematic diagram of a process of training a fatigue detection module for fatigue detection according to an embodiment.

[0021] Figure 5 FIG. shows a flowchart of a method for estimating the fatigue state of an object according to an embodiment.

[0022] Figure 6 FIG. shows a flowchart of a method for training a fatigue detection module for estimating the fatigue state of an object according to an embodiment.

[0023] Figure 7 A block diagram of an apparatus for estimating a fatigue state of an object according to one embodiment is shown.

[0024] Figure 8 A block diagram of an apparatus for estimating a fatigue state of an object according to one embodiment is shown. DETAILED DESCRIPTION

[0025] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that the discussion of these embodiments is only to enable those skilled in the art to better understand and thereby implement the subject matter described herein, and is not a limitation on the scope of protection, applicability, or examples set forth in the claims. Changes may be made to the functions and arrangements of the elements discussed without departing from the scope of the present disclosure. Each example may omit, substitute, or add various processes or components as needed. For example, the methods described may be performed in a different order than described, and each step may be added, omitted, or combined. Additionally, features described relative to some examples may be combined in other examples.

[0026] As used herein, the term "comprising" and its variants denote open terms meaning "including but not limited to". The term "based on" means "at least partially based on". The terms "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below. Unless explicitly specified in the context, the definition of a term is consistent throughout the specification.

[0027] Figure 1 A schematic diagram of collecting physiological signals of an object according to one embodiment is shown.

[0028] In the illustrated embodiment, the collection device may collect physiological signals from the forehead of the object. For example, the sensor of the collection device contacts the forehead at the position 110 shown for collecting the EEG signal of the object. In one embodiment, the object may be a driver driving a vehicle. It can be understood that the fewer the number of signal collection points, the more conducive to the implementation of the collection device, so that it can be used to collect the EEG signal of the driver in real time in a driving environment. In Figure 1In the example shown, the acquisition device acquires EEG signals through a single signal acquisition point 110. Therefore, the acquired EEG signals are referred to as single-channel signals. It can be understood that although the EEG signals are used as examples in this article to illustrate the implementation manners of the present application, the technical solutions of the present application can also adopt other suitable physiological signals for representing the fatigue state of an object.

[0029] Although Figure 1 the acquisition device is not shown in

[0030] Figure 2A FIG. shows a schematic diagram of a device for fatigue detection according to an embodiment. Figure 2B FIG. shows Figure 2A a schematic diagram of a partial processing procedure shown. Figure 2C FIG. shows a schematic diagram of a feature map of an input signal for fatigue detection according to an embodiment.

[0031] As Figure 2A shown, the fatigue detection device 200 includes a feature extraction module 210. The feature extraction module 210 obtains EEG data I and extracts features F from the EEG data I. In one embodiment, the extracted features include PSD and DE. Any suitable method can be used to extract PSD and DE from the EEG data I. For example, in one implementation, the EEG data I can be downsampled to 200 Hz and then filtered using a band-pass filter from 1 to 75 Hz to reduce artifacts and noise. This downsampling and filtering process can be referred to as preprocessing. Then, the short-time Fourier transform (STFT) can be used to calculate the time-frequency features of the EEG signal with a 1-second Hanning window having a 50% overlap rate. Then, based on the output of the STFT, PSD and DE are calculated at a resolution of 2 Hz starting from 1 Hz. Finally, 50 EEG features are extracted from each 1-second window, including 25 PSDs and 25 DEs. Figure 2C FIG. shows the 50 EEG features extracted in this example. Those skilled in the art can understand that the above method for extracting PSD and DE from EEG signals is a method existing in the art. The technical solutions of the present application are not limited thereto, but any suitable feature extraction method can be used to extract the features of EEG signals.

[0032] The fatigue detection device 200 further includes a feature map construction module 220, which forms a feature map FM from the extracted features PSD and DE. For the sake of distinction in terms, in this article, this feature map FM is referred to as the first feature map FM. AsFigure 2C As shown, two types of features, PSD and DE, are constructed into a single-channel feature map FM. The first feature map FM includes 25 PSDs and 25 DEs, and the two together form a 5×10 first feature map. It should be noted that in the existing feature map construction methods for fatigue detection, different types of features are constructed in different feature channels. Taking the features PSD and DE as an example, the existing feature map construction method constructs PSD and DE in two feature channels respectively. For example, 25 PSDs are constructed into a 5×5 feature channel, and 25 DEs are constructed into a 5×5 feature channel. By adopting Figure 2C the shown feature map construction method, different types of features are included in the same feature channel, which can strengthen the connection between different features, thus helping the capsule network module 240 make better use of the vector characteristics of the capsules, and improving the accuracy of the capsule network module 240 in predicting the fatigue state of the object. In Figure 1 the shown single-channel EEG acquisition scheme, in order to reduce the size and structural complexity of the signal acquisition device, the collected single-channel EEG signal has less signal volume compared with the multi-channel EEG signals collected in the existing technology such as the international 10–20 system. Therefore, in the same processing method, the reduced signal volume may lead to a decrease in detection accuracy. By Figure 2C the shown feature map construction method, the accuracy of the capsule network module in predicting the fatigue state can be improved, thus effectively compensating for the problem of reduced detection accuracy caused by the reduction of the collected signal volume.

[0033] In Figure 2C the first feature map exemplified contains 50 features, and the feature map composed of these 50 features can be called a processing unit. The technical solution of the present application is not limited to the specific number of features included in the feature map, Figure 2C the shown feature map can also include other numbers of features, and these features can form a feature map with appropriate number of rows and columns. For example, in Figure 2C the shown example, when the frequency resolution is 1 Hz when calculating PSD and DE, 50 PSDs and 50 DEs can be obtained, thus constructing a 10×10 feature map containing PSD and DE.

[0034] The fatigue detection device 200 further includes a self-attention module 230 and a capsule network module 240. The self-attention module 230 and the capsule network module 240 process the first feature map FM to obtain a result O representing the fatigue state. The self-attention module 230 and the capsule network module 240 can be collectively referred to as a fatigue detection module, and the self-attention module 230 in it performs self-attention processing on the first feature map FM to obtain a second feature map, as Figure 2BAs shown in 201, the output of the self-attention module 230 does not change the dimension of the feature map FM, which remains a 5×10 feature map. The capsule network module 240 processes the second feature map output by the self-attention module 230 to obtain a result representing the fatigue state. For example, the result representing fatigue can be a value between 0 and 1, and its magnitude represents the degree of fatigue.

[0035] In Figure 2A the embodiment shown, the self-attention module 230 includes a position encoding module 2310, a self-attention sub-module 2320, and a residual module 2330. The position encoding module 2310 performs position encoding on the first feature map FM to obtain position information. For example, sinusoidal and cosine encoding can be used to implement position encoding. Alternatively, learnable encoding can be used to implement position encoding. It can be understood that any suitable position encoding method can be used to implement the position encoding of the position encoding module 2310. It can be understood that although not shown in Figure 2A , an embedding module can be included before the position encoding module 2310, which is used to transform the input first feature map FM into the embedding space, where each feature in the transformed feature map FM is represented in the form of an embedding vector. Correspondingly, the position encoding module processes the feature map FM transformed into the embedding space.

[0036] As Figure 2A shown, the position information obtained by the position encoding module 2310 is merged with the feature map FM to obtain a feature map containing position information. The self-attention sub-module 2320 performs self-attention processing on the feature map containing position information to obtain a second feature map. It can be understood that any suitable self-attention operation method can be used to implement the self-attention processing of the self-attention sub-module 2320. The self-attention processing of the self-attention sub-module 2320 can shorten the distance between features with dependencies. In Figure 2C the embodiment of the single-channel feature map shown, by using the self-attention processing of the self-attention sub-module 2320 to shorten the distance between features, it further helps the capsule network module 240 to make full use of the vector characteristics of the capsules, thereby improving the accuracy of the capsule network module 240 in predicting the fatigue state of the object. In this way, in Figure 1 the single-channel EEG acquisition scheme shown, in the case of reduced signal volume, the accuracy of the capsule network module in predicting the fatigue state can be improved, and at the same time, due to the reduced amount of data processed, the demand for processing resources and power can be reduced.

[0037] The self-attention processing of the self-attention sub-module 2320 may weaken the most important part of the feature values while shortening the distance between features. The residual module 2330 can be used to alleviate this problem. The residual module 2330 combines the second feature map output by the self-attention sub-module 2320 and the feature map containing position information before being processed by the self-attention sub-module 2320 to obtain the second feature map AFM finally output by the self-attention module 230. As Figure 2B shown in 201, in Figure 2C the example shown, the second feature map AFM output by the self-attention module 230 has the same dimension as the input feature map, that is, a feature matrix of 5×10. In one embodiment, the residual module 2330 can perform combination and normalization processing, which combines the second feature map output by the self-attention sub-module 2320 and the feature map containing position information before being processed by the self-attention sub-module 2320, and normalizes the combined feature map to obtain the second feature map AFM finally output by the self-attention module 230. The normalization processing can also be called the normalization process. Through this processing, the stability of the neural network processing can be improved, and thus the stability of the training stage of the neural network can be improved.

[0038] In Figure 2A the embodiment shown, the capsule network module 240 includes a first convolutional module 2410, a layer normalization module 2420, a second convolutional module 2430, a capsule processing module 2440, and an output module 2450. The first convolutional module 2410 receives the second feature map AFM output by the self-attention module 230, such as Figure 2B the 5×10 feature map shown in 201. The first convolutional module 2410 may include a plurality of convolutional kernels. For example, in the example shown in the figure, the first convolutional module 2410 includes 8 convolutional kernels of size 3×3, which respectively perform convolutional operations on the 5×10 second feature map AFM with a stride of 1, so as to obtain the 3×8×8 feature map shown in Figure 2B 202.

[0039] The layer normalization (LN, Layer Normalization) module 2420 normalizes the features in the convolved feature map to obtain the normalized feature map shown in Figure 2B 203. It should be noted that since the layer normalization module normalizes the features on the same channel, compared with the commonly used batch normalization (BatchNormalization) module, this layer normalization processing can strengthen the correlation between the features of this layer, so that the features after layer normalization are more suitable for the capsule network module 240 to make full use of the vector characteristics of the capsules, thereby improving the accuracy of the capsule network module 240 in predicting the fatigue state of the object. Furthermore, in Figure 1In the single-channel EEG acquisition scheme shown, the accuracy of predicting the fatigue state by the capsule network module can be improved in the case of signal reduction.

[0040] Such as Figure 2B The normalized feature map shown in 203 is divided into 4 channels, and each channel contains 2D primary capsules. That is, each channel contains a sub-feature map of 3×8×2. Each vector on the dimension of size 2 represents a 2D primary capsule. In other words, a 2D primary capsule represents a vector containing two features. Therefore, each channel contains 3×8×1 primary capsules. The second convolution module 2430 performs convolution processing on the feature map containing 4 channels. In one example, the second convolution module 2430 may include multiple convolutional kernels. For example, in the example shown in the figure, the second convolution module 2430 includes 8 convolutional kernels of size 3×3, which respectively perform convolution operations on the corresponding 3×8 feature sub-maps with a stride of 1, so as to obtain Figure 2B The 1×6×2 feature sub-maps of each of the four channels shown in 204. For another example, the second convolution module 2430 may include 4 convolutional kernels of size 3×3, which respectively perform convolution operations on the 3×8 feature sub-maps of the corresponding channels with a stride of 1. In another example, the second convolution module 2430 may include one convolutional kernel. For example, in the example shown in the figure, the second convolution module 2430 includes 1 convolutional kernel of size 3×3, which performs convolution operations on the corresponding 3×8 feature sub-maps with a stride of 1, so as to obtain Figure 2B The 1×6×2 feature sub-maps of each of the four channels shown in 204. As Figure 2B Shown in 204, after the convolution operation of the second convolution module 2430, 4×6×1 primary capsules are obtained. Each primary capsule represents a vector containing two features, which is called a 2D vector. Those skilled in the art can understand that although this paragraph describes dividing the normalized feature map into 4 channels first and then performing convolution processing on the 3×8 sub-feature maps of each channel, in actual implementation, the above division operation and convolution operation do not actually need to have a sequence. For example, the second convolution module 2430 performs convolution operations on the 3×8×8 normalized feature map to obtain a 1×6×8 feature map, and the 1×6×8 feature map is divided into 4 channels to obtain 1×6×4 primary capsules.

[0041] The capsule processing module 2440 processes multiple primary capsules to obtain a capsule-processed feature map. The capsule processing module 2440 may include multiple capsule processing sub-modules. For example, in Figure 2A And 2B In the example shown, the capsule processing module 2440 includes 10 capsule processing sub-modules, and each capsule processing sub-module receives Figure 2BThe 24 primary capsules (24 two-dimensional vectors) shown in 204 and generate a 1×16 feature sub-map. Correspondingly, 10 capsule processing sub-modules obtain a 10×16 feature map based on the above 24 primary capsules, as shown in Figure 2B shown in 205. Then, the capsule processing module 2440 squeezes the 10×16 feature map into a 10×1 feature map, as shown in Figure 2B shown in 206. The 10 capsule processing sub-modules included in the capsule processing module 2440 can be referred to as high-level capsules. Those skilled in the art can understand that the technical solution of this application is not limited to the above specific implementation manner of the capsule processing module, and the existing implementation manners and their deformations of the capsule processing module in the art can be used in the technical solution of this application.

[0042] The output module 2450 obtains the result O representing the fatigue state based on the capsule-processed feature map. For example, the output module 2450 can be implemented by a fully connected network (FCN, Full Connected Network), which generates the result representing the fatigue state based on the capsule-processed feature map shown in Figure 2B shown in 206. For example, the result can be a value between 0 and 1, which represents the fatigue degree of an object such as a driver.

[0043] The first convolution module 2410, layer normalization module 2420, second convolution module 2430, capsule processing module 2440, and output module 2450 in the above capsule network module 240 can also be referred to as the first convolution layer, layer normalization layer, second convolution layer, capsule processing layer, and output layer, respectively. Together with the self-attention module 230, they constitute a neural network for fatigue detection.

[0044] Figure 3A shows a schematic diagram of a device for fatigue detection according to an embodiment.

[0045] Figure 3A In Figures 2A to 2C the same reference numerals represent the same components. Compared with Figure 2A in Figure 3A the embodiment does not include the residual module 2330. Therefore, the second feature map AFM output by the attention sub-module 2320 is provided to the capsule network module 240.

[0046] Figure 3B shows a schematic diagram of a device for fatigue detection according to an embodiment.

[0047] Figure 3B In Figures 2A to 3A the same reference numerals represent the same components. Compared with Figure 2A in Figure 3BIn the embodiment, the second convolutional module 2430 is not included, so the normalized feature map output by the LN module 2420 is divided into 3×8×4 two-dimensional primary capsules for processing by the capsule processing module 2440.

[0048] Those skilled in the art can understand that obvious deformations can be made to Figures 2A to 3B the embodiments shown to obtain equivalent embodiments, and these embodiments obtained by obvious deformations are included in the protection scope of this application. For example, although in Figures 2A to Figure 3A the embodiments shown, EEG signals collected from the forehead of an object via a single-channel acquisition device are used for fatigue detection, in other embodiments, multi-channel EEG signals collected from the head or forehead of an object by a multi-channel acquisition device can also be used for fatigue detection. In other words, the technical solution of this application can also be applied to multi-channel EEG signals. Another example is that although in Figures 2A to Figure 3A the embodiments shown, the self-attention module 230 and the capsule network module 240 with specific structures are described, in other embodiments, the self-attention module 230 and the capsule network module 240 can be implemented by other neural networks with appropriate structures that can achieve self-attention processing and capsule processing.

[0049] Figure 4 FIG. shows a schematic diagram of a process for training a fatigue detection module for fatigue detection according to an embodiment.

[0050] In the embodiments shown, in order to improve the accuracy of the fatigue detection module, Figure 4 the generative adversarial network (GAN, Generative Adversarial Network) 400 shown is used to generate more training data. In one embodiment, the generative adversarial network 400 can be a conditional generative adversarial network (CGAN, Conditional GAN).

[0051] In Figure 4 the embodiments shown, first, the features of the real training data are normalized. For example, a training data includes Figure 2C the feature map shown and the corresponding label representing the degree of fatigue. Based on the normalized real training data (X r , y), the generator G and the discriminator D in the CGAN 400 are trained. The generator G generates the feature X g corresponding to the label y based on the received Gaussian random noise z and the label y, and the discriminator D is based on the feature X r in the real training data and the label y and the generated feature X gDetermine whether the generated feature is true or false (R / F), and update the CGAN 400 based on the judgment result R / F, and finally obtain the trained generator G. Then, the trained generator G generates the feature X corresponding to the label y based on the Gaussian random noise z and the desired label y g . The generated feature X g and the label form a new training data set together with the features X r and labels of the original real training data. The features X r of the real training data and the features X g of the generated training data both have Figure 2C the format shown, so the accuracy of the fatigue detection model can be further improved by expanding the training data.

[0052] The fatigue detection model can be trained with the new expanded data set. Specifically, the self-attention module 230 in the fatigue detection module performs self-attention processing on the data representing the physiological signal in the new training data set to obtain a second feature map, and the capsule network module 240 in the fatigue detection module processes the second feature map to predict the result representing the fatigue state, and updates the fatigue detection module based on the predicted result representing the fatigue state and the corresponding label representing the fatigue state. The loss value can be determined based on the predicted value and label of the fatigue state, and then the fatigue detection module as a neural network model can be updated based on the loss value. For example, the well-known AdamW optimizer can be used to update the fatigue detection module based on the loss value.

[0053] Figure 5 FIG. shows a flowchart of a method for estimating the fatigue state of an object according to an embodiment.

[0054] In step 510, the physiological signal of the object is obtained. For example, the object may be a driver driving a vehicle. In one embodiment, the physiological signal includes an EEG signal. In one embodiment, the EEG signal of the object collected by a single-channel acquisition device. In one embodiment, the EEG signal is collected through the forehead of the object by a single-channel acquisition device.

[0055] In step 520, features are extracted from the physiological signal of the object and a first feature map is formed using the extracted features. In one embodiment, the features include PSD and DE extracted from the physiological signal, and the first feature map includes a single-channel feature map composed of multiple PSDs and multiple DEs.

[0056] In step 530, the self-attention module in the fatigue detection module performs self-attention processing on the first feature map to obtain a second feature map.

[0057] In step 540, the capsule network module in the fatigue detection module processes the second feature map to obtain a result representing the fatigue state.

[0058] In one embodiment, the self-attention module includes a position encoding module and a self-attention sub-module. In step 530, the position encoding module performs position encoding on the first feature map to obtain position information; the position information is added to the first feature map to obtain a feature map containing position information; the self-attention sub-module performs self-attention processing on the feature map containing position information to obtain the second feature map.

[0059] In one embodiment, the self-attention module includes a position encoding module, a self-attention sub-module, and a residual module. In step 530, the position encoding module performs position encoding on the first feature map to obtain position information; the position information is added to the first feature map to obtain a feature map containing position information; the self-attention sub-module performs self-attention processing on the feature map containing position information to obtain a third feature map; the residual module combines the feature map containing position information and the third feature map to obtain a combined feature map, and performs normalization processing on the combined feature map to obtain the second feature map.

[0060] In one embodiment, the capsule network module includes a first convolutional module, a layer normalization module, a second convolutional module, a capsule processing module, and an output module. In step 540, the first convolutional module performs convolutional processing on the second feature map to obtain a first convolutionally processed feature map; the layer normalization module performs layer normalization processing on the first convolutionally processed feature map to obtain a normalized feature map; the second convolutional module performs convolutional processing on the normalized feature map to obtain a second convolutionally processed feature map, where the second convolutionally processed feature map is divided into multiple primary capsules, and each primary capsule includes a vector, such as Figure 2B the 2D vector shown in 204; the capsule processing module processes the multiple primary capsules to obtain a capsule-processed feature map, such as Figure 2B the feature map shown in 206; the output module obtains the result representing the fatigue state based on the capsule-processed feature map.

[0061] In one embodiment, the first convolutional module includes a plurality of convolutional kernels, the second convolutional module includes one or more convolutional kernels, the capsule processing module includes a plurality of capsule processing sub-modules, and the output module includes a fully-connected network module. In step 540, the plurality of convolutional kernels of the first convolutional module respectively perform convolutional processing on the second feature map to obtain a plurality of first convolution-processed feature sub-maps, and the plurality of first convolution-processed feature sub-maps constitute the first convolution-processed feature map; one or more convolutional kernels of the second convolutional module respectively perform convolutional processing on the plurality of normalized feature sub-maps to obtain a plurality of second convolution-processed feature sub-maps, and the plurality of second convolution-processed feature sub-maps constitute the second convolution-processed feature map; the plurality of capsule processing sub-modules respectively process the plurality of primary capsules to obtain a capsule-processed feature map; and the fully-connected network module obtains the result representing the fatigue state based on the capsule-processed feature map.

[0062] Figure 6 FIG. shows a flowchart of a method for training a fatigue detection module for estimating the fatigue state of an object according to one embodiment.

[0063] In step 610, a first training data set is obtained, and the first training data set includes data representing physiological signals and labels representing fatigue states.

[0064] In step 620, the fatigue detection module is trained based on the first training data set. Step 620 may include steps 6210 to 6230. In step 6210, the self-attention module in the fatigue detection module performs self-attention processing on the data representing physiological signals in the training data in the first training data set to obtain a second feature map; in step 6220, the capsule network module in the fatigue detection module processes the second feature map to predict a result representing the fatigue state; and in step 6230, the fatigue detection module is updated based on the predicted result representing the fatigue state and the corresponding label representing the fatigue state.

[0065] In one embodiment, at step 610, a second training dataset is obtained, and a GAN is trained based on the second training dataset. Each training data in the second training dataset includes data representing a physiological signal and a label representing a fatigue state. The generator in the trained GAN generates a third training dataset, where each training data in the third training dataset includes data representing a physiological signal and a label representing a fatigue state. The first training dataset is generated from at least a part of the second training dataset and at least a part of the third training dataset. In one embodiment, the GAN is a CGAN. In one embodiment, the second training dataset is a real training dataset, and each training data therein includes a feature map and a corresponding label. The feature map is, for example, a single-channel feature map including various types of features as shown in Figure 3, and the label is a numerical value representing a fatigue state. In one embodiment, the physiological signal includes an EEG signal, and the data representing the physiological signal includes PSD and DE extracted from the EEG signal. In one embodiment, each training data in the generated third training dataset includes data representing a physiological signal and a label representing a fatigue state. In one embodiment, the generated data representing the physiological signal is, for example, a single-channel feature map including various types of features as shown in Figure 3.

[0066] In one embodiment, at step 610, training the GAN based on the second training dataset includes: the generator in the GAN generates data representing a physiological signal based on the label of the training data in the second training dataset and random noise; the discriminator in the GAN discriminates the authenticity of the generated data representing a physiological signal based on the generated data representing a physiological signal, the label in the training data, and the corresponding data representing a physiological signal; and the generator and the discriminator of the GAN are updated based on the discrimination result. In step 610, generating the third training dataset by the generator in the trained GAN includes: the trained GAN generates data representing a physiological signal based on the label representing a fatigue state and noise, and uses the label representing a fatigue state and the correspondingly generated data representing a physiological state as the third training dataset. In one embodiment, the label representing fatigue as the input to the trained GAN can be any label, such as a random label value between 0 and 1, and the noise as the input to the trained GAN can be any noise, such as random noise. Since the second training dataset is a real training dataset and the third training dataset is a training dataset generated based on the trained GAN, generating the first training dataset from at least a part of the second training dataset and at least a part of the third training dataset can achieve the augmentation of the training data.

[0067] Figure 7 The block diagram of a device for estimating the fatigue state of an object according to one embodiment is shown.

[0068] The device 700 includes a feature extraction module 710, a feature map construction module 720, and a fatigue detection module 730. The feature extraction module 710 is used to obtain the physiological signal of the object and extract features from the physiological signal. The feature map construction module 720 forms a first feature map using the features. The fatigue detection module 720 includes a self-attention module and a capsule network module. The self-attention module performs self-attention processing on the first feature map to obtain a second feature map, and the capsule network module processes the second feature map to obtain a result representing the fatigue state.

[0069] In one embodiment, the physiological signal includes an EEG signal, and the feature extraction module obtains the EEG signal of the object collected by a single-channel acquisition device. In one embodiment, the single-channel acquisition device collects the electroencephalogram signal through the forehead of the object. In one embodiment, the features include PSD and DE extracted from the physiological signal, and the first feature map includes a single-channel feature map composed of multiple PSDs and multiple DEs.

[0070] In one embodiment, the self-attention module includes a position encoding module and a self-attention sub-module. Among them, the position encoding module performs position encoding on the first feature map to obtain position information. The position information is added to the first feature map to obtain a feature map containing position information. The attention sub-module performs self-attention processing on the feature map containing position information to obtain the second feature map.

[0071] In one embodiment, the self-attention module includes a position encoding module, a self-attention sub-module, and a residual module. Among them, the position encoding module performs position encoding on the first feature map to obtain position information. The position information is added to the first feature map to obtain a feature map containing position information. The attention sub-module performs self-attention processing on the feature map containing position information to obtain a third feature map. The residual module combines the feature map containing position information and the third feature map to obtain a combined feature map, and performs normalization processing on the combined feature map to obtain the second feature map.

[0072] In one embodiment, the capsule network module includes a first convolutional module, a layer normalization module, a second convolutional module, a capsule processing module, and an output module. Among them, the first convolutional module performs convolutional processing on the second feature map to obtain a first convolutionally processed feature map; the layer normalization module performs layer normalization processing on the first convolutionally processed feature map to obtain a normalized feature map; the second convolutional module performs convolutional processing on the normalized feature map to obtain a second convolutionally processed feature map, where the second convolutionally processed feature map is divided into a plurality of primary capsules, and each primary capsule includes a vector; the capsule processing module processes the plurality of primary capsules to obtain a capsule-processed feature map; the output module obtains the result representing the fatigue state based on the capsule-processed feature map.

[0073] In one embodiment, the first convolutional module includes a plurality of convolutional kernels, the second convolutional module includes one or more convolutional kernels, the capsule processing module includes a plurality of capsule processing sub-modules, and the output module includes a fully connected network module. Among them, the plurality of convolutional kernels of the first convolutional module respectively perform convolutional processing on the second feature map to obtain a plurality of first convolutionally processed feature sub-maps, and the plurality of first convolutionally processed feature sub-maps constitute the first convolutionally processed feature map; the one or more convolutional kernels of the second convolutional module respectively perform convolutional processing on the plurality of normalized feature sub-maps to obtain a plurality of second convolutionally processed feature sub-maps, and the plurality of second convolutionally processed feature sub-maps constitute the second convolutionally processed feature map; the plurality of capsule processing sub-modules respectively process the plurality of primary capsules to obtain a capsule-processed feature map; the fully connected network module obtains the result representing the fatigue state based on the capsule-processed feature map.

[0074] Figure 8 FIG. shows a block diagram of an apparatus for estimating the fatigue state of an object according to one embodiment. Apparatus 800 may include one or more control units or processing units 810, and the control unit 810 executes one or more machine-readable instructions stored or encoded in a machine-readable storage medium (i.e., memory 820). Although not shown in Figure 8 FIG., those skilled in the art can understand that the control system 800 may include various other components, such as various communication modules, bus modules, and possibly user interface modules, etc. In one embodiment, apparatus 800 further includes a signal acquisition device for acquiring physiological signals of an object such as a driver. In one embodiment, the signal acquisition device is a single-channel acquisition device for acquiring physiological signals through the forehead of the object. In one embodiment, when executing the program instructions, the processing unit 810 is configured to execute the various operations and functions described above in conjunction with Figures 1 - 7 FIG.

[0075] According to one embodiment, a program product such as a non-transitory machine-readable medium is provided. The non-transitory machine-readable medium may have instructions that, when executed by a processing unit 810, are capable of performing the various operations and functions described above in the respective embodiments of the present application in combination with Figures 1 - 7 the various operations and functions described.

[0076] In one embodiment, to verify the accuracy of the fatigue detection method proposed in the present disclosure, various fatigue detection methods were compared on the publicly available vigilance estimation dataset SEED-VIG. A driving experiment was conducted in a simulated driving system, in which a number of healthy participants with normal vision participated in the driving experiment. During driving, electroencephalogram (EEG) signals at multiple acquisition points on the forehead were recorded at a predetermined sampling rate, thereby collecting multiple single-channel EEG signals. The collected data was labeled by the percentage of eye closure (PERCLOS) recorded by an eye-tracking glasses, which ranges from 0 (high vigilance level, corresponding to low fatigue) to 1 (low vigilance level, corresponding to high fatigue).

[0077] For cross-participant verification, leave-one-subject-out (LOSO) cross-validation was adopted, in which the data of some participants was used for training and the data of the remaining participants was used for testing. To evaluate the performance of the fatigue detection method, the following two metrics were used: root mean square error (RMSE) and Pearson correlation coefficient (PCC): where RMSE represents the error between the true value and the predicted value, and the smaller the value, the better the performance; PCC represents the Pearson correlation coefficient between the true value and the predicted value, and the larger the value, the better the performance; SD represents the standard deviation. RMSE, PCC, and SD are commonly used metrics in the art for measuring performance, and further details thereof will not be elaborated.

[0078] By comparing the experimental results of different fatigue detection methods for processing the input data of all channels (i.e., multi-channel) and the input data of each different single channel, the comparison of RMSE, PCC, and SD based on different fatigue detection methods specifically shows that: according to the method of the first embodiment of the present disclosure, in which the training sample expansion through CGAN, the self-attention module, and the capsule network module are adopted, according to the method of the second embodiment of the present disclosure, in which the training sample expansion through CGAN is not adopted and only the self-attention module and the capsule network module are adopted, according to the method of the third embodiment of the present disclosure, in which the training sample expansion through CGAN and the self-attention module are not adopted and only the capsule network module is adopted, compared with the experimental results of using other methods, it has better overall performance. The fatigue detection method provided by the present disclosure, especially the method of the first embodiment, has higher accuracy compared with other fatigue detection methods. Further based on the experimental results of all channels and multiple single channels of the method of the present disclosure, it is shown that the error for at least one single channel is less than the error for all channels, indicating the feasibility of detecting fatigue based on single-channel electroencephalogram of the forehead.

[0079] By comparing the time required for each test of different fatigue detection methods for single-channel features and multi-channel features, it can be seen that the fatigue detection method provided by the present disclosure is not only superior to other methods in terms of accuracy but also has high real-time performance in terms of processing time, and the result can be detected within about 6 milliseconds for a single channel. In addition, the fatigue detection time for four-channel features significantly exceeds the fatigue detection time for single-channel features. If electroencephalogram features of more channels (usually between 17 and 61 channels) are used, the detection time will increase significantly. This further illustrates the improvement of using single-channel electroencephalogram data for fatigue detection in terms of detection real-time performance.

[0080] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "example" or "exemplary" used throughout this specification means "serving as an example, instance, or illustration", and does not mean "preferred" or "advantageous" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.

[0081] The foregoing description of the content of this application is provided to enable any ordinary person skilled in the art to implement or use the content of this application. Various modifications to the content of this application will be obvious to those of ordinary skill in the art, and also, without departing from the scope of protection of the content of this application, the general principles defined herein can be applied to other variations. Therefore, the content of this application is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.

Claims

1. A method for estimating the fatigue state of an object, comprising: Obtaining a physiological signal of the object; Extracting features from the physiological signal and forming a first feature map using the features; Performing self-attention processing on the first feature map by a self-attention module in a fatigue detection module to obtain a second feature map; Processing the second feature map by a capsule network module in the fatigue detection module to obtain a result representing the fatigue state.

2. The method according to claim 1, wherein The physiological signal includes an electroencephalogram (EEG) signal, and obtaining the physiological signal of the object includes: obtaining the EEG signal of the object collected by a single-channel acquisition device.

3. The method according to claim 2, wherein, The single-channel acquisition device collects the EEG signal through the forehead of the object.

4. The method according to one of claims 1 to 3, wherein The features include power spectral density (PSD) and differential quotient (DE) extracted from the physiological signal, and the first feature map includes a single-channel feature map composed of multiple PSDs and multiple DEs.

5. The method according to claim 4, wherein, The self-attention module includes a position encoding module and a self-attention sub-module. Among them, performing self-attention processing on the first feature map by the self-attention module in the fatigue detection module includes: Performing position encoding on the first feature map by the position encoding module to obtain position information; Adding the position information to the first feature map to obtain a feature map containing position information; Performing self-attention processing on the feature map containing position information by the attention sub-module to obtain the second feature map.

6. The method according to claim 4, wherein, The self-attention module includes a position encoding module, a self-attention sub-module, and a residual module. Among them, performing self-attention processing on the first feature map by the self-attention module in the fatigue detection module includes: Performing position encoding on the first feature map by the position encoding module to obtain position information; Adding the position information to the first feature map to obtain a feature map containing position information; Performing self-attention processing on the feature map containing position information by the attention sub-module to obtain a third feature map; Combining the feature map containing position information and the third feature map by the residual module to obtain a combined feature map, and performing normalization processing on the combined feature map to obtain the second feature map.

7. The method according to claim 4, wherein, The capsule network module includes a first convolution module, a layer normalization module, a second convolution module, a capsule processing module, and an output module. Among them, processing the second feature map by the capsule network module in the fatigue detection module includes: Performing convolution processing on the second feature map by the first convolution module to obtain a first convolved feature map; Performing layer normalization processing on the first convolved feature map by the layer normalization module to obtain a normalized feature map; Performing convolution processing on the normalized feature map by the second convolution module to obtain a second convolved feature map, where the second convolved feature map is divided into multiple primary capsules; Processing the multiple primary capsules by the capsule processing module to obtain a capsule-processed feature map; Obtaining the result representing the fatigue state based on the capsule-processed feature map by the output module.

8. The method according to claim 7, wherein, The first convolution module includes a plurality of convolutional kernels, the second convolution module includes one or more convolutional kernels, the capsule processing module includes a plurality of capsule processing sub-modules, and the output module includes a fully-connected network module. Among them, the convolution processing of the second feature map by the first convolution module includes: the plurality of convolutional kernels of the first convolution module respectively perform convolution processing on the second feature map to obtain a plurality of first convolution-processed feature sub-maps, and the plurality of first convolution-processed feature sub-maps constitute the first convolution-processed feature map. Among them, the convolution processing of the normalized feature map by the second convolution module includes: the one or more convolutional kernels of the second convolution module respectively perform convolution processing on the plurality of normalized feature sub-maps to obtain a plurality of second convolution-processed feature sub-maps, and the plurality of second convolution-processed feature sub-maps constitute the second convolution-processed feature map. Among them, the processing of the plurality of primary capsules by the capsule processing module includes: the plurality of capsule processing sub-modules respectively process the plurality of primary capsules to obtain a capsule-processed feature map. Among them, the output module obtaining the result representing the fatigue state based on the capsule-processed feature map includes: the fully-connected network module obtaining the result representing the fatigue state based on the capsule-processed feature map.

9. A method for training a fatigue detection module for estimating the fatigue state of an object, including: Obtaining a first training data set, where the first training data set includes data representing physiological signals and labels representing fatigue states. Training the fatigue detection module based on the first training data set, including: The self-attention module in the fatigue detection module performs self-attention processing on the data representing physiological signals in the training data in the first training data set to obtain a second feature map. The capsule network module in the fatigue detection module processes the second feature map to predict a result representing the fatigue state. Updating the fatigue detection module based on the predicted result representing the fatigue state and the corresponding label representing the fatigue state.

10. The method according to claim 9, wherein, The obtaining of the first training data set includes: Obtaining a second training data set, training a generative adversarial network (GAN) based on the second training data set, where each training data in the second training data set includes data representing physiological signals and labels representing fatigue states. The generator in the trained GAN generates a third training data set, where each training data in the third training data set includes data representing physiological signals and labels representing fatigue states. Generating the first training data set from at least a part of the second training data set and at least a part of the third training data set.

11. The method according to claim 10, where the generating of the third training data set includes: Obtaining any label and any noise. The generator in the trained GAN generates data representing physiological signals based on the arbitrary label and the arbitrary noise, and the set of the arbitrary label and the generated corresponding data representing physiological signals constitutes the third training dataset.

12. The method according to claim 9, wherein, The physiological signal includes an electroencephalogram (EEG) signal, and the data representing the physiological signal includes the power spectral density (PSD) and the differential quotient (DE) extracted from the physiological signal. Each training data in the training dataset includes a single-channel feature map composed of a plurality of PSDs and a plurality of DEs and a corresponding label representing the fatigue state.

13. The method according to claim 10, wherein, The training of the GAN based on the second training dataset includes: The generator in the GAN generates data representing physiological signals based on the label of the training data in the second training dataset and random noise; The discriminator in the GAN discriminates the authenticity of the generated data representing physiological signals based on the generated data representing physiological signals, the label in the training data, and the corresponding data representing physiological signals; Based on the discrimination result, the generator and the discriminator of the GAN are updated.

14. An apparatus for estimating the fatigue state of an object, comprising: A feature extraction module, configured to obtain the physiological signal of the object and extract features from the physiological signal; A feature map construction module, configured to form a first feature map by using the features; A fatigue detection module, which includes a self-attention module and a capsule network module. The self-attention module performs self-attention processing on the first feature map to obtain a second feature map, and the capsule network module processes the second feature map to obtain a result representing the fatigue state.

15. An apparatus for estimating the fatigue state of an object, comprising: One or more processors; One or more memories, the memories storing computer-executable instructions, and the instructions, when executed, cause the one or more processors to execute the method according to any one of claims 1 to 13.

16. The apparatus according to claim 15, further comprising a signal acquisition device for acquiring the physiological signal of the object.

17. The device according to claim 16, wherein, The signal acquisition device is a single-channel acquisition device for acquiring the physiological signal through the forehead of the object.

18. A machine-readable storage medium storing executable instructions, and the instructions, when executed, cause one or more processors to execute the method according to any one of claims 1 to 13.

19. A computer program product comprising computer-executable instructions, and the instructions, when executed, cause one or more processors to execute the method according to any one of claims 1 to 13.