Fatigue driving detection method, system and device based on multi-modal feature fusion and medium
Through the multimodal feature fusion method, combined with face images, brain waves and steering wheel grip signal, the driver's fatigue state is detected, solving the problems of misjudgment and misjudgment in the existing technology, and improving the accuracy of detection and driving safety.
Patent Information
- Application Number
- CN202510117355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is prone to misjudgment and misjudgment when detecting driver fatigue, which affects the accuracy of detection and driving safety.
The fatigue driving detection method using multimodal feature fusion is used to obtain the driver's face image information, brain wave signals and steering wheel grip strength signals, and the eye characteristics, mouth characteristics, concentration characteristics, relaxation characteristics and grip strength change characteristics are extracted, and the feature fusion is performed, and input it to the pre-trained fatigue state detection model for judgment.
It improves the accuracy of fatigue driving detection, enhances users' driving safety, and reduces the occurrence of misjudgments and misjudgments.
Smart Images

Figure CN119961867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle monitoring technology, and in particular to a multi-modal feature fusion fatigue driving detection method, system, device and medium. Background Art
[0002] Fatigue driving refers to the changes in the driver's psychological function and physiological mechanism after long-term continuous driving, which objectively manifests as a decline in driving skills, drowsiness, slow reaction, limb weakness, lack of concentration and decreased judgment. According to incomplete statistics, 50% of traffic safety accidents are caused by drivers' unconsciousness.
[0003] Fatigue driving can cause great safety risks to the driver. Real-time detection of the driver's fatigue state is a key part of safe car driving. Therefore, when driving a vehicle, it is necessary to detect the driver's fatigue state in real time and issue a timely warning.
[0004] In the existing technology, most of them are based on deep learning technology to detect possible fatigue driving behaviors during driving, including closing eyes, yawning, etc., so as to realize driver fatigue detection. However, this method only makes judgments based on the driver's current facial image, which is prone to misjudgment and missed judgment, affecting the accuracy of fatigue driving detection and the user's driving safety. Summary of the invention
[0005] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0006] To this end, an object of an embodiment of the present invention is to provide a fatigue driving detection method with multimodal feature fusion, which improves the accuracy of fatigue driving detection and the driving safety of users.
[0007] Another object of an embodiment of the present invention is to provide a fatigue driving detection system with multimodal feature fusion.
[0008] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0009] In a first aspect, an embodiment of the present invention provides a method for detecting fatigue driving by fusion of multimodal features, comprising the following steps:
[0010] Acquire facial image information of a driver, and extract eye feature data and mouth feature data of the driver according to the facial image information;
[0011] Acquire brain wave signals of the driver, and extract concentration characteristic data and relaxation characteristic data of the driver according to the brain wave signals;
[0012] Acquire a steering wheel grip force signal of a driver, and extract grip force change characteristic data of the driver according to the steering wheel grip force signal;
[0013] Performing feature fusion on the eye feature data, the mouth feature data, the concentration feature data, the relaxation feature data, and the grip force change feature data to obtain multimodal feature data;
[0014] The multimodal feature data is input into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then whether the driver is driving in fatigue is determined based on the fatigue level detection value.
[0015] Further, in one embodiment of the present invention, the step of acquiring facial image information of the driver and extracting eye feature data and mouth feature data of the driver according to the facial image information specifically includes:
[0016] Acquiring the facial image information through a camera device;
[0017] Performing key point detection on the facial image information to obtain a plurality of facial key points;
[0018] Extracting an eye region image and a mouth region image according to the facial key points;
[0019] Determining the degree of eye closure of the driver according to the eye region image, and generating the eye feature data according to the degree of eye closure corresponding to a plurality of consecutive frames of face images;
[0020] The degree of opening and closing of the driver's mouth is determined according to the mouth area image, and the mouth feature data is generated according to the degree of opening and closing of the mouth corresponding to multiple consecutive frames of facial images.
[0021] Further, in one embodiment of the present invention, the step of acquiring the driver's brain wave signal and extracting the driver's concentration characteristic data and relaxation characteristic data according to the brain wave signal specifically includes:
[0022] Acquiring the brain wave signal through a brain wave sensor;
[0023] Performing signal analysis on the brain wave signal to obtain signal energy proportions of theta band, alpha band, beta band and gamma band;
[0024] Determine the concentration of the driver according to the signal energy proportions of the α band, the β band, and the γ band and a preset first weight coefficient, and generate the concentration feature data according to the concentration corresponding to a plurality of consecutive frames of brain wave signals;
[0025] The driver's relaxation degree is determined according to the signal energy proportion of theta band, alpha band, and beta band and a preset second weight coefficient, and the relaxation degree characteristic data is generated according to the relaxation degree corresponding to multiple frames of continuous brain wave signals.
[0026] Furthermore, in one embodiment of the present invention, the driver's steering wheel grip force signal is obtained, and the driver's grip force change characteristic data is extracted according to the steering wheel grip force signal:
[0027] Acquiring the steering wheel grip force signal through a grip force sensing sensor;
[0028] The steering wheel grip force value of the driver is determined according to the steering wheel grip force signal, and the grip force change characteristic data is generated according to the steering wheel grip force values corresponding to multiple consecutive frames of steering wheel grip force signals.
[0029] Furthermore, in one embodiment of the present invention, the fatigue state detection model is trained by the following steps:
[0030] Obtaining eye feature sample data, mouth feature sample data, concentration feature sample data, relaxation feature sample data, and grip strength change feature sample data of the tester;
[0031] Performing feature fusion on the eye feature sample data, the mouth feature sample data, the concentration feature sample data, the relaxation feature sample data, and the grip force change feature sample data to obtain multimodal feature sample data;
[0032] Determining fatigue level labels of the multimodal feature sample data through manual annotation;
[0033] Inputting the multimodal feature sample data into a pre-built deep learning neural network to obtain a fatigue degree recognition result;
[0034] determining a loss value according to the fatigue level identification result and the fatigue level label;
[0035] The parameters of the deep learning neural network are updated according to the loss value to obtain the trained fatigue state detection model.
[0036] Further, in one embodiment of the present invention, the multimodal feature data is input into a pre-trained fatigue state detection model to obtain a fatigue degree detection value of the driver, and then judging whether the driver is driving fatigued according to the fatigue degree detection value, which specifically includes:
[0037] Inputting the multimodal feature data of a plurality of continuous time periods into the fatigue state detection model respectively, and obtaining the fatigue degree detection values and corresponding confidence levels of the plurality of continuous time periods;
[0038] determining a third weight coefficient of each fatigue degree detection value according to the confidence level, and then performing weighted summation of the fatigue degree detection values of multiple consecutive time periods according to the third weight coefficient to obtain the fatigue degree of the driver;
[0039] When the fatigue level is greater than or equal to a preset first threshold, determining that the driver has fatigue driving behavior;
[0040] When the fatigue level is less than the first threshold, it is determined that the driver does not have fatigue driving behavior.
[0041] Furthermore, in one embodiment of the present invention, the fatigue driving detection method further includes the following steps:
[0042] When the driver is fatigued driving, the vehicle's fragrance system, music system, suspension system and collision warning system are adjusted according to the fatigue level, so that the fragrance system releases preset refreshing gas, the music system plays preset refreshing music, the suspension system is adjusted to the sports mode, and the sensitivity of the collision warning system is enhanced.
[0043] In a second aspect, an embodiment of the present invention provides a fatigue driving detection system with multimodal feature fusion, including:
[0044] A facial feature extraction module, used to obtain facial image information of a driver, and extract eye feature data and mouth feature data of the driver according to the facial image information;
[0045] A brain wave feature extraction module is used to obtain the brain wave signal of the driver, and extract the concentration feature data and relaxation feature data of the driver according to the brain wave signal;
[0046] A grip force feature extraction module, used to obtain a steering wheel grip force signal of a driver, and extract grip force change feature data of the driver according to the steering wheel grip force signal;
[0047] A feature fusion module, used for fusing the eye feature data, the mouth feature data, the concentration feature data, the relaxation feature data and the grip force change feature data to obtain multimodal feature data;
[0048] The fatigue state detection module is used to input the multimodal feature data into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then determine whether the driver is driving fatigued based on the fatigue level detection value.
[0049] In a third aspect, an embodiment of the present invention provides a fatigue driving detection device with multi-modal feature fusion, comprising:
[0050] at least one processor;
[0051] at least one memory for storing at least one program;
[0052] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned fatigue driving detection method of multimodal feature fusion.
[0053] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to execute the above-mentioned fatigue driving detection method based on multimodal feature fusion when executed by the processor.
[0054] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:
[0055] The embodiment of the present invention obtains facial image information of a driver, extracts eye feature data and mouth feature data of the driver based on the facial image information, obtains a brain wave signal of the driver, extracts concentration feature data and relaxation feature data of the driver based on the brain wave signal, obtains a steering wheel grip signal of the driver, extracts grip force change feature data of the driver based on the steering wheel grip signal, performs feature fusion on the eye feature data, mouth feature data, concentration feature data, relaxation feature data and grip force change feature data to obtain multimodal feature data, inputs the multimodal feature data into a pre-trained fatigue state detection model to obtain a fatigue degree detection value of the driver, and then determines whether the driver is driving fatigued based on the fatigue degree detection value. The embodiment of the present invention extracts eye feature data and mouth feature data based on the driver's facial image information, extracts concentration feature data and relaxation feature data based on the driver's brain wave signal, and extracts grip change feature data based on the driver's steering wheel grip signal. Multimodal feature data is obtained by fusing the eye feature data, mouth feature data, concentration feature data, relaxation feature data and grip change feature data. The multimodal feature data is input into a pre-trained fatigue state detection model, so that the driver's fatigue state can be detected from multiple dimensions such as the degree of eye closure, degree of mouth opening, concentration, relaxation and steering wheel grip, so as to judge whether the driver is driving fatigued according to the fatigue degree detection value, thereby improving the accuracy of fatigue driving detection and the driving safety of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0057] Figure 1 A flowchart of the steps of a method for detecting fatigue driving by fusion of multimodal features provided by an embodiment of the present invention;
[0058] Figure 2 A structural block diagram of a multi-modal feature fusion fatigue driving detection system provided in an embodiment of the present invention;
[0059] Figure 3 A structural block diagram of a multi-modal feature fusion fatigue driving detection device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0061] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.
[0062] Reference Figure 1 The embodiment of the present invention provides a fatigue driving detection method based on multi-modal feature fusion, which specifically includes the following steps:
[0063] S101, acquiring facial image information of a driver, and extracting eye feature data and mouth feature data of the driver according to the facial image information;
[0064] S102, obtaining brain wave signals of the driver, and extracting concentration characteristic data and relaxation characteristic data of the driver according to the brain wave signals;
[0065] S103, obtaining a steering wheel grip force signal of the driver, and extracting grip force change characteristic data of the driver according to the steering wheel grip force signal;
[0066] S104, performing feature fusion on the eye feature data, the mouth feature data, the concentration feature data, the relaxation feature data, and the grip strength change feature data to obtain multimodal feature data;
[0067] S105: Input the multimodal feature data into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then determine whether the driver is driving in fatigue according to the fatigue level detection value.
[0068] The embodiment of the present invention extracts eye feature data and mouth feature data based on the driver's facial image information, extracts concentration feature data and relaxation feature data based on the driver's brain wave signal, and extracts grip change feature data based on the driver's steering wheel grip signal. Multimodal feature data is obtained by fusing the eye feature data, mouth feature data, concentration feature data, relaxation feature data and grip change feature data. The multimodal feature data is input into a pre-trained fatigue state detection model, so that the driver's fatigue state can be detected from multiple dimensions such as the degree of eye closure, degree of mouth opening, concentration, relaxation and steering wheel grip, so as to judge whether the driver is driving fatigued according to the fatigue degree detection value, thereby improving the accuracy of fatigue driving detection and the driving safety of users.
[0069] As an optional implementation, facial image information of the driver is obtained, and eye feature data and mouth feature data of the driver are extracted based on the facial image information, which specifically includes:
[0070] S1011, acquiring facial image information through a camera device;
[0071] S1012, performing key point detection on the facial image information to obtain a plurality of facial key points;
[0072] S1013, extracting an eye region image and a mouth region image according to key points of the face;
[0073] S1014, determining the degree of eye closure of the driver according to the eye region image, and generating eye feature data according to the degree of eye closure corresponding to the continuous multiple frames of face images;
[0074] S1015. Determine the degree of opening and closing of the driver's mouth according to the mouth area image, and generate mouth feature data according to the degree of opening and closing of the mouth corresponding to multiple consecutive frames of facial images.
[0075] Specifically, facial image information of the driver is obtained through a camera device, and key point detection is performed on the facial image information to obtain multiple facial key points, such as the beginnings and ends of the eyes, the beginning of the nose, the corners of the mouth, etc. According to the detected facial key points, the eye area contour and the mouth area contour can be identified, thereby extracting the eye area image and the mouth area image; the degree of eye closure of the driver is detected according to the eye area image, for example, the maximum distance between the upper and lower eyelids in the eye area of the driver in a normal state is collected in advance, and the real-time distance between the upper and lower eyelids in the extracted eye area image and the ratio of the maximum distance are used to determine the degree of eye closure. The degree of eye closure of the driver is determined, and then time series data about the degree of eye closure is generated according to the degree of eye closure detected by multiple consecutive frames of face images, so as to obtain eye feature data; the degree of mouth opening and closing of the driver is detected according to the mouth area image, for example, the standard area of the mouth area of the driver in a normal state is collected in advance, and the degree of mouth opening and closing of the driver is determined according to the ratio of the real-time area of the mouth area in the extracted mouth area image and the standard area, and then time series data about the degree of mouth opening and closing is generated according to the degree of mouth opening and closing detected by multiple consecutive frames of face images, so as to obtain mouth feature data.
[0076] As an optional implementation, the driver's brain wave signal is obtained, and the driver's concentration characteristic data and relaxation characteristic data are extracted according to the brain wave signal, which specifically includes:
[0077] S1021, obtaining brain wave signals through a brain wave sensor;
[0078] S1022, performing signal analysis on the brain wave signal to obtain signal energy proportions of theta band, alpha band, beta band, and gamma band;
[0079] S1023, determining the driver's concentration according to the signal energy proportions of the α band, the β band, and the γ band and a preset first weight coefficient, and generating concentration feature data according to the concentration corresponding to the continuous multiple frames of brain wave signals;
[0080] S1024, determining the driver's relaxation degree according to the signal energy proportions of theta band, alpha band, and beta band and a preset second weight coefficient, and generating relaxation degree characteristic data according to the relaxation degrees corresponding to multiple frames of continuous brain wave signals.
[0081] Specifically, brain waves are a physiological signal of human brain activity, and their changes can reflect a person's fatigue status. The human brain continuously produces a variety of rhythmic waves, including theta band (4-8Hz), alpha band (8-12Hz), beta band (12-40Hz), gamma band (40-100Hz) and delta band (0-4Hz). A person's state of consciousness is determined by which band is dominant. Therefore, the driver's concentration and relaxation can be evaluated based on the signal energy proportion of each band, thereby obtaining concentration characteristic data and relaxation characteristic data.
[0082] The brain wave signals collected by the brain wave sensor are subjected to signal analysis to obtain the signal energy proportions of theta band, alpha band, beta band and gamma band in each signal frame, and the signal energy proportions of the alpha band, beta band and gamma band are weighted and summed according to the preset first weight coefficient to obtain the driver's concentration, and time series data is formed according to the concentration corresponding to the continuous multi-frame brain wave signals, which is the concentration feature data, and the signal energy proportions of the theta band, alpha band and beta band are weighted and summed according to the preset second weight coefficient to obtain the driver's relaxation, and time series data is formed according to the relaxation corresponding to the continuous multi-frame brain wave signals, which is the relaxation time series data. It should be noted that the specific values of the first weight coefficient and the second weight coefficient can be determined according to relevant research and experiments on brain waves, and the embodiments of the present invention are not described in detail here.
[0083] As an optional implementation, a steering wheel grip force signal of the driver is obtained, and the driver's grip force change characteristic data is extracted based on the steering wheel grip force signal:
[0084] S1031, obtaining a steering wheel grip force signal through a grip force sensing sensor;
[0085] S1032. Determine the driver's steering wheel grip force value according to the steering wheel grip force signal, and generate grip force change characteristic data according to the steering wheel grip force values corresponding to multiple consecutive frames of steering wheel grip force signals.
[0086] Specifically, when the driver is in a fatigued state, the hands holding the steering wheel will involuntarily relax, resulting in a decrease in grip strength. The steering wheel grip force signal is obtained by a grip force sensing sensor arranged on the steering wheel, so that the driver's real-time steering wheel grip force value can be determined. According to the steering wheel grip force values corresponding to multiple consecutive frames of steering wheel grip force signals, the time series data of the driver's steering wheel grip force value can be determined, which is the grip force change characteristic data.
[0087] In some optional embodiments, after obtaining the eye feature data, mouth feature data, concentration feature data, relaxation feature data and grip change feature data, they are respectively vectorized to obtain eye feature vectors, mouth feature vectors, concentration feature vectors, relaxation feature vectors and grip change feature vectors, and then vector splicing is performed to obtain multimodal feature data after feature fusion.
[0088] As an optional implementation, the fatigue state detection model is trained by the following steps:
[0089] S201, obtaining eye feature sample data, mouth feature sample data, concentration feature sample data, relaxation feature sample data, and grip strength change feature sample data of the tester;
[0090] S202, performing feature fusion on the eye feature sample data, the mouth feature sample data, the concentration feature sample data, the relaxation feature sample data, and the grip strength change feature sample data to obtain multimodal feature sample data;
[0091] S203, determining fatigue level labels of multimodal feature sample data through manual labeling;
[0092] S204, inputting the multimodal feature sample data into a pre-built deep learning neural network to obtain a fatigue degree recognition result;
[0093] S205, determining a loss value according to the fatigue level recognition result and the fatigue level label;
[0094] S206. Update the parameters of the deep learning neural network according to the loss value to obtain a trained fatigue state detection model.
[0095] Specifically, after inputting the multimodal feature sample data into the initialized deep learning neural network, the prediction result of the model output, i.e., the fatigue degree recognition result, can be obtained. The accuracy of the model prediction can be evaluated based on the fatigue degree recognition result and the aforementioned fatigue degree label, thereby updating the parameters of the model. For the fatigue state detection model, the accuracy of the model prediction result can be measured by the loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of a single training data and the prediction result of the model for the training data. In actual training, a training data set has a lot of training data, so the cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the model. For general machine learning models, based on the aforementioned cost function, plus the regularization term that measures the complexity of the model, it can be used as the objective function of the training, and the loss value of the entire training data set can be calculated based on the objective function. There are many types of commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross entropy loss function, etc., which can all be used as loss functions of machine learning models, which will not be elaborated here one by one. In an embodiment of the present invention, any one of the loss functions can be selected to determine the loss value of training. Based on the loss value of training, the back propagation algorithm is used to update the parameters of the model, and a trained fatigue state detection model can be obtained after several rounds of iteration. The specific number of iterations can be set in advance, or the training is considered to be completed when the test set meets the accuracy requirements.
[0096] As an optional implementation, the multimodal feature data is input into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then whether the driver is driving fatigued is determined according to the fatigue level detection value, which specifically includes:
[0097] S1051, respectively inputting the multimodal feature data of the plurality of continuous time periods into the fatigue state detection model to obtain fatigue degree detection values and corresponding confidence levels of the plurality of continuous time periods;
[0098] S1052, determining a third weight coefficient for each fatigue level detection value according to the confidence level, and then performing weighted summation of fatigue level detection values of multiple consecutive time periods according to the third weight coefficient to obtain the fatigue level of the driver;
[0099] S1053: When the fatigue level is greater than or equal to a preset first threshold, it is determined that the driver is in fatigue driving behavior;
[0100] S1054: When the fatigue level is less than the first threshold, it is determined that the driver does not have fatigue driving behavior.
[0101] Specifically, in order to avoid misjudgment and missed judgment in a single detection, an embodiment of the present invention obtains multimodal feature data of multiple continuous time periods, inputs the multimodal feature data of each time period into the fatigue state detection model to obtain the corresponding fatigue degree detection value and confidence (characterizing the confidence level of the fatigue degree detection value), and then determines the third weight coefficient of the fatigue degree detection value of each time period according to the confidence. The higher the confidence, the larger the weight coefficient. Therefore, the fatigue degree of the driver can be obtained by weighted summing the fatigue degree detection values of each time period according to the third weight coefficient, and then judging whether the driver has fatigue driving behavior based on the preset first threshold.
[0102] In some optional embodiments, in order to further accurately determine whether fatigue driving occurs, a comprehensive judgment may be made based on the number of safety warnings from the ADAS, such as ELK warnings, AEB warnings, RCTW warnings, and LKA warnings.
[0103] As an optional implementation, the fatigue driving detection method further includes the following steps:
[0104] S106. When the driver is driving fatigued, the vehicle's fragrance system, music system, suspension system and collision warning system are adjusted according to the degree of fatigue, so that the fragrance system releases preset refreshing gas, the music system plays preset refreshing music, the suspension system is adjusted to the sports mode, and the sensitivity of the collision warning system is enhanced.
[0105] Specifically, the fatigue driving detection results are input as parameters into the vehicle's associated system to execute corresponding alertness and safety countermeasures, including:
[0106] Fragrance system: Use refreshing scents and avoid comfortable scents;
[0107] Music system: Play passionate and refreshing music, and avoid playing soothing light music;
[0108] Suspension system: Adjust to sports mode to avoid excessive comfort;
[0109] Collision warning system: Enhanced sensitivity, that is, warnings are issued even when there is a low probability of a collision;
[0110] ADAS system: It is recommended that users turn on the ADAS system. If they agree, the ACC and lane keeping functions will be turned on.
[0111] The method steps of the embodiment of the present invention are described above. It can be understood that the embodiment of the present invention extracts eye feature data and mouth feature data according to the driver's facial image information, extracts concentration feature data and relaxation feature data according to the driver's brain wave signal, extracts grip change feature data according to the driver's steering wheel grip signal, and obtains multimodal feature data based on the fusion of eye feature data, mouth feature data, concentration feature data, relaxation feature data and grip change feature data. The multimodal feature data is input into a pre-trained fatigue state detection model, and the driver's fatigue state can be detected from multiple dimensions such as eye closure degree, mouth opening degree, concentration, relaxation and steering wheel grip, so as to judge whether the driver is driving fatigued according to the fatigue degree detection value, thereby improving the accuracy of fatigue driving detection and the driving safety of users.
[0112] Reference Figure 2 , an embodiment of the present invention provides a fatigue driving detection system with multi-modal feature fusion, comprising:
[0113] A facial feature extraction module is used to obtain facial image information of the driver, and extract eye feature data and mouth feature data of the driver based on the facial image information;
[0114] The brain wave feature extraction module is used to obtain the driver's brain wave signal, and extract the driver's concentration feature data and relaxation feature data based on the brain wave signal;
[0115] A grip force feature extraction module is used to obtain the driver's steering wheel grip force signal, and extract the driver's grip force change feature data based on the steering wheel grip force signal;
[0116] A feature fusion module is used to fuse eye feature data, mouth feature data, concentration feature data, relaxation feature data, and grip change feature data to obtain multimodal feature data;
[0117] The fatigue state detection module is used to input the multimodal feature data into a pre-trained fatigue state detection model to obtain the driver's fatigue level detection value, and then determine whether the driver is driving fatigued based on the fatigue level detection value.
[0118] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0119] Reference Figure 3 , an embodiment of the present invention provides a fatigue driving detection device with multi-modal feature fusion, comprising:
[0120] at least one processor;
[0121] at least one memory for storing at least one program;
[0122] When the at least one program is executed by the at least one processor, the at least one processor implements the multi-modal feature fusion fatigue driving detection method.
[0123] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0124] An embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned fatigue driving detection method based on multimodal feature fusion.
[0125] A computer-readable storage medium of an embodiment of the present invention can execute a fatigue driving detection method with multimodal feature fusion provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0126] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.
[0127] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0128] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0129] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0130] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0131] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.
[0132] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0133] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0134] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0135] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A fatigue driving detection method based on multimodal feature fusion, characterized in that: The following steps are involved: Acquire facial image information of a driver, and extract eye feature data and mouth feature data of the driver according to the facial image information; Acquire brain wave signals of the driver, and extract concentration characteristic data and relaxation characteristic data of the driver according to the brain wave signals; Acquire a steering wheel grip force signal of a driver, and extract grip force change characteristic data of the driver according to the steering wheel grip force signal; Performing feature fusion on the eye feature data, the mouth feature data, the concentration feature data, the relaxation feature data, and the grip force change feature data to obtain multimodal feature data; The multimodal feature data is input into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then whether the driver is driving in fatigue is determined based on the fatigue level detection value.
2. The method for detecting fatigue driving by multimodal feature fusion according to claim 1, characterized in that: The step of acquiring the facial image information of the driver and extracting the eye feature data and mouth feature data of the driver according to the facial image information specifically includes: Acquiring the facial image information through a camera device; Performing key point detection on the facial image information to obtain a plurality of facial key points; Extracting an eye region image and a mouth region image according to the facial key points; Determining the degree of eye closure of the driver according to the eye region image, and generating the eye feature data according to the degree of eye closure corresponding to a plurality of consecutive frames of face images; The degree of opening and closing of the driver's mouth is determined according to the mouth area image, and the mouth feature data is generated according to the degree of opening and closing of the mouth corresponding to multiple consecutive frames of facial images.
3. The method for detecting fatigue driving by multimodal feature fusion according to claim 1, characterized in that: The step of obtaining the driver's brain wave signal and extracting the driver's concentration characteristic data and relaxation characteristic data according to the brain wave signal specifically includes: Acquiring the brain wave signal through a brain wave sensor; Performing signal analysis on the brain wave signal to obtain signal energy proportions of theta band, alpha band, beta band and gamma band; Determine the concentration of the driver according to the signal energy proportions of the α band, the β band, and the γ band and a preset first weight coefficient, and generate the concentration feature data according to the concentration corresponding to a plurality of consecutive frames of brain wave signals; The driver's relaxation degree is determined according to the signal energy proportion of theta band, alpha band, and beta band and a preset second weight coefficient, and the relaxation degree characteristic data is generated according to the relaxation degree corresponding to multiple frames of continuous brain wave signals.
4. The method for detecting fatigue driving by multimodal feature fusion according to claim 1, characterized in that: The driver's steering wheel grip force signal is obtained, and the driver's grip force change characteristic data is extracted according to the steering wheel grip force signal: Acquiring the steering wheel grip force signal through a grip force sensing sensor; The steering wheel grip force value of the driver is determined according to the steering wheel grip force signal, and the grip force change characteristic data is generated according to the steering wheel grip force values corresponding to multiple consecutive frames of steering wheel grip force signals.
5. The method for detecting fatigue driving by multimodal feature fusion according to claim 1, characterized in that: The fatigue state detection model is trained by the following steps: Obtaining eye feature sample data, mouth feature sample data, concentration feature sample data, relaxation feature sample data, and grip strength change feature sample data of the tester; Performing feature fusion on the eye feature sample data, the mouth feature sample data, the concentration feature sample data, the relaxation feature sample data, and the grip force change feature sample data to obtain multimodal feature sample data; Determining fatigue level labels of the multimodal feature sample data through manual annotation; Inputting the multimodal feature sample data into a pre-built deep learning neural network to obtain a fatigue degree recognition result; determining a loss value according to the fatigue level identification result and the fatigue level label; The parameters of the deep learning neural network are updated according to the loss value to obtain the trained fatigue state detection model.
6. The method for detecting fatigue driving by multimodal feature fusion according to claim 1, characterized in that: The step of inputting the multimodal feature data into a pre-trained fatigue state detection model to obtain a fatigue degree detection value of the driver, and then judging whether the driver is driving fatigued according to the fatigue degree detection value, specifically includes: Inputting the multimodal feature data of a plurality of continuous time periods into the fatigue state detection model respectively, and obtaining the fatigue degree detection values and corresponding confidence levels of the plurality of continuous time periods; determining a third weight coefficient of each fatigue degree detection value according to the confidence level, and then performing weighted summation of the fatigue degree detection values of multiple consecutive time periods according to the third weight coefficient to obtain the fatigue degree of the driver; When the fatigue level is greater than or equal to a preset first threshold, determining that the driver has fatigue driving behavior; When the fatigue level is less than the first threshold, it is determined that the driver does not have fatigue driving behavior.
7. The method for detecting fatigue driving by multimodal feature fusion according to claim 6, characterized in that: The fatigue driving detection method also includes the following steps: When the driver is fatigued driving, the vehicle's fragrance system, music system, suspension system and collision warning system are adjusted according to the fatigue level, so that the fragrance system releases preset refreshing gas, the music system plays preset refreshing music, the suspension system is adjusted to the sports mode, and the sensitivity of the collision warning system is enhanced.
8. A fatigue driving detection system based on multi-modal feature fusion, characterized in that: include: A facial feature extraction module, used to obtain facial image information of a driver, and extract eye feature data and mouth feature data of the driver according to the facial image information; A brain wave feature extraction module is used to obtain the brain wave signal of the driver, and extract the concentration feature data and relaxation feature data of the driver according to the brain wave signal; A grip force feature extraction module, used to obtain a steering wheel grip force signal of a driver, and extract grip force change feature data of the driver according to the steering wheel grip force signal; A feature fusion module, used for fusing the eye feature data, the mouth feature data, the concentration feature data, the relaxation feature data and the grip force change feature data to obtain multimodal feature data; The fatigue state detection module is used to input the multimodal feature data into a pre-trained fatigue state detection model to obtain a fatigue level detection value of the driver, and then determine whether the driver is driving fatigued based on the fatigue level detection value.
9. A fatigue driving detection device based on multi-modal feature fusion, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a fatigue driving detection method using multimodal feature fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to execute a fatigue driving detection method of multimodal feature fusion as described in any one of claims 1 to 7 when executed by the processor.
Citation Information
Cited By
Cloud cabin safety officer fatigue identification early warning method and system under multi-vehicle supervision scene
CN120690004A
Fatigue state recognition method based on learnable filter bank and joint regularization
CN121265057A
Fatigue state recognition method based on learnable filter bank and joint regularization
CN121265057B