Method and system for identifying driver suitable driving credibility based on multi-task learning
By collecting multiple facial images of the driver through vehicle-mounted cameras, constructing multimodal data, and using a multi-task learning system to identify fatigue and distraction, the problem of insufficient accuracy in existing driving warning technologies has been solved, achieving real-time and accurate driving warnings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, the accuracy of detecting driver fatigue and distraction is low, and the accuracy and adaptability of multi-task learning systems are insufficient, resulting in low accuracy of driving warning projects.
Multiple facial images of the driver are collected by the vehicle-mounted camera, and multiple facial features and head shapes are identified to construct multimodal data. A multi-task learning system is used to simultaneously identify fatigue, distraction and emotional states. Based on the driver's current state and past events, a driving suitability confidence value is determined and corresponding driving warning items are triggered.
It improves the accuracy of multimodal data, enhances the accuracy and adaptability of the multi-task learning system, and achieves real-time and accurate driving warning projects.
Smart Images

Figure CN121799407A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of multi-task learning, and in particular to a method and system for identifying driver suitability credibility based on multi-task learning. Background Technology
[0002] With the development of technology, drivers operate vehicles on the road, and during long driving periods, they exhibit various facial expressions and corresponding fatigue states, as well as negative emotions such as road rage. Current technologies collect facial features of drivers and detect fatigue states based on these features, ignoring facial fatigue markers in various facial training data and distraction events of the driver. This affects the accuracy of the multi-task learning system, resulting in low accuracy of driving warning projects. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and system for identifying driver suitability credibility based on multi-task learning.
[0004] This invention provides a method for identifying driver suitability credibility based on multi-task learning, including:
[0005] The vehicle-mounted camera performs visual detection on the driver and captures multiple facial images of the driver. Based on the image recognition of these multiple facial images, multiple facial features are determined.
[0006] Multiple key facial features are determined based on multiple facial features and the driver's head shape. Based on the feature positions, corresponding feature shapes, and real-time facial images of the driver, the driver's multimodal data is determined. The multimodal data simultaneously includes the driver's precise head posture, dynamic geometric information of the eyes and mouth, as well as visual textures and facial expression details of fatigue, distraction, or specific emotions on their face.
[0007] Multiple data combinations are determined based on the identification of multimodal data of drivers. Corresponding facial training data are determined based on the identification of each data combination. A multi-task learning system is constructed based on each facial training data, the corresponding facial fatigue markers, and the distraction behavior events of drivers. The multi-task learning system accurately and synchronously identifies fatigue state, distraction state, and emotional state from multimodal data.
[0008] In the multi-task learning system, the driver's current state is determined based on the multi-task learning system, driver fatigue data, and driving behavior, and the driver's driving suitability confidence value is marked.
[0009] A preset confidence value is determined based on the driver's past driving events, the multi-task learning system, and the driver's current state. The corresponding driving warning item is triggered based on the real-time comparison between the preset confidence value and the driving suitability confidence value.
[0010] This invention provides a driver suitability credibility recognition system based on multi-task learning, which is applied to the aforementioned driver suitability credibility recognition method based on multi-task learning.
[0011] Compared with the prior art, the beneficial effects of the present invention are:
[0012] (1) The vehicle camera performs visual detection on the driver and collects multiple facial images of the driver. Multiple facial features are determined based on the image recognition of multiple facial images. Multiple key facial features are determined based on multiple facial features and the head shape of the driver. Multimodal data of the driver is determined based on the feature position, corresponding feature shape and real-time facial image of the driver. Multiple key facial features are filtered and the feature position, corresponding feature shape and real-time facial image of the driver are introduced to improve the accuracy of multimodal data.
[0013] (2) Multiple data combinations are determined based on the recognition of multimodal data of drivers. The corresponding facial training data is determined based on the recognition of each data combination. A multi-task learning system is constructed based on each facial training data, the corresponding facial fatigue markers and the distraction behavior events of drivers. Multimodal data is fully utilized. The system takes into account each facial training data, the corresponding facial fatigue markers and the distraction behavior events of drivers as a whole, which improves the accuracy and adaptability of the multi-task learning system.
[0014] (3) In the multi-task learning system, the current state of the driver is determined based on the multi-task learning system, the driver's fatigue data and driving behavior, and the driver's driving suitability credibility value is marked; the preset credibility value is determined based on the driver's past driving events, the multi-task learning system and the driver's current state, and the corresponding driving warning item is triggered based on the real-time comparison of the preset credibility value and the driving suitability credibility value. The current state of the driver is introduced and the driver's driving suitability credibility value is output, realizing the real-time comparison of the preset credibility value and the driving suitability credibility value, and ensuring the real-time nature and accuracy of the driving warning item. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the driver suitability confidence identification method based on multi-task learning in an embodiment of the present invention.
[0016] Figure 2 This is a flowchart illustrating step S11 in the driver suitability confidence identification method based on multi-task learning in this embodiment of the invention.
[0017] Figure 3 This is a flowchart illustrating step S12 in the driver suitability confidence identification method based on multi-task learning in this embodiment of the invention.
[0018] Figure 4 This is a flowchart illustrating step S13 in the driver suitability confidence identification method based on multi-task learning in this embodiment of the invention.
[0019] Figure 5 This is a flowchart illustrating step S14 in the driver suitability confidence identification method based on multi-task learning in this embodiment of the invention.
[0020] Figure 6 This is a flowchart illustrating step S15 in the driver suitability confidence identification method based on multi-task learning in this embodiment of the invention.
[0021] Figure 7 This is a schematic diagram of the structural composition of the driver suitability recognition system based on multi-task learning in an embodiment of the present invention.
[0022] Figure 8 This is a flowchart of the driver suitability confidence identification method based on multi-task learning in an embodiment of the present invention.
[0023] Figure 9 This is a multi-task schematic diagram of the driver suitability confidence identification method based on multi-task learning in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0025] Please see Figures 1 to 9 A method for identifying driver suitability credibility based on multi-task learning is proposed and applied to multi-task learning scenarios. The method includes:
[0026] Step S11: The vehicle-mounted camera performs visual detection on the driver and captures multiple facial images of the driver. Multiple facial features are determined based on the image recognition of the multiple facial images.
[0027] Step S12: Determine multiple key facial features based on multiple facial features and the driver's head shape, and determine the driver's multimodal data based on the feature positions, corresponding feature shapes, and real-time facial images of the driver.
[0028] Step S13: Based on the recognition of the driver's multimodal data, determine multiple data combinations, determine the corresponding facial training data according to the recognition of each data combination, and construct a multi-task learning system based on each facial training data, the corresponding facial fatigue markers and the driver's distraction behavior events.
[0029] Step S14: In the multi-task learning system, determine the driver's current state based on the multi-task learning system, driver fatigue data, and driving behavior, and mark the driver's driving suitability confidence value.
[0030] Step S15: Determine a preset confidence value based on the driver's past driving events, the multi-task learning system, and the driver's current state. Trigger the corresponding driving warning item based on the real-time comparison between the preset confidence value and the driving suitability confidence value.
[0031] refer to Figure 2 In step S11, the specific steps are as follows:
[0032] S111: Collect the overall space between the driver and the vehicle camera, and determine the visual detection mode of the vehicle camera based on multiple environmental parameters of the overall space, the corresponding spatial shape and the vehicle camera.
[0033] S112: Based on this visual detection mode, the vehicle camera is triggered to perform visual detection on the driver to acquire multiple facial images of the driver; among the multiple facial images, the corresponding feature regions are determined based on the detection of each facial image, and the corresponding facial features are determined based on the recognition of each feature region to determine multiple facial features.
[0034] In the embodiments of this application, the overall space between the driver and the vehicle-mounted camera is collected, and the visual detection mode of the vehicle-mounted camera is determined based on multiple environmental parameters of the overall space, the corresponding spatial shape, and the vehicle-mounted camera. This approach takes into account multiple environmental parameters of the overall space, the corresponding spatial shape, and the overall consideration of the vehicle-mounted camera, ensuring the accuracy of the visual detection mode of the vehicle-mounted camera.
[0035] At this point, by using the vehicle's light sensor or analyzing the image brightness histogram, the system can accurately identify whether the current light intensity is direct sunlight, normal lighting, low light, or extremely dark. Simultaneously, by analyzing the image's white balance, it determines the warm or cool color temperature of the light source, which is crucial for ensuring accurate skin tone recognition. The system also assesses the dynamic range of the image, paying particular attention to high-contrast scenes such as entering or exiting tunnels, which are prone to localized overexposure or underexposure. In situations with insufficient visible light, the system can seamlessly switch to infrared imaging mode, actively emitting infrared light and capturing reflected signals to obtain clear images.
[0036] The system uses a deep learning model to locate the head region in a two-dimensional image, and then combines this with the preset physical dimensions of the head within the camera to estimate the precise coordinates of the head in three-dimensional space. It uses 3D head key point detection technology to obtain the spatial positions of dozens of feature points such as the tip of the nose, the corners of the eyes, and the corners of the mouth, and then calculates the yaw angle, pitch angle, and roll angle that represent the head posture. Based on this three-dimensional coordinate information, the system can accurately calculate the distance from the driver's face to the camera, and assess the visibility of facial key points in real time, and promptly mark any possible occlusion areas.
[0037] Based on the aforementioned environmental parameters and spatial information, the system dynamically generates the optimal camera control strategy. In terms of exposure control, the system intelligently selects different solutions such as local exposure control, HDR synthesis, or infrared mode according to ambient lighting conditions. The focusing system drives the lens motor for precise focusing based on distance estimation results, and uses digital zoom when necessary to ensure that the facial area occupies sufficient pixel resolution. The color correction module adjusts the white balance according to color temperature parameters to ensure consistent skin tone performance under different light sources. In addition, the system can also dynamically adjust the frame rate and resolution according to the driver's posture changes, achieving the best balance between capturing dynamic details and saving computing power.
[0038] Furthermore, based on this visual detection mode, the vehicle camera is triggered to perform visual detection on the driver to collect multiple facial images of the driver. Among the multiple facial images, the corresponding feature regions are determined based on the detection of each facial image, and the corresponding facial features are determined based on the recognition of each feature region, so as to determine multiple facial features. This overall consideration of recognizing each feature region ensures the accuracy of the corresponding facial features.
[0039] At this point, the system writes the parameters such as exposure time, gain, white balance, and frame rate determined by S111 into the hardware register in real time through the camera driver interface, completing the precise configuration of the camera; the camera continuously outputs image data streams at the set frame rate, and these data are sent to a circular buffer for caching to ensure that the processor always processes the latest frame and avoids delays; each frame of image is precisely timestamped by hardware during acquisition, providing a reliable basis for the timing of subsequent analysis.
[0040] The system employs a lightweight cascaded classifier or CNN model as the first-level detector to quickly scan the entire image and output the boundary coordinates of face candidate boxes. For multiple overlapping candidate boxes that may appear, the system uses a non-maximum suppression algorithm to merge overlapping boxes based on confidence scores, retaining only the optimal face bounding box. To improve efficiency and stability, the system also integrates an object tracking algorithm. After the detector successfully locates a face, the tracker predicts the face position in subsequent frames, and the detector only needs to verify within a small range, which greatly reduces the amount of computation and can maintain position estimation even when the face is briefly occluded.
[0041] The system inputs the cropped facial region image into a high-precision facial key point detection model to obtain a precise coordinate set containing dozens or even hundreds of key points. Based on these coordinates, the system calculates a series of geometric indicators with clear physical meaning in real time, such as the eyelid aspect ratio, mouth opening ratio, and head pose angle. In order to eliminate the influence of differences in head distance and angle, all geometric features are normalized, for example, by dividing the distance between the eyes, to provide a stable and standardized input for subsequent models.
[0042] refer to Figure 3 In step S12, the specific steps are as follows:
[0043] S121: Mark the position of the driver's head. The vehicle camera dynamically captures the position of the driver's head and determines the shape of the driver's head. Based on multiple facial features and the shape of the driver's head, multiple key facial features are determined and the feature shape of each key facial feature is marked.
[0044] S122: Collect real-time facial images of the driver, determine a first set of data based on the real-time facial images of the driver and the feature positions of multiple key facial features, and determine the multimodal data of the driver based on the first set of data and the feature morphology of multiple key facial features.
[0045] In the embodiments of this application, the position of the driver's head is marked, the vehicle camera dynamically captures the position of the driver's head and determines the shape of the driver's head, and multiple key facial features are determined based on multiple facial features and the shape of the driver's head, and the feature shape of each key facial feature is marked. This takes into account the overall consideration of multiple facial features and the shape of the driver's head, and ensures the accuracy of multiple key facial features.
[0046] At this point, advanced tracking algorithms such as Kalman filters or correlation filters are used. After the head is detected for the first time in S112, the system will predict its position and size in subsequent frames. The detector only needs to verify within a small range, which greatly improves efficiency and stability. The Kalman filter can also predict the position of the next frame based on the motion model, effectively dealing with the brief tracking loss caused by vehicle bumps or rapid head turns, smoothing the head movement trajectory, and filtering out high-frequency jitter noise. In multi-person scenarios, the system uses head appearance features to associate identities, ensuring that the driver is always locked.
[0047] The system uses tools such as MediaPipeFaceMesh to obtain the 2D coordinates of hundreds of facial key points. Combined with camera intrinsic parameters, the system uses a perspective n-point algorithm to solve for the coordinates of these points in 3D space. The system adopts a general 3D head mesh model and adjusts its shape and expression parameters through optimization algorithms to make the projection of the 3D key points of the model coincide as much as possible with the 2D key points detected in the image. After successful fitting, the system obtains a dynamic 3D head model specific to the driver and parameterizes its posture as Euler angles and its expression changes as expression coefficients.
[0048] The system predefines a set of geometric features highly related to driving safety, such as eyelid aspect ratio, gaze direction, mouth opening ratio, and mouth corner curvature, and performs real-time calculations based on 3D key points. More importantly, the system analyzes the statistics and trends of each feature through a short-term sliding window and performs dynamic morphological labeling: a persistently low EAR value is labeled as "fatigue squinting", a sudden increase in MAR variance is labeled as "yawning event", a periodically oscillating pitch angle is labeled as "nodding drowsiness", and a gaze direction that is continuously deviated from the front is labeled as "distracted staring".
[0049] Specifically, as the driver began to exhibit moderate fatigue on the monotonous highway, the system used a Kalman filter to stably lock onto the driver's head in consecutive video frames, maintaining smooth tracking even with slight vehicle vibrations. The system acquired 468 3D key points of the driver's face in real time and fitted them to a 3D head model. At this moment, the model showed a pitch angle of -5 degrees, a yaw angle of +30 degrees, and a slight furrowing of the eyebrows. Based on these 3D key points, the system calculated that the driver's EAR value was below the normal eye-opening threshold, and their gaze was directed to the right front. By analyzing data from the past two seconds, the system found that the EAR value was... The low value and small variance were marked as "persistent ptosis"; at the same time, the yaw angle was detected to change rapidly within 1 second and then quickly return to its original position, which was marked as "rapid head turn distraction event"; the output of step S121 is no longer an isolated number, but a set of structured information with semantic labels: {state: fatigue with distraction, feature: [persistent ptosis, rapid head turn distraction event], geometric value: [EAR=0.285, Yaw_max=30°]}, which transforms the driver's raw visual information into a head state report that the machine can directly understand and make decisions, laying a solid foundation for subsequent intelligent warnings.
[0050] Furthermore, real-time facial images of the driver are collected. Based on the real-time facial images of the driver and the feature positions of multiple key facial features, a first-level dataset is determined. Based on the first-level dataset and the feature morphology of multiple key facial features, multimodal data of the driver is determined. This approach takes into account the overall consideration of the first-level dataset and the feature morphology of multiple key facial features, ensuring the accuracy of the multimodal data of the driver. At the same time, multiple key facial features are filtered, and the feature positions, corresponding feature morphologies, and real-time facial images of multiple key facial features are introduced, which improves the accuracy of the multimodal data.
[0051] At this point, the system converts the pixel coordinates of each key facial feature marked in S121 into standardized three-dimensional spatial coordinates or normalized two-dimensional coordinates based on camera intrinsic parameters and head distance to eliminate scale differences. Using the current frame as a reference, the system constructs a sliding time window, which includes not only the instantaneous value of each key feature, but also calculates and integrates statistics such as mean, variance, first / second derivative, and specific event counts within the time window. All feature positions, temporal statistics, and dynamic labels at the current moment are concatenated into a single high-dimensional vector to form a precise geometric description of the driver's physiological and behavioral state.
[0052] The system inputs the real-time facial images acquired in S112 into a pre-trained convolutional neural network, extracts the output of the intermediate layer to obtain a high-dimensional deep image feature vector, which encodes visual information such as texture, skin color, lighting, and subtle changes in facial muscles in the image. Then, the system uses strategies such as simple stitching, attention-weighted fusion, or two-stream network fusion to effectively integrate the first set of data with the deep image feature vector. After fusion, a highly information-dense multimodal data is finally generated. The multimodal data simultaneously contains the driver's precise head posture, dynamic geometric information of the eyes and mouth, as well as visual texture and facial expression details of fatigue, distraction, or specific emotions on their face.
[0053] Specifically, in the first data set construction phase, the system normalizes all the driver's left eye EAR value (0.15), right eye EAR value (0.16), MAR value (0.65, indicating yawning), and head pitch angle (-10 degrees, indicating nodding) at time T. By analyzing the data from the past 2 seconds, the system calculates that the mean EAR value is low and the variance is large, the MAR value has two peaks exceeding 0.5, and the head pitch angle has three downward movements. These values are encapsulated into a 128-dimensional geometric feature vector. In the multimodal data fusion phase, the system inputs the driver's facial image at time T, showing "mouth wide open and eyes squinting," into the MobileNetV2 network and extracts a 256-dimensional depth image feature vector. This vector encodes visual details such as the inside of the mouth exposed due to yawning, the dull skin tone due to fatigue, and the wrinkles at the corners of the eyes caused by squinting.
[0054] The system employs an attention-weighted fusion mechanism, using "yawning events" and "low EAR" labels in the geometric feature vector as queries. It identifies highly relevant "oral texture" and "eye wrinkles" features in the depth image feature vector and assigns them extremely high attention weights. The system generates a multimodal data tensor that integrates precise geometric dynamics and rich visual context. This tensor is passed to the subsequent multi-task learning model, enabling it to simultaneously "understand" the driver's nod amplitude, the quantification of yawning, and the visual expression of fatigue on his face, thus making a more accurate and robust assessment of fatigue status than a single information source.
[0055] refer to Figure 4 In step S13, the specific steps are as follows:
[0056] S131: Real-time monitoring of the driver's multimodal data, determining multiple data combinations based on the identification of the driver's multimodal data, the multiple data combinations covering the eyelid aspect ratio, mouth opening ratio, and eyebrow-eye distance; determining corresponding facial training data based on the identification of each data combination, thereby determining multiple facial training data;
[0057] S132: Match corresponding facial fatigue markers based on the recognition of each facial training data. At the same time, collect distraction behavior events of the driver. Construct a multi-task learning framework based on the driver's distraction behavior events and each facial training data. Construct a multi-task learning system based on the multi-task learning framework, the facial fatigue markers of each facial training data, and the driver's expression data.
[0058] In the embodiments of this application, multimodal data of the driver is monitored in real time, and multiple data combinations are determined based on the identification of the driver's multimodal data. The multiple data combinations cover the aspect ratio of the eyelids, the opening and closing ratio of the mouth, and the distance between the eyebrows and eyes. Based on the identification of each data combination, corresponding facial training data is determined to determine multiple facial training data, which takes into account the overall consideration of the identification of each data combination and ensures the accuracy of the corresponding facial training data.
[0059] At this point, the system maintains a circular buffer to cache the multimodal data tensor sequence output by S12 in real time, with each tensor having a precise timestamp. Although the data is fused, the system can reverse-track and recalculate the core physiological indicators that make it up by recording its construction process. These indicators mainly include the eyelid aspect ratio, mouth opening ratio, eyebrow-eye distance, and head posture angle sequence. For each time point, the system combines these deconstructed indicators into a structured data packet containing a snapshot of the driver's key physiological state at that moment.
[0060] The system defines a fixed-length sliding time window. For the current moment, the system extracts all data combination sequences within the time window. This data combination sequence within the time window, along with the original face image corresponding to the current moment, is encapsulated into a face training data sample. A complete sample contains image data, a geometric temporal data matrix, and metadata such as timestamps and lighting conditions. To improve model robustness, all numerical data in the samples are Z-score normalized. The system also adds slight temporal jitter or noise to the temporal data to generate more training samples.
[0061] Specifically, when drivers begin to show signs of fatigue on the highway, the system receives multimodal data tensors about the driver in real time. At time T, through reverse tracing, the current physiological indicators are deconstructed: EAR value 0.18 (squinting), MAR value 0.12 (normal), eyebrow-eye distance 95% of the standard value (slight frowning), and head pitch angle -5 degrees (slight nodding). These values are combined into a data set at time T. The system then initiates a 3-second sliding time window with time T as the endpoint, extracting all data sets from T-3 seconds to T. Analysis reveals that... The EAR value gradually decreased from 0.35 to 0.18 within 3 seconds; the MAR peaked at 0.6 at T-1.5 seconds (yawning); the head pitch angle dropped more than -5 degrees twice; the system encapsulated the time series data matrix of these 3 seconds, the facial image at time T, and all metadata into a complete facial training data sample and stored it in the training database; this sample not only recorded the driver's squinting state "at this moment", but also recorded the complete dynamic process of how he gradually developed from a conscious state to a fatigued state, as well as the accompanying yawning and nodding behaviors.
[0062] Furthermore, based on the recognition of each facial training data, corresponding facial fatigue markers are matched. Simultaneously, distraction events of the driver are collected. A multi-task learning framework is constructed based on the driver's distraction events and each facial training data. A multi-task learning system is constructed based on the multi-task learning framework, the facial fatigue markers of each facial training data, and the driver's facial expression data. This comprehensive consideration of the multi-task learning framework, the facial fatigue markers of each facial training data, and the driver's facial expression data ensures the accuracy of the multi-task learning system. At the same time, it makes full use of multimodal data, comprehensively considering each facial training data, the corresponding facial fatigue markers, and the driver's distraction events, thereby improving the accuracy and adaptability of the multi-task learning system.
[0063] At this point, for fatigue labeling, the system adopts a multi-level labeling strategy, including a rule engine based on physiological indicators, such as calculating whether PERCLOS (the percentage of time the eyes are closed per unit of time) exceeds a threshold; event detection based on behavioral patterns, identifying yawn waveforms by analyzing MAR sequences; and dynamic evaluation based on time series, analyzing the downward trend and fluctuation frequency of EAR to label fatigue trends; for distraction behaviors, the system analyzes the time series data of head posture angle and gaze direction, uses template matching to identify specific behaviors such as "looking left and right" and "looking down at a mobile phone", or uses threshold triggers to label long-term gaze behavior that deviates from the front.
[0064] A neural network architecture capable of efficiently processing multimodal inputs was designed. The framework comprises two parallel input branches: the image branch extracts convolutional features from face images using an improved MobileNetV2; the geometry branch receives metrics such as EAR, MAR, and head pose angle, and encodes them using a multilayer perceptron. The key innovation of the framework lies in the cross-attention fusion module, which uses geometric features as queries and actively focuses on and enhances the visual parts of the image features most relevant to the current geometric state through a multi-head attention mechanism, achieving precise information complementarity and alignment. The fused features are fed into three parallel fully connected layer branches, corresponding to fatigue classification, distraction classification, and emotion classification tasks, respectively.
[0065] The system defines a joint loss function as a weighted sum of the losses for each task. During training, all samples with fatigue markers, distraction event labels, and facial expression data labels are input into the framework in batches. The gradient of the total loss with respect to all network parameters is calculated using the backpropagation algorithm, and the optimizer is used to update the parameters. Due to the shared backbone and fusion layer, the model is forced to learn a general feature representation that is useful for all tasks. This collaborative learning mechanism greatly improves the model's generalization ability and overall recognition accuracy. After training, the entire network structure and parameters constitute a complete multi-task learning system.
[0066] Specifically, after receiving a sample of facial training data about a driver, the system performs multi-source annotation: analysis of its time-series data reveals a PERCLOS of 45% and detects a yawning waveform, thus assigning a strong "fatigue" label; simultaneously, its head posture highly matches the template of "quickly turning the head to check the rearview mirror," so it is additionally labeled with an event tag of "distraction - turning the head to check"; in the multi-task learning framework, the driver's facial image and geometric indicators are input into two branches respectively, and the cross-attention fusion module successfully activates the "tense" feature in the image by activating the geometric features of "yawning." The system identifies regions associated with "large mouth" and "tired face." During training, this double-labeled sample is used in the training process, and the outputs of its fatigue and distraction classification heads do not match the labels, resulting in significant losses. The system calculates the total loss and backpropagates it to update the parameters of the entire network. By processing thousands of similar samples from drivers or other drivers, this multi-task learning system demonstrates the ability to accurately and synchronously identify fatigue, distraction, and emotional states from multimodal data. The multi-task learning system serves as a powerful, robust, and deployable intelligent early warning model.
[0067] refer to Figure 5 In step S14, the specific steps are as follows:
[0068] S141: Real-time monitoring of the multi-task learning system, determining multiple task items based on the recognition of the multi-task learning system, determining the first-level state coefficient based on the multiple task items and the driver's fatigue data, determining the second-level state coefficient based on the multiple task items and the driver's fatigue behavior, and determining the driver's current state based on the mapping relationship between the first-level state coefficient, the second-level state coefficient and the current state.
[0069] S142: Based on the driver's current state, match the corresponding distraction state, fatigue state, and emotional state, and mark the weights of the distraction state, fatigue state, and emotional state. Perform weighted processing on the weights of the distraction state, fatigue state, and emotional state to determine the driver's driving suitability confidence value.
[0070] In the embodiments of this application, a multi-task learning system is monitored in real time. Multiple task items are identified based on the recognition of the multi-task learning system. A first-level state coefficient is determined based on the multiple task items and the driver's fatigue data. A second-level state coefficient is determined based on the multiple task items and the driver's fatigue behavior. The driver's current state is determined based on the mapping relationship between the first-level state coefficient, the second-level state coefficient and the current state. This approach takes into account the overall consideration of the mapping relationship between the first-level state coefficient, the second-level state coefficient and the current state, ensuring the accuracy of the driver's current state.
[0071] At this point, the system continuously inputs the multimodal data tensor generated in real time in step S12 into the pre-constructed multi-task learning system at a high frequency; the data tensor passes through the shared backbone network and the cross-attention fusion module in sequence, and finally reaches three parallel task classification heads; each classification head outputs a set of probability values through its internal activation function, and the system identifies and separates these raw outputs to form a set of task items including fatigue probability, distraction probability and emotion probability.
[0072] Based on the driver's macro-level fatigue data and micro-level fatigue behavior, the probability of fatigue tasks is dynamically adjusted. The first level of state coefficient, or static weight, is calculated using piecewise linear functions or fuzzy logic rules based on macro-level contextual information such as continuous driving duration, current time period, and frequency of historical fatigue events. It represents the baseline fatigue risk caused by factors such as long-term driving. The second level of state coefficient, or dynamic weight, is calculated using micro-level physiological indicators such as real-time PERCLOS values, yawning frequency, and nodding frequency through event triggering or pattern matching mechanisms. It usually has time decay characteristics and represents the instantaneous risk increment caused by acute fatigue behavior.
[0073] The system applies a two-layer weighted calculation to calculate the probability of fatigue tasks, while applying a single weight to distraction and emotional tasks. The system combines the adjusted probability values of each task into a state vector, which is a quantitative representation of the driver's current state. It not only includes the model's identification results for each state, but also incorporates the contextual dynamic assessment of fatigue risk, making the state description more accurate and comprehensive.
[0074] Specifically, the system inputs the driver's current multimodal data into a multi-task learning system, and the model outputs the original task items: fatigue probability of 0.65, distraction probability of 0.1, and emotional probability of 0.05. In the two-layer state coefficient dynamic evaluation stage, the system evaluates the driver's macro-fatigue data: driving continuously for 3.5 hours and in a period of physiological fatigue, calculating the first-level state coefficient as 1.5. Simultaneously, the system detects that the driver completed a yawn within 10 seconds, triggering the second-level state coefficient to increase by 0.2, becoming 1.2. In the current state vector mapping stage, the system... Weighted calculation is performed: the adjusted fatigue probability is 0.65 multiplied by 1.5 and then by 1.2, reaching 1.17, while the probabilities of distraction and emotion remain unchanged; the system generates the driver's current state vector as [1.17, 0.1, 0.05]; through step S141, the system combines the original fatigue probability of 0.65 with the context of the driver "driving for a long time" and "just yawned", and raises it to an adjusted value of risk level 1.17. This vector will be passed to subsequent steps to calculate the final driving suitability confidence, making the decision basis more sufficient and reliable.
[0075] Furthermore, based on the driver's current state, corresponding distraction state, fatigue state, and emotional state are matched, and the weights of distraction state, fatigue state, and emotional state are marked. The weights of distraction state, fatigue state, and emotional state are weighted to determine the driver's driving suitability confidence value, thus introducing the driver's driving suitability confidence value.
[0076] At this point, the system maintains a driving state scenario library trained from expert knowledge and massive amounts of data. Each scenario represents a typical combination of driving risks, such as "severe fatigue," "fatigue accompanied by distraction," "road rage," or "mild distraction." The system uses a vector similarity algorithm to compare the current state vector input by S141 with the typical vectors of all scenarios in the scenario library, selecting the scenario with the highest similarity as the driver's current driving state. Once a scenario is determined, the system retrieves a pre-set weighting scheme for that scenario. This scheme precisely defines the relative importance of fatigue, distraction, and emotions to overall driving safety under that specific risk combination. For example, in the "fatigue accompanied by distraction" scenario, the weight of fatigue may be assigned a higher value because the superposition of the two risks will produce a synergistic effect.
[0077] Based on the matched scenario, the system assigns corresponding dynamic weight values to the three tasks of fatigue, distraction, and emotion. The sum of these weight values is 1. The system performs a weighted summation operation, with the input value being the probability value adjusted by the double coefficient in S141 or the original model output probability. The system subtracts the risk value obtained by the weighted summation from 1, multiplies it by 100%, and finally outputs a driving suitability confidence value between 0% and 100%. This value is a highly condensed and quantitative risk assessment indicator.
[0078] Specifically, the system receives the driver's current state vector output by S141 and compares it with the scene database. It finds that the vector has the highest similarity to the typical vector of the "fatigue accompanied by distraction" scene. The system then determines that the driver is in this scene and retrieves the corresponding weight allocation scheme: fatigue state weight is 0.6, distraction state weight is 0.3, and emotional state weight is 0.1. In the multi-state weighted fusion stage, the system uses the original model output probability for calculation: the comprehensive risk value equals 0.6 multiplied by the fatigue probability of 0.65, plus 0.3. Multiplying by the distraction probability of 0.4, and adding 0.1 multiplied by the emotional probability of 0.05, the final result is 0.515; the system calculates the driving suitability confidence value as (1-0.515) multiplied by 100%, resulting in 48.5%; through step S142, the system finally quantifies the comprehensive risk state of the driver "fatigue due to long-term driving and distraction due to operation of the central control" into a single confidence value of 48.5%; this value will be passed to subsequent steps for comparison with the preset warning threshold, and once it falls below the threshold, the system will immediately trigger the corresponding warning measures.
[0079] refer to Figure 6 In step S15, the specific steps are as follows:
[0080] S151: Based on the driver's tracing, determine the corresponding past driving events, determine multiple driving fatigue factors based on the identification of past driving events, and determine the preset confidence value based on multiple driving fatigue factors, multi-task learning system and the driver's current state.
[0081] S152: Collect the driver's driving suitability confidence value, compare the driver's driving suitability confidence value with the preset confidence value, and output the corresponding comparison result; determine the corresponding driving warning item based on the mapping relationship between the comparison result and the driving warning item, so as to trigger the driving warning item.
[0082] In the embodiments of this application, the corresponding past driving events are determined based on the tracing of the driver, multiple driving fatigue factors are determined based on the identification of past driving events, and a preset confidence value is determined based on multiple driving fatigue factors, a multi-task learning system, and the current state of the driver. This approach takes into account the overall consideration of multiple driving fatigue factors, a multi-task learning system, and the current state of the driver, ensuring the accuracy of the preset confidence value.
[0083] At this time, the system accesses and integrates multi-source historical data through the vehicle bus and its own data storage unit, including DMS historical data, vehicle dynamics data, driving trip data and human-machine interaction data; the system uses time series analysis and clustering algorithms to process the raw data and aggregate meaningful data points into meaningful past driving events, such as aggregating consecutive low PERCLOS values and frequent yawning events into a "fatigue driving event".
[0084] The system transforms the extracted historical events into calculable quantitative indicators that describe the individual characteristics of drivers. These indicators include: a fatigue tolerance factor, which measures a driver's ability to resist fatigue and is calculated by statistically analyzing the deviation between the average first fatigue warning time and the baseline time; a distraction tendency factor, which measures a driver's susceptibility to distraction and is derived by calculating the average incidence of distracting events per unit time; a time sensitivity factor, which measures a driver's physiological fatigue sensitivity during a specific time period and is derived by analyzing the temporal distribution of historical fatigue events; and a recovery ability factor, which measures a driver's efficiency in recovering their state after a short rest.
[0085] By combining quantified personal factors with real-time status, a dynamic model is used to calculate the most suitable warning threshold for the current moment. The system employs a weighted linear model or a rule-based fuzzy logic engine. The core idea is that the base threshold is primarily modified by personal fatigue factors and then fine-tuned by the current state. The calculation formula is usually the base threshold minus or plus the product of each factor and its weight, and also includes a fine-tuning term based on the current state vector. If the current state indicates that the fatigue risk is rising rapidly, the system can proactively lower the threshold to achieve a more sensitive warning.
[0086] Specifically, during the driving history event tracing phase, the system discovered that drivers triggered fatigue warnings for the first time on average after driving continuously for 4.5 hours. Distraction events were less frequent than average, with fatigue events mostly occurring in the early morning, and recovery was rapid after rest. In the driving fatigue factor quantification phase, the system calculated the driver's fatigue tolerance factor to be +0.12, distraction tendency factor to be -0.05, afternoon sensitivity factor to be -0.03, and recovery ability factor to be +0.08. In the preset confidence value dynamic calculation phase, the system substituted these factors into the model, assuming a baseline threshold of 50%, and calculated a preliminary threshold of 38.5% through weighted calculation. Simultaneously, the system detected that the driver's current state indicated a rapidly increasing fatigue risk, therefore adding a -2% fine-tuning item. The system determined a personalized preset threshold of 36.5% for the driver.
[0087] Furthermore, the driver's driving suitability confidence value is collected, compared with a preset confidence value, and the corresponding comparison result is output. Based on the mapping relationship between the comparison result and the driving warning item, the corresponding driving warning item is determined to trigger the driving warning item. This overall consideration of the mapping relationship between the comparison result and the driving warning item ensures the accuracy of the corresponding driving warning item. At the same time, the driver's current state is introduced, and the driver's driving suitability confidence value is output, realizing real-time comparison between the preset confidence value and the driving suitability confidence value, ensuring the real-time nature and accuracy of the driving warning item.
[0088] At this point, the system obtains the driver's driving competence confidence value from module S142 at a fixed frequency, and obtains the personalized preset threshold for the current moment from module S151. The system executes a hierarchical conditional judgment logic to compare the driving competence confidence value with the threshold. This comparison logic is usually divided into multiple risk levels, such as no risk, level 1 risk (minor), level 2 risk (moderate), and level 3 risk (high), with each level corresponding to a specific numerical range. Based on the comparison result, the system outputs a discrete decision status code, for example, 0 represents normal and 1 represents triggering a level 1 warning. This decision status code becomes the trigger for all subsequent actions.
[0089] The system internally stores a predefined driving warning mapping table, which establishes a one-to-one correspondence between decision status codes and specific sets of driving warning items. This mapping table is designed according to human factors engineering principles, ensuring that the intensity of the warning matches the risk level, avoiding excessive alarm or insufficient warning. Once the mapping relationship is determined, the system sends instructions to the corresponding actuators via the vehicle bus or human-machine interface, driving the vision module to illuminate specific icons, the audio module to play voice prompts or warning sounds, and the tactile module to trigger steering wheel or seat vibrations. In extreme cases, it may even link with the vehicle control module. The driving warning mapping table is shown in Table 1.
[0090] Table 1 Driving Warning Mapping Table
[0091]
[0092] Specifically, in the real-time comparison and decision logic output stage, the system calculates a personalized threshold of 36.5% for the driver; at time T1, the driver's driving suitability confidence value is 42%, which is higher than the threshold, and the system outputs decision status code 0, determining a risk-free state; at time T2, the driver's state deteriorates, and the driving suitability confidence value becomes 39%, falling into the slight range of Level 1 risk, and the system outputs decision status code 1; at time T3, the driver's state further deteriorates, and the driving suitability confidence value drops to 32%, entering the moderate range of Level 2 risk, and the system outputs decision status code 2; in the early warning item mapping and graded intervention triggering stage.
[0093] The system performs corresponding operations based on the status code: at time T1, status code 0 is mapped to "no operation", and the system remains silent; at time T2, status code 1 is mapped to the dashboard icon lighting up, with a yellow coffee cup icon being lit up, giving the driver a gentle visual cues; at time T3, status code 2 is mapped to a combination of icon lighting up and voice prompts, whereby the system keeps the icon lit up while issuing a clear voice warning through the car audio system, advising the driver to rest as soon as possible.
[0094] refer to Figure 8 and Figure 9 In another embodiment of this application, the driver is captured by the vehicle-mounted camera within a 5-second time window. Each image frame is input into a multi-task model. This multi-task model is built on a deep neural network structure and achieves simultaneous recognition of multiple driver states through joint training. Specifically, the model outputs three types of state judgment results simultaneously, corresponding to the driver's fatigue state, distraction state, and emotional state, respectively. Each output is a binary classification probability result calculated by the model, used to represent the probability of the state occurring.
[0095] The output values are between 0% and 100%, with higher values indicating a higher probability of the state occurring. Furthermore, in this embodiment, to comprehensively assess the driver's overall state, the system performs weighted fusion processing on the three state results output by the multi-task model. Specifically, the probability results corresponding to fatigue, distraction, and emotional states are multiplied by preset weight parameters, and the weighted results are summed to obtain the driver's credibility. Within the 5-second time window, statistical analysis is performed on the driver's credibility corresponding to each frame of the image.
[0096] The system calculates the proportion of frames with a driving suitability confidence level below the preset driving suitability threshold within the time window. When this proportion exceeds the set threshold K, the system determines that the driver is in an unsuitable driving state and triggers the voice warning module to output a prompt message to remind the driver to rest or pay attention to driving safety.
[0097] Meanwhile, this embodiment proposes a deep learning-based multi-task driver state recognition model to simultaneously detect and recognize three states of driver fatigue, distraction, and emotion. The model adopts a lightweight network structure, combines key features of the driver's head and face with convolutional neural network features, and achieves feature fusion through a cross-attention mechanism, thereby improving the discrimination ability and generalization performance of multi-task learning.
[0098] Please see Figure 7 , Figure 7 This is a schematic diagram of the structural composition of a driver suitability confidence assessment system based on multi-task learning according to an embodiment of the present invention; the driver suitability confidence assessment system based on multi-task learning includes:
[0099] The facial feature module 21 is used by the vehicle-mounted camera to perform visual detection of the driver and collect multiple facial images of the driver, and determine multiple facial features based on the image recognition of the multiple facial images.
[0100] The multimodal data module 22 is used to determine multiple key facial features based on multiple facial features and the head shape of the driver, and to determine the driver's multimodal data based on the feature position, corresponding feature shape and real-time facial image of the driver.
[0101] The multi-task learning system module 23 is used to determine multiple data combinations based on the recognition of multimodal data of drivers, determine the corresponding facial training data based on the recognition of each data combination, and construct a multi-task learning system based on each facial training data, the corresponding facial fatigue markers and distraction behavior events of drivers.
[0102] The driving suitability confidence value module 24 is used to determine the current state of the driver in the multi-task learning system based on the multi-task learning system, the driver's fatigue data and driving behavior, and to mark the driver's driving suitability confidence value.
[0103] The driving warning module 25 is used to determine a preset confidence value based on the driver's past driving events, multi-task learning system and the driver's current state, and to trigger the corresponding driving warning project based on the real-time comparison between the preset confidence value and the driving suitability confidence value.
[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A method for identifying driver suitability credibility based on multi-task learning, characterized in that, include: The vehicle-mounted camera performs visual detection on the driver and captures multiple facial images of the driver. Based on the image recognition of these multiple facial images, multiple facial features are determined. Multiple key facial features are determined based on multiple facial features and the driver's head shape. Based on the feature positions, corresponding feature shapes, and real-time facial images of the driver, the driver's multimodal data is determined. The multimodal data simultaneously includes the driver's precise head posture, dynamic geometric information of the eyes and mouth, as well as visual textures and facial expression details of fatigue, distraction, or specific emotions on their face. Multiple data combinations are determined based on the recognition of multimodal data of drivers. The corresponding facial training data is determined based on the recognition of each data combination. A multi-task learning system is constructed based on each facial training data, the corresponding facial fatigue markers, and the distraction behavior events of drivers. The multi-task learning system demonstrates the ability to accurately and synchronously identify fatigue, distraction, and emotional states from multimodal data. In the multi-task learning system, the driver's current state is determined based on the multi-task learning system, driver fatigue data, and driving behavior, and the driver's driving suitability confidence value is marked. A preset confidence value is determined based on the driver's past driving events, the multi-task learning system, and the driver's current state. The corresponding driving warning item is triggered based on the real-time comparison between the preset confidence value and the driving suitability confidence value.
2. The method for identifying driver suitability credibility based on multi-task learning according to claim 1, characterized in that, The vehicle-mounted camera performs visual detection on the driver and captures multiple facial images of the driver. Based on image recognition of these multiple facial images, it determines multiple facial features, including: Collect data on the overall space between the driver and the vehicle-mounted camera, and determine the visual detection mode of the vehicle-mounted camera based on multiple environmental parameters of the overall space, the corresponding spatial shape, and the vehicle-mounted camera. Based on this visual detection mode, the vehicle camera is triggered to visually detect the driver, so as to collect multiple facial images of the driver. In the multiple facial images, the corresponding feature regions are determined based on the detection of each facial image, and the corresponding facial features are determined based on the recognition of each feature region, so as to determine multiple facial features.
3. The method for identifying driver suitability credibility based on multi-task learning according to claim 1, characterized in that, The process involves determining multiple key facial features based on multiple facial features and the driver's head shape, and determining the driver's multimodal data based on the feature positions, corresponding feature shapes, and real-time facial images of the driver, including: The vehicle camera dynamically captures the driver's head position and determines the head shape. Based on multiple facial features and the driver's head shape, multiple key facial features are determined and the feature shape of each key facial feature is marked. Real-time facial images of drivers are collected. A first-level data set is determined based on the real-time facial images of drivers and the feature positions of multiple key facial features. Multimodal data of drivers is determined based on the first-level data set and the feature morphology of multiple key facial features.
4. The method for identifying driver suitability credibility based on multi-task learning according to claim 1, characterized in that, The process involves identifying multiple data combinations based on the driver's multimodal data, determining corresponding facial training data based on the identification of each data combination, and constructing a multi-task learning system based on each facial training data, corresponding facial fatigue markers, and the driver's distraction events. This system includes: Real-time monitoring of multimodal data of drivers; identification of multimodal data of drivers to determine multiple data combinations, which include eyelid aspect ratio, mouth opening ratio and eyebrow-eye distance; identification of each data combination to determine corresponding facial training data, thereby determining multiple facial training data.
5. The method for identifying driver suitability credibility based on multi-task learning according to claim 4, characterized in that, The method involves identifying multiple data combinations based on the driver's multimodal data, determining corresponding facial training data based on the identification of each data combination, and constructing a multi-task learning system based on each facial training data, corresponding facial fatigue markers, and the driver's distraction behavior events. It also includes: Based on the recognition of each face training data, corresponding facial fatigue labels are matched. At the same time, distraction behavior events of drivers are collected. A multi-task learning framework is constructed based on the distraction behavior events of drivers and each face training data. A multi-task learning system is constructed based on the multi-task learning framework, the facial fatigue labels of each face training data and the driver's expression data.
6. The method for identifying driver suitability credibility based on multi-task learning according to claim 1, characterized in that, In the multi-task learning system, the driver's current state is determined based on the multi-task learning system, driver fatigue data, and driving behavior, and the driver's driving suitability confidence value is marked, including: The system monitors a multi-task learning system in real time, identifies multiple task items based on the recognition of the multi-task learning system, determines the first-level state coefficient based on the multiple task items and the driver's fatigue data, determines the second-level state coefficient based on the multiple task items and the driver's fatigue behavior, and determines the driver's current state based on the mapping relationship between the first-level state coefficient, the second-level state coefficient and the current state.
7. The method for identifying driver suitability credibility based on multi-task learning according to claim 6, characterized in that, In the multi-task learning system, determining the driver's current state based on the multi-task learning system, driver fatigue data, and driving behavior, and marking the driver's driving suitability confidence value, also includes: Based on the driver's current state, the system matches the corresponding distraction state, fatigue state, and emotional state, and labels the weights of the distraction state, fatigue state, and emotional state. The weights of the distraction state, fatigue state, and emotional state are then weighted to determine the driver's driving suitability confidence value.
8. The method for identifying driver suitability credibility based on multi-task learning according to claim 1, characterized in that, The process involves determining a preset confidence value based on the driver's past driving events, a multi-task learning system, and the driver's current state. Based on a real-time comparison between the preset confidence value and the driving suitability confidence value, corresponding driving warning items are triggered, including: Based on the driver's tracing, corresponding past driving events are identified. Based on the identification of past driving events, multiple driving fatigue factors are determined. Based on multiple driving fatigue factors, a multi-task learning system, and the driver's current state, a preset confidence value is determined.
9. The method for identifying driver suitability credibility based on multi-task learning according to claim 8, characterized in that, The method of determining a preset confidence value based on the driver's past driving events, multi-task learning system, and the driver's current state, and triggering corresponding driving warning items based on real-time comparison of the preset confidence value and the driving suitability confidence value, also includes: Collect the driver's driving suitability confidence value, compare the driver's driving suitability confidence value with the preset confidence value, and output the corresponding comparison result; determine the corresponding driving warning item based on the mapping relationship between the comparison result and the driving warning item, so as to trigger the driving warning item.
10. A driver suitability confidence recognition system based on multi-task learning, characterized in that, The driver competence credibility recognition system based on multi-task learning is applied to the driver competence credibility recognition method based on multi-task learning as described in any one of claims 1-9.