Driving distraction detection method and system based on multi-expert collaboration

CN122821522APending Publication Date: 2026-09-25XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611034409.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明的目的是提供基于多专家协同的驾驶分心检测方法,解决了现有技术中存在的单一模型对复合型分心行为检测不准、误报率高的问题

Benefits of technology

[0016]本发明的有益效果是:现有技术多采用单一模型架构,难以捕捉疲劳、分神、异常操作等多种典型的分心行为,本发明并行工作的三个专家网络分别监控眼部、头部、手部,实现对不同类型分心行为的并行识别,能够覆盖驾驶员分心的主要表现形式,使检测更全面;现有技术多依赖于单帧图像进行判断,容易受到驾驶员瞬时动作(如打喷嚏、眨眼等因素)的干扰,从而产生误报,本发明设置了第四专家网络,采用LSTM分析历史得分序列,能够区分持续性分心与瞬时动作,有效抑制单帧误报,提高报警准确性和鲁棒性,使检测更准确;现有技术通常采用固定报警阈值,无法兼顾不同车速下的安全需求,本发明根据车速动态调整阈值,不仅适应了高速行驶时对安全性的更高的要求,还减少了低速行驶时无效的报警干扰;本发明门控网络利用车辆状态信息(车速、方向盘转角)动态调整专家权重,使系统能自适应高速公路、城市拥堵、转弯等不同驾驶场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821522A_ABST
    Figure CN122821522A_ABST
Patent Text Reader

Abstract

The application discloses a driving distraction detection method based on multi-expert cooperation and belongs to the technical field of driving assistance. The method comprises the following steps: acquiring a driver image and vehicle state information; outputting a state score after processing the driver image; obtaining an expert network fusion weight after processing the vehicle state information; calculating a current frame distraction score; calculating a time sequence consistency score according to a historical state score; and judging whether to alarm after calculating a final alarm score. The driving distraction detection system based on multi-expert cooperation comprises the following: an image acquisition module is connected with a first expert network, a second expert network and a third expert network respectively; a vehicle information acquisition module is connected with a gate network; the first expert network, the second expert network, the third expert network and the gate network are connected with a fourth expert network; and the fourth expert network is connected with an alarm module. The driving distraction detection method and system based on multi-expert cooperation realize parallel identification of different types of distraction behaviors and effectively inhibit single frame false alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of driving assistance technology, and relates to a driving distraction detection method based on multi-expert collaboration, as well as a driving distraction detection system based on multi-expert collaboration. Background Technology

[0002] Distracted driving is one of the main causes of traffic accidents. Existing distraction detection methods are mainly divided into two categories: one is based on vehicle status information (such as lane departure), which has a delayed response; the other is based on vision-based driver monitoring, but most of them are single-model architectures, which are difficult to accurately capture multiple distracting behaviors such as fatigue, gaze deviation, and hand gestures at the same time.

[0003] Furthermore, a single model relies on a single frame image and is easily affected by factors such as sudden changes in lighting and instantaneous actions of the driver (such as sneezing or blinking), which can lead to false alarms. Summary of the Invention

[0004] The purpose of this invention is to provide a driving distraction detection method based on multi-expert collaboration, which solves the problems of inaccurate detection of complex distraction behaviors and high false alarm rate of single models in the prior art.

[0005] Another objective of this invention is to provide a driver distraction detection system based on multi-expert collaboration.

[0006] The technical solution adopted in this invention is a driving distraction detection method based on multi-expert collaboration, comprising: Step 1: Obtain driver image and vehicle status information; Step 2: After extracting the driver's image, input it into the expert network and output the state score respectively; Step 3: Process the vehicle status information to obtain the expert network fusion weights; Step 4: Calculate the distraction score of the current frame in real time based on the state score and the expert network fusion weights; Step 5: Calculate the temporal consistency score based on the historical state score; Step 6: Calculate the final alarm score and then determine whether to trigger an alarm.

[0007] The invention is further characterized by: Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps of the trained first-expert network in processing eye region images and outputting eye state scores are as follows: Step A1: Input the eye region image sequence into the first expert network; Step A2: The first feature extraction layer of the first expert network performs convolution and pooling operations on each frame of the eye region image to extract eye spatial features. The first convolutional layer performs convolution operations on the input image to extract shallow features, which are then input into the first pooling layer for dimensionality reduction. The second convolutional layer extracts mid-level features, which are then input into the second pooling layer for dimensionality reduction. The third convolutional layer extracts high-level semantic features of the eye. Step A3: The first LSTM layer (Long Short-Term Memory Layer) receives the spatial feature map sequence output by the third convolutional layer, analyzes the temporal relationship between multiple consecutive frames, and distinguishes between blinking and continuous eye closing. Step A4: The first fully connected layer receives the temporal feature vector output from the first LSTM layer and maps it to fatigue score and line-of-sight deviation score through a linear transformation. The fatigue score is calculated based on the proportion of time the eyes are closed and the longest duration of eye closure; the higher the score, the more severe the fatigue. The line-of-sight deviation score is calculated based on the angle of deviation between the line of sight and the area in front; a higher score indicates a more severe line-of-sight deviation.

[0008] Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps of the trained second expert network in processing head region images and outputting head state scores are as follows: Step B1: Input the head region image into the second expert network; Step B2: After extracting shallow features from the head layer in the fourth convolutional layer, dimensionality reduction is performed using the third pooling layer; Step B3: After extracting the mid-layer features from the fifth convolutional layer, the features are input into the fourth pooling layer for dimensionality reduction. Step B4: Extract high-level semantic features from the header using the sixth convolutional layer; Step B5: The global average pooling layer reduces the dimensionality of the feature map; Step B6: The second fully connected layer maps the feature vectors to head pose angles; Head attitude angles include yaw angle, pitch angle, and roll angle; Step B7: Calculate the head-down duration and head-turning angle based on the changes in head posture angle over multiple consecutive frames, and output the head posture score after normalization.

[0009] Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps for a trained third-party expert network to process hand region images and output hand state scores are as follows: Step C1: Input the hand region image into the third expert network; Step C2: After extracting shallow features of the hand from the seventh convolutional layer, the fifth pooling layer is used for dimensionality reduction. Step C3: After extracting the mid-layer features of the hand from the eighth convolutional layer, the sixth pooling layer is used for dimensionality reduction. Step C4: Extract high-level semantic features of the hand from the ninth convolutional layer; Step C5: The third fully connected layer outputs the coordinates of the key points of both hands. Based on the key point positions, the average distance between the hands and the center point of the steering wheel and whether the hands have left the steering wheel are calculated. After normalization, the hand position score is output.

[0010] Step 3 includes: Step 3.1: Input the vehicle status information, namely vehicle speed, steering wheel angle, and turn signal status, into the input layer of the gating network; Step 3.2: The hidden layer extracts features from the input and learns the implicit association between vehicle speed, steering wheel angle, turn signal status and driver distraction type. Step 3.3: The output layer outputs three raw weight values; Step 3.4: Normalize the three original weight values ​​to obtain the fusion weights of the first, second, and third expert networks.

[0011] The formula for calculating the distraction score in the current frame is: D_{frame}=w1*e+w2*h+w3*p, Where D_{frame} is the distraction score of the current frame; e is the output score of the first expert network; h is the output score of the second expert network; p is the output score of the third expert network; w1 is the fusion weight of the first expert network; w2 is the fusion weight of the second expert network; and w3 is the fusion weight of the third expert network.

[0012] Step 5 includes: Step 5.1: Input the historical score sequences output by the first, second, and third expert networks within the past T seconds into the fourth expert network; Step 5.2: The second LSTM layer performs temporal modeling on the input sequence, learns the temporal dependencies in the score sequence, and outputs the hidden state vector; The time dependence in the score sequence includes the stability and trend of the scores; Step 5.3: The fourth fully connected layer linearly transforms the hidden state vector into a scalar. Step 5.4: Map the scalar to a timing consistency score.

[0013] The formula for calculating the temporal consistency score is: C=σ(W· +b), Wherein, the temporal consistency score C represents the stability and trend direction of the distracted state in the past T seconds, and σ represents the Sigmoid activation function. denoted as the hidden state vector output by the second LSTM layer; W represents the weight matrix of the fourth fully connected layer; b represents the bias term of the fourth fully connected layer.

[0014] Step 6 includes: Calculate the final alarm score, and trigger an alarm when the final alarm score D_{final} exceeds a preset threshold; The formula for calculating the final alarm score is: D_{final}=D_{frame}*C, Where D_{final} is the final alarm score, C is the timing consistency score, and D_{frame} is the distraction score of the current frame.

[0015] Another technical solution adopted in this invention is a driver distraction detection system based on multi-expert collaboration, including an image acquisition module and a vehicle information acquisition module. The image acquisition module is connected to a first expert network, a second expert network, and a third expert network, respectively. The vehicle information acquisition module is connected to a gating network. The first expert network, the second expert network, the third expert network, and the gating network are connected to a fourth expert network. The fourth expert network is connected to an alarm module, and the alarm module is connected to a threshold generation module.

[0016] The beneficial effects of this invention are as follows: Existing technologies mostly adopt a single model architecture, which makes it difficult to capture various typical distraction behaviors such as fatigue, inattention, and abnormal operation. The three expert networks of this invention, which work in parallel, monitor the eyes, head, and hands respectively, to achieve parallel recognition of different types of distraction behaviors, covering the main manifestations of driver distraction and making the detection more comprehensive. Existing technologies mostly rely on single-frame images for judgment, which are easily interfered with by the driver's instantaneous actions (such as sneezing, blinking, etc.), thus producing false alarms. This invention sets up a fourth expert network and uses LSTM to analyze historical score sequences, which can distinguish between continuous distraction and instantaneous actions, effectively suppressing single-frame false alarms, improving alarm accuracy and robustness, and making the detection more accurate. Existing technologies usually use fixed alarm thresholds, which cannot take into account the safety requirements at different vehicle speeds. This invention dynamically adjusts the threshold according to the vehicle speed, which not only adapts to the higher safety requirements at high speeds, but also reduces invalid alarm interference at low speeds. The gating network of this invention uses vehicle status information (vehicle speed, steering wheel angle) to dynamically adjust expert weights, enabling the system to adapt to different driving scenarios such as highways, urban congestion, and turning. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the internal network structure of the first expert in an embodiment of the present invention; Figure 2 This is a schematic diagram of the internal network structure of the second expert in an embodiment of the present invention; Figure 3 This is a schematic diagram of the internal network structure of the third expert in an embodiment of the present invention; Figure 4 This is a schematic diagram of the internal structure of the gated network in an embodiment of the present invention; Figure 5 This is a schematic diagram of the internal network structure of the fourth expert in an embodiment of the present invention. Detailed Implementation

[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0019] Example 1 This embodiment proposes a driver distraction detection method based on multi-expert collaboration, including: Step 1: Obtain driver image and vehicle status information; Step 2: After extracting the driver's image, input it into the expert network and output the state score respectively; Step 3: Process the vehicle status information to obtain the expert network fusion weights; Step 4: Calculate the distraction score of the current frame in real time based on the state score and the expert network fusion weights; Step 5: Calculate the temporal consistency score based on the historical state score; Step 6: Calculate the final alarm score and then determine whether to trigger an alarm.

[0020] Example 2 Based on Example 1, this example proposes that step 1 includes: The system captures real-time images of the driver via an in-vehicle camera and obtains vehicle status information, including vehicle speed, steering wheel angle, and turn signal status, via the vehicle's CAN (Controller Area Network) bus.

[0021] Example 3 Based on Example 1, this example proposes step 2, which includes: extracting the eye region image, head region image, and hand region image from the driver image respectively; then, the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score.

[0022] Example 4 Based on Example 3, this example proposes the following specific steps for the trained first expert network to process eye region images and output eye state scores: Step A1: Input the eye region image sequence into the first expert network; Step A2: The first feature extraction layer of the first expert network performs convolution and pooling operations on each frame of the eye region image to extract the spatial features of the eye. The first convolutional layer performs convolution operations on the input image to extract shallow features such as eyelid edges and pupil contours. Then, it is input into the first pooling layer for dimensionality reduction. The second convolutional layer extracts mid-level features and inputs them into the second pooling layer for further dimensionality reduction. The third convolutional layer extracts high-level semantic features of the eye, including the degree of eyelid opening and closing, pupil offset, etc. Step A3: Receive the spatial feature map sequence output by the third convolutional layer through the first LSTM layer, analyze the temporal relationship between consecutive frames, and distinguish between blinking and continuous eye closing. Step A4: The first fully connected layer receives the temporal feature vector output from the first LSTM layer and maps it to two values ​​through a linear transformation, namely the fatigue score and the line-of-sight deviation score: Fatigue score: Calculated based on the proportion of time the eyes are closed and the longest time the eyes are closed. The higher the score, the more severe the fatigue. Line of sight deviation score: Calculated based on the deviation angle between the line of sight direction and the area in front. The higher the score, the more serious the line of sight deviation.

[0023] The formula for calculating fatigue score is as follows: , The proportion of eye closure time is calculated as: total eye closure time within the statistical period / total duration of the statistical period. The longest eye closure time is the longest single eye closure time within the statistical period. This represents a preset normalization constant used to normalize the longest eye-closing time to the interval [0, 1]. The weighting coefficients representing the proportion of eye closure time in the fatigue score are determined by network training; The weighting coefficient for the longest duration of eye closure is determined during network training.

[0024] The formula for calculating the line-of-sight deviation score is as follows: , in, The scaling factor representing the gaze deviation score is determined automatically during network training; This indicates the angle of deviation between the line of sight and the area in front. This represents the preset maximum deviation angle, used to normalize the deviation angle to the [0, 1] interval.

[0025] Example 5 Based on Example 3, this example proposes the following specific steps for the trained second expert network to process head region images and output head state scores: Step B1: Input the head region image into the second expert network; Step B2: The fourth convolutional layer extracts shallow head features, including the top contour of the head, the edge of the chin, the shadows on both sides of the nose, and other edge information, and reduces the dimensionality through the third pooling layer. Step B3: The fifth convolutional layer extracts mid-level head features and identifies head components such as hair, nose, eyes, and ears, which are then input into the fourth pooling layer for dimensionality reduction. Step B4: The sixth convolutional layer extracts high-level semantic features of the head, including the relative offset of the nose tip (reflecting the yaw angle), the longitudinal position of the nose tip (reflecting the pitch angle), and the horizontal angle of the eyes (reflecting the roll angle), etc. Step B5: The global average pooling layer reduces the dimensionality of the feature map; Step B6: The second fully connected layer maps the feature vectors into three attitude angles: head yaw angle, pitch angle, and roll angle. Step B7: Calculate the head-down duration and head-turning angle based on the changes in head pose angle over multiple consecutive frames. After normalization, output the head pose score. The higher the score, the more severe the abnormal head pose.

[0026] The formula for calculating head pose score is as follows: , Among them, yaw angle represents the angle at which the head turns left and right; pitch angle represents the angle at which the head turns up and down; roll angle represents the angle at which the head tilts left and right. This represents the maximum preset value for the yaw angle, used for normalization. This represents the maximum preset value for the pitch angle, used for normalization. This represents the maximum preset value for the roll angle, used for normalization. , , The weight coefficients for each attitude angle are automatically learned during the training of the second expert network.

[0027] Example 6 Based on Example 3, this example proposes the following specific steps for a trained third-expert network to process hand region images and output hand state scores: Step C1: Input the hand region image into the third expert network; Step C2: The seventh convolutional layer extracts shallow features of the hand, including finger edges, wrist contours, and skin texture; the fifth pooling layer reduces dimensionality. Step C3: The eighth convolutional layer extracts mid-level features of the hand, including the joint position of individual fingers and the shape of the palm region; the sixth pooling layer further reduces the dimensionality. Step C4: The ninth convolutional layer extracts high-level semantic features of the hand, spatially aggregates the scattered finger and palm features, constructs the geometric constraint relationship of the hand skeleton, and analyzes the relative position and topological structure between key points of the hand. Step C5: Output the coordinates of the key points of both hands in the fully connected layer. Calculate the average distance between the hands and the center point of the steering wheel and whether the hands have left the steering wheel based on the key point positions. After normalization, output the hand position score. The higher the score, the more serious the deviation of the hands from the normal driving position.

[0028] The formula for calculating the hand position score is as follows: , Where d is the average pixel distance between the key points of the hand and the center point of the steering wheel; The preset maximum distance is used for normalization; L indicates whether the hands have left the steering wheel, L=1 indicates they have left, and L=0 indicates they have not left. , These are the weight coefficients, which are automatically learned by the third expert network during training.

[0029] Example 7 Based on Example 1, this example proposes that step 3 includes: Step 3.1: Input the vehicle speed, steering wheel angle, and turn signal status into the input layer of the gating network; Step 3.2: The hidden layer extracts features from the input and learns the implicit association between vehicle speed, steering wheel angle, turn signal status and driver distraction type. Step 3.3: The output layer outputs three raw weight values; Step 3.4: The Softmax layer (normalized exponential function layer) normalizes the three original weight values ​​to obtain the fusion weights w1, w2, and w3 of the first, second, and third expert networks, satisfying w1+w2+w3=1.

[0030] The formula for calculating the fusion weights of an expert network is as follows: , , , Where x is the input vector of the gating network, including vehicle speed, steering wheel angle, and turn signal status; This is the weight matrix of the output layer of the gated network; For bias terms; This represents the original output value of the i-th expert network; This represents the fusion weight of the i-th expert network.

[0031] Example 8 Based on Example 1, this example proposes the following formula for calculating the distraction score of the current frame: D_{frame}=w1*e+w2*h+w3*p, Where D_{frame} is the distraction score of the current frame; e is the output score of the first expert network; h is the output score of the second expert network; p is the output score of the third expert network; w1 is the fusion weight of the first expert network; w2 is the fusion weight of the second expert network; and w3 is the fusion weight of the third expert network.

[0032] Example 9 Based on Example 1, this example proposes that step 5 includes: Step 5.1: Input the historical score sequences output by the first, second, and third expert networks within the past T seconds into the fourth expert network; Step 5.2: The second LSTM layer performs temporal modeling on the input sequence, learning the temporal dependencies in the score sequence, including the stability of the scores (consistently high scores or fluctuating scores) and the trend of the scores (gradually increasing or gradually decreasing), and outputs the hidden state vector. ; , , Where t represents the time step index, t=1 indicates the first frame; T represents the total number of frames in time. Let represent the input vector for frame t, which includes the network output scores of the first, second, and third experts; This represents the hidden state vector of frame t. The hidden state vector of the previous frame; This is the unit vector representing the memory state of the previous frame; This represents the internal computation function of the LSTM unit; Hidden state for the last frame .

[0033] Step 5.3: The fourth fully connected layer hides the state vector. A linear transformation is converted into a scalar; Step 5.4: The Sigmoid activation function maps this scalar to a temporal consistency score C between 0 and 1, calculated as follows: C=σ(W· +b), Wherein, the temporal consistency score C represents the stability and trend direction of the distracted state within the past T seconds, and σ represents the Sigmoid activation function; represents the hidden state vector output by the LSTM layer; W represents the weight matrix of the fully connected layer; b represents the bias term of the fully connected layer. When the historical score remains consistently high, C→1; when the historical score fluctuates (e.g., during a momentary action) or shows a downward trend, C→0.

[0034] Compared with simple statistics (such as mean and variance), LSTM can capture the temporal trend of score changes. For example, when the score gradually rises from 0.3 to 0.7, even if the score of the current frame is only 0.6, LSTM can output a higher C value and give an early warning. In this scenario, the average method will output a lower value (about 0.5), resulting in missed detection. Example 10 Based on Example 1, this example proposes that step 6 includes: Calculate the final alarm score. When the final alarm score D_{final} exceeds the preset threshold, trigger an alarm (such as a vibration alert, voice alert, etc.). The formula for calculating the final alarm score is: D_{final} = D_{frame}*C, Where D_{final} is the final alarm score, C is the timing consistency score, and D_{frame} is the distraction score of the current frame; The alarm threshold is dynamically adjusted based on the vehicle speed signal: preset alarm threshold, first preset speed, and second preset speed. When the vehicle speed exceeds the first preset speed, the alarm threshold is lowered; when the vehicle speed is lower than the second preset speed, the alarm threshold is raised.

[0035] Example 11 This embodiment proposes a driver distraction detection system based on multi-expert collaboration, including an image acquisition module and a vehicle information acquisition module. The image acquisition module is connected to a first expert network, a second expert network, and a third expert network, respectively. The vehicle information acquisition module is connected to a gating network. The first expert network, the second expert network, the third expert network, and the gating network are connected to a fourth expert network. The fourth expert network is connected to an alarm module, and the alarm module is connected to a threshold generation module.

[0036] Image acquisition module: used to acquire images of the driver's face and upper body; Vehicle information acquisition module: used to acquire vehicle speed, steering wheel angle and turn signal status via CAN bus; First expert network: used to process eye images and output eye state scores; Second expert network: used to process head images and output head pose scores; Third expert network: used to process hand region images and output hand position scores; Gated network: used to output the fusion weights of the first, second, and third expert networks based on vehicle status information; The fourth expert network is used to output a temporal consistency score based on the historical score sequence. Threshold generation module: used to dynamically adjust alarm thresholds based on the vehicle status information; Alarm module: Used to calculate the final alarm score based on the distraction score and timing consistency score of the current frame, and trigger an alarm when the threshold is exceeded.

[0037] The first expert network is an eye state recognition network based on CNN (Convolutional Neural Network) and LSTM. Its input is multiple consecutive frames of eye images, and its output is an eye state score calculated based on factors such as the proportion of eye closure time, the longest eye closure duration, and the deviation angle between the gaze direction and the area in front. Figure 1 As shown, the first expert network comprises a first feature extraction layer, a first LSTM layer, and a fully connected layer connected in sequence. The feature extraction layer consists of multiple convolutional layers and pooling layers alternating between each other. The convolutional layers are used to extract local spatial features of the eye image, while the pooling layers are responsible for reducing dimensionality and extracting spatial features of the eye image at each step. The first LSTM layer includes LSTM memory units, used to analyze the temporal relationship between multiple consecutive frames and distinguish between blinking and prolonged eye closure. The first fully connected layer includes input nodes and output nodes, used to output fatigue scores and gaze deviation scores.

[0038] like Figure 2As shown, the second expert network is a lightweight head pose estimation network. Its input is a single-frame head image, and its output is the head yaw angle, pitch angle, and roll angle. The second feature extraction layer consists of alternating depthwise separable convolutional layers and pooling layers, used to extract head image features. The average pooling layer includes a pooling window, compressing the feature map into a one-dimensional feature vector. The second fully connected layer includes input and output nodes, used to output the head yaw angle, pitch angle, and roll angle, and output the head position score.

[0039] like Figure 3 As shown, the third expert network is a hand keypoint detection network. Its input is an image region containing the arm and shoulder, and its output is a hand position score calculated based on the hand keypoint positions and the relative distance between the hand and the steering wheel. The third feature extraction layer includes multiple convolutional and pooling layers distributed alternately to extract hand features such as hand edges and finger joints, and outputs a spatial feature map of the hand keypoints. The third fully connected layer includes input and output nodes, mapping the hand keypoints extracted by the third feature extraction layer to two-dimensional coordinates of multiple keypoints on both hands. The distance between the hand and the center point of the steering wheel is calculated based on the keypoint coordinates and the center coordinates of the steering wheel, and it is determined whether the hand has left the steering wheel. After normalization, this is converted into a hand position score.

[0040] like Figure 4 As shown, the gated network is a multilayer perceptron. Its inputs are vehicle speed, steering wheel angle, and turn signal status. The output, after Softmax normalization, yields the fusion weights of three expert networks. The input layer includes a data interface for receiving information such as vehicle speed, steering wheel angle, and turn signal status. The hidden layer includes fully connected neurons and a ReLU (Rectified Linear Unit) activation function to extract correlations between input features and learn weight allocation patterns under different driving scenarios. The output layer includes fully connected neurons to output three raw weight values. The Softmax layer includes a Softmax normalization function, transforming the three raw weight values ​​output by the output layer into w1, w2, and w3, where w1 + w2 + w3 = 1.

[0041] like Figure 5 As shown, the fourth expert network is a temporal consistency verification network. Its input is the distraction probability sequence output by the first, second, and third expert networks over the past T seconds, and its output is the temporal consistency score. The input layer includes a data interface for receiving the three expert scores for each frame over the past T seconds. The second LSTM layer includes LSTM memory units for learning the stability and trend of the score sequence and outputting a hidden state vector. The fourth fully connected layer includes input and output nodes for storing the hidden state vector. The mapping is to a scalar. The Sigmoid layer includes a Sigmoid activation function, which maps this scalar to a 0-1 temporal consistency score C.

[0042] This invention acquires driver images via an in-vehicle camera, extracting the eye, head, and hand regions. First, second, and third expert networks are used to identify these regions and output scores. Simultaneously, vehicle information such as speed, steering wheel angle, and traffic light status is input into a gating network. This gating network assigns weights to each expert to calculate the distraction score for the current frame. A fourth expert network performs temporal consistency verification on the historical score sequence and outputs a temporal consistency score. Furthermore, this invention dynamically adjusts a threshold based on vehicle information, multiplying the current frame distraction score by the temporal consistency score to obtain the final score, which is then compared with the dynamic threshold to trigger an alarm. Through multi-expert collaboration and vehicle context awareness, this invention significantly improves the accuracy and robustness of driver distraction detection.

[0043] An embodiment of the present invention provides distraction detection in a high-speed cruising scenario: A vehicle is cruising at 100 km / h on a highway, with the steering wheel angle close to 0° and the turn signal off. A preset alarm threshold is T0 = 0.65, a preset first speed is 80 km / h, and a preset second speed is 25 km / h. When the vehicle speed exceeds the first preset speed, the alarm threshold T = T0 - 0.15; when the vehicle speed is below the second preset speed, the alarm threshold T = T0 + 0.15. The gating network inputs vehicle speed 100, steering wheel angle 0, and turn signal off, and outputs weights as follows: eye expert w1 = 0.6, head expert w2 = 0.2, and hand expert w3 = 0.2.

[0044] The driver looked down at his phone for three consecutive seconds: Eye specialist: Visual deviation score 0.85; Head expert: Pitch angle abnormality score 0.80; Hands expert: Score for hands off the steering wheel 0.70; The distraction score for the current frame is D_{frame} = 0.6 * 0.85 + 0.2 * 0.80 + 0.2 * 0.70 = 0.81; The fourth expert network takes as input the score sequence of the past 3 seconds (all around 0.8) and outputs a timing consistency score C=0.95 (consistently high). The final alarm score D_{final}=0.81*0.95=0.77.

[0045] Because the vehicle speed is higher than the first preset value, the system automatically lowers the alarm threshold. T=0.65-0.15=0.50. Finally, the alarm score exceeds the dynamic threshold, and the system triggers an alarm.

[0046] No false alarms for momentary actions (sneezing): When the driver sneezes, the head and eyes momentarily deviate, but the duration is less than 0.5 seconds: the distraction score of the current frame D_{frame}=0.75, but the scores of other frames in the past 3 seconds are all below 0.3, the timing consistency score output by the fourth expert network is C=0.15; the final alarm score D_{final}=0.75*0.15=0.11, which is lower than the dynamic threshold, so no alarm is triggered.

Claims

1. A driving distraction detection method based on multi-expert collaboration, characterized in that, include: Step 1: Obtain driver image and vehicle status information; Step 2: After extracting the driver's image, input it into the expert network and output the state score respectively; Step 3: Process the vehicle status information to obtain the expert network fusion weights; Step 4: Calculate the distraction score of the current frame in real time based on the state score and the expert network fusion weights; Step 5: Calculate the temporal consistency score based on the historical state score; Step 6: Calculate the final alarm score and then determine whether to trigger an alarm.

2. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps of the trained first-expert network in processing eye region images and outputting eye state scores are as follows: Step A1: Input the eye region image sequence into the first expert network; Step A2: The first feature extraction layer of the first expert network performs convolution and pooling operations on each frame of the eye region image to extract eye spatial features. The first convolutional layer performs convolution operations on the input image to extract shallow features, which are then input into the first pooling layer for dimensionality reduction. The second convolutional layer extracts mid-level features, which are then input into the second pooling layer for dimensionality reduction. The third convolutional layer extracts high-level semantic features of the eye. Step A3: The first LSTM layer receives the spatial feature map sequence output by the third convolutional layer, analyzes the temporal relationship between multiple consecutive frames, and distinguishes between blinking and continuous eye closure. Step A4: The first fully connected layer receives the temporal feature vector output from the first LSTM layer and maps it to fatigue score and line-of-sight deviation score through a linear transformation. The fatigue score is calculated based on the proportion of time the eyes are closed and the longest duration of eye closure; the higher the score, the more severe the fatigue. The line-of-sight deviation score is calculated based on the angle of deviation between the line of sight and the area in front; a higher score indicates a more severe line-of-sight deviation.

3. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps of the trained second expert network in processing head region images and outputting head state scores are as follows: Step B1: Input the head region image into the second expert network; Step B2: After extracting shallow features from the head layer in the fourth convolutional layer, dimensionality reduction is performed using the third pooling layer; Step B3: After extracting the mid-layer features from the fifth convolutional layer, the features are input into the fourth pooling layer for dimensionality reduction. Step B4: Extract high-level semantic features from the header using the sixth convolutional layer; Step B5: The global average pooling layer reduces the dimensionality of the feature map; Step B6: The second fully connected layer maps the feature vectors to head pose angles; Head attitude angles include yaw angle, pitch angle, and roll angle; Step B7: Calculate the head-down duration and head-turning angle based on the changes in head posture angle over multiple consecutive frames, and output the head posture score after normalization.

4. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 2 includes: extracting the eye region image, head region image and hand region image from the driver image respectively; the trained first expert network processes the eye region image and outputs the eye state score; the trained second expert network processes the head region image and outputs the head state score; and the trained third expert network processes the hand region image and outputs the hand state score. The specific steps for a trained third-party expert network to process hand region images and output hand state scores are as follows: Step C1: Input the hand region image into the third expert network; Step C2: After extracting shallow features of the hand from the seventh convolutional layer, the fifth pooling layer is used for dimensionality reduction. Step C3: After extracting the mid-layer features of the hand from the eighth convolutional layer, the sixth pooling layer is used for dimensionality reduction. Step C4: Extract high-level semantic features of the hand from the ninth convolutional layer; Step C5: The third fully connected layer outputs the coordinates of the key points of both hands. Based on the key point positions, the average distance between the hands and the center point of the steering wheel and whether the hands have left the steering wheel are calculated. After normalization, the hand position score is output.

5. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 3 includes: Step 3.1: Input the vehicle speed, steering wheel angle, and turn signal status into the input layer of the gating network; Step 3.2: The hidden layer extracts features from the input and learns the implicit association between vehicle speed, steering wheel angle, turn signal status and driver distraction type. Step 3.3: The output layer outputs three raw weight values; Step 3.4: Normalize the three original weight values ​​to obtain the fusion weights of the first, second, and third expert networks.

6. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, The formula for calculating the distraction score of the current frame is: D_{frame}=w1*e+w2*h+w3*p, Where D_{frame} is the distraction score of the current frame; e is the output score of the first expert network; h is the output score of the second expert network; p is the output score of the third expert network; w1 is the fusion weight of the first expert network; w2 is the fusion weight of the second expert network; and w3 is the fusion weight of the third expert network.

7. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 5 includes: Step 5.1: Input the historical score sequences output by the first, second, and third expert networks within the past T seconds into the fourth expert network; Step 5.2: The second LSTM layer performs temporal modeling on the input sequence, learns the temporal dependencies in the score sequence, and outputs the hidden state vector; The time dependence in the score sequence includes the stability and trend of the scores; Step 5.3: The fourth fully connected layer linearly transforms the hidden state vector into a scalar. Step 5.4: Map the scalar to a timing consistency score.

8. The driving distraction detection method based on multi-expert collaboration according to claim 7, characterized in that, The formula for calculating the timing consistency score is as follows: C=σ(W· +b), Wherein, the temporal consistency score C represents the stability and trend direction of the distracted state within the past T seconds, and σ represents the Sigmoid activation function; denoted as the hidden state vector output by the second LSTM layer; W represents the weight matrix of the fourth fully connected layer; b represents the bias term of the fourth fully connected layer.

9. The driving distraction detection method based on multi-expert collaboration according to claim 1, characterized in that, Step 6 includes: calculating the final alarm score, and triggering an alarm when the final alarm score exceeds a preset threshold; The formula for calculating the final alarm score is: D_{final} = D_{frame}*C, Where D_{final} is the final alarm score, C is the timing consistency score, and D_{frame} is the distraction score of the current frame.

10. A driver distraction detection system based on multi-expert collaboration, characterized in that: The method for detecting driver distraction based on multi-expert collaboration as described in any one of claims 1 to 9 includes an image acquisition module and a vehicle information acquisition module. The image acquisition module is connected to a first expert network, a second expert network, and a third expert network, respectively. The vehicle information acquisition module is connected to a gating network. The first expert network, the second expert network, the third expert network, and the gating network are connected to a fourth expert network. The fourth expert network is connected to an alarm module, and the alarm module is connected to a threshold generation module.