An intelligent driving assistance safety management method and system

CN122770751APending Publication Date: 2026-09-18CHONGQING KUPURUI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610911719.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]现有技术中,主流驾驶员监测系统主要依赖摄像头采集驾驶图像,通过视觉算法识别驾驶员手部是否处于方向盘区域,部分方案尝试通过增加方向盘电容传感器或扭矩传感器进行补充判断,但额外硬件不仅增加成本且无法解决视觉注意力监测的根本需求,传感器融合策略也缺乏针对手掌特征缺失场景的系统性补偿机制

Benefits of technology

[0076] 1. Acquire driving images and identify upper limb posture to determine hand features. When hand features are missing, calculate the forearm length based on the upper arm length, and then determine the hand position and calculate the hand distance by combining joint orientation and joint position. When the hand distance is continuously less than the operation threshold, issue an attention warning. In this way, when hand features are missing, the geometric constraints of the upper limb kinematic chain are used to calculate the hand position, realize continuous monitoring of driver attention, and improve the safety of intelligent driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122770751A_ABST
    Figure CN122770751A_ABST
Patent Text Reader

Abstract

The application relates to an intelligent driving auxiliary safety management method and system, and relates to the field of intelligent driving, which comprises the following steps: collecting a driving image in response to a preset intelligent driving instruction; identifying an upper limb posture from the driving image, and determining palm features based on the upper limb posture; when the palm features are empty, identifying an upper arm length from the upper limb posture, and determining a forearm length according to the upper arm length; identifying joint orientations and joint positions from the upper limb posture, determining a palm position in combination with the joint positions, the joint orientations and the forearm length, and determining palm spacings by comparing the palm position; if the palm spacings are smaller than a preset operation threshold, determining a continuous duration based on the palm spacings; and generating and displaying an attention warning in response to the continuous duration. The application has the effects of improving the safety of intelligent driving and accurately identifying the driving state of a driver.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving, and in particular to an intelligent driving assistance safety management method and system. Background Technology

[0002] Assisted safety management technology refers to the use of onboard sensors to monitor the driver's driving status in real time and issue warnings when a deviation in attention is detected to ensure driving safety. It is the core safety defense line of human-machine co-driving in high-level intelligent driving.

[0003] In existing technologies, mainstream driver monitoring systems mainly rely on cameras to collect driving images and use visual algorithms to identify whether the driver's hands are in the steering wheel area. Some solutions attempt to supplement the judgment by adding a steering wheel capacitive sensor or torque sensor, but the additional hardware not only increases costs but also fails to solve the fundamental need for visual attention monitoring. The sensor fusion strategy also lacks a systematic compensation mechanism for scenarios where hand features are missing.

[0004] If the hands are obscured by the posture of holding the steering wheel, skin color features are lost due to wearing gloves, or the interior lighting is insufficient or there is strong backlight, the existing system usually directly judges it as undetectable or simply classifies it as a hands-off state. This either results in frequent false alarms affecting the driving experience or missed alarms causing safety hazards. Summary of the Invention

[0005] To improve the safety of intelligent driving and accurately identify the driver's driving status, this invention provides an intelligent driving assistance safety management method and system.

[0006] In a first aspect, the present invention provides an intelligent driving assistance safety management method, which adopts the following technical solution:

[0007] A method for intelligent driving assistance safety management includes:

[0008] Step 100: Acquire driving images in response to preset intelligent driving commands;

[0009] Step 101: Identify the upper limb posture from the driving image and determine the hand features based on the upper limb posture;

[0010] Step 102: When the palm feature is empty, identify the upper arm length from the upper limb posture, and determine the forearm length based on the upper arm length;

[0011] Step 103: Identify the joint orientation and joint position from the upper limb posture, and determine the palm position by combining the joint position, joint orientation and forearm length, and determine the palm spacing by comparing the palm position;

[0012] Step 104: If the distance between the palms is less than a preset operation threshold, determine the duration based on the distance between the palms;

[0013] Step 105: Generate and display an attention warning in response to the duration.

[0014] By adopting the above technical solution, driving images are acquired and upper limb posture is identified to determine hand features. When the hand features are missing, the forearm length is estimated based on the upper arm length. Then, the hand position is determined by combining the joint orientation and joint position, and the hand distance is calculated. When the hand distance is continuously less than the operation threshold, an attention warning is issued. Thus, when the hand features are missing, the geometric constraints of the upper limb kinematic chain are used to estimate the hand position, thereby realizing continuous monitoring of the driver's attention and improving the safety of intelligent driving.

[0015] Optionally, the method for determining the upper arm length includes:

[0016] Step 106: Identify joint features from the upper limb posture;

[0017] Step 107: Determine the visual length and angle of the upper arm from the upper limb posture based on the joint features;

[0018] Step 108: Determine the visual coefficient based on the upper arm angle;

[0019] Step 109: Determine the upper arm length by combining the upper arm visual length and visual coefficient.

[0020] By adopting the above technical solution, joint features are identified from the upper limb posture, and the visual length and angle of the upper arm are determined based on the joint features. The mapping relationship between the upper arm angle and visual distortion is established by using the camera perspective projection principle to determine the visual coefficient. Then, the visual coefficient is used to compensate for the length measurement error caused by the change of arm orientation in the two-dimensional image to obtain the true upper arm length. This reduces the calculation deviation caused by the driver's arm extending in different directions and improves the accuracy of subsequent palm position calculation.

[0021] Optionally, the method for determining the forearm length includes:

[0022] Step 110: When the palm feature is not empty, determine the visual length of the forearm and the visual length of the palm from the upper limb posture based on the joint feature and the palm feature;

[0023] Step 111: Identify the upper arm angle and forearm angle from the upper limb posture based on the joint features, and determine the compensation coefficient by comparing the upper arm angle and forearm angle;

[0024] Step 112: Determine the forearm compensation length by combining the forearm visual length and compensation coefficient, and determine the palm compensation length by combining the palm visual length and compensation coefficient;

[0025] Step 113: Determine the upper limb proportion by comparing the visual length of the upper arm, the compensated length of the forearm, and the compensated length of the palm;

[0026] Step 114: When the palm feature is empty, determine the forearm ratio based on the upper limb proportion, and determine the forearm length by combining the forearm ratio and the upper arm length.

[0027] By adopting the above technical solution, a two-stage mechanism of establishing a standard in normal state and calling it in abnormal state is used. When the palm features are visible, the visual length of the upper arm, forearm, and palm are collected in advance. The actual length is calculated and the upper limb proportion is established by combining the compensation coefficient determined by the joint angle comparison. When the palm features are missing, the proportion is called to estimate the forearm length. The length distortion caused by the visual projection is compensated by the joint angle, thereby realizing personalized and high-precision estimation of the forearm length and reducing the individual difference error between different drivers caused by using a fixed experience proportion.

[0028] Optionally, it also includes an attention detection method, the attention detection method comprising:

[0029] Step 200: If the distance between the palms is less than a preset operation threshold, determine the distance fluctuation curve based on the distance between the palms, and extract the distance rise curve from the distance fluctuation curve;

[0030] Step 201: Determine the rising distance based on the aforementioned spacing rise curve;

[0031] Step 202: When the rising distance is greater than the preset distance threshold, determine the distance departure time based on the rising distance, and read the initial distance from the spacing rise curve according to the rising distance;

[0032] Step 203: Read the recovery time that is located after the distance-away time and equal to the initial distance from the distance fluctuation curve, and compare the distance-away time and the recovery time to determine the recovery duration;

[0033] Step 204: If the recovery time is less than the preset attention threshold, generate and display an attention warning in response to the duration.

[0034] By adopting the above technical solution, the judgment dimension is expanded from static spatial geometry to dynamic temporal behavior patterns. When the distance between the palms is too small, the rising curve in the distance fluctuation curve is extracted to represent how quickly the hand returns after leaving. When the rising distance is greater than the distance threshold, the distance departure time and the initial distance are recorded. Then, the recovery time is read from the fluctuation curve as the micro-behavioral feature of the driver's high-frequency back-and-forth operation in the palm area. The attention state is judged by calculating the recovery time, thereby accurately identifying the dangerous distraction state of the driver who is overly focused on the local operation area.

[0035] Optionally, the attention detection method further includes:

[0036] Step 205: If the recovery time is less than a preset attention threshold, identify the set of visual focal points from the driving image based on the time of departure, and determine the operation center based on the palm position;

[0037] Step 206: Determine the visual spacing by comparing the set of visual focal points with the operation center;

[0038] Step 207: When the visual distance is less than a preset tracking threshold, determine the tracking time based on the visual distance, and compare the tracking time and the recovery time to determine the tracking difference;

[0039] Step 208: Determine the risk level based on the tracking difference, and match the speed limit according to the risk level;

[0040] Step 209: In response to the speed limit generation, a speed limit command is sent.

[0041] By adopting the above technical solution, through cross-modal temporal pairing analysis of gaze and hand behavior, the set of visual landing points is identified when the recovery time is too short. The operation center is determined according to the palm position and the visual distance between the two is calculated. When the visual distance is less than the tracking threshold, the tracking time is recorded and compared with the recovery time to obtain the tracking difference as the phase difference between the hand recovery time and the gaze tracking time. The tracking difference is used as a quantitative indicator of cognitive load, thereby determining the risk level and matching the vehicle speed limit, thus achieving a refined graded assessment of attention state.

[0042] Optionally, the attention detection method further includes:

[0043] Step 210: When the tracking difference falls into the preset visual range, determine the approach time from the visual landing point set based on the tracking time, and retrieve the approach distance by combining the approach time and the tracking time;

[0044] Step 211: Generate a proximity curve based on the proximity spacing, and identify the proximity rate from the proximity curve;

[0045] Step 212: Analyze the approach rate to determine the approach deceleration, and determine the visual dependence based on the approach deceleration;

[0046] Step 213: Match the dependency coefficient according to the visual dependency, and update the risk level according to the dependency coefficient.

[0047] By adopting the above technical solution, the second derivative of kinematics is introduced as an analytical dimension of cognitive state. When the tracking difference falls into the visual interval, the approach time is determined and the approach distance is retrieved. After generating the distance approach curve, the approach rate is identified from it, and then the approach deceleration is analyzed to determine the visual dependence. The driver's dependence on visual confirmation is quantitatively characterized by the urgency of the gaze approaching the operating area. Finally, the dependence coefficient is matched to update the risk level, thereby establishing a quantitative inference chain from behavioral kinematic parameters to cognitive state, providing a deeper cognitive behavioral indicator for attention risk assessment.

[0048] Optionally, it also includes a confidence calculation method, the confidence calculation method comprising:

[0049] Step 300: If the distance between the palms is less than a preset operation threshold, identify the image brightness, image contrast, local variance and global variance from the driving image, and determine the image sharpness by combining the local variance and global variance;

[0050] Step 301: Determine the brightness confidence level based on the image brightness, and determine the contrast confidence level based on the image contrast.

[0051] Step 302: Determine the image quality confidence level by combining the image sharpness, brightness confidence level, and contrast confidence level, and determine the proportion mean and proportion standard deviation based on the upper limb proportion;

[0052] Step 303: Determine the stability of the proportion by combining the mean and standard deviation of the proportion, and match the stability coefficient according to the stability of the proportion;

[0053] Step 304: Correct the quality confidence score according to the stability coefficient to obtain the image confidence score, and match the confidence score coefficient according to the image confidence score;

[0054] Step 305: Respond to the confidence coefficient update duration to reduce false alarms.

[0055] By adopting the above technical solution, a confidence assessment system is constructed from the dual perspectives of image quality stability and upper limb proportion stability. Brightness, contrast, and sharpness are identified from driving images to determine the quality confidence. Then, the stability is determined based on the mean and standard deviation of the upper limb proportion, and a stability coefficient is matched. This coefficient is used to correct the quality confidence to obtain the image confidence. This automatically relaxes the judgment threshold when the image quality or calculation reliability is low, thereby reducing false alarms caused by temporary interference with the camera or fluctuations in calculation accuracy and improving the robustness of the system.

[0056] Optionally, the confidence calculation method further includes:

[0057] Step 306: When the confidence level of the image is less than a preset confidence threshold, combine the palm position and upper limb posture to generate a complete posture, and identify the posture range from the complete posture;

[0058] Step 307: Determine the multi-source type based on the attitude range, and retrieve multi-source data based on the multi-source type;

[0059] Step 308: Determine the multi-source threshold according to the multi-source type and attitude range, and compare the multi-source data with the multi-source threshold to determine multi-source consistency;

[0060] Step 309: Match multi-source weights based on the multi-source categories, and calculate weighted consistency by combining the multi-source weights and multi-source consistency;

[0061] Step 310: Match the multi-source confidence based on the weighted consistency and update the image confidence in response to the multi-source confidence.

[0062] By adopting the above technical solution, and through cross-validation of multi-source data such as steering wheel torque sensor and capacitance sensor, when the image confidence is lower than the confidence threshold, the hand position and upper limb posture are combined to generate a complete posture and identify the posture range. Based on this, the types of multi-source sensors are determined and the corresponding data are retrieved. After calculating the consistency between the multi-source data and the threshold, the weights are matched to obtain the weighted consistency. Finally, the multi-source confidence is output to correct the image confidence, thereby making up for the deficiency of insufficient confidence in the visual channel, and thus ensuring that an accurate driver state assessment can still be obtained when visual inference is unreliable.

[0063] Optionally, the confidence calculation method further includes:

[0064] Step 311: If the distance between the palms is less than a preset operation threshold, determine the occlusion feature based on the palm position;

[0065] Step 312: When the occlusion feature is not empty, determine the occlusion ratio from the driving image based on the occlusion feature, and determine the confidence weight according to the occlusion ratio;

[0066] Step 313: When the occlusion feature is empty, determine the confidence weight based on the image confidence.

[0067] Step 314: Determine a weighted confidence level by combining the multi-source confidence level, image confidence level, and confidence level weights, and update the image confidence level in response to the weighted confidence level.

[0068] By adopting the above technical solution, occlusion features are determined based on the palm position. When the occlusion features are not empty, the confidence weight is determined according to the occlusion ratio. The difficulty of inference is quantified by identifying the palm occlusion ratio, thereby automatically reducing the visual inference weight when the occlusion is severe. When the occlusion features are empty, the confidence weight is determined according to the image confidence, thereby prioritizing the trust of the visual detection results when there is no occlusion. Finally, the dynamic adaptive allocation of confidence weight is realized. The weighted confidence is calculated by combining multi-source confidence, image confidence, and confidence weight, and the image confidence is updated, thus constructing a complete closed loop from cause identification to weight allocation and then to confidence update.

[0069] Secondly, this application provides an intelligent driving assistance safety management system, which adopts the following technical solution:

[0070] An intelligent driving assistance safety management system includes:

[0071] The acquisition module is used to acquire driving images;

[0072] A memory for storing the program of any of the above-mentioned intelligent driving assistance safety management methods;

[0073] The processor is the unit of memory that allows programs to be loaded and executed by the processor.

[0074] By adopting the above technical solution, driving images are acquired and upper limb posture is identified to determine hand features. When the hand features are missing, the forearm length is estimated based on the upper arm length. Then, the hand position is determined by combining the joint orientation and joint position, and the hand distance is calculated. When the hand distance is continuously less than the operation threshold, an attention warning is issued. Thus, when the hand features are missing, the geometric constraints of the upper limb kinematic chain are used to estimate the hand position, thereby realizing continuous monitoring of the driver's attention and improving the safety of intelligent driving.

[0075] In summary, this application includes at least one of the following beneficial technical effects:

[0076] 1. Acquire driving images and identify upper limb posture to determine hand features. When hand features are missing, calculate the forearm length based on the upper arm length, and then determine the hand position and calculate the hand distance by combining joint orientation and joint position. When the hand distance is continuously less than the operation threshold, issue an attention warning. In this way, when hand features are missing, the geometric constraints of the upper limb kinematic chain are used to calculate the hand position, realize continuous monitoring of driver attention, and improve the safety of intelligent driving.

[0077] 2. Identify joint features from upper limb posture, determine the visual length and angle of the upper arm based on the joint features, establish the mapping relationship between the upper arm angle and visual distortion using the principle of camera perspective projection, determine the visual coefficient, and then compensate for the length measurement error caused by the change of arm orientation in the two-dimensional image through the visual coefficient to obtain the true upper arm length, thereby reducing the calculation deviation caused by the driver's arm extending in different directions and improving the accuracy of subsequent palm position calculation.

[0078] 3. A two-stage mechanism of establishing a standard under normal conditions and calling it in abnormal conditions is adopted. When the palm features are visible, the visual lengths of the upper arm, forearm, and palm are collected in advance. The true length is calculated and the upper limb proportion is established by combining the compensation coefficient determined by the joint angle comparison. When the palm features are missing, the proportion is called to estimate the forearm length. The length distortion caused by the visual projection is compensated by the joint angle, thereby realizing personalized and high-precision estimation of the forearm length and reducing the individual difference error between different drivers caused by using fixed experience proportions. Attached Figure Description

[0079] Figure 1 This is a scenario illustration of an intelligent driving assistance safety management method;

[0080] Figure 2 A flowchart of an intelligent driving assistance safety management method;

[0081] Figure 3 This is a flowchart of the attention detection method. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0083] This application discloses an intelligent driving assistance safety management method.

[0084] Reference Figure 1 and Figure 2 A method for intelligent driving assistance safety management, comprising:

[0085] Step 100: Acquire driving images in response to preset intelligent driving commands.

[0086] Intelligent driving commands are control signals issued when the intelligent driving system is activated or when specific triggering conditions are met, used to initiate driver status monitoring. They can be sent from the intelligent driving domain controller to the image acquisition unit via the vehicle CAN bus.

[0087] Driving images refer to raw image data of the driver's upper limbs and hands, captured by an in-vehicle camera. These images are taken in real time at a rate of 30 frames per second by a near-infrared monocular camera located above the dashboard or near the steering column, and transmitted to an image signal processor for preprocessing via a MIPI interface.

[0088] Step 101: Identify the upper limb posture from the driving image and determine the hand features based on the upper limb posture.

[0089] Upper limb posture refers to the spatial position and angle information of key points such as the driver's shoulder joint, elbow joint and wrist, extracted from driving images by a posture estimation algorithm. A lightweight posture estimation network is used to regress the heat map of the image pixel by pixel, and the two-dimensional coordinates and detection confidence of each joint point are decoded from the peak values ​​of the heat map.

[0090] Hand features refer to the determination of whether there is a clear and identifiable hand outline or key points based on the wrist key points and their extended areas in the upper limb posture using a hand key point detection model. If the confidence of the hand key points output by the model is lower than the preset threshold or no hand detection box is detected, the hand features are determined to be empty.

[0091] Step 102: When the palm feature is empty, identify the upper arm length from the upper limb posture, and determine the forearm length based on the upper arm length.

[0092] An empty palm feature indicates that the driver's hand was not detected, meaning that the driver's hand is not present in the image. In this case, the system has difficulty determining whether the driver is focused on driving.

[0093] Upper arm length refers to the actual three-dimensional spatial distance from the driver's shoulder joint to the elbow joint. When the palm features are empty, the two-dimensional image coordinates of the shoulder joint and elbow joint are extracted from the upper limb posture. The two-dimensional coordinates are mapped to the three-dimensional spatial coordinates under the camera coordinate system by combining the camera intrinsic parameters and depth information. The Euclidean distance between the two joints is calculated as the upper arm length. The detailed process of upper arm length is shown in steps 106 to 109 below.

[0094] Forearm length refers to the actual three-dimensional spatial distance from the driver's elbow joint to the wrist joint. Based on the statistical ratio between the forearm and upper arm in human upper limb kinematics, the upper arm length is multiplied by a preset statistical ratio coefficient to obtain the estimated value of the forearm length. The detailed process of forearm length estimation is shown in steps 110 to 114 below.

[0095] Step 103: Identify the joint orientation and joint position from the upper limb posture, and determine the palm position by combining the joint position, joint orientation and forearm length, and determine the palm spacing by comparing the palm position.

[0096] Joint orientation refers to the vector direction from one joint to an adjacent joint, specifically the three-dimensional direction vectors from the shoulder joint to the elbow joint and from the elbow joint to the wrist joint. This can be obtained by normalizing the difference between the spatial coordinates of the two joints. Joint position refers to the three-dimensional spatial coordinates of the shoulder and elbow joints in the camera coordinate system.

[0097] The palm position refers to the three-dimensional spatial coordinates of the center point of the palm in the camera coordinate system. The spatial position of the wrist joint is calculated with the elbow joint position as the origin and the forearm length as the modulus along the direction of the elbow joint vector. The center position of the palm is then determined by combining the fixed offset from the wrist to the palm.

[0098] The distance between the palms refers to the three-dimensional Euclidean distance between the positions of the left and right palms. The distance between the two hands is obtained by subtracting the calculated position of the left palm from the position of the right palm and then calculating the vector magnitude.

[0099] Step 104: If the distance between the palms is less than a preset operation threshold, determine the duration based on the distance between the palms.

[0100] The operating threshold is a critical value that characterizes the minimum distance between a driver's hands when holding the steering wheel normally. For example, it can be set to 0.15 meters. If the current distance between the palms is less than this threshold, it means that the hands are in an abnormal driving posture that is too close together.

[0101] Duration refers to the cumulative length of time during which the distance between the palms is continuously less than the operation threshold. When the system detects that the distance between the palms is less than the operation threshold for the first time, it starts a timer and records the total duration of the distance being less than the threshold in milliseconds. If the distance recovers to above the threshold during this period, the timer is reset to zero.

[0102] Step 105: Generate and display an attention warning in response to the duration.

[0103] Attention warning refers to visual and auditory warning signals used to prompt drivers to return their hands to a normal driving posture. When the duration exceeds a preset time threshold, a flashing attention warning icon is generated on the in-vehicle display screen, and a voice prompt is issued in conjunction with the voice synthesis module to remind the driver to keep their hands in a normal grip position.

[0104] The system acquires driving images and identifies upper limb postures to determine hand features. When hand features are missing, the system calculates the forearm length based on the upper arm length, and then determines the hand position and calculates the hand-palm distance by combining joint orientation and joint position. When the hand-palm distance is consistently less than the operation threshold, an attention warning is issued. Thus, when hand features are missing, the system uses the geometric constraints of the upper limb kinematic chain to calculate the hand position, enabling continuous monitoring of driver attention and improving the safety of intelligent driving.

[0105] Methods for determining upper arm length include:

[0106] Step 106: Identify joint features from the upper limb posture.

[0107] Joint features refer to the position coordinates and detection confidence scores of the shoulder and elbow joints extracted from the upper limb posture. The index positions corresponding to the shoulder and elbow joints are selected from the set of joint points output by the posture estimation network, and their two-dimensional image coordinates and confidence scores are extracted as joint feature data.

[0108] Step 107: Determine the visual length and angle of the upper arm from the upper limb posture based on the joint features.

[0109] The visual length of the upper arm refers to the pixel distance from the shoulder joint to the elbow joint in a two-dimensional image plane. The visual length in pixels is calculated by substituting the image coordinates of the shoulder joint and elbow joint into the Euclidean distance formula.

[0110] The upper arm angle refers to the deflection angle of the upper arm relative to the optical axis of the camera in three-dimensional space. The upper arm angle is obtained by calculating the spatial angle between the upper arm vector and the optical axis of the camera. The larger the angle, the more the arm is deviating from the positive direction of the camera.

[0111] Step 108: Determine the visual coefficient based on the upper arm angle.

[0112] The visual coefficient is a correction factor used to compensate for the distortion of the perspective projection length caused by the change in the angle between the arm orientation and the camera's optical axis. According to the principle of camera perspective projection, when there is an angle between the upper arm and the optical axis, the projection length of the upper arm on the image plane will be shortened according to the cosine value of the angle. The visual coefficient is obtained by substituting the upper arm angle into the pre-calibrated visual coefficient mapping function, which is the reciprocal of the cosine value of the angle.

[0113] Step 109: Determine the upper arm length by combining the upper arm visual length and visual coefficient.

[0114] Multiply the visual length of the upper arm by a visual coefficient to compensate for the shortening caused by perspective projection and restore the true length of the upper arm in three-dimensional space.

[0115] By identifying joint features from upper limb posture, determining the visual length and angle of the upper arm based on these features, and establishing a mapping relationship between the upper arm angle and visual distortion using the principle of camera perspective projection, visual coefficients are determined. Then, the visual coefficients are used to compensate for the length measurement error caused by changes in the direction of the arm in the two-dimensional image to obtain the true upper arm length. This reduces the estimation deviation caused by the driver's arm extending in different directions and improves the accuracy of subsequent hand position estimation.

[0116] Methods for determining forearm length include:

[0117] Step 110: When the palm feature is not empty, determine the visual length of the forearm and the visual length of the palm from the upper limb posture based on the joint feature and the palm feature.

[0118] Forearm visual length refers to the pixel distance from the elbow joint to the wrist joint in a two-dimensional image plane. The visual length in pixels is calculated by substituting the image coordinates of the elbow and wrist joints into the Euclidean distance formula.

[0119] The visual length of the palm refers to the pixel distance from the wrist joint to the center of the palm in a two-dimensional image plane. The visual length in pixels is calculated by substituting the image coordinates of the wrist joint and the center of the palm into the Euclidean distance formula.

[0120] Step 111: Identify the upper arm angle and forearm angle from the upper limb posture based on the joint features, and determine the compensation coefficient by comparing the upper arm angle and forearm angle.

[0121] The upper arm angle is the angle between the upper arm vector and the camera optical axis, and the forearm angle is the angle between the forearm vector and the camera optical axis. The angles are calculated by extracting the direction vectors of the upper arm and forearm from the upper limb posture and then comparing them with the camera optical axis.

[0122] The compensation coefficient is a correction factor used to correct for forearm visual projection distortion caused by the difference in spatial orientation between the forearm and upper arm. For example, when the forearm angle is greater than the upper arm angle, the forearm projection is shortened more severely, and the compensation coefficient is increased accordingly. The difference between the forearm angle and the upper arm angle is calculated as the relative distortion angle. Then, the relative distortion degree corresponding to the relative distortion angle is looked up from the distortion correspondence table as the compensation coefficient. The distortion correspondence table is a data table that records different relative distortion angles and their corresponding relative distortion degrees.

[0123] The distortion correspondence table is created by placing a standard length calibration object (such as a rigid straight rod with equidistant markers) in the driver's seat and collecting upper limb posture data. When the forearm and the camera optical axis are at different angles (relative distortion angles), the projection length of the same physical length in the two-dimensional image is reduced relative to the actual length in three-dimensional space. The reciprocal of this ratio is used as the degree of relative distortion at that relative distortion angle. Multiple sets of relative distortion angles and their corresponding degrees of relative distortion are recorded as a distortion correspondence table. In practical use, the compensation coefficient for compensating visual projection distortion can be obtained by looking up this table based on the difference between the forearm angle and the upper arm angle.

[0124] The calibration processes for other data tables employ differentiated calibration methods due to their varying physical meanings. Specifically, the risk correspondence table and dependency correspondence table collect driver cognitive load scores under different tracking differences and approach deceleration conditions through simulated driving experiments and fit mapping relationships. The speed limit mapping table presets the upper speed limits corresponding to different risk levels based on traffic regulations and ADAS safety standards. The dependency coefficient mapping table calibrates the amplification factor of different visual dependencies on risk levels through experiments. The stability mapping table, image coefficient mapping table, and multi-source threshold parameter table determine the optimal value range for each parameter based on statistical analysis of real-vehicle test data. The contribution correspondence table assigns weights based on expert experience according to the correlation between the physical characteristics of each sensor and hand state. The multi-source confidence table, area weight mapping table, and mass weight mapping table are trained and fitted with confidence mapping relationships using a large amount of visual and sensor data collected from real-vehicle operations across multiple scenarios. All these data tables collectively constitute a multi-level decision support system combining data-driven and experience-driven approaches.

[0125] Step 112: Determine the forearm compensation length by combining the forearm visual length and compensation coefficient, and determine the palm compensation length by combining the palm visual length and compensation coefficient.

[0126] The forearm compensation length refers to the true length of the forearm after angular distortion compensation. The true spatial distance from the elbow joint to the wrist joint is obtained by multiplying the visual length of the forearm by the compensation coefficient.

[0127] The palm compensation length refers to the true length of the palm after angular distortion compensation. The true spatial distance from the wrist joint to the center of the palm is obtained by multiplying the visual length of the palm by the compensation coefficient.

[0128] Step 113: Determine the upper limb proportion by comparing the visual length of the upper arm, the compensated length of the forearm, and the compensated length of the palm.

[0129] Upper limb proportion refers to the ratio between the length of the upper arm, the length of the forearm, and the length of the palm. The ratio vector obtained after normalizing the visual length of the upper arm, the compensated length of the forearm, and the compensated length of the palm is used as the upper limb proportion. This proportion reflects the relative length relationship of each segment of the upper limb unique to the driver.

[0130] Step 114: When the palm feature is empty, determine the forearm ratio based on the upper limb proportion, and determine the forearm length by combining the forearm ratio and the upper arm length.

[0131] The forearm ratio refers to the ratio of the forearm length to the upper arm length in the upper limb proportion. The forearm ratio is obtained by extracting the forearm length from the upper limb proportion and dividing it by the upper arm length. The forearm length is then obtained by multiplying the currently calculated upper arm length by the forearm ratio.

[0132] A two-stage mechanism is adopted, which establishes the standard under normal conditions and calls it in abnormal conditions. When the palm features are visible, the visual lengths of the upper arm, forearm, and palm are collected in advance. The true length is calculated and the upper limb proportion is established by combining the compensation coefficient determined by the joint angle comparison. When the palm features are missing, the proportion is called to estimate the forearm length. The joint angle is used to compensate for the length distortion caused by the visual projection, thereby realizing personalized and high-precision estimation of the forearm length and reducing the individual difference error between different drivers caused by using fixed experience proportions.

[0133] Reference Figure 3 Attention detection methods include:

[0134] Step 200: If the distance between the palms is less than a preset operation threshold, determine the distance fluctuation curve based on the distance between the palms, and extract the distance rise curve from the distance fluctuation curve.

[0135] The spacing fluctuation curve refers to a continuously changing curve plotted with time as the horizontal axis and the palm spacing value as the vertical axis. The system continuously collects palm spacing data for each frame and stores it in time order to form a time series array as the spacing fluctuation curve.

[0136] The spacing rise curve refers to the curve segment of the palm spacing that is continuously increasing, which is extracted from the spacing fluctuation curve. The spacing rise curve can be extracted by performing a first-order difference operation on the spacing fluctuation curve and marking the intervals where the difference is positive and continuously increasing.

[0137] Step 201: Determine the rising distance based on the aforementioned spacing rise curve.

[0138] The rise distance refers to the increment of the palm spacing from the trough to the current value in the spacing rise curve. The palm spacing value at the starting point of the spacing rise curve is read as the starting spacing, the palm spacing value of the current frame is read as the current spacing, and the difference between the current spacing and the starting spacing is calculated as the rise distance.

[0139] Step 202: When the rising distance is greater than the preset distance threshold, determine the distance departure time based on the rising distance, and read the initial distance from the spacing rise curve according to the rising distance.

[0140] The distance threshold is the minimum displacement required to determine whether the action of moving both hands away from a gathered state is valid; for example, it can be set to 0.05 meters.

[0141] The moment of distance refers to the starting point at which the distance between the palms begins to move outward, that is, the moment corresponding to the starting point of the distance rise curve.

[0142] The initial distance refers to the palm spacing value corresponding to the moment of departure, that is, the palm spacing value at the starting point of the spacing rise curve.

[0143] Step 203: Read the recovery time that is located after the distance-away time and equal to the initial distance from the distance fluctuation curve, and compare the distance-away time and the recovery time to determine the recovery duration.

[0144] The recovery time refers to the moment when the distance between the palms returns from a distant state to the same distance as the initial distance or within the preset tolerance range. The system searches the distance fluctuation curve from the distant time to the first position where the distance is equal to the initial distance or the deviation is less than the tolerance threshold, and records the time corresponding to that position as the recovery time.

[0145] Recovery time refers to the time difference between the time of departure and the time of recovery. Subtracting the time of departure from the time of recovery gives the recovery time in seconds.

[0146] Step 204: If the recovery time is less than the preset attention threshold, generate and display an attention warning in response to the duration.

[0147] The attention threshold is a critical recovery time value used to distinguish whether a driver is excessively focused on a local operating area; for example, it can be set to 0.8 seconds. If the recovery time is less than the attention threshold, it means that the driver has completed the operation cycle of leaving and returning to the operating area in a very short time. The system determines that the driver has distracted behavior of excessively focusing on the operating object near their palm and generates and displays an attention warning.

[0148] The judgment dimension is expanded from static spatial geometry to dynamic temporal behavioral patterns. When the palm spacing is too small, the rising curve in the spacing fluctuation curve is extracted to represent how quickly the hand returns after leaving. When the rising distance is greater than the distance threshold, the distance departure time and the initial distance are recorded. The recovery time is then read from the fluctuation curve as the micro-behavioral feature of the driver's high-frequency back-and-forth operation in the palm area. The attention state is judged by calculating the recovery time, thereby accurately identifying the dangerous distraction state of the driver's excessive focus on the local operation area.

[0149] Attention detection methods also include:

[0150] Step 205: If the recovery time is less than the preset attention threshold, identify the set of visual focal points from the driving image based on the distance time, and determine the operation center according to the palm position.

[0151] The visual focal point set refers to the sequence of gaze points in the scene corresponding to the driver's gaze in the driving image. The system retrieves multiple consecutive frames of images from the driving image that are far away from the time and extracts the focal point position of the driver's gaze in the scene in each frame through the DMS gaze tracking algorithm, forming a focal point coordinate sequence containing timestamps as the visual focal point set.

[0152] The operating center refers to the center point of the driver's hand operating focus area, that is, the midpoint between the two palm positions.

[0153] Step 206: Compare the set of visual focal points with the operation center to determine the visual spacing.

[0154] Visual distance refers to the spatial distance between the visual focal point and the operation center at a certain moment. The coordinates of the focal point at each moment are extracted from the set of visual focal points, and the two-dimensional Euclidean distance between the focal point and the coordinates of the operation center is calculated, thereby representing the degree of spatial deviation between the line of sight and the hand operation area.

[0155] Step 207: When the visual distance is less than the preset tracking threshold, determine the tracking time based on the visual distance, and compare the tracking time and the recovery time to determine the tracking difference.

[0156] The tracking threshold is a critical distance used to determine whether the line of sight has been tracked to the operating area, for example, it is set to 0.05 meters.

[0157] The tracking moment refers to the moment when the visual distance first falls below the tracking threshold. The system iterates through the visual distance sequence in chronological order and marks the timestamp corresponding to the first distance that is less than the tracking threshold as the tracking moment.

[0158] The tracking difference is the time difference between the tracking time and the recovery time. The tracking difference is obtained by subtracting the recovery time from the tracking time.

[0159] Step 208: Determine the risk level based on the tracking difference, and match the speed limit according to the risk level.

[0160] The risk level refers to the degree to which the driver's current attention deviates from the road, which is quantitatively assessed based on the temporal correlation characteristics between the driver's hand movements and eye tracking behavior. A positive tracking difference value indicates that the driver's hand moves for a period of time before moving their eyes, that is, the driver's tactile failure triggers an emergency visual intervention. At this time, the larger the tracking difference value, the more cognitive resources the driver is forced to suddenly withdraw from the driving task and devote to visual search due to the tactile feedback not meeting expectations. The greater the degree to which the driver's current attention deviates from the road, the greater the risk level. A non-positive tracking difference value indicates that there is no tactile failure, and the risk level is 0. The risk level corresponding to the tracking difference value can be found in the risk correspondence table, which is a data table that records different tracking differences and their corresponding risk levels.

[0161] Speed ​​limits refer to the maximum permissible speed of a vehicle matched to the level of risk. The higher the level of risk, the more serious the driver's inattention is, and the more necessary it is to reduce the speed to reduce traffic accidents. The speed limit map table can be used to find the upper limit of the speed corresponding to the level of risk. The speed limit map table is a data table that records different levels of risk and their corresponding speed limits.

[0162] Step 209: In response to the speed limit generation, a speed limit command is sent.

[0163] The speed limit command is a control signal used to limit the maximum speed of a vehicle to the limit speed. This command is generated by the intelligent driving domain controller and sent to the vehicle controller and engine controller or motor controller via the CAN bus to execute the speed limit.

[0164] By performing cross-modal temporal pairing analysis of gaze and hand behavior, the set of visual landing points is identified when the recovery time is too short. The operation center is determined based on the palm position, and the visual distance between the two is calculated. When the visual distance is less than the tracking threshold, the tracking time is recorded and compared with the recovery time to obtain the tracking difference as the phase difference between the hand recovery time and the gaze tracking time. The tracking difference is used as a quantitative indicator of cognitive load, thereby determining the risk level and matching the vehicle speed limit, thus achieving a refined graded assessment of attention state.

[0165] Attention detection methods also include:

[0166] Step 210: When the tracking difference falls into the preset visual range, determine the approach time from the visual landing point set based on the tracking time, and retrieve the approach distance by combining the approach time and the tracking time.

[0167] The visual interval refers to a specific numerical range within which the tracking difference falls. It is used to characterize the dangerous behavior pattern of a driver tracking their gaze to the operating area within a very short time after their hands are brought back together. For example, it can be set to 0 to 0.3 seconds.

[0168] The proximity moment refers to the moment when the visual focus point begins to move closer to the operation center. Based on the tracking time, the moment corresponding to the first continuously monotonically decreasing starting frame in the visual spacing sequence and the cumulative frame count meeting the preset counting threshold is taken as the proximity moment.

[0169] Proximity spacing refers to the visual distance value corresponding to the moment of proximity. The proximity spacing is obtained by reading the corresponding distance value from the visual distance sequence at the moment of proximity.

[0170] Step 211: Generate a proximity curve based on the proximity spacing, and identify the proximity rate from the proximity curve.

[0171] The spacing approach curve refers to a curve segment with time as the horizontal axis and approach spacing as the vertical axis, which captures the process of visual spacing gradually decreasing from the approach moment to the tracking moment.

[0172] The approach rate refers to the slope of the visual distance on the approach curve as a function of time, i.e., the first derivative. The rate value sequence at each moment is obtained by performing a differential operation on the approach curve.

[0173] Step 212: Analyze the approach rate to determine the approach deceleration, and determine the visual dependence based on the approach deceleration.

[0174] Approach deceleration refers to the slope of the approach rate as a function of time, which is the second derivative obtained by differentiating the first derivative with respect to time. The deceleration value at each moment is obtained by differentiating the approach rate sequence.

[0175] Visual dependence refers to the degree of driver reliance on visual confirmation, which is quantified based on proximity deceleration. The larger the absolute value of proximity deceleration, the higher the driver's line of sight approaches the operating area with the higher acceleration, reflecting the stronger the driver's urgent need for visual confirmation, and thus the greater the visual dependence. The visual dependence corresponding to proximity deceleration can be found in the dependence correspondence table, which is a data table that records different proximity decelerations and their corresponding visual dependence.

[0176] Step 213: Match the dependency coefficient according to the visual dependency, and update the risk level according to the dependency coefficient.

[0177] The dependency coefficient is a weighting factor used to correct the degree of risk based on visual dependency. The higher the visual dependency, the more cognitive resources the visual search occupies, and the higher the degree of risk, requiring a larger dependency coefficient. The dependency coefficient corresponding to the current visual dependency can be found in the dependency coefficient mapping table, and the dependency coefficient is multiplied by the current degree of risk to obtain the updated degree of risk, so that the degree of risk more accurately reflects the driver's actual cognitive load. The dependency coefficient mapping table is a data table that records different visual dependencies and their corresponding dependency coefficients.

[0178] By introducing the second derivative of kinematics as an analytical dimension of cognitive state, the approach time is determined when the tracking difference falls into the visual interval, and the approach distance is retrieved. After generating the distance approach curve, the approach rate is identified from it, and then the approach deceleration is analyzed to determine the visual dependence. The driver's dependence on visual confirmation is quantitatively characterized by the urgency of the gaze approaching the operating area. Finally, the dependence coefficient is matched to update the risk level, thereby establishing a quantitative inference chain from behavioral kinematic parameters to cognitive state, providing a deeper cognitive behavioral indicator for attention risk assessment.

[0179] Confidence calculation methods include:

[0180] Step 300: If the distance between the palms is less than a preset operation threshold, identify the image brightness, image contrast, local variance and global variance from the driving image, and determine the image sharpness by combining the local variance and global variance.

[0181] Image brightness refers to the statistical mean of the brightness values ​​of all pixels in a driving image. It is obtained by averaging the grayscale values ​​of all pixels after converting the image to grayscale.

[0182] Image contrast refers to the standard deviation of pixel brightness values ​​in a driving image. It is calculated by dividing the sum of the squares of the deviations of all pixel grayscale values ​​from the mean by the square root of the total number of pixels.

[0183] Local variance refers to the variance of pixel brightness values ​​within each window after dividing the driving image into multiple local windows. The mean of the variances of all windows is taken as the representative value of local variance after calculating the variance of pixels within each window.

[0184] Global variance refers to the variance of the brightness values ​​of all pixels in the entire image.

[0185] Image sharpness refers to the degree to which image details are discernible, based on the ratio of local variance to global variance. The ratio of local variance to global variance is calculated as the sharpness score.

[0186] Step 301: Determine the brightness confidence level based on the image brightness, and determine the contrast confidence level based on the image contrast.

[0187] Brightness confidence is a reliability index that assesses the degree to which an image brightness deviates from the ideal brightness range. It compares the image brightness with the preset lower and upper limits of the ideal brightness range. If the brightness is within the range, the confidence is 1. If it is below the lower limit, it is reduced to 0 linearly. If it is above the upper limit, it is also reduced to 0 proportionally.

[0188] Contrast confidence is a reliability metric based on the magnitude of image contrast. It compares the image contrast with a preset ideal lower limit. If the contrast is higher than the lower limit, the confidence is 1; if it is lower than the lower limit, the confidence is reduced to 0 proportionally.

[0189] Step 302: Determine the image quality confidence level by combining the image sharpness, brightness confidence level and contrast confidence level, and determine the proportion mean and proportion standard deviation based on the upper limb proportion.

[0190] Quality confidence score refers to the overall image quality reliability score obtained by combining three dimensions: sharpness, brightness, and contrast. It is calculated by multiplying image sharpness by a first weight, brightness confidence score by a second weight, and contrast confidence score by a third weight, and then summing the results. The first, second, and third weights are determined based on each dimension's ability to represent whether the image is suitable for upper limb posture recognition and hand estimation. For example, image sharpness is assigned the highest weight because it directly affects the accuracy of upper limb joint positioning; brightness confidence score is assigned a middle weight because vehicles equipped with automatic supplemental lighting systems (such as infrared LED supplemental lighting) have relatively less influence on detection due to ambient brightness; and contrast confidence score is assigned the lowest weight because modern vehicle cameras have a wide dynamic range and the hand area is a naturally high-contrast scene with skin color and dark interior, having the weakest impact on recognition results. These three weights can be set to 0.45, 0.30, and 0.25 respectively, while satisfying the normalization constraint that the sum of the three is 1.

[0191] The proportion mean refers to the statistical average of limb proportion data in history. The system continuously records the upper limb proportion value when the palm feature is detectable. After accumulating enough samples, the arithmetic mean of all historical proportion values ​​is calculated as the proportion mean.

[0192] The standard deviation of proportion refers to the dispersion of limb proportion statistics throughout history. It is calculated as the standard deviation of the deviation of all historical proportion values ​​from the proportion mean.

[0193] Step 303: Determine the stability of the proportion by combining the mean and standard deviation of the proportion, and match the stability coefficient based on the stability of the proportion.

[0194] Proportional stability refers to the degree of stability of the upper limb proportion over time, as reflected by the ratio of the standard deviation of the proportion to the mean of the proportion. The stability value is obtained by dividing the standard deviation of the proportion by the mean of the proportion. The smaller the stability value, the more stable the upper limb proportion.

[0195] The stability coefficient is a weighting factor used to correct the quality confidence based on the proportional stability. The larger the proportional stability, the more reliable the upper limb proportion, and the higher the image quality. The larger the stability coefficient, the more reliable the upper limb proportion. The stability coefficient corresponding to the current proportional stability can be found in the stability mapping table. The stability mapping table is a data table that records different proportional stability and their corresponding stability coefficients.

[0196] Step 304: Correct the quality confidence score according to the stability coefficient to obtain the image confidence score, and match the confidence score coefficient according to the image confidence score.

[0197] Image confidence score is a comprehensive score of the final image quality and the reliability of the inference obtained after correction by the stability coefficient. The image confidence score is obtained by multiplying the quality confidence score by the stability coefficient. If the upper limb proportion is stable, the image confidence score is increased; if it is unstable, the image confidence score is decreased.

[0198] The confidence coefficient is an adjustment factor used to adjust the duration based on the image confidence level. The lower the image confidence level, the less reliable the image is, and the less accurate the hand spacing calculated from the image will be. In this case, the confidence coefficient should be closer to 0 to reduce the duration and thus reduce false alarms. The confidence coefficient corresponding to the current image confidence level can be found in the image coefficient mapping table, and then the product of the original duration and the confidence coefficient can be calculated as the new duration. The image coefficient mapping table is a data table that records different image confidence levels and their corresponding confidence coefficients.

[0199] Step 305: Respond to the confidence coefficient update duration to reduce false alarms.

[0200] A confidence assessment system is constructed from the dual perspectives of image quality stability and upper limb proportion stability. Brightness, contrast, and sharpness are identified from driving images to determine the quality confidence. Then, stability is determined based on the mean and standard deviation of the upper limb proportion, and a stability coefficient is matched. This coefficient is used to correct the quality confidence to obtain the image confidence. This automatically relaxes the judgment threshold when the image quality or inference reliability is low, thereby reducing false alarms caused by temporary interference with the camera or fluctuations in inference accuracy and improving the robustness of the system.

[0201] Confidence calculation methods also include:

[0202] Step 306: When the confidence level of the image is less than a preset confidence threshold, combine the palm position and upper limb posture to generate a complete posture, and identify the posture range from the complete posture.

[0203] The confidence threshold is a lower limit threshold used to determine whether an image's quality is sufficient to support independent judgment; for example, it can be set to 0.6.

[0204] Complete posture refers to the upper limb skeleton data that includes the coordinates of complete joint points of the shoulder, elbow, and wrist, formed by splicing the calculated palm position with the shoulder and elbow joints in the upper limb posture.

[0205] Attitude range refers to the spatial distribution range of each joint in a complete attitude. The minimum bounding box size of all joint coordinates in three-dimensional space is calculated as the attitude range.

[0206] Step 307: Determine the multi-source type based on the attitude range, and retrieve multi-source data based on the multi-source type.

[0207] Multi-source types refer to the sensor types selected based on the posture range that can effectively cover the driver's upper limb activity area. Since different sensors on a vehicle have different detection ranges—for example, steering wheel torque sensors and capacitive sensors can only detect the contact state between the hand and the steering wheel, suitable for judging the state when the upper limbs are close to the steering wheel; seat pressure sensors and seat position sensors can detect changes in the overall position of the driver's torso, suitable for judging the state when the upper limb posture shifts; while the DMS camera and the proximity sensor on the steering wheel are suitable for judging the state when the upper limbs are in the transition area between the steering wheel and the center console—the available sensor types change as the posture range expands or becomes more biased towards a specific area of ​​the cockpit. The system treats the shoulder, elbow, and wrist joints within the posture range as independent spatial sampling points. For each sensor, the system predefines the spatial description parameters of its effective detection range. The system checks each joint point individually to see if it falls within the effective detection range of each sensor. If any of the shoulder, elbow, or wrist joints falls within the effective detection range of a sensor, the sensor is considered to cover the current upper limb activity area and is defined as a valid sensor. The set of all valid sensors constitutes the multi-source type.

[0208] Multi-source data refers to the real-time measurement values ​​collected by each sensor. After determining the currently available and valid multi-source types based on the attitude range, real-time data is read from the corresponding signal channel of the vehicle bus as multi-source data.

[0209] Step 308: Determine the multi-source threshold according to the multi-source type and attitude range, and compare the multi-source data with the multi-source threshold to determine multi-source consistency.

[0210] Multi-source thresholds refer to preset judgment reference values ​​for the normal working boundaries of each sensor in the current multi-source type within its effective detection range. The system first determines the upper limb state presented by the posture range based on the spatial distribution of the shoulder joint, elbow joint and wrist joint in the upper limb posture. For example, when the wrist joint is located in the steering wheel area and the elbow joint angle is within the normal grip range, the posture range presents the state of the hand on the steering wheel. Then, for the multi-source type of sensor, the judgment boundary values ​​of different sensors in the current upper limb state are read from the multi-source threshold parameter table as multi-source thresholds. The multi-source threshold parameter table is a data table that records different multi-source types and upper limb states and their corresponding multi-source thresholds.

[0211] Multi-source consistency refers to whether the upper limb state reflected by the real-time measurement value of a single sensor, after being judged by a threshold, matches the state presented by the posture range. A value of 1 indicates consistency, and a value of 0 indicates inconsistency. For each sensor in the multi-source category, the system performs threshold judgment separately. Taking the steering wheel torque sensor as an example, the system presets the multi-source threshold for contact torque to be 0.5 Nm. The current posture range shows the hand on the steering wheel. If the real-time measurement value of the torque sensor is 1.2 Nm, which is greater than the multi-source threshold, then the state reflected by the torque sensor is that the hand is on the steering wheel, which matches the state presented by the posture range, and the multi-source consistency value of the corresponding sensor is assigned to 1. If the real-time measurement value of the torque sensor is 0.1 Nm, which is less than the multi-source threshold, then the state reflected is that the hand is not in contact with the steering wheel, which does not match the state presented by the posture range, and the multi-source consistency value of the corresponding sensor is assigned to 0.

[0212] Step 309: Match multi-source weights based on the multi-source types, and calculate weighted consistency by combining the multi-source weights and multi-source consistency.

[0213] Multi-source weight refers to the contribution ratio of different sensors in confidence correction. Considering that the steering wheel torque sensor and capacitive sensor can directly reflect whether the hand is in the steering wheel area, their signals have the highest correlation with the hand state, so they are assigned higher weights. For example, the torque sensor weight is set to 0.35 and the capacitive sensor weight is set to 0.30. The seat pressure sensor and seat position sensor reflect the overall posture of the torso and are indirectly related to the hand state, so they are assigned medium weights, for example, set to 0.20 and 0.15 respectively. The multi-source weights corresponding to the multi-source types can be found in the contribution correspondence table, which is a data table that records different multi-source types and their corresponding multi-source weights.

[0214] Weighted consistency refers to a numerical value that characterizes the consistency between multi-source sensor readings and image detection results. It can be obtained by multiplying the multi-source consistency scores of each sensor by their corresponding multi-source weights and then summing them up.

[0215] Step 310: Match the multi-source confidence based on the weighted consistency and update the image confidence in response to the multi-source confidence.

[0216] Multi-source confidence refers to the confidence level of multi-source sensor fusion based on weighted consistency mapping. The higher the weighted consistency, the more the sensor data as a whole supports the visual inference results, indicating that the inferred hand position matches the hand state detected by the sensors, and thus the higher the multi-source confidence. The lower the weighted consistency, the more it indicates a conflict between the visual inference and the sensor data, suggesting that the inference results may be unreliable, and thus the lower the multi-source confidence. The multi-source confidence corresponding to the weighted consistency can be retrieved from the multi-source confidence table, and then the multi-source confidence and image confidence are weighted and fused to obtain the new image confidence. The average of the multi-source confidence and the original image confidence can be calculated as the new image confidence. The multi-source confidence table is a data table that records different weighted consistency and their corresponding multi-source confidence.

[0217] By introducing multi-source data such as steering wheel torque sensor and capacitance sensor for cross-validation, when the image confidence is lower than the confidence threshold, the hand position and upper limb posture are combined to generate a complete posture and identify the posture range. Based on this, the types of multi-source sensors are determined and the corresponding data are retrieved. After calculating the consistency between the multi-source data and the threshold, the weights are matched to obtain the weighted consistency. Finally, the multi-source confidence is output to correct the image confidence, thereby making up for the deficiency of insufficient confidence in the visual channel, and thus ensuring that an accurate driver state assessment can still be obtained when visual inference is unreliable.

[0218] Confidence calculation methods also include:

[0219] Step 311: If the distance between the palms is less than a preset operation threshold, determine the occlusion feature based on the palm position.

[0220] Occlusion features refer to image features extracted based on the palm position and surrounding image region to determine whether the hand is physically occluded. The system extracts a local image region centered on the palm position, extracts the confidence distribution map of each key point output by the hand detection model in this region, and a binary skin color mask generated based on the skin color detection model. The confidence distribution map and the skin color mask are used as occlusion features. If the confidence distribution map shows a discontinuous pattern of alternating local high confidence and local low confidence, or if the skin color mask shows a fragmented distribution in the hand region instead of a complete palm shape, then the occlusion feature is determined to be non-empty.

[0221] Step 312: When the occlusion feature is not empty, determine the occlusion ratio from the driving image based on the occlusion feature, and determine the confidence weight according to the occlusion ratio.

[0222] Occlusion ratio refers to the proportion of the hand area that is occluded. The system counts the number of pixels within the hand detection box whose key point confidence is lower than the preset confidence threshold or whose skin color detection fails, and divides the number of pixels by the total number of pixels in the hand detection box to obtain the occlusion ratio.

[0223] Confidence weight refers to the adaptive weight used to adjust the proportion of image confidence and multi-source confidence in the final confidence synthesis. When the occlusion ratio is larger, the reliability of visual inference is lower. The weight of image confidence in the final synthesis should be reduced and the weight of multi-source confidence should be increased accordingly. For example, when the occlusion ratio is less than 0.2, the confidence weight is 0.85; when the occlusion ratio is between 0.2 and 0.6, the confidence weight is 0.5; and when the occlusion ratio is greater than 0.6, the confidence weight is 0.2. The confidence weight corresponding to the current occlusion ratio can be found in the area weight mapping table, which is a data table that records different occlusion ratios and their corresponding confidence weights.

[0224] Step 313: When the occlusion feature is empty, determine the confidence weight based on the image confidence.

[0225] When the occlusion feature is empty, it means that there is no physical occlusion of the hand, and the visual inference result itself has high reliability. The confidence weight can be directly matched according to the image confidence: the higher the image confidence, the higher the image quality and upper limb inference reliability of the current frame, and the more trustworthy the visual inference result is. The image confidence should play a dominant role in the final confidence synthesis, so the confidence weight should be increased accordingly. Conversely, if the image confidence is low, the weight should be appropriately reduced. The confidence weight corresponding to the current image confidence can be found from the quality weight mapping table. The quality weight mapping table is a data table that records different image confidence and their corresponding confidence weights.

[0226] Step 314: Determine a weighted confidence level by combining the multi-source confidence level, image confidence level, and confidence level weights, and update the image confidence level in response to the weighted confidence level.

[0227] Weighted confidence score is the final confidence score obtained by combining image confidence score and multi-source confidence score according to confidence score weights. It can be calculated according to the formula: Weighted confidence score = Image confidence score * Confidence score weight + Multi-source confidence score * (1 - Confidence score weight), and the weighted confidence score is used as the new image confidence score.

[0228] Occlusion features are determined based on the palm position. When the occlusion features are not empty, the confidence weight is determined according to the occlusion ratio. The difficulty of inference is quantified by identifying the palm occlusion ratio, so that the visual inference weight is automatically reduced when the occlusion is severe. When the occlusion features are empty, the confidence weight is determined according to the image confidence, so that the visual detection result is trusted first when there is no occlusion. Finally, the dynamic adaptive allocation of confidence weight is realized. The weighted confidence is calculated by combining multi-source confidence, image confidence and confidence weight and the image confidence is updated, thus constructing a complete closed loop from cause identification to weight allocation and then to confidence update.

[0229] Based on the same inventive concept, embodiments of the present invention provide an intelligent driving assistance safety management system, comprising:

[0230] The acquisition module is used to acquire driving images;

[0231] A memory for storing the program of any of the above-mentioned intelligent driving assistance safety management methods;

[0232] The processor is the unit of memory that allows programs to be loaded and executed by the processor.

[0233] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0234] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent driving assistance safety management, characterized in that, include: Step 100: Acquire driving images in response to preset intelligent driving commands; Step 101: Identify the upper limb posture from the driving image and determine the hand features based on the upper limb posture; Step 102: When the palm feature is empty, identify the upper arm length from the upper limb posture, and determine the forearm length based on the upper arm length; Step 103: Identify the joint orientation and joint position from the upper limb posture, and determine the palm position by combining the joint position, joint orientation and forearm length, and determine the palm spacing by comparing the palm position; Step 104: If the distance between the palms is less than a preset operation threshold, determine the duration based on the distance between the palms; Step 105: Generate and display an attention warning in response to the duration.

2. The intelligent driving assistance safety management method according to claim 1, characterized in that, The method for determining the upper arm length includes: Step 106: Identify joint features from the upper limb posture; Step 107: Determine the visual length and angle of the upper arm from the upper limb posture based on the joint features; Step 108: Determine the visual coefficient based on the upper arm angle; Step 109: Determine the upper arm length by combining the upper arm visual length and visual coefficient.

3. The intelligent driving assistance safety management method according to claim 2, characterized in that, The method for determining the forearm length includes: Step 110: When the palm feature is not empty, determine the visual length of the forearm and the visual length of the palm from the upper limb posture based on the joint feature and the palm feature; Step 111: Identify the upper arm angle and forearm angle from the upper limb posture based on the joint features, and determine the compensation coefficient by comparing the upper arm angle and forearm angle; Step 112: Determine the forearm compensation length by combining the forearm visual length and compensation coefficient, and determine the palm compensation length by combining the palm visual length and compensation coefficient; Step 113: Determine the upper limb proportion by comparing the visual length of the upper arm, the compensated length of the forearm, and the compensated length of the palm; Step 114: When the palm feature is empty, determine the forearm ratio based on the upper limb proportion, and determine the forearm length by combining the forearm ratio and the upper arm length.

4. The intelligent driving assistance safety management method according to claim 1, characterized in that, It also includes an attention detection method, which includes: Step 200: If the distance between the palms is less than a preset operation threshold, determine the distance fluctuation curve based on the distance between the palms, and extract the distance rise curve from the distance fluctuation curve; Step 201: Determine the rising distance based on the aforementioned spacing rise curve; Step 202: When the rising distance is greater than the preset distance threshold, determine the distance departure time based on the rising distance, and read the initial distance from the spacing rise curve according to the rising distance; Step 203: Read the recovery time that is located after the distance-away time and equal to the initial distance from the distance fluctuation curve, and compare the distance-away time and the recovery time to determine the recovery duration; Step 204: If the recovery time is less than the preset attention threshold, generate and display an attention warning in response to the duration.

5. The intelligent driving assistance safety management method according to claim 4, characterized in that, The attention detection method further includes: Step 205: If the recovery time is less than a preset attention threshold, identify the set of visual focal points from the driving image based on the time of departure, and determine the operation center based on the palm position; Step 206: Determine the visual spacing by comparing the set of visual focal points with the operation center; Step 207: When the visual distance is less than a preset tracking threshold, determine the tracking time based on the visual distance, and compare the tracking time and the recovery time to determine the tracking difference; Step 208: Determine the risk level based on the tracking difference, and match the speed limit according to the risk level; Step 209: In response to the speed limit generation, a speed limit command is sent.

6. The intelligent driving assistance safety management method according to claim 5, characterized in that, The attention detection method further includes: Step 210: When the tracking difference falls into the preset visual range, determine the approach time from the visual landing point set based on the tracking time, and retrieve the approach distance by combining the approach time and the tracking time; Step 211: Generate a proximity curve based on the proximity spacing, and identify the proximity rate from the proximity curve; Step 212: Analyze the approach rate to determine the approach deceleration, and determine the visual dependence based on the approach deceleration; Step 213: Match the dependency coefficient according to the visual dependency, and update the risk level according to the dependency coefficient.

7. The intelligent driving assistance safety management method according to claim 3, characterized in that, It also includes a confidence level calculation method, which includes: Step 300: If the distance between the palms is less than a preset operation threshold, identify the image brightness, image contrast, local variance and global variance from the driving image, and determine the image sharpness by combining the local variance and global variance; Step 301: Determine the brightness confidence level based on the image brightness, and determine the contrast confidence level based on the image contrast. Step 302: Determine the image quality confidence level by combining the image sharpness, brightness confidence level, and contrast confidence level, and determine the proportion mean and proportion standard deviation based on the upper limb proportion; Step 303: Determine the stability of the proportion by combining the mean and standard deviation of the proportion, and match the stability coefficient according to the stability of the proportion; Step 304: Correct the quality confidence score according to the stability coefficient to obtain the image confidence score, and match the confidence score coefficient according to the image confidence score; Step 305: Respond to the confidence coefficient update duration to reduce false alarms.

8. The intelligent driving assistance safety management method according to claim 7, characterized in that, The confidence level calculation method further includes: Step 306: When the confidence level of the image is less than a preset confidence threshold, combine the palm position and upper limb posture to generate a complete posture, and identify the posture range from the complete posture; Step 307: Determine the multi-source type based on the attitude range, and retrieve multi-source data based on the multi-source type; Step 308: Determine the multi-source threshold according to the multi-source type and attitude range, and compare the multi-source data with the multi-source threshold to determine multi-source consistency; Step 309: Match multi-source weights based on the multi-source categories, and calculate weighted consistency by combining the multi-source weights and multi-source consistency; Step 310: Match the multi-source confidence based on the weighted consistency and update the image confidence in response to the multi-source confidence.

9. The intelligent driving assistance safety management method according to claim 8, characterized in that, The confidence level calculation method further includes: Step 311: If the distance between the palms is less than a preset operation threshold, determine the occlusion feature based on the palm position; Step 312: When the occlusion feature is not empty, determine the occlusion ratio from the driving image based on the occlusion feature, and determine the confidence weight according to the occlusion ratio; Step 313: When the occlusion feature is empty, determine the confidence weight based on the image confidence. Step 314: Determine a weighted confidence level by combining the multi-source confidence level, image confidence level, and confidence level weights, and update the image confidence level in response to the weighted confidence level.

10. An intelligent driving assistance safety management system, characterized in that, include: The acquisition module is used to acquire driving images; A memory for storing a program of an intelligent driving assistance safety management method as described in any one of claims 1 to 9; The processor is the unit of memory that allows programs to be loaded and executed by the processor.