Head-eye posture fused driver state grading intervention method and system

By integrating head posture and gaze deviation, combined with vehicle speed and gaze duration, the system dynamically adjusts the gaze area classification and constructs a decision model for tiered intervention. This solves the problems of large gaze point positioning errors and unsuitable intervention strategies in existing technologies, achieving high-precision driver status monitoring and safety assurance.

CN120932208APending Publication Date: 2025-11-11CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510959307.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing driver status monitoring technologies suffer from problems such as large gaze point positioning errors, high false alarm rates, and unsuitable intervention strategies in complex driving scenarios, resulting in insufficient safety and human-centered protection.

Method used

By integrating head posture and gaze deviation, an initial gaze vector is calculated and corrected using facial key points. Combined with vehicle speed and gaze duration, the gaze region classification is dynamically adjusted, and a decision model is constructed for hierarchical intervention.

Benefits of technology

It improves the accuracy of gaze point positioning, reduces interference with normal driving, provides effective early warning during the risk accumulation stage, enhances driving safety and adaptability, and reduces the traffic accident rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932208A_ABST
    Figure CN120932208A_ABST
Patent Text Reader

Abstract

The invention provides a head-eye posture fused driver state grading intervention method and system, and the method comprises the steps: dividing a fixation region which comprises a safety region, an alarm region and a dangerous region based on an in-vehicle three-dimensional coordinate system; calculating an initial line-of-sight vector according to the driver face key points, obtaining a final line-of-sight vector by combining line-of-sight deviation angle correction, and obtaining a fixation point coordinate based on the eye point coordinate and the final line-of-sight vector; based on fixation point coordinates and fixation area matching, the driver fixation area category is judged; and acquiring the current vehicle speed and the fixation duration in the fixation area, inputting the current vehicle speed and the fixation duration into a pre-trained decision model to obtain a graded intervention strategy, and executing an intervention action according to the strategy. The accuracy of the fixation point position is ensured through fusion of the head posture and the sight line deviation; a grading intervention strategy is obtained by combining the fixation point position, the fixation duration and the vehicle speed, interference to normal driving is reduced, effective early warning can be achieved in the risk accumulation stage, and the method has adaptability to complex vehicle use scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of condition monitoring technology, and in particular to a method and system for driver condition classification intervention that integrates head and eye posture. Background Technology

[0002] Driver status monitoring is a core function of intelligent driving assistance systems. By sensing information such as the driver's gaze behavior and head posture, it determines whether there are risks such as distraction or fatigue, and then triggers intervention measures to ensure driving safety. Among these, the accuracy of gaze point positioning and the rationality of the intervention strategy are key to the system's effectiveness. The former needs to accurately map the driver's actual gaze direction, while the latter needs to dynamically adjust the intensity according to the scenario's risks.

[0003] With the development of automotive intelligence, existing driver status monitoring technologies are gradually revealing their limitations. Early monocular vision solutions rely on a single camera to collect eye information, which is susceptible to lighting interference and difficult to accurately convert image coordinates to the three-dimensional space inside the vehicle, resulting in significant gaze point positioning errors. Some solutions determine the gaze direction solely based on head posture, ignoring the inconsistency between head orientation and line of sight, leading to high false alarm rates. Furthermore, existing intervention strategies often employ fixed thresholds, such as uniformly setting a fixed duration for gazing at non-driving areas to trigger an alarm, without dynamically adjusting based on vehicle speed. This results in excessive intervention at low speeds, impacting the driving experience, and delayed response at high speeds, amplifying risks, making it difficult to balance safety and user-friendliness.

[0004] These shortcomings make existing systems unreliable in complex driving scenarios, failing to provide drivers with accurate and appropriate safety guarantees. There is an urgent need for a technical solution that integrates multi-source information and dynamically graded intervention. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a driver state classification intervention method and system that integrates head and eye posture. It employs a method that fuses head posture and gaze deviation, calculating the initial head orientation using facial key points and correcting it with gaze vectors to obtain the fixation point, ensuring the accuracy of the fixation point location. Combining the fixation point location, fixation duration, and vehicle speed, a classification intervention strategy is output based on a decision model. This reduces interference with normal driving and provides effective early warning during risk accumulation stages, demonstrating adaptability to complex driving scenarios.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a driver state classification intervention method integrating head and eye posture, comprising: Based on the in-vehicle three-dimensional coordinate system, the gaze area is divided; the gaze area includes a safe area, an alarm area, and a danger area. Obtain eye point coordinates and calculate the initial gaze vector and gaze offset angle based on the driver's facial key points; use the gaze offset angle to correct the initial gaze vector to obtain the final gaze vector, and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector; The coordinates of the gaze point are matched with the gaze area in the three-dimensional coordinate system inside the vehicle to determine the type of the driver's gaze area; The current vehicle speed and the duration of gaze within the gaze area are acquired in real time, input into a pre-trained decision model to obtain a graded intervention strategy, and intervention actions are executed according to the strategy.

[0007] In one possible implementation, the safety area includes the entire windshield and the area reflected by the exterior rearview mirrors; The alarm area includes the dashboard, central control display screen, rearview mirror body and gear shift area; The danger zones include the seat side adjustment knobs, the interior door handles, the lower half of the steering wheel, the center armrest control panel, and the roof control area.

[0008] In one possible implementation, obtaining the eye point coordinates and calculating the initial gaze vector and gaze offset angle based on the driver's facial key points specifically includes: The pixel coordinates of key facial points are obtained by using two cameras, including the tip of the nose, the inner corner of the left eye, the outer corner of the left eye, the inner corner of the right eye, and the outer corner of the right eye. The average value of the pixel coordinates of the same key point obtained by the two cameras is taken to obtain the final pixel coordinates of the key point. The coordinates of the center of the left pupil are determined based on the coordinates of the inner and outer corners of the left eye, and the coordinates of the center of the right pupil are determined based on the coordinates of the inner and outer corners of the right eye. The eye point pixel coordinates in the image coordinate system are obtained based on the average position of the left pupil center coordinates and the right pupil center coordinates. Transform the eye point coordinates from the image coordinate system to the camera coordinate system, and then to the world coordinate system to obtain the final eye point coordinates. Calculate the initial gaze vector based on the tip of the nose, the center of the left pupil, and the center of the right pupil; Calculate the horizontal pixel distance from the center of each pupil to the inner corner of the eye on the same side, and calculate the line of sight offset angle based on the horizontal pixel distance.

[0009] In one possible implementation, the step of correcting the initial gaze vector using the gaze offset angle to obtain the final gaze vector, and obtaining the gaze point coordinates based on the eye point coordinates and the final gaze vector, specifically includes: Rotate the initial line-of-sight vector around the vertical axis by the line-of-sight offset angle to obtain the final line-of-sight vector in the world coordinate system; Based on the eye point coordinates and the final gaze vector, calculate the gaze point coordinates in the world coordinate system: ; in, Indicates the coordinates of the gaze point; The scalar, representing the line of sight, is used to adjust the length to the surface of the object. Represents the eye point coordinates, indicating Final line-of-sight vector.

[0010] In one possible implementation, the training process of the decision model includes: Collect data on drivers’ natural gaze behavior at various vehicle speeds, and label the gaze region categories and corresponding intervention results; A random forest classifier is constructed, with vehicle speed, gaze duration, and gaze region category as input features, and a hierarchical intervention strategy as output. The model is iteratively trained with the goal of minimizing the false positive rate and false negative rate using a weighted loss function, and finally a well-trained decision model is obtained.

[0011] In one possible implementation, the execution of the intervention action according to the intervention strategy specifically includes: When focusing on the alarm area, if the current vehicle speed and the focusing time meet the first level dynamic threshold, a voice prompt is triggered; the first level dynamic threshold includes a first vehicle speed range corresponding to a first duration, a second vehicle speed range corresponding to a second duration, and a third vehicle speed range corresponding to a third duration. When focusing on a dangerous area, if the current vehicle speed and the focusing time meet the second level dynamic threshold, the seat belt tightening and audible and visual alarms will be triggered. The second level dynamic threshold includes the fourth vehicle speed range corresponding to the fourth duration, the fifth vehicle speed range corresponding to the fifth duration, and the sixth vehicle speed range corresponding to the sixth duration. If no response is received from the driver after a preset time period of focusing on the warning or danger zone, decelerate to a safe stop. Among them, the first speed range is smaller than the second speed range, and the second speed range is smaller than the third speed range; The fourth speed range is smaller than the fifth speed range, and the fifth speed range is smaller than the sixth speed range. The first duration is longer than the second duration, and the second duration is longer than the third duration. The fourth duration is longer than the fifth duration, and the fifth duration is longer than the sixth duration.

[0012] Secondly, the present invention provides a driver state classification intervention system integrating head and eye posture, comprising: The region division module is configured to divide the gaze area based on the in-vehicle three-dimensional coordinate system; the gaze area includes a safe area, an alarm area, and a danger area; The gaze point acquisition module is configured to acquire eye point coordinates and calculate an initial gaze vector and gaze offset angle based on the driver's facial key points; correct the initial gaze vector using the gaze offset angle to obtain a final gaze vector; and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector. The gaze region matching module is configured to match the gaze point coordinates with the gaze region in the vehicle's three-dimensional coordinate system to determine the driver's gaze region category. The graded intervention module is configured to acquire the current vehicle speed and the duration of gaze within the gaze area in real time, input them into a pre-trained decision model to obtain a graded intervention strategy, and execute intervention actions according to the strategy.

[0013] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the driver state classification intervention method fused with head-eye posture as described in the first aspect.

[0014] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the driver state classification intervention method fused with head-eye posture as described in the first aspect.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This invention calculates and corrects gaze vectors based on facial key points, effectively improving the accuracy of gaze point positioning and ensuring the accuracy of region judgment. Simultaneously, the pre-trained decision model, which integrates vehicle speed and gaze duration, can output tiered intervention strategies based on real-time status, avoiding the limitations and lag of traditional intervention methods. This invention can promptly identify dangerous states such as driver distraction and fatigue. Through tiered intervention, it reduces interference with normal driving and provides effective early warning during risk accumulation, significantly improving driving safety and reducing the incidence of traffic accidents caused by abnormal gaze. It provides a more targeted state management solution for intelligent driving assistance systems.

[0016] (2) This invention dynamically adjusts the gaze area classification based on the vehicle's operating status, aligning with the driver's reasonable gaze needs and safety priorities in different scenarios. When driving on the highway, the driver focuses on the road environment, designating the windshield and other areas as safe zones; when parking, due to increased reliance on the central control display area, it is reclassified as a safe zone. This dynamic adaptation ensures that the area division fits the actual driving scenario, making the status monitoring more consistent with the safety logic of different scenarios, ensuring that the area category input to the decision model accurately reflects driving needs, improving the rationality and effectiveness of the graded intervention strategy, and enhancing adaptability to complex driving scenarios.

[0017] (3) This invention employs a fusion of head posture and gaze offset to address the accuracy deficiencies of traditional single-data acquisition of gaze points. An initial gaze vector, i.e., head orientation, is calculated based on facial key points, and gaze offset correction is introduced to eliminate judgment errors caused by head-eye separation. Combined with a binocular vision system, cross-view design and multi-coordinate system transformation improve the positioning accuracy of eye point coordinates and gaze point coordinates. This avoids the problem of head posture methods neglecting independent gaze movement and overcomes the limitation of monocular gaze methods being susceptible to ambient light interference, providing high-precision data support for driver status monitoring and enhancing the reliability of driver assistance systems.

[0018] (4) This invention incorporates vehicle speed for tiered intervention, fully considering the risk differences of the same gaze behavior at different speeds. Through dynamic threshold setting, the higher the vehicle speed, the shorter the corresponding alarm and danger zone gaze duration threshold, and the intervention intensity increases with the risk. From voice prompts to seat belt tightening, sound and light alarms, and then to forced deceleration, it avoids too many false alarms at low speeds and prevents intervention lag at high speeds. Multi-dimensional data joint judgment makes the strategy fit the changes in scenario risk, balances safety and driving experience, maximizes driving safety, and reduces the incidence of traffic accidents.

[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.

[0021] Figure 1 The main flowchart of a driver state classification intervention method that integrates head and eye posture is provided in an embodiment of the present invention. Detailed Implementation

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] Example 1 like Figure 1 As shown, this embodiment discloses a driver state classification intervention method that integrates head and eye posture, including the following steps: S1: Based on the in-vehicle three-dimensional coordinate system, the gaze area is divided; the gaze area includes a safe area, an alarm area, and a danger area. S2: Obtain eye point coordinates and calculate the initial gaze vector and gaze offset angle based on the driver's facial key points; use the gaze offset angle to correct the initial gaze vector to obtain the final gaze vector, and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector; S3: Match the coordinates of the gaze point with the gaze area in the three-dimensional coordinate system inside the vehicle to determine the type of the driver's gaze area; S4: Real-time acquisition of current vehicle speed and gaze duration within the gaze area, input into pre-trained decision model to obtain graded intervention strategies, and execution of intervention actions according to the strategies.

[0024] Next, combined Figure 1 This embodiment provides a detailed description of a driver state classification intervention method that integrates head and eye posture.

[0025] This embodiment first divides the vehicle interior and surrounding environment into a driving zone, an alarm zone, and a danger zone. Images are acquired through cross-viewpoint capture using binocular infrared cameras, and facial key points are extracted using the OpenPose algorithm. The precise gaze point coordinates in the world coordinate system are calculated by fusing head posture and gaze offset. Then, the driver's state is comprehensively judged by combining a dynamic threshold for vehicle speed and the duration of gaze: when the gaze alarm or danger zone duration reaches the threshold at the corresponding vehicle speed, voice prompts, seat belt tightening, and audible and visual alarms are triggered sequentially; if no response is received, the vehicle decelerates to a safe stop. Through the synergy of zoned monitoring, precise positioning, and tiered intervention, dynamic perception and adaptive response to the driver's state are achieved, improving the accuracy and humanization of driving safety assurance.

[0026] In S1, the first step is to construct a three-dimensional spatial coordinate system inside the vehicle, namely the world coordinate system.

[0027] Establish a right-handed Cartesian coordinate system with the center point of the driver's seat as the origin (0,0,0): the X-axis is horizontal to the right (positive direction points to the passenger side), the Y-axis is vertical downward (positive direction points to the ground), and the Z-axis is horizontal forward (positive direction points to the direction of vehicle movement).

[0028] Subsequently, in order to accurately define the impact of different gaze areas on driving safety during driving, safe zones, alarm zones, and danger zones were divided based on the three-dimensional coordinate system inside the vehicle.

[0029] The safe zone is the area of ​​normal driving vision, including the entire area of ​​the windshield and the area reflected in the exterior rearview mirrors; The alarm area refers to the operating area that requires brief attention, including the instrument panel, central control display screen, rearview mirror body, and gear shift area; Dangerous areas are areas that should not be looked at while driving, including the seat side adjustment knobs, the inside door handles, the lower half of the steering wheel, the center console control panel, and the headliner control area.

[0030] Specifically, the "reflected view area" of the exterior rearview mirror refers to the visible area of ​​the mirror surface, that is, the area of ​​the outside scenery that the driver can observe through the mirror reflection, not the physical mirror itself. The physical mirror surface is part of the alarm zone.

[0031] The seat side adjustment knob, in this embodiment, specifically refers to the electric adjustment button group that is usually located on the side of the seat, which requires a large head turn or bending over to operate.

[0032] The central armrest box operating surface refers to the area that the driver needs to look at when operating the armrest box cover (which can be open or closed for retrieving items).

[0033] The aforementioned roof control area refers to the electronic control unit installed in the roof area of ​​the vehicle. It typically includes functional components such as interior lighting control (reading light, ambient light switch), sunroof or sunshade control buttons, emergency alarm button (SOS distress function), microphone array (for voice interaction), and rear seat entertainment system control interface.

[0034] In this embodiment, the safe zone covers the key range of normal driving vision, ensuring the driver obtains core road information; the alarm zone clearly defines functional areas requiring brief operation, facilitating the identification of reasonable but attention-grabbing actions; the danger zone focuses on components or areas requiring significant movement to operate, thereby identifying high-risk distracting behaviors. By making the risk levels of different zones clearly identifiable, subsequent monitoring and decision-making based on the gaze area become more targeted, effectively balancing safety assurance and driving operation needs, and improving the rationality and effectiveness of intervention.

[0035] It should be understood that those skilled in the art can reasonably divide each gaze area according to the actual application scenario and requirements. Based on the established in-vehicle three-dimensional coordinate system, the corresponding three-dimensional coordinate range of each area can be determined. This is something that those skilled in the art can routinely achieve, so it will not be elaborated on here.

[0036] The above describes the zoning for highway driving. Furthermore, considering the differences in drivers' attention needs and safety priorities under different vehicle operating conditions, for example, drivers need to focus on the road environment while driving on the highway, while relying more on the central control display area (such as the reversing camera) when parking.

[0037] Therefore, this embodiment automatically adjusts the gaze area classification according to the current driving mode. In parking mode, the central control display area is reclassified as a safe area.

[0038] In this embodiment, by dynamically dividing the region based on the vehicle's operating status, the gaze area division is deeply adapted to the actual driving scenario, making subsequent driver status monitoring more in line with the safety logic under different scenarios. The region category input by the decision model more accurately reflects driving needs, thereby improving the rationality and effectiveness of the graded intervention strategy, ensuring driving safety under different operating states, and enhancing adaptability to complex vehicle usage scenarios.

[0039] In S2, to address the accuracy deficiencies of traditional single-data fixation point acquisition, this embodiment employs a fusion of head posture and gaze offset. First, an initial gaze vector is calculated based on facial key points to preliminarily characterize the overall head orientation. Then, a gaze offset is introduced to quantify the independent deflection angle of the gaze relative to the head orientation, correcting the initial gaze vector. The corrected gaze vector is then transformed to the world coordinate system to obtain the final gaze vector of the human eye. Finally, the gaze point coordinates in the world coordinate system are calculated by combining the eye point coordinates obtained in advance through a binocular vision system.

[0040] Specifically, two infrared cameras are symmetrically arranged on the top of the dashboard, forming a cross-view design. For example, the left camera is offset to the right by 25°±3°, and the right camera is offset to the left by 35°±3°. This cross-view design expands the effective shooting range of the driver's face, reduces the loss of facial features due to head rotation or changes in posture, improves the stability and accuracy of the binocular vision system in capturing eye points and key facial points, and ensures high-quality image data can be acquired under different driving postures.

[0041] First, the eye-point coordinates in the image coordinate system are obtained from the binocular cameras, then transformed to the camera coordinate system, and finally transformed to the world coordinate system to obtain the final eye-point coordinates. .

[0042] Specifically, the OpenPose algorithm is used to capture the center of the left pupil using dual cameras. and the center of the right pupil The pixel coordinates of the eye point are obtained based on the average position of the centers of both eyes. .

[0043] The OpenPose algorithm is used to extract the pixel coordinates of facial key points, including the tip of the nose. The inner corner of the left eye The outer corner of the left eye The inner corner of the right eye and the outer corner of the right eye It should be understood that by using two cameras to separately capture images and extract key points, and then averaging the coordinates of the same key point obtained from the left and right cameras, the final key point pixel coordinates are obtained. Through dual-view data fusion, measurement errors caused by single-camera biases, uneven lighting, etc., can be offset, improving the stability and accuracy of facial key point coordinates.

[0044] left pupil center Through the inner corner of the left eye and the outer corner of the left eye The midpoint of the line connecting the two pupils is determined; the center of the right pupil is also determined. Through the inner corner of the right eye and the outer corner of the right eye The midpoint of the line is determined.

[0045] It should be understood that the OpenPose algorithm is an open-source human pose estimation algorithm in conventional technology. It outputs the two-dimensional coordinates of multiple facial key points through a pre-trained convolutional neural network. Those skilled in the art can directly call this model to obtain the coordinates of facial key points.

[0046] Transform the eye point coordinates from the image coordinate system to the camera coordinate system: ; In the formula, As the optical center, This refers to the focal length parameter.

[0047] Then transform the eye point coordinates in the camera coordinate system to the eye point coordinates in the world coordinate system. : ; ; In the formula, It is a 3×3 rotation matrix, determined by the camera installation angle; It is a 3×1 translation vector; For equivalent focal length, The baseline distance of the binocular cameras. The difference in the x-coordinates of the matching points in the left and right views.

[0048] Secondly, based on the facial key point, the tip of the nose. left pupil center and the center of the right pupil Calculate the initial line-of-sight vector This refers to the head orientation vector based on the tip of the nose, which can eliminate the influence of individual differences in nasal bridge height. .

[0049] In this embodiment, the initial gaze vector is calculated based on facial key points, eliminating the need for additional devices and reducing driver interference, thus improving driving comfort. Furthermore, utilizing naturally occurring facial key points avoids the need for customized calibration due to individual differences, enhancing versatility. However, head orientation cannot accurately represent the gaze direction, as the head may be facing one direction while the gaze is directed in another. Therefore, this embodiment introduces the calculation of gaze offset, eliminating the bias of judging the gaze direction solely based on head orientation, resolving the gaze point judgment error caused by the "head-eye separation" phenomenon, and ensuring that the final obtained gaze point more closely matches the driver's actual gaze intention.

[0050] Define pupil horizontal offset The horizontal pixel distance from the center of the pupil to the inner corner of the eye on the same side is used as the average of the offsets of the left and right eyes as the line of sight offset. : ; ; .

[0051] In the formula, This represents the left eye offset. This represents the right eye offset. The horizontal coordinate value of the center of the left pupil. The horizontal coordinate value of the inner corner of the left eye; The horizontal coordinate value of the center of the right pupil. This represents the horizontal coordinate value of the inner corner of the right eye.

[0052] Furthermore, calculate the line-of-sight offset angle. : ; Among them, parameters Through calibration experiments, it was determined, for example, that each pixel offset corresponds to an actual line-of-sight deflection of 0.15°.

[0053] The line-of-sight offset angle characterizes the degree of horizontal deflection of the driver's line of sight relative to the direction of the head. It converts the offset distance at the pixel level into a physical angle, providing a quantitative basis for subsequent line-of-sight vector correction, making the correction process more scientific and operable.

[0054] After that, Rotate about the vertical axis Angle, to obtain the final line-of-sight vector in the world coordinate system. : ; Furthermore, the gaze point coordinates are obtained based on the eye point coordinates and the final gaze vector. : ; In the formula, The scalar, representing the line of sight, is used to adjust the length to the surface of the object.

[0055] In this embodiment, by employing a head posture and gaze offset correction mechanism combined with the spatial positioning advantages of binocular vision, the limitations of traditional single head posture or single gaze tracking methods are effectively overcome. This avoids the problems of head posture methods ignoring independent gaze movement and monocular gaze methods being susceptible to ambient light interference. Simultaneously, through precise multi-coordinate system transformation, a mapping from image pixels to gaze coordinates in world space is achieved, providing high-precision data support for subsequent driver status assessment based on gaze position and improving the reliability of the driver assistance system.

[0056] In S3, the coordinates of the gaze point are compared with the coordinate ranges of each gaze area divided in the three-dimensional coordinate system inside the vehicle to determine the category of the driver's gaze area.

[0057] In S4, the current vehicle speed and the duration of gaze within the gaze area are acquired in real time, input into the pre-trained decision model to obtain a graded intervention strategy, and the intervention action is executed according to the strategy.

[0058] Specifically, when focusing on the alarm area, the current vehicle speed and the duration of focus are obtained: if the current vehicle speed and the duration of focus meet the first level dynamic threshold, a voice prompt is triggered; the first level dynamic threshold includes a first vehicle speed range corresponding to a first duration, a second vehicle speed range corresponding to a second duration, and a third vehicle speed range corresponding to a third duration.

[0059] Among them, the first speed range is smaller than the second speed range, and the second speed range is smaller than the third speed range; the first duration is greater than the second duration, and the second duration is greater than the third duration.

[0060] For example, the first vehicle speed range is 0-30 km / h, and the first duration is 5 seconds; the second vehicle speed range is 30-80 km / h, and the second duration is 4 seconds; the third vehicle speed range is above 80 km / h, and the third duration is 3 seconds.

[0061] The voice prompts include tiered voice warnings, for example: when the vehicle is within a first speed range corresponding to a first duration, the system plays "Please note that you have been looking at a non-driving area for an extended period of time. Please focus on driving." When the vehicle is within a second speed range corresponding to a second duration, the system plays "Warning: Do not be distracted at the current speed. Pay attention to road conditions immediately." When the vehicle is within a third speed range corresponding to a third duration, the system plays "Emergency Reminder: While driving at high speed, do not look at irrelevant areas. Ensure safety." By differentiating the voice content and matching the intensity of the prompts according to the vehicle speed and the degree of danger, both excessive interference and effective warnings are avoided.

[0062] When focusing on a dangerous area, obtain the current vehicle speed and the duration of focus: if the current vehicle speed and the focus time meet the second level dynamic threshold, trigger the seat belt tightening and audible and visual alarms; the second level dynamic threshold includes the fourth vehicle speed range corresponding to the fourth duration, the fifth vehicle speed range corresponding to the fifth duration, and the sixth vehicle speed range corresponding to the sixth duration.

[0063] Among them, the fourth speed range is smaller than the fifth speed range, and the fifth speed range is smaller than the sixth speed range; the fourth duration is greater than the fifth duration, and the fifth duration is greater than the sixth duration.

[0064] For example, the fourth speed range is 0-30km / h, and the fourth duration is 4s; the fifth speed range is 30-80km / h, and the fifth duration is 3s; the sixth speed range is above 80km / h, and the sixth duration is 2s.

[0065] The seatbelt tightening forcefully awakens the driver's attention through physical touch. By applying pressure to the driver, it reminds them to immediately turn their gaze back to the driving area, providing a more immediate response than simple auditory cues.

[0066] The sound and light alarm includes a red warning light on the dashboard that flashes rapidly (2-3 times per second), while the speakers in the vehicle emit a high-frequency buzzing sound (the volume increases by 5-10 dB as the vehicle speed increases), and the interval between the buzzing sounds gradually shortens (from 1 second / time to 0.5 seconds / time). Through the dual strong stimulation of sight and hearing, the driver's distracted state is quickly broken.

[0067] If the driver does not respond after a preset time period of focusing on the warning or danger zone, the vehicle will decelerate to a safe stop. This may be due to the driver being unaware (e.g., fatigued or drowsy), experiencing a sudden physical condition, or intentionally ignoring the warning.

[0068] The deceleration process involves the vehicle system automatically activating its hazard lights and smoothly decelerating at a certain acceleration, while simultaneously sending an "emergency deceleration" signal to surrounding vehicles via the vehicle-to-everything (V2X) network. When the speed drops below 30 km / h, the vehicle gradually activates its right turn signal and slowly moves to the emergency lane or the right side of the road. Upon coming to a complete stop, the vehicle automatically unlocks the doors, dials an emergency contact number, and uploads its location information to the backend system to ensure driver safety. Vehicle speed and gaze duration can be obtained through the vehicle's central processing unit.

[0069] As one implementation method, the training process of the decision model includes: Collect data on drivers’ natural gaze behavior at various vehicle speeds, and label the gaze region categories and corresponding intervention results; A random forest classifier is constructed, with vehicle speed, gaze duration, and gaze region category as input features, and a hierarchical intervention strategy as output. The model is iteratively trained with the goal of minimizing the false positive rate and false negative rate using a weighted loss function, and finally a well-trained decision model is obtained.

[0070] In this embodiment, the decision model uses a random forest classifier, which, compared to single threshold judgment, can learn complex driving scenario patterns through multi-feature combination and improve the reliability of the system by combining cross-validation to optimize parameters.

[0071] This embodiment considers multi-dimensional data including vehicle speed, driver's gaze area, and gaze duration, enabling a comprehensive characterization of the risk level of driver distraction. At different vehicle speeds, the degree of danger varies significantly for the same gaze duration; furthermore, at the same vehicle speed, the risk of gazing into a dangerous area is higher than that of a warning area. By jointly considering vehicle speed factors beyond the driver's state, the intervention strategy better reflects the dynamic changes in risk in actual driving scenarios, avoiding problems such as excessive false alarms at low speeds or delayed intervention at high speeds caused by traditional static thresholds that rely solely on gaze duration.

[0072] Meanwhile, the tiered intervention logic, from voice prompts to forced deceleration, ensures driving safety while respecting the driver's operational autonomy to the greatest extent, avoiding excessive or premature intervention, and balancing safety and driving experience.

[0073] This specific embodiment improves the positioning accuracy of eye point coordinates and initial gaze vectors by employing a binocular cross-view design and integrating the OpenPose algorithm, combined with multi-coordinate system transformation, thus resolving the error issues associated with monocular vision and single head posture judgment. Simultaneously, it introduces gaze offset quantification to quantify head-eye separation deviation and accurately captures the actual gaze direction through an initial gaze vector correction mechanism. Furthermore, it constructs a dynamic hierarchical intervention strategy based on vehicle speed, gaze area, and duration, matching the scenario's risk level. This achieves high-precision gaze point positioning while balancing safety and user experience through adaptive intervention, providing more reliable core technological support for intelligent driving assistance and driving the upgrade of driver state monitoring towards precision and intelligence.

[0074] Example 2 This embodiment provides a driver state classification intervention system that integrates head and eye posture, including: The region division module is configured to divide the gaze area based on the in-vehicle three-dimensional coordinate system; the gaze area includes a safe area, an alarm area, and a danger area; The gaze point acquisition module is configured to acquire eye point coordinates and calculate an initial gaze vector and gaze offset angle based on the driver's facial key points; correct the initial gaze vector using the gaze offset angle to obtain a final gaze vector; and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector. The gaze region matching module is configured to match the gaze point coordinates with the gaze region in the vehicle's three-dimensional coordinate system to determine the driver's gaze region category. The graded intervention module is configured to acquire the current vehicle speed and the duration of gaze within the gaze area in real time, input them into a pre-trained decision model to obtain a graded intervention strategy, and execute intervention actions according to the strategy.

[0075] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the driver state classification intervention method fused with head and eye posture as described in Embodiment 1 above.

[0076] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the driver state classification intervention method fused with head and eye posture as described in Embodiment 1 above.

[0077] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0078] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A driver state classification intervention method integrating head-eye posture, characterized in that, include: Based on the in-vehicle three-dimensional coordinate system, the gaze area is divided; the gaze area includes a safe area, an alarm area, and a danger area. Obtain eye point coordinates and calculate the initial gaze vector and gaze offset angle based on the driver's facial key points; use the gaze offset angle to correct the initial gaze vector to obtain the final gaze vector, and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector; The coordinates of the gaze point are matched with the gaze area in the three-dimensional coordinate system inside the vehicle to determine the type of the driver's gaze area; The current vehicle speed and the duration of gaze within the gaze area are acquired in real time, input into a pre-trained decision model to obtain a graded intervention strategy, and intervention actions are executed according to the strategy.

2. The driver state classification intervention method integrating head and eye posture as described in claim 1, characterized in that, The safe zone includes the entire area of ​​the windshield and the area reflected in the exterior rearview mirrors; The alarm area includes the dashboard, central control display screen, rearview mirror body and gear shift area; The danger zones include the seat side adjustment knobs, the interior door handles, the lower half of the steering wheel, the center armrest control panel, and the roof control area.

3. The driver state classification intervention method integrating head and eye posture as described in claim 1, characterized in that, The process of obtaining eye point coordinates and calculating the initial gaze vector and gaze offset angle based on the driver's facial key points specifically includes: The pixel coordinates of key facial points are obtained by using two cameras, including the tip of the nose, the inner corner of the left eye, the outer corner of the left eye, the inner corner of the right eye, and the outer corner of the right eye. The average value of the pixel coordinates of the same key point obtained by the two cameras is taken to obtain the final pixel coordinates of the key point. The coordinates of the center of the left pupil are determined based on the coordinates of the inner and outer corners of the left eye, and the coordinates of the center of the right pupil are determined based on the coordinates of the inner and outer corners of the right eye. The eye point pixel coordinates in the image coordinate system are obtained based on the average position of the left pupil center coordinates and the right pupil center coordinates. Transform the eye point coordinates from the image coordinate system to the camera coordinate system, and then to the world coordinate system to obtain the final eye point coordinates. Calculate the initial gaze vector based on the tip of the nose, the center of the left pupil, and the center of the right pupil; Calculate the horizontal pixel distance from the center of each pupil to the inner corner of the eye on the same side, and calculate the line of sight offset angle based on the horizontal pixel distance.

4. The driver state classification intervention method integrating head and eye posture as described in claim 1, characterized in that, The process of correcting the initial gaze vector using the gaze offset angle to obtain the final gaze vector, and obtaining the fixation point coordinates based on the eye point coordinates and the final gaze vector, specifically includes: Rotate the initial line-of-sight vector around the vertical axis by the line-of-sight offset angle to obtain the final line-of-sight vector in the world coordinate system; Based on the eye point coordinates and the final gaze vector, calculate the gaze point coordinates in the world coordinate system: ; in, Indicates the coordinates of the gaze point; The scalar, representing the line of sight, is used to adjust the length to the surface of the object. Represents the eye point coordinates, indicating Final line-of-sight vector.

5. The driver state classification intervention method integrating head and eye posture as described in claim 1, characterized in that, The training process of the decision model includes: Collect data on drivers’ natural gaze behavior at various vehicle speeds, and label the gaze region categories and corresponding intervention results; A random forest classifier is constructed, with vehicle speed, gaze duration, and gaze region category as input features, and a hierarchical intervention strategy as output. The model is iteratively trained with the goal of minimizing the false positive rate and false negative rate using a weighted loss function, and finally a well-trained decision model is obtained.

6. The driver state classification intervention method integrating head and eye posture as described in claim 1, characterized in that, The execution of intervention actions according to the intervention strategy specifically includes: When focusing on the alarm area, if the current vehicle speed and the focusing time meet the first level dynamic threshold, a voice prompt is triggered; the first level dynamic threshold includes a first vehicle speed range corresponding to a first duration, a second vehicle speed range corresponding to a second duration, and a third vehicle speed range corresponding to a third duration. When focusing on a dangerous area, if the current vehicle speed and the focusing time meet the second level dynamic threshold, the seat belt tightening and audible and visual alarms will be triggered. The second level dynamic threshold includes the fourth vehicle speed range corresponding to the fourth duration, the fifth vehicle speed range corresponding to the fifth duration, and the sixth vehicle speed range corresponding to the sixth duration. If no response is received from the driver after a preset time period of focusing on the warning or danger zone, decelerate to a safe stop. Among them, the first speed range is smaller than the second speed range, and the second speed range is smaller than the third speed range; The fourth speed range is smaller than the fifth speed range, and the fifth speed range is smaller than the sixth speed range. The first duration is longer than the second duration, and the second duration is longer than the third duration. The fourth duration is longer than the fifth duration, and the fifth duration is longer than the sixth duration.

7. A driver state classification intervention system integrating head and eye posture, characterized in that, include: The region division module is configured to divide the gaze area based on the in-vehicle three-dimensional coordinate system; the gaze area includes a safe area, an alarm area, and a danger area; The gaze point acquisition module is configured to acquire eye point coordinates and calculate an initial gaze vector and gaze offset angle based on the driver's facial key points; correct the initial gaze vector using the gaze offset angle to obtain a final gaze vector; and obtain the gaze point coordinates based on the eye point coordinates and the final gaze vector. The gaze region matching module is configured to match the gaze point coordinates with the gaze region in the vehicle's three-dimensional coordinate system to determine the driver's gaze region category. The graded intervention module is configured to acquire the current vehicle speed and the duration of gaze within the gaze area in real time, input them into a pre-trained decision model to obtain a graded intervention strategy, and execute intervention actions according to the strategy.

8. The driver state classification intervention system integrating head and eye posture as described in claim 7, characterized in that, The safe zone includes the entire area of ​​the windshield and the area reflected in the exterior rearview mirrors; The alarm area includes the dashboard, central control display screen, rearview mirror body and gear shift area; The danger zones include the seat side adjustment knobs, the interior door handles, the lower half of the steering wheel, the center armrest control panel, and the roof control area.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the driver state classification intervention method that integrates head and eye posture as described in any one of claims 1-6.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the driver state classification intervention method that integrates head and eye posture as described in any one of claims 1-6.

Citation Information

Cited By

  • Video image analysis method and system based on AI large model

    CN121553151A