An active voice guidance method and system for automatic driving takeover engagement state
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JILIN UNIVERSITY
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-07
AI Technical Summary
现有接管提醒技术大多忽视了这一核心认知恢复过程,导致即便驾驶员接收到提醒信号,其空间定向、操作协调性及反应速度仍可能无法满足安全接管的要求
首先,在接管引导的精准性与接管成功率方面,本发明突破了传统方案以提醒输出为核心的局限,转而以驾驶员的接管准备形成为控制目标。通过基于感觉统合恢复规律将接管参与过程划分为多个连续状态,并针对每个状态迁移阶段设计专属的主动发声策略(如任务占用剥离、前向空间锚定、动作准备诱导等),本发明能够分步、定向地引导驾驶员从非驾驶沉浸态逐步迁移至接管就绪态,从而显著提高接管成功率,解决了现有技术中无差别或固定分级提醒所导致的核心缺陷。
Smart Images

Figure CN122186216B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human-computer interaction technology in intelligent cockpits, and particularly relates to an active voice guidance method and system for autonomous driving takeover participation. Background Technology
[0002] With the rapid development of autonomous driving technology, Level 2+ / Level 3 autonomous driving systems have been gradually applied to mass-produced vehicles. These systems can continuously perform lateral and longitudinal control of the vehicle in scenarios such as highways and urban expressways, significantly reducing the driver's workload. However, limited by the system's perception capabilities and processing limits, in situations such as road construction, complex traffic participant intervention, system malfunctions, or exceeding the designed operating range, the vehicle still needs to issue a takeover request to the driver, requiring them to regain control of the vehicle within a limited time. Therefore, the response efficiency of the takeover request and the quality of the driver's takeover preparation directly affect the safety and usability of the autonomous driving system in real-world road environments.
[0003] In existing technologies, alert solutions for autonomous driving takeover scenarios mainly revolve around the combination of multimodal alert methods and the graded control of alert intensity. Common approaches include: configuring corresponding prompts for different state transition points of the autonomous driving system; issuing audio-visual prompts based on the driver's gaze recognition results; or using multimodal stimuli such as visual, auditory, and tactile stimuli for graded adjustment. The core idea of these solutions is to encourage the driver to complete the takeover operation as quickly as possible by enhancing the salience of the alerts or enriching the combination of alert channels. However, in real-world driving environments, drivers in autonomous driving mode may be engaged in tasks unrelated to driving, such as watching videos, making or receiving phone calls, reading messages, or even resting with their eyes closed. Their attention resources, spatial orientation ability, and operational readiness are all outside of driving mode. In such cases, simply increasing the alert intensity or layering multiple alert methods often fails to stably and efficiently guide the driver to complete the takeover preparation, and may instead trigger startled reactions, incorrect operations, or delayed takeover.
[0004] Furthermore, existing technologies generally lack modeling and phased guidance for the gradual recovery process of drivers from a non-driving state to a takeover-ready state. Different drivers exhibit significant differences in their attention recovery, visual regression, and hand preparation behaviors during the takeover process, and the recovery pace of the same driver also varies in different scenarios. Existing solutions often employ one-time or fixed-level reminder strategies, making it difficult to dynamically adjust the reminder method and pace according to the driver's real-time takeover readiness status, easily leading to insufficient or excessive reminders.
[0005] Meanwhile, current mass-produced vehicles are generally equipped with driver monitoring systems, steering wheel grip sensors, in-vehicle media controllers, and bus data acquisition capabilities, enabling real-time acquisition of information such as driver head posture, gaze direction, steering wheel grip status, and in-vehicle media playback status. However, existing technologies rarely utilize this real-time data for closed-loop assessment and adaptive intervention of takeover readiness, making it difficult to effectively verify whether the driver has truly completed takeover readiness and to cope with complex scenarios such as drastic fluctuations or rapid recovery of driver status.
[0006] Furthermore, from the perspective of human perception and motor control, in autonomous driving mode, the driver's sensory integration mode is in a non-driving-specific mode, with vision, vestibular sense, and proprioception in a decoupled or relaxed state. Takeover driving, however, requires the driver to quickly switch to a driving coordination mode, achieving coordination between vision and vestibular sense, and activation of proprioception and touch. Most existing takeover warning technologies neglect this core cognitive recovery process, resulting in situations where, even if the driver receives a warning signal, their spatial orientation, operational coordination, and reaction speed may still fail to meet the requirements for safe takeover. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides an active voice guidance method and system for autonomous driving takeover participation. Based on consideration of the driver's sensory integration recovery patterns, a progressive guidance mechanism for the takeover preparation state is constructed using observable information from onboard sensors.
[0008] Specifically, the technical solution provided by this invention is as follows: A method for proactively guiding autonomous driving takeover participation includes the following steps: S1. Respond to takeover requests and obtain multi-dimensional status information, including takeover request information, driver status information, in-vehicle media status information, and vehicle operation status information; S2. Calculate the takeover participation proxy variables based on the multi-dimensional state information, including non-driving task occupancy, forward visual regression, and hand takeover readiness; the non-driving task occupancy, forward visual regression, and hand takeover readiness together constitute a quantitative representation of the driver's sensory integration recovery degree, which is used to reflect the process of the driver transitioning from a non-driving state of visual-vestibular-proprioceptive decoupling to a driving coordination state. S3. Based on the takeover participation agent variable, determine the driver's takeover participation state as one of multiple consecutive states, including non-driving immersion state, task disengagement state, forward orientation recovery state, action preparation state, and takeover ready state. S4. Based on the migration relationship between the current determined takeover participation state and the target takeover participation state, generate the corresponding active voice control strategy. When it is determined that the takeover participation state needs to migrate from the current state to the next state, generate the corresponding active voice segment. S5. Output an active sound signal according to the active sound control strategy, and re-collect driver status information after a preset time, update the takeover participation agent variable, re-determine the takeover participation status, until the takeover is completed.
[0009] Furthermore, the non-driving task occupancy rate is calculated based on one or more of the following information: the probability that the driver's gaze is directed at the non-driving display area, the activity status of video or entertainment content, the call status, and the driver's closed-eye or resting state; the forward visual regression rate is calculated based on one or more of the following information: the probability of looking at the forward road, the degree of head straightening, and the duration of continuous forward gaze; the hand takeover readiness rate is calculated based on one or more of the following information: the steering wheel grip status, the distance between the hand and the steering wheel, and the degree of torso leaning forward.
[0010] Non-driving task occupancy (NDO) quantifies the degree to which the visual-cognitive pathway in the driver's sensory integration is detached from the driving task, from the perspective of visual attention and cognitive resource occupancy. In one optional implementation, its calculation formula is as follows: NDO = a 1·P_screen + a 2·M_media + a 3·C_call + a 4·S_rest Here, NDO represents the non-driving task occupancy rate; P_screen is the probability that the driver's gaze is directed towards the non-driving display area. This probability is obtained by capturing the driver's eyes and gaze direction in real time through the in-vehicle driver monitoring system's camera and calculating the proportion of time the gaze falls on the non-driving related screen area within a unit time window; M_media is the media content occupancy intensity, dynamically determined based on the playback status information of the entertainment host; C_call is the call status occupancy factor, obtained by determining whether the driver is in a call state through the in-vehicle communication module or Bluetooth connection status; S_rest is the closed-eye or rest state factor, obtained by analyzing the degree of eye opening and closing and the duration of eye closing through the driver monitoring system. a 1. a 2. a 3. a 4 represents the pre-defined weighting coefficient; Forward visual regression (FVR) quantifies the degree of visual-vestibular coordinated reconstruction of a driver from the perspective of spatial orientation recovery. In one optional implementation, its calculation formula is as follows: FVR = b 1·P_road + b 2·A_head + b 3·T_front Wherein, FVR represents forward visual regression; P_road is the probability of forward road gaze, which represents the proportion of the total time the driver's gaze falls on the windshield and the road area in front of them within the observation time window; A_head is the degree of head alignment, which quantifies whether the head is facing forward based on the driver's head yaw angle; and T_front is the normalized value of the duration of continuous forward gaze, used to measure the driver's sustained stability in keeping their gaze on the road ahead. b 1. b 2. b 3 represents the pre-defined weighting coefficient; Hand-to-hand readiness (HPR) quantifies the degree of proprioceptive activation and hand-eye coordination readiness of the driver from the perspective of somatic motor readiness. In one optional implementation, its calculation formula is as follows: HPR = d 1·G_wheel + d 2·D_hand-wheel -1 + d 3·L_torso Among them, HPR represents the readiness of hand takeover; G_wheel represents the steering wheel grip state; D_hand-wheel is the straight-line distance between the driver's hands and the steering wheel; L_torso is the degree of torso lean, which is obtained through skeletal key point detection or seat pressure distribution by the driver monitoring system; d 1. d 2. d 3 represents the pre-defined weighting coefficient.
[0011] Furthermore: The non-driving immersion state refers to a state in which the driver is completely immersed in a non-driving task and has not begun preparations for takeover. In this state, the driver's vision, vestibular sense, and proprioception are decoupled, meaning that visual attention is deviated from the driving scene, the vestibular spatial reference frame is not activated, and the proprioceptive operation loop is dormant. The determination criteria are that at least two of the following conditions are met: non-driving task occupancy is not less than a first threshold, forward visual regression is not greater than a second threshold, and hand takeover readiness is not greater than a third threshold. The task disengagement state refers to a state in which the driver has begun to disengage from non-driving tasks but has not yet established forward orientation. In this state, the driver's visual-cognitive coupling begins to disengage from non-driving tasks, but visual-vestibular spatial coordination has not yet been established. The determination criteria are: the occupancy of non-driving tasks decreases by more than a predetermined amount, and the forward visual regression degree does not reach the fourth threshold. The forward orientation recovery state indicates that the driver's vision has returned to the road ahead, but the hands are not yet fully prepared. In this state, the driver's visual-vestibular spatial reference system has re-established coordination, and vision has returned to the driving scene and formed a stable forward spatial orientation, but the proprioceptive-tactile operational loop has not yet been fully activated. The determination condition is that the forward visual return degree reaches the fifth threshold, and the hand takeover readiness degree is lower than the sixth threshold. The action-ready state indicates that the driver has established forward orientation and begun preparing for hand and torso operations. In this state, while maintaining visual-vestibular spatial orientation, the proprioceptive channel is activated, the somatic motor system enters the operation-ready mode, and the visual-vestibular-proprioceptive system begins to enter a driving coordination mode. The determination criteria are: forward visual regression is higher than the fifth threshold, and hand takeover readiness continuously increases and exceeds the seventh threshold without triggering a takeover confirmation operation. The takeover ready state indicates that the driver has completed takeover preparation and is able to safely hand over control. In this state, the driver's vision, vestibular sense, and proprioception have completed the switching of driving coordination mode, and sensory integration is fully adapted to the needs of manual driving, allowing for safe takeover of vehicle control. The determination criteria are that at least two of the following are met: steering wheel grip is established, forward visual regression reaches a threshold, hand takeover readiness is higher than the eighth threshold, and the driver triggers a takeover confirmation operation.
[0012] Preferably, the active sound control strategy includes generating active sound control parameters, which include one or more of the following: media volume reduction ratio, sound source location, sound image migration trajectory, frequency band distribution, rhythm parameters, duration, semantic type, and output stage sequence.
[0013] Furthermore, the active sound segments sequentially include a task occupancy stripping sound segment, a forward space anchoring sound segment, an action preparation induction sound segment, and a takeover confirmation sound segment; The task occupancy stripping sound segment is used to reduce non-driving task occupancy, including: lowering the volume of the current in-vehicle media, outputting a short pulse prompt in the current attention focus area, and shifting the sound image towards the front of the vehicle; the forward space anchoring sound segment is used to establish forward space orientation, including: outputting a continuous and stable sound anchor point at the speaker position corresponding to the central axis of the vehicle's windshield; the action preparation induction sound segment is used to induce the hands and torso to enter an operation preparation state, including: outputting local short pulses in the steering wheel area or the near-field speakers in the front cabin, inducing the hands to move closer to the steering wheel through rhythmic changes; the takeover confirmation sound segment is used to confirm the completion of takeover, including: outputting a short confirmation sound, ending the forward anchoring and action induction sound segments, and triggering control switchover confirmation; The design of the four active sound segments follows the natural sequence of sensory integration recovery in the human body: First, significant auditory stimulation interrupts the continuous occupation of the visual-cognitive channel by the current non-driving task (task occupancy stripping sound segment); second, a stable forward sound anchor guides the reconstruction of spatial coordination between vision and vestibular sense (forward spatial anchoring sound segment); then, local sound impulses in the steering wheel area activate proprioception and induce hand preparation (motor preparation induction sound segment); finally, positive feedback from confirming sounds consolidates the established driving coordination pattern (takeover confirmation sound segment). The entire process smoothly connects the various recovery stages of the sensory-motor system, avoiding perceptual confusion or startle reactions caused by jumping between stages.
[0014] Each sound segment corresponds to a state transition: the task occupancy stripping sound segment corresponds to the transition from non-driving immersion state to task disengagement state, the forward spatial anchoring sound segment corresponds to the transition from task disengagement state to forward orientation recovery state, the action preparation induction sound segment corresponds to the transition from forward orientation recovery state to action preparation state, and the takeover confirmation sound segment corresponds to the transition from action preparation state to takeover ready state; the corresponding sound segments are output sequentially according to the order of state transitions.
[0015] Furthermore, it also includes a rapid recovery detection step: when the rate of change of the takeover participation proxy variable is detected to exceed a preset rapid recovery threshold, at least one of the following operations is performed: the preset duration of the active sound segment is shortened to a duration shorter than the standard duration; the active sound segment corresponding to at least one intermediate state between the current takeover participation state and the target takeover participation state is skipped; during the output of the current sound segment, if it is detected that the proxy variable has reached the threshold corresponding to the next state, the current sound segment is terminated in advance, and the minimum output duration of each sound segment is not lower than a preset lower limit.
[0016] Furthermore, it also includes a status confirmation delay mechanism: when the fluctuation range of the takeover participant agent variable exceeds a threshold within a preset time window, the confirmation time for status determination is extended.
[0017] Furthermore, if it is determined that the takeover participation state transition is unsuccessful within the preset detection window, at least one of the following actions will be performed: increase the salience of the current sound segment; extend the duration of the current sound segment; switch to a higher priority parameter set; reduce the media volume; or trigger a minimum risk maneuver when the remaining takeover time is insufficient, automatically decelerate and pull over to the side of the road.
[0018] An active sound-guiding system based on the above method, the system comprising the following modules: The takeover requirement acquisition module is used to acquire takeover countdown, risk level, and vehicle operating status; The status acquisition module is used to collect the driver's head posture, eye movement, screen viewing status, steering wheel grip status, torso posture, and cockpit media status. The proxy variable calculation module is used to calculate the proxy variables involved in takeover, including non-driving task occupancy, forward vision regression degree, and hand takeover readiness degree. The status determination module is used to determine the takeover participation status of the driver based on the takeover participation agent variable. The active voice strategy generation module is used to generate corresponding active voice control parameters based on the migration relationship between the currently determined takeover participation state and the target takeover participation state. The media control module is used to suppress, pause, or avoid frequencies of the media according to the active sound control parameters. The sound field output module is used to control the in-vehicle speakers to output corresponding active sound signals; The closed-loop update module is used to adjust the subsequent active sound control strategy based on the driver's response. The modules communicate with each other via vehicle bus or Ethernet to collaboratively complete the closed-loop guidance process from the issuance of the takeover request to the completion of the takeover.
[0019] Compared with the prior art, the present invention has at least the following beneficial effects: Firstly, regarding the accuracy and success rate of takeover guidance, this invention breaks through the limitations of traditional solutions that focus on reminder output, instead taking the driver's takeover readiness as the control objective. By dividing the takeover process into multiple continuous states based on sensory integration recovery principles, and designing exclusive active vocalization strategies for each state transition stage (such as task occupancy stripping, forward spatial anchoring, and action preparation induction), this invention can guide the driver step-by-step and directionally from a non-driving immersion state to a takeover-ready state, thereby significantly improving the takeover success rate and solving the core defects caused by indiscriminate or fixed-level reminders in existing technologies.
[0020] Secondly, this invention significantly reduces safety risks during takeover. By employing a gradual wake-up strategy, it avoids issuing sudden, high-intensity alarms when the driver is resting with their eyes closed or deeply engaged in non-driving tasks, effectively reducing the probability of startle reactions and accidental physical actions (such as incorrect foot pedaling or operating without establishing visual awareness). Simultaneously, the active sound design follows the natural sequence of sensory integration recovery: first interrupting the current mode, then establishing spatial coordination, and finally activating operational readiness. This allows for a smooth transition between the driver's perception and motor systems, reducing perceptual confusion and spatial disorientation, thereby improving the overall safety of the takeover process.
[0021] Furthermore, this invention achieves closed-loop control and strong adaptive capabilities. By updating the proxy variables for takeover in real time and verifying the state transition results in a closed loop, the system can continuously evaluate the driver's actual response and dynamically adjust the proactive audio strategy to avoid over- or under-prompting alerts. In addition, the system also incorporates enhanced mechanisms such as rapid recovery detection, dynamic detection window adjustment, rapid state transition path, early termination of audio segments, historical state memory, and false trigger suppression. These mechanisms can adapt to complex scenarios such as individual driver differences, rapid recovery of attention, or drastic state fluctuations, significantly reducing ineffective interventions and improving the user experience.
[0022] Meanwhile, this invention features low cost and strong adaptability to mass production. The state determination and proxy variable calculation of this invention are entirely based on the Driver Monitoring System (DMS), steering wheel capacitive sensor, media controller, and vehicle bus data already standard on mass-produced vehicles. It requires no additional medical-grade or high-cost equipment such as EEG or vestibular evoked responses, and can be directly adapted to existing mass-produced models, facilitating rapid promotion and application. Furthermore, the calculation formulas and weighting coefficients of each proxy variable can be determined through standardized real-vehicle calibration tests, ensuring the repeatability and universality of the solution across different vehicle models and driver groups.
[0023] Furthermore, this invention exhibits excellent scalability. Its core state transition control logic does not rely on a specific form of alert, and can seamlessly integrate with multimodal takeover alert schemes such as visual cues (e.g., dashboard icon flashing, HUD light strips), steering wheel vibration, and seat tactile cues, forming a multimodal collaborative takeover guidance system. This open architecture allows the invention to be flexibly integrated into existing or future smart cockpit interaction systems, providing a foundational framework for higher-level autonomous driving human-machine interaction design. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0025] Figure 1 This is a flowchart illustrating the active sound-guiding method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the active sound guidance system provided in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of the present invention.
[0027] Example 1 This embodiment provides an autonomous driving takeover active voice signaling method based on takeover participation state transition control. The core concept is to utilize existing driver monitoring systems, media status data, and vehicle bus data from mass-produced vehicles to calculate multiple observable takeover participation proxy variables. Based on these proxy variables, the driver's takeover participation process is divided into multiple continuous states. According to the transition relationship between the current state and the target state, active voice signals for different functional segments are generated to gradually guide the driver from a non-driving state to a takeover-ready state safely and efficiently. Throughout the process, a closed-loop update mechanism continuously evaluates the driver's response, dynamically adjusts the voice signaling strategy, and can adapt to complex scenarios such as rapid driver recovery and state fluctuations.
[0028] like Figure 1 As shown, the method mainly includes the following steps: I. Obtaining multi-dimensional status information In response to a takeover request from the autonomous driving system, the following types of information are acquired simultaneously: (1) Takeover requirements information: including at least the remaining takeover time, takeover risk level, takeover triggering reason (such as road construction, sensor failure, exceeding the design operating range, etc.), and whether the minimum risk maneuver countdown has been entered.
[0029] (2) Vehicle operating status information: including at least vehicle speed, longitudinal acceleration, lateral acceleration, yaw rate, and future short-term trajectory direction information (which can be provided by navigation or lane line recognition).
[0030] (3) Driver status information: collected through the in-vehicle driver monitoring system (DMS), including at least eye opening and closing status, gaze direction classification (such as forward, central control screen, rearview mirror, side window, etc.), head posture (pitch, yaw, roll angle), torso posture (forward, backward, side tilt), hand position, and steering wheel grip status (which can be detected by steering wheel capacitive sensor or torque sensor).
[0031] (4) In-vehicle media status information: including at least audio / video playback status, current media volume, call status (whether it is a Bluetooth call or a car call), current display content status (the type of content displayed on the central control screen, passenger screen, and rear screen), mobile phone interconnection status, and headphone connection status (such as whether the driver is wearing a Bluetooth headphone).
[0032] The above information can be obtained in real time through existing hardware such as the vehicle's CAN bus, Ethernet, DMS camera, and infotainment system, without the need for additional sensors.
[0033] II. Calculate the proxy variables involved in the takeover. This step calculates multiple proxy variables for takeover participation based on the aforementioned multidimensional information. Each proxy variable corresponds to an observable dimension in the driver's sensory integration recovery process. This embodiment calculates at least Non-Driving Task Occupancy (NDO), Forward Visual Regression (FVR), and Hand Takeover Readiness (HPR), and may supplement this with Auditory Takeover Response (ATR) and Motor Consistency Response (MCR) to improve the accuracy of state determination.
[0034] (1) Non-driving task occupancy (NDO) This variable reflects the degree to which the driver is currently occupied by tasks unrelated to driving, i.e., the observable progress of task disengagement. Its calculation is based on one or more of the following information: the probability that the driver's gaze is directed towards a non-driving display area (center console screen, mobile phone, passenger screen, etc.), the activity status of video or entertainment content, call status, and whether the driver's eyes are closed or resting. An exemplary calculation formula is as follows: NDO = a 1·P_screen + a 2·M_media + a 3·C_call + a 4·S_rest P_screen represents the probability that the driver's gaze is directed towards a non-driving display area. This probability can be obtained by capturing the driver's eyes and gaze direction in real time through the camera of the in-vehicle driver monitoring system (DMS), and by statistically analyzing the proportion of time the gaze falls on the non-driving related screen area within a unit time window (e.g., 1 second). The value ranges from 0 to 1.
[0035] M_media represents the intensity of media content occupancy. It comprehensively reflects the type of media currently playing in the vehicle and its potential attraction to the driver's attention. It can be dynamically determined based on the playback status information of the infotainment system: if video content (such as movies or short videos) is playing, M_media takes a higher value (usually 0.8 to 1.0); if only audio (such as music or podcasts) is playing, it takes a medium value (0.4 to 0.7); if no media is playing, it takes 0.
[0036] C_call is the call status occupancy factor, which determines whether the driver is in a call based on the vehicle communication module or Bluetooth connection status. If in a call, it is set to 0.8 to 1.0 (the specific value can be fine-tuned according to the hands-free / handheld call mode), and if not in a call, it is set to 0.
[0037] S_rest is the eye-closed or resting state factor, which is analyzed by the driver monitoring system to determine the degree of eye opening and closing and the duration of eye closure: when the driver's eyes are closed for more than 2 seconds without rapid eye movement, it is considered to be in a resting state and is set to 0.9; if only a brief blink occurs (less than 0.5 seconds), a low value between 0 and 0.2 is used; and 0 is used in the normal open eye state.
[0038] Each of the four indicators mentioned above is multiplied by a pre-defined weighting coefficient. a 1. a 2. a 3. a 4. Summing the results yields the non-driving task occupancy rate. The weighting coefficients all range from 0 to 1 and satisfy the following conditions: a 1+ a 2+ a 3+ a 4 = 1, for example: a 1 = 0.4 a 2 = 0.3 a 3 = 0.2 a 4=0.1. These coefficients were optimized through real vehicle calibration tests to ensure that the calculation results can accurately reflect the degree to which the driver is actually occupied by non-driving tasks.
[0039] (2) Forward Visual Regression (FVR) This variable reflects whether the driver's vision returns to the road ahead and the driving-related area, i.e., the observable degree of spatial orientation recovery. Its calculation is based on one or more of the following information: probability of forward road gaze, head alignment degree (normalized value of the deviation angle between the head yaw angle and the centering position), and duration of continuous forward gaze. An exemplary calculation formula is as follows: FVR = b 1·P_road + b 2·A_head + b 3·T_front P_road is the forward road gaze probability, which represents the proportion of the total time the driver's gaze is on the windshield and the road ahead relative to a certain observation time window (e.g., the most recent 1 second). This value can be calculated in real time by the gaze tracking algorithm in the driver monitoring system, and the value ranges from 0 to 1. The higher the value, the more the driver pays attention to the road conditions ahead.
[0040] A_head represents the degree of head alignment. This indicator quantifies whether the head is facing forward based on the driver's head yaw angle (i.e., the angle of head rotation around the vertical axis). When the absolute value of the head yaw angle is less than 15 degrees, the head is considered to be basically aligned, and A_head is set to 1. When the absolute value of the yaw angle is greater than 30 degrees, the head is considered to be significantly deviating from forward, and the value is 0. For yaw angles between 15 and 30 degrees, a linear interpolation method is used to calculate the intermediate value. This angle data is also provided by the head attitude estimation module of the driver monitoring system.
[0041] T_front is a normalized value for the duration of continuous forward gaze, which measures the driver's sustained stability in keeping their gaze directed toward the road ahead. It detects the duration for which the driver continuously gazes at the area in front (the area of the windshield and the road). If the duration is 2 seconds or more, it is considered that a stable forward gaze has been established, and T_front is set to 1. If the duration of continuous gaze is less than 0.2 seconds (e.g., only a brief saccade), it is set to 0. For durations between 0.2 seconds and 2 seconds, it can be normalized to the 0~1 range using a linear or nonlinear mapping (e.g., 0.25 for a duration of 0.5 seconds).
[0042] The three indicators mentioned above are each multiplied by a pre-defined weighting coefficient. b 1. b 2. b The summation of the three values yields the forward visual regression score. The weighting coefficients range from 0 to 1, and the sum of the three values is typically 1. For example... b 1 = 0.5 b 2 = 0.3 b =3=0.2, where the probability of forward gaze has the highest weight, reflecting the dominant role of gaze direction in determining spatial orientation recovery. The degree of head alignment and duration of sustained gaze serve as auxiliary indicators, jointly reflecting the completeness of the driver's visual attention returning from forward gaze. All input data can be collected and calculated in real time through the driver monitoring system standard on mass-produced vehicles, without the need for additional hardware.
[0043] (3) Hands-on readiness (HPR) This variable reflects the driver's hand and body readiness to reclaim control, i.e., the observable progress of action preparation. Its calculation is based on one or more of the following information: steering wheel grip, distance between hands and steering wheel (measured via DMS or hand-tracking camera; a distance less than 5cm is considered close), and torso lean (angle between torso and seat back; higher value is taken when leaning forward). An exemplary calculation formula is as follows: HPR = d 1·G_wheel + d 2·D_hand-wheel -1 +d 3·L_torso G_wheel represents the steering wheel grip status. This is a binary indicator. It takes a value of 1 when the driver has at least one hand gripping the steering wheel (usually detected by a steering wheel capacitive sensor or torque sensor to detect effective contact or torque), and a value of 0 otherwise.
[0044] D_hand-wheel represents the straight-line distance between the driver's hand and the steering wheel, measured in centimeters (cm). It is measured in real-time by a hand-tracking camera or depth sensor in the driver monitoring system, representing the shortest distance from the hand closest to the steering wheel to the rim or spokes of the steering wheel. Since a larger value indicates the hand is further away from the steering wheel and the lower the readiness to take over, its reciprocal (D_hand-wheel) is used in the formula. -1 To construct a positive correlation, for example, the reciprocal is 0.2 when the distance is 5cm, 0.5 when the distance is 2cm, and 1.0 when the distance is 1cm. Then, the reciprocal result is normalized to the 0~1 range (usually the maximum effective distance is set to 10cm, and the normalized value of the reciprocal is 0 when the distance is greater than 10cm).
[0045] L_torso represents the degree of torso forward tilt, reflecting whether the driver's upper body is leaning forward from the backrest position to approach the steering wheel. This indicator can be obtained through skeletal key point detection or seat pressure distribution by the driver monitoring system: when the angle between the torso and the vertical direction (i.e., the plumb line) is less than 10 degrees, the driver is considered to be in a significantly forward tilted posture, and the value is 1; when the angle is greater than 45 degrees, the driver is considered to be in a leaning or relaxed state, and the value is 0; for angles between 10 degrees and 45 degrees, linear interpolation can be used to calculate the intermediate value (e.g., 0.5 for an angle of 27.5 degrees).
[0046] The three indicators mentioned above are each multiplied by a pre-defined weighting coefficient. d 1. d 2. d The three factors are then summed to obtain the hand-to-hand control readiness level. The weighting coefficients range from 0 to 1, and the sum of the three is usually 1. For example... d 1 = 0.5 d 2 = 0.3 d 3=0.2, with the steering wheel grip state having the highest weight because it is the most direct evidence of hand takeover readiness. Hand distance and torso lean forward serve as auxiliary indicators, jointly depicting the driver's progress from body posture to action execution. All input data can be acquired in real time through the capacitive steering wheel, driver monitoring camera, and seat sensors standard in mass-produced vehicles.
[0047] (4) Auditory takeover response (ATR) This variable reflects the degree of observable behavioral change in a driver after receiving active audio prompts, i.e., the driver's sensitivity to auditory guidance signals. Its calculation is based on factors including the speed of head return to center after active audio triggering, the speed of gaze shift, and the quantitative value of the behavior of ceasing to look at the screen. An exemplary calculation formula is as follows: ATR = c 1·R_head + c 2·R_eye + c 3. Q_screen-off R_head is the normalized value of the head yaw rate, which reflects how quickly the driver's head turns forward after the active sound trigger. Specifically, it is obtained by continuously tracking the rate of change of the head yaw angle through the driver monitoring system: within a preset time window (e.g., 200ms to 600ms) after the sound trigger, the maximum angular velocity (in degrees / second) of the absolute value of the head yaw angle is calculated, and this angular velocity is normalized relative to an empirical maximum reference value (e.g., 200 degrees / second) so that its value falls between 0 and 1. The larger the angular velocity, the faster the driver responds to the spatial orientation guidance of the sound signal.
[0048] R_eye represents the forward gaze shift speed, which characterizes the rate at which the driver's gaze point moves from a non-frontal area (such as the central control screen, side window, mobile phone, etc.) towards the windshield and road. Using an eye-tracking algorithm, within the same time window after the sound is emitted, the rate of change of the angle between the gaze direction vector and the forward axis is calculated and normalized to the range of 0 to 1. The higher the value, the more agile the driver's visual regression under auditory guidance.
[0049] Q_screen-off is a quantitative value for stopping screen-looking behavior. This indicator is specifically used to assess whether the driver actively stops looking at the screen (central control screen, mobile phone, or other non-driving display areas) due to the sound signal when the driver is looking at the screen before the sound is given. It is obtained by: first, determining whether the driver's gaze is on the screen at the most recent time point before the sound (e.g., within 200ms before the sound). If so, then detecting whether the driver's gaze leaves the screen within 500ms after the sound and remains away for at least 200ms. If the conditions are met, Q_screen-off takes a higher value (e.g., 0.9 or 1.0). If the driver did not look at the screen before the sound or did not stop looking at the screen in time after the sound, a lower value (e.g., 0 to 0.2) is taken.
[0050] The three indicators mentioned above are each multiplied by a pre-defined weighting coefficient. c 1. c 2. c 3. (The three values range from 0 to 1, and are usually summed to 1) Then sum them to obtain the auditory takeover response. Weighting coefficients. c 1. c 2.c 3. It can be obtained through optimization through real vehicle calibration tests. For example, by collecting response data of different drivers to various active sounds, the optimal coefficient can be determined with the goal of maximizing the takeover success rate and minimizing the reaction time.
[0051] (5) Motion Consistency Response (MCR) This variable reflects whether the driver has re-established driving coordination related to vehicle movement, i.e., the level of multi-sensory integrated coordination. Its calculation is based on the coordination between head movement and vehicle yaw changes, the consistency between the line of sight and the forward trajectory of the road, and the coordination between torso response and changes in vehicle longitudinal and lateral acceleration. An exemplary calculation formula is as follows: MCR = e 1·C_head-yaw + e 2. C_gaze-path + e 3·C_torso-acc C_head-yaw is the correlation coefficient between head yaw rate and vehicle yaw rate. This index is used to assess whether the driver's head rotation is coordinated with the vehicle's lateral sway trend. Specifically, it is obtained by continuously sampling the driver's head yaw rate (in degrees / second) through the driver monitoring system within a preset time window (e.g., 1 to 2 seconds), and simultaneously acquiring the vehicle's yaw rate (in degrees / second) within the same time window from the vehicle bus. Then, the Pearson correlation coefficient between the two sets of time series data is calculated. The correlation coefficient ranges from -1 to 1. A positive value indicates that the head rotation direction is consistent with the vehicle's yaw direction (e.g., the head turns left when the vehicle turns left), and the closer the absolute value is to 1, the higher the coordination. In actual driver takeover scenarios, the absolute value is usually taken, or the positively correlated part is normalized to obtain a C_head-yaw value between 0 and 1. A higher value indicates that the driver's spatial orientation perception is more synchronized with the vehicle's actual movement.
[0052] C_gaze-path represents the matching degree between the driver's gaze direction and the curvature of the road ahead. It measures whether the driver's gaze point moves along the expected trajectory of the road ahead. First, the curvature of the current lane line or the bending direction and radius of the future short-term path are obtained through a forward-facing camera or navigation map. Simultaneously, the horizontal angle of the driver's gaze and its changing trend are obtained through an eye-tracking algorithm. Then, within the same time window, the consistency between the angle of the gaze direction relative to the vehicle's longitudinal axis and the angle of the tangent direction of the road ahead is calculated. This is typically quantified using the cosine of the difference between the two angles or a probability-based matching score. For example, when the driver's gaze smoothly follows the inside of the curve, the matching degree is close to 1; if the gaze is fixed or looking towards the outside of the curve, the matching degree is close to 0.
[0053] C_torso-acc represents the synergy between torso roll and vehicle lateral acceleration, reflecting the degree to which the driver's upper body posture naturally adapts to the vehicle's lateral dynamics. Using a torso posture estimation algorithm in the driver monitoring system, the torso roll angle or roll angular velocity (typically the angle between the shoulder line and the horizontal plane) is obtained, while lateral acceleration signals are acquired from the vehicle bus. The temporal correlation between the two is calculated (using a cross-correlation function or synchronous phase difference). When the vehicle accelerates to the right during a turn, the driver's torso naturally tilts to the left to resist centrifugal force; this reverse synergy is also an adaptive response. Therefore, the absolute value of the correlation between the torso roll direction and the lateral acceleration direction, or a sign-adjusted synergy coefficient, is typically calculated and normalized to the 0-1 range.
[0054] The three indicators mentioned above are each multiplied by a pre-defined weighting coefficient. e 1. e 2. e 3 (values range from 0 to 1, usually summing to 1) are summed to obtain the motion consistency response. This proxy variable comprehensively reflects whether the driver switches from passive occupant mode to active driving cooperative mode, and is an important reference indicator for judging the depth of sensory integration recovery. All input data can be acquired in real time through the driver monitoring system (cameras, gyroscopes, etc.) and vehicle bus of mass-produced vehicles, without the need for additional sensors.
[0055] III. Determining the Current Status of Takeover Participation Based on the preset state determination rules, and according to one or more takeover participation agent variables calculated above, the driver's current takeover participation state is uniquely determined to be one of the following five consecutive states: Non-driving immersion state (S0): The driver is completely immersed in a non-driving task, and vision, vestibular system, and proprioception are decoupled. Judgment rules: at least two of the following must be met: NDO ≥ first threshold (e.g., 0.7); FVR ≤ second threshold (e.g., 0.3); HPR ≤ third threshold (e.g., 0.3).
[0056] Task Disengagement State (S1): The driver begins to disengage from a non-driving task but has not yet established forward orientation. Judgment Criteria: NDO decreases by more than a predetermined amount (e.g., more than 30% decrease compared to the initial value); FVR has not yet reached the fourth threshold (e.g., 0.6) or ATR is higher than the set threshold (e.g., 0.4).
[0057] Forward orientation recovery state (S2): The driver's vision has returned to the road ahead, but the hands are not yet fully prepared. Judgment rule: FVR reaches the fifth threshold (e.g., 0.7); HPR is below the sixth threshold (e.g., 0.4).
[0058] Action Readiness State (S3): The driver has established forward orientation and begun preparations for hand and torso maneuvers, but the takeover completion conditions have not yet been fully met. Judgment Rules: FVR is higher than the fifth threshold; HPR continues to rise and is higher than the seventh threshold (e.g., 0.6); and the takeover confirmation operation has not yet been triggered.
[0059] Takeover Ready State (S4): The driver has completed takeover preparation and can safely hand over control. Judgment Criteria: At least two of the following must be met: Steering wheel grip is established; Forward gaze duration reaches a threshold (e.g., 1 second); HPR is higher than the eighth threshold (e.g., 0.8); The driver triggers a takeover confirmation operation (e.g., pressing the steering wheel confirmation button or applying steering torque).
[0060] In practical applications, state determination can employ rule-based models (such as the threshold rule mentioned above), probabilistic models (such as Bayesian classifiers), or machine learning models (such as support vector machines and lightweight neural networks). Rule-based models are preferred to ensure real-time performance and interpretability.
[0061] To ensure repeatability and universality across different vehicle models and driver groups, the weighting coefficients and state determination thresholds of the aforementioned proxy variables need to be determined through standardized real-vehicle calibration tests. The optimal calibration process is as follows: (1) Recruit test drivers: cover different ages (20-60 years old), different driving experience (novice to experienced driver), and different genders, with a sample size of no less than 50 people, and build a calibration sample library covering the mainstream passenger car user group.
[0062] (2) Design calibration scenarios: covering typical non-driving scenarios such as watching movies (central control screen / mobile phone), making calls (Bluetooth / in-vehicle), resting with eyes closed, and distracted screen viewing (such as operating navigation), as well as takeover events of low urgency (remaining takeover time > 5 seconds), medium urgency (3~5 seconds), and high urgency (< 3 seconds). Tests were conducted on real vehicles or high-fidelity driving simulators to collect full data on the driver takeover process and takeover results (takeover success rate, takeover reaction time, startle reaction rate, etc.) in each scenario.
[0063] (3) Optimize weight coefficients: With the optimization objectives of "highest takeover success rate, shortest takeover reaction time, and lowest startle reaction rate", multiple linear regression or genetic algorithm is used to solve for the optimal values of the weight coefficients of each proxy variable. The optimization objective requires that the matching degree between the calculated results of the proxy variables and the actual takeover readiness state of the driver is not less than 90%.
[0064] (4) Determine the state judgment threshold: Based on the statistical distribution of proxy variables in different takeover stages in the calibration sample, the percentile method or ROC curve is used to determine the optimal segmentation threshold between each state.
[0065] The above calibration ensures that the accuracy of state determination in different scenarios is no less than 95%.
[0066] IV. Generating Active Sound Control Strategies This step generates corresponding active sound control parameters based on the migration relationship between the currently determined takeover participation state and the target takeover participation state. Active sound control parameters include at least one or more of the following: media suppression ratio, sound source location, sound image migration trajectory, frequency band distribution, rhythm parameters, duration, semantic type, and output stage sequence.
[0067] This embodiment designs the active sound generation as four functional sound segments, each corresponding to a specific state transition. The corresponding sound segments are output sequentially according to the state transition order, or some sound segments are skipped according to the fast recovery path.
[0068] (1) Task occupies the stripped sound segment This sound segment is used to reduce non-driving task occupancy, corresponding to the state transition S0→S1. Its design aims to interrupt the current non-driving task occupancy and activate the driver's auditory attention network through sudden changes in sound intensity and sound image transition.
[0069] Specific control parameters include: reducing the volume of current in-vehicle media (such as music and video audio tracks) by 6dB to 15dB (preferably 10dB); outputting a short pulse cue (mid-to-high frequency, duration 100-200ms) in the driver's current focus area (e.g., towards the center console screen); shifting the sound image to the center of the front cabin within 300ms to 700ms; and using short, prominent mid-to-high frequency sound waves (such as 2kHz-4kHz pulse trains). If no media is playing in the vehicle, this sound segment is output directly.
[0070] (2) Forward spatial anchoring sound segment This sound segment is used to establish forward spatial orientation, corresponding to the state transition S1→S2. Its design purpose is to guide the driver to reconstruct the forward spatial reference frame, using the human ear's ability to locate the sound source in front to guide the head back to center and the line of sight forward.
[0071] Specific control parameters include: outputting a continuous and stable sound anchor point (such as continuous low-intensity broadband noise or gentle rhythmic sound) at the speaker position corresponding to the center axis of the vehicle's windshield (such as the center of the center console or behind the instrument panel); superimposing a weak lateral auxiliary sound image according to the direction of risk (for example, if the risk comes from the right front, an extremely low-intensity auxiliary sound can be superimposed near the right A-pillar), but the lateral sound image is not used as the main sound source; using non-verbal short rhythms (such as a short "tap" sound every 0.8 seconds) or extremely short confirmation words (such as "ahead"). The duration of this sound segment is generally 800~1500ms until the FVR reaches the threshold.
[0072] (3) Action preparation induction segment This sound segment is used to induce the hands and torso into an operational readiness state, corresponding to the state transition S2→S3. Its design purpose is to activate the approach movement of the hands to the steering wheel, using auditory-spatial positioning to induce hand movement through local sound pulses in the steering wheel area.
[0073] Specific control parameters include: outputting short local pulses (frequency 1kHz~3kHz, pulse width 50~100ms, interval 200~400ms) in the steering wheel area (e.g., the center of the steering wheel or near the steering column) or in the cabin near-field speakers (e.g., on both sides of the instrument panel); gradually shortening the rhythm (e.g., from an interval of 500ms to 200ms) and local acoustic localization to induce the hands to move closer to the steering wheel; optionally, outputting low-frequency pulses (80~150Hz) in the footwell area to assist in activating proprioception. If the driver's torso is not leaning forward, the pulse interval can be shortened to enhance the sense of urgency.
[0074] (4) Takeover confirmation audio segment This sound segment is used to confirm the completion of the takeover, corresponding to the state transition from S3 to S4. Its design purpose is to serve as a positive feedback signal to reinforce the established operational readiness state.
[0075] Specific control parameters include: outputting a short confirmation sound (such as "ding" or "takeover confirmation"); ending the forward anchoring and motion guidance sound segment; triggering a control handover confirmation (such as a prompt "Please drive carefully"). This sound segment typically lasts 200~400ms.
[0076] Based on the transition relationship between the current state and the target state, the aforementioned sound segments are output sequentially. For example, if the current state is S0 and the target state is S1, then the task-occupying stripping sound segment is generated; if the current state is S1 and the target state is S2, then the forward spatial anchoring sound segment is generated, and so on. If the transition fails within the predetermined time, an upgrade process is executed.
[0077] V. Output active sound and perform closed-loop update. The active sound control strategy generated according to the aforementioned steps outputs active sound signals through the in-vehicle speaker system. After each sound segment is output, the system re-acquires driver state information within a preset detection window (typically 600ms~1000ms), updates the takeover participation proxy variables from the aforementioned steps, and re-determines the takeover participation state. If the determination result reaches the target state (e.g., transitioning from S0 to S1), the next stage of transition continues; if unsuccessful, at least one of the following processes is executed: Increase the prominence of the current sound segment: for example, increase the sound pressure level by 3-6 dB, or increase the pulse repetition frequency; Extend the duration of the current audio segment: for example, from 800ms to 1200ms; Switch to a higher priority parameter set: for example, use a more prominent voiceprint or add voice prompts; Further reduce the media volume: for example, lower it by another 5dB or pause media playback; If there is insufficient time remaining before takeover (e.g., less than 1.5 seconds), trigger a minimum risk maneuver (e.g., automatically decelerate and pull over).
[0078] To address complex scenarios such as rapid recovery of driver attention or drastic fluctuations in driver state, this embodiment further employs the following enhancement strategies: (1) Adjustment of quick recovery detection and dynamic detection window The system monitors the rate of change of the proxy variables involved in the takeover process in real time, such as ΔFVR / Δt (change in forward visual regression per second) or ΔHPR / Δt. When the rate of change exceeds a preset rapid recovery threshold (e.g., FVR increases by more than 0.5 per second), it is determined that the driver is rapidly recovering on their own. At this time, at least one of the following operations is performed: the detection window is shortened to 200ms~400ms (the standard window is 600~1000ms); at least one active sound segment corresponding to an intermediate state between the current state and the target state is skipped; if the proxy variable is detected to have reached the threshold corresponding to the next state during the output of the current sound segment, the current sound segment is terminated in advance. It should be noted that a minimum output duration is set for each sound segment, preferably 150ms~200ms, to avoid the sound being too fragmented and losing its guiding effect.
[0079] (2) Fast state transition path The system includes a preset standard path (S0→S1→S2→S3→S4) and multiple fast paths (such as S0→S2→S4, S1→S3→S4). When the change pattern of the proxy variable meets the rapid recovery characteristic (e.g., FVR jumps from 0.2 to 0.7 in a short time while NDO decreases synchronously), the system automatically selects the fast path, skipping the sound segment corresponding to the intermediate state, thereby shortening the overall takeover guidance time.
[0080] (3) Historical state memory and false trigger inhibition Record the state sequence within a recent time window (e.g., 5 seconds). If the driver is detected to enter and exit a non-driving immersion state multiple times in a short period of time (e.g., S0→S2→S0→S2), reduce the trigger sensitivity of the state transition, for example, by increasing the threshold required for the state transition by 0.05~0.1, or by extending the confirmation time, to avoid ineffective intervention caused by the driver's brief distraction and return to the correct state.
[0081] (4) Status confirmation delay mechanism When the fluctuation of the proxy variable involved in the takeover is detected to exceed the threshold within a preset time window (e.g., FVR jumps from 0.3 to 0.7 and then falls back to 0.4 within 300ms), the confirmation time for state determination is extended. For example, the confirmation window is extended from 600ms to 1200ms, and the state transition is performed only after the fluctuation converges, thereby avoiding misjudgments caused by transient noise or unconscious actions of the driver.
[0082] The following examples illustrate the implementation of the present invention in several specific application scenarios.
[0083] Application Scenario 1: Moderate emergency takeover during movie viewing Initial conditions: The vehicle is driving in Level 3 autonomous driving mode. Road construction is underway ahead, and the system issues a takeover request. The takeover time remains at 4.0 seconds. The driver is watching a video on the central control screen, with both hands off the steering wheel and leaning back in the seat.
[0084] Initial data collection and calculation: screen viewing probability P_screen = 0.9, forward gaze probability P_road = 0.05, media playback state M_media = 0.9, steering wheel not held G_wheel = 0, torso not leaning forward L_torso = 0.1. Calculated NDO = 0.78, FVR = 0.18, HPR = 0.12. According to the judgment rule, satisfying NDO ≥ 0.7, FVR ≤ 0.3, and HPR ≤ 0.3, it is judged as S0 non-driving immersion state.
[0085] Phase 1 (S0→S1): The task-occupied audio segment is stripped. The video volume is reduced by 10dB, and two short mid-to-high frequency pulses (3kHz frequency, 80ms pulse width, 150ms interval) are output to the central control screen area. Simultaneously, the audio image shifts towards the windshield's central axis within 500ms. After 700ms, re-acquisition occurs: NDO drops to 0.51 (a 34% decrease), ATR rises to 0.42, and FVR rises to 0.26. The system is determined to have entered the S1 task-detached state.
[0086] Phase Two (S1→S2): Generating a forward spatial anchoring sound segment. A continuous, stable sound anchor point (gentle, broadband noise superimposed with a short "click" sound every 0.6 seconds) is output from the central speaker on the windshield. Since the risk direction is directly forward, lateral auxiliary sounds are not superimposed. After 900ms, data is re-acquired: FVR increases to 0.63, ATR is 0.57, and NDO decreases to 0.24, indicating entry into the S2 forward directional recovery state.
[0087] Phase 3 (S2→S3): Generating the action preparation induction sound segment. A local double pulse (frequency 2kHz, pulse width 60ms, interval 300ms) is output in the steering wheel area. Since the torso has not yet leaned forward significantly, the pulse interval is gradually shortened to 200ms. After 800ms, the sound is re-acquired: HPR rises to 0.68, the torso lean angle decreases, and the hands approach the steering wheel (approximately 3cm away), indicating entry into the S3 action preparation state.
[0088] Phase 4 (S3→S4): Output a confirmation tone (short "ding" sound). The steering wheel capacitive sensor then detects the driver's grip, and the driver applies a slight steering torque, completing the control transfer. Ultimately, FVR=0.74, HPR=0.89, control transfer successful, total time approximately 2.4 seconds.
[0089] Application Scenario 2: High-Emergency Takeover in a Closed-Eyes-Rest State Initial conditions: The driver is resting with their eyes closed, and there are only 1.8 seconds left before the system takes over. The system must avoid a startling reaction caused by a single sharp alarm.
[0090] Initial data collection: Blindness factor S_rest=0.95, NDO=0.82, FVR=0.05, HPR=0.10, classified as S0. Due to the extremely short remaining takeover time, the conventional gradual strategy is skipped, and a high-urgency mode is adopted: First, a gradually increasing low- to mid-frequency composite warning sound is output (frequency linearly increases from 200Hz to 2kHz, duration 0.5 seconds, sound pressure level gradually increases from 50dB to 75dB), prioritizing auditory and head responses; then, a single sound anchor point (high-intensity short pulse) is quickly established on the central axis of the windshield within 100ms; when the DMS detects that the head has been raised but the hands are still not on the wheel, local sound induction of the steering wheel is immediately initiated (pulse interval 200ms, sound pressure level 70dB); if no hand is detected on the wheel within 1.5 seconds, the system triggers a minimum-risk maneuver, automatically decelerates and pulls over to the side of the road.
[0091] Application Scenario 3: Takeover guidance during a call Initial conditions: The driver is making a call via in-vehicle Bluetooth, the remaining takeover time is 3.0 seconds, the central control screen is not playing video, and the call occupancy factor C_call=0.9.
[0092] Initial calculations: NDO=0.65 (mainly from calls), FVR=0.22, HPR=0.15, determined as S0.
[0093] S0→S1 phase: Output a short pulse of medium to high frequency on the non-talking ear side (e.g., the headrest speaker on the driver's left ear), reduce the call volume by 6dB, so that the driver can perceive the need to take over but without interrupting the call.
[0094] S1→S2 phase: Set the acoustic anchor point in the forward area on the same side as the talking ear (e.g., the left front windshield area), and use auditory localization to guide the head back to the center. The call continues but the volume remains low.
[0095] S2→S3 stage: Output local pulses in the steering wheel area to guide hand preparation.
[0096] Ultimately, the driver took over while on a call, and the system did not disconnect the call, only briefly lowered the volume, resulting in a good user experience.
[0097] Application Scenario 4: Driver's rapid self-recovery scenario Initial conditions: 3.5 seconds remaining in the takeover time; the driver only glanced at the central control screen briefly (about 0.3 seconds) before automatically looking up and straightening up.
[0098] Processing flow: t=0ms: Takeover triggered, NDO=0.52, FVR=0.42, at the S0 / S1 boundary. System setting confirmation delay is 100ms.
[0099] t=80ms: Reacquisition, NDO=0.38, FVR=0.56, ΔFVR / Δt = (0.56-0.42) / 0.08 =1.75 / second, exceeding the fast recovery threshold (default is 1.0 / second).
[0100] t=80ms: Trigger the fast recovery mode, cancel the task-occupied stripped sound segment, directly determine it as S2 (forward directional recovery state), and output a shortened forward spatial anchoring sound segment (only 300ms).
[0101] t=380ms: FVR rises to 0.71, HPR rises to 0.63, judged as S3, output action preparation induction sound segment (standard duration 800ms but terminated early).
[0102] t=680ms: HPR rises to 0.86, grip is established, and takeover is complete. The total time is 680ms, saving about 60% of the time compared to the standard path (about 1800~2200ms).
[0103] Application Scenario 5: Delay in State Confirmation under Dramatic Fluctuations of Proxy Variables Initial conditions: The driver frequently observes the rearview mirror and the central control screen, and the FVR fluctuates rapidly between 0.3 and 0.6 (the change exceeds 0.2 every 0.5 seconds).
[0104] Processing flow: If the fluctuation amplitude is detected to exceed the preset threshold (0.15), the status confirmation window is extended from 600ms to 1200ms. The moving average of FVR is continuously observed within 1200ms. Only after the fluctuation converges to 0.55~0.65 and stabilizes for more than 300ms is the process considered to enter S2. This avoids the mistaken assumption that a momentary increase in FVR has been corrected, thus preventing the necessary guiding sound segment from being skipped.
[0105] Example 2 Based on the above method, this embodiment provides an active voice guidance system for autonomous driving takeover participation. For example... Figure 2 As shown, the system mainly consists of a takeover request acquisition module, a status acquisition module, a proxy variable calculation module, a status determination module, an active sound generation strategy module, a media control module, a sound field output module, and a closed-loop update module. The modules communicate with each other at high speed via vehicle bus (such as CAN bus) or vehicle Ethernet, and work together to complete the entire process of guiding the driver to safely take over the vehicle from the issuance of the takeover request.
[0106] The takeover request acquisition module is responsible for monitoring the status of the autonomous driving system in real time. When the system detects that the driver needs to take over the vehicle, this module immediately acquires key takeover request information, including the takeover countdown (i.e., the remaining time available for the driver to prepare for takeover), the risk level of the current takeover event (e.g., low, medium, high urgency), and the real-time operating status of the vehicle, such as vehicle speed, longitudinal and lateral acceleration, and yaw rate. This information is not only used for subsequent calculation of proxy variables, but also provides a basis for the urgency and parameter selection of the active communication strategy. The status acquisition module continuously collects raw data related to driver behavior and the cabin environment from the existing sensor network in the vehicle. Specifically, this includes head posture, eye opening and closing status, gaze direction classification, screen viewing status (whether the driver is looking at the central control screen or mobile phone and other non-driving display areas), hand position and movement trajectory obtained through the driver monitoring system camera; steering wheel grip status obtained through the steering wheel capacitive or torque sensor; torso posture (e.g., leaning forward, leaning back, tilting sideways) obtained through the seat or camera; and cabin media status (e.g., audio or video playback content, volume, call status, etc.) obtained from the entertainment system. The outputs of these two modules together form the basic data source for all subsequent calculations and control.
[0107] The proxy variable calculation module receives data from the state acquisition module and the takeover request acquisition module. Following pre-calibrated formulas, it calculates at least three core takeover participation proxy variables in real time: non-driving task occupancy, forward visual regression, and hand takeover readiness. It can also optionally supplement the calculation with auditory takeover response and kinematic consistency response to further improve the robustness of state determination. Each proxy variable is calculated using a weighted summation method, where the weighting coefficients and thresholds have been determined through the standardized real-vehicle calibration tests described earlier, ensuring that the calculation results accurately reflect the driver's current sensory integration recovery progress. The state determination module, based on the multiple proxy variable values output by the proxy variable calculation module and combined with preset determination rules (such as threshold comparison and decline magnitude determination), uniquely determines the driver's takeover participation state as one of five consecutive states: non-driving immersion, task disengagement, forward orientation recovery, action readiness, or takeover ready. This module incorporates a state confirmation delay mechanism and historical state memory function, which can extend the confirmation window when proxy variables fluctuate drastically, avoiding misjudgments caused by transient noise.
[0108] After obtaining the current state, the active sound strategy generation module dynamically generates corresponding active sound control parameters based on the transition relationship between the current state and the target state (i.e., the next desired state). Specifically, for the transition from a non-driving immersion state to a task disengagement state, the module generates control parameters for the task-occupied stripping sound segment, including the media reduction ratio, the sound source location of the short pulse output, the trajectory and duration of the forward migration of the sound image, etc. For the transition from the task disengagement state to the forward-oriented recovery state, the module generates forward spatial anchoring sound segment parameters, specifying the continuous sound anchor point output of the windshield center axis speaker, the interval of non-verbal rhythms, and whether to superimpose lateral auxiliary sound images. For the transition from the forward-oriented recovery state to the action preparation state, the module generates action preparation induction sound segment parameters, specifying the frequency, pulse width, and rhythm change pattern of the local short pulses of the steering wheel or front cabin near-field speakers. For the transition from the action preparation state to the takeover ready state, the module generates takeover confirmation sound segment parameters, outputting a short confirmation sound and triggering a control handover prompt. This module also supports automatic selection of the fast recovery path. When the rate of change of the proxy variable exceeds the fast recovery threshold, it can skip the audio segment corresponding to the intermediate state and directly generate higher-order guidance parameters.
[0109] The media control module works closely with the active sound strategy generation module. When an output task requires a stripped-off sound band, this module dynamically lowers the volume of the currently playing in-vehicle media (such as music, video audio tracks, or call volume) according to the reduction ratio in the control parameters. If necessary, it can pause or perform frequency band avoidance (e.g., only lowering the frequency band overlapping with the active sound band) to ensure that the active sound signal can be clearly transmitted to the driver without being drowned out by the media content. For other sound bands, the media control module can maintain the media's reduced volume or gradually restore the original volume as needed. The sound field output module is responsible for converting the control parameters output by the active sound strategy generation module into specific speaker drive signals and controlling the various speaker units arranged in the vehicle to emit sound as needed. The system supports speakers including center console speakers (usually located in the center of the dashboard), front cabin speakers (such as near the front door panels or A-pillars), headrest speakers (integrated into the driver's seat headrest), and footwell speakers (located near the driver's foot area). By driving these speakers independently or in combination, the sound field output module can accurately achieve spatial positioning and migration of sound images. For example, it can use headrest speakers to output asymmetrical cues to the driver's ears, or use front cabin speakers to establish stable acoustic anchor points on the central axis of the windshield.
[0110] The closed-loop update module is the core of the adaptive control of the entire system. It continuously monitors the driver's response to the actively emitted sounds. After each actively emitted sound segment is completed, within a preset detection window (standardly 600 to 1000 milliseconds, but can be dynamically adjusted to 200 to 400 milliseconds depending on the rapid recovery state), the closed-loop update module triggers the state acquisition module and the surrogate variable calculation module to reacquire the driver's state and update the surrogate variables. The updated surrogate variables are then fed back to the state determination module to re-determine the current takeover participation state. If the determination result indicates a successful state transition (e.g., successfully transitioning from a non-driving immersion state to a mission disengagement state), the closed-loop update module instructs the actively emitted sound strategy generation module to enter the next state transition phase. If the transition fails within the predetermined detection window, the module will perform a series of upgrade processes, such as increasing the salience of the current sound segment (increasing the sound pressure level or changing the rhythm), extending the output duration of the current sound segment, switching to a higher priority parameter set, or directly triggering a minimum risk maneuver when the remaining takeover time is severely insufficient. In addition, the closed-loop update module also has a built-in fast recovery detection function. When the rate of change of the proxy variable exceeds the preset threshold, it automatically shortens the detection window, skips intermediate state segments, or terminates the current segment in advance (while ensuring that the minimum output duration of each segment is not less than 150 milliseconds), thereby achieving efficient fast path guidance.
[0111] Through the coordinated work of the above modules, the entire system takes the driver's takeover participation state as the direct control target and active voice as the guiding means, forming a closed-loop control loop of data collection, calculation, judgment, voice, and re-collection, and continuously iterating until the driver reaches the takeover ready state and completes the switch of control.
[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for proactively guiding autonomous driving takeover in an autonomous driving mode, characterized in that, Including the following steps: S1. Respond to takeover requests and obtain multi-dimensional status information, including takeover request information, driver status information, in-vehicle media status information, and vehicle operation status information; S2. Calculate the takeover participation proxy variables based on the multi-dimensional state information, including non-driving task occupancy, forward visual regression, and hand takeover readiness; the non-driving task occupancy, forward visual regression, and hand takeover readiness together constitute a quantitative representation of the driver's sensory integration recovery degree, which is used to reflect the process of the driver transitioning from a non-driving state of visual-vestibular-proprioceptive decoupling to a driving coordination state. S3. Based on the takeover participation agent variable, determine the driver's takeover participation state as one of multiple consecutive states, including non-driving immersion state, task disengagement state, forward orientation recovery state, action preparation state, and takeover ready state. S4. Based on the migration relationship between the current determined takeover participation state and the target takeover participation state, generate the corresponding active voice control strategy. When it is determined that the takeover participation state needs to migrate from the current state to the next state, generate the corresponding active voice segment. S5. Output an active sound signal according to the active sound control strategy, and collect the driver status information again after a preset time, update the takeover participation agent variable, and re-determine the takeover participation status until the takeover is completed. The non-driving task occupancy rate is calculated based on one or more of the following information: the probability that the driver's gaze is directed at the non-driving display area, the activity status of video or entertainment content, the call status, and the driver's closed eyes or resting status; the forward visual regression rate is calculated based on one or more of the following information: the probability of looking at the forward road, the degree of head alignment, and the duration of continuous forward gaze. The hand takeover readiness is calculated based on one or more of the following information: steering wheel grip status, distance between hand and steering wheel, and degree of torso leaning forward. The formula for calculating the occupancy rate of non-driving tasks is as follows: NDO = a 1·P_screen + a 2·M_media + a 3·C_call + a 4·S_rest Here, NDO represents the non-driving task occupancy rate; P_screen is the probability that the driver's gaze is directed towards the non-driving display area. This probability is obtained by capturing the driver's eyes and gaze direction in real time through the in-vehicle driver monitoring system's camera and calculating the proportion of time the gaze falls on the non-driving related screen area within a unit time window; M_media is the media content occupancy intensity, dynamically determined based on the playback status information of the entertainment host; C_call is the call status occupancy factor, obtained by determining whether the driver is in a call state through the in-vehicle communication module or Bluetooth connection status; S_rest is the closed-eye or rest state factor, obtained by analyzing the degree of eye opening and closing and the duration of eye closing through the driver monitoring system. a 1. a 2. a 3. a 4 represents the pre-defined weighting coefficient; The formula for calculating the forward visual regression degree is: FVR = b 1·P_road + b 2·A_head + b 3·T_front Wherein, FVR represents forward visual regression; P_road is the probability of forward road gaze, which represents the proportion of the total time the driver's gaze falls on the windshield and the road area in front of them within the observation time window; A_head is the degree of head alignment, which quantifies whether the head is facing forward based on the driver's head yaw angle; and T_front is the normalized value of the duration of continuous forward gaze, used to measure the driver's sustained stability in keeping their gaze on the road ahead. b 1. b 2. b 3 represents the pre-defined weighting coefficient; The formula for calculating the readiness of the hand takeover is: HPR = d 1·G_wheel + d 2·D_hand-wheel -1 + d 3·L_torso Among them, HPR represents the readiness of hand takeover; G_wheel represents the steering wheel grip state; D_hand-wheel is the straight-line distance between the driver's hands and the steering wheel; L_torso is the degree of torso lean, which is obtained through skeletal key point detection or seat pressure distribution by the driver monitoring system; d 1. d 2. d 3 represents the pre-defined weighting coefficient.
2. The active sound-guiding method as described in claim 1, characterized in that: The non-driving immersion state refers to the state in which the driver is completely immersed in a non-driving task and has not begun to prepare to take over. The determination condition is that at least two of the following conditions are met: the non-driving task occupancy is not less than the first threshold, the forward visual regression is not greater than the second threshold, and the hand takeover readiness is not greater than the third threshold. The task disengagement state indicates that the driver has begun to disengage from non-driving tasks but has not yet established forward orientation. The determination condition is that the non-driving task occupancy rate decreases by more than a predetermined amount and the forward visual regression rate does not reach the fourth threshold. The forward orientation recovery state refers to a state in which the driver's vision has returned to the road ahead but the hands are not yet fully prepared. The determination condition is that the forward vision recovery degree reaches the fifth threshold and the hand takeover preparation degree is lower than the sixth threshold. The action preparation state indicates that the driver has established forward orientation and has begun to prepare for hand and torso operations. The determination condition is that the forward visual regression degree is higher than the fifth threshold, the hand takeover preparation degree continues to rise and is higher than the seventh threshold, and the takeover confirmation operation has not yet been triggered. The takeover ready state indicates that the driver has completed the takeover preparation and is able to safely hand over control. The determination condition is that at least two of the following conditions are met: the steering wheel is in place, the forward visual regression reaches the threshold, the hand takeover readiness is higher than the eighth threshold, and the driver triggers the takeover confirmation operation.
3. The active vocal guidance method as described in claim 1, characterized in that, The active sound control strategy includes generating active sound control parameters, which include one or more of the following: media volume reduction ratio, sound source location, sound image migration trajectory, frequency band distribution, rhythm parameters, duration, semantic type, and output stage sequence.
4. The active vocal guidance method as described in claim 1, characterized in that, The active sound segments include, in sequence, the task occupancy stripping sound segment, the forward space anchoring sound segment, the action preparation induction sound segment, and the takeover confirmation sound segment. The task occupancy stripping sound segment is used to reduce non-driving task occupancy, including: lowering the volume of the current in-vehicle media, outputting a short pulse prompt in the current attention focus area, and shifting the sound image towards the front of the vehicle; the forward space anchoring sound segment is used to establish forward space orientation, including: outputting a continuous and stable sound anchor point at the speaker position corresponding to the central axis of the vehicle's windshield; the action preparation induction sound segment is used to induce the hands and torso to enter an operation preparation state, including: outputting local short pulses in the steering wheel area or the near-field speakers in the front cabin, inducing the hands to move closer to the steering wheel through rhythmic changes; the takeover confirmation sound segment is used to confirm the completion of takeover, including: outputting a short confirmation sound, ending the forward anchoring and action induction sound segments, and triggering control switchover confirmation; Each sound segment corresponds to a state transition: the task occupancy stripping sound segment corresponds to the transition from non-driving immersion state to task disengagement state, the forward spatial anchoring sound segment corresponds to the transition from task disengagement state to forward orientation recovery state, the action preparation induction sound segment corresponds to the transition from forward orientation recovery state to action preparation state, and the takeover confirmation sound segment corresponds to the transition from action preparation state to takeover ready state; the corresponding sound segments are output sequentially according to the order of state transitions.
5. The active sound-guiding method as described in claim 4, characterized in that, It also includes a rapid recovery detection step: when the rate of change of the takeover participant variable is detected to exceed a preset rapid recovery threshold, at least one of the following operations is performed: the preset duration of the active sound segment is shortened to a duration shorter than the standard duration; Skip the active voice segment corresponding to at least one intermediate state between the current takeover participation state and the target takeover participation state; during the output of the current voice segment, if it is detected that the proxy variable has reached the threshold corresponding to the next state, terminate the current voice segment in advance, and ensure that the minimum output duration of each voice segment is not lower than the preset lower limit.
6. The active vocal guidance method as described in claim 4, characterized in that, It also includes a status confirmation delay mechanism: when the fluctuation of the takeover participant variable exceeds a threshold within a preset time window, the confirmation time for status determination is extended.
7. The active vocal guidance method as described in claim 4, characterized in that, If the takeover participation state transition is determined to be unsuccessful within the preset detection window, at least one of the following actions will be performed: increase the salience of the current sound segment; Extend the duration of the current audio segment; switch to a higher priority parameter set; reduce the media volume; If there is insufficient time remaining before takeover, trigger a minimum risk maneuver, automatically decelerate, and pull over to the side of the road.
8. An active sound-emitting guidance system based on the method of any one of claims 1 to 7, characterized in that, The system includes the following modules: The takeover requirement acquisition module is used to acquire takeover countdown, risk level, and vehicle operating status; The status acquisition module is used to collect the driver's head posture, eye movement, screen viewing status, steering wheel grip status, torso posture, and cockpit media status. The proxy variable calculation module is used to calculate the proxy variables involved in takeover, including non-driving task occupancy, forward vision regression degree, and hand takeover readiness degree. The status determination module is used to determine the takeover participation status of the driver based on the takeover participation agent variable. The active voice strategy generation module is used to generate corresponding active voice control parameters based on the migration relationship between the currently determined takeover participation state and the target takeover participation state. The media control module is used to suppress, pause, or avoid frequencies of the media according to the active sound control parameters. The sound field output module is used to control the in-vehicle speakers to output corresponding active sound signals; The closed-loop update module is used to adjust the subsequent active sound control strategy based on the driver's response. The modules communicate with each other via vehicle bus or Ethernet to collaboratively complete the closed-loop guidance process from the issuance of the takeover request to the completion of the takeover.
Citation Information
Patent Citations
Generative auxiliary driving takeover prompting method and system based on dynamic scene response
CN118205574A
Vehicle take-over reminding method and device
CN121799440A