A phased automatic driving driver takeover reaction time prediction method and system
By breaking down driver takeover behavior into three stages—physical recovery, perceptual recovery, and cognitive recovery—and combining it with a multimodal fusion prediction model, the problem of difficulty in locating bottleneck stages and insufficient cross-scenario generalization ability in existing technologies is solved, achieving accurate prediction of takeover reaction time and improved safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-02-14
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies struggle to distinguish key stages in the process of autonomous driving driver takeover, such as action preparation, perception regression, and cognitive execution. They are unable to pinpoint bottleneck stages and implement targeted interventions. They rely on limited signal sources or a small number of features, and their generalization ability across drivers and scenarios is insufficient. They also struggle to continuously update the remaining reaction time during takeover and align it with the timing of active safety interventions. Furthermore, the lack of unified and reproducible calculation rules for stage labels affects the consistency of model training and evaluation.
The driver takeover behavior is decoupled into three stages: physical recovery, perceptual recovery, and cognitive recovery. Standardized three-stage reaction time labels are generated. By combining static individual attributes, quasi-static environmental features, and dynamic temporal representations through a multimodal fusion prediction model, a prediction model based on a time sliding window mechanism is constructed to achieve phased prediction and online assessment of takeover reaction time.
This approach achieves structural decoupling of the takeover process, improves the model's predictive accuracy and generalization ability, ensures the interpretability and diagnosability of the takeover assessment, reduces takeover risk, increases safety margin, and mitigates consistency issues between model training and evaluation.
Smart Images

Figure CN122154435A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of driver takeover behavior modeling and prediction technology, and in particular to a phased autonomous driving driver takeover reaction time prediction method and system. Background Technology
[0002] In Level 2 (L2) hybrid assisted driving, the autonomous driving system needs to issue a takeover prompt to the driver when approaching operational boundaries or identifying potential risks. The driver must then transition from a non-driving state to an effective control state within a short period. The control handover phase often involves multiple coupled processes, including attention return, action preparation, and control decision-making. Insufficient driver preparation or delayed reaction can easily lead to safety risks such as rear-end collisions, lane departures, and conflicts with vulnerable road users. Therefore, predicting the takeover response time during the control handover process and conducting online assessments of takeover capabilities are crucial foundations for optimizing prompting strategies, implementing proactive safety triggers, and designing safety redundancy.
[0003] However, existing takeover capability assessment technologies generally suffer from the following shortcomings: First, they often estimate takeover reaction time as a single overall indicator, lacking a structured mechanistic characterization of the takeover process. This makes it difficult to distinguish key stages such as action preparation, perceptual regression, and cognitive execution, thus failing to pinpoint bottleneck stages and implement targeted interventions. Second, takeover behavior is influenced by individual driver differences, contextual scenarios, and temporal fluctuations in physiological and behavioral states. However, existing methods often rely on limited signal sources or a small number of features, lacking a mechanism to uniformly integrate static individual attributes, quasi-static environmental conditions, and dynamic physiological behavioral sequences, resulting in insufficient generalization capabilities across drivers and scenarios. Third, stage labels lack unified and reproducible calculation rules, especially when multiple parts of the body are involved in parallel actions, easily leading to repeated timing errors and affecting the consistency of model training and evaluation. Fourth, most methods rely on one-time estimation, making it difficult to continuously update the remaining reaction time during the takeover process and align it with the timing of active safety interventions.
[0004] Chinese invention patent CN114882477A discloses a driver takeover time prediction scheme based on driver eye movement information. The scheme includes: collecting driver eye movement data during driving and simultaneously measuring the driver's reaction time during steering and braking operations; classifying the driver's eye movement data according to saccade angles; constructing a mapping model between eye movement data and driving operations, using which the predicted driver's reaction time during autonomous driving takeover can be obtained. This invention collects driver eye movement data during driving, obtains driving eye movement data in the driver's viewpoint image coordinate system, and substitutes it into a regression equation to obtain the predicted autonomous driving takeover time. The invention collects eye movement features such as driver fixation point, fixation area distribution, gaze shift, blinking, eyelid opening and closing, and pupil changes; preprocesses and extracts features from the eye movement data; and inputs these features into a prediction model to output the driver takeover time. However, there are still problems such as difficulty in distinguishing key links such as action preparation, perception regression and cognitive execution, inability to locate bottleneck stages and implement targeted interventions, reliance on limited signal sources or a small number of features, insufficient generalization ability across drivers and scenarios, difficulty in rolling updates of remaining reaction time and aligning with the timing of active safety intervention during takeover, and lack of unified and reproducible calculation rules for stage labels, which affect the consistency of model training and evaluation.
[0005] In summary, there is currently a lack of a phased driver takeover reaction time prediction method and system for autonomous driving to solve or partially solve the above problems. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a phased driver takeover reaction time prediction method and system for autonomous driving. This method aims to solve or partially solve the problems of difficulty in distinguishing key links such as action preparation, perception regression and cognitive execution, inability to locate bottleneck stages and implement targeted interventions, reliance on limited signal sources or a small number of features, insufficient generalization ability across drivers and scenarios, difficulty in updating the remaining reaction time during takeover and aligning it with the timing of active safety intervention, and the lack of unified and reproducible calculation rules for stage labels, which affects the consistency of model training and evaluation.
[0007] The objective of this invention can be achieved through the following technical solutions: According to one aspect of the present invention, a phased driver takeover reaction time prediction method for autonomous driving is provided, specifically including: S1. Perform structured analysis and phased decoupling of driver takeover behavior, decoupling driver takeover behavior into three stages: physical recovery, sensory recovery, and cognitive recovery. S2. Obtain the action elements and calculation rules corresponding to the reaction time of each stage of the driver takeover behavior, and generate standardized three-stage reaction time labels; S3. Based on the three-stage reaction time tags, hierarchical encoding is performed on the dynamic temporal physiological and behavioral signals before and after takeover to obtain a dynamic temporal representation that simultaneously exhibits short-term disturbances and long-term trends. S4. Using static individual attributes, quasi-static environmental characteristics, and the dynamic temporal representation as inputs, and the advance prediction results, instant prediction results, and continuous prediction results of the three-stage reaction time as outputs, construct a multimodal fusion prediction model based on the time sliding window mechanism. S5. Utilize the trained prediction model to perform advance, real-time, and continuous prediction of takeover reaction time, and deploy it on the human-in-the-loop simulation platform and the autonomous vehicle terminal to realize the prediction of the autonomous driver's staged takeover reaction time and drive the online evaluation of takeover capability.
[0008] As a preferred technical solution, the action elements and calculation rules corresponding to the reaction time of each stage specifically include calculable stage boundary events. These stage boundary events include completing basic driving posture preparation, returning the gaze to the driving area, and completing the first effective control operation. The stage boundary events are given based on a decision function of multimodal action signals. It manifests as: In the formula, For multimodal data at time The observation, For threshold or rule parameters, For indicator functions, For the stage The decision function, This indicates the first phase of attitude preparation. This indicates a return to focus in the second phase. This indicates that the third stage marks the first successful completion of effective control operations.
[0009] As a preferred technical solution, the process of generating the standardized three-stage reaction time tag includes: The reaction time for a given stage is obtained by overlapping and merging the duration intervals of action elements within a stage, after deduplication, as the total effective action duration. The action elements include hand, foot, eye, head, and full-body posture movements. The start and end points of these action elements are identified using multimodal acquisition devices, including eye-tracking, head posture, steering wheel / pedal interaction, and seat posture / pressure acquisition. satisfy: In the formula, For the first The set of action elements included in a stage For the first The start and end ranges of each action element This indicates the total duration of the interval.
[0010] As a preferred technical solution, the hierarchical coding adopts a hierarchical multi-receptive-field temporal coding structure, which simultaneously models short-term disturbances and long-term trends in dynamic time series, and obtains dynamic time series representation by combining a global dependency modeling mechanism. The hierarchical coding process satisfies the following: The hierarchical encoding yields a dynamic temporal representation with enhanced global dependencies. This manifests as: In the formula, For the first The dynamic temporal feature representation of the output of hierarchical coding. For layer index, The total number of layers in the hierarchical coding. Indicates the first Layer-time convolution or dilation convolution operations, Indicates downsampling aggregation, Represents a nonlinear mapping. This represents a global dependency modeling module that includes self-attention and gated aggregation.
[0011] As a preferred technical solution, the multimodal fusion prediction model embeds the driver's static individual attributes and quasi-static environmental features as modulation conditions into the dynamic temporal representation, enabling the model to capture dynamic physiological behavioral features while fusing individual differences and scene context. The static individual attributes and quasi-static environmental features are respectively encoded into condition vectors: The dynamic time series representation is fused using a conditional modulation function: In the formula, This is a conditional vector of static individual attributes. For static individual attributes, The conditional vector for quasi-static environmental characteristics. Characteristics of a quasi-static environment. This is a characterization based on the fusion of dynamic and static coupling. and It consists of a multilayer perceptron, an embedding layer, and a normalization layer. This indicates cross-attention conditional injection, which is then combined with residuals and gating for fusion.
[0012] As a preferred technical solution, the multimodal fusion prediction model couples dynamic and static information through a multimodal interactive fusion mechanism. The interactive fusion mechanism includes conditional modulation and temporal interactive aggregation of dynamic temporal representations, and outputs the reaction times of posture recovery, perception recovery and cognition recovery by stage-specific decoders respectively, so as to realize multi-task joint prediction in stages.
[0013] As a preferred technical solution, the immediate prediction result is generated using data from a fixed window before the takeover notification; the early prediction result is generated by introducing an early prediction window before the takeover notification; and the remaining reaction time is dynamically updated by moving the prediction window backward to generate a continuous prediction result. The fixed window... It manifests as: The advance prediction window It manifests as: The backward prediction window It manifests as: In the formula, The time when the takeover notification is issued, For fixed window length, To predict the offset in advance, To continuously update the offset.
[0014] As a preferred technical solution, the online assessment of takeover capability includes: aligning the remaining takeover reaction time obtained from the continuous prediction results with the timing of active safety intervention; when the prediction results indicate that the driver cannot complete effective takeover before the latest intervention time of the active safety system, triggering active safety braking or control intervention.
[0015] As the preferred technical solution, the reaction times at each stage are as follows: In the formula, This refers to the reaction time during the physical recovery phase. The reaction time is the time required for the sensory recovery phase. This refers to the reaction time during the cognitive recovery phase. For the overlap merging operator, , and These are the number of movement elements for the hands, feet, and whole body, respectively. and For the first The start and end times of each hand movement and For the first The start and end times of each foot movement. and For the first The start and end times of each posture adjustment action. and The number of head and eye movement elements, respectively. It represents the union of sets.
[0016] According to another aspect of the present invention, a phased autonomous driving driver takeover reaction time prediction system is provided. The system is used to execute the above-described autonomous driving driver phased takeover reaction time prediction method. The system includes: a phased decoupling module, a reaction time generation module, a hierarchical coding module, a prediction module, and an output module. The phased decoupling module is used to decouple the driver takeover behavior into three stages: physical recovery stage, sensory recovery stage, and cognitive recovery stage. The reaction time generation module is used to calculate the reaction time of the three stages; The hierarchical coding module is used to perform hierarchical multi-receptive field temporal coding and global dependency modeling on the driver's dynamic temporal data to generate dynamic temporal representations. The prediction module is used to encode static individual attributes and quasi-static environmental features into conditional vectors, inject the conditional vectors into the dynamic temporal representation through cross-attention to generate a dynamic-static fusion representation, perform temporal aggregation on the dynamic-static fusion representation, and predict the takeover response time through the decoding head. The output module is used to output the prediction results generated by the prediction module.
[0017] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) Clear phased mechanism and identifiable bottleneck: This invention decouples the takeover process into three phases: physical recovery, sensory recovery and cognitive recovery, and sets calculable phase boundary events. This solves the problem of difficulty in distinguishing key links such as action preparation, sensory return and cognitive execution, and the inability to locate bottleneck phases and implement targeted interventions. It achieves automatic determination of phase start and end and location of bottleneck phases. Compared with the scheme that only outputs a single total takeover time, it can clearly locate the specific bottleneck phase and significantly improve the interpretability and diagnosability of takeover assessment.
[0018] (2) Improve model accuracy and prediction accuracy: This invention integrates static individual attributes, quasi-static scene context and dynamic physiological behavior temporal representation into a multimodal fusion prediction model, which solves the problem that existing methods rely on limited signal sources or a small number of features and have insufficient generalization ability across drivers and scenes. While maintaining sensitivity to short-term disturbances and long-term trends, it significantly improves prediction accuracy and robustness under cross-driver and cross-scene conditions.
[0019] (3) Improve driving safety: Based on the sliding window, this invention realizes the advance prediction, real-time prediction and continuous rolling update of the three-stage reaction time, and aligns with the latest intervention time of active safety. When the prediction indicates that the driver cannot complete the effective takeover within the safe time limit, the control intervention is triggered and the takeover prompt optimization is linked, forming a closed-loop application of prediction-evaluation-intervention. This solves the problem that existing methods are mostly based on one-time estimation and it is difficult to continuously update the remaining reaction time and align it with the intervention time of active safety during the takeover process. This achieves the technical effect of improving safety margin and reducing takeover risk.
[0020] (4) Unified and reproducible stage labels: This invention generates time intervals within a stage by using elements as the basis, and achieves deduplication timing by merging intervals to avoid repeated timing errors caused by parallel actions of multiple parts. This results in reproducible and standardized three-stage reaction time labels, solving the problem of the lack of unified and reproducible calculation rules for stage labels, especially when there are parallel actions of multiple parts, which easily leads to repeated timing errors and affects the consistency of model training and evaluation. This achieves the technical effect of improving the reliability and comparability of model supervised training and performance evaluation. Attached Figure Description
[0021] Figure 1 A schematic diagram illustrating the definition of the three-stage takeover response time and the action elements of MTM-TO; Figure 2 This is a schematic diagram illustrating the real-time, early, and continuous prediction mechanism for the three-stage takeover response time; Figure 3 This is a schematic diagram of the dynamic and static factor fusion framework in the three-stage takeover response time prediction model; Figure 4 This is a schematic diagram showing the distribution characteristics of the three-stage reaction time in the embodiment under different total pipeline reaction time levels; Figure 5 This is a schematic diagram showing the comparison between the predicted reaction times of the three stages and the total control unit reaction time in the embodiment. Figure 6 This is a schematic diagram illustrating the visualization of AEB secondary intervention determination and the results of the intervention experiment in the embodiment; Figure 7 This is a schematic diagram showing the distribution of safety margins for each sample under the early intervention mechanism in the embodiment. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] Example 1 To address the problems existing in the prior art, this embodiment provides a phased method for predicting driver takeover reaction time in autonomous driving, specifically including: S1. Perform structured analysis and phased decoupling of driver takeover behavior, decoupling driver takeover behavior into physical recovery stage, sensory recovery stage and cognitive recovery stage respectively.
[0024] Furthermore, step S1 specifically includes the following sub-steps: S11: Structured analysis of takeover behavior. This involves completing physiological data collection and time alignment, and obtaining the trigger time of the takeover cues. Physiological data observation sequences were collected before, during, and after the takeover process. The sequence is then synchronized, resampled, and missing data is filled to form an observation sequence on a unified time axis. This sequence is then sliced into prediction window inputs to obtain... .
[0025] S12: Decoupling of Takeover Behavior in Stages. Based on the multi-resource theory, the takeover process simultaneously occupies multiple resource channels, exhibiting a staged structure of action preparation - perceptual return - cognitive execution. Accordingly, the takeover behavior is decoupled into three stages: physical recovery, perceptual recovery, and cognitive recovery, thus forming a staged structure and supporting bottleneck localization.
[0026] S13: Stages of the takeover process. The calculable evolutionary boundary events for the three stages include at least: completion of basic driving posture preparation, return of gaze to the driving area, and completion of the first effective control operation. These can be given by a unified form of rule-based judgment: In the formula, For multimodal data at time The observation, For threshold or rule parameters, For indicator functions, For the stage The decision function, This indicates the first phase of attitude preparation. This indicates a return to focus in the second phase. This indicates that the third stage marks the first successful completion of effective control operations.
[0027] The above processing decouples the takeover process and determines the boundaries of each stage by collecting multimodal behavioral action data, thereby preparing for the calculation of the duration of each stage.
[0028] S2. Obtain the action elements and calculation rules corresponding to the reaction time of each stage of the driver takeover behavior, and generate standardized three-stage reaction time labels.
[0029] Furthermore, step S2 specifically includes the following sub-steps: S21: Define the motion elements corresponding to the reaction time of each stage. Configure a set of motion elements for each stage, which should include at least hand, foot, eye, head, and whole-body posture-related movements; and identify the start and end nodes of the motion elements through multimodal acquisition equipment.
[0030] S22: Reaction time calculation rules for each stage. This involves generating a set of action duration intervals for each action element. Identify its start and end nodes and generate action ranges. and in the stage Internally formed action range set .
[0031] S23: Standardized three-stage reaction time label. To avoid duplicate timing due to parallel actions, this embodiment uses interval union timing to obtain the total effective action time after deduplication. The stage reaction time is defined as: In the formula, For the first The set of action elements included in a stage For the first The start and end ranges of each action element This indicates the total duration of the interval.
[0032] This yields the standardized three-stage label vector: It serves as the target output and performance evaluation benchmark for supervised training of the prediction model.
[0033] S3. Based on the three-stage reaction time tags, hierarchical encoding is performed on the dynamic temporal physiological and behavioral signals before and after takeover to obtain a dynamic temporal representation that simultaneously exhibits short-term disturbances and long-term trends.
[0034] Furthermore, step S3 specifically includes the following sub-steps: S31: Standardize the processing of dynamic temporal physiological and behavioral signals. Organize the dynamic physiological and behavioral sequences within the fixed window before the takeover prompt and the optional takeover process sliding window into a unified input tensor. Then, it is standardized, discrete variable embedded, and anomaly truncated to obtain a representation that can be directly input into a timing encoder: In the formula, To standardize preprocessing operators.
[0035] S32: Multi-scale temporal representation. A hierarchical multi-receptive-field encoder is constructed, which progressively expands the temporal receptive field through multi-layer temporal convolution and downsampling aggregation to extract multi-scale representations that take into account both short-term perturbations and long-term trends. In the formula, For the first The dynamic temporal feature representation of the output of hierarchical coding. For layer index, The total number of layers in the hierarchical coding. Indicates the first Layer-time convolution or dilation convolution operations, Indicates downsampling aggregation, This represents a nonlinear mapping.
[0036] S33: Robust dynamic feature input. After multi-scale local encoding, a global dependency modeling module is introduced to integrate long-range dependencies, including self-attention and gated aggregation, to obtain dynamic temporal representations: In the formula, This indicates a global dependency modeling module that includes self-attention and gated aggregation, enabling the output dynamic representation to possess both local sensitivity and long-range consistency.
[0037] S4. Using static individual attributes, quasi-static environmental characteristics, and the dynamic temporal representation as inputs, and the advance prediction results, immediate prediction results, and continuous prediction results of the three-stage reaction time as outputs, construct a multimodal fusion prediction model based on a time sliding window mechanism.
[0038] Furthermore, step S4 specifically includes the following sub-steps: S41: Encoding of static individual attributes and quasi-static environmental features. This involves encoding static individual attributes... Quasi-static environment characteristics Input the conditional encoder separately to obtain the conditional vector: In the formula, This is a conditional vector of static individual attributes. For static individual attributes, The conditional vector for quasi-static environmental characteristics. Characteristics of a quasi-static environment. and It is composed of a multilayer perceptron, an embedding layer and a normalization layer, which puts it in a potential space where it can be integrated with dynamic representation.
[0039] S42: Multimodal fusion. Static and quasi-static conditions are embedded as modulation terms into the dynamic representation, forming a coupled dynamic-static representation. A unified form of conditional modulation and interactive aggregation is adopted: In the formula, This is a characterization based on the fusion of dynamic and static coupling. This represents cross-attention conditional injection, combined with residual and gating for fusion, specifically manifested as: in, For attention mapping, This is an equivalent fusion method combining splicing and linear mapping. In the first-path cross-attention, dynamic temporal representation will be used. The linear mapping is a trainable projection matrix of the query. In the second-path cross-attention, the intermediate representations are... The linear mapping is a trainable projection matrix of the query. , In the first-way cross-attention, the individual condition vector is... A linear mapping is a trainable projection matrix of Key and Value. In the second-path cross-attention, the environmental condition vector is... The linear mapping is the trainable projection matrix of the Query.
[0040] in , Depend on or The copy is extended to the time dimension.
[0041] Through this mechanism, the model can explicitly incorporate individual differences and contextual information while capturing dynamic physiological behavioral changes, thereby improving prediction accuracy and generalization ability.
[0042] S43: Stage-specific decoding and multi-task joint output. This represents the fused representation. The execution sequence aggregation, i.e., attention convergence, yields a global representation u, and the stage-specific prediction heads output the reaction times for each of the three stages: In the formula, For the fusion of time series representation Temporal aggregation operators that perform aggregation along the time dimension The reaction time is divided into three stages. For the first Each stage has its own dedicated prediction head.
[0043] This enables phased, multi-task joint prediction, rather than simply regressing on a single total takeover time.
[0044] S5. Utilize the trained prediction model to perform advance, real-time, and continuous prediction of takeover reaction time, and deploy it on the human-in-the-loop simulation platform and the autonomous vehicle terminal to realize the prediction of the autonomous driver's staged takeover reaction time and drive the online evaluation of takeover capability.
[0045] S51: Model training. Using the labels generated in step S2. To monitor the signal, the model constructed in steps S3-S4 is trained. The training objective uses a weighted combination of multi-task losses, and a regularization term can be introduced to improve robustness. In the formula, The regression loss function includes mean absolute error and mean squared error. For stage weights, These are regularization terms or consistency constraints. For model parameters, The weight coefficients for the regularization term.
[0046] S52: Model Deployment. During the deployment phase, real-time prediction, advance prediction, and continuous prediction are constructed based on a time-sliding window, corresponding to a fixed window before prompts, an advance prediction window, and a shifted prediction window, respectively.
[0047] Fixed window before prompt It manifests as: Early prediction window It manifests as: Shift the prediction window It manifests as: In the formula, The time when the takeover notification is issued, For fixed window length, To predict the offset in advance, To continuously update the offset.
[0048] Extract for each window Repeat steps S3-S4 to output the corresponding prediction results, forming an online update closed loop.
[0049] S53: Safety Intervention Control. Align the continuously predicted remaining takeover capacity with the safety margin to determine the appropriate action. Execute safety intervention control when the triggering conditions are met. The triggering logic can be summarized as follows: In the formula, For the predicted results, For quantities related to risk or safety margin, i.e., distance margin. For the fusion decision function, For the threshold, This is an indicator function used to map the fusion determination result into a binary trigger signal; when It can perform active safety braking in real time and can adjust the intensity, form or semantic information of the prompts in conjunction with the prompts to achieve a closed-loop application of prediction-assessment-intervention.
[0050] S54: Deployment and Verification Methods. The method can be trained based on multimodal data collected from a human-in-the-loop driving simulation platform and deployed on an autonomous vehicle for online inference verification, forming online assessment and safety intervention capabilities for real-world vehicle applications.
[0051] In this embodiment, steps S1-S2 are used to perform structured analysis and phased decoupling of the driver takeover behavior during the control handover period, forming automatically identifiable three-stage boundaries and reproducible stage reaction time labels, thereby characterizing the stage evolution of the takeover process and locating potential bottleneck links.
[0052] To achieve structured analysis and phased decoupling of takeover actions, a Methods-Time Measurement for Take-Over (MTM-TO) motion timing method was developed to decompose and time driver takeover operations at the action level. The core of MTM-TO lies in transforming the driver's behavior during takeover from an overall process that cannot be directly measured into a set of detectable and timeable basic motion elements. This allows for the capture of key time points for each motion element based on multimodal sensor observations, forming a unified basis for phase boundary identification and phase reaction time calculation.
[0053] Specifically, MTM-TO breaks down takeover operations into basic motion elements related to body parts such as hands, feet, eyes, head, and overall posture. Each motion element corresponds to a clear start-end point determination rule and can be detected and identified by at least one sensor signal. Table 1 provides the definitions, determination thresholds / rules, and corresponding detection devices and identification methods for these motion elements to ensure the feasibility and reproducibility of motion node identification.
[0054] Table 1 Definition of Basic Motion Elements of MTM-TO Next, to clarify the action elements and calculation rules corresponding to the reaction time in each stage, Figure 1 The diagram illustrates the correspondence between the three-stage reaction time division and the action elements of MTM-TO. The takeover process exhibits a phased structure of action preparation, sensory recovery, and cognitive execution in terms of resource channel occupancy. Based on this, the takeover behavior is decoupled into three stages: physical recovery, sensory recovery, and cognitive recovery, with each stage corresponding to a combination of action elements for different body parts.
[0055] (1) First stage, posture recovery stage: from receiving the takeover prompt to completing the recovery of posture such as hands, feet and sitting posture and having a basic driving posture; the action elements include hand actions of grasping, releasing and positioning, foot actions of lifting and moving, and whole body posture adjustment actions such as turning and adjusting sitting posture.
[0056] (2) The second stage, the perception recovery stage: from the moment the driver raises his head to initiate the shift of his gaze until the gaze point stabilizes again in the driving-related area; the action elements include the shifting, scanning and gazing of eye movements, and head movements such as raising, turning and tilting the head.
[0057] (3) The third stage, cognitive recovery stage: from the time the driver leaves the non-driving task until the driver completes the first effective control operation, such as the steering wheel or pedals, which are valid inputs that meet the threshold conditions; the action elements include hand control actions, foot control actions, whole body stabilization actions and decision-related eye actions.
[0058] The start and end times of action elements within a stage are identified and interval sets are formed. This is then achieved through an overlap merging operator. The total duration of the interval union is used as the stage reaction time. Figure 1 This illustrates parallel and overlapping scenarios.
[0059] First stage of physical recovery time: In the formula, This refers to the reaction time during the physical recovery phase. For the overlap merging operator, and For the first The start and end times of each hand movement and For the first The start and end times of each foot movement. and For the first The start and end times of each posture adjustment action. , and The number of movement elements for the hands, feet, and whole body, respectively.
[0060] Second-stage sensory recovery time: In the formula, The reaction time is the time required for the sensory recovery phase. and The number of head and eye movement elements, respectively.
[0061] Third stage cognitive recovery time: Finally, based on the three-stage structure of change, bottleneck stages are identified, and intervention is provided as a basis. For example... Figure 4 As shown, the samples can be grouped according to the total take-off reaction time (TRT), and the absolute duration and relative proportion of the three stages in each group can be calculated as follows: The differences between groups were statistically tested (Kruskal-Wallis test) to verify whether the contribution of each stage changed systematically with the TRT level. When the proportion of the third stage increased significantly with the increase of TRT, the cognitive recovery stage could be identified as the key bottleneck affecting the overall difference, providing a stage-specific optimization basis for subsequent prompting strategies and safety interventions.
[0062] In this embodiment, steps S3-S4 are used to hierarchically encode the dynamic temporal physiological and behavioral signals before and after takeover to obtain a robust dynamic representation that simultaneously represents short-term disturbances and long-term trends; step S4 is used to fuse static individual attributes, quasi-static environmental characteristics and dynamic physiological temporal representations into a model, and output the immediate, advance and continuous prediction results of the three-stage reaction time based on the time sliding window mechanism. Figure 2 The diagram illustrates how to construct three types of prediction output windows. Figure 3 This demonstrates a framework for integrating dynamic and static factors. Figure 5 The results show a comparison of the predicted reaction time for the three-stage reaction and the overall take-off reaction time.
[0063] First, to fully utilize the multi-source information on people, vehicles, and scenarios before takeover, this embodiment adopts a static-quasi-static-dynamic multi-channel input organization method. Its core logic is: static and quasi-static features provide individual background and contextual constraints, while dynamic features reflect the real-time evolution of the takeover readiness state. The two complement each other to improve prediction accuracy and generalization ability. Three types of inputs are defined as follows: (1) Static channel, i.e., individual attribute layer: In the formula The length of the vector formed by encoding individual attributes, such as age, driving experience, driving style, gender, driving knowledge, etc. (2) Quasi-static channel, i.e., driving environment layer: In the formula The length of the encoded vector for the driving environment, such as non-driving task type, prompting strategy, scene type, traffic flow variables, or take-off lead, is constant within the sample but varies between samples. (3) Dynamic channel, i.e., sequence layer: In the formula The time window length, Multimodal feature dimensions for each moment; dynamic features aligned with the stage action elements of MTM-TO, such as hand-to-steering wheel distance, foot-to-pedal distance and posture pressure characterizing body readiness, head pitch and eye movement parameters.
[0064] Secondly, to address the modeling contradiction of significant short-term abrupt changes and implicit long-term trends in dynamic signals, this embodiment employs a hierarchical multi-receptive-field temporal encoder. Multi-scale representations are extracted, and global dependency modeling is further introduced to obtain robust dynamic representations.
[0065] In summary, based on the above dynamic coding and dynamic-static fusion structure, this embodiment achieves the following results on the test set: (1) Phased real-time forecasting is better than total time (TRT) forecasting.
[0066] Compared to models that target a single TRT (Tracking Time Reduction), the three-stage model exhibits lower overall error and higher fit in real-time prediction: the first stage and the third stage... The improvement is more significant, and the verification of phased decoupling can reduce the error accumulation caused by variable coupling and provide a finer-grained process characterization, as shown in Table 2.
[0067] Table 2 Comparison of Three-Stage Reaction Time and Main Take-Off Reaction Time (2) The fusion of dynamic and static data is superior to the baselines of machine learning and traditional deep learning.
[0068] Compared to baselines such as XGBoost, LightGBM, LSTM, and DIPNet, the dynamic-static fusion model in this embodiment demonstrates superior overall accuracy across the three stages, showcasing its advantages in collaborative modeling of short-term mutations combined with long-term dependencies, as well as individual differences combined with scene context, as shown in Table 3.
[0069] Table 3 Comparison of prediction results for three-stage reaction time using different models (3) Early prediction is more stable and can obtain earlier available predictions.
[0070] Figure 5 The results show that the three-stage model has a more concentrated error distribution and a more sufficient lead time at the stable prediction time, and can achieve earlier and more reliable prediction output before the takeover warning is issued, thus gaining time margin for warning strategy optimization and proactive safety intervention. Compared with the TRT model, phased modeling can effectively reduce error accumulation and lag in early prediction.
[0071] In this embodiment, step S5, which is used to complete model training, deployment, and verification, means that: on the one hand, multimodal data is collected on the high-fidelity in-loop platform and the model is trained; on the other hand, continuous prediction is performed during the takeover process and online assessment of takeover capability is driven, further aligning with the timing of proactive security intervention to form a closed-loop application.
[0072] First, a human-centered, real-vehicle driving simulation experimental platform was built, forming a multimodal, non-intrusive data acquisition link. The platform uses a real-vehicle cockpit linked with the simulation system; takeover prompts are triggered by the central control screen; steering wheel and pedal operations are acquired through sensors; and non-intrusive data such as eye movements, head movements, hand and foot movements, and posture are collected from multiple sources and uniformly aligned. To ensure that multimodal signals can be used for motion element recognition and stage label generation, all sensor data are synchronized to the same time reference frame and packaged for recording, ensuring cross-modal temporal consistency and reproducibility. Then, a takeover scenario and experimental variable system covering typical high-risk conditions were constructed to obtain representative training samples. In the embodiment, two typical high-risk takeover scenarios were selected: unprotected left-turn conflicts in urban areas and sudden construction avoidance on highways. Parameters such as traffic flow density level and vehicle speed level were set to form a sub-scenario set; simultaneously, three key variables were systematically set: non-driving task type, takeover lead time, and prompting strategy, to cover the differences in takeover under different levels of immersion and warning levels. Furthermore, a multi-factor experimental design, namely a mixed orthogonal design, is adopted to achieve balanced coverage of variable combinations, and the order effect is controlled by Latin squares, etc., to finally form a dataset that can be used for training and evaluation.
[0073] Simultaneously, the continuous forecast results will be used for online assessment of takeover capabilities, and aligned with the timing of proactive safety intervention to form a closed loop, such as... Figure 6 Figure 7 As shown. Taking the secondary intervention of Automatic Emergency Braking (AEB) as an example, let's assume the latest intervention time of the active safety system is... The remaining time of the third stage obtained through continuous prediction is denoted as A defining criterion is provided: if it is predicted that the driver will not be able to complete effective takeover within the remaining time, even near the latest intervention time, early intervention is triggered. Specifically, a remaining time greater than 0 in the third stage indicates that the driver has not yet entered the effective operation phase, and the system needs to intervene with braking or control in advance to increase the safety margin. To quantify the benefits of early intervention, a safety margin can be defined. for: In the formula, The timing of intervention given for the early intervention mechanism.
[0074] Figure 6 A visualization of the secondary intervention determination and an illustration of the intervention results are provided. Figure 7 The distribution characteristics of the safety margin of samples under the early intervention mechanism are presented, which can be used to evaluate the gain of the closed-loop strategy on the safety margin.
[0075] Finally, to verify the deployability of the project, the model can be deployed in a real vehicle test track for virtual-real fusion testing. In the embodiment, real-time access to driver physiological behavior data can be achieved through data middleware, and information on the vehicle and surrounding traffic participants can be acquired simultaneously as quasi-static and dynamic inputs; the model outputs real-time, advance, and continuous prediction results of the three-stage reaction time online, and is linked with human-machine interface prompts (HMI) and active safety control modules.
[0076] Example 2 In view of the phased autonomous driving driver takeover reaction time prediction method provided in the foregoing embodiments, this embodiment provides a phased autonomous driving driver takeover reaction time prediction system for executing the above-mentioned phased autonomous driving driver takeover reaction time prediction method.
[0077] The system includes: a phased decoupling module, a reaction time generation module, a hierarchical coding module, a prediction module, and an output module.
[0078] 1. A phased decoupling module is used to decouple driver takeover behavior into three phases: physical recovery, sensory recovery, and cognitive recovery. 2. A reaction time generation module, used to calculate the reaction time of the three stages; 3. Hierarchical coding module, used to perform hierarchical multi-receptive field temporal coding and global dependency modeling on the driver's dynamic temporal data to generate dynamic temporal representation; 4. A prediction module, used to encode static individual attributes and quasi-static environmental features into conditional vectors respectively, inject the conditional vectors into the dynamic temporal representation through cross-attention to generate a dynamic-static fusion representation, perform temporal aggregation on the dynamic-static fusion representation, and predict the takeover response time through a decoding head; 5. Output module, used to output the prediction results generated by the prediction module.
[0079] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A phased method for predicting driver takeover reaction time in automated driving, characterized in that, The method specifically includes: S1. Perform structured analysis and phased decoupling of driver takeover behavior, decoupling driver takeover behavior into three stages: physical recovery, sensory recovery, and cognitive recovery. S2. Obtain the action elements and calculation rules corresponding to the reaction time of each stage of the driver takeover behavior, and generate standardized three-stage reaction time labels; S3. Based on the three-stage reaction time tags, hierarchical encoding is performed on the dynamic temporal physiological and behavioral signals before and after takeover to obtain a dynamic temporal representation that simultaneously exhibits short-term disturbances and long-term trends. S4. Using static individual attributes, quasi-static environmental characteristics, and the dynamic temporal representation as inputs, and the advance prediction results, instant prediction results, and continuous prediction results of the three-stage reaction time as outputs, construct a multimodal fusion prediction model based on the time sliding window mechanism. S5. Utilize the trained prediction model to perform advance, real-time, and continuous prediction of takeover reaction time, and deploy it on the human-in-the-loop simulation platform and the autonomous vehicle terminal to realize the prediction of the autonomous driver's staged takeover reaction time and drive the online evaluation of takeover capability.
2. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The action elements and calculation rules corresponding to the reaction time of each stage specifically include calculable stage boundary events. These stage boundary events include completing basic driving posture preparation, returning the gaze to the driving area, and completing the first effective control operation. The stage boundary events are given based on a decision function of multimodal action signals. It manifests as: In the formula, For multimodal data at time The observation, For threshold or rule parameters, For indicator functions, For the stage The decision function, This indicates the first phase of attitude preparation. This indicates a return to focus in the second phase. This indicates that the third stage marks the first successful completion of effective control operations.
3. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The process of generating the standardized three-stage reaction time tag includes: The reaction time for a given stage is obtained by overlapping and merging the duration intervals of action elements within a stage, after deduplication, as the total effective action duration. The action elements include hand, foot, eye, head, and full-body posture movements. The start and end points of these action elements are identified using multimodal acquisition devices, including eye-tracking gaze acquisition, head posture acquisition, steering wheel pedal interaction acquisition, and seat posture and pressure acquisition. satisfy: In the formula, For the first The set of action elements included in a stage For the first The start and end ranges of each action element This indicates the total duration of the interval.
4. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The hierarchical coding adopts a hierarchical multi-receptive-field temporal coding structure, which simultaneously models short-term disturbances and long-term trends in dynamic time series, and obtains dynamic time series representations by combining a global dependency modeling mechanism. The hierarchical coding process satisfies the following: The hierarchical encoding yields a dynamic temporal representation with enhanced global dependencies. This manifests as: In the formula, For the first The dynamic temporal feature representation of the output of hierarchical coding. For layer index, The total number of layers in the hierarchical coding. Indicates the first Layer-time convolution or dilation convolution operations, Indicates downsampling aggregation, Represents a nonlinear mapping. This represents a global dependency modeling module that includes self-attention and gated aggregation.
5. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The multimodal fusion prediction model embeds the driver's static individual attributes and quasi-static environmental features as modulation conditions into the dynamic temporal representation, enabling the model to capture dynamic physiological behavioral features while fusing individual differences and scene context. The static individual attributes and quasi-static environmental features are respectively encoded into condition vectors: The dynamic time series representation is fused using a conditional modulation function: In the formula, This is a conditional vector of static individual attributes. For static individual attributes, The conditional vector representing quasi-static environmental characteristics. Characteristics of a quasi-static environment. This is a characterization based on the fusion of dynamic and static coupling. and It consists of a multilayer perceptron, an embedding layer, and a normalization layer. This indicates cross-attention conditional injection, which is then combined with residuals and gating for fusion.
6. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The multimodal fusion prediction model couples dynamic and static information through a multimodal interactive fusion mechanism. The interactive fusion mechanism includes conditional modulation and temporal interactive aggregation of dynamic temporal representations, and outputs the reaction times of posture recovery, perception recovery and cognition recovery by stage-specific decoders, respectively, to achieve multi-task joint prediction in stages.
7. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The immediate prediction result is generated using data from a fixed window before the takeover notification. An earlier prediction window is introduced before the takeover notification to generate the earlier prediction result. The remaining reaction time is dynamically updated by shifting the prediction window backward to generate a continuous prediction result. The fixed window... It manifests as: The advance prediction window It manifests as: The backward prediction window It manifests as: In the formula, The time when the takeover notification is issued, For fixed window length, To predict the offset in advance, To continuously update the offset.
8. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The online assessment of takeover capability includes: aligning the remaining takeover reaction time obtained from the continuous prediction results with the timing of active safety intervention; when the prediction results indicate that the driver cannot complete effective takeover before the latest intervention time of the active safety system, active safety braking or control intervention is triggered.
9. The method for predicting driver takeover reaction time in a phased manner for automated driving according to claim 1, characterized in that, The reaction times at each stage are as follows: In the formula, The reaction time during the physical recovery phase is... The reaction time is the time required for the sensory recovery phase. This refers to the reaction time during the cognitive recovery phase. For the overlap merging operator, , and These are the number of movement elements for the hands, feet, and whole body, respectively. and For the first The start and end times of each hand movement. and For the first The start and end times of each foot movement. and For the first The start and end times of each posture adjustment action. and The number of head and eye movement elements, respectively. It represents the union of sets.
10. A phased autonomous driving driver takeover reaction time prediction system, characterized in that, The system is used to execute the autonomous driving driver phased takeover reaction time prediction method according to any one of claims 1-9, and the system includes: a phased decoupling module, a reaction time generation module, a hierarchical coding module, a prediction module, and an output module; The phased decoupling module is used to decouple the driver takeover behavior into three stages: physical recovery stage, sensory recovery stage, and cognitive recovery stage. The reaction time generation module is used to calculate the reaction time of the three stages; The hierarchical coding module is used to perform hierarchical multi-receptive field temporal coding and global dependency modeling on the driver's dynamic temporal data to generate a dynamic temporal representation. The prediction module is used to encode static individual attributes and quasi-static environmental features into condition vectors, inject the condition vectors into the dynamic temporal representation through cross-attention to generate a dynamic-static fusion representation, perform temporal aggregation on the dynamic-static fusion representation, and predict the takeover response time through the decoding head. The output module is used to output the prediction results generated by the prediction module.