Vehicle control method and system based on head posture of driver

By combining a head-mounted device and an in-vehicle vision acquisition device with an inertial sensor to acquire the driver's head posture information, a mapping rule is established to achieve natural driver input, which solves the problem of driver distraction in existing technologies and improves driving safety and the accuracy of human-machine collaborative decision-making.

CN121947514APending Publication Date: 2026-05-01AI TUER
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AI TUER
Filing Date
2026-01-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing human-computer interaction methods require drivers to distract their visual attention or perform multiple operations in driving scenarios, which affects driving safety. Existing vehicle control systems cannot accurately understand the driver's intentions.

Method used

By setting an inertial measurement unit and an in-vehicle vision acquisition device on the head-mounted device, and combining binocular stereo vision and inertial sensing data, the six-degree-of-freedom attitude information of the driver's head is obtained, and a mapping rule between head attitude and vehicle control commands is established to achieve natural driver operation input.

Benefits of technology

It increases drivers' trust and acceptance of intelligent driving systems, ensuring that drivers can enjoy the convenience of autonomous driving while maintaining a sense of control and participation in the vehicle, thus improving driving safety and the accuracy of human-machine collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121947514A_ABST
    Figure CN121947514A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle control, in particular to a vehicle control method and system based on the head posture of a driver, and the method comprises the following steps: obtaining first inertial data of the head of the driver, and synchronously collecting image data containing optical mark points on a head-mounted device; based on the image data, calculating to obtain a visual estimation posture of the head of the driver; based on the first inertial data, calculating to obtain an inertial estimation posture of the head of the driver; performing data fusion on the visual estimation attitude and the inertial estimation attitude to obtain six-degree-of-freedom attitude information of the target; acquiring head posture information and a driving instruction of a driver, and establishing a mapping rule between the head posture information and the driving instruction information; and according to a mapping rule, converting the target six-degree-of-freedom attitude information into a corresponding vehicle control instruction, and sending the vehicle control instruction to a vehicle execution mechanism to control vehicle actions, so that the vehicle driving safety of a driver can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle control technology, specifically to a vehicle control method and system based on the driver's head posture. Background Technology

[0002] As intelligent driving technology develops towards higher levels of automation (L2+ and above), the shared control mode is gradually becoming the mainstream. In this mode, the vehicle control system needs to understand the driver's intentions more accurately and naturally in order to achieve a smooth, safe, and expected handover of control and collaborative operation. The existing human-machine interaction methods for controlling the vehicle are clearly insufficient in driving scenarios. The existing methods mostly rely on physical buttons, touch screens, or voice commands for function control, which often require the driver to distract their visual attention or perform multiple operations in driving scenarios, affecting the safety of the driver.

[0003] Therefore, we propose a vehicle control method and system based on driver head posture to solve the above problems. Summary of the Invention

[0004] The purpose of this invention is to provide a vehicle control method and system based on driver head posture to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a vehicle control method and system based on driver head posture, the method comprising the following steps: The driver's head is first inertial data is acquired by an inertial measurement unit installed on the head-mounted device, and image data containing optical markers on the head-mounted device is simultaneously acquired by a vision acquisition device installed in the vehicle. Based on the image data, the visual estimated posture of the driver's head is calculated; based on the first inertial data, the inertial estimated posture of the driver's head is calculated. The visually estimated attitude and the inertial estimated attitude are fused to obtain the target six-degree-of-freedom attitude information of the driver's head. The driver's head posture information and driving commands are acquired, and a mapping rule between the head posture information and driving command information is established. According to the mapping rule, the target six-degree-of-freedom posture information is converted into corresponding vehicle control commands and sent to the vehicle actuators to control the vehicle's actions.

[0006] Preferably, the step of acquiring first inertial data of the driver's head through an inertial measurement unit mounted on the head-mounted device, and simultaneously acquiring image data containing optical markers on the head-mounted device through a vision acquisition device mounted in the vehicle includes: The optical marker array on the control head-mounted device emits optical signals according to a preset encoding mode, wherein at least some of the markers have features that can be uniquely identified by the visual system, and the preset encoding mode is a hybrid encoding mode based on the time dimension and / or frequency dimension. During the illumination of the optical marker, at least one pair of left and right eye images containing the optical marker are simultaneously acquired by a vision acquisition device installed inside the vehicle. While acquiring the left and right eye images, inertial data reflecting the driver's head movements are acquired by an inertial measurement unit mounted on the head-mounted device. A unified time reference timestamp is added to the left eye image, the right eye image, and the inertial data to achieve spatiotemporal synchronization of multi-source sensor data.

[0007] Preferably, the step of calculating the visually estimated pose of the driver's head based on the image data includes: Acquire image data containing optical markers on the driver's head; The imaging area of ​​the optical markers contained in the image data is detected; based on the preset coding features of the optical markers, the unique identifier of each detected marker is distinguished and identified; and the two-dimensional coordinates of each identified optical marker in the image coordinate system are determined. Based on the two-dimensional coordinates of the optical marker and the calibration parameters of the visual acquisition system, the observation coordinates of the optical marker in three-dimensional space are reconstructed; the reconstructed three-dimensional observation coordinates are matched with the corresponding known three-dimensional reference coordinates on the head-mounted device; based on the successfully matched three-dimensional coordinate pairs, the visual estimated posture of the driver's head relative to the reference coordinate system is obtained by solving the spatial rigid body transformation relationship.

[0008] Preferably, the step of calculating the estimated inertial attitude of the driver's head based on the first inertial data includes: Based on the first inertial data, the angular motion information reflecting the head rotation state and the linear motion information reflecting the head translation state are calculated respectively. The calculated angular motion information and linear motion information are combined to construct a complete motion state description of the driver's head relative to the initial moment or reference coordinate system; Based on the complete motion state description, inertial estimated attitude data containing changes in head orientation and position in three-dimensional space is generated.

[0009] Preferably, the step of fusing the visually estimated attitude with the inertial estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head includes: Receive and synchronously acquire the visually estimated pose and the inertial estimated pose; Obtain the visual confidence level corresponding to the visually estimated pose and the inertial confidence level corresponding to the inertially estimated pose; wherein the visual confidence level is determined based on the residual or the number of available feature points in the visual solution process, and the inertial confidence level is determined based on the noise level or integration time of the inertial data. Based on the visual confidence and the inertial confidence, the respective weights of the visually estimated pose and the inertial estimated pose in the fusion process are determined; Based on their respective weights, the visually estimated pose and the inertial estimated pose are weighted and fused to generate the fused target six-degree-of-freedom pose information.

[0010] Preferably, the step of determining the respective weights of the visually estimated pose and the inertial estimated pose in the fusion process based on the visual confidence and the inertial confidence includes: In response to the visual confidence being higher than a first threshold and the inertial confidence being lower than a second threshold, a weight higher than that of the inertial estimated posture is assigned to the visually estimated posture; In response to the visual confidence being lower than the first threshold and the inertial confidence being higher than the second threshold, a weight higher than that of the visual estimated posture is assigned to the inertial estimated posture; In response to the fact that both the visual confidence and the inertial confidence are within a preset normal range, the weights are dynamically allocated according to the ratio of their confidence or a preset proportional relationship.

[0011] Preferably, the step of acquiring the driver's head posture information and driving commands, and establishing a mapping rule between the head posture information and driving command information includes: The driver's head posture information is acquired, and the head posture information is divided into several time segments based on time units. A corresponding posture segment window is created for each time segment. For each pose segment window, feature parsing is performed to extract the pose contour features and pose change trend features of the current segment. Based on the pose contour features, the command is initially screened and a mapping index with possible control commands is established. The selection path of the mapping index is updated based on the posture change trend characteristics to determine the candidate control instruction set corresponding to the current segment; Segments with the same pose contour features among multiple time segments are grouped into the same pose grouping container, and the temporal identification node in the container is determined according to the pose change trend features of each segment. Based on the posture grouping container and the timing identification node, an instruction decision container is constructed for each segment, and the instruction decision container is mapped to a vehicle-executable instruction script through a pre-compilation mechanism to obtain the association mapping relationship between head movements and driving operation instructions.

[0012] Preferably, the steps of performing feature parsing on each pose segment window, extracting the pose contour features and pose change trend features of the current segment, performing initial screening of commands based on the pose contour features, and establishing a mapping index with possible control commands include: For each pose segment window, a feature parsing chain is constructed, which includes a parsing node, a resource allocation node, and a resource locking node; The feature parsing chain is used to obtain the localization dataset and temporal change data in the pose segment window; the parsing node is used to extract features from the localization dataset and temporal change data to obtain pose contour features and pose change trend features. The posture contour features are compared with a predefined instruction posture library to filter out a set of possible control instructions that match the current posture segment window. A mapping index node is created for each selected possible control instruction to form a preliminary instruction mapping index.

[0013] Preferably, the step of converting the target six-degree-of-freedom attitude information into corresponding vehicle control commands according to the mapping rules and sending them to the vehicle actuators to control the vehicle's movements includes: Acquire the six-degree-of-freedom attitude information of the driver's head; A time-series analysis is performed on the continuously acquired target six-degree-of-freedom posture information to identify the head movement patterns contained therein; The identified head movement patterns are matched with the predefined head movements in the mapping rules; When a match is successful, the vehicle control command corresponding to the head action is extracted according to the mapping rules; the vehicle control command is sent to the vehicle actuator to control the vehicle to perform the corresponding action.

[0014] A vehicle control system based on driver head posture, applied to any one of the above-described vehicle control methods based on driver head posture, includes: The image acquisition module is used to acquire the first inertial data of the driver's head through the inertial measurement unit set on the head-mounted device, and to simultaneously acquire image data containing optical markers on the head-mounted device through the visual acquisition device set in the vehicle. The attitude calculation module is used to calculate the visual estimated attitude of the driver's head based on the image data; and to calculate the inertial estimated attitude of the driver's head based on the first inertial data. The attitude fusion module is used to fuse the visually estimated attitude with the inertial estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head. The vehicle control module is used to acquire the driver's head posture information and driving commands, establish a mapping rule between the head posture information and driving command information, and convert the target six-degree-of-freedom posture information into corresponding vehicle control commands according to the mapping rule, and send them to the vehicle actuators to control the vehicle's actions.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. By fusing binocular stereo vision with inertial sensing data, using active infrared coded markers, and combining infrared visual acquisition, the system improves stability under various lighting conditions. It utilizes the driver's natural head movements during driving as control input, enabling the vehicle to predict the driver's operational intentions based on the driver's visual focus. This provides more intelligent assistance, improves the accuracy of human-machine collaborative decision-making, and allows the driver to maintain a sense of control and participation in the vehicle through natural movements while enjoying the convenience of autonomous driving. This increases the driver's trust and acceptance of the intelligent driving system, thereby improving the safety of the driver. 2. The head pose is segmented into temporal segments, and the head pose is collected and analyzed sequentially. During the analysis, the possible vehicle control commands corresponding to each pose segment are screened out. When determining the final head pose, the corresponding vehicle control command is selected. This allows for preparation based on the possible vehicle control commands during the head pose analysis, thereby improving the efficiency of determining the vehicle control command based on the head pose. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a system structure block diagram of the present invention; Figure 3 This is a schematic diagram of the system hardware of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] For examples, please refer to Figures 1 to 3 This invention provides a vehicle control method and system technical solution based on driver head posture: The vehicle control method based on driver head posture includes the following steps: S1: First inertial data of the driver's head is acquired by an inertial measurement unit installed on the head-mounted device, and image data containing optical markers on the head-mounted device is simultaneously acquired by a vision acquisition device installed in the vehicle. The steps of acquiring first inertial data of the driver's head through an inertial measurement unit (IMU) mounted on a head-mounted device and simultaneously acquiring image data containing optical markers on the head-mounted device through a vision acquisition device installed in the vehicle include: controlling a set of infrared LEDs integrated on the head-mounted device to emit infrared light according to a preset coding modulation mode through a microprocessor built into the head-mounted device, wherein the pulse frequency and / or emission sequence of the infrared light are encoded according to the preset unique ID of the LED; simultaneously acquiring left and right eye images covering the area of ​​the driver's head through a binocular camera module installed inside the vehicle; the binocular camera module is equipped with an infrared transmittance filter that matches the emission band of the infrared LEDs; while acquiring the left and right eye images, acquiring the first inertial data through an IMU on the head-mounted device, and acquiring second inertial data characterizing the vehicle body motion through an automotive-grade IMU integrated into the binocular camera module or a fixed position on the vehicle body; and adding a timestamp with a unified time reference to the left eye image, the right eye image, the first inertial data, and the second inertial data based on a precision clock protocol or hardware trigger signal to complete the spatiotemporal synchronization of multi-source data.

[0020] Specifically, the optical marker array on the head-mounted device emits optical signals according to a preset encoding mode, wherein at least some markers have features that can be uniquely identified by the vision system. The emission pulse frequency and / or emission duty cycle of the optical markers are dynamically adjusted according to the ambient light intensity or the operating status of the vision acquisition device. During the emission of the optical markers, at least one pair of left-eye and right-eye images containing the optical markers are simultaneously acquired by a vision acquisition device installed inside the vehicle. While acquiring the left-eye and right-eye images, inertial data reflecting the driver's head movement is acquired by an inertial measurement unit installed on the head-mounted device. A timestamp with a unified time reference is added to the left-eye image, right-eye image, and inertial data to achieve spatiotemporal synchronization of multi-source sensor data. The preset encoding mode is a hybrid encoding mode based on the time dimension and / or frequency dimension, specifically including: controlling different markers to emit light within alternating time windows; and / or controlling different markers to emit light with pulse signals of different frequencies. The synchronization... The data acquisition is achieved through the following methods: sending a hardware trigger signal generated by a common clock source to the visual acquisition device and the head-mounted device; in response to the hardware trigger signal, the visual acquisition device controls its image sensor to start exposure, while the head-mounted device controls the optical marker to emit light within the exposure cycle; wherein, the exposure cycle is matched with the emission pulse cycle of the optical marker; second inertial data characterizing the vehicle's own motion is synchronously acquired through the inertial measurement unit built into the visual acquisition device; a timestamp with the same time reference as the left and right eye images is added to the second inertial data; the specific method of acquiring inertial data through the inertial measurement unit is as follows: continuously acquiring angular velocity data output by the gyroscope and acceleration data output by the accelerometer at a rate higher than the image sampling frequency of the visual acquisition device; caching the acquired raw inertial data in the local memory of the head-mounted device; in response to receiving a timestamp signal synchronized with image acquisition, associating the corresponding inertial data packet with the timestamp and sending it.

[0021] The optical markers on the head-mounted device are actively emitting infrared LEDs, and the visual acquisition device is a camera equipped with an infrared filter. The image data distinguishes and tracks each marker by identifying the encoded information of the infrared LEDs. The encoding modulation mode includes time division multiplexing (TDM), frequency division multiplexing (FDM), or a hybrid mode of TDM and FDM. In TDM mode, LEDs with different IDs emit light in different time windows. In FDM mode, LEDs with different IDs emit light with pulses of different frequencies. The visual acquisition device uniquely identifies and calculates the three-dimensional spatial position of each infrared LED in multiple frames of images based on the encoded modulation information. The visual acquisition device is a binocular camera module containing an automotive-grade IMU. The binocular camera module uses a global shutter sensor, and its exposure time is synchronized with the emission pulse of the infrared LEDs through the PTP protocol or hardware trigger signal to reduce image motion blur and ensure that at least one complete LED emission pulse is captured within the exposure cycle.

[0022] S2: Based on the image data, calculate the visual estimated posture of the driver's head; based on the first inertial data, calculate the inertial estimated posture of the driver's head. The steps for calculating the visually estimated pose of the driver's head based on the image data include: acquiring image data containing optical markers on the driver's head; identifying and extracting the position information of the optical markers from the image data; calculating the visually estimated pose of the driver's head based on the position information of the optical markers and their known spatial distribution on the head-wearing device; identifying the optical markers includes: distinguishing the emission timing and / or emission frequency of different markers in the image sequence based on a preset emission coding pattern; and uniquely determining the marker identity corresponding to each light spot in a single frame or multiple frames of images according to the emission coding pattern.

[0023] The specific content of identifying and extracting the location information of the optical markers from the image data includes: processing the image data to detect the imaging area containing the optical markers; distinguishing and identifying the unique identifier of each detected marker based on the preset coding features of the optical markers; and determining the two-dimensional coordinates of each identified optical marker in the image coordinate system. Based on the positional information of the optical markers and their known spatial distribution on the head-mounted device, the specific details of the driver's estimated head posture are calculated: Based on the two-dimensional coordinates of the optical markers and the calibration parameters of the visual acquisition system, the observation coordinates of the optical markers in three-dimensional space are reconstructed; the reconstructed three-dimensional observation coordinates are matched with the corresponding known three-dimensional reference coordinates on the head-mounted device; based on the successfully matched three-dimensional coordinate pairs, the estimated head posture relative to the reference coordinate system is obtained by solving the spatial rigid body transformation relationship; the reconstruction of the three-dimensional observation coordinates specifically involves: using the principle of binocular stereo vision, triangulation is performed on the two-dimensional coordinates of a marker point in the middle of the left and right views to solve... Calculate its coordinates in three-dimensional space; or, estimate its coordinates in three-dimensional space by combining monocular vision with known marker size or depth sensor information; solve the spatial rigid body transformation relationship including: based on at least three pairs of non-collinear matching coordinate pairs, calculate the rotation and translation parameters required to transform the known three-dimensional reference coordinates to the three-dimensional observation coordinates; when there are more than three pairs of matching coordinate pairs, use an optimization method to calculate the rotation and translation parameters that minimize the overall transformation error, that is, calculate the visually estimated pose of the driver's head relative to the vehicle coordinate system based on the three-dimensional spatial coordinates of at least three non-collinear LED markers in the current frame and the known fixed geometric relationship of the LED markers in the head-mounted device coordinate system.

[0024] The left and right eye images are processed separately. Based on the encoding and modulation mode of the infrared LEDs, the two-dimensional pixel coordinates of each infrared LED marker with a unique ID in each frame are identified and extracted. Based on the pre-calibrated intrinsic, extrinsic, and distortion parameters of the binocular camera module, stereo matching and triangulation are performed on LED markers with the same ID in the left and right eye images to calculate the three-dimensional spatial coordinates of each LED marker in the vehicle coordinate system. Based on the three-dimensional spatial coordinates of at least three non-collinear LED markers in the current frame and the known fixed geometric relationships of the LED markers in the head-mounted device coordinate system, the visually estimated posture of the driver's head relative to the vehicle coordinate system is calculated. The optical markers are actively emitting infrared LEDs, and different markers are distinguished and tracked by identifying the specific encoding information of the infrared LEDs.

[0025] It should be noted that the step of calculating the visually estimated posture of the driver's head relative to the vehicle coordinate system based on the three-dimensional spatial coordinates of at least three non-collinear LED markers in the current frame, and the known fixed geometric relationships of the LED markers in the head-mounted device coordinate system, includes: obtaining the fixed three-dimensional coordinates of each LED marker in the head-mounted device's body coordinate system according to the mechanical design drawings or factory calibration data of the head-mounted device, thus forming a set of known reference points. ,in And at least three points are not collinear; the three-dimensional observation coordinates of the LED marker points corresponding to the reference point set in the vehicle coordinate system, obtained by stereo vision calculation at the current moment, constitute the observation point set. Construct an optimization problem that minimizes the reprojection error, with the objective function being: ,in, The rotation matrix is ​​the desired one (representing the orientation of the driver's head coordinate system relative to the vehicle coordinate system, describing the head's orientation in three-dimensional space, such as yaw, pitch, and roll). The translation vector is the desired vector (representing the position of the origin of the driver's head coordinate system relative to the origin of the vehicle coordinate system, and the three-dimensional coordinates of the head in the vehicle coordinate system). This represents the total number of LED markers that were successfully matched and used for calculation in the current frame. The index number of the LED marker point ( =1, 2, ..., n), For the first The known fixed three-dimensional coordinates of each LED marker point in the coordinate system of the head-mounted device For the first The three-dimensional coordinates of each LED marker point in the vehicle coordinate system are obtained in real time through binocular vision stereo matching and triangulation. These are observation values ​​that change with the driver's head movement. This represents the three-dimensional spatial distance between two points. This is a multiplication operation between a rotation matrix and a vector, the result of which is to multiply the points... Rotate from the head coordinate system to the same direction as the vehicle coordinate system. To point The theoretical predicted position is obtained by completely transforming the head coordinate system to the vehicle coordinate system; the above optimization problem is solved using Direct Linear Transformation (DLT), SVD decomposition, or Iterative Closest Point (ICP) algorithm to obtain the optimal rotation matrix. Translation vector ,like A robust estimation algorithm is employed, specifically: during the iterative solution process, the estimation is based on the reprojection residual of each LED marker point. Dynamic weight allocation is used to assign lower weights to markers with excessively large residuals or to treat them as outliers and remove them, thereby improving the robustness of pose calculation when some markers are temporarily occluded or mismatched. No. The reprojection residuals of each LED marker point are used to detect outliers (occlusion, mismatches, and other anomalies); Convert to quaternion Or Euler angles, combined Together, they constitute the visually estimated pose of the driver's head relative to the vehicle coordinate system. ,in, Indicates the first The reprojection residual vector of a point represents the difference between the observed position and the theoretically predicted position of the LED marker point. Let the sum of squared residuals be the sum of squared residuals at all points. Find the value that minimizes this sum. and , Represents the rotation matrix Translation vector Perform a minimization optimization.

[0026] The steps for calculating the inertial estimated attitude of the driver's head based on the first inertial data include: calculating angular motion information reflecting the head's rotation state and linear motion information reflecting the head's translation state based on the first inertial data; combining the calculated angular motion information and linear motion information to construct a complete motion state description representing the driver's head relative to the initial moment or reference coordinate system; generating inertial estimated attitude data containing the changes in the head's orientation and position in three-dimensional space based on the complete motion state description; the angular motion information is obtained by processing gyroscope data in the inertial measurement unit, and the linear motion information is obtained by processing accelerometer data in the inertial measurement unit; the generated inertial estimated attitude data includes at least one of the following: the rotation angle or quaternion representation of the head relative to the reference coordinate system, the position coordinates of the head relative to the reference coordinate system, the angular velocity of the head movement, and the linear velocity of the head movement.

[0027] S3: The visually estimated attitude and the inertial estimated attitude are fused to obtain the target six-degree-of-freedom attitude information of the driver's head; The steps of fusing the visually estimated attitude and the inertially estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head include: receiving and synchronously acquiring the visually estimated attitude and the inertially estimated attitude; acquiring the visual confidence score corresponding to the visually estimated attitude and the inertial confidence score corresponding to the inertially estimated attitude; wherein, the visual confidence score is determined based on the residual or the number of available feature points in the visual solution process, and the inertial confidence score is determined based on the noise level or integration time of the inertial data; determining the respective weights of the visually estimated attitude and the inertially estimated attitude in the fusion process according to the visual confidence score and the inertial confidence score; and performing weighted fusion of the visually estimated attitude and the inertially estimated attitude based on their respective weights to generate the fused target six-degree-of-freedom attitude information. The weights are determined based on the visual confidence level and the inertial confidence level, specifically including: in response to the visual confidence level being higher than a first threshold and the inertial confidence level being lower than a second threshold, assigning a weight higher than the inertial estimated posture to the visually estimated posture; in response to the visual confidence level being lower than the first threshold and the inertial confidence level being higher than the second threshold, assigning a weight higher than the visually estimated posture to the inertial estimated posture; and in response to both the visual confidence level and the inertial confidence level being within a preset normal range, dynamically allocating weights based on the confidence ratio of the two or a preset proportional relationship.

[0028] Visually estimated pose and inertially estimated pose are fused to obtain high-precision six-DOF pose information of the driver's head. The weights of each sensor source in the fusion process are dynamically adjusted based on their real-time operating conditions. Visually estimated pose data is continuously received from the vision processing module. and its corresponding timestamp and the inertial estimated attitude from the inertial processing module. and its corresponding timestamp The For rotation matrix, It is a translation vector. For attitude quaternions, and These are angular velocity and acceleration, respectively. First, the time synchronization management unit is invoked, based on the Precision Clock Protocol (PTP) or hardware-triggered timing, to... and Perform time alignment. The specific method is as follows: If Less than the preset synchronization threshold If the data arrives at the same time, it is determined that the two are synchronized data and enters the fusion queue; otherwise, motion interpolation or prediction methods are used to extrapolate the timestamp of the first arriving data to the other data, ensuring that the fused attitude information corresponds to the same moment. The confidence scores of the visually estimated attitude and the inertially estimated attitude are calculated in parallel as the basis for subsequent weight allocation decisions. The average reprojection residual output by the visual solution module... Number of infrared LED markers that can be effectively tracked and total Current ambient light intensity The following formula is used for comprehensive calculation. The corresponding formula is: ,in, This is an adjustable coefficient. This is the illumination effect function, which outputs a larger value in strong backlighting or extremely dark environments. This formula ensures that when image quality is high and feature points are stable... The value approaches 1 when the residual is large, the number of feature points is small, or the lighting is poor. Approaching 0; Raw data noise level of the inertial measurement unit (IMU) Integral duration since the last effective visual correction Current IMU temperature ; Use formula Comprehensive calculation ,in, This is an adjustable coefficient. This is the temperature compensation function. This formula reflects the characteristic of IMU data: high short-term accuracy but errors accumulate over time. The longer the integration time, the more accurate the data becomes. The more significant the decay, the better. Maintain a dynamic weight allocation state machine based on real-time calculations. and With preset threshold , , , The comparison results determine the fusion strategy; in the vision-dominated mode, the trigger condition is... and The scenario is characterized by good lighting, unobstructed faces, and minimal vehicle vibration, but the IMU integration error has accumulated significantly, indicating a high weighting of the pose estimation in the visual estimation. ), assigning low weights to inertial attitude estimation ( The system primarily relies on visual data to provide an absolute pose reference, while utilizing IMU data for smoothing; in inertial-dominated mode, the triggering condition is... and The scenario involves the driver blinking, their face being completely obscured, and the camera experiencing temporary blindness when entering or exiting a tunnel. However, the IMU has just completed visual correction, resulting in high short-term accuracy. Therefore, a high weight is assigned to the inertial estimation of attitude. ), assigning low weights to visual pose estimation ( Or set to 0). The system switches to pure inertial navigation (INS) mode, using IMU data for short-term pose prediction until the visual signal is recovered; in dynamic weighted fusion mode, and All are within their respective normal ranges. and If the scenario involves most normal operating conditions, then dynamic calculation is performed based on the confidence ratio of the two or a preset ratio. One implementation method is... , According to the determined weights and The attitude information is then fused. For the rotation component, quaternion spherical linear interpolation (SLERP) or weighted averaging followed by normalization is typically used; for the translation component, a weighted average is directly applied. Ultimately, the resulting target six-DOF attitude information is obtained. The data is then fed into the subsequent intent recognition module. Simultaneously, the weight allocation status and confidence value of this fusion process are recorded, achieving intelligent complementarity between the advantages of visual and inertial sensors. When the visual signal is reliable, absolute attitude accuracy is ensured; when the visual signal briefly fails, it seamlessly switches to inertial prediction to guarantee continuity; under normal conditions, optimal weighting is performed to balance accuracy and dynamic performance, without relying on a single fixed algorithm, exhibiting high adaptability.

[0029] S4: Acquire the driver's head posture information and driving commands, establish a mapping rule between head posture information and driving command information; according to the mapping rule, convert the target six-degree-of-freedom posture information into corresponding vehicle control commands, and send them to the vehicle actuators to control vehicle actions; The steps for acquiring driver head posture information and driving commands, and establishing a mapping rule between head posture information and driving command information include: acquiring driver head posture information; dividing the head posture information into several time-series segments based on time units, and creating a corresponding posture segment window for each time-series segment; performing feature analysis on each posture segment window, extracting the posture contour features and posture change trend features of the current segment, performing initial screening of commands based on the posture contour features, and establishing a mapping index with possible control commands; updating the selection path of the mapping index based on the posture change trend features, and determining the candidate control command set corresponding to the current segment; grouping segments with the same posture contour features among multiple time-series segments into the same posture grouping container, and determining the time-series identification node in the container based on the posture change trend features of each segment; constructing a command decision container for each segment based on the posture grouping container and the time-series identification node, and mapping the command decision container into a vehicle-executable command script through a pre-compilation mechanism to obtain the association mapping relationship between head movements and driving operation commands.

[0030] It should be noted that the steps for obtaining the driver's head posture information, dividing the head posture information into several time segments based on time units, and creating a corresponding posture segment window for each time segment are as follows: acquiring a continuous image sequence of the driver's head through an in-vehicle camera, extracting the three-dimensional coordinates of key points of the head based on each frame of the image, and constructing a posture localization dataset; dividing the continuous action into several posture segments according to a preset time unit, with each segment corresponding to a posture segment window, and setting the window lifecycle. The steps of performing feature parsing on each attitude segment window, extracting the attitude contour features and attitude change trend features of the current segment, performing initial screening of commands based on the attitude contour features, and establishing a mapping index with possible control commands include: constructing a feature parsing chain for each attitude segment window, the feature parsing chain including parsing nodes, resource allocation nodes, and resource locking nodes; obtaining the positioning dataset and temporal change data in the attitude segment window through the feature parsing chain; using the parsing nodes to extract features from the positioning dataset and temporal change data to obtain attitude contour features and attitude change trend features; comparing the attitude contour features with a predefined command attitude library to filter out a set of possible control commands that match the current attitude segment window; and establishing a mapping index node for each filtered possible control command to form a preliminary command mapping index. The workflow of the feature parsing chain includes: requesting computing resources from the resource pool through the resource allocation node to extract the posture contour features and posture change trend features; if resources are sufficient, activating the parsing node and starting feature extraction; if resources are insufficient, locking the current feature parsing chain through the resource locking node and waiting for the resources to be released before activating the parsing node; the resource pool is a shared computing resource pool in the vehicle processing unit or cloud server, and the resource allocation supports dynamic scheduling and priority allocation; Specifically, the index structure is a multi-layered tree-graph hybrid structure, including: a root node layer, which classifies based on posture contour features; a branch node layer, which constructs temporal paths based on posture change trend features; and a leaf node layer, which stores specific vehicle control commands. Nodes are connected by directed edges to form a dynamic mapping network from posture features to control commands. Each node contains the following attribute set: node identifier, posture feature encoding vector, associated control command label, confidence score, response time threshold, resource occupancy identifier, and priority weight. Each temporal path in the branch node layer is assigned a weight value, which is dynamically adjusted based on at least one of the following factors: historical recognition accuracy, user habit data, and vehicle environmental context information. The system also includes a node state machine, supporting node transitions between the following states: active state: can participate in current command matching; pending state: waiting for further feature confirmation; locked state: suspended when resources are occupied or paths conflict; and invalid state: temporarily invalid after recognition failure or user disabling. It supports a context-aware index switching mechanism, capable of selecting different index subgraphs based on at least one of the following conditions: vehicle driving status, driver physiological state, environmental perception data, and system resource load. It adopts a three-layer tree structure of "root node-branch node-leaf node," where the root node corresponds to the posture contour feature category, branch nodes correspond to the temporal change pattern, and leaf nodes correspond to specific control commands. Directed edges are introduced between branch nodes and leaf nodes to represent the dynamic mapping relationship between posture change paths and commands, forming a graph structure that supports multi-path mapping. Each node contains the following attributes: node ID, posture feature encoding, command label, confidence level, response time threshold, resource occupancy identifier, and priority weight. Node states are divided into: active, pending, locked, and inactive, supporting dynamic state transitions and resource reclamation. Temporal weights are assigned to each branch path, with weights dynamically adjusted based on historical recognition accuracy, user habit data, and environmental context. Path pruning and merging are supported to improve indexing efficiency and adaptability. The extraction of attitude change trend features includes: using continuous inter-frame displacement and angle change sequences from temporal change data; extracting dynamic change patterns based on Long Short-Term Memory (LSTM) networks or Temporal Convolutional Networks (TCNs) to output attitude change trend feature vectors; the instruction attitude library includes multiple predefined standard attitude templates, each template associated with one or more vehicle control instructions; the mapping index nodes are nodes in a tree structure or graph structure, each node recording a possible control instruction and its corresponding attitude contour feature similarity, confidence, and execution priority information; it also includes an index node update mechanism: when attitude change trend features become more available, the mapping index nodes are dynamically adjusted based on temporal matching results, including node merging, deletion, or path reselection; The steps of updating the selection path of the mapping index based on the attitude change trend features and determining the candidate control instruction set corresponding to the current segment include: selecting the most matching instruction path from multiple possible instructions corresponding to the attitude contour features through time-series matching based on the attitude change trend features; using the matching result as the candidate control instruction for the current segment and updating the execution path of the mapping index. The steps of grouping segments with the same attitude contour features into the same attitude grouping container and determining the timing identification node of each segment within the container based on the attitude change trend features of each segment include: grouping multiple time segments with the same attitude contour features into the same attitude grouping container; and establishing timing identification nodes within each attitude grouping container based on the attitude change trend features of each segment to distinguish the timing change features of different commands within the same container.

[0031] Based on the attitude grouping container and timing identification node, the steps of constructing an instruction decision container for each segment and mapping the instruction decision container to a vehicle-executable instruction script through a pre-compilation mechanism include: allocating a container control thread for each instruction decision container, performing thread-container pre-compilation and container-segment pre-compilation; generating instruction scripts through thread-container pre-compilation and caching them in the vehicle control unit; and passing the compilation result of the first segment to other segments in the same container through container-segment pre-compilation to achieve instruction script reuse and preloading.

[0032] Specifically, continuous head pose information is segmented into multiple temporal segments based on time units, and independent "pose segment windows" are created to achieve structured slicing and independent management of pose data, facilitating subsequent segment-by-segment parsing and real-time processing. Within each segment window, pose contour features (static spatial features) and pose change trend features (dynamic temporal features) are extracted simultaneously. A two-layer decision path of "contour → initial screening → change → fine screening" is established to improve the accuracy and efficiency of command mapping. A preliminary command mapping index is established based on the pose contour features, and then the index path is dynamically updated according to the pose change trend features. The system now employs a progressive selection process from "possible instruction set" to "best candidate instruction" to enhance its adaptability to complex attitude sequences. It groups time segments with the same initial attitude into the same "attitude grouping container," and further distinguishes the temporal evolution patterns of different instructions within the container through "time sequence identification nodes." This enables the classification, aggregation, and fine-grained identification of attitude data. The system also introduces a two-level compilation mode: "thread-container pre-compilation" and "container-segment pre-compilation." This pre-compiles the identification results into executable instruction scripts and caches them in the vehicle control unit, supporting rapid instruction invocation and batch loading, and reducing real-time processing latency. Through a two-layer recognition mechanism of "initial screening based on contour features + fine screening based on change trends," the system effectively distinguishes commands with similar initial postures but different subsequent actions (such as "nodding to confirm" versus "looking down at the instrument panel"), reducing the false recognition rate. Employing temporal fragmentation processing and a pre-compilation caching mechanism, the system can perform command prediction and compilation preparation before the posture is fully completed, significantly shortening the overall latency from posture recognition to command execution. Through a "posture grouping container" structure, different commands with the same initial posture can be parsed in parallel within the same container, and the compilation results of the same command can be reused across fragments, improving system resource utilization and processing efficiency. The modular container structure and configurable mapping index mechanism facilitate the addition or modification of command types, making it suitable for personalized command extensions for different vehicle models and driving scenarios. By leveraging cloud collaboration, dynamic resource scheduling, and compilation result caching, the system reduces real-time computation, making it suitable for automotive embedded platforms with limited computing power and enhancing system usability and deployment flexibility. Pre-compiled instruction scripts are delivered to the vehicle control execution module, which selects the corresponding instruction script to execute vehicle control operations based on the driver's real-time head posture. The head posture is segmented into time-series fragments, and head postures are collected and analyzed sequentially. During the analysis, possible vehicle control instructions corresponding to each posture fragment are identified, and the corresponding vehicle control instruction is selected when the final head posture is determined. This allows for preparation based on possible vehicle control instructions during head posture analysis, thereby improving the efficiency of determining vehicle control instructions based on head posture.

[0033] It also includes: setting corresponding triggering conditions and safety activation conditions for each associated mapping relationship, and constructing a complete mapping rule set including head action definition, driving operation command, triggering conditions and safety activation conditions; the triggering conditions include at least one of the following: the duration of the head action reaches a preset threshold, the amplitude of the head action exceeds a preset range, and the matching degree between the movement trajectory of the head action and the preset reference trajectory meets the requirements; the safety activation conditions include at least one of the following: the current driving speed of the vehicle is within the speed limit allowed for executing the driving operation command, the current driving scenario meets the safety premise for executing the driving operation command, and executing the driving operation command does not conflict with other currently activated vehicle control commands; establishing associated mapping relationships includes: using a one-to-one mapping method to map a head action to a driving operation command; or using a one-to-many mapping method to map a head action to a set of sequentially executed driving operation command sequences.

[0034] The steps of converting the target six-degree-of-freedom attitude information into corresponding vehicle control commands according to the mapping rules and sending them to the vehicle actuators to control vehicle actions include: acquiring the target six-degree-of-freedom attitude information of the driver's head; performing time-series analysis on the continuously acquired target six-degree-of-freedom attitude information to identify the head movement patterns contained therein; matching the identified head movement patterns with predefined head movements in the mapping rules; when a match is successful, extracting the vehicle control commands corresponding to the head movements according to the mapping rules; and sending the vehicle control commands to the vehicle actuators to control the vehicle to perform the corresponding actions. Specifically, the temporal analysis includes: determining whether the change in head posture meets a predefined starting condition; after meeting the starting condition, tracking the continuous change in head posture until it meets a predefined ending condition; recognizing the sequence of head posture changes between the starting and ending conditions as a complete head action; the matching process includes: calculating the similarity between the identified head action pattern and each predefined head action in the mapping rule; selecting the predefined head action with the highest similarity as the matching result; wherein, the similarity is calculated based on at least one dimension of the action's spatial trajectory, duration, and speed; the execution parameters corresponding to the vehicle control command include at least one of the following: the execution intensity of the vehicle control command, the execution duration of the vehicle control command, and the execution target position or angle of the vehicle control command.

[0035] Specifically, a multi-sensor observation model is designed to jointly optimize visual reprojection error and IMU pre-integration residual; an adaptive noise covariance matrix is ​​introduced to dynamically adjust fusion weights based on sensor confidence; when visual failure (such as blinking or occlusion) is detected, the system automatically switches to an IMU-dominated prediction mode. A rule engine based on a finite state machine (FSM) is established to parse head action sequences: "turn right + look right" for 1.5 seconds → trigger lane change assist; "nod twice" → answer a call; "look up at the HUD area" → switch instrument display mode; "turn left and gaze" → activate the left surround-view camera; user-defined action commands are supported, and mapping rules are updated via OTA; for example, during intelligent lane change on a highway, if the driver turns their head 45° to the right and holds it for 2 seconds, the system recognizes this as a lane change intention; the system immediately calls on the lateral millimeter-wave radar and camera to detect vehicles behind the target lane; if safety conditions are met, the turn signal is automatically activated, and the EPS is controlled to perform a smooth lane change; at the same time, virtual lane lines and safety prompts are displayed in the AR-HUD. For example, when automatically navigating narrow roads, the driver continuously turns their head to the left and gazes at the left rearview mirror area. The system recognizes this as a need to navigate narrow roads; it automatically activates the 360° surround view system to synthesize a top-down view of the vehicle's surroundings; combined with ultrasonic radar data, it automatically adjusts the steering wheel angle to assist the vehicle in navigating narrow roads; after navigating, it automatically exits the surround view mode and restores the default view. This head posture perception system, which can operate stably in complex environments and possesses high precision and low latency, can map posture data to vehicle control commands in real time and accurately, achieving a natural driving experience of "what you see is what you control".

[0036] A vehicle control system based on driver head posture, applied to any one of the above-described vehicle control methods based on driver head posture, includes: The image acquisition module is used to acquire the first inertial data of the driver's head through the inertial measurement unit set on the head-mounted device, and to simultaneously acquire image data containing optical markers on the head-mounted device through the visual acquisition device set in the vehicle. The attitude calculation module is used to calculate the visual estimated attitude of the driver's head based on the image data; and to calculate the inertial estimated attitude of the driver's head based on the first inertial data. The attitude fusion module is used to fuse the visually estimated attitude with the inertial estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head. The vehicle control module is used to acquire the driver's head posture information and driving commands, establish a mapping rule between the head posture information and driving command information, and convert the target six-degree-of-freedom posture information into corresponding vehicle control commands according to the mapping rule, and send them to the vehicle actuators to control the vehicle's actions.

[0037] This invention improves stability under various lighting conditions (including strong backlight and nighttime) by fusing visual (binocular stereo vision) and inertial sensing data, using active infrared coded markers combined with infrared visual acquisition. Through dynamically weighted fusion, it seamlessly switches to inertial-dominated mode when vision temporarily fails, ensuring continuous perception. It utilizes natural head movements of the driver during driving (such as observing rearview mirrors and side windows) as control input, eliminating the need for additional learning or complex operating procedures, significantly reducing the cognitive load of interaction. This allows the vehicle to "understand" the driver's visual focus and operational intentions, providing more intelligent assistance (such as automatically detecting the safety of the target lane and assisting steering when the driver intends to change lanes), improving the accuracy of human-machine collaborative decision-making. While enjoying the convenience of autonomous driving, the driver can still maintain a sense of control and participation through natural actions, increasing trust and acceptance of the intelligent driving system, thereby enhancing driver safety.

[0038] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0039] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A vehicle control method based on driver head posture, characterized in that, Includes the following steps: The driver's head is first inertial data is acquired by an inertial measurement unit installed on the head-mounted device, and image data containing optical markers on the head-mounted device is simultaneously acquired by a vision acquisition device installed in the vehicle. Based on the image data, the visual estimated posture of the driver's head is calculated; Based on the first inertial data, the estimated inertial attitude of the driver's head is calculated; The visually estimated attitude and the inertial estimated attitude are fused to obtain the target six-degree-of-freedom attitude information of the driver's head. The driver's head posture information and driving commands are acquired, and a mapping rule between the head posture information and driving command information is established. According to the mapping rule, the target six-degree-of-freedom posture information is converted into corresponding vehicle control commands and sent to the vehicle actuators to control the vehicle's actions.

2. The vehicle control method based on driver head posture according to claim 1, characterized in that: The step of acquiring first inertial data of the driver's head through an inertial measurement unit mounted on the head-mounted device, and simultaneously acquiring image data containing optical markers on the head-mounted device through a vision acquisition device mounted in the vehicle, includes: The optical marker array on the control head-mounted device emits optical signals according to a preset encoding mode, wherein at least some of the markers have features that can be uniquely identified by the visual system, and the preset encoding mode is a hybrid encoding mode based on the time dimension and / or frequency dimension. During the illumination of the optical marker, at least one pair of left and right eye images containing the optical marker are simultaneously acquired by a vision acquisition device installed inside the vehicle. While acquiring the left and right eye images, inertial data reflecting the driver's head movements are acquired by an inertial measurement unit mounted on the head-mounted device. A unified time reference timestamp is added to the left eye image, the right eye image, and the inertial data to achieve spatiotemporal synchronization of multi-source sensor data.

3. The vehicle control method based on driver head posture according to claim 1, characterized in that: The step of calculating the visually estimated pose of the driver's head based on the image data includes: Acquire image data containing optical markers on the driver's head; The imaging area of ​​the optical markers contained in the image data is detected; based on the preset coding features of the optical markers, the unique identifier of each detected marker is distinguished and identified; and the two-dimensional coordinates of each identified optical marker in the image coordinate system are determined. Based on the two-dimensional coordinates of the optical marker and the calibration parameters of the visual acquisition system, the observation coordinates of the optical marker in three-dimensional space are reconstructed; the reconstructed three-dimensional observation coordinates are matched with the corresponding known three-dimensional reference coordinates on the head-mounted device; based on the successfully matched three-dimensional coordinate pairs, the visual estimated posture of the driver's head relative to the reference coordinate system is obtained by solving the spatial rigid body transformation relationship.

4. The vehicle control method based on driver head posture according to claim 1, characterized in that: The step of calculating the estimated inertial attitude of the driver's head based on the first inertial data includes: Based on the first inertial data, the angular motion information reflecting the head rotation state and the linear motion information reflecting the head translation state are calculated respectively. The calculated angular motion information and linear motion information are combined to construct a complete motion state description of the driver's head relative to the initial moment or reference coordinate system; Based on the complete motion state description, inertial estimated attitude data containing changes in head orientation and position in three-dimensional space is generated.

5. The vehicle control method based on driver head posture according to claim 1, characterized in that: The step of fusing the visually estimated attitude with the inertial estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head includes: Receive and synchronously acquire the visually estimated pose and the inertial estimated pose; Obtain the visual confidence level corresponding to the visually estimated pose and the inertial confidence level corresponding to the inertially estimated pose; wherein the visual confidence level is determined based on the residual or the number of available feature points in the visual solution process, and the inertial confidence level is determined based on the noise level or integration time of the inertial data. Based on the visual confidence and the inertial confidence, the respective weights of the visually estimated pose and the inertial estimated pose in the fusion process are determined; Based on their respective weights, the visually estimated pose and the inertial estimated pose are weighted and fused to generate the fused target six-degree-of-freedom pose information.

6. The vehicle control method based on driver head posture according to claim 5, characterized in that: The step of determining the respective weights of the visually estimated pose and the inertial estimated pose in the fusion process based on the visual confidence and the inertial confidence includes: In response to the visual confidence being higher than a first threshold and the inertial confidence being lower than a second threshold, a weight higher than that of the inertial estimated posture is assigned to the visually estimated posture; In response to the visual confidence being lower than the first threshold and the inertial confidence being higher than the second threshold, a weight higher than that of the visual estimated posture is assigned to the inertial estimated posture; In response to the fact that both the visual confidence and the inertial confidence are within a preset normal range, the weights are dynamically allocated according to the ratio of their confidence or a preset proportional relationship.

7. The vehicle control method based on driver head posture according to claim 1, characterized in that: The steps of acquiring the driver's head posture information and driving commands, and establishing a mapping rule between the head posture information and driving command information include: The driver's head posture information is acquired, and the head posture information is divided into several time segments based on time units. A corresponding posture segment window is created for each time segment. For each pose segment window, feature parsing is performed to extract the pose contour features and pose change trend features of the current segment. Based on the pose contour features, the command is initially screened and a mapping index with possible control commands is established. The selection path of the mapping index is updated based on the posture change trend characteristics to determine the candidate control instruction set corresponding to the current segment; Segments with the same pose contour features among multiple time segments are grouped into the same pose grouping container, and the temporal identification node in the container is determined according to the pose change trend features of each segment. Based on the posture grouping container and the timing identification node, an instruction decision container is constructed for each segment, and the instruction decision container is mapped to a vehicle-executable instruction script through a pre-compilation mechanism to obtain the association mapping relationship between head movements and driving operation instructions.

8. The vehicle control method based on driver head posture according to claim 7, characterized in that: The steps of performing feature parsing on each pose segment window, extracting the pose contour features and pose change trend features of the current segment, performing initial screening of commands based on the pose contour features, and establishing a mapping index with possible control commands include: For each pose segment window, a feature parsing chain is constructed, which includes a parsing node, a resource allocation node, and a resource locking node; The feature parsing chain is used to obtain the localization dataset and temporal change data in the pose segment window; the parsing node is used to extract features from the localization dataset and temporal change data to obtain pose contour features and pose change trend features. The posture contour features are compared with a predefined instruction posture library to filter out a set of possible control instructions that match the current posture segment window. A mapping index node is created for each selected possible control instruction to form a preliminary instruction mapping index.

9. The vehicle control method based on driver head posture according to claim 8, characterized in that: The step of converting the target six-degree-of-freedom attitude information into corresponding vehicle control commands according to the mapping rules and sending them to the vehicle actuators to control the vehicle's actions includes: Acquire the six-degree-of-freedom attitude information of the driver's head; A time-series analysis is performed on the continuously acquired target six-degree-of-freedom posture information to identify the head movement patterns contained therein; The identified head movement patterns are matched with the predefined head movements in the mapping rules; When a match is successful, the vehicle control command corresponding to the head action is extracted according to the mapping rules; the vehicle control command is sent to the vehicle actuator to control the vehicle to perform the corresponding action.

10. A vehicle control system based on driver head posture, applied to the vehicle control method based on driver head posture as described in any one of claims 1-9, characterized in that, include: The image acquisition module is used to acquire the first inertial data of the driver's head through the inertial measurement unit set on the head-mounted device, and to simultaneously acquire image data containing optical markers on the head-mounted device through the visual acquisition device set in the vehicle. The attitude calculation module is used to calculate the visually estimated attitude of the driver's head based on the image data. Based on the first inertial data, the estimated inertial attitude of the driver's head is calculated; The attitude fusion module is used to fuse the visually estimated attitude with the inertial estimated attitude to obtain the target six-degree-of-freedom attitude information of the driver's head. The vehicle control module is used to acquire the driver's head posture information and driving commands, establish a mapping rule between the head posture information and driving command information, and convert the target six-degree-of-freedom posture information into corresponding vehicle control commands according to the mapping rule, and send them to the vehicle actuators to control the vehicle's actions.