Driving state detection method, vehicle and computer readable storage medium

By using multimodal weighted fusion to process vehicle driving status, steering wheel and visual signals, and dynamically adjusting feature weights, the accuracy and robustness issues of existing fatigue driving detection technologies in complex environments are solved, achieving accurate fatigue driving detection and improved safety.

CN120974429APending Publication Date: 2025-11-18CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511185508.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing fatigue driving detection technologies are not very accurate under changing ambient light conditions and complex road conditions, and lack effective road interference compensation mechanisms, resulting in poor system robustness and applicability.

Method used

By acquiring vehicle driving status signals, steering wheel signals, and visual perception signals, multimodal weighted fusion processing is performed to generate a fused feature vector. The target state detection model is then used for prediction, and the feature weights are dynamically adjusted to adapt to different driving scenarios.

Benefits of technology

It accurately detects fatigue driving in various complex driving scenarios, improving detection accuracy and system robustness, and enhancing driving safety and the effectiveness of the early warning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974429A_ABST
    Figure CN120974429A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a driving state detection method, a vehicle and a computer readable storage medium, and relates to the technical field of vehicles. The method comprises the following steps: acquiring a driving state signal, a steering wheel signal and a visual perception signal of a vehicle during driving; performing feature extraction based on the steering wheel signal to obtain driving behavior features; performing feature extraction based on the visual perception signal to obtain visual features; according to the driving state signals, multi-modal weighted fusion processing is carried out on the driving behavior features and the visual features, fusion feature vectors are generated, and the driving state signals are used for determining fusion weights corresponding to the driving behavior features and fusion weights corresponding to the visual features respectively; and performing reasoning prediction on the fusion feature vector by using a target state detection model, and determining driving state information corresponding to the vehicle. The technical problems of low fatigue driving detection precision, poor system stability and poor scene applicability in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and more specifically, to a driving state detection method, a vehicle, and a computer-readable storage medium. Background Technology

[0002] With the development of autonomous driving technology, ensuring the driver's focus and alertness in semi-autonomous or manual driving modes has become an urgent need. Especially during long-distance driving, night driving, or complex road conditions, fatigue driving detection systems can effectively prevent traffic accidents and protect the safety of people's lives and property.

[0003] Existing fatigue driving detection technologies still suffer from significant drawbacks: vision-based solutions are susceptible to changes in ambient light and driver-related accessories, leading to unstable detection rates; behavioral signal-based solutions rely on vehicle dynamic data, but experience high false alarm rates and struggle to accurately assess performance in specific scenarios such as road curvature variations; furthermore, existing technologies are overly simplistic in feature extraction and lack effective road interference compensation mechanisms, reducing the overall robustness and applicability of the system. These shortcomings limit the effectiveness and reliability of fatigue driving detection systems in real-world driving environments.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a driving state detection method, a vehicle, and a computer-readable storage medium to at least solve the technical problems of low accuracy in fatigue driving detection, poor system stability, and poor applicability to various scenarios in related technologies.

[0006] According to one aspect of the embodiments of this application, a driving state detection method is provided, comprising: acquiring a driving state signal of a vehicle in motion, a steering wheel signal, and a visual perception signal, wherein the steering wheel signal is used to record rotation data corresponding to multiple steering wheel rotation events, and the display content corresponding to the visual perception signal includes the target part of the driver; performing feature extraction based on the steering wheel signal to obtain driving behavior features; performing feature extraction based on the visual perception signal to obtain visual features; performing multimodal weighted fusion processing on the driving behavior features and visual features according to the driving state signal to generate a fused feature vector, wherein the driving state signal is used to determine the fusion weights corresponding to the driving behavior features and the fusion weights corresponding to the visual features; and using a target state detection model to infer and predict the fused feature vector to determine the driving state information corresponding to the vehicle.

[0007] Furthermore, feature extraction based on steering wheel signals yields driving behavior features, including: preprocessing and extracting data from the steering wheel signals to obtain multiple data samples, where each data sample includes angle data and angular velocity data corresponding to a steering wheel rotation event; and performing feature analysis and calculation based on multiple data samples to obtain driving behavior features.

[0008] Furthermore, driving behavior characteristics include: fine operation entropy characteristics, significant steering frequency characteristics, stationary duration characteristics, operation interval fluctuation characteristics, steering distribution pattern characteristics, and operation rate correlation characteristics. Based on multiple data samples, feature analysis and calculation are performed to obtain driving behavior characteristics including: calculating the angle change of multiple data samples to determine the fine operation entropy characteristics and significant steering frequency characteristics. The fine operation entropy characteristics are used to characterize the information entropy corresponding to data samples whose steering wheel angle change is less than a first threshold within a first time window; the significant steering frequency characteristics are used to characterize the proportion of data samples whose steering wheel angle change is greater than a second threshold within a second time window. For multiple data samples... Angular velocity was calculated from the samples to determine the stationary duration and operation interval fluctuation characteristics. The stationary duration characteristic was used to characterize the duration during which the steering wheel angular velocity was less than the third threshold within the third time window, and the operation interval fluctuation characteristic was used to characterize the degree of fluctuation in the interval between consecutive steering actions within the fourth time window. Angle time series analysis was performed on multiple data samples to determine the steering distribution morphology characteristics, which were used to characterize the kurtosis of the steering wheel angle sequence within the fifth time window. Angle velocity time series analysis was performed on multiple data samples to determine the operation rate correlation characteristics, which were used to characterize the autocorrelation coefficient of the steering wheel angular velocity sequence within the sixth time window.

[0009] Furthermore, the visual perception signal includes video stream data collected by the vehicle's infrared camera. Based on the visual perception signal, feature extraction is performed to obtain visual features, including: using a visual recognition algorithm to identify and determine the key point position sequence of the target part from continuous image frames of the video stream data; and performing morphological change frequency analysis of the target part based on the key point position sequence to obtain visual features.

[0010] Furthermore, the target area is the eye, and the visual features include the temporal features of eye closure. Based on the key point position sequence, the morphological change frequency analysis of the target area is performed to obtain the visual features, which include: using the key point position sequence, the distance between multiple key point pairs of the target area is calculated to obtain eye morphological parameters, wherein the eye morphological parameters are used to characterize the degree of eye opening and closing; based on the eye morphological parameters and the seventh time window, the temporal features of eye closure are obtained by calculating the frequency of eye morphological changes.

[0011] Furthermore, based on the driving state signal, multimodal weighted fusion processing is performed on driving behavior features and visual features to generate a fused feature vector, including: performing vehicle dynamic feature analysis based on the driving state signal to determine the driving scenario type of the vehicle; determining the first weight corresponding to the driving behavior features and the second weight corresponding to the visual features based on the driving scenario type; splicing the driving behavior features and visual features to obtain a spliced ​​feature sequence; and generating a fused feature vector based on the spliced ​​feature sequence, the first weight, and the second weight.

[0012] Furthermore, the first weight includes: the first sub-weight corresponding to the fine operation entropy feature, the second sub-weight corresponding to the significant steering frequency feature, the third sub-weight corresponding to the stationary duration feature, the fourth sub-weight corresponding to the operation interval fluctuation feature, the fifth sub-weight corresponding to the steering distribution pattern feature, and the sixth sub-weight corresponding to the operation rate correlation feature. Based on the driving scenario type, the first weight corresponding to the driving behavior feature is determined as follows: when the driving scenario type is a sharp turn scenario, the first sub-weight is set to 0.3, the second sub-weight to 0.1, the third sub-weight to 0.8, the fourth sub-weight to 0.2, and the fifth sub-weight to 0.3. The first sub-weight is set to 0.4, and the sixth sub-weight is set to 0.1. When the driving scenario type is a congestion scenario, the first sub-weight is set to 0.9, the second sub-weight to 0.2, the third sub-weight to 0.7, the fourth sub-weight to 0.6, the fifth sub-weight to 0.8, and the sixth sub-weight to 0.5. When the driving scenario type is any other than a sharp turn scenario or a congestion scenario, the first sub-weight is set to 1, the second sub-weight is set to 1, the third sub-weight is set to 1, the fourth sub-weight is set to 1, the fifth sub-weight is set to 1, and the sixth sub-weight is set to 1.

[0013] Furthermore, the target state detection model includes a pre-trained transformation component and a probability mapping component. The target state detection model is used to infer and predict the fused feature vector to determine the driving state information, including: using the transformation component to perform inference analysis on the fused feature vector to generate inference results; and using the probability mapping component to perform probability prediction mapping on the inference results to obtain driving state information, wherein the driving state information is used to characterize the probability that the driver is in a state of fatigue driving.

[0014] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the driving state detection method of any one of the above-mentioned methods when it runs.

[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the storage medium is located to execute the driving state detection method of any one of the above.

[0016] In this embodiment, driving status signals, steering wheel signals, and visual perception signals of the vehicle during operation are acquired. The steering wheel signals record rotation data corresponding to multiple steering wheel rotation events, and the visual perception signals display content including the target parts of the driver. Driving behavior features are extracted based on the steering wheel signals; visual features are extracted based on the visual perception signals; and multimodal weighted fusion processing is performed on the driving behavior features and visual features according to the driving status signals to generate a fused feature vector. The driving status signals are used to determine the fusion weights corresponding to the driving behavior features and the visual features, respectively. A target state detection model is used to infer and predict the fused feature vector to determine the vehicle's corresponding driving state information. It is noteworthy that by dynamically fusing feature data from different modalities, this embodiment can automatically adjust the weight ratio of steering wheel behavior and visual features according to the vehicle's real-time driving state, thereby effectively overcoming the problem of decreased reliability of single features due to environmental factors (such as lighting and road conditions). Furthermore, fusing visual features and driving behavior features not only improves the accuracy of fatigue detection but also enhances the system's robustness and adaptability. This allows the system to reliably and accurately detect driver fatigue in various complex driving scenarios, ensuring driving safety. The multimodal weighted fusion mechanism in this embodiment ensures that even in harsh environments, a high fatigue driving detection rate can be maintained through the enhancement of data from another modality. This significantly improves the effectiveness and safety of the driver warning system and enhances the user's driving experience.

[0017] In other words, the embodiments of this application achieve the goal of accurately detecting fatigue driving under various driving conditions through an adaptive multimodal weighted fusion mechanism, thereby improving the accuracy and robustness of fatigue driving detection and enhancing user driving safety. This solves the technical problems of low accuracy of fatigue driving detection, poor system stability and scenario applicability in related technologies. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0019] Figure 1 This is a hardware structure block diagram of an optional computing terminal for implementing a driving state detection method according to an embodiment of this application;

[0020] Figure 2 This is a flowchart of a driving state detection method according to an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of an optional driving state detection system architecture according to an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of an optional multimodal timing fusion scheme according to an embodiment of this application;

[0023] Figure 5 This is a structural block diagram of a driving state detection device according to an embodiment of this application. Detailed Implementation

[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] According to an embodiment of this application, a method embodiment for driving state detection is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, steps shown or described may be executed in a different order than that shown here.

[0027] First, the operating environment of the above method embodiments will be described by way of example. Figure 1 This is a hardware structure block diagram of an optional computing terminal for implementing a driving state detection method according to an embodiment of this application, such as... Figure 1 As shown, the computing terminal 10 (e.g., a computer terminal, a mobile smart terminal, a vehicle terminal, or a cloud computing virtual terminal) may include: one or more processors 102, a memory 104 for storing data, and a transmission device 106 for implementing communication functions. Each processor 102 may include, but is not limited to, a processing component such as a microprocessor (MCU) or a field programmable gate array (FPGA).

[0028] The aforementioned computing terminal 10 may further include: a display device 110, an input / output device 108, a Universal Serial Bus (USB) port (which can be used as one of the ports of a computer bus, not shown in the figure), a network interface (not shown in the figure), a power supply (not shown in the figure), and a camera (not shown in the figure). Those skilled in the art will understand that... Figure 1 The structure of the computing terminal 10 shown is for illustrative purposes only and does not impose strict limitations on the structure of the computing terminal 10 described above. For example, the computing terminal 10 may also include components that are larger than... Figure 1 The more or fewer components shown, or the computing terminal 10 may have the same Figure 1 The components are shown in different categories.

[0029] It should be noted that one or more processors 102 and / or other data processing circuits in the aforementioned computing terminal 10 may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be wholly or partially integrated into any other element in the vehicle terminal 10 (or mobile device).

[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the driving state detection method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned driving state detection method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the vehicle terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0031] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the vehicle terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0032] Under the above operating environment, the embodiments of this application provide the following: Figure 2 The driving status detection method shown is as follows: Figure 2 This is a flowchart of a driving state detection method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following implementation steps S201 to S205.

[0033] Step S201: Acquire the vehicle's driving status signal, steering wheel signal, and visual perception signal while the vehicle is in motion. The steering wheel signal is used to record rotation data corresponding to multiple steering wheel rotation events, and the display content corresponding to the visual perception signal includes the target part of the driver.

[0034] The aforementioned driving status signals refer to data that reflect the current operating status of the vehicle, such as vehicle speed, acceleration, yaw rate, lane departure, etc. These signals are usually collected in real time through the vehicle controller area network (CAN) bus or inertial measurement unit (IMU).

[0035] The aforementioned steering wheel signal is used to record the driver's steering wheel manipulation behavior. This steering wheel signal may include rotational data such as steering wheel angle and angular velocity. Specifically, the sampling frequency of the steering wheel signal can be set to 100Hz to capture subtle driving operation characteristics.

[0036] The aforementioned visual perception signal is acquired by an infrared camera deployed in the cockpit to capture video stream data containing the driver's face (especially target areas such as the face, eyes, and mouth, which are preset as key perception areas). This visual perception signal can work stably under various lighting conditions, ensuring the reliable extraction of subsequent visual features.

[0037] Step S202: Feature extraction is performed based on the steering wheel signal to obtain driving behavior features.

[0038] By refining and mining the steering wheel signals, multiple fatigue-sensitive driving behavior features are extracted. For example, in an application scenario, the angle and angular velocity signals in the steering wheel signals are first filtered and segmented to identify valid steering wheel rotation events. Then, based on these events and a sliding window mechanism, various typical driving behavior features are calculated at multiple time scales, including but not limited to micro-motion entropy, large-motion ratio, stationary duration, operation interval fluctuation, steering distribution pattern, and operation rate correlation features. These driving behavior features can characterize driving behavior from multiple dimensions, such as operation randomness, movement amplitude, stationary behavior, movement rhythm, distribution characteristics, and temporal correlation, significantly improving the expressive power and distinguishability of fatigue-related information in the steering wheel signal.

[0039] Step S203: Extract features based on visual perception signals to obtain visual features.

[0040] For visual perception signals, computer vision algorithms are used to analyze image sequences of target areas (such as a driver's face) to extract visual features closely related to fatigue. For example, in an application scenario, firstly, face detection and keypoint localization algorithms are used to accurately identify the eye region in each frame of the image. Then, based on the distance changes of key points in the eye region (such as the upper and lower eyelids), the eyelid closure ratio (PERCLOS) is calculated, and the frequency of changes in the eyelid closure ratio is statistically analyzed within a continuous time window to form an eye closure temporal feature. This eye closure temporal feature serves as an example of a visual feature. This eye closure temporal feature can effectively characterize the frequency of drowsiness and the duration of continuous eye closure in drivers, exhibiting high discriminative power.

[0041] Step S204: Based on the driving state signal, perform multimodal weighted fusion processing on the driving behavior features and visual features to generate a fused feature vector, wherein the driving state signal is used to determine the fusion weights corresponding to the driving behavior features and the fusion weights corresponding to the visual features, respectively.

[0042] Step S204 described above is a crucial step in achieving scene-adaptive fusion in this application. First, the current driving scene, such as a straight road, a curve, or a congested section, is identified based on driving state signals (e.g., vehicle lateral acceleration, turn signal status, GPS road information). The reliability and discriminative ability of each feature modality differ significantly across different scenes. For example, in a curve, steering wheel behavior is easily affected by road curvature, reducing its reliability, while visual features are less affected by the environment. Therefore, the system dynamically assigns different fusion weights to driving behavior features and visual features according to the scene type, and then concatenates the weighted and scaled multimodal features to form a unified fusion feature vector. This multimodal weighted fusion processing mechanism effectively suppresses the interference of the road environment on feature discriminative power and improves the robustness of the system in complex scenes.

[0043] Step S205: Use the target state detection model to infer and predict the fused feature vector to determine the driving state information corresponding to the vehicle.

[0044] An end-to-end inference analysis of the aforementioned fused feature vectors is performed using a pre-trained target state detection model. This target state detection model may include a deep learning classifier (such as a fully connected network), receives fused feature vectors of a specific dimension (such as 80 dimensions), mines high-order interaction relationships between features through multi-layer nonlinear transformations, and finally outputs a continuous probability value between 0 and 1 after passing through an activation function (such as the sigmoid function). This probability value represents the degree of risk of the driver being fatigued. Driving state information can be used to trigger tiered warnings, such as audible alerts and seat vibrations, thereby effectively enhancing driving safety and preventing traffic accidents caused by fatigued driving.

[0045] Based on steps S201 to S205 above, by dynamically fusing feature data from different modalities, this embodiment of the application can automatically adjust the weight ratio of steering wheel behavior and visual features according to the real-time driving status of the vehicle, thereby effectively overcoming the problem of decreased reliability of a single feature due to environmental factors (such as lighting and road conditions). Furthermore, fusing visual features and driving behavior features not only improves the accuracy of fatigue detection but also enhances the robustness and adaptability of the system, enabling the system to stably and accurately detect the driver's fatigue state in various complex driving scenarios, ensuring driving safety. The multimodal weighted fusion mechanism in this embodiment of the application ensures that even in harsh environments, a high fatigue driving detection rate can be maintained through the enhancement of another modality's data, thereby significantly improving the effectiveness and safety of the driver warning system and enhancing the user's driving experience.

[0046] In other words, the embodiments of this application achieve the goal of accurately detecting fatigue driving under various driving conditions through an adaptive multimodal weighted fusion mechanism, thereby improving the accuracy and robustness of fatigue driving detection and enhancing user driving safety. This solves the technical problems of low accuracy of fatigue driving detection, poor system stability and scenario applicability in related technologies.

[0047] As an optional implementation, step S202 above, which involves extracting features based on the steering wheel signal to obtain driving behavior features, may further include the following execution steps:

[0048] Step S221: Preprocess the steering wheel signal to extract data and obtain multiple data samples. Each data sample includes angle data and angular velocity data corresponding to a steering wheel rotation event.

[0049] Step S222: Perform feature analysis and calculation based on multiple data samples to obtain driving behavior features.

[0050] In step S221, preprocessing data extraction is performed on the steering wheel signal. Specifically, the system first filters and reduces noise from the raw steering wheel angle signal acquired from the vehicle's CAN bus, and identifies discrete steering wheel rotation events based on angular velocity information. Each steering wheel rotation event represents a complete operation intention of the driver and corresponds to a data sample. Each data sample contains a high-resolution angle sequence and angular velocity sequence within the duration of the event, and the sampling rate can be set to 100Hz, thus providing high-quality input for subsequent refined feature calculations.

[0051] Furthermore, based on the multiple data samples obtained in step S221, multi-dimensional feature analysis and calculation are performed to extract driving behavior features that are highly sensitive to fatigue driving.

[0052] For example, the driving behavior features extracted in the embodiments of this application may include the following six typical features.

[0053] The first type is fine-machining entropy features: these are obtained by calculating the information entropy of minute angular changes (e.g., absolute values ​​of angular changes less than 0.5°) over a short period (e.g., within 1 second). Fine-machining entropy features characterize the randomness and disorder of driver micro-operations. They are also known as micro-motion entropy.

[0054] The second category is significant steering frequency characteristics: these are obtained by statistically analyzing the proportion of samples where the absolute value of the steering wheel angle change exceeds 3° within a certain time window. Significant steering frequency characteristics can be characterized by the macro-motion ratio, used to capture the frequency changes of large-amplitude maneuvers.

[0055] The third type is the stationary duration feature: this is obtained by calculating the maximum continuous time during which the absolute value of the angular velocity remains below a set threshold (e.g., 0.3 rad / s) within a rolling time window. The stationary duration feature can also be characterized by zero-operation duration, used to quantify the interval without steering wheel operation.

[0056] The fourth category is the operation interval fluctuation characteristic: this is obtained by identifying the peak point of angular velocity to determine the time of action, and then calculating the standard deviation of the time interval between adjacent actions. The operation interval fluctuation characteristic can be characterized by action interval variability, which is used to assess the stability of the operation rhythm.

[0057] The fifth category is steering distribution morphology: obtained by calculating the kurtosis of the steering wheel angle sequence. Steering distribution morphology can be characterized by angular kurtosis, which describes the sharpness or flatness of the angle distribution.

[0058] The sixth category is the operation rate correlation feature: obtained by calculating the autocorrelation coefficient of the angular velocity sequence under a fixed time delay (e.g., 10 time points corresponding to 0.1 seconds, sampling rate 100Hz). The operation rate correlation feature can be characterized by the velocity autocorrelation parameter and is used to evaluate the temporal consistency of operation habits.

[0059] Through steps S221 to S222 described above, this embodiment of the application achieves a deep transformation of steering wheel signals from raw data to high-discrimination features. The above technical solution, through refined event segmentation and multi-dimensional feature engineering, significantly improves the expressive power of driving behavior features and the sensitivity to fatigue states, providing a reliable data foundation for subsequent multimodal fusion and high-precision fatigue determination.

[0060] As an optional implementation, the driving behavior characteristics include: fine operation entropy characteristics, significant steering frequency characteristics, stationary duration characteristics, operation interval fluctuation characteristics, steering distribution pattern characteristics, and operation rate correlation characteristics. In step S222 above, the driving behavior characteristics are obtained by performing feature analysis and calculation based on multiple data samples, and may also include the following execution steps:

[0061] Step S2221: Calculate the angle change for multiple data samples to determine the fine operation entropy feature and the significant steering frequency feature. The fine operation entropy feature is used to characterize the information entropy of data samples whose steering wheel angle change is less than a first threshold within the first time window. The significant steering frequency feature is used to characterize the proportion of data samples whose steering wheel angle change is greater than a second threshold within the second time window.

[0062] Step S2222: Angular velocity is calculated for multiple data samples to determine the stationary duration feature and the operation interval fluctuation feature. The stationary duration feature is used to characterize the duration during which the steering wheel angular velocity is less than the third threshold within the third time window, and the operation interval fluctuation feature is used to characterize the degree of fluctuation in the interval time of continuous steering actions within the fourth time window.

[0063] Step S2223: Perform angle time series analysis on multiple data samples to determine the steering distribution morphology characteristics, wherein the steering distribution morphology characteristics are used to characterize the distribution kurtosis of the steering wheel angle sequence within the fifth time window.

[0064] Step S2224: Perform angular velocity time series analysis on multiple data samples to determine the operation rate correlation characteristics, wherein the operation rate correlation characteristics are used to characterize the autocorrelation coefficient of the steering wheel angular velocity sequence within the sixth time window.

[0065] In step S2221, the system calculates continuous angle differences from the data sample sequence to obtain an angle change sequence. The fine-grained operation entropy feature focuses on subtle operations performed by the driver under unconscious or fatigued conditions. It is calculated by filtering out all samples with absolute angle changes less than a first threshold (e.g., 0.5°) within a first time window (e.g., 5 seconds) and calculating the information entropy of these micro-changes. A higher information entropy value indicates greater randomness and disorder in the micro-operations, and is positively correlated with fatigue. The significant steering frequency feature focuses on conscious, large-amplitude operations. It calculates the proportion of samples with absolute angle changes exceeding a second threshold (e.g., 3°) within a second time window (e.g., 10 seconds) out of the total sample count; this is the macro-motion ratio, used to quantify the frequency of large-amplitude steering operations.

[0066] In step S2222, calculations are performed based on the angular velocity signal. The stationary duration feature is used to capture situations where the driver's hands are off the steering wheel or remain stationary for an extended period. It is calculated as the maximum continuous time period within a rolling third time window (e.g., 30 seconds) where the absolute value of the angular velocity remains below a third threshold (e.g., 0.3 rad / s), i.e., zero-operation duration. The operation interval fluctuation feature is used to assess the stability of the driver's operating rhythm. The system first identifies the peak point of the angular velocity to determine the time of the operation, and then calculates the standard deviation of the time interval between adjacent actions within a fourth time window (e.g., 60 seconds), i.e., action interval variance. Greater fluctuations indicate a more irregular operating rhythm.

[0067] In step S2223, the steering distribution morphology characteristics are achieved by performing high-order statistical analysis on the steering wheel angle sequence within the fifth time window (e.g., 15 seconds), specifically calculating the kurtosis of the steering wheel angle sequence. Kurtosis reflects the difference between the angle value distribution pattern and the normal distribution; a higher value indicates a sharper distribution (i.e., the operation is concentrated at certain specific angles), while a lower value indicates a flatter distribution (the operation angles are more dispersed).

[0068] In step S2224, the operation rate correlation feature is extracted by analyzing the temporal correlation of the angular velocity sequence. The autocorrelation coefficient of the angular velocity sequence within a sixth time window (e.g., 20 seconds) at a fixed time lag (e.g., 10 time points corresponding to 0.1 seconds, assuming a sampling rate of 100Hz) is calculated. This autocorrelation coefficient reflects the self-similarity and consistency of the driver's operating habits over a short period; a decrease in the autocorrelation coefficient may indicate a decline in attention and operational coherence.

[0069] Through steps S2221 to S2224 described above, this embodiment of the application implements a multi-dimensional and refined steering wheel behavior feature extraction process. The above technical solution overcomes the limitations of traditional methods that rely solely on simple statistical quantities (such as standard deviation), effectively capturing subtle changes in deep behavioral patterns such as operational finesse, rhythm, and consistency under fatigued driving conditions. This provides highly discriminative input features for subsequent deep fusion with visual features, significantly improving the accuracy of fatigue state detection and early warning capabilities.

[0070] As an optional implementation, the visual perception signal includes video stream data collected by the vehicle's infrared camera. In step S203 above, feature extraction is performed based on the visual perception signal to obtain visual features, which may further include the following execution steps:

[0071] Step S231: Using a visual recognition algorithm, identify and determine the key point position sequence of the target part from the continuous image frames of the video stream data;

[0072] Step S232: Analyze the frequency of morphological changes of the target part based on the key point position sequence to obtain visual features.

[0073] In step S231, a visual recognition algorithm deployed on the vehicle-mounted computing unit processes the continuous video stream data acquired by the infrared camera. This visual recognition algorithm first locates the driver's facial region using a face detection model, and then employs a keypoint regression network to accurately identify the key point locations of target areas (such as the eyes). For example, for the eyes, key points may include the midpoint of the upper eyelid, the midpoint of the lower eyelid, and the corner of the eye. Specifically, the system performs the above recognition operation on each frame of the image, thereby outputting a keypoint location sequence that changes over time. This keypoint location sequence accurately records the dynamic changes of the target area in the video stream.

[0074] In step S232, the morphological changes of the target area are quantitatively analyzed based on the keypoint position sequence to obtain visual features. For example, when the target area is the eye, the visual feature is the eye closure timing feature, which is calculated as the percentage of eye closure time over the pupil (PERCLOS). Specifically, firstly, eye morphological parameters are calculated based on the keypoint position sequence. For example, the degree of eye opening and closing is characterized by calculating the normalized distance between the upper and lower eyelid keypoints. Then, within the time window, the proportion of frames with eye morphological parameters below a set threshold (determined as closed state) is counted out of the total number of frames, thus obtaining the PERCLOS value. This PERCLOS value sequence (e.g., one data point per second) constitutes the final eye closure timing feature.

[0075] Through steps S231 to S232 described above, this embodiment of the application realizes an automatic extraction process from the raw video stream to standardized, quantifiable visual features. The above implementation utilizes an infrared camera to overcome interference from changes in ambient lighting, and through precise key point localization and temporal frequency analysis, stably extracts visual features, providing reliable and effective input for subsequent multimodal fusion with steering wheel behavior features, significantly enhancing the accuracy and robustness of the fatigue driving detection system in complex real-world environments.

[0076] As an optional implementation, the target area is the eye, and the visual features include the temporal features of eye closure. In step S232 above, the morphological change frequency analysis of the target area is performed based on the key point position sequence to obtain the visual features. The following execution steps may also be included:

[0077] Step S2321: Using the key point position sequence, the distance between multiple key point pairs of the target part is calculated to obtain the eye morphology parameters, wherein the eye morphology parameters are used to characterize the degree of opening and closing of the eye.

[0078] Step S2322: Based on the eye morphology parameters and the seventh time window, the frequency of eye morphology changes is calculated to obtain the temporal characteristics of eye closure.

[0079] In step S2321, the eye keypoint position sequence obtained from consecutive image frames is used for precise calculation. For each image frame, multiple predefined keypoint pairs are selected, such as the keypoint pair consisting of the midpoint of the upper eyelid and the midpoint of the lower eyelid, and the pixel distance between the keypoint pairs is calculated. This distance value is normalized to obtain the eye morphology parameter (EyeAperture Ratio). The eye morphology parameter is a standardized scalar, with a value range typically between 0 and 1, where 0 represents the eye being completely closed and 1 represents the eye being completely open. Therefore, the eye morphology parameter can accurately quantify the real-time opening and closing degree of the eye.

[0080] In step S2322, based on the eye morphology parameter sequence calculated in step S2321, statistical analysis is performed within a set seventh time window (e.g., 60 seconds) to calculate the frequency of eye morphology changes, ultimately obtaining the eye closure timing feature. Exemplarily, this embodiment uses the percentage of eye closure time per unit time as the core eye closure timing feature. Specifically, within the seventh time window, the proportion of frames where the eye morphology parameter value is consistently below a set threshold (determined as a closed state) is counted out of the total number of frames; this proportion is the PERCLOS value. Further, the seventh time window is slid at fixed time intervals (e.g., per second) and a PERCLOS value is output, thereby forming a continuous eye closure timing feature sequence.

[0081] Through steps S2321 to S2322 described above, this embodiment of the application implements a standardized and quantifiable method for extracting visual fatigue features. The above technical solution, by accurately calculating eye morphological parameters from key points and further transforming them into eye closure timing features with clear physiological significance, effectively overcomes the interference of changes in ambient lighting and individual differences, providing stable and reliable visual input for subsequent multimodal fusion, and significantly improving the accuracy and objectivity of fatigue state discrimination.

[0082] As an optional implementation, step S204 above, which involves performing multimodal weighted fusion processing on driving behavior features and visual features based on the driving state signal to generate a fused feature vector, may further include the following execution steps:

[0083] Step S241: Perform vehicle dynamic feature analysis based on driving status signals to determine the driving scenario type of the vehicle.

[0084] Step S242: Determine the first weight corresponding to the driving behavior features and the second weight corresponding to the visual features based on the driving scenario type.

[0085] Step S243: The driving behavior features and visual features are spliced ​​together to obtain a spliced ​​feature sequence;

[0086] Step S244: Generate a fused feature vector based on the spliced ​​feature sequence, the first weight, and the second weight.

[0087] In step S241, vehicle dynamic characteristics are analyzed based on driving status signals acquired from the vehicle's CAN bus or inertial measurement unit. Driving status signals include, but are not limited to, vehicle lateral acceleration, yaw rate, turn signal status, and road curvature information provided by Global Positioning System (GPS). By analyzing characteristic patterns in the driving status signals—for example, high lateral acceleration combined with a large steering angle may indicate a sharp turn, while low speed accompanied by frequent starts and stops may indicate a congested scene—the type of driving scenario the vehicle is currently in can be accurately determined.

[0088] In step S242, appropriate fusion weights, namely the first weight and the second weight, are assigned to the driving behavior features and visual features respectively according to the driving scenario type. The above weight calculation mechanism is an important part of this application to achieve scene adaptive fusion and improve anti-interference capability.

[0089] In step S243, the driving behavior features (e.g., a six-dimensional feature vector) and visual features (e.g., a 1-dimensional PERCLOS temporal feature encoded by encoding) are concatenated to form a concatenated feature sequence with a higher dimension.

[0090] Furthermore, in step S244, the first weight and the second weight are applied to the corresponding driving behavior feature part and visual feature part in the spliced ​​feature sequence, respectively, to achieve weighted scaling of features of different modalities. Finally, the weighted feature parts are recombined to generate the final fused feature vector, which emphasizes the most reliable feature information in the current driving scenario.

[0091] Through steps S241 to S244 described above, this application embodiment implements a scene-adaptive multimodal feature fusion method. The above technical solution effectively suppresses the interference of external environmental factors such as road curvature and traffic conditions on the discriminative power of single-modal features by intelligently analyzing driving states and dynamically adjusting fusion weights. This significantly improves the robustness and representativeness of the fused feature vectors, laying the foundation for ultimately achieving high-precision and high-reliability fatigue driving state determination.

[0092] As an optional implementation, the first weight includes: a first sub-weight corresponding to the fine operation entropy feature, a second sub-weight corresponding to the significant steering frequency feature, a third sub-weight corresponding to the stationary duration feature, a fourth sub-weight corresponding to the operation interval fluctuation feature, a fifth sub-weight corresponding to the steering distribution pattern feature, and a sixth sub-weight corresponding to the operation rate correlation feature. In step S242 above, determining the first weight corresponding to the driving behavior feature based on the driving scenario type may further include at least one of the following execution steps:

[0093] Step S2421: When the driving scenario type is a sharp turn scenario, the first sub-weight is determined to be 0.3, the second sub-weight is determined to be 0.1, the third sub-weight is determined to be 0.8, the fourth sub-weight is determined to be 0.2, the fifth sub-weight is determined to be 0.4, and the sixth sub-weight is determined to be 0.1.

[0094] Step S2422: When the driving scenario type is a congestion scenario, the first sub-weight is determined to be 0.9, the second sub-weight is determined to be 0.2, the third sub-weight is determined to be 0.7, the fourth sub-weight is determined to be 0.6, the fifth sub-weight is determined to be 0.8, and the sixth sub-weight is determined to be 0.5.

[0095] Step S2423: When the driving scenario type is any other than a sharp turn scenario and a congestion scenario, the first sub-weight is set to 1, the second sub-weight is set to 1, the third sub-weight is set to 1, the fourth sub-weight is set to 1, the fifth sub-weight is set to 1, and the sixth sub-weight is set to 1.

[0096] In steps S2421 to S2423, the system refines the six sub-weights constituting the first weight based on the identified driving scenario type. This is a crucial mechanism for achieving road interference compensation in this application. Based on prior knowledge of the reliability of steering wheel behavior characteristics under different scenarios, this mechanism achieves adaptive fusion by reducing the weight of susceptible features and increasing the importance of stable features.

[0097] For example, this application pre-defines experimentally verified weight configuration schemes for two types of scenarios with significant interference: sharp turn scenarios and congestion scenarios.

[0098] In sharp turn scenarios, large steering maneuvers (corresponding to significant steering frequency characteristics) and steering rhythm (corresponding to steering interval fluctuation characteristics and steering rate correlation characteristics) are highly susceptible to road curvature, significantly reducing reliability. Therefore, the system assigns a low second sub-weight (e.g., 0.1) to significant steering frequency characteristics, a low fourth sub-weight (e.g., 0.2) to steering interval fluctuation characteristics, and a low sixth sub-weight (e.g., 0.1) to steering rate correlation characteristics. Conversely, whether the driver takes their hands off the steering wheel for an extended period (corresponding to the idle time characteristic) is a relatively stable indicator, so a high third sub-weight (e.g., 0.8) is assigned to the idle time characteristic. Fine-grained steering entropy characteristics and steering distribution morphology characteristics are partially disturbed, therefore the first and fifth sub-weights are set to 0.3 and 0.4, respectively.

[0099] In congested traffic scenarios, vehicles frequently start and stop, and normal operation inherently involves numerous micro-operations and periods of stationary waiting. Therefore, the fine-grained operation entropy feature, which effectively characterizes the increased disorder of micro-operations under fatigue conditions, is assigned the highest first sub-weight (e.g., 0.9); the operation interval fluctuation feature, which characterizes the stability of operation rhythm, is assigned a relatively high fourth sub-weight (e.g., 0.6); simultaneously, the stationary duration feature and the steering distribution pattern feature also have high discriminative power, and are assigned weights of 0.7 and 0.8 respectively. Large-amplitude operations (corresponding to significant steering frequency features) are less frequent and have weak discriminative power in this scenario, therefore the second sub-weight is relatively low (e.g., 0.2).

[0100] In other normal scenarios (such as scenarios with gentle road curvature and smooth traffic), all steering wheel behavior features are less affected. The system assigns a baseline value of 1 to all six sub-weights, which means that all features will be included in the fusion process equally without the need for specific suppression or enhancement.

[0101] Through steps S2421 to S2423 described above, this embodiment of the application implements a scene-adaptive feature fusion strategy with finer granularity and higher accuracy. This strategy, by configuring differentiated weights for different sub-features within the driving behavior features, achieves more refined compensation for road interference, maximizing the retention of effective information and suppressing noise interference. This significantly improves the purity and discriminative power of the multimodal fusion features, providing a guarantee for ultimately achieving high-precision and highly robust fatigue state detection.

[0102] As an optional implementation, the target state detection model includes a pre-trained transformation component and a probability mapping component. In step S205 above, the target state detection model is used to infer and predict the fused feature vector to determine the driving state information. This step may also include the following execution steps:

[0103] Step S251: Use the transformation component to perform inference analysis on the fused feature vector and generate inference results;

[0104] Step S252: The probability mapping component is used to perform probability prediction mapping on the inference results to obtain driving state information, wherein the driving state information is used to characterize the probability that the driver is in a state of fatigue driving.

[0105] In step S251, the pre-trained transformation component in the target state detection model is used to perform deep inference analysis on the fused feature vector. The transformation component typically consists of one or more fully connected layers, responsible for receiving the high-dimensional fused feature vector (e.g., 80-dimensional) and mining complex high-order interactions and hidden patterns between different features through a series of nonlinear transformations. The above process maps the input fused feature vector to a lower-dimensional, more discriminative feature space, generating an abstract inference result rich in semantic information.

[0106] In step S252, the probability mapping component in the target state detection model is used to make a final decision on the inference result. The probability mapping component is typically an output layer that uses a sigmoid function as the activation function. This component maps the inference result to a continuous probability value between 0 and 1. This probability value represents the final driving state information, intuitively indicating the confidence level or risk level of whether the driver is currently fatigued. For example, the closer the probability value is to 1, the higher the risk of determining that the driver is fatigued; the closer the probability value is to 0, the greater the likelihood that the driver is awake.

[0107] For example, the target state detection model in this application embodiment can adopt a multilayer perceptron (MLP) structure. The transformation component can consist of a fully connected layer containing 32 neurons, used to perform nonlinear transformation and feature compression on the 80-dimensional fused feature vector. The probability mapping component is a fully connected layer with an output dimension of 1, followed by a sigmoid function, which converts the 32-dimensional inference result into a fatigue probability value between 0 and 1.

[0108] Through steps S251 to S252 described above, this embodiment of the application completes the end-to-end inference process from multimodal fusion features to final fatigue state determination. The above implementation automatically learns the complex mapping relationship between fusion features and fatigue state through a pre-trained neural network model, outputting a continuous probability value with clear physical meaning and easily set thresholds for early warning, significantly improving the objectivity, accuracy, and overall reliability of driving state detection.

[0109] In an exemplary application scenario, embodiments of this application provide, as follows: Figure 3 The diagram shows an optional driving state detection system architecture, such as... Figure 3 As shown, the system data flow originates from two signal acquisition terminals: an infrared camera and a steering wheel sensor. The infrared camera captures infrared images of the driver's face and transmits the data to the visual feature extraction module for processing, extracting fatigue-related biometrics (i.e., visual features). The steering wheel sensor acquires raw steering wheel rotation signals in real time via the vehicle's CAN bus. These signals first enter a signal preprocessing module for filtering and event segmentation, and then a six-dimensional feature calculation module (i.e., driving behavior feature extraction) calculates a refined steering wheel behavior feature vector. Simultaneously, the vehicle status module (i.e., driving status signal acquisition and analysis) continuously acquires vehicle dynamic information to determine the current driving scenario. The dynamic weighting module (i.e., multimodal weighted fusion processing) receives the scenario judgment results from the vehicle status module and assigns appropriate fusion weights to the visual and driving behavior features. The weighted features are then fed into a multimodal fusion model for concatenation and deep inference. The model output enters the fatigue state decision module (i.e., the target state detection model) to generate the final fatigue probability value. Ultimately, the graded alarm system triggers the corresponding level of early warning based on the fatigue probability value, completing the closed loop from state perception to safety intervention.

[0110] Through such Figure 3 The driving state detection system architecture shown in this application, when applied to specific scenarios, enables fully automated processing from multi-source signal acquisition, refined feature extraction, scene-adaptive weighted fusion to final state decision-making and alarm triggering. This system architecture effectively integrates the advantages of both visual and behavioral modalities and suppresses road interference through a dynamic weight allocation mechanism, significantly improving the accuracy, robustness, and practicality of fatigue state detection, thus providing a highly reliable driver state monitoring solution for intelligent connected vehicles.

[0111] In an exemplary application scenario, embodiments of this application provide, as follows: Figure 4 One possible multimodal timing fusion scheme is shown, such as Figure 4 As shown, the scheme receives two parallel feature streams as input. One stream is a PERCLOS sequence (i.e., temporal features of eye closure), which represents the PERCLOS values ​​calculated within a continuous time window (e.g., 60 seconds). The PERCLOS sequence is input into a Bidirectional Gated Recurrent Unit (BiGRU) network containing 32 hidden units. Figure 4BiGRU temporal modeling is performed using a method labeled BIGRU-32 to fully capture the long-term contextual dependence and forward / backward variation patterns of eye-closing frequency, outputting a high-order 64-dimensional temporal feature representation. Another path is the real-time computed six-dimensional feature vector (i.e., driving behavior features), which is passed through a fully connected (dense layer) containing 16 neurons. Figure 4 The data (labeled as Dense-16) undergoes nonlinear transformation and feature compression to obtain a 16-dimensional high-order representation. Subsequently, the two feature streams are fused in a feature concatenation step, combining the 64-dimensional visual temporal features with the 16-dimensional behavioral features to form an 80-dimensional fused feature vector. This fused feature vector is finally input into a sigmoid function unit for final sigmoid classification, outputting a scalar probability value between 0 and 1, directly representing the risk level of driver fatigue.

[0112] Through such Figure 4 The multimodal temporal fusion scheme illustrated in this application, when applied to specific scenarios, enables deep collaboration and information complementarity between visual and behavioral modalities at the temporal level. This scheme not only effectively models the temporal evolution of fatigue biosignatures using a BiGRU network, but also mines deep patterns of behavioral features through a fully connected network. Finally, efficient feature fusion and state decision-making are achieved through splicing and nonlinear transformation, significantly improving the accuracy of fatigue state discrimination and the ability to capture temporal dynamics, thereby enhancing the overall performance and reliability of the system in real-world driving environments.

[0113] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0114] It should be noted that, for the sake of simplicity, the technical solutions in the above method embodiments are described as a series of actions. However, those skilled in the art should understand that this application is not limited to the order of actions in the described action combination, because according to this application, some of the above steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this application specification are preferred embodiments, and the actions and modules involved are not necessarily essential for implementing the technical solutions of this application.

[0115] According to embodiments of the present invention, a device embodiment for a driving state detection apparatus is also provided. This driving state detection apparatus is used to implement the above-described method embodiment and various optional implementations of the method embodiment. The technical content already described above will not be repeated in the device embodiment. It should be noted that, in the following related descriptions of the device embodiment, a "module" can be software, hardware, or a combination of software and hardware used to implement a specified function.

[0116] Figure 5 This is a structural block diagram of a driving state detection device according to an embodiment of this application, such as... Figure 5 As shown, the driving state detection device includes: an acquisition module 501, used to acquire driving state signals, steering wheel signals, and visual perception signals of the vehicle during driving, wherein the steering wheel signals are used to record rotation data corresponding to multiple steering wheel rotation events, and the display content corresponding to the visual perception signals includes the target parts of the driver; a first extraction module 502, used to perform feature extraction based on the steering wheel signals to obtain driving behavior features; a second extraction module 503, used to perform feature extraction based on the visual perception signals to obtain visual features; a processing module 504, used to perform multimodal weighted fusion processing on the driving behavior features and visual features according to the driving state signals to generate a fused feature vector, wherein the driving state signals are used to determine the fusion weights corresponding to the driving behavior features and the fusion weights corresponding to the visual features; and a detection module 505, used to use a target state detection model to infer and predict the fused feature vector to determine the driving state information corresponding to the vehicle.

[0117] Optionally, the first extraction module 502 is further configured to: perform preprocessing data extraction on the steering wheel signal to obtain multiple data samples, wherein each data sample includes angle data and angular velocity data corresponding to a steering wheel rotation event; and perform feature analysis calculation based on the multiple data samples to obtain driving behavior features.

[0118] Optionally, the driving behavior features include: fine operation entropy features, significant steering frequency features, stationary duration features, operation interval fluctuation features, steering distribution pattern features, and operation rate correlation features. The first extraction module 502 is further used to: calculate the angle change of multiple data samples to determine the fine operation entropy features and significant steering frequency features, wherein the fine operation entropy features are used to characterize the information entropy of data samples whose steering wheel angle change is less than a first threshold within a first time window, and the significant steering frequency features are used to characterize the proportion of data samples whose steering wheel angle change is greater than a second threshold within a second time window in the multiple data samples; and to calculate the angular velocity of the multiple data samples. The system calculates the characteristics of stationary duration and operation interval fluctuation. The stationary duration characteristic is used to characterize the duration during which the steering wheel angular velocity is less than the third threshold within the third time window, and the operation interval fluctuation characteristic is used to characterize the degree of fluctuation in the interval time of continuous steering actions within the fourth time window. Angle time-series analysis is performed on multiple data samples to determine the steering distribution morphology characteristics, which are used to characterize the kurtosis of the steering wheel angle sequence within the fifth time window. Angle velocity time-series analysis is performed on multiple data samples to determine the operation rate correlation characteristics, which are used to characterize the autocorrelation coefficient of the steering wheel angular velocity sequence within the sixth time window.

[0119] Optionally, the visual perception signal includes video stream data collected by the vehicle's infrared camera. The second extraction module 503 is further configured to: use a visual recognition algorithm to identify and determine the key point position sequence of the target part from the continuous image frames of the video stream data; and perform morphological change frequency analysis of the target part based on the key point position sequence to obtain visual features.

[0120] Optionally, the target area is the eye, and the visual features include the eye closure timing features. The second extraction module 503 is further used to: calculate the distance between multiple key point pairs of the target area using the key point position sequence to obtain eye morphology parameters, wherein the eye morphology parameters are used to characterize the degree of eye opening and closing; and obtain the eye closure timing features based on the eye morphology parameters and the seventh time window to calculate the frequency of eye morphology changes.

[0121] Optionally, the above processing module 504 is further configured to: perform vehicle dynamic feature analysis based on driving state signals to determine the driving scenario type in which the vehicle is located; determine the first weight corresponding to the driving behavior features and the second weight corresponding to the visual features according to the driving scenario type; perform splicing processing on the driving behavior features and visual features to obtain a spliced ​​feature sequence; and generate a fused feature vector based on the spliced ​​feature sequence, the first weight and the second weight.

[0122] Optionally, the first weight includes: a first sub-weight corresponding to the fine operation entropy feature, a second sub-weight corresponding to the significant steering frequency feature, a third sub-weight corresponding to the stationary duration feature, a fourth sub-weight corresponding to the operation interval fluctuation feature, a fifth sub-weight corresponding to the steering distribution pattern feature, and a sixth sub-weight corresponding to the operation rate correlation feature. The processing module 504 is further configured to: when the driving scenario type is a sharp turn scenario, determine the first sub-weight as 0.3, the second sub-weight as 0.1, the third sub-weight as 0.8, the fourth sub-weight as 0.2, and the fifth sub-weight as 0. 4. The sixth sub-weight is set to 0.1; when the driving scenario type is a congestion scenario, the first sub-weight is set to 0.9, the second sub-weight to 0.2, the third sub-weight to 0.7, the fourth sub-weight to 0.6, the fifth sub-weight to 0.8, and the sixth sub-weight to 0.5; when the driving scenario type is any other than a sharp turn scenario and a congestion scenario, the first sub-weight is set to 1, the second sub-weight to 1, the third sub-weight to 1, the fourth sub-weight to 1, the fifth sub-weight to 1, and the sixth sub-weight to 1.

[0123] Optionally, the target state detection model includes a pre-trained transformation component and a probability mapping component. The detection module 505 is further used to: use the transformation component to perform inference analysis on the fused feature vector to generate inference results; and use the probability mapping component to perform probability prediction mapping on the inference results to obtain driving state information, wherein the driving state information is used to characterize the probability that the driver is in a fatigued driving state.

[0124] It should be noted that the above-mentioned acquisition module 501, first extraction module 502, second extraction module 503, processing module 504 and detection module 505 correspond to steps S201 to S205 in the method embodiment, respectively. The five modules are the same as the instances and application scenarios implemented by the corresponding steps, but are not limited to the content disclosed in the above method embodiment.

[0125] It should be noted that the modules mentioned in the above device embodiments can be implemented by software, hardware, or a combination of both. For example, when the modules are implemented by hardware, they can be placed in the same processor, or they can be placed in different processors in any combination. As another example, the modules can be hardware or software components stored in memory and processed by one or more processors; they can also run as part of a computing terminal.

[0126] Embodiments of this application also provide a vehicle, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the driving state detection method in various embodiments of this application during runtime.

[0127] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to execute the driving state detection method of various embodiments of this application.

[0128] Optionally, the aforementioned computer storage media may include, but are not limited to: hard disk drives (HDDs), solid state drives (SSDs), USB flash drives, optical discs, memory cards, cloud storage media, and network attached storage (NAS) devices.

[0129] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the driving state detection method in various embodiments of this application.

[0130] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the driving state detection method in various embodiments of this application.

[0131] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.

[0132] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0135] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0137] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A driving state detection method, characterized in that, include: The vehicle acquires driving status signals, steering wheel signals, and visual perception signals while in motion. The steering wheel signals are used to record rotation data corresponding to multiple steering wheel rotation events, and the display content corresponding to the visual perception signals includes the target parts of the driver. Driving behavior features are obtained by extracting features from the steering wheel signals. Visual features are obtained by extracting features based on the visual perception signals. Based on the driving state signal, the driving behavior features and the visual features are subjected to multimodal weighted fusion processing to generate a fused feature vector, wherein the driving state signal is used to determine the fusion weights corresponding to the driving behavior features and the fusion weights corresponding to the visual features, respectively. The target state detection model is used to infer and predict the fused feature vector to determine the driving state information corresponding to the vehicle.

2. The driving state detection method according to claim 1, characterized in that, Based on the steering wheel signal, feature extraction is performed to obtain the driving behavior features, including: The steering wheel signal is preprocessed and data is extracted to obtain multiple data samples, wherein each data sample includes angle data and angular velocity data corresponding to a steering wheel rotation event; The driving behavior features are obtained by performing feature analysis and calculation based on the multiple data samples.

3. The driving state detection method according to claim 2, characterized in that, The driving behavior characteristics include: fine operation entropy characteristics, significant steering frequency characteristics, stationary duration characteristics, operation interval fluctuation characteristics, steering distribution pattern characteristics, and operation rate correlation characteristics. Based on the multiple data samples, feature analysis and calculation are performed to obtain the driving behavior characteristics, which include: The angle change is calculated for the multiple data samples to determine the fine operation entropy feature and the significant steering frequency feature. The fine operation entropy feature is used to characterize the information entropy of the data sample whose steering wheel angle change is less than a first threshold in the first time window. The significant steering frequency feature is used to characterize the proportion of the data sample whose steering wheel angle change is greater than a second threshold in the multiple data samples in the second time window. Angular velocity is calculated on the multiple data samples to determine the stationary duration feature and the operation interval fluctuation feature. The stationary duration feature is used to characterize the duration during which the steering wheel angular velocity is less than a third threshold within a third time window, and the operation interval fluctuation feature is used to characterize the degree of fluctuation in the interval time of continuous steering actions within a fourth time window. Angle time series analysis is performed on the multiple data samples to determine the steering distribution morphology characteristics, wherein the steering distribution morphology characteristics are used to characterize the distribution kurtosis of the steering wheel angle sequence within the fifth time window; An angular velocity time series analysis is performed on the multiple data samples to determine the operation rate correlation feature, wherein the operation rate correlation feature is used to characterize the autocorrelation coefficient of the steering wheel angular velocity sequence within the sixth time window.

4. The driving state detection method according to claim 1, characterized in that, The visual perception signal includes video stream data captured by the vehicle's infrared camera. Feature extraction is performed based on the visual perception signal to obtain the visual features, which include: Using a visual recognition algorithm, the key point position sequence of the target part is identified and determined from consecutive image frames of the video stream data; Based on the key point location sequence, the morphological change frequency analysis of the target part is performed to obtain the visual features.

5. The driving state detection method according to claim 4, characterized in that, The target area is the eye, and the visual features include eye closure timing features. Based on the key point position sequence, a morphological change frequency analysis of the target area is performed to obtain the visual features, which include: Using the key point location sequence, distance calculations are performed on multiple key point pairs of the target area to obtain eye morphology parameters, wherein the eye morphology parameters are used to characterize the degree of opening and closing of the eye. The eye closure timing characteristics are obtained by calculating the frequency of eye morphological changes based on the eye morphology parameters and the seventh time window.

6. The driving state detection method according to claim 1, characterized in that, Based on the driving state signal, the driving behavior features and the visual features are subjected to multimodal weighted fusion processing to generate the fused feature vector, including: Based on the driving status signal, perform vehicle dynamic feature analysis to determine the driving scenario type of the vehicle. Based on the driving scenario type, determine the first weight corresponding to the driving behavior feature and the second weight corresponding to the visual feature; The driving behavior features and the visual features are spliced ​​together to obtain a spliced ​​feature sequence. The fused feature vector is generated based on the spliced ​​feature sequence, the first weight, and the second weight.

7. The driving state detection method according to claim 6, characterized in that, The first weight includes: a first sub-weight corresponding to the fine operation entropy feature, a second sub-weight corresponding to the significant steering frequency feature, a third sub-weight corresponding to the stationary duration feature, a fourth sub-weight corresponding to the operation interval fluctuation feature, a fifth sub-weight corresponding to the steering distribution pattern feature, and a sixth sub-weight corresponding to the operation rate correlation feature. The first weight corresponding to the driving behavior feature is determined based on the driving scenario type, including: When the driving scenario type is a sharp turn scenario, the first sub-weight is determined to be 0.3, the second sub-weight is determined to be 0.1, the third sub-weight is determined to be 0.8, the fourth sub-weight is determined to be 0.2, the fifth sub-weight is determined to be 0.4, and the sixth sub-weight is determined to be 0.

1. When the driving scenario type is a congestion scenario, the first sub-weight is determined to be 0.9, the second sub-weight is determined to be 0.2, the third sub-weight is determined to be 0.7, the fourth sub-weight is determined to be 0.6, the fifth sub-weight is determined to be 0.8, and the sixth sub-weight is determined to be 0.

5. When the driving scenario type is a scenario other than the sharp turn scenario and the congestion scenario, the first sub-weight is determined to be 1, the second sub-weight is determined to be 1, the third sub-weight is determined to be 1, the fourth sub-weight is determined to be 1, the fifth sub-weight is determined to be 1, and the sixth sub-weight is determined to be 1.

8. The driving state detection method according to claim 1, characterized in that, The target state detection model includes a pre-trained transformation component and a probability mapping component. The target state detection model is used to infer and predict the fused feature vector to determine the driving state information, including: The transformation component is used to perform inference analysis on the fused feature vector to generate inference results; The probability mapping component is used to perform probability prediction mapping on the inference result to obtain the driving state information, wherein the driving state information is used to characterize the probability that the driver is in a state of fatigued driving.

9. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program executes the driving state detection method according to any one of claims 1 to 8 when it runs.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the driving state detection method according to any one of claims 1 to 8.