Method, system and electronic device for detecting wheel spin and storage medium
By combining video frame data and multi-frame image sequence analysis, a mask map is generated and the key points of the hand are identified, which solves the problem that traditional sensors are easily bypassed and achieves more accurate steering wheel off-hand detection.
Patent Information
- Application Number
- CN202411504312.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-10-25
AI Technical Summary
In existing technologies, sensor solutions are easily circumvented when detecting whether the driver has taken their hands off the wheel, resulting in insufficient reliability and accuracy of detection, which poses a safety hazard, especially in intelligent assisted driving environments.
By acquiring video frame data of the steering wheel area, locating images of the steering wheel and hands, generating a mask map and identifying key hand position information, detecting single-frame images by combining preset thresholds, and making judgments by integrating feature vectors from multiple image sequences, the accuracy and stability of detection are improved.
It effectively reduces detection errors caused by factors such as hand posture or gloves, improves the accuracy and anti-interference ability of steering wheel off-hand detection, and ensures driving safety.
Smart Images

Figure CN119380413B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automobiles, in particular to a steering wheel hand-off detection method, a steering wheel hand-off detection system, an electronic device and a computer readable storage medium. BACKGROUND
[0002] With the gradual popularization of L2-level intelligent assisted driving, the safety of vehicle driving is increasingly concerned. Intelligent assisted driving, as an intermediate stage of the development of autonomous driving technology, still requires the participation of the driver. Therefore, the monitoring of the state of the driver is particularly important to ensure the safe driving of the vehicle. The driver monitoring system (DMS) is born for this purpose, which improves driving safety and comfort by analyzing the behavior and physiological state of the driver in real time.
[0003] In the existing driver monitoring system, fatigue and distraction detection technology has been developed. By installing fixed-position cameras in the cockpit, the system can capture the driver's facial image in real time, and analyze the facial features through artificial intelligence (AI) algorithms to determine the fatigue and distraction state of the driver. These cameras are usually installed in the middle of the A-pillar or above the steering column to ensure that the driver's face can be monitored by the camera throughout the normal driving state.
[0004] Steering wheel hand-off detection, as a key technology to ensure that the driver maintains control of the vehicle, its traditional solution mainly relies on steering wheel sensors, which are used to monitor the position of the driver's hands. By monitoring the contact between the hands and the steering wheel, the system can determine in real time whether the driver has a hand-off behavior, and issue a warning when a hand-off situation is detected to remind the driver to hold the steering wheel again.
[0005] Currently, the mainstream steering wheel sensors include the following three types:
[0006] 1. Torque induction sensor: This type of sensor determines whether the driver is controlling the vehicle by measuring the torque change of the steering wheel. When the driver holds the steering wheel, the steering wheel will be subjected to a certain torque force, and the sensor can accurately sense the change of this torque force. If the torque force of the steering wheel decreases or disappears, the system will determine that the driver may have a hand-off behavior.
[0007] 2. Grip induction sensor: This sensor is usually installed on the inner side of the steering wheel, which determines whether the driver continuously holds the steering wheel by sensing the pressure or grip degree of the driver's hands. If the grip decreases or disappears, the system will issue an alarm to prompt the driver to hold the steering wheel tightly again.
[0008] 3. Capacitive sensor: This type of sensor detects hand contact by measuring changes in capacitance between the driver's hand and the steering wheel. When the driver touches the steering wheel, a small change in capacitance occurs, which the sensor can detect to determine whether the driver is holding the steering wheel. Capacitive sensors have the advantages of high sensitivity and strong environmental adaptability, and can effectively identify slight touch or disengagement.
[0009] Although these detection schemes have certain effects in practical applications, they still have certain limitations. For example, hand gestures or sensor sensitivity may affect the accuracy of detection. With the popularization of intelligent assisted driving technology, some drivers fix water bottles or other objects on the steering wheel, or even use metal stickers to cover the capacitive sensors of the steering wheel, in order to evade detection by the sensors. These behaviors cause the driver to be out of hand for a long time in the assisted driving mode, increasing the safety hazards during driving.
[0010] In view of the limitations of traditional steering wheel hand-off detection schemes, a new generation of steering wheel hand-off detection technology begins to combine more sensor types and intelligent algorithms to improve the accuracy and reliability of detection. The goal is to be able to quickly and accurately issue warning signals when potential dangerous driving behavior by the driver is detected.
[0011] The current process of steering wheel hand-off detection generally includes the following steps:
[0012] S1: When the vehicle starts driving or reaches a certain speed, the system starts and the steering wheel sensor begins to operate.
[0013] S2: The steering wheel sensor feeds back the detection signal to the vehicle system.
[0014] S3: The vehicle system determines whether the driver has a hand-off behavior according to internal logic and pre-set calibration data.
[0015] S4: If a hand-off behavior is detected, the system issues a prompt tone or other warning signal; otherwise, continue detection.
[0016] Since the current hand-off detection system relies entirely on sensor feedback signals, it is easy to be evaded by operations such as interfering with the work of the sensor by external means. Therefore, the hand-off detection system relying only on sensors has certain reliability problems in practical applications, and needs to be combined with more accurate sensors and algorithm optimization to further improve driving safety. SUMMARY
[0017] In view of the above problems, the embodiments of the present application are proposed in order to provide a steering wheel hand-off detection method, a steering wheel hand-off detection system, an electronic device and a computer readable storage medium that overcome the above problems or at least partially solve the above problems.
[0018] To solve the above problems, the embodiment of the present application discloses a steering wheel off-hand detection method, the method comprising: acquiring video frame data of a steering wheel region; locating a steering wheel image and a hand image from a single frame image of the video frame data; acquiring a mask image of the steering wheel image and key point position information of the hand image; detecting the single frame image according to the mask image, the key point position information and a preset determination threshold to obtain a single frame image detection result; determining whether there is a steering wheel off-hand condition based on the single frame image detection result; in the case where it is determined that there is no steering wheel off-hand condition according to the single frame image detection result, extracting a feature vector of a multi-frame image sequence of the video frame data; and comprehensively detecting whether there is a steering wheel off-hand condition according to the feature vector of the multi-frame image sequence and the single frame image detection result.
[0019] Optionally, the detecting the single frame image according to the mask image, the key point position information and a preset determination threshold to obtain a single frame image detection result comprises: calculating an occupying ratio information of a finger key point in the mask image in the single frame image according to the mask image and the key point position information; and comparing the occupying ratio information with the determination threshold to obtain the single frame image detection result.
[0020] Optionally, the determining whether there is a steering wheel off-hand condition based on the single frame image detection result comprises: in the case where the determination threshold is an extreme ratio of the finger key point outside the mask image, determining that there is a steering wheel off-hand condition based on the single frame image detection result indicating that the occupying ratio information is greater than the determination threshold; and determining that there is no steering wheel off-hand condition based on the single frame image detection result indicating that the occupying ratio information is less than or equal to the determination threshold.
[0021] Optionally, the comprehensively detecting whether there is a steering wheel off-hand condition according to the feature vector of the multi-frame image sequence and the single frame image detection result comprises: merging the feature vector and normalized data of the key point position information into a feature vector matrix; classifying the feature vector matrix to obtain a classification result of the multi-frame image sequence; in the case where the classification result of the multi-frame image sequence is greater than a preset classification threshold, determining whether there is a single frame image detection result satisfying a preset steering wheel off-hand condition in the multi-frame image sequence; and in the case where there is a single frame image detection result satisfying the preset steering wheel off-hand condition in the multi-frame image sequence, determining that there is a steering wheel off-hand condition.
[0022] Optionally, after the classification of the feature vector matrix obtains the classification result of the multi-frame image sequence, the method further comprises: in the case that the classification result of the multi-frame image sequence is less than or equal to the preset classification threshold, determining that there is no steering wheel hand-off situation.
[0023] Optionally, after the classification of the feature vector matrix obtains the classification result of the multi-frame image sequence, the method further comprises: in the case that the classification result of the multi-frame image sequence is greater than the preset classification threshold, and there is no single-frame image detection result satisfying the preset steering wheel hand-off condition in the multi-frame image sequence, determining that there is no steering wheel hand-off situation.
[0024] Optionally, the obtaining of the mask graph of the steering wheel image and the key point position information of the hand image comprises: performing edge detection processing and ellipse matching processing on the steering wheel image to obtain the mask graph; and rotating the hand image through corresponding corner points, inputting the hand image in a positive hand posture to a BlazeHand model, and outputting the key point position information.
[0025] The embodiment of the application further discloses a steering wheel hand-off detection system, which comprises: a video frame data acquisition module configured to acquire video frame data of a steering wheel region; an image positioning module configured to position a steering wheel image and a hand image from a single-frame image of the video frame data; a mask graph and key point position acquisition module configured to acquire a mask graph of the steering wheel image and key point position information of the hand image; a single-frame image detection result determination module configured to determine a single-frame image detection result by detecting the single-frame image according to the mask graph, the key point position information and a preset determination threshold; a hand-off situation determination module configured to determine whether there is a steering wheel hand-off situation based on the single-frame image detection result; a feature vector extraction module configured to extract a feature vector of a multi-frame image sequence of the video frame data in the case that it is determined that there is no steering wheel hand-off situation according to the single-frame image detection result; and a hand-off situation detection module configured to comprehensively detect whether there is a steering wheel hand-off situation according to the feature vector of the multi-frame image sequence and the single-frame image detection result.
[0026] Optionally, the single-frame image detection result determination module comprises: a proportion information calculation module configured to calculate proportion information of a finger key point in the mask graph in the single-frame image according to the mask graph and the key point position information; and a proportion threshold comparison module configured to compare the proportion information with the determination threshold to obtain the single-frame image detection result.
[0027] Optionally, the hand-off condition determining module is configured to determine that the steering wheel is in a hand-off condition when the ratio of the maximum finger key point to the mask graph is greater than the determination threshold; and determine that the steering wheel is not in the hand-off condition when the ratio of the maximum finger key point to the mask graph is less than or equal to the determination threshold.
[0028] Optionally, the hand-off condition determining module is configured to include a feature vector merging module configured to merge the feature vector and the normalized data of the key point position information into a feature vector matrix; a vector matrix classification module configured to classify the feature vector matrix to obtain a multi-frame image sequence classification result; a single-frame result judging module configured to determine whether the single-frame image detection result meeting the preset steering wheel hand-off condition exists in the multi-frame image sequence when the multi-frame image sequence classification result is greater than a preset classification threshold; and a steering wheel hand-off determining module configured to determine that the steering wheel is in the hand-off condition when the single-frame image detection result meeting the preset steering wheel hand-off condition exists in the multi-frame image sequence.
[0029] Optionally, the system further includes a steering wheel non-hand-off determining module configured to determine that the steering wheel is not in the hand-off condition when the multi-frame image sequence classification result is less than or equal to the preset classification threshold after the vector matrix classification module classifies the feature vector matrix to obtain the multi-frame image sequence classification result.
[0030] Optionally, the steering wheel non-hand-off determining module is further configured to determine that the steering wheel is not in the hand-off condition when the multi-frame image sequence classification result is greater than the preset classification threshold and the single-frame image detection result meeting the preset steering wheel hand-off condition does not exist in the multi-frame image sequence after the vector matrix classification module classifies the feature vector matrix to obtain the multi-frame image sequence classification result.
[0031] Optionally, the mask graph and key point position obtaining module includes a mask graph obtaining module configured to perform edge detection processing and ellipse matching processing on the steering wheel image to obtain the mask graph; and a key point obtaining module configured to rotate the hand image through a corresponding corner point, input the hand image in a positive hand posture to a BlazeHand model, and output the key point position information.
[0032] The embodiment of the present application further discloses an electronic device, including one or more processors, and one or more machine readable media having instructions stored thereon that, when executed by the one or more processors, cause the electronic device to perform the method for detecting hand-off of a steering wheel as described above.
[0033] The application also discloses a computer readable storage medium, which stores a computer program enabling a processor to execute the steering wheel hand-off detection method.
[0034] The application has the following advantages:
[0035] The steering wheel hand-off detection scheme provided by the application obtains video frame data of a steering wheel region. Then, a steering wheel image and a hand image are located from a single frame image of the video frame data. Furthermore, a mask image of the steering wheel image and key point position information of the hand image are obtained. Next, the single frame image is detected according to the mask image, the key point position information and a preset determination threshold to obtain a single frame image detection result, and whether there is a steering wheel hand-off condition is determined based on the single frame image detection result. When it is determined according to the single frame image detection result that there is no steering wheel hand-off condition, feature vectors of a multi-frame image sequence of the video frame data are extracted, and then whether there is a steering wheel hand-off condition is comprehensively detected according to the feature vectors of the multi-frame image sequence and the single frame image detection result.
[0036] The application not only relies on the detection result of the single frame image, but also introduces the feature vectors of the multi-frame image sequence for comprehensive detection. This combination of single frame image and multi-frame image detection makes up for the accidental errors that may be caused by single frame detection. Compared with the method of relying only on a single sensor or simple image recognition, the application is more robust and can effectively improve the accuracy of detection. By obtaining the mask image of the steering wheel image and the key point position information of the hand image, the hand state of the driver can be more accurately recognized. The mask image is an image processing technology that highlights the steering wheel region, and in combination with the hand key point position information, it can be determined whether the driver's hand is off the steering wheel. This is more accurate than relying on sensor detection of contact force, torque changes and other physical signals, and reduces the detection errors caused by the driver's hand posture or gloves and other factors. In the detection process, a preset determination threshold is introduced, that is, a reasonable index is set to judge the hand-off condition. The determination mechanism combining the single frame image detection result and the feature vectors of the multi-frame image sequence makes the hand-off detection more intelligent. By extracting the feature vectors of the multi-frame image sequence, the operation state of the driver in a period of time can be continuously monitored, rather than relying only on a single frame image at a certain moment. This continuous image sequence analysis can capture the dynamic characteristics of the hand-off behavior and avoid false judgments caused by detection errors at a certain moment.
[0037] In summary, the application introduces multi-frame image sequence analysis, mask image combined with hand key points, intelligent threshold judgment and other technical means, not only overcomes the shortcomings of traditional sensor solutions, but also improves the accuracy, stability and safety of the overall detection. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 is a step flow chart of a steering wheel hand-off detection method according to an embodiment of the present application;
[0039] Figure 2 is an image diagram collected in a multi-scenario steering wheel hand-off detection scheme of multi-model joint inter-frame checking according to an embodiment of the present application;
[0040] Figure 3 is an algorithm flow diagram of a multi-scenario steering wheel hand-off detection scheme of multi-model joint inter-frame checking according to an embodiment of the present application;
[0041] Figure 4a is a structure diagram of a hand key point position according to an embodiment of the present application;
[0042] Figure 4b is a structure diagram of a hand key point position after normalization operation according to an embodiment of the present application;
[0043] Figure 5 is a whole structure diagram of a GRU model according to an embodiment of the present application;
[0044] Figure 6 is a structure block diagram of a steering wheel hand-off detection system according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the above objectives, characteristics and advantages of the present application more apparent, comprehensible and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0046] The steering wheel hand-off detection scheme according to the embodiments of the present application detects whether the driver is hand-off through video frame data acquisition and processing, and makes more accurate judgment by combining single frame image analysis and multi-frame image sequence analysis. Specifically, the steering wheel image and the hand image are located from the single frame image of the video frame data, the mask diagram of the steering wheel image is generated, and the key point position information of the hand image is recognized. Then, the single frame image is detected according to the mask diagram, the key point position information and the preset judgment threshold, and whether there is a hand-off situation is judged. In the case where the single frame image does not exist the hand-off situation, the feature vector of the multi-frame image sequence is extracted, and the single frame image detection result and the feature vector of the multi-frame image sequence are comprehensively determined to finally determine whether there is a hand-off situation. The accuracy, reliability and anti-interference ability of the hand-off detection are improved.
[0047] Reference Figure 1 is a step flow chart of a steering wheel hand-off detection method according to an embodiment of the present application. The steering wheel hand-off detection method can be applied to a steering wheel hand-off detection system, which is referred to as a system. The steering wheel hand-off detection method can specifically include the following steps:
[0048] Step 101, acquire video frame data of the steering wheel area.
[0049] In embodiments of the present application, the system captures video frame data of the driver and the steering wheel area in real time through a camera installed in the vehicle. Video frame data refers to the image at each instant in the video stream, which records dynamic scenes by continuously acquiring frames. The acquisition operation here can be real-time collection, meaning that the system can continuously and synchronously capture video frame data without delay, reflecting the driver's behavior changes in a timely manner.
[0050] For example, assume a car equipped with an intelligent auxiliary driving system starts driving. After the system starts, the camera continuously records the driver's hand and the steering wheel area. If the frame rate of the camera is 30 frames per second (fps), the system will collect 30 images per second for analysis. These images can accurately capture the dynamic changes of the driver's hands, whether they are holding the steering wheel, lifting their hands, or moving their hand position. The system can quickly respond through these images.
[0051] Step 102, locate the steering wheel image and hand image from the single frame image of the video frame data.
[0052] In embodiments of the present application, each video frame data contains a single frame image, and the system locates the steering wheel and the driver's hands from these single frame images. Locating refers to finding the location of the target object in the image through algorithms, which here refers to the area of the steering wheel and the hand. Object detection algorithms based on deep learning are usually used to achieve this, such as Convolutional Neural Networks (CNN), which can automatically detect and label specific objects in images.
[0053] For example, after the system receives a certain frame image, it first needs to determine the location of the steering wheel and the hand. For example, the system analyzes the circular structure (steering wheel) and the shape of the human hand (such as fingers, palm) to lock their areas. If a frame image shows that the driver's hands are at the 10 o'clock and 2 o'clock positions of the steering wheel, the system will accurately label these two positions in the image. Even if the driver's hand posture changes slightly (for example, the hand slides a few centimeters), the system can still dynamically track these changes.
[0054] Step 103, acquire the mask image of the steering wheel image and the key point position information of the hand image.
[0055] In embodiments of the present application, the mask image is a binary image used to mark specific regions in the image, which marks the steering wheel region with white or other colors, while other parts are covered. By generating a mask image of the steering wheel, the system can more clearly highlight the outline of the steering wheel. The key points of the hand image refer to specific joints or skeletal nodes of the hand, such as finger tips, joint positions, etc., which are marked by pose estimation algorithms. The key point position information of the hand can help the system understand the specific pose of the hand and its accurate position relative to the steering wheel.
[0056] For example, assuming that the system detects the steering wheel region, it will generate a mask image to mark the position and shape of the steering wheel, which may be a circular ring area representing the region where the steering wheel is located. At the same time, the system uses hand pose estimation algorithms to analyze the hand image and mark the key positions of the finger joints and palm. For example, the system may mark the top of the index finger, the joint of the thumb, and the center of the palm, etc., which can help the system accurately determine whether the hand is holding the steering wheel.
[0057] Step 104, detecting a single frame image according to the mask image, the key point position information, and a preset determination threshold to obtain a single frame image detection result.
[0058] In embodiments of the present application, in this step, the system combines the mask image of the steering wheel image and the key point position information of the hand image, and uses a preset determination threshold to determine whether the driver is holding the steering wheel or the steering wheel is out of hand. The threshold refers to a certain condition or standard set by the system, which will trigger a specific response when the data exceeds or falls below this value. Here, the determination threshold may include the distance between the hand and the steering wheel, the contact area, the grip strength, etc.
[0059] For example, assuming that the system sets a determination threshold that requires some key points of the driver's hand to be located within the mask area of the steering wheel, otherwise it is determined to be out of hand. If the system finds that the key points of the driver's hand are outside the mask in a single frame image, or the contact area of the hand is less than the preset determination threshold, the system will preliminarily determine that it may be an out-of-hand behavior, and record the single frame image detection result of the single frame image as out of hand.
[0060] Step 105, determining whether there is a steering wheel out-of-hand situation based on the single frame image detection result.
[0061] In embodiments of the present application, in a single frame image, if the single frame image detection result shows that the contact between the hand and the steering wheel does not meet the preset conditions (i.e., the hand does not contact the steering wheel or the contact is insufficient), the system will directly determine that the driver has an out-of-hand behavior, i.e., there is a steering wheel out-of-hand situation. The system will immediately issue a prompt or warning to remind the driver to hold the steering wheel tightly.
[0062] For example, if the hand key points of the driver are far away from the steering wheel in a single frame image, and there are not enough finger key points in the mask image, the system will directly determine that there is a steering wheel hand-off situation according to the detection result of the single frame image. Then, the system will issue a warning sound or display a warning message on the instrument panel, prompting the driver that the hand is not holding the steering wheel, and requiring him to hold the steering wheel tightly to ensure safety.
[0063] Step 106, extracting the feature vector of the multi-frame image sequence of the video frame data, when it is determined that there is no steering wheel hand-off situation according to the single frame image detection result.
[0064] In an embodiment of the present application, if it is determined that no hand-off behavior occurs, i.e., there is no steering wheel hand-off situation, according to the single frame image detection result, the system will further analyze the multi-frame image sequence. The multi-frame image sequence refers to the image data extracted from multiple consecutive video frames, and the system will analyze the long-time interaction between the hand and the steering wheel through these multi-frame image sequences. The system extracts a feature vector from these multi-frame image sequences, which is a set of numbers representing the features in the multi-frame image sequence, such as the motion trajectory and position change of the hand.
[0065] For example, if the single frame image detection result indicates that there is no steering wheel hand-off situation, the system will extract the hand position and motion trajectory in the next few seconds (e.g., 10 frames of images) as the feature vector. By analyzing whether the hand is in contact with the steering wheel or has displacement during this period of time, the system can more comprehensively determine whether the driver has let go of the steering wheel.
[0066] Step 107, comprehensively detecting whether there is a steering wheel hand-off situation according to the feature vector of the multi-frame image sequence and the single frame image detection result.
[0067] In an embodiment of the present application, the system comprehensively analyzes the single frame image detection result and the feature vector of the multi-frame image sequence. If both the single frame image detection result and the feature vector of the multi-frame image sequence indicate that the driver's hand is not in contact with the steering wheel, the system will finally determine that there is a steering wheel hand-off situation. Through comprehensive judgment, the system can reduce false positives or false negatives, ensuring the accuracy of the detection result.
[0068] For example, if the single frame image detection result shows that the driver's hand temporarily leaves the steering wheel, but the feature vector of the multi-frame image sequence indicates that the hand has recovered contact within a short period of time. Therefore, the system will comprehensively determine that there is no steering wheel hand-off situation, avoiding misjudgment as hand-off due to a short action. Conversely, if the feature vector of the multi-frame image sequence indicates that the hand leaves the steering wheel, the system will finally determine that there is a steering wheel hand-off situation, and trigger the corresponding warning.
[0069] The detection scheme for steering wheel hand-off provided by the embodiment of the present application acquires video frame data of a steering wheel region. Then, a steering wheel image and a hand image are located from a single frame image of the video frame data. Furthermore, a mask image of the steering wheel image and key point position information of the hand image are acquired. Next, the single frame image is detected according to the mask image, the key point position information, and a preset determination threshold to obtain a single frame image detection result, and whether there is a steering wheel hand-off condition is determined based on the single frame image detection result. In a case where it is determined according to the single frame image detection result that there is no steering wheel hand-off condition, feature vectors of a multi-frame image sequence of the video frame data are extracted, and then whether there is a steering wheel hand-off condition is comprehensively detected according to the feature vectors of the multi-frame image sequence and the single frame image detection result.
[0070] The embodiment of the present application not only relies on the detection result of the single frame image, but also introduces the feature vectors of the multi-frame image sequence for comprehensive detection. This detection method combining single frame image and multi-frame image makes up for the accidental errors that may be caused by single frame detection. Compared with the method relying on only a single sensor or simple image recognition, the present application is more robust and can effectively improve the accuracy of detection. By acquiring the mask image of the steering wheel image and the key point position information of the hand image, the hand state of the driver can be more accurately recognized. The mask image is an image processing technology that highlights the steering wheel region, and in combination with the hand key point position information, it can be determined whether the driver's hand is off the steering wheel. This is more accurate than relying on sensor detection of contact force, torque change, and other physical signals, and reduces the detection errors caused by the driver's hand posture or gloves and other factors. In the detection process, a preset determination threshold is introduced, that is, a reasonable index is set to judge the hand-off condition. The determination mechanism combining the single frame image detection result and the feature vectors of the multi-frame image sequence makes the hand-off detection more intelligent. By extracting the feature vectors of the multi-frame image sequence, the operation state of the driver in a period of time can be continuously monitored, rather than relying on only a single frame image at a moment. This continuous image sequence analysis can capture the dynamic characteristics of the hand-off behavior and avoid false judgments caused by detection errors at a moment.
[0071] In summary, the embodiment of the present application overcomes the shortcomings of the traditional sensor scheme by introducing multi-frame image sequence analysis, mask image combined with hand key points, intelligent threshold judgment, and other technical means, and improves the accuracy, stability, and safety of the overall detection.
[0072] In an exemplary embodiment of the present application, one implementation of detecting a single frame image to obtain a single frame image detection result according to the mask map, key point position information and a preset determination threshold is to calculate the proportion information of the key points of the driver's hand in the mask map in the single frame image according to the mask map and the key point position information; and compare the proportion information with the determination threshold to obtain the single frame image detection result. By calculating the proportion information of the key points of the driver's hand in the mask map of the steering wheel, it is determined whether the driver has maintained control of the steering wheel. If the proportion is lower than the set determination threshold, the system will determine that the driver may have released the steering wheel, thereby further triggering a warning or other safety measures.
[0073] In actual application, the mask map of the steering wheel needs to be determined. The mask map is a binary image used to mark a specific region in an image, where the mask region corresponds to the position of the steering wheel, and other parts are masked. In the steering wheel release detection system, the mask map is usually generated by a computer vision algorithm that can identify the steering wheel in the image and mark its contour to form a mask map. The generation of the mask map depends on the image segmentation technology, which can accurately distinguish the steering wheel from the background and other irrelevant image elements.
[0074] At the same time, the system will obtain the key point position information of the driver's hand through a hand pose estimation algorithm. These key points usually include the tips of the fingers, joints and the center position of the palm. The pose estimation algorithm is to detect the human skeleton in the image and locate the key point position information of the hand through a deep learning model such as a convolutional neural network. For each frame of image, the system will extract these key points of the hand in real time, so as to understand the specific position and pose of the hand in space.
[0075] After obtaining the mask map and the key point information of the hand, the system will calculate the proportion of the key points of the fingers in the mask map. This proportion represents the distribution of the key points of the hand in the steering wheel region, and the specific steps are as follows:
[0076] The system will check the coordinates of each finger key point to determine whether these coordinate points are located within the mask map of the steering wheel. Each pixel point in the mask map has a marked value (usually 0 or 1), and the region marked with 1 is the steering wheel region. Therefore, the system only needs to check the pixel point value corresponding to each hand key point, and if the value is 1, it means that the hand key point is located in the steering wheel region.
[0077] Then, the system calculates the number of key points located in the mask map among all finger key points, and calculates the proportion with the total number of hand key points. This proportion is called proportion information.
[0078] For example, if a driver's hand has 5 key points (one for each finger), and 3 of them are inside the mask, while the other 2 are outside, the occupancy information for that image is 60% (3 / 5).
[0079] In the detection system, the decision threshold is a pre-set standard value that the system will use to compare with the occupancy information to determine whether the driver is holding the steering wheel. The threshold is generally set by the developer based on the actual driving scene requirements and safety requirements. For example, a reasonable threshold might be 80%, that is, when the proportion of hand key points located in the steering wheel area is less than 80%, the system will consider that the driver does not hold the steering wheel completely, and a hand-off behavior may occur.
[0080] The threshold can be adjusted according to the specific application scenario. For example, on the highway, the threshold can be set higher (such as 90%) to ensure that the driver holds the steering wheel tightly with both hands. While driving in the city at low speed, the threshold can be relatively low (such as 70%) to allow the driver to perform some slight hand movements while maintaining control.
[0081] The system compares the calculated occupancy information with the decision threshold to obtain the detection result of a single frame image. If the occupancy information is higher than or equal to the decision threshold, the system will consider that the driver's hand in the frame image holds the steering wheel normally; if the occupancy information is lower than the decision threshold, the system will consider that the driver may have a hand-off behavior.
[0082] For a specific example: in a certain frame of image, assuming the calculated occupancy information is 75%, and the decision threshold is set to 80%. Since 75% is lower than 80%, the system will judge that the driver's hand in the frame image does not hold the steering wheel completely, and a hand-off behavior may occur. Conversely, if the occupancy information is 85%, the system will judge that there is no hand-off, and the driver is still in control.
[0083] Compared with the traditional detection method relying only on the steering wheel sensor, the hand-off detection using image data can more intuitively and accurately reflect the hand movement of the driver. By analyzing the specific position of the fingers, the system can more accurately determine whether the driver actually hands off, rather than just relying on the contact pressure of the hand or the torque change of the steering wheel. The traditional steering wheel sensor is easily disturbed by external interference, for example, some drivers cheat the system by placing heavy objects on the steering wheel. The image analysis-based method can effectively prevent such cheating behavior by directly capturing the hand behavior through vision. The determination threshold can also be dynamically adjusted according to different driving scenarios and the personal habits of the driver, providing a more flexible detection scheme. For example, in different driving modes (high speed, low speed, curve, etc.), the system can adaptively adjust the hand-off determination standard to ensure that the driver can maintain proper control in each case.
[0084] In an exemplary embodiment of the present application, one implementation of determining whether there is a steering wheel hand-off situation based on single-frame image detection results is that when the determination threshold is the maximum ratio of the finger key points outside the mask graph, if the single-frame image detection result indicates that the proportion information is greater than the determination threshold, it is determined that there is a steering wheel hand-off situation; if the single-frame image detection result indicates that the proportion information is less than or equal to the determination threshold, it is determined that there is no steering wheel hand-off situation. In the steering wheel hand-off detection system, single-frame image detection is a key step to determine whether the driver has a hand-off behavior by detecting whether the driver's hand is located in the steering wheel area. This implementation determines whether there is a hand-off behavior based on the proportion of the finger key points outside the mask graph. This method sets a determination threshold (maximum ratio) and compares it with the proportion information in the image. If the proportion information of the hand key points exceeds the determination threshold, the system considers that the driver has handed off.
[0085] In actual application, after obtaining the mask graph and the hand key points, the proportion of the finger key points outside the mask graph is calculated. The specific operation steps are as follows:
[0086] The system checks the coordinates of each hand key point one by one to determine whether they are within the steering wheel area marked by the mask graph. This is done by checking whether the pixel value corresponding to the key point is the marked value (usually 1) in the mask graph. If the pixel value is 1, it means that the key point is located within the steering wheel area; if it is 0, it means that the key point is located outside the steering wheel.
[0087] The system calculates the number of finger key points located outside the mask graph and compares it with the total number of finger key points to obtain a proportion information. This proportion information can be used to represent the degree of the driver's hand leaving the steering wheel.
[0088] For example, assume the system detects 5 key points of the driver's hand, 3 of which are outside the steering wheel mask. The system will calculate the ratio information of this image as 60% (3 / 5).
[0089] The decision threshold is defined as the maximum ratio of finger key points outside the mask image. The system sets a standard, if the ratio of finger key points in a frame image exceeds this ratio, it is considered that the driver may have let go of the steering wheel.
[0090] In single-frame image detection, the system compares the ratio information of finger key points outside the steering wheel mask with the preset decision threshold. If the ratio information is greater than or equal to the decision threshold, the system will determine that the driver has let go of the steering wheel; otherwise, the system will determine that there is no steering wheel let go.
[0091] Specifically, assume that the ratio of finger key points in a frame image is 75%, and the system's decision threshold is set to 70%. Because 75% is greater than 70%, the system will determine that the driver has let go of the steering wheel in this image frame. If the ratio is 65%, the system will consider that the driver has not let go of the steering wheel and continue to monitor the contact state of the driver's hand and the steering wheel.
[0092] Assume that a driver is driving on a highway, and the system monitors the driver's hand state in real time through a camera. At a certain time, the driver's hand temporarily leaves the steering wheel, possibly to adjust the in-vehicle equipment. The system detects that 4 of the 5 key points of the driver's hand are outside the mask image (ratio 80%), and the system's decision threshold is 70%. Therefore, the system determines that the driver may have let go of the steering wheel and issues a warning signal to remind the driver to hold the steering wheel tightly.
[0093] Conversely, if the driver performs a similar action, only 2 of the hand key points are outside the mask image (ratio 40%), the system will consider that the driver's hand is still in a controllable state and will not issue any warnings.
[0094] The embodiment of the present application can more accurately reflect the actual behavior of the driver's hand through image analysis. It not only relies on physical contact, but also analyzes the state of the hand through spatial position. The system has strong tolerance for the driver's slight hand movement. When the driver adjusts the gesture, the system can judge whether the driver has really let go of the steering wheel through the overall ratio of key points, reducing false positives. The decision threshold can be adjusted according to different driving environments or driver habits. For example, a higher threshold can be set for high-speed driving, while the threshold can be slightly lower for low-speed driving, enhancing the adaptability of the system.
[0095] In an exemplary embodiment of the present application, one implementation of detecting whether the steering wheel is off-hand based on the feature vector of the multi-frame image sequence and the single-frame image detection result is to combine the feature vector and the normalized data of the key point position information into a feature vector matrix; classify the feature vector matrix to obtain the classification result of the multi-frame image sequence; in the case that the classification result of the multi-frame image sequence is greater than a preset classification threshold, determine whether there is a single-frame image detection result meeting the preset steering wheel off-hand condition in the multi-frame image sequence; in the case that there is a single-frame image detection result meeting the preset steering wheel off-hand condition in the multi-frame image sequence, determine that there is a steering wheel off-hand condition.
[0096] Normalization is a data processing technique used to adjust data of different dimensions to the same range for more accurate comparison. In practical application, the system normalizes the feature vector and the hand key point position information and combines them into a feature vector matrix. The matrix includes multiple dimensions of features such as hand trajectory and position change.
[0097] For example, suppose the coordinate values of the hand key points are large, and the motion trajectory change in the feature vector is small. By normalizing them to the same range (e.g. between 0 and 1), the differences between different features can be eliminated. Then, the system integrates these normalized data into a feature matrix for further analysis. For example, each column in the matrix represents the hand state in a frame image, and each row represents different features such as hand position and motion direction.
[0098] Next, the feature vector matrix is classified to obtain the classification result of the multi-frame image sequence. The system analyzes the feature vector matrix through a classification algorithm. Common classification algorithms include Support Vector Machine (SVM) and neural network, which can classify the data in the feature matrix to determine whether the driver is off-hand. The classification result usually includes two categories: off-hand or not off-hand.
[0099] For example, suppose the system uses a neural network to classify the feature matrix. The trained model analyzes the state of the hand in the multi-frame images and determines that the driver's hand has been away from the steering wheel for multiple frames. At this time, the classification algorithm marks the result as off-hand. If the hand is in contact with the steering wheel most of the time, the classification result is not off-hand.
[0100] To improve the accuracy of off-hand detection, the system not only relies on single-frame image detection results, but also analyzes multi-frame image sequences over a period of time. A multi-frame image sequence refers to a series of images of the driver's hand and the steering wheel continuously captured by the system over a period of time.
[0101] In each single frame image, the system extracts the key point information of the hand and the mask map of the steering wheel, and combines the single frame detection algorithm mentioned earlier to determine whether the hand has left the steering wheel. However, since the driver's hand may have a short movement during normal driving, such as adjusting the seat, turning on the air conditioner, etc., relying on single frame detection can easily cause false positives. Therefore, the system also makes a comprehensive judgment based on the image frames within a certain period of time, which is called multi-frame image sequence analysis.
[0102] Image classification technology classifies the feature vectors of each frame image through a trained deep learning model. Feature vectors are numerical representations extracted from images to describe images, usually including hand key points, hand contact area, finger position, etc. The model generates a classification result based on these feature vectors, indicating whether the frame image may have a hand-off behavior.
[0103] The system classifies multi-frame image sequences and combines the classification results of each frame image to generate a multi-frame image sequence classification result. This classification result can be understood as a general judgment of multiple image frame detection, reflecting whether the driver has a hand-off behavior during this period.
[0104] The classification threshold is a key parameter for judging multi-frame image sequences. The classification threshold can be a preset value that defines which frame image classification results indicate that the driver may have a hand-off in a multi-frame image sequence. For example, the system may set a classification threshold, if more than 50% of the image classification results in a multi-frame image sequence show a hand-off state, then the multi-frame image sequence will be marked as dangerous.
[0105] By setting the classification threshold, the system can ensure that false alarms are not easily triggered by single frame image misjudgments. Especially in the case of the driver's hand leaving the steering wheel for a short time, such as adjusting the air conditioner or controlling the wiper, the system can determine whether the driver has truly lost control of the steering wheel through the continuity of multi-frame images.
[0106] For example, assume the system's classification threshold is set to 60%. In a certain period of time, 10 frames of images are collected, and if 6 of them have image classification results showing that the driver has a hand-off behavior, then the classification result of the multi-frame image sequence will be considered to exceed the threshold, indicating a high probability of a hand-off situation.
[0107] Once the classification result of the multi-frame image sequence exceeds the preset classification threshold, the system will further check the single frame image detection results in these image sequences to ensure that these detection results are consistent with the actual situation.
[0108] In this process, the system needs to determine whether a single frame or multiple frames of image detection results meet the steering wheel hand-off condition. For example, some frames of images can show that the driver's hands are completely off the steering wheel, or even that multiple finger key points are outside the mask graph. This situation will trigger a single-frame image hand-off alarm.
[0109] If at least one frame of image detection results in a multi-frame image sequence indicates that the driver does indeed have a hand-off behavior, the system will determine that the driver has a hand-off situation by combining the overall judgment of the classification results of the multi-frame image sequence.
[0110] By combining the classification results of the multi-frame image sequence and the single-frame image detection results, the system can more accurately determine whether the driver has a steering wheel hand-off behavior.
[0111] For example, suppose the system captures 30 frames of images in 5 seconds, and 20 frames of images have classification results showing that the driver's hands are off the steering wheel. At the same time, single-frame detection results show that several frames of images have finger key points that have completely exceeded the steering wheel area. Combining these two pieces of evidence, the system will determine that the driver has a hand-off behavior and timely issue a warning signal.
[0112] The embodiments of the present application significantly improve the accuracy and robustness of hand-off detection by combining multi-frame image classification results and single-frame detection results. This avoids false positives that can occur with single image detection, and also reduces the situation of unnecessary warnings triggered by the driver's brief hand movement. The driver's hands cannot remain completely still during daily driving, and this multi-frame detection method can tolerate the driver's hand movements and will not misjudge as hand-off due to some irrelevant actions.
[0113] In an exemplary embodiment of the present application, after classifying the feature vector matrix to obtain the classification results of the multi-frame image sequence, one implementation is to determine that there is no steering wheel hand-off situation when the classification results of the multi-frame image sequence are less than or equal to a preset classification threshold. If the classification results of the multi-frame image sequence are less than or equal to the preset classification threshold, the system determines that the driver does not have a hand-off behavior. This means that during this period, the system has not found that the driver's hands have left the steering wheel. This determination effectively reduces the possibility of false positives. For example, when the driver may briefly move his hands to other positions for operation, the system will not trigger an alarm, thereby reducing the driver's sense of tension and discomfort.
[0114] The embodiment of the present application sets a reasonable classification threshold, so that the system can avoid misjudging the state of the driver due to accidental factors. When the driver's hands are only temporarily away from the steering wheel, the system can still confirm that the driver has not truly let go through analysis of multiple frames of images. This feature is particularly important in actual driving, as it can reduce the driver's distrust of the system and allow them to remain relaxed during driving.
[0115] In an exemplary embodiment of the present application, after classifying the feature vector matrix to obtain the classification result of the multi-frame image sequence, one implementation is that, when the classification result of the multi-frame image sequence is greater than the preset classification threshold, and there is no single-frame image detection result in the multi-frame image sequence that meets the preset steering wheel hand-off condition, it is determined that there is no steering wheel hand-off situation.
[0116] In the specific implementation process, it is assumed that the system sets the multi-frame image sequence to be 20 consecutive frames, and the classification threshold is set to 0.7 (representing a 70% probability of considering that there is a hand-off behavior). In a certain detection, the system analyzes the 20 frames of images and obtains a classification result of 0.8, which exceeds the preset threshold of 0.7. According to the traditional method, this may be directly determined as a hand-off behavior. However, the present embodiment further requires further checking each frame in the 20 frames.
[0117] Further analysis shows that in the 20 frames, none of the single-frame image detection results meets the steering wheel hand-off condition. This may be because the driver has slight hand movements (such as adjusting the grip or immediately holding back after a short release), resulting in a high multi-frame classification result, but the driver is actually always in control of the vehicle.
[0118] In this case, although the classification result of the multi-frame image sequence exceeds the threshold, since there is no single-frame image showing a clear hand-off state, the system ultimately determines that there is no steering wheel hand-off situation. This method effectively avoids false positives for temporary and safe hand movements.
[0119] The embodiment of the present application greatly reduces the possibility of false positives by considering both multi-frame sequences and single-frame results. This double verification mechanism can better distinguish between true hand-off behavior and normal driving actions. It has stronger tolerance for temporary hand movements (such as quickly adjusting the steering wheel position) and does not trigger an alarm due to transient hand-off, thereby improving the stability and reliability of the system. By reducing false positives, unnecessary warnings can be avoided, improving the driver's trust in the system and the comfort of use.
[0120] In an exemplary embodiment of the present application, one implementation of locating the steering wheel image and hand images from a single frame of video frame data is to input the single frame image into YOLOX (YOLOX is an acronym for “You Only Look Once X”. It is an improved version of the YOLO (You Only Look Once) series of object detection algorithms), and output the steering wheel image and hand images.
[0121] In practical application, the system first extracts a single frame image from the video stream. This single frame image contains the entire cockpit scene, including the steering wheel area and the driver's hands. Next, this single frame image is input into the pre-trained YOLOX model. The YOLOX model quickly locates and identifies the positions of the steering wheel and hands in the image and outputs images of these two key areas.
[0122] For example, there is a cockpit video stream with 1080p resolution (1920x1080 pixels). The system extracts 30 frames of images per second from this video stream for analysis. For one of the frames, the YOLOX model completes processing in about 20 milliseconds and outputs results similar to the following:
[0123] 1. Steering wheel image: position (x=800, y=500), size 200x200 pixels
[0124] 2. Left hand image: position (x=750, y=600), size 100x150 pixels
[0125] 3. Right hand image: position (x=950, y=600), size 100x150 pixels
[0126] These output results not only contain the position information of the objects, but also include the confidence score, which represents the model's certainty of the detection result.
[0127] In the embodiment of the present application, YOLOX is a single-stage object detection algorithm that can complete object localization and classification simultaneously in one forward propagation. This means it can complete the analysis of the entire image in a very short time (usually milliseconds), which is very suitable for the needs of real-time monitoring systems. Compared with the earlier versions of YOLO, YOLOX has significant improvements in small object detection and bounding box localization. This is particularly important for accurately capturing the position of the hands, as at certain angles, the hands may only occupy a small part of the image. YOLOX has strong feature extraction capabilities and can maintain good detection performance under complex conditions such as different light conditions and hand posture changes. This ensures the stability of the system in various driving environments.
[0128] In an exemplary embodiment of the present application, one implementation of obtaining the mask image of the steering wheel image and the key point position information of the hand image is to perform edge detection processing and ellipse matching processing on the steering wheel image to obtain the mask image; and the hand image is rotated at the corresponding corner points and input to the BlazeHand model (a lightweight hand posture estimation model) in a forward hand posture to output the key point position information.
[0129] In actual application, the process of generating the steering wheel mask image is:
[0130] 1. Edge detection processing: the system first performs edge detection on the steering wheel image. Common edge detection algorithms include the Canny edge detector or the Sobel operator. This step can highlight the outline of the steering wheel.
[0131] 2. Ellipse matching processing: since most steering wheels are elliptical, the system uses an ellipse matching algorithm (such as Hough transform) to fit the detected edges to obtain the best matching elliptical shape.
[0132] 3. Generate mask image: based on the matched ellipse, the system creates a binary image where the steering wheel area is 1 (white) and other areas are 0 (black).
[0133] The hand key point detection process is:
[0134] 1. Hand image rotation: the system first rotates the hand image to present a forward posture. This step helps improve the accuracy of subsequent detection.
[0135] 2. Input BlazeHand model: the rotated hand image is input into the pre-trained BlazeHand model.
[0136] 3. Output key point position information: the BlazeHand model outputs the key point positions of the hand, usually including 21 key points representing the main joint positions of the palm and fingers.
[0137] For example, assume there is a 1280x720 pixel steering wheel area image. First, perform Canny edge detection on this image to obtain a binary image highlighting the outline. Then, use Hough transform for ellipse matching to find the best matching ellipse parameters (center point coordinates, major axis, minor axis, and rotation angle). Based on these parameters, generate a 1280x720 binary mask image where the steering wheel area is white (1) and other areas are black (0).
[0138] For hand images, it is assumed that the hand in the original image is at a 45-degree angle. The system first rotates the image by -45 degrees to present the hand in a positive pose. Then, this rotated image of 224x224 pixels (the standard input size for BlazeHand) is input to the BlazeHand model. The model outputs the coordinates of 21 key points.
[0139] The embodiments of the present application can accurately locate the steering wheel through edge detection and ellipse matching, and maintain good results even in complex backgrounds or lighting conditions. The ellipse matching algorithm can adapt to steering wheels of different shapes and sizes, making the system applicable to various vehicle models. The BlazeHand model is designed for mobile and embedded devices and can complete hand pose estimation within milliseconds, meeting the needs of real-time monitoring. Compared with simple hand detection, key point information provides a detailed description of the hand pose, providing rich data for subsequent grip analysis. By rotating the hand image for preprocessing, the system can handle various hand poses, improving the robustness of detection.
[0140] In an exemplary embodiment of the present application, one implementation of classifying the feature vector matrix to obtain the classification result of the multi-frame image sequence is that a time series classification model based on a Gated Recurrent Unit (GRU) is used to classify the feature vector matrix to obtain the classification result of the multi-frame image sequence. The GRU is a variant of a Recurrent Neural Network (RNN) specially designed for processing sequence data. Through its special gating mechanism, the GRU can effectively capture long-term dependencies while avoiding the gradient vanishing problem in traditional RNNs.
[0141] In this implementation, the system first constructs a feature vector matrix that contains the feature information of each frame in the multi-frame image sequence. Then, the feature vector matrix is input into the GRU-based time series classification model for processing. The model finally outputs the classification result of the multi-frame image sequence, i.e., whether there is a steering wheel hand-off behavior in the time sequence.
[0142] For example, assume there is a 30-frame (about 1 second) image sequence. For each frame, the following features have been extracted:
[0143] 1. Hand key point positions (21 points, each point has x and y coordinates)
[0144] 2. Related features of the steering wheel mask image (such as steering wheel center coordinates, semi-major axis, semi-minor axis, etc.)
[0145] 3. Other possible related features (such as the overlap area ratio of the hand and the steering wheel)
[0146] Thus, each frame can have a 50-dimensional feature vector. For a 30-frame sequence, a 30x50 feature matrix is obtained.
[0147] Next, this 30x50 matrix is input into the GRU model. Assume the GRU model has the following structure:
[0148] 1. Input layer: accepts the 30x50 feature matrix
[0149] 2. GRU layer 1: 64 hidden units
[0150] 3. GRU layer 2: 32 hidden units
[0151] 4. Fully connected layer: 16 neurons
[0152] 5. Output layer: 2 neurons (representing the probabilities of the two states of hand-off and non-hand-off)
[0153] The model processes this 30-frame sequence and finally outputs a binary classification result, such as [0.05, 0.95], indicating a 95% probability of considering that there is a hand-off behavior in this sequence.
[0154] In the embodiments of the present application, GRU can effectively capture the time-dependent relationship in sequence data. This is crucial for analyzing the hand movements of drivers, as some actions (such as temporarily releasing the steering wheel to adjust the grip) may appear to be hand-off in a single frame, but are actually safe in the entire sequence. Compared with traditional RNN, GRU is better at capturing long-term dependencies. This allows the model to consider driving behavior patterns over a longer time range, improving the accuracy of classification. The gating mechanism of GRU allows it to selectively remember important information and forget irrelevant information. This improves the model's resistance to noise and short-term interference, making the classification result more stable.
[0155] Based on the above related description of an embodiment of a steering wheel hand-off detection method, a multi-model joint inter-frame verification multi-scenario steering wheel hand-off detection scheme (referred to as a detection scheme) is introduced below. The detection scheme can be applied to a driver detection system. The detection scheme can include the following two parts: 1. Multi-model joint detection and judgment; 2. Inter-frame verification.
[0156] In the detection scheme, the camera deployment position is on the A-pillar. Due to different vehicle models, the A-pillar angle and height are different, and the selection of the camera is based on the field of view (FOV) that can observe the complete steering wheel. Specifically, it can be an infrared camera (IR). The collected image is as shown in Figure 2 .
[0157] The learning of visual tasks based on deep learning models cannot be separated from image data and corresponding labels. The collection scheme adopted by the present detection scheme is as follows:
[0158] 1) Collection personnel: the number of drivers is specified to be 20 or more, the height distribution is 160-190 cm, and the normal distribution is generally presented, meeting the diversity requirements.
[0159] 2) Collection equipment: wide-angle IR cameras are located in the A-pillar, and the field of view range can present the facial features and complete steering wheel of the driver as a whole.
[0160] 3) Collection data format: continuous frame short video, convenient for frame extraction and subsequent multi-frame verification.
[0161] 4) Collection process:
[0162] a. After the personnel get on the car, adjust the seat to the appropriate driving position, hold the steering wheel and start the collection preparation.
[0163] b. Drive normally and turn left for one round, repeat 5 times;
[0164] c. Drive normally and turn right for one round, repeat 5 times;
[0165] d. Drive with one hand on the left side for one round, repeat 5 times;
[0166] e. Drive with one hand on the right side for one round, repeat 5 times;
[0167] f. Hold the steering wheel with one hand and make a phone call with the other hand, repeat 10 times or exchange hands.
[0168] g. Move both hands away from the steering wheel and make various gestures at different heights, repeat 10 times.
[0169] h. Move both hands away from the steering wheel, make a phone call with one hand, and make specified actions with the other hand, repeat 20 times, exchange left and right.
[0170] i. Move the steering wheel position, adjust the seat to the appropriate position again, and repeat b to h once.
[0171] Since the current data collection is performed according to specified actions, video labeling can be directly completed by different actions. Hand and steering wheel need to rely on labeling tools to complete the adjustment and calibration of corresponding labeling boxes based on pre-labeling.
[0172] The data is divided into image data and corresponding target detection labels, continuous frame short video data and corresponding video labels. The former participates in the training of the target detection model, and the trained model has the ability to recognize the positions of the image of the opponent and the steering wheel in the current picture. The latter data participates in the training of the behavior recognition model, and the trained model has the ability to judge the behavior of the current continuous frame, that is, whether it is driving with the steering wheel off.
[0173] Referring to Figure 3 , an algorithm flow diagram of a multi-model joint inter-frame verification multi-scene steering wheel off detection scheme is shown.
[0174] The overall flow of single frame detection is as follows:
[0175] 1) Real-time single frame pictures are obtained on the camera side, and the IR camera real-time pictures are preprocessed to meet the format requirements of the model input.
[0176] 2) 2D target detection is performed by the trained 2D target detection model (which can effectively recognize the steering wheel and driver's hand features under the current camera angle), accurate positioning of the steering wheel and driver's hands in the picture is performed, and the positions are intercepted to obtain the steering wheel position image and the driver's hand position image.
[0177] 3) The steering wheel position image is subjected to edge detection in the classic image algorithm and the current picture is captured by elliptical matching to obtain the accurate mask image of the steering wheel, i.e. the contour map.
[0178] 4) The hand position image is rotated according to the corresponding corner points, and the 21 key point position information under the current hand posture is obtained by inputting the positive hand posture to the key point detection model for hand key point detection, as shown in Figure 4a , and normalization operation is performed for the next step of information fusion, as shown in Figure 4b .
[0179] 5) Single frame image off-hand judgment is performed by the current hand key point position and the current steering wheel position in the picture: the percentage of the finger key points in the steering wheel mask image is used as the judgment threshold, and if the 12 key points of the finger part meet the condition that more than 7 key points are away from the mask area of the steering wheel, it is directly judged as off-hand and off-hand reminder is performed; otherwise, the joint judgment process is entered, and the error is reduced by joint judgment calculation.
[0180] The overall flow of multi-frame detection is as follows:
[0181] 1) According to the running efficiency of the real vehicle algorithm, 10 frames are temporarily adopted as a queue, and every two frames are separated, that is, 20 frames are accumulated every time for sliding window reading, and the model input is 10 frames of image sequence every time, and the continuous multi-frame feature vector to be detected is extracted through the Backbone model.
[0182] 2) The 10-frame feature vector is combined, and the hand key point normalized data is fused into an information vector to form a 10-dimensional feature vector matrix.
[0183] 3) The feature vector matrix is input into the continuous frame detection model (such as the Long Short-Term Memory (LSTM) model) for multi-frame judgment, and the 10-dimensional feature vector matrix is classified to obtain whether the hand is off. Note that the current hand-off result is the result of feedback after synthesizing the previous 10 frames of images, and the result is locked to the last frame.
[0184] Joint judgment process:
[0185] 1) If the single-frame image detection result of the current frame directly determines that the hand is off, it is determined that the current frame is the steering wheel hand-off, and the alarm information is reported.
[0186] 2) If the single-frame image detection result of the current frame is not off, the multi-frame image sequence classification result is continuously counted and saved in the queue.
[0187] 3) The multi-frame image sequence classification result is jointly determined with the single-frame image detection result, if the multi-frame image sequence classification result is more than 50% off, and there is an off result in the single-frame image detection result, it is determined that the hand is off, otherwise it is determined that the hand is not off.
[0188] Three deep learning models are used in the overall process:
[0189] 1) Hand and steering wheel target detection model:
[0190] YOLOX is selected as the main target detection model. YOLOX performs well in target detection, especially in real-time application scenarios. In order to improve the detection efficiency and frame rate of the model, YOLOX is optimized. First, the input data is converted from multiple channels to a single channel to reduce the complexity of data processing. In addition, in order to improve the detection speed, the Backbone of YOLOX is replaced by MobileNetV2. The lightweight structure and efficient convolution operation of MobileNetV2 enable it to maintain high detection speed and accuracy in limited computing resources. The model will be trained using specific scene data collected at present to better adapt to the target characteristics in actual application.
[0191] 2) Hand keypoint detection model:
[0192] The BlazeHand model is used for hand keypoint detection. BlazeHand is a model optimized for hand detection, which can accurately identify and locate multiple key points of the hand. The model input is single-channel hand images, which are preprocessed to adapt to the input requirements of the BlazeHand model.
[0193] The BlazeHand model provides detailed information about hand gestures and actions, which provides more accurate finger positions and states for subsequent steering wheel hand-off analysis.
[0194] The BlazeHand model uses keypoint regression to predict the position of hand key points. For each hand key point k (such as finger joints), the model outputs a predicted position which can be represented as:
[0195]
[0196] where x is the input image, f(x) is the feature map extracted by the convolutional neural network (CNN), and MLP represents the multi-layer perceptron used to regress the specific coordinates of each key point.
[0197] The loss function is mean squared error (MSE):
[0198]
[0199] where N is the number of key points, is the predicted position of the i-th key point, p i is the corresponding true position.
[0200] The key point positions are then standardized. The standardization formula is:
[0201]
[0202] where is the standardized key point position, μ and σ are the mean and standard deviation of the position data, respectively. The predicted results are restored to the original scale by the inverse standardization formula:
[0203]
[0204] 3) Continuous frame hand-off scene classification model:
[0205] A GRU-based time series classification model is used to process continuous frame data, and the overall structure of the GRU model is as shown in Figure 5
[0206] The GRU model is good at processing time series data, and it can analyze continuous frame data by capturing dynamic changes over time.
[0207] The core is the control of its gate, update gate
[0208] z t = σ(W z · [h t"1 , x t ] + b z )
[0209] Where σ is the sigmoid activation function, W z is the weight matrix of the update gate, b z is the bias term, h t-1 is the hidden state at the previous moment, and x t is the current input. The update gate z t controls the degree of preservation of the hidden state h t-1 at the previous moment in the current moment t. It determines the weighted proportion of the hidden state h t at the current moment and the hidden state h t-1 at the previous moment, controlling the transmission and forgetting of information. The weight matrix W z and the bias term b z generate a value between 0 and 1 through the σ function, representing the degree of preservation.
[0210] Reset gate
[0211] r t = σ(W r · [h t"1 , x t ] + b r )
[0212] Where W r is the weight matrix of the reset gate, and b r is the bias term. The reset gate r t determines the influence of the hidden state h t-1 at the previous moment when calculating the candidate hidden state. It controls the combination of the current input x t and the hidden state h t-1 at the previous moment, affecting how to "reset" the previous memory information.
[0213] The final hidden state formula is as follows:
[0214]
[0215] The final hidden state h t is the hidden state at the current moment, which is the hidden state h t-1and a weighted combination of the candidate hidden state. The update gate z t The proportions of the previous hidden state and the candidate hidden state in the current hidden state are determined, so as to balance between long-term dependence and short-term change.
[0216] In order to enrich the information of a single frame, the detection feature vector of the image data and the hand key point information are normalized and combined on the input side. This combination not only enhances the expression ability of the single frame data, but also helps the GRU model to better understand the changes of the hand and the steering wheel in the time dimension. Finally, the GRU model will classify the hand-off scene according to the time sequence data, effectively improving the accuracy and real-time response ability of the scene classification.
[0217] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the action sequence described, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the present application.
[0218] Referring to Figure 6 , a structural block diagram of a steering wheel hand-off detection system according to an embodiment of the present application is shown. The steering wheel hand-off detection system can specifically include the following modules.
[0219] A video frame data acquisition module 61 is configured to acquire video frame data of a steering wheel region;
[0220] An image positioning module 62 is configured to position a steering wheel image and a hand image from a single frame image of the video frame data;
[0221] A mask image and key point position acquisition module 63 is configured to acquire a mask image of the steering wheel image and key point position information of the hand image;
[0222] A single frame image detection result determination module 64 is configured to detect the single frame image according to the mask image, the key point position information, and a preset determination threshold to obtain a single frame image detection result;
[0223] A hand-off condition determination module 65 is configured to determine whether there is a steering wheel hand-off condition based on the single frame image detection result;
[0224] A feature vector extraction module 66 is configured to extract a feature vector of a multi-frame image sequence of the video frame data when it is determined that there is no steering wheel hand-off condition according to the single frame image detection result;
[0225] The hand-off condition detection module 67 is configured to detect whether the steering wheel is in the hand-off condition according to the feature vectors of the plurality of image sequences and the single-frame image detection result.
[0226] In an example embodiment of the present application, the single-frame image detection result determination module 64 comprises:
[0227] The proportion information calculation module is configured to calculate the proportion information of the finger key points in the mask graph according to the mask graph and the key point position information.
[0228] The proportion threshold comparison module is configured to compare the proportion information with the determination threshold to obtain the single-frame image detection result.
[0229] In an example embodiment of the present application, the hand-off condition determination module 65 is configured to determine that the steering wheel is in the hand-off condition based on the single-frame image detection result indicating that the proportion information is greater than the determination threshold, and determine that the steering wheel is not in the hand-off condition based on the single-frame image detection result indicating that the proportion information is less than or equal to the determination threshold.
[0230] In an example embodiment of the present application, the hand-off condition detection module 67 comprises:
[0231] The feature vector merging module is configured to merge the feature vectors and the normalized data of the key point position information into a feature vector matrix.
[0232] The vector matrix classification module is configured to classify the feature vector matrix to obtain a plurality of image sequence classification results.
[0233] The single-frame result judgment module is configured to determine whether the single-frame image detection result satisfying the preset steering wheel hand-off condition exists in the plurality of image sequences based on the classification result of the plurality of image sequences being greater than a preset classification threshold.
[0234] The steering wheel hand-off determination module is configured to determine that the steering wheel is in the hand-off condition based on the single-frame image detection result satisfying the preset steering wheel hand-off condition existing in the plurality of image sequences.
[0235] In an example embodiment of the present application, the system further comprises:
[0236] The steering wheel non-hand-off determination module is configured to determine that the steering wheel is not in the hand-off condition based on the classification result of the plurality of image sequences being less than or equal to the preset classification threshold after the vector matrix classification module classifies the feature vector matrix to obtain the plurality of image sequence classification results.
[0237] In an example embodiment of the present application, the steering wheel off-hand determination module is further configured to determine that there is no steering wheel off-hand situation when the classification result of the multi-frame image sequence is greater than the preset classification threshold, and there is no single-frame image detection result in the multi-frame image sequence that satisfies the preset steering wheel off-hand condition.
[0238] In an example embodiment of the present application, the image positioning module 62 is configured to input the single-frame image into YOLOX to output the steering wheel image and the hand image.
[0239] In an example embodiment of the present application, the mask map and key point position acquisition module 63 comprises:
[0240] The mask map acquisition module is configured to perform edge detection processing and ellipse matching processing on the steering wheel image to obtain the mask map.
[0241] The key point acquisition module is configured to rotate the hand image through the corresponding corner points, input the hand image in a forward hand posture into the BlazeHand model, and output the key point position information.
[0242] In an example embodiment of the present application, the vector matrix classification module is configured to classify the feature vector matrix by using a GRU-based time series classification model to obtain the classification result of the multi-frame image sequence.
[0243] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts are described in the method embodiment.
[0244] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0245] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0246] The embodiments of the present application are described with reference to the flowchart illustrations and / or block diagrams of the methods, terminal devices (systems) and computer program products according to the embodiments of the present application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing terminal devices to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal devices, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0247] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0248] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operational steps are carried out on the computer or other programmable terminal devices to produce a computer implemented process so that the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0249] Although preferred embodiments of the present application have been described, those skilled in the art will be able to make additional modifications and variations to these embodiments without departing from the scope of the present application. Accordingly, the appended claims are intended to encompass all such modifications and variations as falling within the scope of the present application.
[0250] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprising", "including", or any other closure, are intended to cover the non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements does not include those elements alone but can include other elements not expressly listed or even include elements inherent in such process, method, article, or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0251] The above describes in detail the method for detecting steering wheel release and the system for detecting steering wheel release provided by the present application. The principles and implementation modes of the present application are described by applying specific examples. The above description of the examples is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, the specific implementation modes and application ranges will be changed according to the idea of the present application. In summary, the content of the present description should not be understood as a limitation of the present application.
Claims
1. A method for detecting steering wheel slippage, characterized in that, The method includes: Acquire video frame data for the steering wheel area; The steering wheel image and hand image are located from a single frame of the video frame data; Obtain the mask image of the steering wheel image and the key point location information of the hand image; The single-frame image is detected based on the mask image, the key point location information, and the preset judgment threshold to obtain the single-frame image detection result; The step of detecting the single-frame image based on the mask image, the key point location information, and a preset judgment threshold to obtain the single-frame image detection result includes: The proportion of finger key points in the single frame image located in the mask image is calculated based on the mask image and the key point location information. The single-frame image detection result is obtained by comparing the proportion information with the determination threshold. Based on the detection results of the single-frame image, determine whether the steering wheel has been slipped from the hands; If it is determined from the single-frame image detection result that there is no case of the steering wheel being taken off, the feature vector of the multi-frame image sequence of the video frame data is extracted; Based on the feature vectors of the multi-frame image sequence and the detection results of the single-frame image, it is determined whether the steering wheel has been slipped from the hands. The step of detecting whether a steering wheel has been released from the hands based on the feature vectors of the multi-frame image sequence and the detection results of the single-frame image includes: The feature vectors and the normalized data of the key point location information are combined into a feature vector matrix; The classification results of the multi-frame image sequence are obtained by classifying the feature vector matrix. The step of obtaining the mask image of the steering wheel image and the key point location information of the hand image includes: The mask image is obtained by performing edge detection and ellipse matching processing on the steering wheel image; The hand image is rotated at the corresponding corner points and input into the BlazeHand model with a positive hand pose, and the key point position information is output.
2. The method according to claim 1, characterized in that, The step of determining whether a steering wheel has been released from the hands based on the detection results of the single-frame image includes: If the determination threshold is the maximum ratio of finger key points located outside the mask image, and the percentage information indicated by the single-frame image detection result is greater than the determination threshold, it is determined that there is a situation where the steering wheel is taken off the hand; if the percentage information indicated by the single-frame image detection result is less than or equal to the determination threshold, it is determined that there is no situation where the steering wheel is taken off the hand.
3. The method according to claim 1, characterized in that, After classifying the feature vector matrix to obtain the classification result of the multi-frame image sequence, the method further includes: If the classification result of the multi-frame image sequence is less than or equal to a preset classification threshold, it is determined that there is no situation where the steering wheel is taken off the hands.
4. The method according to claim 1, characterized in that, After classifying the feature vector matrix to obtain the classification result of the multi-frame image sequence, the method further includes: If the classification result of the multi-frame image sequence is greater than a preset classification threshold, it is determined that the steering wheel has been released from the hands.
5. A system for detecting steering wheel slippage, characterized in that, The system includes: The video frame data acquisition module is used to acquire video frame data of the steering wheel area; An image localization module is used to locate the steering wheel image and the hand image from a single frame image of the video frame data; The mask and key point location acquisition module is used to acquire the mask of the steering wheel image and the key point location information of the hand image; The single-frame image detection result determination module is used to detect the single-frame image based on the mask image, the key point position information and the preset judgment threshold to obtain the single-frame image detection result; The single-frame image detection result determination module includes: The proportion information calculation module is used to calculate the proportion information of finger key points in the single frame image located in the mask image based on the mask image and the key point position information; The proportion threshold comparison module is used to compare the proportion information with the judgment threshold to obtain the single-frame image detection result; The hand-off situation determination module is used to determine whether there is a situation where the steering wheel has been removed from the hands based on the detection results of the single-frame image; The feature vector extraction module is used to extract the feature vectors of the multi-frame image sequence of the video frame data when it is determined from the single-frame image detection result that there is no steering wheel slippage. The hands-off detection module is used to detect whether there is a situation where the steering wheel has been removed from the hands based on the feature vectors of the multi-frame image sequence and the detection results of the single-frame image. The release detection module includes: The feature vector merging module is used to merge the feature vectors with the normalized data of the key point location information into a feature vector matrix. The vector matrix classification module is used to classify the feature vector matrix to obtain classification results for multi-frame image sequences; The mask and key point location acquisition module includes: a mask acquisition module, used to perform edge detection processing and ellipse matching processing on the steering wheel image to obtain the mask; and a key point acquisition module, used to rotate the hand image through corresponding corner points, input it into the BlazeHand model with a positive hand pose, and output the key point location information.
6. An electronic device, characterized in that, include: One or more processors; and One or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform the steering wheel release detection method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The stored computer program causes the processor to execute the method for detecting steering wheel slippage as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for detecting hand release of steering wheel
CN113920310A