Driving state recognition method and device, electronic equipment and vehicle

CN122531073APending Publication Date: 2026-08-07CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
Filing Date
2026-04-01
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请旨在提出一种驾驶状态识别方法、装置、电子设备及车辆,解决当前疲劳状态检测方案存在误报率高、微表情检测精度不足、硬件成本与功耗偏高、补光干扰用户体验,以及多状态检测能力有限等问题,具体技术方案如下:

Benefits of technology

[0014]The driving state recognition method provided in this application only collects 2D image frames containing the driver's face, overcoming the problem that active monitoring schemes that rely on infrared and red-green-blue RGB dual cameras and matching supplementary lighting devices not only have high power consumption, increasing hardware complexity and cost, but also that the use of near-infrared supplementary lighting may cause discomfort to some drivers and affect user experience. By identifying keyframes in 2D image frames, an initial facial model point cloud of the driver is generated. From this point cloud, an initial 3D Gaussian ellipsoid set is generated. The initial 3D Gaussian ellipsoid set is then projected onto a plane to obtain a 2D rendered image. This 2D rendered image is compared with the keyframes to obtain the comparison difference results. The initial 3D Gaussian ellipsoid set is then optimized based on the comparison difference results to obtain a target 3D facial model point cloud of the driver. This allows for precise capture of subtle facial expression changes, reducing errors in detecting driver's facial micro-movements. The target 3D facial model point cloud identifies the driver's driving state characteristics, and these characteristics are used to determine the driver's actual driving state, enabling more accurate identification of the driver's driving state. Furthermore, this model can detect multiple types of dangerous conditions, meeting the increasingly complex multi-dimensional safety monitoring needs. In summary, this application constructs a target 3D facial model point cloud using 2D image frames to identify the driver's driving state. It is not affected by external environmental factors such as road conditions and does not rely on complex equipment. This not only reduces false alarm rate, cost, and power consumption, but also improves detection accuracy, multi-state detection capabilities, and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531073A_ABST
    Figure CN122531073A_ABST
Patent Text Reader

Abstract

The application provides a driving state recognition method and device, electronic equipment and vehicle, the method comprising: collecting a 2D image frame containing the face of a driver; identifying a key frame in the 2D image frame, and generating an initial face model point cloud of the driver through the key frame; generating an initial 3D Gaussian ellipsoid set through the initial face model point cloud; performing plane projection on the initial 3D Gaussian ellipsoid set to obtain a 2D rendering image, and comparing the 2D rendering image with the key frame to obtain a comparison difference result; optimizing the initial 3D Gaussian ellipsoid set through the comparison difference result to obtain a target 3D face model point cloud of the driver; identifying a driving state feature of the driver through the target 3D face model point cloud, and identifying the actual driving state of the driver through the driving state feature, thereby reducing the false positive rate, cost and power consumption, and improving the detection accuracy, multi-state detection capability and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a driving state recognition method, device, electronic device, and vehicle. Background Technology

[0002] Currently, the industry mainly relies on two technical approaches to detect driver fatigue: passive monitoring and active monitoring. Passive monitoring solutions indirectly infer the driver's condition by analyzing vehicle operating parameters (such as continuous driving time, lane departure, etc.); active monitoring solutions use dual red, green, and blue (RGB) cameras and near-infrared cameras to collect facial images of the driver, analyze the driver's physiological characteristics through image analysis, and then infer the driver's condition.

[0003] However, while passive monitoring solutions are simple to implement, they are easily affected by external environmental factors such as road conditions, resulting in a high false alarm rate and impacting user experience. Active monitoring solutions, based on two-dimensional image analysis, struggle to accurately capture subtle facial expression changes, especially in detecting micro-movements such as eyelid closure, leading to high errors and an inability to reliably identify early signs of fatigue. Furthermore, active monitoring solutions rely on infrared and RGB dual cameras and accompanying supplementary lighting equipment, resulting in high power consumption, increased hardware complexity and cost, and the use of near-infrared supplementary lighting may cause discomfort to some drivers, affecting user experience. In addition, existing solutions have performance limitations when simultaneously detecting multiple hazardous conditions, making it difficult to meet the increasingly complex multi-dimensional safety monitoring needs. Summary of the Invention

[0004] In view of this, this application aims to propose a driving state recognition method, device, electronic device, and vehicle to solve the problems of high false alarm rate, insufficient accuracy of micro-expression detection, high hardware cost and power consumption, interference of supplementary lighting with user experience, and limited multi-state detection capability of current fatigue state detection schemes. The specific technical solution is as follows: According to a first aspect of this application, a driving state recognition method is provided, the method comprising: Acquire 2D image frames containing the driver's face; Identify keyframes in the 2D image frame and generate an initial facial model point cloud of the driver based on the keyframes; An initial 3D Gaussian ellipsoid set is generated using the initial facial model point cloud; The initial 3D Gaussian ellipsoid set is projected onto a plane to obtain a 2D rendered image, and the 2D rendered image is compared with the key frame to obtain the comparison difference result; The initial 3D Gaussian ellipsoid set is optimized based on the comparison difference results to obtain a point cloud of the target 3D facial model of the driver. The driver's driving state features are identified by the point cloud of the target 3D facial model, and the driver's actual driving state is identified by the driving state features.

[0005] Optionally, identifying keyframes in the 2D image frame includes: Feature extraction is performed on the 2D image frame to obtain the feature points of the 2D image frame; Identify the time point and camera viewpoint when the 2D image frames are acquired; Candidate keyframes are identified from the 2D image frames using the feature points, the time points, and the camera viewpoint. Identify the common viewpoints and comprehensive information of the candidate keyframes, and identify keyframes from the candidate keyframes using the common viewpoints and comprehensive information.

[0006] Optionally, identifying candidate keyframes from the 2D image frame using the feature points, the time points, and the camera viewpoint includes: Feature points of adjacent 2D image frames are matched to obtain feature point matching pairs; If the number of feature point matching pairs is greater than the first number threshold, then all adjacent 2D image frames are determined as candidate keyframes. Identify the uniformity of the distribution and the total number of feature points within the 2D image frame; The information content score of the 2D image frame is determined by the distribution uniformity and the total number. If the information content score is greater than the score threshold, then the 2D image frame is determined as a candidate keyframe; If the camera viewpoint of the 2D image frame is greater than the acquisition viewpoint threshold, then the 2D image frame is determined as a candidate keyframe. Identify the actual time point of the candidate keyframe, and determine the time point after adding a preset duration to the actual time point as the target time point; If no new candidate keyframes are detected between the actual time point and the target time point, the 2D image frame at the nearest time point after the target time point is determined as a candidate keyframe.

[0007] Optionally, identifying the common viewpoints and comprehensive information of the candidate keyframes, and identifying keyframes from the candidate keyframes using the common viewpoints and comprehensive information, includes: Insert the candidate keyframes into a preset candidate keyframe window; The candidate keyframes located at the preset positions in the preset candidate keyframe window are marked as candidate keyframes to be processed; Identify the common viewpoints of the candidate keyframes to be processed; If the number of shared viewpoints is less than the second threshold, the candidate keyframe to be processed will be removed from the preset candidate keyframe window. Identify the visual entropy, geometric entropy, and illumination robustness value of the remaining candidate keyframes within the preset candidate keyframe window; The combined information content of the remaining candidate keyframes is determined by the visual entropy, geometric entropy, and illumination robustness value. If the total amount of information is greater than the preset information value, then the remaining candidate keyframes will be determined as keyframes.

[0008] Optionally, after identifying keyframes from the candidate keyframes using the common viewpoint and the comprehensive information, the method further includes: Insert the keyframes sequentially into the preset keyframe window; If the preset keyframe window is full, then when a new keyframe appears, the information gain value of the preset keyframe window is inserted after the new keyframe is inserted. If the information gain value is greater than 0, the new keyframe is inserted into the preset keyframe window, and keyframes that meet the preset conditions are removed from the preset keyframe window.

[0009] Optionally, generating an initial facial model point cloud of the driver using the keyframes includes: Match the feature points of any pair of keyframes to obtain target matching pairs; By combining the target matching pairs with the eight-point method, the camera relative extrinsic parameters of any two key frames are determined. The target matching pairs are triangulated using the camera's relative extrinsic parameters to obtain a sparse point cloud; Input the keyframe into the monocular depth estimation branch and output the first depth estimation map of the smooth face region in the keyframe. The keyframe and the camera relative extrinsic parameters are input into the multi-view depth estimation branch, and a second depth estimation map of the face target region in the keyframe is output. The first depth estimation map and the second depth estimation map are fused to obtain a fused depth map; The fused depth map is back-projected using the camera relative extrinsic parameters of the keyframes to obtain a dense point cloud; The sparse point cloud and the dense point cloud are merged to obtain an initial facial model point cloud containing the driver.

[0010] Optionally, after identifying the driver's actual driving state through the driving state features, the method further includes: If the actual driving condition is slightly abnormal, an audible prompt will be given; If the actual driving condition is moderately abnormal, then an audible and vibration alert will be issued; If the actual driving condition is severely abnormal, the vehicle will be slowed down and pulled over to the side of the road, and a distress message will be sent.

[0011] According to a second aspect of this application, a driving state recognition device is provided, the device comprising: The acquisition module is used to acquire 2D image frames containing the driver's face; The first recognition module is used to identify key frames in the 2D image frame and generate an initial facial model point cloud of the driver based on the key frames. The generation module is used to generate an initial 3D Gaussian ellipsoid set from the initial facial model point cloud; The comparison module is used to perform planar projection on the initial 3D Gaussian ellipsoid set to obtain a 2D rendered image, and compare the 2D rendered image with the key frame to obtain the comparison difference result; The optimization module is used to optimize the initial 3D Gaussian ellipsoid set based on the comparison difference results to obtain a point cloud of the target 3D facial model of the driver. The second recognition module is used to recognize the driver's driving state features through the point cloud of the target 3D facial model, and to recognize the driver's actual driving state through the driving state features.

[0012] According to another aspect of this application, an electronic device is also provided, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the driving state recognition method as described above.

[0013] According to another aspect of this application, a vehicle is also provided, including the aforementioned driving state recognition device.

[0014] The driving state recognition method provided in this application only collects 2D image frames containing the driver's face, overcoming the problem that active monitoring schemes that rely on infrared and red-green-blue RGB dual cameras and matching supplementary lighting devices not only have high power consumption, increasing hardware complexity and cost, but also that the use of near-infrared supplementary lighting may cause discomfort to some drivers and affect user experience. By identifying keyframes in 2D image frames, an initial facial model point cloud of the driver is generated. From this point cloud, an initial 3D Gaussian ellipsoid set is generated. The initial 3D Gaussian ellipsoid set is then projected onto a plane to obtain a 2D rendered image. This 2D rendered image is compared with the keyframes to obtain the comparison difference results. The initial 3D Gaussian ellipsoid set is then optimized based on the comparison difference results to obtain a target 3D facial model point cloud of the driver. This allows for precise capture of subtle facial expression changes, reducing errors in detecting driver's facial micro-movements. The target 3D facial model point cloud identifies the driver's driving state characteristics, and these characteristics are used to determine the driver's actual driving state, enabling more accurate identification of the driver's driving state. Furthermore, this model can detect multiple types of dangerous conditions, meeting the increasingly complex multi-dimensional safety monitoring needs. In summary, this application constructs a target 3D facial model point cloud using 2D image frames to identify the driver's driving state. It is not affected by external environmental factors such as road conditions and does not rely on complex equipment. This not only reduces false alarm rate, cost, and power consumption, but also improves detection accuracy, multi-state detection capabilities, and user experience.

[0015] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of this application more easily understood, specific embodiments of this application are given below. Attached Figure Description

[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart of the steps of a driving state recognition method provided in this application; Figure 2 yes Figure 1 The flowchart shown is a process for determining candidate keyframes in a driving state recognition method provided in this application; Figure 3 yes Figure 1 The flowchart shown is a process for determining keyframes in a driving state recognition method provided in this application; Figure 4This is a schematic diagram of the structure of a driving state recognition device provided in this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and with various variations and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0018] Currently, driver monitoring systems (DMS) are used to detect driver fatigue. There are two main types of detection methods. The first is passive monitoring, which indirectly determines the driver's state by analyzing vehicle operation data (such as continuous driving time and lane departure). The second relies on a multi-sensor system, including infrared and RGB dual cameras and accompanying supplementary lighting, to collect facial images of the driver and analyze facial features such as eyelid closure frequency and yawning. However, the first method is easily affected by external factors such as road conditions, resulting in a high false alarm rate and severely impacting user experience. The second method uses 2D images for state recognition, which has significant measurement errors in detecting micro-movements such as eyelid closure, making it difficult to accurately capture the early physiological characteristics of fatigued driving. Furthermore, multi-sensor systems are too expensive and have limited installation locations, consuming excessive computing power on the vehicle's chip, which is detrimental to the range optimization of new energy vehicles. In addition, infrared monitoring may cause driver discomfort due to inappropriate supplementary lighting intensity. Based on these problems, this application proposes a driving state recognition method. (Refer to...) Figure 1 The diagram illustrates a flowchart of a driving state recognition method provided in this application, the method including: Step 101: Acquire 2D image frames containing the driver's face.

[0019] This application utilizes a monocular camera to capture 2D image frames containing the driver's face. Before capturing the 2D image frames, this application also obtains vehicle status information via the Controller Area Network (CAN) bus to determine whether the vehicle has entered a driving state. When the vehicle is detected to be in a driving state, a start command is sent to the monocular camera installed on the driver's side A-pillar. After completing a self-test based on the start command, the monocular camera enters standby mode, waiting for an image capture command. Because a driver's poor driving state only poses a driving risk when the vehicle is in motion, this setup avoids capturing images when the vehicle is stationary, thus avoiding wasting resources.

[0020] This application triggers image acquisition via a timer, with an acquisition frequency set to 10-30 frames / second, preferably 15 frames / second. The resolution of each acquired image is set to 1920×1080 pixels to ensure the clarity of the driver's facial features. The acquired images are then transmitted in real-time to the vehicle processing unit via a USB 3.0 interface. If three consecutive acquisition attempts fail, the monocular camera is automatically restarted. If acquisition still fails after restarting, the driver is prompted to check the camera status via the vehicle's display screen. Simultaneously, an anomaly log is recorded in the vehicle system, which may include the timestamp of the anomaly, error code, and possible causes of the malfunction.

[0021] When this application acquires images using a monocular camera, the images not only contain the driver's face but also the background (such as the seat, windows, etc.) where the driver is located. Therefore, in order to obtain 2D image frames containing the driver's face, preprocessing of the acquired images is required. The preprocessing steps include using a bilateral filtering algorithm to denoise the 2D image frames. The denoising formula is shown in (1): (1) in, Represents pixels in the original 2D image frame pixel values, Represents the pixels in a denoised 2D image frame pixel values, It is a normalization factor. It is a pixel value similarity weight. It is the spatial neighborhood weight. is the radius of the filter window, i is the pixel offset in the x-direction, and j is the pixel offset in the y-direction. Represents the number of pixels in the original 2D image frame. The pixel values ​​after offsetting i in the x direction and j in the y direction.

[0022] Then, face region detection is performed on the denoised image frames. A cascaded classifier based on Haar features is used to locate the face region, resulting in a basic rectangular detection box. The formula for the cascaded classifier based on Haar features is shown in (2): (2) in, This represents the basic rectangular detection box obtained by locating the face region. These are denoised 2D image frames. (.) is a cascaded classifier function based on Haar features.

[0023] To ensure complete capture of facial features and considering potential minor deviations in the detection bounding box, this application expands the basic rectangular detection bounding box of the facial region by a certain percentage of pixels, generating a larger rectangular bounding box that includes redundant areas. Then, the expanded rectangular bounding box is cropped to obtain a 2D image frame containing the driver's face. At this point, most irrelevant background information from the original image is removed in this cropping result, and the 2D image frame containing the driver's face includes the driver's face and a small amount of surrounding background. The 2D image frame containing the driver's face is then normalized and adjusted to a standard size. This adjusted 2D image frame containing the driver's face is denoted as... .

[0024] Step 102: Identify keyframes in the 2D image frame and generate an initial facial model point cloud of the driver using the keyframes.

[0025] Before identifying keyframes from 2D image frames, this application first identifies candidate keyframes from 2D image frames, and then identifies keyframes from candidate keyframes. This application identifies candidate keyframes based on feature points, time points, and camera viewpoints of 2D image frames. Therefore, it is necessary to use the Feature from Accelerated Segment Test (FAST) corner detection and Oriented FAST and Rotated BRIEF (ORB) feature detection algorithm to extract features from 2D image frames and obtain feature points of 2D image frames. The acquisition of feature points is shown in formula (3): (3) in, A standard-sized 2D image frame containing the driver's face, where D is from... The set of feature points extracted from it. yes Feature detection algorithm.

[0026] Specifically, when calculating feature points of a 2D image frame using the ORB feature detection algorithm, this application first uses the oFAST algorithm in the ORB feature detection algorithm to quickly locate the position of the feature points in the 2D image frame, and then uses the rBRIEF algorithm in the ORB feature detection algorithm to calculate a binary "descriptor" for each detected feature point for subsequent feature matching and recognition.

[0027] After identifying candidate keyframes based on feature points, time points, and camera viewpoints, this application identifies the common viewpoints and comprehensive information content of the candidate keyframes, and identifies keyframes from the candidate keyframes based on this information. The common viewpoint refers to the point with the smallest Hamming distance among the feature points obtained after brute-force matching of the feature points of the candidate keyframes. Based on the above, step 102, "identifying keyframes in 2D image frames," specifically includes the following sub-steps: Sub-step 1021: Extract features from the 2D image frame to obtain the feature points of the 2D image frame.

[0028] Sub-step 1022: Identify the time point and camera viewpoint of the acquired 2D image frames.

[0029] Sub-step 1023 identifies candidate keyframes from 2D image frames using feature points, time points, and camera viewpoints.

[0030] Sub-step 1024: Identify the common viewpoints and comprehensive information of candidate keyframes, and identify keyframes from the candidate keyframes using the common viewpoints and comprehensive information.

[0031] The aforementioned sub-steps introduce multi-dimensional triggering conditions such as feature points, camera viewpoints, and time points in the candidate keyframe identification stage. This effectively overcomes the redundancy or omission problems that may arise from relying solely on feature matching, ensuring that the identified candidate keyframes are reasonably distributed in time, comprehensively covered in viewpoints, and rich and evenly distributed in content features. By introducing common viewpoints and comprehensive information content as higher-order indicators for final screening, the optimal balance between spatial coverage, content integrity, and visual quality of the selected keyframes is further ensured. This provides highly representative and stable inputs for subsequent tasks such as 3D reconstruction and state recognition, significantly improving robustness and accuracy under dynamic scenes or complex viewpoint changes.

[0032] The above content of this application proposes that candidate keyframes can be identified from 2D image frames using feature points, time points, and camera viewpoints. Specifically, firstly, adjacent 2D image frames are determined based on the acquisition order, feature point matching is performed on adjacent 2D image frames, and the number of feature point matches is calculated using the following formula (4): (4) in, It is the set of feature points of the i-th 2D image frame. It is the set of feature points of the j-th 2D image frame, where N is the total number of feature points. It represents the number of feature point matches between the i-th and j-th 2D image frames, and match(.) is the matching function.

[0033] If the number of matched feature points exceeds a set first threshold, then all adjacent 2D image frames are recorded as candidate keyframes. If the number of matched feature points is less than or equal to the set first threshold, then the process continues to iterate through subsequent adjacent 2D image frames. This continues until all 2D image frames have been traversed.

[0034] Because there may be a rather extreme special case, that is, after traversing all 2D image frames, no adjacent 2D image frames have a number of matching feature points greater than the set first number threshold. In this case, it is impossible to obtain candidate keyframes, which is obviously not feasible. Therefore, this application also sets other methods for determining candidate keyframes based on quality triggering, motion triggering and time triggering.

[0035] When determining candidate keyframes based on quality triggering, the identification is still based on feature points. However, what is identified now is the distribution uniformity and total number of feature points within each 2D image frame. After weighted summation of the distribution uniformity and total number, an information content score is obtained. The calculation formula (5) for the information content score is as follows: (5) in, , These are weighting coefficients. It is the uniformity of the distribution of feature points within the j-th frame of the 2D image. It is the total number of feature points within the j-th frame of the 2D image. It is the information content score of the j-th 2D image frame.

[0036] When the information content score of 2D image frame A is detected to be greater than the set score threshold, 2D image frame A is identified as a candidate keyframe. When determining candidate keyframes based on motion triggering, the camera viewpoint of the 2D image frame at the time of acquisition is first identified. If the camera viewpoint of a certain 2D image frame B is greater than the set acquisition viewpoint threshold, that is, the relative pose translation and rotation exceed a certain range, then 2D image frame B is identified as a candidate keyframe.

[0037] When determining candidate keyframes based on time triggering, at least one candidate keyframe is first identified using the method described above. Then, the actual time point at which the candidate keyframe is acquired is identified. If no new candidate keyframe is generated after a certain time interval from the actual time point, the first 2D image frame acquired after that time interval is identified as a candidate keyframe. For example, if a candidate keyframe C has been identified, the actual acquisition time is 12:00:00, and the preset duration is set to 20 seconds, then if another candidate keyframe D is acquired at 12:00:15, because the time interval between candidate keyframes C and D is 15 seconds, a candidate keyframe will not be determined based on time triggering. If no new candidate keyframe is acquired between 12:00:00 and 12:00:20, then the first 2D image frame acquired after 12:00:20 according to the acquisition frequency is identified as a candidate keyframe. The first quantity threshold, scoring threshold, acquisition viewpoint threshold, and preset duration can be set to specific values ​​as needed; this application does not impose specific limitations on them. Based on the above, sub-step 1023 specifically includes the following steps, such as... Figure 2 As shown: Step 01: Match feature points of adjacent 2D image frames to obtain feature point matching pairs.

[0038] Step 02: If the number of feature point matching pairs is greater than the first number threshold, then all adjacent 2D image frames are determined as candidate keyframes.

[0039] Step 03: Identify the uniformity and total number of feature points within a 2D image frame.

[0040] Step 04: Determine the information content score of the 2D image frame by the distribution uniformity and total number.

[0041] Step 05: If the information content score is greater than the score threshold, then the 2D image frame is determined as a candidate keyframe.

[0042] Step 06: If the camera viewpoint of the 2D image frame is greater than the acquisition viewpoint threshold, then the 2D image frame is determined as a candidate keyframe.

[0043] Step 07: Identify the actual time point of the candidate keyframe and determine the time point after adding a preset duration to the actual time point as the target time point.

[0044] Step 08: If no new candidate keyframes are detected between the actual time point and the target time point, then the 2D image frame of the nearest time point after the target time point is determined as a candidate keyframe.

[0045] The above steps, by combining four dimensions—feature matching quantity, intra-frame feature quality, camera viewpoint variation, and temporal uniformity—ensure that the selected candidate keyframes possess rich visual information (high matching degree, uniform feature distribution, and sufficient quantity) while covering important viewpoint changes and time nodes. This effectively avoids redundancy, omissions, or uneven distribution problems that may arise from traditional single-index screening. It can automatically generate a candidate keyframe set that achieves an optimal balance in information content, viewpoint diversity, and temporal coherence even in complex scenarios such as facial movement, lighting changes, and unstable acquisition intervals, significantly improving the robustness and accuracy of subsequent state recognition.

[0046] After obtaining candidate keyframes, this application can further filter them to determine the final keyframes. For this purpose, a preset candidate keyframe window is first set up. The filtered candidate keyframes are inserted into this window sequentially according to the acquisition time. The window size can be set according to needs, for example, it can be set to 10 frames. The five newly inserted frames in the window are used as candidate keyframes to be processed, and brute-force matching is performed. That is, each feature point in the five frames is compared with all other points. During the comparison, the binary "descriptor" of the feature point is compared, and the Hamming distance is calculated. The point with the smallest Hamming distance is taken as the common viewpoint. Here, the Hamming distance refers to the difference in the number of bits when comparing the binary representations of two feature points bit by bit. For example, a feature point in frame A has the descriptor 10110101. Three feature points in frame B have descriptors B-1: 10110101, B-2: 10110111, and B-3: 00101100. Calculating the Hamming distances between the feature point in frame A and the three feature points in frame B, we find that the Hamming distance with B-1 is 0, with B-2 is 1, and with B-3 is 4. Therefore, B-1 and the feature point in frame A are identified as a co-viewpoint. It's important to note that when determining co-viewpoints, a ratio is determined based on the minimum and second-minimum Hamming distances. If the ratio is less than a certain threshold, the match is considered good, and the point corresponding to the minimum Hamming distance is identified as the co-viewpoint. Otherwise, the match is considered ambiguous. A co-viewpoint refers to a corresponding two-dimensional feature point that can be observed in multiple two-dimensional images from different perspectives, representing the same three-dimensional point in space. By brute-force matching, we find feature point pairs that successfully match in different candidate keyframes to be processed. These successfully matched feature points are the common viewpoints.

[0047] After identifying common viewpoints in all candidate keyframes undergoing brute-force matching, the number of common viewpoints within each candidate keyframe is determined. If the number of common viewpoints is less than a second threshold, the candidate keyframe is discarded, i.e., removed from the preset candidate keyframe window. The second threshold can be adjusted based on the environment; it can be reduced when external lighting changes drastically.

[0048] After filtering candidate keyframes using common viewpoints, this application further sets up a method to calculate the information quality of the remaining candidate keyframes within a preset candidate keyframe window, and further determine the keyframes based on this. At this point, it is necessary to identify the visual entropy, geometric entropy, and illumination robustness value of each remaining candidate keyframe. The visual entropy can be obtained by calculating the mean variance of the ORB descriptor using an 8x8 grid and normalizing it to [0,1]. The geometric entropy refers to the entropy obtained by calculating the distribution of reprojection errors between the current frame and the last frame of the window. The illumination robustness value can be obtained using formula (6): (6) Where EV_t is the exposure value of the candidate keyframe, EV_ref is the exposure value of the reference frame, and L is the illumination robustness value of the candidate keyframe. The reference frame is usually selected as the earliest frame in the preset candidate keyframe window, or a keyframe known to be under "standard" lighting. EV_range refers to the normalization factor, representing the maximum range of exposure variation that the system expects or can handle. It is a preset empirical constant used to scale the difference to the [0,1] interval.

[0049] After assigning weights to the visual entropy, geometric entropy, and illumination robustness value, the weighted sum is used to obtain the comprehensive information content of each remaining candidate keyframe. The formula for calculating the comprehensive information content (7) is as follows: (7) in, , These are weighting coefficients. It is the visual entropy of the candidate keyframe. L is the geometric entropy of the candidate keyframe, L is the illumination robustness value of the candidate keyframe, and M is the total information content of the candidate keyframe.

[0050] If the total information content of a candidate keyframe is greater than a preset information value (which can be set to 0.55 or 0.6, etc.), then the candidate keyframe is determined as a keyframe. Therefore, sub-step 1024 specifically includes the following steps, such as... Figure 3 As shown: Step 10: Insert the candidate keyframes into the preset candidate keyframe window.

[0051] Step 11: Mark the candidate keyframes located at the preset positions in the preset candidate keyframe window as candidate keyframes to be processed.

[0052] Step 12: Identify the common viewpoints of the candidate keyframes to be processed.

[0053] Step 13: If the number of common viewpoints is less than the second threshold, the candidate keyframes to be processed will be removed from the preset candidate keyframe window.

[0054] Step 14: Identify the visual entropy, geometric entropy, and illumination robustness value of the remaining candidate keyframes within the preset candidate keyframe window.

[0055] Step 15: Determine the total information content of the remaining candidate keyframes using visual entropy, geometric entropy, and illumination robustness value.

[0056] Step 16: If the total amount of information is greater than the preset information value, then the remaining candidate keyframes are determined as keyframes.

[0057] The candidate keyframes at the preset positions can be the aforementioned 5 newly inserted frames, or 4 newly inserted frames, etc., adjusted according to requirements; this application does not impose specific limitations here. When setting weights for visual entropy, geometric entropy, and illumination robustness value, the weights can be dynamically adjusted according to the actual situation. For example, when the illumination changes drastically, the weight of the illumination robustness value can be increased, and the weight of the visual entropy can be decreased. When the average visual entropy of the remaining candidate keyframes is greater than a certain value, the weight of the visual entropy can be increased.

[0058] The above steps, through establishing a sliding window mechanism for candidate keyframes and introducing multi-stage fine-tuning, achieve efficient purification of high-quality keyframes. Specifically, window management ensures the continuity and real-time nature of processing; initial filtering based on the number of shared viewpoints effectively removes frames with excessive viewpoint overlap and redundant information, improving the viewpoint diversity of the frame set; and a comprehensive information content evaluation integrating visual entropy, geometric entropy, and illumination robustness values ​​selects the most representative and robust keyframes from three dimensions: visual content, structural features, and anti-interference capability. This lays a solid data foundation for subsequent tasks such as 3D reconstruction, state estimation, or model optimization.

[0059] This application also sets up a preset keyframe window and inserts keyframes into this window. When inserting keyframes, this application sets up a preset keyframe window maintenance mechanism. Its main purpose is to ensure uniform spatiotemporal distribution and maximum information content within a fixed maximum window capacity. If the preset keyframe window is not full, keyframes are directly inserted in descending order; if the preset keyframe window is full, an incremental elimination process is executed. At this point, it is necessary to determine whether inserting a new keyframe into the preset keyframe window will increase the overall comprehensive information content within the preset keyframe window. If so, the new keyframe is inserted into the preset keyframe window, and simultaneously, a keyframe that meets preset conditions is eliminated from the preset keyframe window. The specific execution steps include: Insert keyframes sequentially into the preset keyframe window; If the preset keyframe window is full, then when a new keyframe appears, the information gain value of the preset keyframe window is inserted after the new keyframe is inserted. If the information gain value is greater than 0, the new keyframe will be inserted into the preset keyframe window, and keyframes that meet the preset conditions will be removed from the preset keyframe window.

[0060] The information gain value of the preset keyframe window is calculated as follows: the total information content within the preset keyframe window after inserting a new keyframe - the total information content within the preset keyframe window before inserting a new keyframe plus the change in common viewpoint × preset coefficient. A keyframe that meets the preset conditions is a keyframe determined by combining factors such as total information content, acquisition time, and distance from the center frame. The elimination principle is that the smaller the total information content, the easier it is to eliminate; the earlier the acquisition time, the easier it is to eliminate; and the farther the distance from the center frame, the easier it is to eliminate. The formula (8) for determining keyframes that meet the preset conditions is as follows: (8) in, , , These are weighting coefficients. It is the total information content of the i-th keyframe. It is the insertion time of the new keyframe. It is the insertion time range for all frames within the preset keyframe window. It is the insertion time of the i-th keyframe. It is the distance between the i-th keyframe and the center frame within the preset keyframe window. It is the maximum distance between keyframes within the preset keyframe window and the center frame within the preset keyframe window. It is a rejection function. It is a keyframe that meets the preset conditions.

[0061] The above steps ensure that a set of keyframes with the highest information density and strongest complementarity is always maintained within the preset keyframe window. When a new keyframe arrives, its insertion is no longer based on simple temporal replacement, but strictly on the information gain it brings. Only when a new frame can contribute new and distinctive visual information to the existing set is it included in the window. At the same time, old frames that meet preset conditions are removed. This mechanism continuously improves the overall information content and diversity of the keyframe set within a limited capacity, effectively avoiding data redundancy and information saturation. It ensures that subsequent calculations such as sparse point cloud reconstruction and dense mapping are always based on the most representative and discriminative visual data, thereby significantly improving the efficiency and accuracy of the entire processing flow.

[0062] To prevent window aging, this application will also mark the 3 frames with the lowest overall information content as loop closure candidate frames and send them to the loop closure matching thread for judgment every 10 frames inserted (that is, when the preset keyframe window is filled). If extreme texture loss occurs, the window will be temporarily expanded to 1.5 times and then restored to the original window size after it stabilizes.

[0063] In this application, after obtaining the keyframes, an initial facial model point cloud is constructed based on the keyframes. During the construction process, since the feature points of the 2D image frames have already been identified, and the keyframes are obtained after multiple screenings of the 2D image frames, the feature points in the keyframes do not need to be re-identified. They can be directly accessed based on the feature points of their corresponding original 2D image frames. To facilitate differentiation, the set of feature points within the keyframes is represented using... The keyframes of any two frames are matched for feature points to obtain target matching pairs between the frames. The target matching pairs contain the correspondence of feature points (e.g., the i-th point of frame A corresponds to the j-th point of frame B). Then, the eight-point method is used to analyze the target matching pairs to obtain the fundamental matrix describing the epipolar geometric relationship between the two frames. The calculation formula (9) is as follows: (9) in, It is the set of feature points within the i-th keyframe. It is the set of feature points within the j-th keyframe. () is the calculation function for the fundamental matrix using the eight-point method, where F is the fundamental matrix between the i-th frame and the j-th keyframe.

[0064] By decomposing the fundamental matrix, the relative pose between the two keyframes corresponding to the target matching pair can be obtained (i.e., the camera relative extrinsic parameters of the two keyframes, including the rotation matrix R and the translation matrix t). By triangulating the corresponding target matching pair with the camera relative extrinsic parameters, a 3D map point can be obtained. By triangulating all target matching pairs in the above manner, a sparse point cloud can be obtained. After obtaining the sparse point cloud, the 3D point coordinates of the camera relative extrinsic parameters and the sparse point cloud can be optimized by using the bundle adjustment method to minimize the reprojection error. The calculation formula (10) for minimization is as follows: (10) in, It is a projection function. It is the rotation matrix of the i-th keyframe. It is the coordinate of the j-th 3D point within the i-th keyframe. It is the translation matrix of the i-th keyframe. It represents the pixel coordinates of the j-th 3D point in the i-th keyframe. It is the square norm of the reprojection error. () is a function that minimizes the reprojection error.

[0065] After obtaining the sparse point cloud, this application generates a dense point cloud based on the original keyframe image. Traditional methods for constructing dense point clouds use column-mapping (COLMAP) feature matching algorithms. However, large areas of facial skin are smooth, lacking sufficient texture features for traditional feature matching algorithms, resulting in sparse, hollow, or noisy reconstructed point clouds. Furthermore, the subcutaneous scattering effect of skin and the assumption of reflectivity consistency lead to inaccurate matching cost calculations. Therefore, this application designs a lightweight, dual-branch network structure. One branch is a monocular depth estimation branch, which estimates the depth of large smooth areas of facial skin; the other is a multi-view depth estimation branch, which estimates the depth of target areas such as the corners of the eyes, nose, and mouth. An adaptive attention mechanism fusion module then fuses the outputs of both branches to obtain a fused depth map of the facial region. A dense point cloud is then generated based on this fused depth map. Finally, the sparse and dense point clouds are merged to obtain an initial facial model point cloud containing the driver. Based on the above, step 102, "generating an initial facial model point cloud of the driver using keyframes," specifically includes the following sub-steps: Sub-step 1025: Match feature points of any two keyframes to obtain target matching pairs.

[0066] Sub-step 1026: Determine the relative extrinsic parameters of the camera for any pair of keyframes by combining target matching pairs with the eight-point method.

[0067] Sub-step 1027: Triangulate the target matching pairs using the camera relative to extrinsic parameters to obtain sparse point clouds.

[0068] Sub-step 1028: Input the keyframe into the monocular depth estimation branch and output the first depth estimation map of the smooth face region in the keyframe.

[0069] Sub-step 1029 inputs the keyframe and camera relative extrinsic parameters into the multi-view depth estimation branch and outputs a second depth estimation map of the face target region in the keyframe.

[0070] Sub-step 10210: The first depth estimation map and the second depth estimation map are fused to obtain a fused depth map.

[0071] Sub-step 10211: Back-project the fused depth map using the camera relative extrinsic parameters of the keyframe to obtain a dense point cloud.

[0072] Sub-step 10212 merges the sparse point cloud and the dense point cloud to obtain an initial facial model point cloud containing the driver.

[0073] The Eight-Point Algorithm is a classic algorithm for solving the fundamental matrix. It linearly estimates the epipolar geometry (i.e., the camera-relative extrinsic parameters of each pair of keyframes) using at least eight pairs of matched feature points. The target region can be areas such as the corners of the eyes, nose, and mouth of a face. The pixel values ​​in the depth estimation map are the depth values ​​of the scene points corresponding to those pixels in the camera coordinate system. When back-projecting the fused depth map using the camera-relative extrinsic parameters of the keyframes, the pixel coordinates (e.g., the coordinates of pixel A) are first back-projected onto the normalized camera plane using the camera-relative extrinsic parameters, resulting in a 3D direction vector. This direction vector is then multiplied by the depth value of pixel A to obtain the 3D coordinates of pixel A in the current camera coordinate system. Finally, the camera-relative extrinsic parameters are used to transform pixel A from the current camera coordinate system to a global or reference coordinate system, yielding its final 3D coordinates. After performing the above operations on all pixels in the fused depth map, a dense point cloud is obtained.

[0074] The above steps first utilize sparse point clouds generated based on feature matching and motion reconstruction methods to provide accurate constraints on key facial structures and camera pose. Simultaneously, the introduced monocular depth estimation and multi-view depth estimation branches respectively compensate for the lack of data in smooth regions of the sparse point cloud and improve the depth accuracy of the target region by utilizing multi-view consistency. Subsequently, the dense and sparse point clouds are merged to form a facial model, which not only significantly improves the spatial coverage and geometric detail restoration capabilities of the point cloud but also particularly enhances the modeling effect on smooth facial regions lacking texture features.

[0075] It should be noted that this application also designs a loss function during the fusion process, placing greater trust on the geometric branch in texture-rich regions and greater trust on the learned branch in texture-weak regions, thereby encouraging branch complementarity. This results in a more accurate fused depth map of the facial region. The formula (11) for the loss function is as follows: (11) in, , , , These are weighting coefficients. It is the fusion depth estimation loss. It is the monocular depth estimation loss. It is the geometric depth estimation loss. It is the loss for diversity depth estimation. It is the overall depth estimation loss. Determined by the first depth estimation map and Determined by the second depth estimation map Determined by merging depth maps.

[0076] This application can also perform consistency checks according to the criterion of "the ratio of baseline distance to shooting distance is less than 5%", retaining only depth points with reprojection error less than one pixel and geometric consistency exceeding 75%. Then, a local weighted average is performed using a small voxel grid with a side length of 0.5 mm to obtain a dense point cloud without color. Finally, the surface of the point cloud is reconstructed using the Poisson reconstruction algorithm. The main features are setting the maximum octree depth to 11, the selection weight to 4, and the minimum number of faces to 10. First, a dense grid is generated, and then fake faces are removed using a confidence threshold of 0.25 to obtain a 3D mesh model of the scene. During the point cloud reconstruction process, parameters are dynamically adjusted according to different scene characteristics: when the ORB feature count of a single image is detected to be >8000, the scene is considered rich. At this time, the candidate propagation window for batch matching is enlarged from 7x7 to 11x11, increasing the effective matching number by approximately 40%. For weak texture areas, a depth estimation method based on photometric consistency is used.

[0077] Step 103: Generate an initial 3D Gaussian ellipsoid set using the initial facial model point cloud.

[0078] In this application, after obtaining the initial facial model point cloud, it is first segmented into different point cloud regions (e.g., eye region, nose region, mouth region, face region, background region, etc.). The point density distribution characteristics of each point cloud region (number of 3D points in the point cloud region / number of 3D points in the initial facial model point cloud) and its corresponding center position (mean of the coordinates of all 3D points in the point cloud region) are calculated. For each point cloud region, the covariance matrix is ​​calculated using its center position and the coordinates of each 3D point. An initial opacity value is determined based on the point density distribution characteristics, and the color information of the 3D points in the point cloud region is obtained from the original keyframes corresponding to the point cloud region. Different initial 3D Gaussian ellipsoids are defined based on the center position, covariance matrix, opacity value, and color information of different point cloud regions. An initial 3D Gaussian ellipsoid set is constructed using all the initial 3D Gaussian ellipsoids.

[0079] Step 104: Project the initial 3D Gaussian ellipsoid set onto a plane to obtain a 2D rendered image, and compare the 2D rendered image with the keyframes to obtain the comparison difference results.

[0080] In this application, the construction of the 3D Gaussian ellipsoid is divided into two steps. The first step is to define the geometric parameters of the 3D Gaussian ellipsoid, including defining the probability density shape of the 3D Gaussian ellipsoid through formula (12): (12) in, It is the probability density shape of a 3D Gaussian ellipsoid. These are the coordinates of three-dimensional points within the point cloud region. It is the center of the point cloud region. It is the covariance matrix of the point cloud region. It is the inverse of the covariance matrix of the point cloud region. Then, by decomposing the covariance matrix, the shape of the 3D Gaussian ellipsoid is defined, as shown in formula (13) below: (13) Where R is the rotation matrix, which determines the orientation of the 3D Gaussian ellipsoid, and S is the scale matrix, which determines the radius of the 3D Gaussian ellipsoid on each axis. It is the transpose of the scaling matrix. It is the transpose of the rotation matrix. It is the covariance matrix.

[0081] The second step defines the appearance parameters of the 3D Gaussian ellipsoid. At this point, this application performs a planar projection on each initial 3D Gaussian ellipsoid within the initial 3D Gaussian ellipsoid set. During projection, the covariance matrix of the initial 3D Gaussian ellipsoid is first converted into a 2D covariance matrix, and the projection area is determined based on the 2D covariance matrix. The formula (14) for the projected 2D covariance matrix is ​​as follows: (14) Where J is the projection Jacobian matrix and W is the view transformation matrix. It is the covariance matrix. It is the transpose of the view transformation matrix. It is the transpose of the projective Jacobian matrix. It is the projected 2D covariance matrix.

[0082] Projecting its 3D center position yields the 2D center coordinates of the projection area. Within this projection area, for each pixel coordinate, the transparency contribution of the ellipsoid to that pixel is calculated. For example, taking pixel u of the i-th 3D Gaussian ellipsoid as an example, the formula (15) for calculating the transparency contribution of the ellipsoid to that pixel is as follows: (15) in, It is the transparency contribution value of the i-th 3D Gaussian ellipsoid at pixel u. It is the fundamental opacity of the i-th 3D Gaussian ellipsoid. It is a pixel. coordinates These are the 2D center coordinates of the i-th 3D Gaussian ellipsoid projected onto the 2D image. It is the 2D covariance matrix of the i-th 3D Gaussian ellipsoid. It is a pixel. Squared Mahalanobis distance to the 2D center coordinates.

[0083] Then, the depth values ​​of all Gaussian ellipsoids covering the pixel u are calculated. Based on the depth values, the projections of the Gaussian ellipsoids are sorted from near to far. Then, Alpha blending is performed from front to back to obtain the cumulative transmittance of the pixel u. The formula for calculating the transmittance (16) is as follows: (16) in, It is the transmittance of the i-th 3D Gaussian ellipsoid at pixel u. It is the transparency contribution value of the j-th 3D Gaussian ellipsoid at pixel u.

[0084] For each Gaussian ellipsoid involved in the sorting, its color at pixel u is obtained, multiplied by the transparency contribution value at pixel u and the current transmittance, and then summed to generate the corresponding pixel color in the final 2D rendered image. The formula for calculating the pixel color of the final 2D rendered image (17) is as follows: (17) in, It is the transmittance of the i-th 3D Gaussian ellipsoid at pixel u. It is the transparency contribution value of the i-th 3D Gaussian ellipsoid at pixel u. It is the color of the i-th 3D Gaussian ellipsoid at pixel u. It is the pixel color after pixel u is projected onto the 2D rendered image, and N is the total number of Gaussian ellipsoids participating in the sorting.

[0085] During the projection process, different projection strategies are adopted for Gaussian ellipsoids at different distances: more detailed information is preserved for nearby Gaussian ellipsoids, and high-resolution projection is used, such as facial features, i.e., pupils, mouth, etc.; for distant Gaussian ellipsoids, downsampling is performed, such as background areas that may be involved, like seats, to improve rendering efficiency.

[0086] After obtaining the 2D rendered image, this application compares it with each keyframe to identify pixel differences.

[0087] Step 105: Optimize the initial 3D Gaussian ellipsoid set by comparing the difference results to obtain the point cloud of the target 3D facial model of the driver.

[0088] This application sets a pixel difference threshold. When the pixel difference exceeds the threshold, spatial hashing is used to accelerate retrieval, identifying which specific Gaussian ellipsoid projections the difference region intersects with. Then, the depth values ​​of the Gaussian ellipsoids involved in the projection are adjusted to dynamically maintain occlusion priority. Furthermore, relevant parameters of the Gaussian ellipsoids involved in the projection (including center position, covariance matrix, opacity value, and color information) can be adjusted until the pixel difference is less than or equal to the pixel difference threshold, at which point the adjustment ends. Based on the parameter-optimized Gaussian ellipsoid set, a 3D facial model point cloud of the driver is obtained.

[0089] Step 106: Identify the driver's driving state features through the point cloud of the target 3D facial model, and identify the driver's actual driving state through the driving state features.

[0090] The driving state features in this application may include facial driving state features, such as yawning frequency (frequency of yawning within 1 minute), eye gaze direction (time spent looking forward within 1 minute), eye opening and closing degree (number of times the eyes close and duration of eye closure within 1 minute), head posture (time spent looking away from the center within 1 minute), head nodding frequency (number of times the head nods within 1 minute), and mouth opening degree (number of times the mouth opens within 1 minute). It may also include body movement driving state features, such as the frequency and actions of the driver's steering wheel operations and the degree of body tilt. Facial driving state features can be identified through point cloud recognition of a target 3D facial model. Body movement driving state features can be identified through a monocular camera.

[0091] The aforementioned features are input into the state feature processing module, which includes a feature extraction layer, a feature fusion layer, and a feature mapping layer. These layers fuse features to determine the driver's fatigue level, dangerous driving level, and level of inattention, thereby determining the driver's driving state. The driver's driving state can be set to four levels: normal driving state, slightly abnormal driving state, moderately abnormal driving state, and severely abnormal driving state. This application sets different control mechanisms based on different driving states, vehicle data, and engineering experience. Specific implementation steps include: If the actual driving condition is slightly abnormal, an audible prompt will be given; If the actual driving condition is moderately abnormal, then an audible and vibration alert will be given. If the actual driving condition is seriously abnormal, control the vehicle to slow down and pull over to the side of the road and send a distress message.

[0092] For example, if the driver's driving state is detected to be normal, then the driving state is determined to be normal. At this time, the current driving state (lighting, sound, and suitable temperature, etc.) can be maintained, and music (frequently listened to, soothing music, and custom music, etc.) can be played, and fragrance (frequently used fragrance, refreshing fragrance, and custom fragrance, etc.) can be released. If the driver is detected to have their eyes closed for more than 2 seconds continuously within 1 minute, or their head deviating from the normal angle by more than 15 degrees within 1 minute, a warning for a slightly abnormal driving state is triggered (this can also be determined based on other driving characteristics, indicating a slightly abnormal driving state). In this case, audio prompts can be given, such as a gentle reminder through a voice assistant to help the driver adjust their attention, or the voice assistant can proactively chat with the driver to alleviate driver fatigue. If the driver is detected to be frequently distracted, dozing off, lethargic, or unresponsive to Level 1 abnormal alerts, a warning for a moderate abnormal driving state is triggered. In this case, not only should audio prompts be given, such as reminders through a voice assistant and continuous guidance for safe driving, but vibration prompts should also be given, such as setting up steering wheel vibration, seat vibration, and seatbelt tightening. When a driver is detected to be in a moderately abnormal driving state, exhibiting symptoms such as unresponsiveness, persistent drowsiness, or incapacitation, a warning for a severely abnormal driving state is triggered. In this case, the vehicle will automatically reduce its speed, pull over, and periodically honk its horn to alert following vehicles. It will also call the user center, relatives' numbers, and emergency numbers for assistance. Furthermore, after confirming an abnormal driving state, a detailed report of the abnormal situation will be generated, including timestamps, different driving states, and specific behavioral indicators. The output interface supports both CAN bus and wireless transmission, allowing connection to in-vehicle display systems and cloud monitoring platforms. The detailed abnormal situation report is encapsulated in a structure format, including fields such as status code, confidence level, and recommended actions, helping rescue personnel to analyze the driver's condition promptly and provide appropriate assistance.

[0093] The above steps employ only audible alerts for mild abnormal conditions, serving as a warning while avoiding excessive interference. For moderate abnormal conditions, both audible and vibration alerts are activated simultaneously, using multi-sensory feedback to enhance the warning and prompt the driver to adjust their state promptly. In cases of severe abnormal driving, the vehicle automatically decelerates and pulls over, sending a distress signal and directly intervening to maximize safety. This progressive design balances driving autonomy and safety, guiding the driver to self-correct through gentle prompts in the early stages of risk, while also taking proactive control measures in emergencies, effectively reducing the risk of accidents caused by abnormal driver behavior.

[0094] The driving state recognition method provided in this application only collects 2D image frames containing the driver's face, overcoming the problem that active monitoring schemes that rely on infrared and red-green-blue RGB dual cameras and matching supplementary lighting devices not only have high power consumption, increasing hardware complexity and cost, but also that the use of near-infrared supplementary lighting may cause discomfort to some drivers and affect user experience. By identifying keyframes in 2D image frames, an initial facial model point cloud of the driver is generated. From this point cloud, an initial 3D Gaussian ellipsoid set is generated. The initial 3D Gaussian ellipsoid set is then projected onto a plane to obtain a 2D rendered image. This rendered image is compared with the keyframes to obtain the comparison difference results. The initial 3D Gaussian ellipsoid set is then optimized based on the comparison difference results to obtain a target 3D facial model point cloud of the driver. This method accurately captures subtle facial expression changes, reducing errors in detecting driver facial micro-movements. The target 3D facial model point cloud identifies the driver's driving state characteristics, and these characteristics are used to determine the driver's actual driving state, enabling more accurate identification of the driver's driving state. Furthermore, this model can detect multiple types of dangerous states, meeting the increasingly complex multi-dimensional safety monitoring needs. In summary, this application uses 2D image frames to construct a target 3D facial model point cloud to identify the driver's driving state. This method is unaffected by external environmental factors such as road conditions and does not rely on complex equipment. It not only reduces false alarm rates, costs, and power consumption but also improves detection accuracy, multi-state detection capabilities, and user experience.

[0095] Reference Figure 4 The diagram shows a structural schematic of a driving state recognition device provided in this application, the device comprising: The acquisition module 201 is used to acquire 2D image frames containing the driver's face.

[0096] The first recognition module 202 is used to recognize key frames in 2D image frames and generate an initial facial model point cloud of the driver based on the key frames.

[0097] The generation module 203 is used to generate an initial 3D Gaussian ellipsoid set from the initial facial model point cloud.

[0098] The comparison module 204 is used to perform planar projection on the initial 3D Gaussian ellipsoid set to obtain a 2D rendered image, and compare the 2D rendered image with the key frame to obtain the comparison difference result.

[0099] The optimization module 205 is used to optimize the initial 3D Gaussian ellipsoid set by comparing the difference results, so as to obtain the point cloud of the target 3D facial model of the driver.

[0100] The second recognition module 206 is used to recognize the driver's driving state features through the point cloud of the target 3D facial model, and to recognize the driver's actual driving state through the driving state features.

[0101] Optionally, the first identification module 202 includes: The feature extraction submodule is used to extract features from 2D image frames to obtain feature points of the 2D image frames.

[0102] The first recognition submodule is used to identify the time point and camera viewpoint of the acquired 2D image frames.

[0103] The second recognition submodule is used to identify candidate keyframes from 2D image frames by using feature points, time points, and camera viewpoints.

[0104] The third identification submodule is used to identify the common viewpoints and comprehensive information of candidate keyframes, and to identify keyframes from the candidate keyframes through the common viewpoints and comprehensive information.

[0105] Optionally, the second identification submodule includes: The matching unit is used to match feature points of adjacent 2D image frames to obtain feature point matching pairs.

[0106] The first determining unit is used to determine adjacent 2D image frames as candidate keyframes if the number of feature point matching pairs is greater than a first quantity threshold.

[0107] The first recognition unit is used to identify the uniformity of the distribution and the total number of feature points within a 2D image frame.

[0108] The second determining unit is used to determine the information content score of a 2D image frame by means of distribution uniformity and total quantity.

[0109] The third determining unit is used to determine the 2D image frame as a candidate keyframe if the information content score is greater than the score threshold.

[0110] The fourth determining unit is used to determine the 2D image frame as a candidate keyframe if the camera angle of the 2D image frame is greater than the acquisition angle threshold.

[0111] The second identification unit is used to identify the actual time point of the candidate keyframe and determine the time point after adding a preset duration to the actual time point as the target time point.

[0112] The fifth determining unit is used to determine the 2D image frame at the nearest time point after the target time point as a candidate key frame if no new candidate key frame is detected between the actual time point and the target time point.

[0113] Optionally, the third identification submodule includes: The insertion unit is used to insert candidate keyframes into a preset candidate keyframe window.

[0114] The marking unit is used to mark the candidate keyframes located at the preset position in the preset candidate keyframe window as candidate keyframes to be processed.

[0115] The third identification unit is used to identify the common viewpoints of the candidate keyframes to be processed.

[0116] The elimination unit is used to eliminate the candidate keyframe to be processed from the preset candidate keyframe window if the number of co-viewpoints is less than the second number threshold.

[0117] The fourth identification unit is used to identify the visual entropy, geometric entropy, and illumination robustness value of the remaining candidate keyframes within the preset candidate keyframe window.

[0118] The sixth determining unit is used to determine the comprehensive information content of the remaining candidate keyframes through visual entropy, geometric entropy, and illumination robustness value.

[0119] The seventh determining unit is used to determine the remaining candidate keyframes as keyframes if the total amount of information is greater than the preset information value.

[0120] Optionally, the driving status recognition device also includes: The Insert module is used to insert keyframes sequentially into a preset keyframe window.

[0121] The judgment module is used to determine the information gain value of the preset keyframe window after inserting the new keyframe when the preset keyframe window is full.

[0122] The rejection module is used to insert new keyframes into the preset keyframe window if the information gain value is greater than 0, and to reject keyframes that meet preset conditions from the preset keyframe window.

[0123] Optionally, the first identification module 202 further includes: The matching submodule is used to match feature points of any two keyframes to obtain target matching pairs.

[0124] The extrinsic parameter determination submodule is used to determine the camera relative extrinsic parameters of any pair of keyframes by combining target matching pairs with the eight-point method.

[0125] The triangulation submodule is used to triangulate the target matching pairs using the camera's relative extrinsic parameters to obtain sparse point clouds.

[0126] The first input / output submodule is used to input keyframes into the monocular depth estimation branch and output a first depth estimation map of the smooth face region in the keyframe.

[0127] The second input / output submodule is used to input keyframes and camera relative extrinsic parameters into the multi-view depth estimation branch and output a second depth estimation map of the face target region in the keyframe.

[0128] The fusion submodule is used to fuse the first depth estimation map and the second depth estimation map to obtain a fused depth map.

[0129] The backprojection submodule is used to backproject the fused depth map using the camera relative extrinsic parameters of keyframes to obtain a dense point cloud.

[0130] The merge submodule is used to merge sparse and dense point clouds to obtain an initial facial model point cloud containing the driver.

[0131] Optionally, the driving status recognition device also includes: The first prompt module is used to provide an audible prompt if the actual driving condition is slightly abnormal.

[0132] The second prompt module is used to provide sound and vibration prompts if the actual driving condition is moderately abnormal.

[0133] The control module is used to control the vehicle to slow down and pull over to the side of the road and send a distress message if the actual driving condition is seriously abnormal.

[0134] The driving state recognition device provided in this application only collects 2D image frames containing the driver's face, overcoming the problem that active monitoring schemes that rely on infrared and red-green-blue RGB dual cameras and matching supplementary lighting equipment not only have high power consumption, increasing hardware complexity and cost, but also that the use of near-infrared supplementary lighting may cause discomfort to some drivers and affect user experience. By identifying keyframes in 2D image frames, an initial facial model point cloud of the driver is generated. From this point cloud, an initial 3D Gaussian ellipsoid set is generated. The initial 3D Gaussian ellipsoid set is then projected onto a plane to obtain a 2D rendered image. This 2D rendered image is compared with the keyframes to obtain the comparison difference results. The initial 3D Gaussian ellipsoid set is then optimized based on the comparison difference results to obtain a target 3D facial model point cloud of the driver. This allows for precise capture of subtle facial expression changes, reducing errors in detecting driver's facial micro-movements. The target 3D facial model point cloud identifies the driver's driving state characteristics, and these characteristics are used to determine the driver's actual driving state, enabling more accurate identification of the driver's driving state. Furthermore, this model can detect multiple types of dangerous conditions, meeting the increasingly complex multi-dimensional safety monitoring needs. In summary, this application constructs a target 3D facial model point cloud using 2D image frames to identify the driver's driving state. It is not affected by external environmental factors such as road conditions and does not rely on complex equipment. This not only reduces false alarm rate, cost, and power consumption, but also improves detection accuracy, multi-state detection capabilities, and user experience.

[0135] Reference Figure 5 This application also provides an electronic device, which may be, but is not limited to, an in-vehicle terminal. For example... Figure 5 As shown, it includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304. Processor 301, memory 303 for storing processor-executable instructions; The processor 301 is configured to execute the instructions to implement the driving state recognition method as described above: Acquire 2D image frames containing the driver's face; Identify keyframes in the 2D image frame and generate an initial facial model point cloud of the driver based on the keyframes; An initial 3D Gaussian ellipsoid set is generated using the initial facial model point cloud; The initial 3D Gaussian ellipsoid set is projected onto a plane to obtain a 2D rendered image, and the 2D rendered image is compared with the key frame to obtain the comparison difference result; The initial 3D Gaussian ellipsoid set is optimized based on the comparison difference results to obtain a point cloud of the target 3D facial model of the driver. The driver's driving state features are identified by the point cloud of the target 3D facial model, and the driver's actual driving state is identified by the driving state features.

[0136] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0137] The communication interface is used for communication between the aforementioned terminal and other devices.

[0138] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0139] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0140] In another embodiment provided in this application, a vehicle is also provided, which may specifically include the above-mentioned driving state recognition device.

[0141] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0142] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0143] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0144] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A driving state recognition method, characterized in that, The method includes: Acquire 2D image frames containing the driver's face; Keyframes in the 2D image frame are identified, and an initial facial model point cloud of the driver is generated based on the keyframes. An initial 3D Gaussian ellipsoid set is generated using the initial facial model point cloud; The initial 3D Gaussian ellipsoid set is projected onto a plane to obtain a 2D rendered image, and the 2D rendered image is compared with the key frame to obtain the comparison difference result; The initial 3D Gaussian ellipsoid set is optimized based on the comparison difference results to obtain a point cloud of the target 3D facial model of the driver. The driver's driving state features are identified by the point cloud of the target 3D facial model, and the driver's actual driving state is identified by the driving state features.

2. The method according to claim 1, characterized in that, The identification of keyframes in the 2D image frame includes: Feature extraction is performed on the 2D image frame to obtain the feature points of the 2D image frame; Identify the time point and camera viewpoint when the 2D image frames are acquired; Candidate keyframes are identified from the 2D image frames using the feature points, the time points, and the camera viewpoint. Identify the common viewpoints and comprehensive information of the candidate keyframes, and identify keyframes from the candidate keyframes using the common viewpoints and comprehensive information.

3. The method according to claim 2, characterized in that, The step of identifying candidate keyframes from the 2D image frame using the feature points, the time points, and the camera viewpoint includes: Feature points of adjacent 2D image frames are matched to obtain feature point matching pairs; If the number of feature point matching pairs is greater than the first number threshold, then all adjacent 2D image frames are determined as candidate keyframes. Identify the uniformity of the distribution and the total number of feature points within the 2D image frame; The information content score of the 2D image frame is determined by the distribution uniformity and the total number. If the information content score is greater than the score threshold, then the 2D image frame is determined as a candidate keyframe; If the camera viewpoint of the 2D image frame is greater than the acquisition viewpoint threshold, then the 2D image frame is determined as a candidate keyframe. Identify the actual time point of the candidate keyframe, and determine the time point after adding a preset duration to the actual time point as the target time point; If no new candidate keyframes are detected between the actual time point and the target time point, the 2D image frame at the nearest time point after the target time point is determined as a candidate keyframe.

4. The method according to claim 2, characterized in that, The step of identifying the common viewpoints and comprehensive information of the candidate keyframes, and identifying keyframes from the candidate keyframes using the common viewpoints and comprehensive information, includes: Insert the candidate keyframes into a preset candidate keyframe window; The candidate keyframes located at the preset positions in the preset candidate keyframe window are marked as candidate keyframes to be processed; Identify the common viewpoints of the candidate keyframes to be processed; If the number of shared viewpoints is less than the second threshold, the candidate keyframe to be processed will be removed from the preset candidate keyframe window. Identify the visual entropy, geometric entropy, and illumination robustness value of the remaining candidate keyframes within the preset candidate keyframe window; The combined information content of the remaining candidate keyframes is determined by the visual entropy, geometric entropy, and illumination robustness value. If the total amount of information is greater than the preset information value, then the remaining candidate keyframes will be determined as keyframes.

5. The method according to claim 2, characterized in that, After identifying keyframes from the candidate keyframes using the common viewpoint and the comprehensive information, the method further includes: Insert the keyframes sequentially into the preset keyframe window; If the preset keyframe window is full, then when a new keyframe appears, the information gain value of the preset keyframe window is inserted after the new keyframe is inserted. If the information gain value is greater than 0, the new keyframe is inserted into the preset keyframe window, and keyframes that meet the preset conditions are removed from the preset keyframe window.

6. The method according to claim 2, characterized in that, The step of generating an initial facial model point cloud of the driver using the keyframes includes: Match the feature points of any pair of keyframes to obtain target matching pairs; By combining the target matching pairs with the eight-point method, the camera relative extrinsic parameters of any two key frames are determined. The target matching pairs are triangulated using the camera's relative extrinsic parameters to obtain a sparse point cloud; Input the keyframe into the monocular depth estimation branch and output the first depth estimation map of the smooth face region in the keyframe. The keyframe and the camera relative extrinsic parameters are input into the multi-view depth estimation branch, and a second depth estimation map of the face target region in the keyframe is output. The first depth estimation map and the second depth estimation map are fused to obtain a fused depth map; The fused depth map is back-projected using the camera relative extrinsic parameters of the keyframes to obtain a dense point cloud; The sparse point cloud and the dense point cloud are merged to obtain an initial facial model point cloud containing the driver.

7. The method according to claim 1, characterized in that, After identifying the driver's actual driving state through the driving state features, the method further includes: If the actual driving condition is slightly abnormal, an audible prompt will be given; If the actual driving condition is moderately abnormal, then an audible and vibration alert will be given. If the actual driving condition is severely abnormal, the vehicle will be slowed down and pulled over to the side of the road, and a distress message will be sent.

8. A driving state recognition device, characterized in that, The device includes: The acquisition module is used to acquire 2D image frames containing the driver's face; The first recognition module is used to identify key frames in the 2D image frame and generate an initial facial model point cloud of the driver based on the key frames. The generation module is used to generate an initial 3D Gaussian ellipsoid set from the initial facial model point cloud; The comparison module is used to perform planar projection on the initial 3D Gaussian ellipsoid set to obtain a 2D rendered image, and compare the 2D rendered image with the key frame to obtain the comparison difference result; The optimization module is used to optimize the initial 3D Gaussian ellipsoid set based on the comparison difference results to obtain a point cloud of the target 3D facial model of the driver. The second recognition module is used to recognize the driver's driving state features through the point cloud of the target 3D facial model, and to recognize the driver's actual driving state through the driving state features.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the driving state recognition method as described in any one of claims 1 to 7.

10. A vehicle, characterized in that, include: The driving state recognition device according to claim 8.