Cabin feature fusion binaural tracking method based on layered depth estimation
By using a method based on layered depth estimation and cockpit feature fusion, monocular camera and Kalman filtering technology, the existing ear positioning methods are solved, and high precision and stable ear positioning in complex environments are achieved.
Patent Information
- Application Number
- CN202510458242.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-26
AI Technical Summary
The existing ear positioning methods rely on multi-eye vision or expensive depth sensors, resulting in high cost, complex deployment, and poor robustness under conditions such as complex lighting in the cockpit and diverse occupants' postures. They ignore visual features such as inherent geometric background in the vehicle cockpit, limiting the improvement of accuracy.
The cockpit feature fusion binaural tracking method based on layered depth estimation is adopted to collect video streams through a monocular camera to generate multi-scale depth maps. Combining the edge features of the fixed color area of the cockpit background and the geometric proportional relationship between the face key points, the ear three-dimensional coordinates are derived using the PnP algorithm, and real-time smooth tracking is performed through Kalman filtering to reduce the impact of noise.
It realizes low-cost, high-precision and high-root three-dimensional positioning of the ear, suitable for complex lighting and occupant posture changes scenarios, reducing hardware costs and improving real-time and stability of positioning.
Smart Images

Figure CN120543633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent cockpit perception method, and in particular to an occupant binaural targeted tracking method based on multi-scale depth estimation and vehicle cockpit characteristics. The method can be widely used in vehicle-mounted human-computer interaction systems, driver status monitoring systems, in-vehicle sound field control, and active noise control systems. Background Art
[0002] With the rapid development of intelligent vehicle technology, in-vehicle human-computer interaction has placed higher demands on user experience. Ear positioning and tracking, as an important module for in-vehicle multimodal perception, has a wide range of applications in areas such as in-vehicle voice pickup, active noise reduction, and driver status monitoring. However, existing ear positioning methods mainly rely on multi-viewing or expensive depth sensors, resulting in high costs and complex deployment. At the same time, they are less robust under complex lighting conditions and diverse occupant postures in the cabin. In addition, these methods often ignore the inherent constraints of visual features such as the inherent geometric background in the vehicle cabin, and rarely consider the actual vehicle application scenario information, which also limits the improvement of accuracy. Summary of the Invention
[0003] Purpose of the Invention: To overcome the shortcomings of existing technologies, this invention proposes a binaural tracking method based on hierarchical depth estimation and cabin feature fusion. By integrating depth estimation, facial geometry, and cabin background geometry constraints, this method achieves low-cost, high-precision, and highly robust 3D ear localization.
[0004] Technical solution: The cockpit feature fusion binaural tracking method based on layered depth estimation includes the following steps:
[0005] S1. A monocular camera is used to capture a video stream of the occupants and generate a multi-scale depth map. The multi-scale depth estimation uses the CabinDepth model to construct a scale pyramid. The input image is gradually downsampled to generate feature maps at multiple resolution levels. The multi-scale network outputs a depth map as follows:
[0006] D=f(I)
[0007] Where I represents the input image and f represents the estimation function of CabinDepth.
[0008] S2. Extract edge features of fixed color areas in the cockpit background and construct geometric background constraints. The specific method is to use the Canny edge detection algorithm, set thresholds T1 and T2 to filter low-texture areas in the background, and extract clear edge information.
[0009] S3. Detect facial key points using a multi-feature cascade joint detector, including the corners of the eyes, the tip of the nose, and the corners of the mouth, and calculate their geometric proportion relationship; the geometric proportion relationship is calculated using the following formula:
[0010]
[0011] Among them, P i , P j is the two-dimensional coordinate of the key point, P i -P j It represents the Euclidean distance between two points, and P1-P2 is the reference distance (such as interocular distance).
[0012] S4. Combining the geometric model of facial key points with the depth map, the PnP algorithm is used to derive the three-dimensional coordinates of the ear. The equation solved by the PnP algorithm is:
[0013] s·M=K·[R|t]·W
[0014] Where s is the scaling factor, M is the pixel coordinate, K is the camera intrinsic parameter matrix, R|t is the rotation and translation matrix, and W is the three-dimensional space point.
[0015] S5. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. The Euler angle parameter of the head posture estimation is θ x ,θ y ,θ z , the Kalman filter update formula is as follows:
[0016] X t =A·X t-1 +B·U+w t Z t =H·X t +v t
[0017] where X t is the ear position state vector, Z t is the observation quantity, A is the state transfer matrix, H is the observation matrix, w t , v t are process noise and observation noise, respectively.
[0018] Furthermore, in the step S1, a multi-scale depth map is generated by a CabinDepth model, and the resolution levels of its scale pyramid include 1 / 2, 1 / 4, and 1 / 8 of the original image resolution, so as to balance detail preservation and computational efficiency.
[0019] Furthermore, in step S2, the fixed color edge features of the cockpit background are extracted using the Canny algorithm based on the depth map output by DpethFM, and the threshold ranges T1 and T2 are set.
[0020] Furthermore, in the S3 step, facial key point detection uses the Haar feature classifier and the Dlib key point detector for preliminary detection, and combines the MTCNN algorithm to accurately extract and optimize the feature points. The three-level detection network P-Net, R-Net, and O-Net extracts key points and ensures that the key point detection accuracy reaches a certain level.
[0021] Furthermore, in step S4, the ear localization accuracy is improved by a multimodal weighting strategy that integrates background features, facial geometric features, and depth information; the weighting formula is:
[0022] P ear =ω1P deapth +ω2P face +ω3P bg
[0023] Where ω1, ω2, ω3 are the weights of depth information, facial geometric features and background constraints, respectively, satisfying:
[0024] ω1+ω2+ω3=1
[0025] Furthermore, in the step S5, the head posture is estimated by detecting the two-dimensional image coordinates M of the facial key points. 2D With the standard 3D model point M 3D The matching relationship between them is solved by using the PnP (Perspective-n-Point) algorithm to solve the rotation matrix R of the head. pose and the translation vector T pose , the formula is as follows:
[0026] [R pose , T pose ]=PnP(M 3D , M 2D , K)
[0027] Among them, M 3D are predefined standard 3D model points, including key points such as the corners of the eyes, the tip of the nose, and the corners of the mouth; M 2D The two-dimensional image points detected by the multi-feature cascade joint detector, K is the intrinsic parameter matrix of the camera;
[0028] Rotate the head by angle R pose Convert to Euler angles (α, β, γ) to represent the pitch, yaw, and roll angles of the head;
[0029] Based on the results of head posture estimation, the rotation matrix R is solved by combining the PnP algorithm pose and the translation vector T pose , the initial three-dimensional coordinates of the ear Perform dynamic compensation;
[0030] In order to further smooth the time series data of ear position and reduce the impact of noise on the positioning results, the Kalman filter algorithm is introduced to track the position of the occupant's two ears in real time.
[0031] Compared with the prior art, the present invention has the following advantages:
[0032] (1) Efficient Depth Estimation: This paper introduces a multi-scale depth estimation method based on the CabinDepth model. This method uses a pyramid-structured downsampling method to improve the depth map's resolution and edge feature preservation in complex backgrounds. Compared to traditional single-scale depth estimation models, this method significantly improves computational efficiency and meets real-time requirements.
[0033] (2) Robustness of ear localization: This method integrates the visual features of the cabin seat with the geometric key points of the occupant's face, integrates the depth estimation results, and constructs a multimodal feature weighted model. High-precision ear localization can be achieved under complex lighting conditions (such as alternating light and dark scenes), background interference (such as reflections or texture changes), and posture changes (such as head deflection or tilt).
[0034] (3) Practicality of hardware deployment: Using a monocular camera for depth estimation avoids the high cost and complex deployment of multi-camera vision systems or laser sensors. This method not only reduces hardware costs but can also be directly integrated into existing vehicle camera systems, making it suitable for smart cockpit scenarios.
[0035] (4) Real-time tracking stability: By introducing head posture estimation compensation and Kalman filtering algorithms, the present invention can effectively smooth the time series data of the monocular camera's ear positioning, reduce jitter, and improve tracking stability and real-time performance. Even in the case of rapid occupant movement or drastic posture changes, the positioning error remains within an acceptable range. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Flow chart of the method of the present invention;
[0037] Figure 2 Layout diagram of occupant binaural positioning and tracking test;
[0038] Figure 3 correspond Figure 2 Original image, multi-scale depth map estimated by the CabinDepth model, including: (a): occupant head pose forward depth map (without hat) (b): occupant head pose side-turn depth map (without hat) (c): occupant head pose forward depth map (with hat) (d): occupant head pose side-turn depth map (with hat);
[0039] Figure 4Logitech C920Pro monocular camera calibration test internal and external parameter matrix, where (a) is the camera external parameter matrix with the camera origin fixed (b) is the camera external parameter matrix with the calibration plate origin fixed;
[0040] Figure 5 Reprojection error of the Logitech C920Pro monocular camera calibration results;
[0041] Figure 6 The binaural 3D spatial region bounding box of the occupant after head pose compensation and Kalman filtering smoothing. DETAILED DESCRIPTION
[0042] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0043] The test embodiment adopts a simplified structure of the main driver's part of the vehicle cockpit, arranges and builds the test bench and position, including a simplified car seat (the rear part contains the visual features of the seat leather used in the model) and the driver's seat, focusing on the visual tracking detection effect to carry out verification, as shown in the attached figure. Figure 2 As shown in the figure, a cockpit feature fusion binaural tracking method based on layered depth estimation is built, which includes the following steps:
[0044] The cockpit feature fusion binaural tracking method based on layered depth estimation of this embodiment includes the following steps:
[0045] S1. A monocular camera is used to capture a video stream of the occupants and generate a multi-scale depth map. The multi-scale depth estimation uses the CabinDepth model to construct a scale pyramid. The input image is gradually downsampled to generate feature maps at multiple resolution levels. The multi-scale network outputs a depth map as follows:
[0046] D=f(I)
[0047] Where I represents the input image and f represents the estimation function of CabinDepth.
[0048] Specifically, the CabinDepth model is used for multi-scale depth estimation. Its core method is based on flow matching, and pyramid downsampling and auxiliary normal loss (Surface Normal Loss) are used to improve the resolution and accuracy of the depth map. The steps include:
[0049] S11. Input the occupant video frame I captured by the monocular camera and convert the input into latent space features z through the encoder of the CabinDepth model.
[0050] S12, depth estimation is optimized by the following flow matching objective function:
[0051]
[0052] where φ t (x0) is the latent feature path at time step t, and x0 and x1 are the latent codes of the input image and target depth.
[0053] S13. To further improve the geometric accuracy of depth estimation, normal loss is introduced:
[0054]
[0055] Where n is the true normal, is the estimated depth-derived normal, M(n) is the confidence mask, and π(t) is the loss weight function.
[0056] S 14. Through multi-scale pyramid output, the CabinDepth model performs pyramid downsampling on the input image to generate a multi-resolution depth map:
[0057] D k =f(I k ), k∈{1, 1 / 2, 1 / 4, 1 / 8}
[0058] Among them I k are input images at different resolutions.
[0059] S2. Extract edge features of the fixed color area of the cabin background and construct geometric background constraints to improve the robustness of ear position calculation. The steps include:
[0060] S21, use the Canny algorithm to extract the background edge, set the high and low thresholds T1 and T2, and filter the low texture area;
[0061] S22. The extracted edge information constructs geometric constraints for a fixed background area to reduce the impact of illumination changes and background interference on ear position calculation.
[0062] S3. Detect facial key points using a multi-feature cascade joint detector, including the corners of the eyes, the tip of the nose, and the corners of the mouth, and calculate their geometric proportions. The geometric proportions are calculated using the following formula:
[0063]
[0064] Among them, P i , P j is the two-dimensional coordinate of the key point, P i -P jIt represents the Euclidean distance between two points, and P1-P2 is the reference distance (such as interocular distance).
[0065] S31 and Haar cascade classifiers are used to quickly detect occupant facial regions, and possible occupant facial regions are extracted from grayscale images through a pre-trained model.
[0066] S32. Based on the face region defined by the Haar classifier, Dlib performs preliminary facial landmark location. Dlib uses a 68-point model to initially calibrate facial landmarks, including the outlines of facial features such as the eyes, nose, and mouth.
[0067] S33. Considering that Haar detection is fast but significantly affected by changes in illumination and angle, resulting in high rates of false detection and missed detection, the MTCNN algorithm is used to optimize and correct the key points detected by Dlib to further improve detection accuracy. The MTCNN algorithm consists of three layers: P-Net, R-Net, and O-Net. Each layer of the network gradually refines the location of the occupant's facial key points. The core loss function of the MTCNN network consists of the following three parts:
[0068] (1) Calculate the classification loss to detect whether it is the passenger's facial area:
[0069]
[0070] Where yi is the label of the sample (1 is a face, 0 is a non-face), is the network prediction value.
[0071] (2) Calculate the bounding box regression loss and adjust the detection box:
[0072]
[0073] t i is the real frame coordinate, Predicted box coordinates.
[0074] (3) Calculate key point loss and optimize feature point positions:
[0075]
[0076] in To predict the key point coordinates, The real key point coordinates, K is the number of key points.
[0077] S34. By combining the optimization results of the Haar classifier, Dlib detector and MTCNN, the final positions of 68 feature points are generated.
[0078] S35. From the optimized 68 facial feature points, select key points related to ear position derivation. Using these feature points as input, preliminarily calculate the two-dimensional pixel position of the occupant's ear. The specific calculation is as follows:
[0079] Ear estimation based on eye corner position:
[0080]
[0081] P eye-left and P eye-right is the left and right eye coordinates, d offse tThe horizontal offset of the ear relative to the eye corner. According to statistical data, the offset of the passenger in this example is 15 pixels.
[0082] Ear estimation based on jaw and mouth corner positions:
[0083] P ear2 =k·(P jaw -P mouth )+b
[0084] Among them, P jaw is the position matrix of the mandibular feature points including left and right 2 and 14, P mouth is the position matrix of the left and right mouth corner feature points, k is the scale factor matrix, and b is the offset vector representing the basic position deviation of the ear, which depends on the occupant's facial size and posture (calculated based on the pre-captured video stream).
[0085] S4. Combining the geometric model of facial key points with the depth map, the PnP algorithm is used to derive the three-dimensional coordinates of the ear. The equation solved by the PnP algorithm is:
[0086] s·M=K·[R|t]·w
[0087] Where s is the scaling factor, M is the pixel coordinate, K is the camera intrinsic parameter matrix, R|t is the rotation and translation matrix, and W is the three-dimensional space point.
[0088] The ear localization accuracy is improved by integrating background features, facial geometry features and depth information using a multimodal weighting strategy. The weighting formula is:
[0089] P ear =ω1P deapth +ω2P face +ω3P bg
[0090] Where ω1, ω2, ω3 are the weights of depth information, facial geometric features and background constraints, respectively, satisfying ω1+ω2+ω3=1
[0091] In this example, ω1, ω2, ω3 are set to [0.28, 0.47, 0.25] after optimization calculation.
[0092] S5. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. The steps are as follows:
[0093] S51, first perform head pose estimation, head pose estimation by detecting the two-dimensional image coordinates M of facial key points 2D With the standard 3D model point M 3D The matching relationship between them is solved by using the PnP (Perspective-n-Point) algorithm to solve the rotation matrix R of the head. pose and the translation vector T pose , the formula is as follows:
[0094] [R pose , T pose ]=PnP(M 3D , M 2D , K)
[0095] Among them, M 3D are predefined standard 3D model points, including key points such as the corners of the eyes, the tip of the nose, and the corners of the mouth; M 2D The two-dimensional image points detected by the multi-feature cascade joint detector, K is the intrinsic parameter matrix of the camera.
[0096] S52, rotate the head by an angle R pose Converted into Euler angles (α, β, γ), representing the pitch angle (Pitch), yaw angle (Yaw) and roll angle (Roll) of the head.
[0097] S53, based on the results of head posture estimation, combined with the PnP algorithm to solve the rotation matrix R pose and the translation vector T pose , the initial three-dimensional coordinates of the ear Perform dynamic compensation.
[0098] S54. To further smooth the time series data of ear positions and reduce the impact of noise on the positioning results, the present invention introduces a Kalman filter algorithm to track the position of the occupant's ears in real time. The Euler angle parameter of the head posture estimation is θ x ,θ y ,θ z , the Kalman filter update formula is as follows:
[0099] X t =A·X t-1 +B·U+w t Z t =H·X t +v t
[0100] where X t is the ear position state vector, Z t is the observation quantity, A is the state transfer matrix, H is the observation matrix, w t , v t are process noise and observation noise, respectively.
[0101] The Kalman filter process is divided into the following two steps:
[0102] (1) State prediction:
[0103] X t|t-1 =A·X t-1|t-1 +B·U t
[0104] (2) Status update:
[0105] X t|t =X t|t-1 +K t ·(Z t -H·X t|t-1 )
[0106] (3) Kalman gain calculation:
[0107] K t =P t|t-1 ·H T ·(H·P t|t-1 ·H T +R) -1
[0108] The current X t Ear position state vector, Z t Through the observation value after posture compensation, A is the state transfer matrix, which describes the dynamic change of ear position, H is the observation matrix, which is used to map the state vector to the observation space, K t is the Kalman gain, which controls the balance between the predicted value and the observed value, P t|t-1 The covariance matrix of the state prediction, R is the covariance matrix of the observation noise.
[0109] By combining state prediction with observation updates, Kalman filtering effectively reduces jitter and noise interference in the time series of ear position. After dynamic compensation and Kalman filtering are completed, the occupant's ear positions are updated in real time at a rate of 30 frames per second, and the final three-dimensional spatial coordinates of the ears are output. This real-time performance ensures the applicability of this invention in complex lighting conditions and scenarios with rapid posture changes.
[0110] S55. Select the active control area of the left and right ears of the occupant, set the facial features of the ears with the same x, y but different depths as the starting point of the selection, extend the selection space rectangular area, and set the size to 5×10×10 (cm).
[0111] The embodiments of the present invention are not limited to the above one. Without violating the spirit and scope of the present invention, those skilled in the art may make various changes and modifications to the present invention, but these changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
[0112] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. A binaural tracking method for cockpit features fusion based on layered depth estimation, characterized in that: The following steps are involved: S1. A monocular camera is used to capture a video stream of the occupants and generate a multi-scale depth map. The multi-scale depth estimation uses the CabinDepth model to construct a scale pyramid. The input image is gradually downsampled to generate feature maps at multiple resolution levels. The multi-scale network outputs a depth map as follows: D=f(I) Where I represents the input image, f represents the estimation function of CabinDepth; S2. Extract edge features of the fixed color area of the cabin background and construct geometric background constraints; The specific method is to use the Canny edge detection algorithm to set thresholds T1 and T2 to filter the low-texture area in the background and extract clear edge information; S3. Detect facial key points using a multi-feature cascade joint detector and calculate their geometric proportion relationship; the geometric proportion relationship is calculated using the following formula: Among them, P i , P j is the two-dimensional coordinate of the key point, P i -P j Represents the Euclidean distance between two points, P1-P2 is the reference distance; S4. Combining the geometric model of facial key points with the depth map, the PnP algorithm is used to derive the three-dimensional coordinates of the ear. The equation solved by the PnP algorithm is: s·M=K·[R|t]·W Where s is the scaling factor, M is the pixel coordinate, K is the camera intrinsic parameter matrix, R|t is the rotation and translation matrix, and W is the three-dimensional space point; S5. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. Dynamically compensate the ear position by estimating the head posture, and use Kalman filtering for real-time smooth tracking. The Euler angle parameter of the head posture estimation is θ x ,θ y ,θ z , the Kalman filter update formula is as follows: X t =A·X t-1 +B·U+w t Z t =H·X t +v t where X t is the ear position state vector, Z t is the observation quantity, A is the state transfer matrix, H is the observation matrix, w t , v t are process noise and observation noise, respectively.
2. The binaural tracking method for cockpit feature fusion based on layered depth estimation according to claim 1, characterized in that: In the step S1, a multi-scale depth map is generated by the CabinDepth model, and the resolution levels of its scale pyramid include 1 / 2, 1 / 4, and 1 / 8 of the original image resolution to balance detail preservation and computational efficiency.
3. The binaural tracking method for cockpit feature fusion based on layered depth estimation according to claim 2, characterized in that: In step S2, the fixed color edge features of the cockpit background are extracted based on the depth map output by DpethFM using the Canny algorithm, and the threshold ranges T1 and T2 are set. The Canny algorithm is used to suppress noise and enhance edges.
4. The binaural tracking method for cockpit feature fusion based on layered depth estimation according to claim 3, characterized in that: In the S3 step, facial key point detection uses the Haar feature classifier and the Dlib key point detector for preliminary detection, and combines the MTCNN algorithm to accurately extract and optimize feature points. The three-level detection network P-Net, R-Net, and O-Net extract key points and ensure that the key point detection accuracy reaches a certain level.
5. The binaural tracking method for cockpit feature fusion based on layered depth estimation according to claim 4, characterized in that: In the step S4, the ear positioning accuracy is improved by fusing background features, facial geometric features and depth information using a multimodal weighting strategy; The weighted formula is: P ear =ω1P deapth +ω2P face +ω3P bg Where ω1, ω2, ω3 are the weights of depth information, facial geometric features and background constraints, respectively, satisfying: ω1+ω2+ω3=1.
6. The binaural tracking method for cockpit feature fusion based on layered depth estimation according to claim 5, characterized in that: In the step S5, the head posture is estimated by detecting the two-dimensional image coordinates M of the facial key points. 2D With the standard 3D model point M 3D The matching relationship between them is solved by using the PnP (Perspective-n-Point) algorithm to solve the rotation matrix R of the head. pose and the translation vector T pose , the formula is as follows: [R pose ,T pose ]=PnP(M 3D ,M 2D ,K) Among them, M 3D These are predefined standard 3D model points, including key points such as the corners of the eyes, the tip of the nose, and the corners of the mouth; M 2D The two-dimensional image points detected by the multi-feature cascade joint detector, K is the intrinsic parameter matrix of the camera; Rotate the head by angle R pose Convert to Euler angles (α, β, γ) to represent the pitch, yaw, and roll angles of the head; Based on the results of head posture estimation, the rotation matrix R is solved by combining the PnP algorithm pose and the translation vector T pose , the initial three-dimensional coordinates of the ear Perform dynamic compensation; In order to further smooth the time series data of ear position and reduce the impact of noise on the positioning results, the Kalman filter algorithm is introduced to track the position of the occupant's two ears in real time.