Ear hole positioning and tracking method based on depth camera profile scanning

By combining depth camera contour scanning with point cloud registration and filter tracking algorithms, the accuracy and privacy issues of ear hole positioning in active noise-canceling headrests when the head moves are solved, achieving efficient and accurate ear hole positioning and tracking.

CN122134807APending Publication Date: 2026-06-02NANJING UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-02-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing active noise-canceling headrests suffer from insufficient accuracy and privacy risks in ear canal positioning methods when the head moves, especially during translation, rotation, and short-term obstruction. Furthermore, existing technologies are not convenient for users to use.

Method used

A depth camera-based contour scanning method is adopted, combined with a tracking algorithm that integrates point cloud registration, kernel correlation filter and Kalman filter weighted fusion, to obtain the real-time position coordinates of the ear canal. Through depth image processing and feature extraction, efficient localization and tracking of the ear and ear canal are achieved.

Benefits of technology

The system achieves high-precision positioning and tracking of the ear canal even under conditions of head translation, rotation, and short-term occlusion, thus avoiding facial privacy leaks. The system has a simple structure and is highly practical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134807A_ABST
    Figure CN122134807A_ABST
Patent Text Reader

Abstract

This invention discloses a method for ear canal localization and tracking based on depth camera contour scanning. The method includes: first, acquiring a depth image of the side of the face using a depth camera, marking the ear region and ear canal position in the image, and converting the depth image into a 3D point cloud, which is then saved as a template point cloud; the user maintains the same position and posture as when the template was acquired to obtain the side target point cloud; then, registering the template point cloud with the target point cloud; finally, based on the ear position features and ear canal coordinates determined by the registration and localization, continuous tracking is performed using a tracking algorithm. This invention features a simple system structure and strong practicality, effectively tracking both translation and rotation of the target head in three-dimensional space while avoiding privacy leaks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of active noise control, and its application purpose is to accurately locate the ear canals of an active noise-canceling headrest. Background Technology

[0002] Local active noise control (ANC) systems can create a relatively quiet area in noisy environments. However, the range of this quiet area is usually limited, especially for single-channel control systems, where the diameter of the 10dB quiet zone typically does not exceed one-tenth of the wavelength of the sound wave in the noise field. Active noise-canceling headrests are local ANC systems that place secondary source speakers near the head to create a quiet area around the ears. They can be applied in various settings such as cars, airplanes, and homes. When an individual's head moves, the ears may deviate from this quiet area, leading to a significant reduction in noise suppression effectiveness. Therefore, real-time monitoring of the position of both ears becomes crucial. By switching different control filters in a timely manner, the position of the quiet zone can be adjusted to ensure that both ears remain within the quiet area even when the head moves or rotates, thus maintaining ideal noise suppression.

[0003] To improve the robustness of active noise-canceling headrests in situations involving head movement, existing auxiliary active noise-canceling headrest ear canal localization technologies mainly include the following:

[0004] The first method is head positioning based on infrared sensors. A thermal imaging module using infrared technology is installed on the aircraft seat to determine the position of the passenger's head. However, the positioning accuracy achieved through temperature sensing is sensitive to the environment and limited by the sensor's image resolution.

[0005] The second method utilizes wearable devices to assist in positioning. This includes: (1) calculating the movement trajectory of the head and ears by wearing two headphones; (2) attaching two microphones to a hat or shoulder strap so that the error microphone can move with the head movement; and (3) using a laser Doppler vibrometer and a lightweight reflective chip near the ear canal to determine the sound in the ear in real time. However, the above three technologies require users to wear additional items such as chips, headphones, or hats, which is inconvenient in practical applications.

[0006] The third method uses a camera to track head position. Using a Kinect camera to track head position, the obtained head position information is used to update the control and observation filters in the remote sensing technology. This research focuses on sound frequencies below 1 kHz and can only be used for head translation, not head rotation. Another method uses a combination of cameras and depth cameras to track the human ear in three-dimensional space. However, the color camera and computer vision color image processing technology used in this method pose a risk of user facial privacy breaches.

[0007] In summary, existing binaural localization methods used to improve the robustness of active noise-canceling headrests under head movement conditions have certain limitations. They primarily focus on changes in the position of the ears and ear canals during head translation, paying less attention to changes in the position of the ears and ear canals during head rotation or even short-term occlusion. Therefore, designing a simple and practical binaural localization method for active noise-canceling headrest systems that can address head translation, rotation, and short-term occlusion scenarios is of great significance. Summary of the Invention

[0008] To address the problems existing in the prior art, this invention provides a method for ear canal localization and tracking based on depth camera contour scanning. This method can obtain the position coordinates of the ear canal in real time under conditions of head translation, rotation, and short-term occlusion.

[0009] To achieve the above-mentioned objectives, the present invention employs the following technical solution: a method for ear canal localization and tracking based on depth camera contour scanning, comprising the following steps: Step S1, using a depth camera to acquire depth contour information of the user's side face in real time and outputting it as a depth image; Step S2, enhancing the depth image, marking the ear region and ear canal position in the image, thereby generating a side template point cloud and saving it; Step S3, the user maintains the same position and posture as when the image was acquired in Step S1, and remains still, using a depth camera to acquire depth contour information and converting it into a side target point cloud; Step S4, combining the side template point cloud with... The side target point cloud is registered to achieve initial localization of the ear and ear canal; Step S5: Based on the registration result of Step S4, ear features are extracted and a tracker with weighted fusion of kernel correlation filter and Kalman filter is initialized. The tracking algorithm continuously tracks the ear movement and ear canal position changes; Step S6: If the target is lost during tracking, the user is prompted that the tracking has failed and the tracking is paused; After a 3-second pause, the re-detection stage is entered, and ear features are detected in the global field of view of the depth camera for re-localization; Step S7: If the re-detection is unsuccessful, Step S6 is repeated; Step S8: When localization and tracking are successful, the coordinates of the ear canal in three-dimensional space are output in real time.

[0010] Compared with existing technologies, this invention proposes a real-time ear canal localization and tracking method based on depth camera contour scanning technology, combined with point cloud registration, kernel correlation filter, and Kalman filter weighted fusion tracking. This method features simple structure, strong practicality, and no leakage of facial privacy. It can efficiently and accurately locate and track ear canals under various conditions, including head translation, rotation, and short-term ear occlusion, demonstrating excellent performance. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the method of the present invention.

[0012] Figure 2 This is a schematic diagram of an experimental device for locating and tracking the ear canal of an artificial head model using a depth camera in an embodiment of the present invention, wherein: (a) is a top view and (b) is a rear view.

[0013] Figure 3 This is a schematic diagram of the spatial sampling point division and experimental device for ear hole positioning and tracking in an embodiment of the present invention.

[0014] Figure 4 This is a grid diagram showing the spatial sampling points for ear canal positioning and tracking in an embodiment of the present invention.

[0015] Figure 5 The images show the results of the adaptive foreground and background segmentation algorithm, where: (a) the original depth map, and (b) the processed result.

[0016] Figure 6 The images show the results of histogram equalization and sharpening, including: (a) the adaptive foreground and background segmentation image, and (b) the processing result.

[0017] Figure 7 Example diagrams of spatial filtering with threshold T of 3 and L of 5; (a) pixel distribution before spatial filtering, (b) pixel distribution after spatial filtering.

[0018] Figure 8 The images show the spatial filtering results, where (a) and (c) are the original head depth images taken from different angles, and (b) and (d) are the spatial filtering results of (a) and (c), respectively.

[0019] Figure 9 The results of ear localization are coarse registration with random sampling consistency and fine registration with iterative nearest point. (a), (b), (c), (d) and (e) are the ear localization results of the left side of the face when the head is tilted to the left by 60°, 30°, no tilt, 30° and 60° respectively.

[0020] Figure 10 The images show the visual positioning and tracking effect of the left ear canal during the artificial head rotation process (a), (b), and (c) in the embodiments of the present invention.

[0021] Figure 11 This is a comparison between the position of the left ear canal and the actual position during the translation and rotation of the artificial head in this embodiment of the invention, where (a) represents translation and (b) represents rotation.

[0022] Figure 12 This is a schematic diagram (rear view) of the experimental device for locating and tracking the ear canal of a real person's left side face using a depth camera, as described in an embodiment of the present invention.

[0023] Figure 13The image shows the visual positioning and tracking effect of the depth camera on the left ear canal during the translation and rotation of a real person's head in an embodiment of the present invention. In the image, (a) is the initial detection, (b) and (c) are the tracking results of the head moving forward and backward, (d) and (e) are the tracking results of the head moving left and right, (f) and (g) are the tracking results of the head tilting, and (h) and (i) are the tracking results of the head rotating left and right. Detailed Implementation

[0024] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only used to illustrate the solution of the present invention and are not intended to limit its scope. Based on the disclosure of the present invention, any equivalent modifications made by those skilled in the art should fall within the scope defined by the claims of this application.

[0025] This invention provides a method for ear canal localization and tracking based on depth camera contour scanning. For a saved template of the target object, the subsequent localization and tracking process only requires acquiring its depth image and converting it into a 3D point cloud. Since only depth information is used, the privacy risks associated with RGB images are completely avoided. The ear canal registration, localization, and real-time tracking stages include: coarse registration using the Random Sample Consensus Algorithm (RANSAC), followed by fine registration using the Iterative Closest Point Algorithm (ICP), aligning the target point cloud with the template point cloud to determine the precise location of the ear and ear canal; initialization using a kernelized correlation filter (KCF) based on the registered ear position, and preliminary tracking using histogram of oriented gradients (HOG) features; and subsequent weighted fusion of the tracking result with a Kalman filter tracking method to form a more stable tracking mechanism with re-detection capabilities.

[0026] like Figure 1 As shown, the workflow for ear canal localization and tracking based on a depth camera and combined with registration and positioning algorithms and tracking algorithms in this embodiment is as follows:

[0027] Step S1: A depth camera is used to acquire the depth contour information of the user's side profile in real time and output it as a depth image. During depth camera acquisition, the user keeps their head upright, ensuring their ears are at the same horizontal level as the depth camera. The depth camera is then activated to acquire a real-time depth image of the target's side profile. The depth camera used in this embodiment is manufactured by ORBBEC, and its model is Gemini 335.

[0028] Step S2: Enhance the depth image, mark the ear region and ear canal position in the image, generate a side template point cloud based on this and save it, then turn off the depth camera.

[0029] The depth image enhancement process sequentially employs adaptive depth thresholding foreground and background segmentation, histogram equalization, image sharpening, and spatial filtering image processing algorithms. Each algorithm is described below:

[0030] Adaptive depth thresholding for foreground and background segmentation first requires mean normalization of the grayscale values ​​in the depth image. Then, the result of the Sigmoid function calculation is weighted and fused with the grayscale standard deviation to generate the background segmentation threshold. Specifically:

[0031]

[0032] In the formula I i Let μ be the grayscale value at pixel i. non To normalize the grayscale mean, μ non The weight μ is obtained by inputting the sigmoid function. w

[0033]

[0034] To more reasonably segment the background depth image, the mapped mean and standard deviation are weighted and combined, μ w The weight is set to 0.8, and σ is the standard deviation of the grayscale values, with a weight of 0.2.

[0035]

[0036] Here, k and b are the mean coefficient and offset in the sigmoid function, respectively. By setting different values ​​of k and b, foreground and background segmentation can be achieved. This adaptive depth thresholding segmentation method based on sigmoid function weighting, compared to existing techniques that typically use a fixed depth threshold or simple Otsu thresholding for background removal, utilizes the mean and standard deviation of image grayscale values ​​to dynamically calculate the weight μ using the sigmoid function. w This allows the algorithm to more gently and adaptively segment the foreground (side profile) and background based on the facial depth distribution characteristics of different users, reducing the "one-size-fits-all" error caused by a fixed threshold.

[0037] Histogram equalization calculates the probability density function of grayscale distribution based on the normalized histogram by statistically analyzing the pixel distribution of grayscale values, and then uses the cumulative distribution function to map the original grayscale values ​​to new grayscale values, thereby achieving uniform grayscale distribution.

[0038] Image sharpening works by calculating the difference between a pixel and its surrounding pixels, amplifying areas of rapid grayscale changes while suppressing areas of gradual change. This requires designing a suitable sharpening convolution kernel K.

[0039] (4)

[0040] The parameter 'a' controls the sharpening intensity, the center value of K is positive to maintain the value of the current pixel, and the surrounding negative values ​​are used to reduce the influence of surrounding pixels, thereby enhancing the edges.

[0041] A convolution operation is performed on the image I(x, y) and the convolution kernel K to obtain the sharpened image I′(x, y).

[0042]

[0043] The spatial filtering process requires first determining a zero pixel I and creating a filtering window centered on the candidate pixel. Then, the number of non-zero pixels in the filtering window is counted, and the size of the filtering window is L. 2 For each pixel, when the number of pixels counted exceeds the user-defined threshold T, the value of the candidate pixel will become the average value V of the non-zero pixels; otherwise, it will remain unchanged.

[0044] The annotation was done manually, and the saved side template point cloud was used for subsequent registration.

[0045] The depth image with labeled ear region and ear canal location is converted into a 3D point cloud and saved as a side template point cloud. Then the depth camera is turned off.

[0046] Step S3: Restart the depth camera. The user maintains the same position and posture as when the image was acquired in Step S1 and remains still. Use the depth camera to acquire depth contour information and convert it into a side target point cloud. Generate a template after remaining still for 15 to 30 seconds. No further calibration is required.

[0047] Step S4: Register the side template point cloud saved in step S2 with the side target point cloud in step S3 to achieve the initial positioning of the ear and ear canal.

[0048] This step first employs a coarse registration method based on random sampling consistency, followed by fine registration using an iterative nearest-point algorithm to align the template point cloud with the target point cloud. Based on this, the positions of the ear and ear canal within the target point cloud are calculated, completing the initial localization. Specifically:

[0049] In point cloud registration, random uniform consistency sampling coarse registration and iterative nearest point fine registration are two complementary algorithms. The goal of coarse registration is to provide an initial transformation that roughly aligns the source point cloud with the target point cloud. Fine registration optimizes the transformation based on this to improve the registration accuracy.

[0050] In this embodiment, the parameters for coarse and fine registration are set as follows: number of samples N. max = 40000 iterations, interior point threshold ϵ = 10, minimum interior point ratio τ = 50%, maximum number of iterations K for iterating to the nearest point max = 50, convergence threshold δ < 10 -6 Then, four point pairs are randomly selected from the matching point pair set M. First, the fast point feature histogram feature descriptors of the template point cloud and the target point cloud are calculated, and a feature matching point pair set is established between the two point clouds. Secondly, for the three-dimensional rigid transformation, random uniform sampling randomly selects four pairs of points from the template point cloud and the target point cloud each time to determine the rotation matrix R and translation vector t, and uses the least squares method to calculate the rigid transformation matrix T(R, t); for the randomly sampled point pairs... The goal of solving the rigid transformation matrix T is to find the optimal R and t that minimize the error function.

[0051]

[0052] Finally, calculate the distance between the template point and the target point. If the distance... If the number of interior points in the current model is greater than that in the previous model, then the current transformation T is recorded as the optimal transformation. The above steps are repeated until the maximum number of iterations is reached or the number of interior points meets a certain threshold.

[0053] Iterative nearest-point fine registration uses random, uniform, and consistent sampling of the coarse registration result T0(R0, t) o Starting from this point, for each point p in the template point cloud... k Find the nearest point q in the target point cloud. k

[0054]

[0055] Based on point pairs (p k , q k Singular value decomposition (SVD) can decompose any m × n matrix A into the form of equation (8) to solve for T(R, t) with the minimum error. The calculation method is as follows:

[0056] (8)

[0057] Where U is an m × n orthogonal matrix, called the left singular matrix; Σ is an m × n diagonal matrix whose diagonal elements are non-negative real numbers, called singular values; and V is an n × n orthogonal matrix, called the right singular matrix.

[0058] After obtaining the new T(R, t), apply it to the template point cloud, update the position of the point cloud, and repeat the above steps of iterative nearest point registration until the error change is less than the threshold or the maximum number of iterations is reached.

[0059] Step S5: Based on the registration results of step S4, extract ear features and initialize a tracker that combines kernel correlation filter and Kalman filter weighted fusion. Continuously track ear movement and ear canal position changes through a tracking algorithm, and output the three-dimensional spatial coordinates of the ear canal.

[0060] The tracking algorithm used in this invention primarily employs a kernel correlation filter (KCR) and secondarily a Kalman filter, forming a weighted fusion of the two. Based on the ear and ear canal positions determined during initialization, the KCR is initialized and learned separately, and the Kalman filter is also initialized.

[0061] Kernel correlation filter (KCR) tracking is a high-efficiency target tracking algorithm. Its main idea is to train a filter to maximize the difference between the target and the background, thereby achieving target localization. In the initialization phase of the KCR algorithm, the positive region of the target needs to be determined first. This region is obtained using an ear localization algorithm (random sampling consistency coarse registration and iterative nearest-point fine registration method). Cosine weighting is applied to the samples within the positive region, and a 31-dimensional histogram of directional gradients is calculated. Each dimension's feature can be viewed as an M × N sample, denoted as x1, x2, x3, …, x… 31 A training label matrix y with the same sample size M × N is constructed using a two-dimensional Gaussian function. The kernel function k(x, z) is then used.

[0062]

[0063] Where x represents the features of the training samples, and z represents the features of the predicted samples. : indicates inverse Fourier transform, and These represent the training sample and the predicted sample at the th... Components on each feature channel This represents the bandwidth parameter of the Gaussian kernel.

[0064] The original features are mapped to a high-dimensional space, making the target features more linearly separable in that space. A nonlinear regressor model is initialized and trained using the input histogram of oriented gradients.

[0065]

[0066] in, This is the training label matrix, and λ is the regularization parameter. Represents kernel correlation function The Fourier transform results are obtained. In the next frame, the region of interest (ROI) from the previous frame is cosine-weighted, and the histogram of directional gradients is calculated to obtain z1, z2, z3, …, z 31 k(x, z) is obtained using formula (10), and then obtained using the following formula.

[0067]

[0068] Calculate the response matrix in the Fourier domain In the matrix The system finds the location with the largest response. If the response value exceeds a threshold, that location is the target location in the current frame; otherwise, a weighted fusion of the Kalman filter prediction and the current kernel correlation filter tracking is used.

[0069] (12)

[0070] in, KCF It is the positive region predicted by the kernel correlation filter. KF It is the positive region predicted by Kalman filter tracking. ROI This is the final fused positive region. In this process, γ1 and γ2 are the fusion weights for kernel correlation filter tracking prediction and Kalman filter tracking prediction, respectively; let γ1 be 0.8 and γ2 be 0.2. During the nonlinear regressor model update process, corresponding samples are selected based on the newly determined target location, and the calculated model is denoted as α′. When processing the next frame, the model used is the interpolation result between the current calculated model and the initial model.

[0071]

[0072] Where m is the learning rate, and in this embodiment, the value of m ranges from [0.0001, 0.001].

[0073] The above steps describe the traditional kernel correlation filter algorithm's tracking and detection process. However, in real-world scenarios, the target being tracked is not discretely changing but continuously changing with the movement of the depth camera or the target. Therefore, the core idea of ​​this improved kernel correlation filter algorithm is to treat the scale parameter s as a continuous variable and solve for the optimal scale by maximizing the response function r(s). Assuming the target position in the current frame is (x0, y0), the scale parameter is s, and the candidate image patch is obtained by scaling the original image patch and extracting its directional gradient histogram features, the response function is...

[0074]

[0075] Where x(s) represents the scaled image patch features, Scale factor Scaled candidate sample features With target template features The kernel correlation vector between them is represented in the Fourier frequency domain. A second-order Taylor expansion of r(s) is performed at the initial scale s0.

[0076] (15)

[0077] The first and second derivatives are obtained through finite difference calculations. Based on this, the optimal scale is solved using Newton's iterative method. By maximizing r(s) and setting its derivative to zero, we have...

[0078]

[0079] Solve the iterative update formula and iterate until convergence or the maximum number of iterations is reached.

[0080]

[0081] Where k is the number of iterations. Compared to existing discrete scale pyramid searches, this invention can capture minute scale changes, and the tracking of minute forward and backward head movements (depth changes) is smoother and more accurate.

[0082] The core idea of ​​Kalman filtering is to recursively update the target's state estimate by combining a linear system model with observation data, thereby minimizing the estimation error. Traditional Kalman filter tracking algorithms typically use Euclidean distance to measure the similarity between detected and predicted bounding boxes during data association. However, this distance metric is sensitive to scenarios such as target deformation and occlusion. To improve the robustness of this distance metric, this invention introduces the intersection-union ratio (IUGR) as an association criterion. IUGR more directly reflects the spatial overlap of bounding boxes, improving matching robustness. The IUGR function is:

[0083]

[0084] Among them, roi pre These are the predicted bounding boxes of the Kalman filter, roi det This is the detection box of the kernel correlation filter tracking algorithm, where IOU is the intersection-to-union ratio exponent. For tracker T... i and detection D j Calculate the joint cost matrix C ij

[0085]

[0086] in, Let be the prediction error covariance matrix of the Kalman filter. For the first The observation state vector of each detection box. For the first One tracker in Prior state estimation at time 10:00 This is the observation matrix, used to map the state space to the observation space. To observe the noise covariance matrix, The covariance matrix, The weighting coefficients for distance and IOU are given. The Hungarian algorithm is used to find the optimal association, and Kalman filter detection boxes with an IOU less than the IOU threshold of 0.8 are removed. When the IOU is greater than the weighted fusion threshold of 0.8, the detection boxes tracked by the kernel correlation filter algorithm and the predicted boxes by the Kalman filter algorithm are weighted and fused, while the observation noise R is dynamically adjusted.

[0087]

[0088]

[0089] Where R0 is the basic observation noise, α is the attenuation coefficient, and the higher the IOU exponent, the lower the noise and the more reliable observations. express The Kalman gain matrix at time t. This is determined by applying R... k With the adjustment, tracking boxes with higher IOU exponents will receive greater gain weights, enhancing the Kalman filter's observations' correction of the current tracking state.

[0090] When the IOU exponent is less than the threshold, the maximum response value of the kernel correlation filter tracking algorithm is compared with the response value threshold of 0.2. If the maximum response value is greater than 0.2, the Kalman filter algorithm's parameters are updated and reset according to the tracking frame of the kernel correlation filter tracking algorithm. This invention uses the intersection-to-union ratio (IOU) to construct the cost matrix. This mechanism, which replaces simple distance measurement, significantly improves robustness when ear deformation or partial occlusion occurs.

[0091] Step S6: If the maximum response value is less than 0.2 due to target loss or other issues during tracking, the user is prompted that tracking has failed and tracking is paused for 3 seconds. After 3 seconds, all regions of each frame's depth image are re-detected. When the maximum response value of the re-detected bounding box is greater than 0.3, tracking is restarted and the parameters of the Kalman filter tracking algorithm are reset. Many existing methods simply switch or use a single tracker. This invention achieves smooth weighted fusion through an IOU threshold and combines it with an automatic reset mechanism after long-term loss, solving the drift and loss problems in long-term tracking.

[0092] Step S7: If the re-detection fails, repeat step S6 until the detection succeeds or the system restarts.

[0093] Step S8: Upon successful positioning and tracking, output the coordinates of the ear canal in three-dimensional space in real time.

[0094] Example

[0095] The following, in conjunction with the accompanying drawings, demonstrates through experiments the accuracy of the ear canal localization and tracking method based on depth camera contour scanning proposed in this invention in locating and tracking the ear canal during translation or rotation of artificial head models and real human heads.

[0096] 1. Experiment on ear canal localization and tracking of artificial head model

[0097] like Figure 2 The diagram shows the setup for a depth camera to perform ear canal localization and tracking experiments on an artificial head model. The artificial head is supported by a bracket and placed on a horizontal table.

[0098] In the specific experimental setup, the center position of the artificial head model was moved within a 13 × 10 two-dimensional grid plane, with each grid cell spaced 2.00 cm apart. The artificial head model was also translated in a direction perpendicular to this plane, with a vertical spacing of 4.00 cm. The initial position was defined as the point where the center of the artificial head model coincided with the center of the grid. The coordinate system establishment of the depth camera and the artificial head model in three-dimensional space is as follows. Figure 3 As shown.

[0099] With the above settings, the artificial head can translate in the X, Y, and Z directions respectively to conduct translational positioning and tracking experiments of the artificial head in three-dimensional space; in addition, the artificial head can rotate around the rotation axis of the support frame at different angles to conduct positioning and tracking experiments under rotation conditions.

[0100] Grid settings visualization diagram as follows Figure 4 As shown, X is the direction perpendicular to the left side of the artificial head model, ranging from 0 to 20.00 cm; Z is the direction relative to the front and back of the artificial head model, ranging from 0 to 26.00 cm; and Y is the direction perpendicular to the ground, ranging from -4.00 to 4.00 cm.

[0101] In this embodiment, the results of adaptive depth threshold foreground and background segmentation are shown in Figure 5, and the results of histogram equalization and image sharpening are shown in Figure 6. Figure 6 As shown, an example of a spatial filtering image processing algorithm is as follows: Figure 7 As shown, the results of the spatial filtering image processing algorithm are as follows: Figure 8 As shown. The point cloud registration and localization results for the ear and ear canal are as follows. Figure 9 As shown.

[0102] The tracking system's tracking effect on the rotation of the artificial head model is as follows: Figure 10 As shown, the real-time tracking algorithm provides a relatively accurate visual tracking effect during the rotation of the artificial head.

[0103] Figure 11 (a) shows the comparison between the estimated spatial position of the ear canal and its corresponding real position under translation conditions. The real position coordinates are the specific coordinate values ​​of the ear canal of the artificial head model on the coordinate table. Blue dots represent experimental tracking points, and orange stars represent the actual positions of the ear canals of the artificial head. Experimental tests show that as the artificial head model moves away from its initial position, its position within the camera's field of view changes significantly, gradually moving towards the edge of the image. The estimated positions of the ear and ear canal may deviate to some extent from the initial estimation, resulting in a difference from the real position. During the experiment, there was a certain measurement error in the actual translation position of the artificial head model, which was approximately 0.50 cm. The experimental results of the translation measurement experiment of the artificial head model showed that the error in the x, y, and z dimensions was less than 1.10 cm, and the total error was less than 1.40 cm. Detailed error data are shown in Table 1.

[0104] Table 1: Errors during the translation of the artificial head model (test error: 0.50 cm)

[0105] Because head rotation causes dynamic changes in the position and outline of the ear, as the rotation angle increases, the ear outline captured by the depth camera gradually becomes blurred, and obvious data-free holes begin to appear. Therefore, the maximum limit for left rotation is 60°, and the maximum limit for right rotation is 30°. Figure 11 (b) shows a comparison between the spatial position of the ear canal estimated by the tracking system and its corresponding actual position under rotation. During the experiment, there was a certain measurement error in the actual rotation position of the artificial head model, which was approximately 0.50 cm. The experimental results show that the errors in the X, Y, and Z dimensions are all less than 1.80 cm, and the total error is less than 1.90 cm. Detailed error data are shown in Table 2.

[0106] Table 2: Errors during the horizontal rotation of the artificial head model (test error: 0.50 cm)

[0107] 2. Real-person left ear canal localization and tracking experiment

[0108] like Figure 12 The diagram shows the apparatus used in this invention for locating and tracking the left ear canal of a real person. The depth camera is mounted at the same height as the center of the person's head, with a vertical distance of 25.00 cm. A depth data stream is recorded using this depth camera, with parameters set as follows: resolution 848 × 480, frame rate 30 fps, and duration 30 seconds. The positions of the ear and ear canal in each frame are manually labeled as ground truth. Finally, the aforementioned positioning and tracking algorithm framework is used to process the depth data stream, and the algorithm output is labeled as the predicted value for subsequent accuracy evaluation.

[0109] Figure 13 This image shows the visual localization and tracking effect of a depth camera on the left ear canal during the translation and rotation of a human head. The green rectangle represents the rectangular area used to detect the ear's position, indicating the result of the fusion tracking algorithm combining kernel correlation filter and Kalman filter. When the human head remains stable and the ear is clearly visible, the tracking algorithm combining kernel correlation filter and Kalman filter achieves good tracking results. However, when the human head rotates or translates, causing the ear to become blurred in the depth image and difficult to distinguish from the background, the tracking effect deteriorates significantly.

[0110] The output tracking box and ear hole positions represent the tracking results. The tracking error is calculated and compared with the true value. Test results show a tracking success rate of 90.2%. In the depth image, the horizontal ear hole tracking error is 5 pixels, which translates to a 3D actual distance error of approximately 1.00 cm; the vertical ear hole tracking error is 6 pixels, which translates to a 3D actual distance error of approximately 1.20 cm; and the total ear hole tracking error is 8 pixels, which translates to a total 3D actual distance error of approximately 1.60 cm. The real-time tracking algorithm has a processing speed of approximately 55 FPS. Furthermore, the contour scanning method based on a depth camera does not compromise user privacy. Compared to existing active noise-canceling headrests that combine head tracking, this invention achieves a simpler system structure and greater practicality. While avoiding privacy leaks, it effectively tracks the translation and rotation of the target head in 3D space.

[0111] Finally, it should be noted that the entire processing of the above examples used Python 3.8.10 as the programming language and version, and the hardware configuration was a personal laptop computer with an AMD Ryzen 7 4800H processor and 16 GB of RAM. The above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for ear canal localization and tracking based on depth camera contour scanning, characterized in that, The method includes the following steps: Step S1: Use a depth camera to collect the depth contour information of the user's face side in real time and output it in the form of a depth image. Step S2: Enhance the depth image by marking the ear region and ear canal position in the image, and generate and save the side template point cloud accordingly. Step S3: The user maintains the same position and posture as when the image was acquired in step S1, and remains still. The depth camera is used to acquire depth contour information and convert it into a side target point cloud. Step S4: Register the side template point cloud with the side target point cloud to achieve initial positioning of the ear and ear canal; Step S5: Based on the registration results of step S4, extract ear features and initialize a tracker that combines kernel correlation filter and Kalman filter weighted fusion. Continuously track ear movement and ear canal position changes through a tracking algorithm. Step S6: If the target is lost during tracking, the user is prompted that the tracking has failed and the tracking is paused; after a 3-second pause, the re-detection stage begins, where ear features are detected in the global field of view of the depth camera for re-localization; Step S7: If the re-detection fails, repeat step S6. Step S8: Upon successful positioning and tracking, output the coordinates of the ear canal in three-dimensional space in real time.

2. The ear canal localization and tracking method based on depth camera contour scanning according to claim 1, characterized in that, After saving the side template point cloud in steps S1 and S2, the user can directly execute step S3 when using it again, without repeating steps S1 and S2.

3. The ear canal localization and tracking method based on depth camera contour scanning according to claim 1, characterized in that, In step 2, the enhancement processing of the depth image includes sequentially employing adaptive depth threshold foreground and background segmentation, histogram equalization, image sharpening, and spatial filtering image processing algorithms.

4. The ear canal localization and tracking method based on depth camera contour scanning according to claim 3, characterized in that, The adaptive depth threshold foreground and background segmentation algorithm is as follows: First, the gray values ​​of the depth image are normalized by mean. Then, the calculation result of the Sigmoid function is weighted and fused with the gray standard deviation to generate the background segmentation threshold.

5. The ear canal localization and tracking method based on depth camera contour scanning according to claim 4, characterized in that, The background segmentation threshold is: , where μ w σ is the weight, and σ is the standard deviation of the grayscale values.

6. The ear canal localization and tracking method based on depth camera contour scanning according to claim 1, characterized in that, In step S4, registering the side template point cloud with the side target point cloud specifically involves: first, coarse registration is performed using a random sampling consistency method, and then fine registration is performed using an iterative nearest point algorithm. After aligning the side template point cloud with the side target point cloud, the positions of the ear and ear canal in the side target point cloud are calculated accordingly to complete the initial positioning.

7. The ear canal localization and tracking method based on depth camera contour scanning according to claim 1, characterized in that, In step S5, the tracking algorithm primarily uses kernel correlation filtering (KCR) and secondarily uses Kalman filtering (KL), and the two are weighted and fused. Specifically, during the KCR tracking phase, the position with the largest response is found. If the response value exceeds a threshold, this position is the target position in the current frame; otherwise, the weighted fusion result of the KL tracking prediction and the current KCR tracking is used. in, KCF It is the positive region predicted by kernel correlation filter tracking. KF It is the positive region predicted by Kalman filter tracking. ROI This is the positive region of final fusion, let γ1 be 0.8 and γ2 be 0.

2.

8. The ear canal localization and tracking method based on depth camera contour scanning according to claim 7, characterized in that, In the kernel correlation filtering algorithm, the scale parameter s is treated as a continuous variable, and the optimal scale is solved by maximizing the response function r(s). Assuming the target position in the current frame is (x0, y0), the scale parameter is s, and the candidate image patch is obtained by scaling the original image patch and extracting its directional gradient histogram features, the response function is... Where x(s) represents the scaled image patch features, Scale factor Scaled candidate sample features With target template features The representation of the kernel correlation vector between them in the Fourier frequency domain; The response function r(s) is expanded in second order Taylor series at the initial scale s0. The first and second derivatives are obtained through finite difference calculations. Based on this, the optimal scale is solved using Newton's iteration method. By maximizing r(s) and setting its derivative to zero, we have... Solve the iterative update formula Where k is the number of iterations; iteration continues until convergence or the maximum number of iterations is reached.

9. The ear canal localization and tracking method based on depth camera contour scanning according to claim 7, characterized in that, In the Kalman filter algorithm, the intersection-to-union ratio (CUIR) is introduced as an association criterion. Among them, roi pre These are the predicted bounding boxes of the Kalman filter, roi det It is the detection box of the kernel correlation filter algorithm, and IOU is the intersection-union ratio exponent; When the Intersection over Union (IOU) index is greater than the weighted fusion threshold of 0.8, the detection boxes tracked by the kernel correlation filter algorithm and the predicted boxes tracked by the Kalman filter algorithm are weighted and fused, while the values ​​are dynamically adjusted. Observation noise R at time k : Where R0 is the basic observation noise, and α is the attenuation coefficient. Let be the prediction error covariance matrix of the Kalman filter. For the observation matrix, express The Kalman gain matrix at time t.

10. The ear canal localization and tracking method based on depth camera contour scanning according to claim 9, characterized in that, When the Intersection over Union (IOU) index is less than the weighted fusion threshold of 0.8, the maximum response value of the kernel correlation filter algorithm is compared with the response value threshold of 0.

2. If the maximum response value is greater than 0.2, the tracking box of the kernel correlation filter algorithm is updated and the parameters of the Kalman filter algorithm are reset. If the maximum response value is less than 0.2, tracking is stopped for 3 seconds. After 3 seconds, all regions of the depth image of each frame are re-detected. When the maximum response value of the re-detected box is greater than 0.3, tracking is restarted and the parameters of the Kalman filter algorithm are reset.