Driver and passenger ear side position positioning and tracking method for in-vehicle sound field reconstruction

By constructing facial feature triangles and abnormal exposure and redundant identification processing, combined with the spatial vector relationship between the nose and the ear, high-precision positioning and tracking of the driver and passenger's ear position is achieved, solving the problem of inaccurate positioning of the ear position in the reconstruction of the in-vehicle sound field, and improving the robustness and practicality of the sound field reconstruction.

CN120708260APending Publication Date: 2025-09-26TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510787078.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider the positioning and tracking of the ear positions of drivers and passengers in the reconstruction of the in-vehicle sound field, especially under abnormal lighting conditions and redundant recognition. This results in low accuracy in ear position tracking, and the high cost of depth cameras limits their application.

Method used

By constructing facial feature triangles and using a monocular camera to obtain the pixel coordinates of the driver and passenger's facial feature points, abnormal exposure and redundant identification processing are performed. Combined with the spatial vector relationship between the nose and the ear, the three-dimensional coordinates of the controlled object's two ears are calculated to achieve accurate positioning and tracking of the ear side position.

Benefits of technology

It improves the accuracy of ear side position tracking, improves facial feature matching under extreme lighting and special driving behaviors, ensures the stability and reliability of sound field reconstruction, reduces hardware costs, and overcomes the application defects of easy ear occlusion and depth cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708260A_ABST
    Figure CN120708260A_ABST
Patent Text Reader

Abstract

The invention relates to a driver and passenger ear side position positioning and tracking method for in-vehicle sound field reconstruction, and the method comprises the steps: obtaining the pixel coordinates of the feature points of the five sense organs of a driver and passenger according to a camera monitoring image and a facial feature template, and constructing a facial feature triangle which maps the positions of the eyes and the nose; whether abnormal exposure and redundant recognition exist in the facial feature triangle or not is judged, if yes, a camera monitoring image is processed, the facial feature triangle mapping the positions where the binocular and the nose are located is reconstructed, and if not, a numerical value transformation mechanism of the geometric features of the facial feature triangle and the displacement of the controlled object is further determined; and three-dimensional coordinates of the positions of the two ears of the controlled object are calculated on the basis of the space vector for connecting the nose part and the ear sides. Compared with the prior art, the method can improve the adverse effects of abnormal exposure and redundant recognition on face recognition, improves the tracking precision of the ear side positions of the driver and passengers, and facilitates the improvement of the efficiency and accuracy of sound field reconstruction in the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of acoustic control technology, and in particular to a method for locating and tracking the ear side position of a driver and passenger for reconstructing a sound field in a vehicle. Background Art

[0002] As vehicle occupants increasingly demand higher quality in-vehicle auditory experience, creating a diverse and private acoustic space is crucial to the development of intelligent vehicle cockpits. Sound field reconstruction technology, by planning and solving speaker drive signals, can control the audio delivered to specific areas within the vehicle.

[0003] Limited by the uncertainty of the driver's and occupants' positional changes during driving, the performance of in-vehicle sound field reconstruction is easily affected by random changes in the controlled area, resulting in degradation. Although depth cameras can directly obtain the spatial coordinates of facial landmarks, direct location of ear landmarks is difficult due to the easy occlusion of the ears. In addition, their high cost also restricts practical application. Therefore, accurately locating and tracking the driver's and occupants' ears through monocular camera images is crucial for the robust implementation of in-vehicle sound field reconstruction.

[0004] In existing research, the journal papers "Train driver's head posture estimation based on ASM local positioning and feature triangles [J]. Journal of the China Railway Society, 2016." and "Face detection and head posture estimation fusion algorithm based on SSD model [J]. Journal of Jiangsu University (Natural Science Edition), 2019." both used computer vision technology to calculate and estimate the head posture and position of drivers and passengers.

[0005] In addition, invention patent CN 115052225 A proposes an active control method for the sub-regional sound field in a vehicle, which can ensure the performance balance of acoustic energy contrast and sound field reconstruction error between regions in the vehicle, under the premise that the individual speaker driving signal is within the linear operating range.

[0006] However, the above solution still has the following disadvantages:

[0007] 1. Although existing technologies can numerically estimate changes in the driver's and passenger's head posture, they do not fully consider various abnormal lighting conditions during vehicle driving and redundant recognition in feature recognition.

[0008] 2. Most existing acoustic control methods are aimed at constant control areas within the controlled environment, and no framework for reconstructing the in-vehicle sound field is established for locating and tracking the ear positions of the driver and passengers. Summary of the Invention

[0009] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for locating and tracking the ear position of the driver and passenger for reconstructing the in-vehicle sound field, which can improve the adverse effects of abnormal exposure and redundant recognition on facial recognition and improve the accuracy of tracking the ear position of the driver and passenger.

[0010] The object of the present invention can be achieved by the following technical solution: A method for locating and tracking the ear position of a driver and passenger for reconstructing the in-vehicle sound field, comprising the following steps:

[0011] S1. Obtain pixel coordinates of the facial features of the driver and passenger based on the camera monitoring image and the facial feature template, and construct a facial feature triangle that maps the location of the binoculars and nose;

[0012] S2. Determine whether there is abnormal exposure and redundant recognition of the facial feature triangle. If so, process the camera monitoring image and return to step S1. Otherwise, execute step S3.

[0013] S3. Determine the numerical transformation mechanism between the geometric features of the facial feature triangle and the displacement of the controlled object, and calculate the three-dimensional coordinates of the positions of the two ears of the controlled object based on the spatial vector connecting the nose and the ear side.

[0014] Furthermore, the step S1 includes the following process:

[0015] S11, using the facial feature template to identify and match the camera monitoring image, and determine whether the camera monitoring image is valid. If it is valid, execute step S13; if not, execute step S12;

[0016] S12, rotating the camera monitoring image so that the facial area of ​​the subject in the image is in a vertical state, and then executing step S13;

[0017] S13, further dividing the face area of ​​the subject under test in the camera monitoring image, performing a preliminary frame selection on the eyes and nose of the subject under test based on the numerical matrix of the facial feature template, and obtaining corresponding pixel coordinates with the geometric center of the frame selection area as a reference position;

[0018] S14. Connect the pixel coordinates of the eye and nose feature points to construct a facial feature triangle that maps the posture change of the object under test, thereby achieving a geometric mapping from the camera monitoring image to the facial features of the object under test.

[0019] Furthermore, step S11 specifically uses the facial feature template to identify and match the camera monitoring image to determine whether the head of the subject in the camera monitoring image is tilted. If the head is tilted, the camera monitoring image is judged to be invalid, otherwise it is judged to be valid.

[0020] Furthermore, the step S12 specifically uses the center of the camera monitoring image as the rotation origin, and rotates the image clockwise / counterclockwise at set angles to make the facial area of ​​the subject in the image in a vertical state.

[0021] Furthermore, the process of processing the camera monitoring image in step S2 includes abnormal exposure processing and redundant identification processing.

[0022] Furthermore, the abnormal exposure processing is specifically as follows:

[0023] Considering the extreme lighting conditions of driving, the camera assesses the degree of exposure abnormality in the monitored image based on the grayscale mean value of the facial area. If the mean value is above the overexposure threshold or below the underexposure threshold, the image is considered abnormally exposed.

[0024] Targeted grayscale correction is performed on overexposed and underexposed images, and the grayscale value near each feature point is corrected using the grayscale change function to improve the possibility and accuracy of feature point recognition.

[0025] Furthermore, the redundancy identification process is specifically as follows:

[0026] Taking into account special driving behaviors, the facial detection frame is divided into six equal parts in the horizontal and vertical directions to generate grid areas. Based on the position constraints of the physiological characteristics of the eyes and nose, redundant recognition results that deviate from the typical distribution range are eliminated.

[0027] Furthermore, the displacement of the controlled object in step S3 includes translational displacement and rotational displacement of the controlled object.

[0028] Furthermore, step S3 includes the following process:

[0029] S31. Based on the pinhole imaging model, derive the quantitative relationship between the physical size of the feature and the imaging pixel size. Combined with the camera position in the monitoring environment, establish a mathematical relationship model between the base length of the facial feature triangle and the measured depth.

[0030] S32, calculating the depth of the subject according to the length of the base of the facial feature triangle, and calculating the translational displacement in different directions according to the spatial coordinate change based on the initial coordinates of the subject's head;

[0031] S33, performing simulation and calculation based on the three-dimensional human head model, fitting the functional mapping relationship between the rotation angle of the head of the subject and the geometric features of the facial feature triangle, and performing inverse solution to obtain the rotation displacement of the head of the subject in different directions;

[0032] S34. Apply the calculated translational displacement to the initial coordinates of the head of the controlled object, and convert the rotational displacement to a space vector connecting the nose and the ear side, thereby obtaining the binaural positions of the controlled object.

[0033] Furthermore, the translational displacement is specifically:

[0034]

[0035] The rotation displacement is specifically:

[0036]

[0037] Among them, d x d y d z is the translational displacement along the longitudinal, transverse and vertical directions, x0, y0 and z0 are the initial coordinate positions of the head of the measured object, and θ, β and α are the rotational displacements along yaw, pitch and roll.

[0038] Compared with the prior art, the present invention has the following advantages:

[0039] The present invention obtains the pixel coordinates of the driver's facial features based on camera monitoring images and facial feature templates, thereby constructing a facial feature triangle that maps the location of the eyes and nose. By determining whether the facial feature triangle has abnormal exposure and redundant recognition, the system processes camera monitoring images that have abnormal exposure and redundant recognition. Furthermore, by determining the numerical transformation mechanism between the geometric features of the facial feature triangle and the displacement of the controlled object, and based on the spatial vector connecting the nose and the ear, the three-dimensional coordinates of the controlled object's ears can be calculated. This method can mitigate the adverse effects of abnormal exposure and redundant recognition on facial recognition in scenarios where the vehicle driver and occupant's posture changes, and improve the accuracy of tracking the driver's ear position.

[0040] Through targeted image processing (abnormal exposure processing and redundant recognition processing), the present invention can improve the interference of abnormal exposure and redundant recognition on facial feature extraction, improve the matching degree of facial feature templates for ear side position tracking, enhance the facial feature matching accuracy in extreme lighting environments (such as tunnels, at night) and special driving behaviors (such as turning the head to observe), ensure the stability of ear side position tracking, provide continuous and reliable positioning data for sound field reconstruction, and break through the application defects of easy occlusion of the ears and high cost of depth cameras.

[0041] Based on the geometric features and displacement transformation mechanism of facial feature triangles, the present invention infers the position of both ears through the spatial vector relationship between the nose and the ear side, converts complex head motion solutions into geometric parameter iterations, improves the efficiency of ear side positioning, and provides a high-precision spatial reference for in-vehicle sound field reconstruction, ensuring the practicality and reliability of the sound field reconstruction solution in real-world vehicle scenarios. With the low hardware cost of a monocular camera and accurate position estimation of the controlled object, the ear side position positioning and tracking of the driver and passenger can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of the method flow of the present invention;

[0043] Figure 2 Schematic diagram of the application process of the embodiment;

[0044] Figure 3a Schematic diagram of overexposure grayscale correction function;

[0045] Figure 3b Schematic diagram of underexposure grayscale correction function;

[0046] Figure 4 Schematic diagram of facial region division for redundant / erroneous recognition in an embodiment;

[0047] Figure 5 Schematic diagram of the triangle fitting of the head rotation displacement and facial features of the subject in the embodiment;

[0048] Figure 6 Schematic diagram of the triangle fitting function between the head rotation displacement and facial features of the subject in the embodiment;

[0049] Figure 7 is the ear side position of the measured object facing the in-car environment in the embodiment;

[0050] Figure 8 is the moving trajectory of the measured object facing the in-vehicle environment in the embodiment;

[0051] Figure 9 is the spatial coordinate of the object to be measured facing the in-vehicle environment in the embodiment;

[0052] Figure 10 Schematic diagram of the reconstructed sound pressure level change at the main driving position in the embodiment;

[0053] Figure 11 Schematic diagram of the reconstructed sound pressure level change at the co-pilot position in the embodiment. DETAILED DESCRIPTION

[0054] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] Example

[0056] like Figure 1 As shown, a method for locating and tracking the ear position of a driver and passenger for reconstructing the in-vehicle sound field includes the following steps:

[0057] S1. Obtain pixel coordinates of the facial features of the driver and passenger based on the camera monitoring image and the facial feature template, and construct a facial feature triangle that maps the location of the binoculars and nose;

[0058] S2. Determine whether there is abnormal exposure and redundant recognition of the facial feature triangle. If so, process the camera monitoring image and return to step S1. Otherwise, execute step S3.

[0059] S3. Determine the numerical transformation mechanism between the geometric features of the facial feature triangle and the displacement of the controlled object, and calculate the three-dimensional coordinates of the positions of the two ears of the controlled object based on the spatial vector connecting the nose and the ear side.

[0060] This embodiment applies the above solution. To verify the effectiveness of this solution, after obtaining the three-dimensional coordinates of the binaural positions of the controlled object, the results are applied to the in-vehicle sound field reconstruction scenario, and the changes in the sound field reconstruction performance before and after are compared and verified.

[0061] like Figure 2 As shown, the main contents of this embodiment are:

[0062] 1. Based on the camera image and facial feature template, the pixel coordinates of the driver's facial features are obtained, and the facial feature triangles are constructed to map the location of the eyes and nose.

[0063] Specifically, different facial feature templates are used to identify and match the camera monitoring image. If no target facial features are detected, the subject is judged to have a head tilt (i.e., the image is invalid). Otherwise, the pixel coordinates of the facial feature points are extracted based on the facial feature template. Based on the numerical matrix of the facial feature template, the eyes and nose of the subject are preliminarily selected, and the corresponding pixel coordinates are obtained with the geometric center of the selected area as the reference position;

[0064] After determining that the image is invalid, the camera monitoring image of the subject's head with a side tilt is rotated clockwise or counterclockwise at intervals of 5 degrees to make the subject's face in the camera monitoring image vertical. Then, based on the preliminary selection of the facial area, the facial features such as the eyes and nose are extracted and marked based on the facial feature template;

[0065] The facial area of ​​the subject is further divided, and the recognition range of the target organ is narrowed to improve the success rate of feature point recognition. The geometric center of the feature frame selection area is used as the reference position to obtain the corresponding pixel coordinates;

[0066] The pixel coordinates of the eye and nose feature points are connected to construct a facial feature triangle that maps the posture changes of the object being measured, thereby realizing the geometric mapping from the camera monitoring image to the facial features of the object being measured.

[0067] Second, targeted image processing is used to mitigate the negative impact of abnormal exposure and redundant recognition on facial recognition, and to improve the matching of facial feature templates for ear position tracking in extreme lighting environments and special driving behaviors.

[0068] Specifically, for monitoring images that cannot accurately construct facial feature triangles, targeted image processing is implemented based on two problems: abnormal exposure (overexposure / underexposure) and redundant recognition (false detection / repeated detection). Among them, redundant recognition caused by extreme lighting environments and special driving behaviors during vehicle driving needs to be processed in parallel. Abnormal exposure is processed through targeted grayscale correction, and redundant recognition is processed by eliminating incorrect / redundant feature positions.

[0069] In this embodiment, the exposure level of the occupant's face in the input camera monitoring image is evaluated based on the grayscale average value, and the grayscale value range of 0-255 is divided into three equal parts: 0-85 is underexposure, 85-170 is normal exposure, and 170-255 is overexposure.

[0070] Then, according to the grayscale changes of the pixels, the over-exposed and under-exposed images are corrected in a targeted manner (e.g. Figure 3a and 3b As shown in the figure), the grayscale change function is used to correct the grayscale value near each feature point to improve the recognition possibility and recognition accuracy of the feature point;

[0071] Divide the face area of ​​the subject into six equal parts horizontally and vertically (e.g. Figure 4 As shown in the figure, based on the position constraints of physiological features such as eyes and nose, the incorrect / redundant recognition that does not conform to the distribution characteristics of the face is eliminated to correct the recognition results;

[0072] Finally, the corrected monitoring image is used to complete the geometric construction of the facial feature triangle, laying the foundation for the subsequent numerical conversion of the mapping relationship between the triangle geometric features and the head rotation.

[0073] 3. Clarify the geometric features of the facial feature triangle and the numerical transformation mechanism of the translational and rotational displacements of the controlled object, and infer the position of the controlled object's ears based on the spatial vector connecting the nose and the side of the ear;

[0074] Specifically, using the coordinate transformation relationship from the pixel coordinate system to the camera coordinate system, we can obtain the functional expression of the base length of the facial pixel feature triangle and the measured depth. Assume that there is a point in the real world, which can be expressed as follows in the camera coordinate system, image coordinate system, and pixel coordinate system respectively:

[0075]

[0076] Among them, P c is the camera coordinate position, P i is the image coordinate position, P p is the pixel coordinate position. Generally speaking, the camera coordinate system uses the camera optical center as the origin, the image coordinate system uses the image center as the origin, and the pixel coordinate system uses the upper left corner of the image as the origin.

[0077] In this embodiment, the pinhole model is applied to the conversion between the image coordinate system and the camera coordinate system. The corresponding similar triangle relationship is expressed as follows:

[0078]

[0079] Among them, f is the focal length of the camera, and the rest are the coordinate values ​​in the image coordinate system and the camera coordinate system respectively.

[0080] The conversion relationship between the image coordinate system and the camera coordinate system is:

[0081]

[0082] The above formula can be expressed as a homogeneous equation:

[0083]

[0084] Similarly, the conversion relationship between image coordinates and pixel coordinates is expressed in the form of a homogeneous equation as follows:

[0085]

[0086] Among them, 1 / dx and 1 / dy are used for unit conversion between millimeters and pixels, and u0 and v0 are the coordinate values ​​of the image origin in the pixel coordinate system.

[0087] Combining the above two homogeneous equations, the conversion relationship between the pixel coordinate system and the camera coordinate system is directly derived as follows:

[0088]

[0089] The coefficient matrix of the above formula is usually defined as K, which is also called the camera intrinsic parameter matrix.

[0090] The length of the base of the facial pixel feature triangle is inversely proportional to the measured depth, and this numerical relationship is determined by the camera intrinsic parameter matrix. Its approximate expression can be expressed as follows:

[0091]

[0092] Among them, l0 and l p Represents the distance between two points in the camera coordinate system and the pixel coordinate system respectively. The former is in millimeters, while the latter is in pixels.

[0093] Then, the depth of the object under test is calculated based on the length of the base of the facial feature triangle, and the translational displacement in different directions is calculated by the change in spatial coordinates, as shown in the following formula:

[0094]

[0095] Among them, d x , d y , d z is the translational displacement along the longitudinal, transverse and vertical directions, and x0, y0, z0 are the predetermined initial coordinate positions.

[0096] Then, simulation and calculation are performed based on the three-dimensional head model to fit the functional mapping relationship between the head rotation angle of the subject and the geometric features of the facial feature triangle. For the three rotation displacements of the occupant's head, namely roll, roll and pitch, the rotation information of the subject's head is inversely calculated by combining the geometric features of the feature triangle in the image, thus avoiding the complexity of directly measuring the rotation angle. The correlation calculation of the three rotation displacements is obtained based on the three-dimensional geometric feature transformation of the feature triangle under the pixel coordinate (such as Figure 5 shown).

[0097] By rotating the three-dimensional human head, the mapping function between different rotation degrees under each degree of freedom and the corresponding triangle feature transformation is fitted.

[0098] Based on the physiological characteristics of the human body, the different rotational displacements of the subject's head are all analyzed with the connection point between the head and neck as the rotation center, and the pixel feature triangle geometric transformation is analyzed with the center point of the line connecting the two eyes as the rotation center.

[0099] In order to accurately describe the relationship between the geometric feature changes of the feature triangle in the pixel coordinate system and the head rotation, this embodiment uses a smooth curve to perform parameter fitting on the head rotation behavior under different degrees of freedom, thereby constructing a mapping relationship between each rotation angle and the feature triangle parameters (such as Figure 6 shown).

[0100] The reverse solution obtains the rotational displacement of the head of the measured object in different directions as follows:

[0101]

[0102] Among them, θ, β, α are the rotational displacements along yaw, pitch, and roll.

[0103] Finally, the calculated translational displacement is applied to the initial coordinates of the controlled object's head, and the rotational displacement is converted into a spatial vector connecting the nose and the ear side to obtain the binaural position of the controlled object.

[0104] Fourth, the driver and passenger ear positioning and tracking framework is loaded into the in-vehicle sound field reconstruction model. The effectiveness of this solution is evaluated by comparing the changes in sound field reconstruction performance before and after.

[0105] Specifically, based on the geometric characteristics of the vehicle cabin interior space, the driver and passenger head activity area is divided into several independent sound field reconstruction areas, and the corresponding expected sound pressure level parameters are set for each reconstruction area;

[0106] In the sound field reconstruction, the reconstruction area that receives the target audio is defined as the bright area, and the reconstruction area that blocks the target audio is defined as the dark area. In this embodiment, the main driving area is defined as the sound field reconstruction bright area, and its spatial coordinate position can be expressed in matrix form as follows:

[0107]

[0108] Among them, D b It is the three-dimensional space coordinate matrix composed of the main driving position control points.

[0109] Similarly, the co-pilot area is defined as the dark area for sound field reconstruction, and its spatial coordinate position can be expressed in matrix form as follows:

[0110]

[0111] Among them, D d is the three-dimensional space coordinate matrix composed of the co-pilot position control points.

[0112] In this embodiment, the expected sound pressure in the bright area (main driving area) is defined as 0.1 Pa, and the corresponding expected sound pressure level is 73.98 dB.

[0113] In this embodiment, the difference in sound pressure levels between the bright area and the dark area is set to 15 dB, thereby defining the expected sound pressure level of the dark area (passenger vehicle area) as 58.98 dB.

[0114] Then, the physical coordinates of the vehicle speakers and the distribution of control points in the reconstruction area are combined to construct an acoustic transfer function for sound field reconstruction, and the initial mathematical model for sound field reconstruction is completed.

[0115] This embodiment defines the acoustic transfer function of the bright area and expresses it in matrix form as follows:

[0116]

[0117] Among them, G b is the acoustic transfer function matrix from l loudspeakers to n control points in the bright area of ​​the reconstruction area.

[0118] The acoustic transfer function of the dark zone is defined and expressed in matrix form as follows:

[0119]

[0120] Among them, G d is the acoustic transfer function matrix from l loudspeakers to m control points in the dark area of ​​the reconstruction area.

[0121] Then, under the constraint of a constant driver and passenger head position as a static reconstruction condition, the driving signals of the on-board speakers that meet the target sound pressure level in each area are solved, and the sound field reconstruction performance of different reconstruction areas is evaluated.

[0122] In this embodiment, the speaker driving signal is obtained by solving the minimum sound field reconstruction error in the reconstruction area. The minimum sound field reconstruction error at the driver and co-driver positions is used as the optimization target to obtain a reconstructed sound field that meets the given requirements. The mathematical expression is:

[0123] u(f)=min{abs{P b -E b}+abs{P d -E d}}

[0124] Where P and E are the reconstructed sound field and the expected sound field in the reconstruction area, respectively, u is the loudspeaker driving signal obtained by numerical solution, and f is the reconstruction frequency.

[0125] Based on the speaker driving signal and the corresponding acoustic transfer function, the sound pressure level of the reconstruction area is calculated. In this embodiment, the sound pressure level distribution at the main driving position is:

[0126] P b (f) l×1 =G b (f) l×n u(f) n×1

[0127] The sound pressure level distribution at the co-pilot position is:

[0128] P d (f) l×1 =G d (f) l×m u(f) m×1

[0129] Then, based on the real-time acquired 3D coordinates of the driver and passenger's ear feature points, the spatial boundary definition and expected sound pressure level parameters of the sound field reconstruction area are dynamically adjusted to generate a speaker drive signal that matches the current posture.

[0130] In this embodiment, a 5-second video of the position change of the object facing the car environment is used as the processing object to obtain its binaural position (such as Figure 7 As shown), movement trajectory (as shown Figure 8 ) and spatial coordinates (as shown in Figure 9 shown).

[0131] Finally, based on the spatial coordinates of the driver's and passenger's ears obtained by tracking, the scope of the sound field reconstruction area is adjusted, the updated acoustic driving signal is loaded into the vehicle speakers, and the sound field reconstruction performance before and after the update is evaluated.

[0132] This embodiment constructs a reconstruction area change scheme for a total of 14 seconds at a 2-second interval at the main driving position, wherein the 2nd to 8th seconds are +0.1m translation displacements in different directions, and the 8th to 14th seconds are +15deg rotation displacements in different directions.

[0133] Using a 1kHz sinusoidal audio signal as the excitation source, the sound pressure level changes in the main driving area and the co-pilot area before and after the update and reconstruction area are observed (e.g. Figure 10 and Figure 11 shown).

[0134] Combined with the deteriorating impact of occupant posture activities on sound field reconstruction, it can be found that the in-vehicle sound field control based on ear position tracking can always maintain the sound pressure in the bright area at around 73.98dB, and the dark area can also be better maintained near the desired sound pressure level of 58.98dB.

[0135] In summary, this solution proposes a method for ear position localization and tracking that synergistically mitigates the negative effects of abnormal exposure and redundant recognition on facial recognition, and demonstrates good compatibility with the uncertainty of the control area for in-vehicle sound field reconstruction. Firstly, the pixel coordinates of the driver's facial features are obtained from camera monitoring images and facial feature templates. This is used to construct a facial feature triangle that maps the location of the binoculars and nose. The numerical transformation mechanism between the geometric characteristics of the facial feature triangle and the translational and rotational displacements of the controlled object is clarified. This method uses low-cost monocular camera hardware and precise controlled object position to infer the driver's ear position, achieving ear position localization and tracking. Secondly, image preprocessing is used to mitigate the negative effects of abnormal exposure and redundant recognition on facial recognition, improving the matching of the facial feature template with ear position tracking. This makes the technology highly adaptable to abnormal lighting conditions and redundant recognition during vehicle driving. When applied in practice, this solution provides a more efficient solution for in-vehicle sound field reconstruction while ensuring accurate ear position tracking.

Claims

1. A method for locating and tracking the ear position of a driver and passenger for reconstructing the in-vehicle sound field, characterized in that: The following steps are involved: S1. Obtain pixel coordinates of the facial features of the driver and passenger based on the camera monitoring image and the facial feature template, and construct a facial feature triangle that maps the location of the binoculars and nose; S2. Determine whether there is abnormal exposure and redundant recognition of the facial feature triangle. If so, process the camera monitoring image and return to step S1. Otherwise, execute step S3. S3. Determine the numerical transformation mechanism between the geometric features of the facial feature triangle and the displacement of the controlled object, and calculate the three-dimensional coordinates of the positions of the two ears of the controlled object based on the spatial vector connecting the nose and the ear side.

2. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 1, characterized in that: The step S1 includes the following process: S11, using the facial feature template to identify and match the camera monitoring image, and determine whether the camera monitoring image is valid. If it is valid, execute step S13; if not, execute step S12; S12, rotating the camera monitoring image so that the facial area of ​​the subject in the image is in a vertical state, and then executing step S13; S13, further dividing the face area of ​​the subject under test in the camera monitoring image, performing a preliminary frame selection on the eyes and nose of the subject under test based on the numerical matrix of the facial feature template, and obtaining corresponding pixel coordinates with the geometric center of the frame selection area as a reference position; S14. Connect the pixel coordinates of the eye and nose feature points to construct a facial feature triangle that maps the posture change of the object under test, thereby achieving a geometric mapping from the camera monitoring image to the facial features of the object under test.

3. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 2, characterized in that: The step S11 specifically uses the facial feature template to identify and match the camera monitoring image to determine whether the head of the subject in the camera monitoring image is tilted. If the head is tilted, the camera monitoring image is judged to be invalid, otherwise it is judged to be valid.

4. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 2, characterized in that: The step S12 specifically uses the center of the camera monitoring image as the rotation origin, and rotates the image clockwise / counterclockwise at set angles to make the facial area of ​​the subject in the image in a vertical state.

5. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 2, characterized in that: The process of processing the camera monitoring image in step S2 includes abnormal exposure processing and redundant identification processing.

6. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 5, characterized in that: The abnormal exposure processing is specifically as follows: Considering the extreme lighting conditions of driving, the camera assesses the degree of exposure abnormality in the monitored image based on the grayscale mean value of the facial area. If the mean value is above the overexposure threshold or below the underexposure threshold, the image is considered abnormally exposed. Targeted grayscale correction is performed on overexposed and underexposed images, and the grayscale value near each feature point is corrected using the grayscale change function to improve the possibility and accuracy of feature point recognition.

7. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 5, characterized in that: The redundancy identification process is specifically as follows: Taking into account special driving behaviors, the facial detection frame is divided into six equal parts in the horizontal and vertical directions to generate grid areas. Based on the position constraints of the physiological characteristics of the eyes and nose, redundant recognition results that deviate from the typical distribution range are eliminated.

8. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 1, characterized in that: The displacement of the controlled object in step S3 includes translational displacement and rotational displacement of the controlled object.

9. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 8, characterized in that: Step S3 The following processes are included: S31. Based on the pinhole imaging model, derive the quantitative relationship between the physical size of the feature and the imaging pixel size. Combined with the camera position in the monitoring environment, establish a mathematical relationship model between the base length of the facial feature triangle and the measured depth. S32, calculating the depth of the subject according to the length of the base of the facial feature triangle, and calculating the translational displacement in different directions according to the spatial coordinate change based on the initial coordinates of the subject's head; S33, performing simulation and calculation based on the three-dimensional human head model, fitting the functional mapping relationship between the rotation angle of the head of the subject and the geometric features of the facial feature triangle, and performing inverse solution to obtain the rotation displacement of the head of the subject in different directions; S34. Apply the calculated translational displacement to the initial coordinates of the head of the controlled object, and convert the rotational displacement to a space vector connecting the nose and the ear side, thereby obtaining the binaural positions of the controlled object.

10. The method for locating and tracking the ear position of a driver and passenger for in-vehicle sound field reconstruction according to claim 9, characterized in that: The translation displacement is specifically: The rotation displacement is specifically: Among them, d x d y d z is the translational displacement along the longitudinal, transverse and vertical directions, x0, y0 and z0 are the initial coordinate positions of the head of the measured object, and θ, β and α are the rotational displacements along yaw, pitch and roll.