Feature point selection for ear tracking
By identifying and processing user image feature points and determining the three-dimensional positioning of the ear, the problem of lack of physical scale of camera images in the prior art is solved, and the correction ability and user experience of the audio system are improved.
Patent Information
- Application Number
- CN202411810539.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-06
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-13
AI Technical Summary
The existing camera-based head and ear tracking systems have difficulty in accurately correcting sounds due to the lack of physical scale of camera images, which affects the user experience during the size changes between individuals.
By acquiring the user's image, processing the image to identify the face feature points, selecting the feature point pair based on the head posture, determining the three-dimensional positioning of the user's ear, and processing the audio signal to generate the processed audio signal.
Improved camera-based head and ear tracking accuracy, providing better noise cancellation, crosstalk cancellation and 3D audio listening experience.
Smart Images

Figure CN120151741A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 608,529, filed on December 11, 2023, entitled "LANDMARK SELECTION BASED ON HEAD AND EAR TRACKING". The subject matter of this related application is hereby incorporated herein by reference. Technical Field
[0003] This application relates to systems and methods for head and ear tracking, and more particularly, to feature point selection for ear tracking. Background Art
[0004] Headrest audio systems, seat or chair audio systems, soundbars, vehicle audio systems, and other personal and / or near - field audio systems are becoming increasingly popular. However, when listeners move their heads, even slightly, the sound experienced by users of personal and / or near - field audio systems can change significantly ( For example , 3 dB to 6 dB or another value). In the example of a headrest audio system, the variation from one person using the audio system to another can also vary significantly depending on how the user is positioned in the seat and how the headrest is adjusted. This amount of sound pressure level (SPL) variation makes it difficult to adjust the audio system. Additionally, when rendering spatial audio through headrest speakers, this variation causes features such as crosstalk cancellation to fail.
[0005] One way to correct the audio of personal and / or near - field audio systems is camera - based head tracking. Due to the increasing adoption of driver monitoring systems (DMS), consumer devices including imaging devices, and webcameras, camera - based head tracking is increasingly incorporated into vehicles, personal computer games, and home theaters. Using available cameras, camera - based head tracking systems acquire image data and process the images to estimate the position of the head in the environment. Based on the estimated head position, camera - based head tracking systems can estimate the position of the user's ears in space, enabling the audio system to modify its characteristics to account for the position of the user's ears.
[0006] However, one drawback of existing camera - based head and / or ear tracking is that camera images do not have a physical scale. Camera images consist of pixels that generally cannot be dimensioned or scaled for the objects represented therein because depending on the seating position, standing position, and other user positions, the head or face in the image may be closer or farther. In systems that include scaling, conventional methods for performing the scaling of pixels to distance units involve using speculative dimensions, such as the average inter - eye distance ( For example, the interpupillary distance (the distance between the centers of the user's pupils). However, there are significant variations in the interocular distance between individuals; for example, the pupillary distance of adults can typically vary between 54 mm and 68 mm. Other estimated dimensions include the average horizontal visible iris diameter (HVID), which can range from 11.6 mm to 12.0 mm for approximately 50% of the population. Some estimates indicate a greater variability of 12% in the average HVID. Unfortunately, the accuracy of this variability is similar to that of the interocular distance, and thus the tolerance is not high enough to produce accurate image scaling. The tolerance generated by using the estimated dimensions is unacceptable for personal and / or near-field audio systems (such as headrest speaker systems, etc.) that need to account for the user's head movement.
[0007] As shown previously, what is needed in the art is improved camera-based tracking for audio systems and the like. Summary of the Invention
[0008] Various embodiments disclose a computer-implemented method for audio processing based on a user's head pose, the method comprising: acquiring one or more images of the user; processing the one or more images to identify a plurality of facial feature points representing positions on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; determining a three-dimensional localization of the user's ears based on the set of feature point pairs; and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional localization of the ears.
[0009] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the accuracy of camera-based head and / or ear tracking is improved. The improved camera-based tracking provides an improved noise cancellation, improved crosstalk cancellation, and other improved three-dimensional audio listening experiences for users of personal and / or near-field audio systems such as headrest audio systems, seat / chair audio systems, soundbars, vehicle audio systems, etc. The described technology also enables tracking of user-specific scaled ear localization in three dimensions using images from a single standard imaging camera. These technical advantages represent one or more technical improvements over prior art methods. Brief Description of the Drawings
[0010] In order to be able to understand in detail the manner in which the above-described features of the various embodiments are realized, the inventive concept briefly outlined above can be described more specifically by reference to the various embodiments, some of which are shown in the drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be considered in any way to limit the scope, and there are other equally effective embodiments.
[0011] Figure 1is a schematic diagram showing a computing system according to various embodiments;
[0012] Figure 2 is according to various embodiments showing Figure 1 an operation diagram of a tracking application;
[0013] Figure 3 is according to various embodiments showing Figure 1 a diagram of exemplary sets of feature point pairs included in a general head geometry;
[0014] Figure 4 is according to various embodiments showing Figure 1 a diagram of an exemplary set of feature point coordinates of a series of facial feature points;
[0015] Figure 5 is a flowchart of method steps for providing a user - specific ear position to an audio application according to various embodiments; and
[0016] Figure 6 is a flowchart of method steps for generating a user - specific ear position according to various embodiments. DETAILED DESCRIPTION
[0017] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.
[0018] Figure 1 is a schematic diagram showing a computing system 100 according to various embodiments. As shown, the computing system 100 includes, but is not limited to, a computing device 110, one or more cameras 150, and one or more speakers 160. The computing device 110 includes, but is not limited to, one or more processing units 112 and one or more memories 114. In various embodiments, an interconnect bus (not shown) connects the processing unit 112, the memory 114, the camera 150, the speaker 160, and / or any other components of the computing device 110. The memory 114 stores, but is not limited to, a tracking application 120, an audio application 130, one or more facial feature points 140, an intrinsic matrix 142, a general head geometry 144, and a registered head geometry 146. The tracking application 120 includes, but is not limited to, a face detection and feature point estimation model 122, a head pose estimation module 124, a scaling module 126, and an ear estimation module 128. Although shown as components of the tracking application 120, the head pose estimation module 124, the scaling module 126, and / or the ear estimation module 128 may include executable instructions that work in cooperation with the tracking application 120 as sub - components and / or separate software modules.
[0019] In operation, the computing system 100 processes two-dimensional (2D) image data 152 captured using one or more cameras 150 and performs user-specific image scaling to determine the position of the user's ears within the environment. The two-dimensional image data 152 includes one or more images of the user captured by the camera 150. The camera 150 continuously captures images of the user over time. In some embodiments, the tracking application 120 uses a face detection and feature point estimation model 122 to successfully detect a face from the two-dimensional image data 152. The face detection and feature point estimation model 122 includes a machine learning model, a rule-based model, or another type of model that takes the two-dimensional image data 152 as input and generates face feature points 140 and / or a face detection status. The head pose estimation module 124 uses the face feature points 140 to estimate the head pose and the depth of the head relative to the camera 150 within the environment. The scaling module 126 selects one or more sets of feature point pairs from the face feature points 140 and uses the selected feature point pairs to scale the head and convert the pixel coordinates of the multiple sets of feature point pairs within the 2D image data 152 into three-dimensional coordinates within the environment. The ear estimation module 128 uses the three-dimensional coordinates of the multiple sets of feature point pairs to estimate the three-dimensional coordinates of the user's ears within the environment. The tracking application 120 can transmit the estimated position of the user's ears to the audio application 130 for generating a processed audio signal and / or sound field.
[0020] In various embodiments, the computing device 110 can be included in a vehicle system, a home theater system, a soundbar, etc. In some embodiments, the computing device 110 is included in one or more devices such as consumer products ( For example , portable speakers, gaming products, etc.), vehicles ( For example , head units of cars, trucks, vans, etc.), smart home devices ( For example , smart lighting systems, security systems, digital assistants, etc.), communication systems ( For example , conference call systems, video conferencing systems, speaker amplification systems, etc.), etc. In various embodiments, the computing device 110 is located in various environments, including but not limited to indoor environments ( For example , living rooms, conference rooms, convention halls, home offices, etc.) and / or outdoor environments ( For example , patios, rooftops, gardens, etc.). The computing device 110 is also capable of providing an audio signal ( For example , generated using the audio application 130) to one or more speakers 160 to generate a sound field providing various audio effects.
[0021] One or more processing units 112 can be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), and / or any other type of processing unit or combination of different processing units, such as a CPU and / or DSP configured to operate in conjunction with a GPU. The processing unit 112 generally includes one or more programmable processors that execute program instructions to manipulate input data. In some embodiments, the processing unit 112 can include any number of processing cores, memories, and other modules for facilitating program execution. When executing program instructions, the processing unit 112 communicates with a user and / or one or more external devices via an I / O interface and one or more I / O devices (not shown). When executing an application program, the processing unit 112 can also exchange messages with one or more remote devices via a network interface (not shown).
[0022] The memory 114 can include random access memory (RAM) modules, flash memory cells, or any other type of memory cells or a combination thereof. The processing unit 112 is configured to read data from and write data to the memory 114. In various embodiments, the memory 114 includes non-volatile memory, such as an optical drive, a magnetic drive, a flash drive, or other storage. In some embodiments, a separate data repository, such as an external data repository (“cloud storage”) included in a network, can supplement the memory 114. One or more memories 114 store, but are not limited to, the tracking application 120, the audio application 130, and the generic head geometry 144 and the registered head geometry 146. The tracking application 120 within the memory 114 can be executed by the processing unit 112 to implement the overall functionality of the computing device 110 and thus coordinate the operation of the computing system 100 as a whole.
[0023] The memory 114 stores executable programs (including the tracking application 120 and the audio application 130) and data (such as the intrinsic matrix 142, the generic head geometry 144, and the registered head geometry 146). The intrinsic matrix 142 includes data values associated with the camera 150 that generates the 2D image data 152. The data values can include, for example, the sensor size of the camera, the focal length of the camera, the principal point ( For example , reference position), etc. The generic head geometry 144 is a model of a human head that includes the three-dimensional positions of feature points representing the various parts of a human face. The registered head geometry 146 is a model of a specific human head that includes the three-dimensional positions of feature points representing the various parts of a specific face. The positions of the feature points on the registered head geometry 146 have been modified from the generic head geometry 146 based on a registration process that scales and adjusts the three-dimensional positions of the feature points.
[0024] The tracking application 120 processes 2D image data 152 received from one or more cameras 150 and generates data including an estimate of the user's ear position and / or head pose. The estimated ear position and / or the estimated head pose can be used by one or more other applications. For example, the audio application 130 can use the estimated ear position and / or the estimated head pose to modify an audio signal that, when output by one or more speakers 160, generates a sound field within the environment. For example, the tracking application 120 can process a given frame of the 2D image data 152 to identify a plurality of facial feature points 140 and determine the number of pixels between pairs of feature points included in the plurality of facial feature points 140. The tracking application 120 can then generate a depth estimate and use the depth estimate and the facial feature points 140 to generate a scaled, user-specific, and three-dimensional estimated ear position of the user.
[0025] The tracking application 120 provides the estimated ear position and, in some embodiments, the estimated head pose to the audio application 130. Thus, the tracking application 120 performs camera-based head and / or ear positioning tracking based on two-dimensional image data 152 captured using one or more cameras 150. The audio application 130 uses the estimated ear position (and optionally, the estimated head pose), speaker configuration data, and / or an input audio signal to generate a set of modified and / or processed audio signals that affect the sound field and / or provide various adaptive audio effects. The adaptive audio effects can include, for example, noise cancellation, crosstalk cancellation, spatial / positional audio effects, etc., where the adaptive audio effects adapt to the user's ear positioning. For example, the audio application 130 can identify one or more head-related transfer functions (HRTFs) based on the ear position, head orientation, and / or speaker configuration. In some examples, the HRTF is ear-specific for each ear. The audio application 130 modifies one or more speaker-specific audio signals based on the HRTF to maintain a desired audio effect that dynamically adapts to the ear positioning. In some embodiments, the audio application 130 generates a set of modified or processed audio signals corresponding to a set of speakers 160. In various embodiments, the tracking application 120 provides the scaled user-specific ear position in real time ( For example , within 10 ms) or near real time ( For example , within 100 ms).
[0026] One or more cameras 150 include various types of cameras for capturing two-dimensional images of a user. The one or more cameras 150 include cameras for a driver monitoring system (DMS) in a vehicle, a sound bar, a web camera, and the like. In some embodiments, the one or more cameras 150 include only a single standard two-dimensional imager without stereo or depth capabilities. In some embodiments, in addition to the one or more cameras 150, the computing system 100 further includes other types of sensors for obtaining information about the acoustic environment. Other types of sensors include, but are not limited to, motion sensors (such as an accelerometer or an inertial measurement unit (IMU)( For example , a three-axis accelerometer, a gyroscope sensor, and / or a magnetometer)), a pressure sensor, and the like. Additionally, in some embodiments, the sensors may include wireless sensors (including radio frequency (RF) sensors( For example , sonar, and radar)) and / or wireless communication protocols (including Bluetooth, Bluetooth Low Energy (BLE), cellular protocols, and / or Near Field Communication (NFC)).
[0027] One or more speakers 160 include various speakers for outputting audio to form a sound field or various audio effects near the user. In some embodiments, the speakers 160 include two or more speakers located in the headrest of a seat (such as a vehicle seat or a gaming chair), or another user-specific speaker group that is connected or positioned for use by a single user (such as a personal and / or near-field audio system). In some embodiments, the speakers 160 are associated with a speaker configuration stored in the memory 114. The speaker configuration indicates the position and / or orientation of the speakers 160 in three-dimensional space and / or relative to each other and / or relative to the vehicle, the vehicle seat, the gaming chair, the position of the camera 150, and the like. The audio application 130 can retrieve or otherwise identify the speaker configuration of the speakers 160.
[0028] Figure 2 is a diagram showing Figure 1 the operation 200 of the tracking application 120 according to various embodiments. As shown, the operation 200 includes, but is not limited to, 2D image data 152, the tracking application 120, an intrinsic matrix 142, a generic head geometry 144, a registered head geometry 146, and user-specific ear positions 210. The tracking application 120 includes and / or utilizes, but is not limited to, one or more two-dimensional feature point coordinates 202, a head pose estimation 204, one or more feature point depth estimations 206, and one or more three-dimensional (3D) feature point coordinates 208.
[0029] In operation, tracking application 120 receives and / or retrieves and processes 2D image data 152. In various embodiments, tracking application 120 processes 2D image data 152 as a series of frames, where tracking application 120 generates user-specific ear positions 210 for each frame. Additionally or alternatively, tracking application 120 can process 2D image data 152 in real time and continuously update the user-specific ear positions 210 for each frame included in the 2D image data 152. In various embodiments, when generating the user-specific ear positions 210, tracking application 120 can produce intermediate data including but not limited to 2D feature point coordinates 202, head pose estimation 204, feature point depth estimation 206, and 3D feature point coordinates 208. Tracking application 120 generates data including but not limited to user-specific ear positions 210.
[0030] In various embodiments, tracking application 120 uses a face detection and feature point estimation model 122 to generate 2D feature point coordinates 202 for a given frame of 2D image data 152. In such cases, tracking application 120 can use the face detection and feature point estimation model 122 and / or a general head geometry to generate the 2D feature point coordinates 202 as pixel coordinates within the frame. The 2D feature point coordinates 202 are the two-dimensional positions of one or more pixels of one or more feature points of various parts of the head. Such feature points can include but not limited to one or more eye feature points ( For example , center, outer point, inner point, etc.), one or more eyebrow feature points ( For example , outer point, inner point, midpoint, etc.), one or more nose feature points ( For example , bridge of the nose, tip of the nose, bottom of the nose, root of the nose / radix, glabella, etc.), one or more mouth feature points ( For example , left point, right point, upper midpoint, lower midpoint, etc.), one or more jaw feature points ( For example , left point, right point, upper midpoint, lower midpoint, etc.), and one or more ear feature points ( For example , left ear canal, right ear canal, etc.). In some embodiments, the face detection and feature point estimation model 122 does not provide one or more ear feature points. The face detection and feature point estimation model 122 stores the 2D feature point coordinates 202 or otherwise provides the 2D feature point coordinates to a head pose estimation module 124. In some embodiments, the face detection and feature point estimation model 122 and a registered head geometry 146 generate feature point pairs (in two-dimensional space and three-dimensional space respectively), the feature point pairs including but not limited to the edges of two eyes ( For example , outer edges), a nose bridge to chin feature point pair, a nose bridge to right jaw feature point pair, a left eye inner edge to left jaw feature point pair, and a right eyebrow inner edge to left jaw pair.
[0031] The head pose estimation module 124 generates a head pose estimation 204 based on the 2D feature point coordinates 202. In some embodiments, the head pose estimation module 124 also uses the generic head geometry 144 to generate the head pose estimation. The generic head geometry 144 includes the three-dimensional positions of the feature points that correspond to the two-dimensional positions of the 2D feature point coordinates 202. Thus, the head pose estimation module 124 can analyze the 2D feature point coordinates 202 and the generic head geometry 144 to generate a head pose estimation 204, such as a three-dimensional orientation vector. The orientation vector enables the scaling module 126 to more accurately identify the feature point depth estimation 206. In some embodiments, the orientation vector can be expressed as values for three axes that represent the pitch angle, yaw angle, and roll angle relative to a reference orientation. In such cases, the head pose estimation module 124 can select multiple feature points to determine the head pose estimation 204.
[0032] For example, the head pose estimation module 124 can select six feature points from the generic head geometry 144 and determine the 2D feature point coordinates 202 of the selected feature points within the 2D image data. The head pose estimation module 124 can then use the three-dimensional coordinates of the feature points in the generic head geometry 144 to estimate the head pose along three axes. In this way, the head pose estimation module 124 can select six feature points to represent the six degrees of freedom (DOF) of the head pose. In various embodiments, the head pose estimation module 124 can perform various head pose estimation algorithms, such as the Perspective-n-Point (PnP) method or model registration.
[0033] In various embodiments, the scaling module 126 generates a feature point depth estimation 206 for the corresponding feature point coordinates and / or feature point coordinate pairs in the 2D feature point coordinates 202. In some embodiments, the feature point pairs can include a bridge of nose to chin pair, a glabella to chin pair, a glabella to nasal base pair, or another pair that is predominantly vertical ( For example , having the largest difference between the coordinates in the vertical dimension). The feature point pairs can include an eye to eye pair, a jaw to jaw pair, or another pair that is predominantly horizontal ( For example , having the largest difference between the coordinates in the horizontal dimension). Any set of feature point pairs can be used to generate the head pose estimation 204 and / or the feature point depth estimation 206. Greater accuracy is achieved for feature point pairs that have a greater distance between them. Thus, in some embodiments, the bridge of nose to chin pair or the glabella to chin pair can provide greater accuracy.
[0034] In some embodiments, the tracking application 120 generates a scaled registered head geometry based on one or more registered images selected from the two-dimensional image data 152. The tracking application 120 can include a machine learning model, a rule-based model, or another type of model that takes the two-dimensional image data 152 as input and generates the registered head geometry 146. In some embodiments, the input to the tracking application for generating the registered head geometry 146 can be the intrinsic matrix 142, which includes data such as the sensor size of the camera, the focal length of the camera, etc. Once registration is performed, the tracking application 120 uses the registered head geometry 146 and the "live" or most recently captured two-dimensional image data 152 to identify the user-specific ear positions.
[0035] In various embodiments, the tracking application 120 performs a non-real-time registration process to generate a user-specific registered head geometry 146. During registration, the tracking application 120 selects one or more registered images from the 2D image data 152 as a subgroup of images associated with a particular head orientation or range of head orientations. In some embodiments, the tracking application 120 includes a rule-based module or program that detects or identifies the face orientation, such as the face orientation state ( For example , centered and / or facing forward, facing up, facing down, facing left, facing right). Registration includes a relatively short time period ( For example , within 30 seconds, within 45 seconds, or within 1 minute) that does not need to be performed in real time. The registered head geometry 146 includes the three-dimensional positions of one or more feature points ( For example , positioning or positions in three dimensions). Feature points include but are not limited to each eye ( For example , center, outer point, inner point, etc.), each eyebrow ( For example , outer point, inner point, midpoint, etc.), the nose ( For example , bridge of the nose, tip of the nose, bottom of the nose, root of the nose / radix, glabella, etc.), the mouth ( For example , left point, right point, upper midpoint, lower midpoint, etc.), the mandible ( For example , left point, right point, upper midpoint, lower midpoint, etc.), each ear ( For example , ear canal, etc.), etc. Relative to using the generic head geometry 144, the generated registered head geometry 146 provides improved accuracy of the ear positions. The improved ear position accuracy provides a significant improvement in the user experience, especially for personal and / or near-field audio systems, such as headrest audio systems.
[0036] The scaling module 126 generates a feature point depth estimate 206 using one or more of the 2D feature point coordinates 202, the head pose estimate 204, the registered head geometry 146, and / or the intrinsic matrix 142. The feature point depth estimate 206 can be considered a scaling factor for scaling the 3D feature point coordinates 208. The scaling module 126 generates the feature point depth estimate 206 based on Equation (1).
[0037]
[0038] In Equation (1), the focal length of the camera 150 is denoted as f. The 3D distance between a pair of feature point coordinates in the environment is denoted as W. The distance between a pair of 2D feature point coordinates 202 in the image ( For example , the two-dimensional image data 152) is denoted as w. The feature point depth estimate 206 for a particular 2D feature point coordinate 202 or a pair of feature points is denoted as d est . The focal length of the camera 150 is determined during a non-real-time camera calibration process and is included in the intrinsic matrix 142. The distance w in the image frame can be indicated in terms of the number of pixels, and / or can be generated by multiplying the number of pixels by the physical width of each pixel. In some embodiments, the physical width of each pixel can be included in the intrinsic matrix 142. Equation (1) considers an example where the line connecting a pair of 2D feature point coordinates 202 is orthogonal to the direction the camera 150 is facing. However, by accounting for any relationship between the 2D feature point coordinates 202, the scaling module 126 can use the head pose estimate 204 to improve the accuracy of each feature point depth estimate 206.
[0039] In some embodiments, the scaling module 126 can use the registered head geometry 146 to generate the feature point depth estimate 206. Additionally or alternatively, in some embodiments, the scaling module 126 selects a particular pair of feature points and generates the feature point depth estimate 206. In such cases, the scaling module 126 can select the pair of feature points based on the head pose estimate 206. For example, the scaling module 126 can use the head pose estimate 206 ( For example , an orientation vector classified as "looking left") to select one or more visible pairs of feature points. The scaling module 126 can then generate the feature point depth estimate for each feature point included in the selected pair of feature points.
[0040] The scaling module 126 uses the 2D feature point coordinates 202 and the corresponding feature point depth estimate 206 to generate the 3D feature point coordinates 208. In some embodiments, the 3D feature point coordinates 208 are generated using Equations (2)-(4).
[0041] X = (X img - P x ) * d est(2)
[0042] Y = (Y img - P y ) * d est (3)
[0043] Z = d est (4)
[0044] In equations (2)-(4), X, Y, and Z are 3D feature point coordinates 208 corresponding to three-dimensional feature points in three-dimensional space. X img and Y img are 2D feature point coordinates 202 generated by the face detection and feature point estimation model 122. P x and P y respectively represent the principal points in the X-axis and Y-axis and serve as reference points within the environment. In various embodiments, the camera calibration process determines the values of P x and P y and the determined values are included in the intrinsic matrix 142.
[0045] The ear estimation module 128 generates a user-specific ear position 210 based on the 3D feature point coordinates 208. The ear estimation module 128 transforms one or more 3D feature point coordinates 208 into a user-specific ear position 210 by applying an ear relationship 220 to the 3D feature point coordinates 208 indicated as the starting point of the ear relationship 220. The ear relationship 220 includes a set of three-dimensional ear relationship vectors, one vector for each ear. Each ear relationship 220 includes a magnitude and a three-dimensional direction. The ear estimation module 128 calculates the ear position by setting the initial or starting point of the ear relationship 220 at the 3D feature point coordinates 208 of a specific feature point in the registered head geometry 146. The ear estimation module 128 identifies the user-specific ear position 210 as the three-dimensional coordinates at the position of the end point or terminus of the ear relationship 220.
[0046] In some embodiments, the ear estimation module 128 also uses the head pose estimation 204. For example, the ear estimation module 128 can rotate the ear relationship 220 around a predetermined point in three-dimensional space. If the user is looking left (or right), the ear relationship 220 is different from the case where the user is looking straight ahead. In cases where the 2D feature point coordinates 202 and the 3D feature point coordinates 208 include ear feature points, no feature point-to-ear transformation is performed. Instead, the tracking application 120 uses the 3D feature point coordinates 208 of the ear feature points as the user-specific ear position 210. Although a general ear position e can be generated by applying a general ear relationship to a general head geometry 144, the result is not scaled for the user. In contrast, the ear relationship 220 and / or the 3D feature point coordinates 208 can be scaled for the user.
[0047] Figure 3 According to various embodiments, Figure 1 FIG300 is a diagram of exemplary multiple sets of feature point pairs included in a general head geometry of FIG302. As shown, the diagram includes an image of a head 302, a plurality of facial feature points 140, and a plurality of feature point pairs 310-350. The plurality of feature point pairs include a nose bridge to chin pair 310, an eyebrow to eyebrow pair 320, a first eye inner edge to jaw pair 330, a nose bridge to jaw pair 340, and a second eye inner edge to jaw pair 350.
[0048] In various embodiments, the tracking application 120 selects one or more feature point pairs for use in generating the head pose estimate 204 and / or the user-specific ear locations 210. For example, the head pose estimation module 124 may select multiple feature point pairs to generate the feature point depth estimate 206. In such a case, the head pose estimation module 124 may select three feature point pairs, including a primary vertical feature point pair ( For example , nose bridge to chin pair 310), mainly horizontal pair ( For example , eyebrow to eyebrow pair 320) and / or another feature point pair visible in the frame of the 2D image data 152 ( For example , first eye inner edge to jaw pair 330). Similarly, the scaling module 126 may select five feature point pairs, such as each of the feature point pairs 310-350.
[0049] In various embodiments, the head pose estimation module 124 and / or the scaling module 126 may select feature point pairs from the plurality of facial feature points 140 based on criteria to increase the accuracy of scaling. For example, the criteria may include whether each of the facial feature points 140 in a given feature point pair is in the field of view of the camera 150. When the tracking application 120 does see a given facial feature point 140, either the tracking application 120 cannot detect the facial feature point 140 or the camera cannot see the facial feature point 140. In such cases, the positioning error when using the feature point pair will be higher. Therefore, the head pose may affect the accuracy of facial feature point detection and subsequent use. Another criterion is to use the maximum possible separation distance between the facial feature points 140 in the feature point pair. In various embodiments, any error in feature point detection directly affects the error of subsequent estimation. For example, when the tracking application 120 marks the facial feature point 140 with an error of one pixel, a longer feature point pair separation distance may produce a lower percentage error. Therefore, the tracking application 120 may apply a preference to sort one or more feature point pairs by separation distance and select feature point pairs with a high separation distance or a separation distance exceeding a threshold.
[0050] In various embodiments, the tracking application 120 can select feature point pairs to track and process based on the user's head pose. In such cases, the tracking application 120 can use the head pose estimation 206 to select one or more feature point pairs corresponding to the head pose. For example, the tracking application 120 can generate a mapping table that identifies specific feature point pairs for each of the multiple head poses ( For example , classification of the head pose for a given orientation vector). The head pose can be classified based on the overall orientation of the head, such as "looking left", "looking right", "looking up", "looking down", "looking straight ahead". These head pose classifications can be within a specific angular range ( For example , within 45° of the principal point). In some embodiments, the tracking application 120 tracks a larger range of head movement ( For example , within 20° of the pitch angle along the principal point, within 45° of the yaw angle along the principal point, etc.). In such cases, the tracking application can classify additional head poses associated with additional orientation angle ranges.
[0051] For example, the tracking application 120 can select a single feature point pair based on the head pose. In such cases, the tracking application can generate a mapping table that specifies one of the feature point pairs 310 - 350 for each head pose. In some embodiments, the tracking application can update the mapping table with the 2D feature point coordinates of the corresponding facial feature points 140 included in the feature point pair. For example, Table 1 shows a single feature point pair corresponding to multiple head poses.
[0052] Head pose Feature point pair Look left Bridge of nose to chin pair 340 Look right Inner edge of eye to chin pair 330 Look up Inner edge of eye to chin pair 350 Look down Eyebrow to eyebrow pair 320 Look straight ahead Bridge of nose to chin pair 310
[0053] Table 1: Single Feature Point Pairs for Different Head Poses
[0054] In some embodiments, the scaling module 126 selects more than one feature point pair for each head pose. In such cases, the multiple feature point pairs reduce the error introduced by averaging the feature point distances by the scaling module 126. Because the tracking application 120 can introduce errors when determining the feature point distances by mislabeling the facial feature points 140 as incorrect pixels (resulting in pixel positions that deviate by one pixel in X or Y, similar to jitter). The scaling module 126 can average one or more scaling factors derived from multiple pairs of feature point pairs, thereby reducing or eliminating jitter and producing a more accurate scaling factor. Therefore, the increased accuracy of the scaling factor results in higher accuracy when the tracking application 120 tracks the user's head pose and ears. For example, Table 2 shows multiple feature point pairs corresponding to each of the multiple head poses.
[0055]
[0056]
[0057] Table 2: Multiple feature point pairs for different head poses
[0058] Additionally or alternatively, in various embodiments, the tracking application 120 applies a confidence score to each facial landmark 140. In such cases, the head pose estimation module 124 and / or the scaling module 126 may rank or sort the landmark pairs based on the scores of the included facial landmarks 140 and select the highest ranked or sorted landmark pair. In some embodiments, the tracking application may generate a confidence score for a given facial landmark 140 based on various criteria, such as the amount of movement of the facial landmark 140 detected from a previous frame, the confidence scores of neighboring facial landmarks 140, and the like.
[0059] Figure 4 According to various embodiments, Figure 1 FIG. 4 is a diagram of an exemplary set of feature point coordinates 400 of a series of facial feature points. As shown, the set of feature point coordinates 400 includes a common feature point pair 410, a set of feature point regions 422-444, and the set of feature point regions includes a nose feature point region 422, a mouth feature point region 424, eye feature point regions 432 and 434, and eyebrow feature point regions 442 and 444.
[0060] In various embodiments, the tracking application 120 can select one or more feature point pairs for processing independently of the head pose estimate 206. For example, the tracking application 120 can specify that a universal feature point pair 410 of facial feature points 140 including the bridge of the nose (coordinate 27) and the chin (coordinate 8) can be used. In such cases, the tracking application 120 can use the universal feature point pair 410 without determining the head pose estimate 206. In some embodiments, the tracking application 120 selects the universal feature point pair 410 based on criteria such as visibility of the included facial feature points 140 in multiple head poses, accurate labeling of the included facial feature points 140, and / or distances between facial feature points 140.
[0061] In some embodiments, the tracking application 120 selects a single group of multiple feature point pairs for all head poses. In such cases, the tracking application 120 averages the selected feature point pairs included in the group to estimate the user-specific ear position 210. The tracking application 120 can select each feature point pair to be included in the group based on criteria such as the visibility of the included facial feature points 140 across multiple head poses, the accurate labeling of the included facial feature points 140, and / or the distance between the facial feature points 140. For example, the tracking application 120 can generate a mapping table that includes multiple feature point pairs representing the same feature point region. For example, Table 3 shows a group of feature point pairs that include dissimilar feature point pairs between the eyebrow feature point region 442 ( For example , coordinates 17 - 21) and a portion of the mouth feature point region 424 ( For example , coordinates 48 - 67).
[0062]
[0063] Table 3: Feature Point Pairs of Method 5
[0064] Figure 5 is a flowchart of method steps for providing a user-specific ear position to an audio application according to various embodiments. Although the method steps are shown in sequence, those skilled in the art will understand that some method steps can be performed in a different order, repeated, omitted, and / or performed by components other than those described in Figure 5 . Although the method steps are described with respect to the Figures 1 to 4 system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the various embodiments.
[0065] As shown, method 500 begins at step 502, where the tracking application 120 retrieves the intrinsic matrix 142. In various embodiments, the tracking application 120 retrieves the intrinsic matrix 142 from the memory 114. The intrinsic matrix 142 includes one or more data values associated with the camera 150 used to acquire the 2D image data 152. For example, the computing device 110 or another device can perform one or more camera calibration methods to determine the focal length and / or principal point for generating the feature point depth estimate 206. In such cases, the values of the focal length and / or principal point can be included in the intrinsic matrix 142 and retrieved by the tracking application 120 when performing real-time estimation of the user's head pose and / or ear position.
[0066] At step 504, tracking application 120 optionally performs registration on the user's head to generate a registered head geometry 146. In various embodiments, tracking application 120 performs a non-real-time registration process to generate a scaled registered head geometry based on one or more registration images selected from 2D image data 152. During registration, tracking application 120 selects one or more registration images from 2D image data 152 as a subgroup of images associated with a particular head orientation or range of head orientations. The registered head geometry 146 includes the three-dimensional positions ( For example , localizations or positions in three dimensions) of one or more facial feature points 140. Relative to using a generic head geometry 144, the generated registered head geometry 146 provides improved accuracy of ear position. Once registration has been performed, tracking application 120 uses the registered head geometry 146 and "live" or most recently captured two-dimensional image data 152 to identify the user-specific ear position.
[0067] At step 506, tracking application 120 obtains 2D image data 152. In various embodiments, tracking application 120 retrieves 2D image data 152 from memory 114 and / or receives 2D image data 152 from camera 150. The two-dimensional image data 152 can include one or more images captured using camera 150. Tracking application 120 receives and / or retrieves updated two-dimensional image data 152 over time to provide a dynamically real-time updated user-specific ear position 210.
[0068] At step 508, tracking application 120 identifies 2D feature point coordinates 202. In various embodiments, tracking application 120 uses a face detection and feature point estimation model 122 to generate 2D feature point coordinates 202 based on 2D image data 152. The 2D feature point coordinates 202 are the two-dimensional positions of one or more facial feature points 140 on the user's head. For example, the facial feature points 140 can include one or more feature points at the positions of eyes, eyebrows, nose, mouth, jaw, ears, cheeks, etc.
[0069] At step 510, the tracking application 120 determines the head pose. In various embodiments, the head pose estimation module 124 of the tracking application 120 may process one or more of the 2D feature point coordinates 202 to generate a head pose estimate 204. In some embodiments, the head pose estimation module 124 also uses the generic head geometry 144 to generate the head pose estimate. The generic head geometry 144 includes the three-dimensional positions of the feature points, which correspond to the two-dimensional positions of the 2D feature point coordinates 202. Thus, the head pose estimation module 124 can analyze the 2D feature point coordinates 202 and the generic head geometry 144 to generate a head pose estimate 204, such as a three-dimensional orientation vector. The orientation vector enables the scaling module 126 to more accurately identify the feature point depth estimate 206. In some embodiments, the orientation vector can be expressed as values for three axes, which represent the pitch angle, yaw angle, and roll angle relative to a reference orientation. In such cases, the head pose estimation module 124 may select multiple feature points to determine the head pose estimate 204.
[0070] At step 512, the tracking application 120 selects two or more 2D feature point coordinates 202. In various embodiments, the scaling module 126 of the tracking application 120 selects one or more feature point pairs for estimating the user's head pose and / or ear position. In such cases, the scaling module 126 may retrieve the 2D coordinates corresponding to each of the facial feature points 140 included in each of the selected feature point pairs. In some embodiments, the scaling module 126 determines the feature point depth estimate 206 for each of the facial feature points 140 included in the selected feature point pairs. In such cases, the scaling module 126 generates the feature point depth estimate 206 for the corresponding 2D feature point coordinates 202 based on the registered head geometry 146. The feature point depth estimate 206 can be considered a scaling factor for scaling the 3D feature point coordinates 208.
[0071] At step 514, the tracking application 120 generates a user-specific ear position 210 based on the 3D feature point coordinates 208. In various embodiments, the scaling module 126 converts the 2D feature point coordinates 202 into scaled and user-specific 3D feature point coordinates 208 for the user. The generated scaled 3D feature point coordinates 208 are based on the 2D feature point coordinates 202 and the corresponding feature point depth estimates 206. In such cases, the scaling module 126 can use information from the intrinsic matrix 142 (such as, the focal length value) to accurately generate the 3D feature point coordinates 208. The ear estimation module 128 of the tracking application then generates a user-specific ear position 210 based on the 3D feature point coordinates 208. The ear estimation module 128 transforms one or more 3D feature point coordinates 208 into an ear position by applying the ear relationship 220 to the 3D feature point coordinates 208 indicated in the ear relationship 220. However, in cases where the 2D feature point coordinates 202 and the 3D feature point coordinates 208 include ear feature points, the tracking application 120 identifies the 3D feature point coordinates 208 corresponding to the ear feature points and uses these 3D feature point coordinates 208 as the user-specific ear position 210.
[0072] At step 516, the audio application 130 generates a processed audio signal based on the user-specific ear position 210. In some embodiments, the audio application 130 further generates a processed audio signal based on the head pose estimate 204. For example, the audio application 130 identifies one or more HRTFs based on the user-specific ear position 210, the head pose estimate 204, and the speaker configuration. The audio application 130 generates a processed audio signal based on the HRTFs to maintain a desired audio effect that dynamically adapts to the user-specific ear position 210. The processed audio signal is generated to create a sound field and / or provide various adaptive audio effects, such as noise cancellation, crosstalk cancellation, spatial / location audio effects, etc., where the adaptive audio effects adapt to the user-specific ear position 210 of the user. In some embodiments, the computing system 100 includes one or more microphones, and the audio application 130 uses the microphone audio signal and / or other audio signals to generate the processed audio signal. In some embodiments, the audio application 130 further generates a processed audio signal based on the speaker configuration of a set of speakers 160.
[0073] At step 518, the audio application 130 provides the processed audio signal to the speaker 160. The speaker 160 generates a sound field based on the processed audio signal. Thus, the sound field includes one or more audio effects that adapt to the user in real time and dynamically based on the user-specific ear position 210 and the head pose estimate 204. The process returns to step 506 so that the sound field dynamically adapts based on the updated user-specific ear position 210 identified using the updated two-dimensional image data 152.
[0074] Figure 6 is a flowchart of method steps for generating a user-specific ear position according to various embodiments. Although the method steps are shown in sequence, those skilled in the art will understand that some method steps may be performed in a different order, repeated, omitted, and / or performed by components other than those described in Figure 6 Although the method steps are described with respect to the Figures 1 to 4 system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the various embodiments.
[0075] As shown, method 600 begins at step 602, where the tracking application 120 retrieves the intrinsic matrix 142. In various embodiments, the tracking application 120 retrieves the intrinsic matrix 142 from the memory 114. The intrinsic matrix 142 includes one or more data values associated with the camera 150 used to acquire the 2D image data 152. For example, the computing device 110 or another device may perform one or more camera calibration methods to determine the focal length and / or principal point for generating the feature point depth estimate 206. In such cases, the values of the focal length and / or principal point may be included in the intrinsic matrix 142 and retrieved by the tracking application 120 when performing real-time estimation of the user's head pose and / or ear position.
[0076] At step 604, the tracking application 120 determines whether to select a generic feature point pair. In various embodiments, the tracking application 120 may select one or more feature point pairs for processing, where the selection is independent of the head pose estimate 206. For example, the tracking application 120 may specify using the generic feature point pair 410 to estimate the user's ear position. When the tracking application 120 determines to use the generic feature point pair to be used, the tracking application proceeds to step 606. Otherwise, the tracking application 120 proceeds to step 610.
[0077] At step 606, tracking application 120 retrieves the coordinates of general feature point pair 410. In various embodiments, the scaling module 126 of tracking application 120 can use general feature point pair 410 without determining head pose estimate 206. In some embodiments, tracking application 120 selects general feature point pair 410 based on criteria such as the visibility of included face feature points 140 across multiple head poses, the accurate labeling of included face feature points 140, and / or the distance between face feature points 140. In such cases, scaling module 126 retrieves the 2D feature point coordinates 202 of each face feature point among the face feature points 140 included in general feature point pair 410.
[0078] At step 606, tracking application 120 determines whether to select feature points based on head pose. In various embodiments, tracking application 120 selects a single group of multiple feature point pairs for all head poses. In such cases, tracking application 120 averages the selected feature point pairs included in the group to estimate user-specific ear position 210 without the need to determine the user's head pose. When tracking application 120 determines not to select face feature points 140 based on head pose, tracking application 120 proceeds to step 612. Otherwise, tracking application 120 determines to select face feature points 140 based on head pose and proceeds to step 620.
[0079] At step 612, tracking application 120 retrieves the coordinates of the multiple feature points included in a single group of multiple feature point pairs. In various embodiments, the scaling module 126 of tracking application 120 can select each feature point pair to be included in the single group based on criteria such as the visibility of included face feature points 140 across multiple head poses, the accurate labeling of included face feature points 140, and / or the distance between face feature points 140. After selecting the multiple feature point pairs, scaling module 126 retrieves the 2D feature point coordinates 202 corresponding to each face feature point among the corresponding face feature points 140 included in the group.
[0080] At step 620, tracking application 120 determines whether to apply a confidence score to the feature point pairs. In various embodiments, tracking application 120 can determine a confidence score for each face feature point 140. In such cases, head pose estimation module 124 and / or scaling module 126 can rank or order the feature point pairs based on the confidence scores of the included face feature points 140. When tracking application 120 determines to apply a confidence score, tracking application 120 proceeds to step 622. Otherwise, tracking application 120 determines not to apply a confidence score and proceeds to step 624.
[0081] At step 622, the tracking application 120 applies confidence scores to the corresponding facial feature points 140. In such cases, the head pose estimation module 124 and / or the scaling module 126 may select the highest ranked or sorted feature point pairs. In some embodiments, the tracking application may generate confidence scores for a given facial feature point 140 based on various criteria such as the amount of movement of the facial feature point 140 detected from a previous frame, the confidence scores of adjacent facial feature points 140, and the like.
[0082] At step 624, the tracking application 120 estimates the user's head pose. In various embodiments, the head pose estimation module 124 of the tracking application 120 may generate a head pose estimation based on the 2D feature point coordinates 202 and / or the values included in the intrinsic matrix 142. In some embodiments, the head pose estimation module 124 also uses the generic head geometry 144 to generate the head pose estimation 204. The generic head geometry 144 includes the three-dimensional positions of the feature points that correspond to the two-dimensional positions of the 2D feature point coordinates 202. Thus, the head pose estimation module 124 can analyze the 2D feature point coordinates 202 and the generic head geometry 144 to generate a head pose estimation 204, such as a three-dimensional orientation vector. The orientation vector enables the scaling module 126 to more accurately identify the feature point depth estimation 206. In some embodiments, the orientation vector may be expressed as values for three axes that represent the pitch angle, yaw angle, and roll angle relative to a reference orientation. In such cases, the head pose estimation module 124 may select multiple feature points to determine the head pose estimation 204.
[0083] At step 630, the tracking application 120 determines whether to select multiple feature point pairs based on the head pose. In various embodiments, the tracking application 120 selects one or more feature point pairs for use in generating the head pose estimation 204 and / or the user-specific ear position 210. For example, the tracking application 120 may determine whether to select the feature point pairs to be tracked and processed based on the user's head pose. In such cases, the tracking application 120 may use the head pose estimation 206 to select one or more feature point pairs corresponding to the head pose. In some embodiments, the scaling module 126 selects more than one feature point pair for each head pose. In such cases, the multiple feature point pairs reduce the estimation generated by averaging the feature point distances by the scaling module 126. When the tracking application 120 determines to select multiple feature point pairs, the tracking application 120 proceeds to step 632. Otherwise, the tracking application 120 determines to select a single feature point pair based on the head pose estimation 204 and proceeds to step 634.
[0084] At step 632, tracking application 120 retrieves the coordinates of multiple pairs of feature points corresponding to the head pose. In various embodiments, scaling module 126 retrieves the 2D feature point coordinates 202 of each corresponding facial feature point 140 of the selected pairs of feature points. In such cases, scaling module 126 may average one or more scaling factors derived from multiple pairs of feature points, thereby reducing or eliminating jitter and producing a more accurate scaling factor. Thus, the increased accuracy of the scaling factor results in higher accuracy when tracking application 120 tracks the user's head pose and ears.
[0085] At step 634, tracking application 120 retrieves the coordinates of the selected pairs of feature points corresponding to the head pose. In various embodiments, scaling module 126 retrieves the 2D feature point coordinates 202 of the corresponding facial feature points 140 of the selected pairs of feature points. In some embodiments, tracking application 120 may generate a mapping table that identifies a specific pair of feature points for each of the multiple head poses ( For example , classification of the head pose for a given orientation vector). In such cases, tracking application 120 may generate a mapping table that designates one of the pairs of feature points for each head pose. Tracking application 120 may then retrieve the 2D feature point coordinates 202 of the corresponding facial feature points 140 included in the pair of feature points designated in the mapping table.
[0086] At step 640, tracking application 120 determines the user-specific ear position 210 based on the 3D feature point coordinates 208. In some embodiments, scaling module 126 determines a feature point depth estimate 206 for each facial feature point among the facial feature points 140 included in one or more selected pairs of feature points. In such cases, scaling module 126 generates a feature point depth estimate 206 of the corresponding 2D feature point coordinates 202 based on the registered head geometry 146. The feature point depth estimate 206 may be considered a scaling factor for scaling the 3D feature point coordinates 208. Scaling module 126 then uses information from the intrinsic matrix 142 (such as the focal length value) to accurately generate the 3D feature point coordinates 208. The ear estimation module 128 of the tracking application then generates the user-specific ear position 210 based on the 3D feature point coordinates 208. The ear estimation module 128 transforms one or more 3D feature point coordinates 208 into an ear position by applying the ear relationship 220 to the 3D feature point coordinates 208 indicated in the ear relationship 220.
[0087] In summary, techniques are disclosed for implementing image scaling using selected feature point pairs representing positions on a user's face. The techniques relate to a method that includes obtaining one or more images of a user and processing the one or more images to determine a three-dimensional localization of the user's ears. Tracking application 120 may use head geometry and values associated with a camera that captures two-dimensional images to estimate the position of the user's ears. In some embodiments, the tracking application selects one or more feature point pairs based on the estimated head pose of the user. Each head pose is mapped to one or more corresponding feature point pairs, where the corresponding feature point pairs may be in the view of the camera and accurately marked in the captured two-dimensional image. The tracking application determines the three-dimensional coordinates of the facial feature points and uses the three-dimensional coordinates to estimate the position of the user's ears relative to the facial feature points. The estimate of the ear position is transmitted to an audio device that processes one or more audio signals to generate one or more processed audio signals based on the three-dimensional position of the ears.
[0088] At least one technical advantage of the disclosed techniques over the prior art is that, using the disclosed techniques, the accuracy of camera-based head and / or ear tracking is improved. The improved camera-based tracking provides an improved noise cancellation, improved crosstalk cancellation, and other improved three-dimensional audio listening experiences for users of personal and / or near-field audio systems such as headrest audio systems, seat / chair audio systems, soundbars, vehicle audio systems, etc. The described techniques also enable tracking of user-specific scaled ear localization in three dimensions using a single standard imaging camera. These technical advantages represent one or more technical improvements over prior art methods.
[0089] 1. In various embodiments, a computer-implemented method for audio processing based on a user's head pose includes: obtaining one or more images of the user; processing the one or more images to identify a plurality of facial feature points representing positions on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; determining a three-dimensional localization of the user's ears based on the set of feature point pairs; and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional localization of the ears.
[0090] 2. The computer-implemented method of clause 1, wherein the set of feature point pairs includes at least two feature point pairs.
[0091] 3. The computer-implemented method of any one of clauses 1 or 2, further comprising averaging the three-dimensional localization of the ears.
[0092] 4. The computer-implemented method according to any one of clauses 1 to 3, further comprising: for each facial feature point included in the plurality of facial feature points, determining a confidence score associated with the position of the facial feature point in the one or more images of the user; sorting the set of one or more feature point pairs based on the confidence score to generate a sorted set of feature point pairs, wherein the one or more feature point pairs are selected based on the sorted set of feature point pairs.
[0093] 5. The computer-implemented method according to any one of clauses 1 to 4, wherein determining the three-dimensional positioning of the ear based on the set of feature point pairs of the user comprises: using a face detection model to generate two-dimensional feature point coordinates of the plurality of facial feature points based on the one or more images; and generating a feature point depth estimate of the two-dimensional feature point coordinates based on the estimated head pose; generating three-dimensional feature point coordinates based on the two-dimensional feature point coordinates and the feature point depth estimate, wherein the three-dimensional positioning of the ear is based on the three-dimensional feature point coordinates.
[0094] 6. The computer-implemented method according to any one of clauses 1 to 5, further comprising: determining a head pose vector based on the two-dimensional feature point coordinates of the plurality of facial feature points; and determining the feature point depth estimate based on the head pose vector.
[0095] 7. The computer-implemented method according to any one of clauses 1 to 6, further comprising: determining a head pose based on the two-dimensional feature point coordinates of the plurality of facial feature points, wherein the head pose is within 45 degrees of the principal point; and determining the feature point depth estimate based on the head pose vector.
[0096] 8. The computer-implemented method according to any one of clauses 1 to 7, wherein determining the three-dimensional positioning of the ear of the user based on one or more relationships in a registered head geometry, wherein the one or more relationships correlate the three-dimensional feature point coordinates with the three-dimensional positioning of the ear.
[0097] 9. The computer-implemented method according to any one of clauses 1 to 8, wherein the plurality of facial feature points includes one or more of the following: eye feature points, eyebrow feature points, nose feature points, glabella feature points, mouth feature points, chin feature points, or mandibular feature points.
[0098] 10. The computer-implemented method according to any one of clauses 1 to 9, wherein the one or more feature point pairs include at least one of the following: bridge of nose to chin feature point pair, bridge of nose to mandibular feature point pair, or eye edge to mandibular feature point pair.
[0099] 11. A computer-implemented method as described in any one of clauses 1 to 10, wherein one or more speakers generate an audio output from the processed audio signal; and the one or more speakers include one or more of the following: headrest speakers, gaming chair speakers, or soundbar speakers.
[0100] 12. A computer-implemented method as described in any one of clauses 1 to 11, wherein the one or more processed audio signals apply one or more audio effects to the one or more audio signals, and the audio effects include one or more of the following: spatial audio effects, noise cancellation, or crosstalk cancellation.
[0101] 13. In various embodiments, one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform audio processing based on a user's head pose by performing the following steps: obtaining one or more images of the user; processing the one or more images to identify a plurality of facial feature points representing positions on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; determining a three-dimensional localization of the user's ears based on the set of feature point pairs; and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional localization of the ears.
[0102] 14. The one or more non-transitory computer-readable media as described in clause 13, wherein the set of feature point pairs includes at least two feature point pairs.
[0103] 15. The one or more non-transitory computer-readable media as described in clause 13 or 14, further comprising averaging the three-dimensional localization of the ears.
[0104] 16. The one or more non-transitory computer-readable media as described in any one of clauses 13 to 15, further comprising: for each facial feature point included in the plurality of facial feature points, determining a confidence score associated with the position of the facial feature point in the one or more images of the user; sorting the set of one or more feature point pairs based on the confidence scores to generate a sorted set of feature point pairs, wherein the one or more feature point pairs are selected based on the sorted set of feature point pairs.
[0105] 17. One or more non-transitory computer-readable media as described in any one of clauses 13 to 16, wherein determining the three-dimensional positioning of the ear based on the set of feature point pairs of the user includes: using a face detection model to generate two-dimensional feature point coordinates of the plurality of face feature points based on the one or more images; generating a feature point depth estimate of the two-dimensional feature point coordinates based on the estimated head pose; and generating three-dimensional feature point coordinates based on the two-dimensional feature point coordinates and the feature point depth estimate, wherein the three-dimensional positioning of the ear is based on the three-dimensional feature point coordinates.
[0106] 18. One or more non-transitory computer-readable media as described in any one of clauses 13 to 17, further including: determining a head pose vector based on the two-dimensional feature point coordinates of the plurality of face feature points; and determining the feature point depth estimate based on the head pose vector.
[0107] 19. One or more non-transitory computer-readable media as described in any one of clauses 13 to 18, wherein the three-dimensional positioning of the user's ear is determined based on one or more relationships in the registered head geometry, and the one or more relationships correlate the three-dimensional feature point coordinates with the three-dimensional positioning of the ear.
[0108] 20. In various embodiments, a system includes: one or more speakers; a camera that captures one or more images of a user; a memory that stores instructions; and one or more processors that, when executing the instructions, are configured to perform audio processing based on a user's head pose by performing the following steps: obtaining one or more images of the user; processing the one or more images to identify a plurality of face feature points representing positions on the user's head; selecting a set of one or more feature point pairs from the plurality of face feature points based on the estimated head pose of the user; determining a three-dimensional positioning of the user's ear based on the set of feature point pairs; and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positioning of the ear.
[0109] Any and all combinations in any form of any of the claim elements described in any of the claims and / or any of the elements described in this application fall within the intended scope of the invention and protection.
[0110] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the embodiments.
[0111] Aspects of the present implementation can be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware implementation, an entirely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software aspects with hardware aspects, all of which may generally be referred to herein as a "module", "system", or "computer". Additionally, any hardware and / or software technologies, processes, functions, components, engines, modules, or systems described in the present disclosure may be implemented as a circuit or a collection of circuits. Moreover, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0112] Any combination of one or more computer-readable media may be utilized. A computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following media: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing media. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0113] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine. These instructions, when executed by the processor of the computer or other programmable data processing apparatus, enable the functions / actions specified in one or more blocks of the flowchart and / or block diagram to be implemented. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently depending upon the functionality involved, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0115] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which is determined by the claims that follow.
Claims
1. A computer-implemented method for performing audio processing based on a user's head posture, the computer-implemented method comprising: Obtain one or more images of the user; processing the one or more images to identify a plurality of facial landmarks representing locations on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; Determining a three-dimensional location of the user's ear based on the set of feature point pairs; as well as One or more audio signals are processed to generate one or more processed audio signals based on the three-dimensional positioning of the ear. 2 . The computer-implemented method of claim 1 , wherein the set of feature point pairs comprises at least two feature point pairs.
3. The computer-implemented method of claim 2, further comprising averaging the three-dimensional locations of the ears.
4. The computer-implemented method of claim 1 , further comprising: determining, for each facial feature point included in the plurality of facial feature points, a confidence score associated with a location of the facial feature point in the one or more images of the user; The set of one or more feature point pairs is sorted based on the confidence score to generate an sorted set of feature point pairs, wherein selecting the one or more feature point pairs is based on the sorted set of feature point pairs.
5. The computer-implemented method of claim 1 , wherein determining the three-dimensional location of the ear based on the set of feature point pairs of the user comprises: generating two-dimensional feature point coordinates of the plurality of facial feature points based on the one or more images using a face detection model; as well as generating a feature point depth estimate of the two-dimensional feature point coordinates based on the estimated head pose; Three-dimensional feature point coordinates are generated based on the two-dimensional feature point coordinates and the feature point depth estimates, wherein the three-dimensional location of the ear is based on the three-dimensional feature point coordinates.
6. The computer-implemented method of claim 5, further comprising: determining a head posture vector based on the two-dimensional feature point coordinates of the plurality of facial feature points; as well as The feature point depth estimate is determined based on the head pose vector.
7. The computer-implemented method of claim 5, further comprising: determining a head pose based on the two-dimensional feature point coordinates of the plurality of facial feature points, wherein the head pose is within 45 degrees of a principal point; as well as The feature point depth estimate is determined based on the head pose vector.
8. A computer-implemented method as described in claim 5, wherein the three-dimensional location of the ear of the user is determined based on one or more relationships in the registered head geometry, wherein the one or more relationships relate the three-dimensional feature point coordinates to the three-dimensional location of the ear.
9. The computer-implemented method of claim 1, wherein the plurality of facial feature points comprises one or more of the following: eye feature points, eyebrow feature points, nose feature points, forehead feature points, mouth feature points, chin feature points, or mandibular feature points.
10. The computer-implemented method of claim 1, wherein the one or more feature point pairs include at least one of: a nose bridge to chin feature point pair, a nose bridge to jaw feature point pair, or an eye edge to jaw feature point pair.
11. The computer-implemented method of claim 1 , wherein: one or more speakers generating an audio output from the processed audio signal; and The one or more speakers include one or more of: a headrest speaker, a gaming chair speaker, or a sound bar speaker.
12. The computer-implemented method of claim 1, wherein the one or more processed audio signals have one or more audio effects applied to the one or more audio signals, wherein the audio effects include one or more of: spatial audio effects, noise cancellation, or crosstalk cancellation.
13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform audio processing based on a user's head gesture by performing the following steps: Obtain one or more images of the user; processing the one or more images to identify a plurality of facial landmarks representing locations on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; Determining a three-dimensional location of the user's ear based on the set of feature point pairs; as well as One or more audio signals are processed to generate one or more processed audio signals based on the three-dimensional positioning of the ear.
14. The one or more non-transitory computer-readable media of claim 13, wherein the set of feature point pairs includes at least two feature point pairs.
15. The one or more non-transitory computer-readable media of claim 14, further comprising averaging the three-dimensional locations of the ears.
16. The one or more non-transitory computer-readable media of claim 13, further comprising: determining, for each facial feature point included in the plurality of facial feature points, a confidence score associated with a location of the facial feature point in the one or more images of the user; The set of one or more feature point pairs is sorted based on the confidence score to generate an sorted set of feature point pairs, wherein selecting the one or more feature point pairs is based on the sorted set of feature point pairs.
17. The one or more non-transitory computer-readable media of claim 13, wherein determining the three-dimensional location of the ear based on the set of feature point pairs of the user comprises: generating two-dimensional feature point coordinates of the plurality of facial feature points based on the one or more images using a face detection model; generating a feature point depth estimate of the two-dimensional feature point coordinates based on the estimated head pose; as well as Three-dimensional feature point coordinates are generated based on the two-dimensional feature point coordinates and the feature point depth estimates, wherein the three-dimensional location of the ear is based on the three-dimensional feature point coordinates.
18. The one or more non-transitory computer readable media of claim 17, further comprising: determining a head posture vector based on the two-dimensional feature point coordinates of the plurality of facial feature points; as well as The feature point depth estimate is determined based on the head pose vector.
19. One or more non-transitory computer-readable media as described in claim 17, wherein the three-dimensional position of the ear of the user is determined based on one or more relationships in the registered head geometry, wherein the one or more relationships relate the three-dimensional feature point coordinates to the three-dimensional position of the ear.
20. A system comprising: one or more speakers; a camera that captures one or more images of a user; a memory storing instructions; and One or more processors, which when executing the instructions are configured to perform audio processing based on the user's head posture by performing the following steps: Obtain one or more images of the user; processing the one or more images to identify a plurality of facial landmarks representing locations on the user's head; selecting a set of one or more feature point pairs from the plurality of facial feature points based on the estimated head pose of the user; Determining a three-dimensional location of the user's ear based on the set of feature point pairs; as well as One or more audio signals are processed to generate one or more processed audio signals based on the three-dimensional positioning of the ear.