Head and ear tracking using image scaling with emotion detection

By processing images and audio signals based on the user's 3D head geometry and emotions, generating emotions-specific 3D ear positions is solved, and the problems of inaccurate image scaling and no user emotions are not taken into account in the prior art, achieving more accurate head and ear tracking and a better audio experience.

CN120147113APending Publication Date: 2025-06-13HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411820834.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-06
Filing Date
2024-12-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing camera-based head and ear tracking system has problems such as inaccurate image scaling and no user mood consideration, which makes it difficult to adjust the sound experience changes and feature failure.

Method used

By acquiring the user's image, the image is processed to generate emotion-specific 3D ear positions based on the user's 3D head geometry and emotions, and the audio signal is processed to generate the processed audio signal based on the 3D position of the ear.

Benefits of technology

Improved accuracy of camera-based head and ear tracking, providing better noise cancellation, crosstalk cancellation and 3D audio listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147113A_ABST
    Figure CN120147113A_ABST
Patent Text Reader

Abstract

Techniques for image scaling based on emotion detection are described. In some embodiments, the technique includes: acquiring one or more images of a user; processing the one or more images to generate an emotion-specific 3D ear position of the user based on a three-dimensional (3D) head geometry of the user and an emotion, wherein the emotion is identified based on the one or more images of the user; and processing the one or more audio signals to generate one or more processed audio signals based on the three-dimensional position of the ear. Further embodiments include a system and a non-transitory computer readable medium that perform the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 608,524, filed on December 11, 2023, entitled "HEAD AND EAR TRACKING USING IMAGE SCALING WITH EMOTION DETECTION". The subject matter of this related application is hereby incorporated herein by reference. Technical Field

[0003] This application relates to systems and methods for head and ear tracking, and more particularly to head and ear tracking using image scaling with emotion detection. Background Art

[0004] Headrest audio systems, seat or chair audio systems, soundbars, in - vehicle audio systems, and other personal and / or near - field audio systems are becoming increasingly popular. However, when a listener moves their head, even slightly, the sound experienced by the user of a personal and / or near - field audio system can change significantly ( For example , 3 dB to 6 dB or another value). In the example of a headrest audio system, it can also vary significantly for different people depending on how the user is positioned in the seat and how the headrest is adjusted. This degree of sound pressure level (SPL) change is difficult to adjust the audio system. Additionally, when rendering spatial audio through headrest speakers, such changes cause features such as crosstalk cancellation to fail. One way to correct the audio of a personal and / or near - field audio system is camera - based head tracking. Due to the popularity of driver monitoring systems (DMS), consumer gaming (and other) soundbars including imaging devices and webcams, camera - based head tracking is making its way into vehicles, personal computer gaming, and home theaters.

[0005] However, one disadvantage of existing camera - based head and / or ear tracking is that camera images do not have a physical scale. Camera images consist of pixels where the size or scale of the objects represented therein cannot typically be set because the head or face in the image can be closer or farther depending on the seat position, standing position, and the positioning of another user. In systems that include scaling, conventional methods for performing the scaling of pixels to distance units involve using presumed sizes, such as the average inter - pupil distance ( For example, the distance between the pupils). However, the interpupillary distance of adults varies significantly between 54 mm and 68 mm. Other estimated dimensions include the average horizontal visible iris diameter (HVID), which can be between 11.6 mm and 12.0 mm, with an average value of 11.8 + / - 0.2 mm for 50% of the population. Some estimates indicate that the average HVID size is 11.6 + / - 12%, unfortunately, its accuracy is similar to the interpupillary distance, and thus the tolerance is not high enough to produce accurate image scaling. To account for the user's head movement, the tolerance generated by using the estimated dimensions is unacceptable for personal and / or near-field audio systems (such as headrest speaker systems, etc.).

[0006] Another disadvantage of existing camera-based head and / or ear tracking is that existing systems do not account for the user's mood ( For example , facial expressions, and / or movements). For example, when a user expresses emotion using their face, the distances between facial feature points change, which often results in errors in traditional systems. Additionally, when a user speaks, some of the lengths between feature points increase or decrease. When a user speaks, they may appear to move closer or farther away from the camera, when in reality they are not moving relative to the camera. Traditional head tracking systems do not take these considerations into account.

[0007] As shown previously, there is a need in the art for improved camera-based audio system tracking, etc. SUMMARY OF THE INVENTION

[0008] One embodiment of the present disclosure describes a method that includes: obtaining one or more images of a user; processing the one or more images to generate a mood-specific 3D ear position of the user based on the user's three-dimensional (3D) head geometry and mood, where the mood is identified based on the one or more images of the user; and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional position of the ear. Further embodiments include a system and a non-transitory computer-readable medium that perform the steps of the method.

[0009] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the accuracy of camera-based head and / or ear tracking is improved. The improved camera-based tracking provides an improved noise cancellation, improved crosstalk cancellation, and an additionally improved three-dimensional audio listening experience for users of personal and / or near-field audio systems (such as headrest audio systems, seat / chair audio systems, soundbars, vehicle audio systems, etc.). The described technology also enables the tracking of mood-specific and user-specific scaled ear positions in three dimensions using a single standard imaging camera. These technical advantages represent one or more technical improvements over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to understand in detail the manner in which the above - described features of the various embodiments are implemented, the inventive concept briefly outlined above can be described more specifically by referring to the various embodiments, some of which are illustrated in the drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be considered to limit the scope in any way, and there are other equivalent embodiments.

[0011] Figure 1 is a schematic diagram showing a computing system according to various embodiments;

[0012] Figure 2 is a diagram showing the Figure 1 registration module according to various embodiments;

[0013] Figure 3 is a diagram showing the emotion - specific ear - position module according to various embodiments;

[0014] Figure 4 is a flowchart of method steps for generating a registered head geometry according to various embodiments;

[0015] Figure 5 is a flowchart of method steps for generating an ear - tracking sound field using emotion - specific geometry according to various embodiments; and

[0016] Figure 6 is a flowchart of method steps for generating an ear - tracking sound field using an emotion - specific depth - scaling factor according to various embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.

[0018] Figure 1FIG. 0 is a schematic diagram showing a computing system 100 according to various embodiments. As shown, the computing system 100 includes, but is not limited to, a computing device 110, one or more cameras 150, and one or more speakers 160. The one or more computing devices 110 include, but are not limited to, one or more processing units 112 and one or more memories 114. In various embodiments, an interconnect bus (not shown) connects the one or more processing units 112, the one or more memories 114, the speakers 160, the one or more cameras 150, and any other components of the computing device 110. The one or more memories 114 store (but are not limited to) an emotion-aware tracking application 120, an audio application 122, and one or more general head geometries 140 and one or more registered head geometries 142. The emotion-aware tracking application 120 includes or utilizes (but is not limited to) a registration module 124, a face detection model 126, a head geometry determination model 130, and an emotion-aware ear position module 132. Although shown separately from the emotion-aware tracking application 120, the registration module 124, the face detection model 126, and the head geometry determination model 130 may include executable instructions that work in concert with the emotion-aware tracking application 120 as sub-modules and / or separate software modules.

[0019] In operation, the computing system 100 processes two-dimensional image data 152 in real time to perform emotion-specific image scaling. The emotion-aware tracking application 120 tracks emotion-specific ear positions based on one or more general head geometries 140 and / or one or more registered head geometries 142. In some embodiments, the emotion-aware tracking application 120 uses a single head geometry ( For example , a general head geometry 140 or a registered head geometry 142) to identify an emotion-aware depth scaling factor and uses the emotion-aware depth scaling factor to generate an emotion-specific ear position. In other embodiments, the emotion-aware tracking application 120 uses multiple different emotion-specific head geometries ( For example , a general head geometry 140 or a registered head geometry 142) to generate an emotion-specific ear position.

[0020] The one or more processing units 112 can be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), and / or any other type of processing unit or combination of different processing units, such as a CPU configured to operate in conjunction with a GPU and / or a DSP. Generally, the processing unit 112 can be any technically feasible hardware unit capable of processing data and / or executing software applications.

[0021] The memory 114 may include random access memory (RAM) modules, flash memory cells, or any other type of memory cells or a combination thereof. The processing unit 112 is configured to read data from and write data to the memory 114. In various embodiments, the memory 114 includes non-volatile memory, such as an optical drive, a magnetic drive, a flash drive, or other storage devices. In some embodiments, a separate data storage device (such as an external data storage device (a "cloud storage device") included in a network) may supplement the memory 114. One or more memories 114 store (but are not limited to) the emotion-aware tracking application 120, the audio application 122, as well as the general head geometry 140 and the registered head geometry 142. The emotion-aware tracking application 120 within one or more memories 114 may be executed by one or more processing units 112 to implement the overall functions of one or more computing devices 110 and thus coordinate the operation of the computing system 100 as a whole.

[0022] One or more cameras 150 include various types of cameras for capturing two-dimensional images of a user. One or more cameras 150 include cameras for a DMS in a vehicle, a soundbar, a webcam, etc. In some embodiments, one or more cameras 150 include only a single standard two-dimensional imager without stereoscopic or depth capabilities. In some embodiments, the computing system 100 includes other types of sensors in addition to one or more cameras 150 to obtain information about the acoustic environment. Other types of sensors include (but are not limited to) motion sensors (such as an accelerometer or an inertial measurement unit (IMU)( For example , a triaxial accelerometer, a gyroscope sensor, and / or a magnetometer)), a pressure sensor, etc. Additionally, in some embodiments, the sensors may include wireless sensors (including radio frequency (RF) sensors( For example , sonar, and radar)) and / or wireless communication protocols (including Bluetooth, Bluetooth Low Energy (BLE), cellular protocols, and / or near field communication (NFC)).

[0023] The speaker 160 includes various speakers for outputting audio to create a sound field or various audio effects near the user. In some embodiments, the speaker 160 includes two or more speakers located in the headrest of a seat (such as a vehicle seat or a gaming chair), or another user-specific speaker group (such as a personal and / or near-field audio system) that is connected or positioned for use by a single user. In some embodiments, the speaker 160 is associated with a speaker configuration stored in the memory 114. The speaker configuration indicates the position and / or orientation of the speaker 160 in three-dimensional space and / or the position and / or orientation relative to each other and / or relative to the vehicle, vehicle seat, gaming chair, camera 150, etc. The audio application 122 can retrieve or otherwise identify the speaker configuration of the speaker 160.

[0024] The two-dimensional image data 152 includes one or more images of the user captured by the camera 150. The camera 150 continuously captures images of the user over time. Some embodiments of the emotion perception tracking application 120 involve a registration module 124 that performs registration to generate one or more registered head geometries 142. During registration, the registration module 124 selects one or more registration images based on certain criteria. The registration generates one or more registered head geometries 142 that are user-specific for the user. Some embodiments of the registration module 124 generate a single registered head geometry 142 corresponding to a neutral emotion. If a single registered head geometry 142 is used, the emotion perception ear position module 132 identifies an emotion perception depth scaling factor to account for various different emotions, such as speaking, facial expressions, and other facial movements. Other embodiments of the registration module 124 generate a set of emotion-specific registered head geometries 142, and the emotion perception ear position module 132 selects an emotion-specific registered head geometry 142 to generate an emotion-specific ear position. The set of emotion-specific registered head geometries 142 can include a neutral emotion registered head geometry 142 and registered head geometries 142 for other facial emotions. In an example where the registration module 124 generates a set of emotion-specific registered head geometries 142, the selection criteria specify a corresponding set of emotions to be identified in the two-dimensional image data 152. The selection criteria can specify one or more facial orientations captured for each emotion. If a single registered head geometry 142 is to be generated, the selection criteria can specify one facial orientation or multiple facial orientations. To this end, in some examples, the registration module 124 generates an audio ( For example , using the speaker 160) user instruction and / or a visual ( For example , using a display device, not shown) user instruction, instructing the user to make a specific facial expression or emotion, such as speaking, laughing, smiling, etc.

[0025] In some embodiments, based on successful detection of a face using the face detection model 126, an enrollment image is further selected from the two-dimensional image data 152. The face detection model 126 includes a machine learning model, a rule-based model, or another type of model that takes the two-dimensional image data 152 as input and generates two-dimensional feature points and / or a face detection status. Additionally or alternatively, the enrollment image is selected from the two-dimensional image data 152 as an image subset that is associated with a specific orientation direction or a range of orientation directions identified by the head orientation detector of the enrollment module 124. The enrollment module 124 includes a rule-based module or program that detects or identifies a face orientation, such as a face orientation status ( For example , centered and / or facing forward, facing up, facing down, facing left, facing right). Enrollment includes a relatively short time period that does not need to be performed in real time ( For example , within 30 seconds, within 45 seconds, or within 1 minute). Since enrollment is relatively fast, in some embodiments, ear positions are not provided during the relatively short enrollment period. In an example where enrollment is not completed, the emotion-aware tracking application 120 is able to provide emotion-specific ear positions based on one or more general head geometries 140 and the two-dimensional image data 152 to improve the pre-enrollment performance of the system. In some examples, enrollment is not performed and the emotion-specific ear positions are not user-specific. Even in examples where the emotion-specific ear positions are not user-specific positions, the emotion-aware aspect of the system has a significant advantage in terms of accuracy over traditional systems.

[0026] The enrollment module 124 provides one or more enrollment images as input to the head geometry determination model 130. The head geometry determination model 130 generates a scaled enrollment head geometry 142 based on the one or more enrollment images selected from the two-dimensional image data 152. The head geometry determination model 130 may include a machine learning model, a rule-based model, or another type of model that takes the two-dimensional image data 152 as input and generates a three-dimensional enrollment head geometry 142. In some embodiments, the input to the head geometry determination model 130 includes camera data, such as the sensor size of the camera, the focal length of the camera, etc. Once enrollment is performed, the emotion-aware tracking application 120 uses the enrollment head geometry 142 and the 'live' or most recently captured two-dimensional image data 152 to identify user-specific ear positions.

[0027] The emotion-aware tracking application 120 identifies pairs of two-dimensional feature points in the two-dimensional image data 152 ( For example, the number of pixels between (identified using the face detection model 126), and generates a depth estimate using the registered head geometry 142. The emotion-aware tracking application 120 uses this information to generate scaled, user-specific, and three-dimensional ear positions. In some embodiments, the emotion-aware tracking application 120 also provides head orientation data to the audio application 122.

[0028] The emotion-aware tracking application 120 provides the ear positions (and in some embodiments, the head orientation data) to the audio application 122 that modifies the audio signal. Thus, the emotion-aware tracking application 120 performs camera-based head and / or ear position tracking based on the two-dimensional image data 152 captured using the camera 150. The audio application 122 uses the emotion-specific ear positions and in some examples uses the head orientation data, speaker configuration data, and / or the input audio signal to generate a set of modified and / or processed audio signals. The modified and / or processed audio signals affect the sound field and / or provide various adaptive audio effects, such as noise cancellation, crosstalk cancellation, spatial / positional audio effects, etc., where the adaptive audio effects adapt to the user's ear position. For example, the audio application 122 can identify one or more head-related transfer functions (HRTFs) based on the ear position, head orientation, and speaker configuration. In some examples, the HRTF is ear-specific for each ear. The audio application 122 modifies one or more speaker-specific audio signals based on the HRTF to maintain the desired audio effects that dynamically adapt to the emotion-specific ear positions. In some embodiments, the audio application 122 generates a set of modified or processed audio signals corresponding to a set of speakers 160. In the examples where registration is performed, the emotion-aware tracking application 120 provides the scaled user-specific and emotion-specific ear positions in real time ( For example , within 10 ms) or near real time ( For example, within 100 ms).

[0029] In various embodiments, one or more computing devices 110 are included in a vehicle system, a home theater system, a soundbar, etc. In some embodiments, one or more computing devices 110 are included in one or more devices such as consumer products ( For example , portable speakers, gaming products, etc.), vehicles ( For example , head units of cars, trucks, vans, etc.), smart home devices ( For example , smart lighting systems, security systems, digital assistants, etc.), communication systems ( For example , conference call systems, video conferencing systems, speaker amplification systems, etc.), etc. In various embodiments, one or more computing devices 110 are located in various environments, including but not limited to indoor environments ( For example, living rooms, meeting rooms, conference halls, home offices, etc.) and / or outdoor environments ( For example , patios, rooftops, gardens, etc.). The computing device 110 is also capable of providing an audio signal ( For example , an audio signal generated using the audio application 122) to the speaker 160 to generate a sound field that provides various audio effects.

[0030] Figure 2 is a diagram showing the Figure 1 registration module 124 according to various embodiments. As shown, the registration module 124 includes, but is not limited to, an image selection module 202 and a head geometry determination model 130. The registration module 124 processes two-dimensional image data 152 as input to generate or output a registered head geometry 142. The image selection module 202 includes and / or utilizes (but is not limited to) a face detection model 126.

[0031] In operation, the registration module 124 processes a specific subset of the two-dimensional image data 152 to perform the registration of the emotion perception tracking application 120. This registration generates one or more user-specific registered head geometries 142. In some embodiments, the registration module 124 is an emotion perception registration module that generates a set of user-specific and emotion-specific registered head geometries 142. In other embodiments, the registration module 124 is emotion-neutral or agnostic and generates a single registered head geometry 142. The registration module 124 also identifies a registration scaling factor that corresponds to the ratio between a first distance between a first pair of feature points and a second distance between a second pair of feature points, where the first pair of feature points and the second pair of feature points share an intermediate point corresponding to a shared feature point. The registration scaling factor acts as a baseline scaling factor. In a further embodiment, the emotion perception tracking application 120 does not include the registration module 124. If no registration is performed, the baseline scaling factor is a default value, such as the minimum value (or other metric) of the first distance plus the second distance calculated over time using multiple images of the two-dimensional image data 152.

[0032] In one example of operation, the registration module 124 receives the two-dimensional image data 152 captured using the camera 150. For example, the emotion perception tracking application 120 can obtain the two-dimensional image data 152 from the camera 150 and provide the two-dimensional image data 152 to the registration module 124. The registration module 124 receives the two-dimensional image data 152 as one or more two-dimensional images. The image selection module 202 selects one or more of the images as the registration image 204. In some embodiments, a single registration image 204 is selected. However, the image selection module 202 selects any specific number of registration images 204 from the two-dimensional image data 152 according to one or more criteria. The image selection module 202 analyzes the two-dimensional image data 152 ( For example, one or more images) to confirm that each registered image 204 meets one or more criteria. In an example where the registration generates a set of emotion-specific registered head geometries 142, the image selection module 202 determines whether the two-dimensional image data 152 corresponds to one or more emotions specified in one or more criteria. In some examples, the registration module 124 generates an audio ( For example , using the speaker 160) user instruction and / or visual ( For example , using a display device, not shown) user instruction, instructing the user to make a specific facial expression or emotion, such as speaking, laughing, smiling, etc. The user instruction corresponds to one or more criteria and increases the likelihood that the two-dimensional image data 152 includes an acceptable registered image.

[0033] One or more criteria may specify that the image selection module 202 selects a registered image 204 for which the face detection model 126 provides facial emotion, face detection status, and / or other data. In some embodiments, the image selection module 202 uses the detection of facial feature points as data indicating that a face has been detected in the image. The image selection module 202 includes rules for identifying whether two-dimensional facial feature points correspond to emotions ( For example , opening the mouth, closing the mouth, smiling, frowning, crying, or any combination thereof). In some embodiments, the criteria also specify that the image selection module 202 selects a registered image 204 for which a specific facial orientation state is identified. The image selection module 202 includes a rule-based module or program that identifies the facial orientation and generates a facial orientation state for a specific image ( For example , centered and / or facing forward, facing upward, facing downward, facing left, facing right). Alternatively, the image selection module 202 refers to the head orientation identified by comparing or analyzing two-dimensional feature point coordinates ( For example , from the face detection model 126) based on the position of the corresponding feature points in the generic head geometry 140. The criteria specify a specific head orientation or range of head orientations. In some embodiments, the criteria may specify that the image selection module 202 selects multiple registered images 204 corresponding to a set of different facial and / or head orientations.

[0034] The image selection module 202 provides the registered image 204 to the head geometry determination model 130. The head geometry determination model 130 generates one or more registered head geometries 142, which include a three-dimensional representation of the user's head. In some embodiments, one or more registered head geometries 142 include a set of user-specific and emotion-specific registered head geometries 142. In other embodiments, a single registered head geometry 142. The head geometry determination model 130 is trained and / or configured to generate one or more registered head geometries 142 that include the three-dimensional positions of one or more feature points (For example in a three - dimensional position (position / location)). Feature points include but are not limited to each eye ( For example , center, outer points, inner points etc. ), each eyebrow ( For example , outer points, inner points, mid - points etc. ), the nose ( For example , bridge of the nose, tip of the nose, base of the nose, root / radix of the nose, between the eyebrows, etc.), the mouth ( For example , left point, right point, upper mid - point, lower mid - point etc. ), the jawline ( For example , left point, right point, upper mid - point, lower mid - point etc. ), each ear ( For example , ear canal etc. ), etc. The head geometry determination model 130 may take a longer time period ( For example , relative to the face detection model 126) to process the two - dimensional image data 152 and generate the three - dimensional registered head geometry 142. As a result, the registration process is generally not a real - time process. However, the registered head geometry 142 generated using the head geometry determination model 130 provides improved ear position accuracy. Especially for personal and / or near - field audio systems (such as headrest audio systems), the improved ear positioning accuracy provides a significant improvement in the user experience.

[0035] To account for the lack of ear feature points in the face detection model 126, the head geometry determination model 130 generates and / or includes ear relationships 206. The ear relationships 206 include a set of three - dimensional relationship vectors. The ear relationships 206 associate one or more feature points generated by the face detection model 126 (and the head geometry determination model 130) with the ear positions. In one example, the feature - point ear relationships 206 associate two ears with a single feature point (such as a chin feature point or a nose feature point). Alternatively, each ear is associated with a different feature point. In some embodiments, the head geometry determination model 130 does not generate ear feature points, and the ear relationships 206 are static relationships, such as predefined or pre - configured relationships. The ear relationships 206 may include the relationships between the ears and the feature points, which are incorporated into the registered head geometry 142 (or the general head geometry 140).

[0036] The runtime of the face detection model 126 is shorter than the runtime of the head geometry determination model 130, enabling the emotion-aware tracking application 120 to use the face detection model 126 in combination with the registered head geometry 142 and head orientation data to generate three-dimensional ear positions in real-time or near real-time. The face detection model 126 generates two-dimensional feature points based on two-dimensional image data 152. The face detection model 126 is trained and / or configured to generate two-dimensional positions corresponding to one or more feature points ( For example , positions in two dimensions). Feature points include but are not limited to each eye ( For example , center, outer point, inner point etc. ), each eyebrow ( For example , outer point, inner point, midpoint, etc.), nose ( For example , bridge of the nose, tip of the nose, base of the nose, root / radix of the nose, glabella etc. ), mouth ( For example , left point, right point, midpoint of the upper lip, midpoint of the lower lip etc. ), jawline ( For example , left point, right point, upper midpoint, lower midpoint, etc.), each ear ( For example , ear canal, etc.). In some embodiments, the face detection model 126 does not provide two-dimensional feature point positions of the ears. In some embodiments, the face detection model 126 and the registered head geometry 142 generate feature point pairs (in two-dimensional space and three-dimensional space respectively), which include but are not limited to the edges of both eyes ( For example , outer edges), nose bridge to chin feature point pair, nose bridge to right jawline feature point pair, left eye inner edge to left jawline feature point pair, and right eyebrow inner edge to left jawline pair.

[0037] Figure 3 are diagrams showing the emotion-aware ear position module 132 according to various embodiments. As shown, the emotion-aware ear position module 132 includes and / or utilizes (but is not limited to) the face detection model 126, the emotion module 300, the emotion-aware depth estimator 302, the feature point conversion module 306, and the feature point to ear transformation module 306. The emotion-aware ear position module 132 receives and / or retrieves and processes two-dimensional image data 152. The ear localization process generates intermediate data, which includes but is not limited to two-dimensional feature point coordinates 308, head orientation 310, emotion 312, emotion-scaled feature point depth estimates 314, and emotion-scaled three-dimensional feature point coordinates 316. The emotion-aware ear position module 132 generates data including but not limited to user-specific ear positions 318. The emotion-aware ear position module 132 also uses one or more general head geometries 140, one or more registered head geometries 142, camera data 336, and ear relationships 206.

[0038] The emotion-aware ear position module 132 uses the face detection model 126 to generate two-dimensional feature point coordinates 308. The face detection model 126 generates the two-dimensional feature point coordinates 308 based on the two-dimensional image data 152. The two-dimensional feature point coordinates 308 are the two-dimensional positions of one or more feature points, and the one or more feature points include but are not limited to one or more eye feature points ( For example such as , center, outer point, inner point etc. ), one or more eyebrow feature points ( For example , outer point, inner point, midpoint etc. ), one or more nose feature points ( For example , bridge of the nose, tip of the nose, bottom of the nose, root of the nose, glabella etc. ), one or more mouth feature points ( For example , left point, right point, midpoint of the upper lip, midpoint of the lower lip etc. ), one or more jawline feature points ( For example , left point, right point, upper midpoint, lower midpoint etc. ) and two or more ear feature points ( For example , left ear canal, right ear canal, etc.). However, in some embodiments, the face detection model 126 does not provide ear feature points. The face detection model 126 stores the two-dimensional feature point coordinates 308 or otherwise provides them to the emotion-aware depth estimator 302.

[0039] The emotion module 300 analyzes the two-dimensional feature point coordinates 308 to detect the emotion 312 and the head orientation 310. The emotion module 300 includes rules for determining the emotion 312 based on the distances between the feature points indicated by the two-dimensional feature point coordinates 308. In some examples, the emotion 312 corresponds to an identifier and / or label of the relationship between the feature point pairs. The emotion module 300 generates the head orientation 310 based on the two-dimensional feature point coordinates 308 and the registered head geometry 142 or the general head geometry 140. The registered head geometry 142 (or the general head geometry 140) includes three-dimensional positions corresponding to the two-dimensional positions of the two-dimensional feature point coordinates 308. In some embodiments, the head orientation 310 includes a three-dimensional orientation vector.

[0040] The emotion-aware depth estimator 302 identifies the emotion-scaled feature point depth estimates 314, which can be expressed as To identify the emotion-scaled feature point depth estimates 314, the emotion-aware depth estimator 302 generates an initial depth estimate based on the head orientation 310, which is based on the registered head geometry 142 or the general head geometry 140. The emotion-aware depth estimator 302 can identify the initial depth estimate based on Equation (1).

[0041]

[0042] In Equation (1), the focal length of camera 150 is denoted as f. The distance between a pair of two-dimensional feature point coordinates 308 is denoted as the capital letter W. The distance between a pair of two-dimensional feature point coordinates 308 in the image ( For example , the two-dimensional image data 152) is denoted as the lowercase letter w. The emotion-scaled feature point depth estimate value 314 of a pair of two-dimensional feature point coordinates 308 or the feature point pair is denoted as d est . The focal length of camera 150 is included in the camera data 336. The distance w in the image can be indicated by the number of pixels, and / or can be generated by multiplying the number of pixels by the physical width of each pixel. The physical width of each pixel can be included in the camera data 336. Equation (1) considers an example in which the line connecting the pair of two-dimensional feature point coordinates 308 is orthogonal to the direction that camera 150 faces. However, the emotion-aware depth estimator 302 can use the head orientation 310 to improve the accuracy of each emotion-scaled feature point depth estimate value 314, where the two-dimensional feature point coordinates 308 are at any angle relative to the direction that camera 150 faces ( For example , based on the head orientation 310).

[0043] In an embodiment where the emotion-aware depth estimator 302 uses a set of emotion-specific registered head geometries 142 and / or emotion-specific generic head geometries 140, the emotion 312 is mapped to a specific emotion-specific head geometry. The emotion-aware depth estimator 302 uses the head orientation 310 and the emotion-specific head geometry to identify the emotion-scaled feature point depth estimate value using the value of d calculated using the emotion-specific head geometry and Equation (1) est . The result is emotion-scaled based on the emotion-specific registered head geometry 142 and / or the emotion-specific generic head geometry 140.

[0044] However, in an embodiment where the emotion-aware depth estimator 302 uses a single registered head geometry 142 or a single generic head geometry 140, the emotion-scaled feature point depth estimate value 314 is subjected to emotion scaling. The emotion scaling is performed based on Equations (2) to (6). The emotion 312 is related to one or more distances between multiple pairs of two-dimensional feature point coordinates 308. For example, the main feature point distance LM dist can be given by Equation (2). Equation (2) decomposes the total distance LM of the main feature point pair dist into two sub-distances by adding an intermediate point between the two feature points.

[0045] LM dist = d1 + d2 (2)

[0046] The emotion module 300 decomposes the total distance LM between two feature points into two sub - distances, for example, according to formula (2), by adding an intermediate point between the two feature points. Although formula (2) shows an example where all feature points are on a straight line, the intermediate feature point can also be offset so that the intermediate feature point is not on the line between the two original feature points. dist In the example shown in, LM corresponds to the distance between a pair of main feature points (such as the nose - bridge feature point and the chin feature point). The intermediate point corresponds to the base of the nose. The distance d1 is the distance corresponding to the pair of feature points including the nose - bridge feature point and the base - of - nose feature point. The distance d2 is the distance corresponding to the pair of feature points including the base - of - nose feature point and the chin feature point. The scaling factor SF is given by formula (3):

[0047] In Figure 3 the example shown, LM dist corresponds to the distance between a pair of main feature points (such as the nose - bridge feature point and the chin feature point). The intermediate point corresponds to the base of the nose. The distance d1 is the distance corresponding to the pair of feature points including the nose - bridge feature point and the base - of - nose feature point. The distance d2 is the distance corresponding to the pair of feature points including the base - of - nose feature point and the chin feature point. The scaling factor SF is given by formula (3):

[0048]

[0049] In some embodiments, the registration scaling factor SF is identified during registration. ENR If the registration scaling factor SF is not identified, ENR then another baseline scaling factor corresponding to a metric (such as the minimum of d1 / d2) is used instead of SF. ENR The emotion module 300 also identifies d1 / d2 in real - time based on the two - dimensional image data 152. The real - time value of d1 / d2 is the real - time scaling factor SF. RT In the example of SF, each of d1 and d2 corresponds to the real - time feature - point - pair distance. RT

[0050] The emotion module 300 identifies a threshold distance between a selected pair of secondary two - dimensional feature - point coordinates 308 and compares the threshold distance with the real - time or most recently identified distance between the selected pair of two - dimensional feature - point coordinates 308. The secondary feature - point pair acts as a trigger to determine whether to scale the initial depth - estimate value. As a result, the secondary feature - point pair can be referred to as a trigger pair. The trigger pair typically involves feature points that move during speech and / or facial expressions. For example, the emotion module 300 can use formula (4) to identify the mouth - opening value MO, which is the distance corresponding to the pair of feature points including the mid - point of the upper lip and the mid - point of the lower lip.

[0051] MO = mid - point of the upper lip−mid - point of the lower lip (4) The emotion module 300 identifies the real - time mouth - opening MO in real - time RT and compares MO RT with the mouth - opening threshold MO. THR MO THR is the threshold distance between the mid - point of the upper lip and the mid - point of the lower lip ( For example, the magnitude of the distance). Although Equation (4) involves the distance between the midpoint of the upper lip and the midpoint of the lower lip, the keypoint pair distance between any real-time keypoint pairs can be compared with a threshold to determine whether to further scale the initial depth estimate. In some examples, the keypoints of distance d2 are used as the trigger keypoint pair.

[0052] Although Equation (4) is indicated as a linear or one-dimensional subtraction, two-dimensional keypoint coordinates 308 can be used to calculate the mouth-opening value MO. If the real-time keypoint pair distance MO between the midpoint of the upper lip and the midpoint of the lower lip RT is greater than (or equal to) the threshold distance value MO THR , the emotion module 300 detects the 'open mouth' emotion 312 or another emotion 312. Other distances can also be used to identify the emotion 312, such as the distance between the bottom of the nose and the chin, the distance between the left point of the mouth and the right point of the mouth, etc. The emotion 312 and the head orientation 310 enable the emotion-aware depth estimator 302 to more accurately identify the emotion-scaled keypoint depth estimate 314. The emotion module 300 provides the emotion 312 and the head orientation 310 to the emotion-aware depth estimator 302.

[0053] The emotion-aware depth estimator 302 generates an emotion-scaled keypoint depth estimate 314 for the corresponding two-dimensional keypoint coordinates and / or two-dimensional keypoint coordinate pairs in the two-dimensional keypoint coordinates 308. The emotion-aware depth estimator 302 uses one or more of the head orientation 310, the two-dimensional keypoint coordinates 308, the registered head geometry 142, and the camera data 336 to generate the emotion-scaled keypoint depth estimate 314. The emotion-scaled keypoint depth estimate 314 can be expressed as In an implementation where a single head geometry is used, if the emotion-aware depth estimator 302 receives or identifies an emotion 312 indicating MO RT ≥MO THR , the depth estimator 302 determines according to Equation (5)

[0054]

[0055] As indicated in Equation (5), the emotion-scaled keypoint depth estimate can be calculated as the initial depth estimate d est , scaled by the ratio between the registered scaling factor SF ENR and the real-time scaling factor SF RT . However, if the emotion-aware depth estimator 302 receives or identifies an emotion 312 indicating MO RT <MO THR , the depth estimator 302 determines according to Equation (6)

[0056]

[0057] As used in formulas (5) and (6), d est is the 'initial' or pre-emotion scaled depth estimate because, in these formulas, a single neutral emotion head geometry is used to generate d est .

[0058] The feature point transformation module 304 uses the two-dimensional feature point coordinates 308 and the corresponding emotion-scaled feature point depth estimates 314 ( For example , ) to generate user-specific emotion-scaled three-dimensional feature point coordinates 316. In some embodiments, the emotion-scaled three-dimensional feature point coordinates 316 are generated using formulas (7) to (9).

[0059] X = (X img - P x ) * d est (7)

[0060] Y = (Y img - P y ) * d est (8)

[0061] Z = d est (9)

[0062] In formulas (7) to (9), X, Y, and Z are the emotion-scaled three-dimensional feature point coordinates 316 corresponding to the three-dimensional feature points in three-dimensional space. X img and Y img are the two-dimensional feature point coordinates 308 generated by the face detection model 126.

[0063] The feature point to ear transformation module 306 generates emotion-specific ear positions 318 based on the emotion-scaled three-dimensional feature point coordinates 316. The feature point to ear transformation module 306 transforms one or more emotion-scaled three-dimensional feature point coordinates 316 into emotion-specific ear positions 318 by applying the ear relationship 206 to the emotion-scaled three-dimensional feature point coordinates 316 indicated as the starting point of the ear relationship 206. The ear relationship 206 includes a set of three-dimensional relationship vectors, one three-dimensional relationship vector for each ear. While one example corresponds to the nose bridge to chin pair of two-dimensional feature point coordinates 308, other pairs may include the glabella to chin pair, the glabella to nasal base pair, or another pair that is predominantly vertical ( For example , the difference between the coordinates is greatest in the vertical dimension). A pair of two-dimensional feature point coordinates 308 may include a pair of inter-ocular pairs, a pair of inter-maxillary pairs, or another pair that is predominantly horizontal ( For example, the difference between coordinates is greatest in the horizontal dimension). Any pair of feature points can be used. The greater the distance between the pair of feature points, the higher the accuracy. As a result, in some embodiments, the bridge of the nose to the chin pair or the glabella to the chin pair may provide higher accuracy.

[0064] Each ear relationship 206 includes a magnitude and a three-dimensional orientation. The feature point to ear transformation module 306 calculates the ear position by setting the initial or starting point of the ear relationship 206 at the emotion-scaled three-dimensional feature point coordinates 316 of a specific feature point in the registered head geometry 142. The feature point to ear transformation module 306 identifies the emotion-specific ear position 318 as the three-dimensional coordinates at the end or terminal position of the ear relationship 206. In some embodiments, the feature point to ear transformation module 306 also uses the head orientation 310, for example, by rotating the ear relationship 206 around a predetermined point in three-dimensional space. For example, if the user looks left (or right), the ear position derived from the ear relationship 206 is different from the case where the user looks straight ahead because the feature point starting point and the head orientation are different. In instances where the two-dimensional feature point coordinates 308 and the emotion-scaled three-dimensional feature point coordinates 316 include ear feature points, the feature point to ear transformation is not performed. Instead, the emotion-aware ear position module 132 utilizes the emotion-scaled three-dimensional feature point coordinates 316 of the ear feature points as the emotion-specific ear position 318. In examples where one or more registered head geometries 142 are used, the emotion-specific ear position 318 is also user-specific and scaled according to the user.

[0065] Figure 4 is a flowchart of method steps for generating a registered head geometry according to various embodiments. Although the method steps are shown in sequence, those skilled in the art should understand that some method steps can be performed in a different order, repeated, omitted, and / or performed by Figure 4 components other than those described in. Although the method steps are described with respect to the Figures 1 to 3 system, those skilled in the art should understand that any system configured to perform the method steps in any order falls within the scope of the various embodiments.

[0066] As shown, method 400 begins at step 402, where the registration module 124 receives two-dimensional image data 152 captured using the camera 150. In some embodiments, the emotion-aware tracking application 120 obtains the two-dimensional image data 152 from the camera 150 and provides the two-dimensional image data 152 to the registration module 124. The registration module 124 receives the two-dimensional image data 152 as one or more two-dimensional images.

[0067] At step 404, the registration module 124 selects one or more registration images 204 for the emotion 312. The emotion 312 may include a 'neutral' emotion, speaking, laughing, smiling, etc. The registration module 124 selects one or more registration images 204 corresponding to the emotion from the two-dimensional image data 152. In some embodiments, a single registration image 204 is selected. However, the registration module 124 selects any particular number of registration images 204 from the two-dimensional image data 152. The registration module 124 analyzes the two-dimensional image data 152( For example , one or more images) to confirm that each registration image 204 meets one or more criteria.

[0068] The criteria may specify that the registration module 124 selects a registration image 204 for which the face detection model 126 indicates that a face is detected. In some embodiments, the criteria also specify that the registration module 124 selects a registration image 204 for a particular head orientation or range of head orientations for a particular emotion 312. In some embodiments, the criteria may specify that the registration module 124 selects multiple registration images 204 corresponding to a set of different head orientations for the emotion 312. In some embodiments, the registration module 124 generates audio user instructions and / or visual user instructions instructing the user to make the appropriate emotion 312.

[0069] At step 406, the registration module 124 generates a registered head geometry 142 based on the registration images 204 of the emotion 312. The image selection module 202 provides the emotion-specific registration images 204 to the head geometry determination model 130. The head geometry determination model 130 generates a registered head geometry 142, which includes the three-dimensional positions of one or more emotion-scaled three-dimensional feature point coordinates 316, and a three-dimensional representation of the user's head. In some embodiments, the registered head geometry 142 includes an ear relationship 206. The ear relationship 206 includes a set of three-dimensional relationship vectors associating one or more feature points with ear positions. In some embodiments, the ear position is the emotion-scaled three-dimensional feature point coordinates 316.

[0070] At step 408, the registration module 124 provides the registered head geometry 142 to the emotion-aware tracking application 120. For example, the registration module 124 stores the registered head geometry 142 in a memory accessible to the emotion-aware tracking application 120. In some embodiments, the image selection module 202 transmits a message indicating that the registered head geometry 142 is available to the emotion-aware tracking application 120.

[0071] At step 410, the registration module 124 determines whether to capture an additional emotion 312 for an additional emotion-specific registration head geometry 142. Some implementations of the registration module 124 generate a single registration head geometry 142, e.g., corresponding to a neutral emotion 312 identified in one or more images of the two-dimensional image data 152. Other implementations of the registration module 124 generate a set of emotion-specific registration head geometries 142. In an example where the registration module 124 generates a set of emotion-specific registration head geometries 142, a selection criterion specifies a corresponding set of emotions 312 to be identified in the two-dimensional image data 152. The registration module 124 determines whether to generate and store a registration head geometry 142 for a respective emotion among the emotions 312 specified in the selection criterion. If a registration head geometry 142 corresponding to an emotion 312 has been generated and stored, the process ends. However, if a registration head geometry 142 for one or more emotions 312 has not been generated and stored, the process moves to step 402.

[0072] Figure 5 is a flowchart of method steps for generating user-specific ear positions according to various embodiments. Although the method steps are shown in sequence, those skilled in the art should understand that some method steps may be performed in a different order, repeated, omitted, and / or performed by components other than those described in Figure 5 Although the method steps are described with respect to the Figures 1 to 3 system, those skilled in the art should understand that any system configured to perform the method steps in any order falls within the scope of the various embodiments. Figure 5 Examples are provided where the emotion-aware tracking application 120 uses a set of emotion-specific head geometries, such as emotion-specific generic head geometries 140 or emotion-specific (and user-specific) registration head geometries 142.

[0073] As shown, method 500 begins at step 502, where the emotion-aware tracking application 120 receives or obtains two-dimensional image data 152. The emotion-aware tracking application 120 retrieves the most recently captured two-dimensional image data 152 from the memory 114 and / or receives two-dimensional image data 152 from the camera 150. The two-dimensional image data 152 may include one or more images captured using the camera 150. The emotion-aware tracking application 120 receives and / or retrieves two-dimensional image data 152 that is updated over time to provide real-time dynamic updates of emotion-specific ear positions 318.

[0074] At step 504, the emotion perception tracking application 120 generates two-dimensional feature point coordinates 308. The emotion perception tracking application 120 uses the face detection model 126 to generate the two-dimensional feature point coordinates 308 based on the two-dimensional image data 152. The two-dimensional feature point coordinates 308 are the two-dimensional positions of one or more feature points, and the one or more feature points include but are not limited to one or more eye feature points ( For example , center, outer point, inner point etc. ), one or more eyebrow feature points ( For example , outer point, inner point, midpoint etc. ), one or more nose feature points ( For example , bridge of the nose, tip of the nose, bottom of the nose, root of the nose, between the eyebrows etc. ), one or more mouth feature points ( For example such as , left point, right point, upper midpoint, lower midpoint etc. ), one or more jawline feature points ( For example , left point, right point, upper midpoint, lower midpoint etc. ), and two or more ear feature points ( For example , left ear canal, right ear canal, etc.). However, in some embodiments, the face detection model 126 does not provide ear feature points.

[0075] At step 506, the emotion perception tracking application 120 determines the emotion 312. The emotion module 300 includes rules for identifying the emotion 312 based on the distances between feature point pairs calculated using the two-dimensional feature point coordinates 308. In one example, the emotion perception tracking application 120 determines that a first distance between the upper lip midpoint and the lower lip midpoint is greater than a first threshold distance, and a second distance between the left mouth point and the right mouth point is greater than a second threshold distance. The emotion perception tracking application 120 determines that the user is laughing based on the first distance being greater than the first threshold distance and the second distance being greater than the second threshold distance.

[0076] At step 508, the emotion perception tracking application 120 identifies an emotion-specific head geometry. The emotion-specific head geometry can be the emotion-specific general head geometry 140 or the emotion-specific registered head geometry 142. The emotion perception tracking application 120 maps the emotion 312 to the emotion-specific general head geometry 140 or registered head geometry 142 corresponding to the emotion 312. In some embodiments, the emotion perception tracking application 120 also generates a head orientation 310 based on an analysis between the two-dimensional feature point coordinates 308 and the emotion-specific head geometry.

[0077] At step 510, the emotion perception tracking application 120 determines the feature point depth estimate 314 of the emotion scale. The emotion perception tracking application 120 uses the head orientation 310 to identify the feature point depth estimate 314 of the emotion scale. The emotion perception tracking application 120 identifies the feature point depth estimate 314 of the emotion scale for the corresponding two-dimensional feature point coordinates and / or pairs of two-dimensional feature point coordinates in the two-dimensional feature point coordinates 308. For example, the emotion perception tracking application 120 uses the value of d est to determine as described with respect to Figure 3 . Formula (1) considers an example where the line connecting the pair of two-dimensional feature point coordinates 308 is orthogonal to the direction the camera 150 is facing. However, the emotion perception tracking application 120 can use the head orientation 310 to improve the accuracy of the emotion scale feature point depth estimate 314, where the two-dimensional feature point coordinates 308 are at any angle relative to the direction the camera 150 is facing ( For example , based on the head orientation 310).

[0078] At step 512, the emotion perception tracking application 120 converts the two-dimensional feature point coordinates 308 into three-dimensional feature point coordinates 316 of the emotion scale, which are scaled according to the user and / or user-specific. The emotion perception tracking application 120 uses the feature point conversion module 304 to generate the three-dimensional feature point coordinates 316 of the emotion scale based on the two-dimensional feature point coordinates 308 and the corresponding feature point depth estimate 314 of the emotion scale, for example, using formulas (7) to (9).

[0079] At step 514, the emotion perception tracking application 120 generates a user-specific ear position based on the three-dimensional feature point coordinates 316 of the emotion scale. The emotion perception tracking application 120 uses the feature point to ear transformation module 306 to generate the emotion-specific ear position 318 based on the three-dimensional feature point coordinates 316 of the emotion scale. The feature point to ear transformation module 306 transforms one or more three-dimensional feature point coordinates 316 of the emotion scale into ear positions by applying the ear relationship 206 to the three-dimensional feature point coordinates 316 of the emotion scale indicated in the ear relationship 206. However, in instances where the two-dimensional feature point coordinates 308 and the three-dimensional feature point coordinates 316 of the emotion scale include ear feature points, the emotion perception tracking application 120 identifies the three-dimensional feature point coordinates 316 of the emotion scale corresponding to the ear feature points and uses these three-dimensional feature point coordinates 316 of the emotion scale as the emotion-specific ear position 318.

[0080] At step 516, audio application 122 generates a processed audio signal based on the emotion-specific ear position 318. In some embodiments, audio application 122 further generates a processed audio signal based on the head orientation 310. For example, audio application 122 identifies one or more HRTFs based on the emotion-specific ear position 318, head orientation 310, and speaker configuration. Audio application 122 generates a processed audio signal based on the HRTFs to maintain the desired audio effects that dynamically adapt to the emotion-specific ear position 318. The processed audio signal is generated to create a sound field and / or provide various adaptive audio effects such as noise cancellation, crosstalk cancellation, spatial / positional audio effects, etc., where the adaptive audio effects adapt to the user's emotion-specific ear position 318. In some embodiments, computing system 100 includes one or more microphones, and audio application 122 uses the microphone audio signals and / or other audio signals to generate the processed audio signal. In some embodiments, audio application 122 further generates the processed audio signal based on the speaker configuration of the set of speakers 160.

[0081] At step 518, audio application 122 provides the processed audio signal to speaker 160. Speaker 160 generates a sound field based on the processed audio signal. As a result, the sound field includes one or more audio effects that dynamically adapt to the user in real time based on the emotion-specific ear position 318 and head orientation 310. In embodiments where one or more registered head geometries 142 are used, the emotion-specific ear position 318 is also user-specific. The process returns to step 502 such that the sound field is dynamically adjusted based on the updated emotion-specific ear position 318 identified using the updated two-dimensional image data 152.

[0082] Figure 6 is a flowchart of method steps for generating a user-specific ear position according to various embodiments. Although the method steps are shown in sequence, those skilled in the art should understand that some method steps can be executed in a different order, repeated, omitted, and / or performed by components other than those described in Figure 6 Although the method steps are described with respect to the Figures 1 to 3 system, those skilled in the art should understand that any system configured to execute the method steps in any order falls within the scope of the various embodiments. Figure 6 An example is provided where the emotion-aware tracking application 120 uses a single head geometry such as the generic head geometry 140 or the registered head geometry 142.

[0083] As shown in the figure, method 600 begins at step 602, where emotion-aware tracking application 120 receives or obtains two-dimensional image data 152. Emotion-aware tracking application 120 retrieves the most recently captured two-dimensional image data 152 from memory 114 and / or receives two-dimensional image data 152 from camera 150. The two-dimensional image data 152 may include one or more images captured using camera 150. Emotion-aware tracking application 120 receives and / or retrieves updated two-dimensional image data 152 over time to provide real-time dynamic updates of emotion-specific ear positions 318.

[0084] At step 604, emotion-aware tracking application 120 generates two-dimensional feature point coordinates 308. Emotion-aware tracking application 120 uses face detection model 126 to generate two-dimensional feature point coordinates 308 based on two-dimensional image data 152. The two-dimensional feature point coordinates 308 are the two-dimensional positions of one or more feature points, and the one or more feature points include but are not limited to one or more eye feature points ( For example , center, outer point, inner point etc. ), one or more eyebrow feature points ( For example , outer point, inner point, midpoint etc. ), one or more nose feature points ( For example , bridge of nose, tip of nose, bottom of nose, root of nose, between eyebrows etc. ), one or more mouth feature points ( For example such as , left point, right point, upper midpoint, lower midpoint etc. ), one or more jawline feature points ( For example , left point, right point, upper midpoint, lower midpoint etc. ), and two or more ear feature points ( For example , left ear canal, right ear canal etc. ), etc. However, in some embodiments, face detection model 126 does not provide ear feature points.

[0085] At step 606, emotion-aware tracking application 120 determines emotion 312. Emotion module 300 includes rules for identifying emotion 312 based on the distances between feature point pairs calculated using two-dimensional feature point coordinates 308, as discussed with respect to formulas (2) through (4). In some examples, emotion-aware tracking application 120 also determines head orientation 310. Emotion-aware tracking application 120 detects emotion 312 based on a threshold distance between trigger feature point pairs. In one example, the trigger pair includes the upper lip midpoint and the lower lip midpoint. Emotion-aware tracking application 120 detects emotion 312 based on the distance between the trigger feature point pair being greater than the threshold distance for the emotion. Although a single distance is indicated, emotion-aware tracking application 120 evaluates any relationship with any number of threshold distances ( For example, any number of trigger feature point pairs (greater than or less than any number of threshold distances) to identify the emotion 312.

[0086] The emotion perception tracking application 120 also uses the real-time feature point pair distance to determine the real-time scaling factor SF RT . The real-time scaling factor SF RT is the real-time value of the ratio between the first distance between the first feature point pair and the second distance between the second feature point pair, where the first feature point pair and the second feature point pair share an intermediate point corresponding to the shared feature point, as discussed with respect to formula (3). In some embodiments, the emotion perception tracking application 120 also generates the head orientation 310 based on the analysis between the two-dimensional feature point coordinates 308 and a single head geometry (such as the generic head geometry 140 or the user-specific registered head geometry 142).

[0087] At step 608, the emotion perception tracking application 120 determines the emotion-scaled feature point depth estimate value 314 by performing emotion scaling on the initial feature point depth estimate. The emotion perception tracking application 120 generates the initial depth estimate based on the head orientation 310, as discussed with respect to formula (1). If the emotion perception tracking application 120 receives or identifies the emotion 312, indicating that the distance between the trigger feature point pairs is greater than (or alternatively less than) the threshold distance of the trigger feature point pairs ( For example , opening the mouth, speaking, laughing, etc.), then the emotion perception tracking application 120 determines the emotion-scaled feature point depth estimate value 314 by scaling the initial depth estimate by the ratio between the baseline or registered scaling factor and the real-time scaling factor SF RT , as discussed with respect to formula (5). However, if the emotion perception tracking application 120 receives or identifies the emotion 312, indicating that the distance between the trigger feature point pairs is less than (or alternatively greater than) the threshold distance of the trigger feature point pairs, then the emotion perception tracking application 120 sets the emotion-scaled feature point depth estimate value 314 to the initial depth estimate, as discussed with respect to formula (6).

[0088] At step 610, the emotion perception tracking application 120 converts the two-dimensional feature point coordinates 308 into emotion-scaled three-dimensional feature point coordinates 316, which are scaled according to the user and / or user-specific for emotion. The emotion perception tracking application 120 generates the emotion-scaled three-dimensional feature point coordinates 316 based on the two-dimensional feature point coordinates 308 and the corresponding emotion-scaled feature point depth estimate value 314, for example, using formulas (7) to (9).

[0089] At step 612, the emotion-aware tracking application 120 generates a user-specific ear position based on the emotion-scaled three-dimensional feature point coordinates 316. The emotion-aware tracking application 120 uses the feature point to ear transformation module 306 to generate an emotion-specific ear position 318 based on the emotion-scaled three-dimensional feature point coordinates 316. The feature point to ear transformation module 306 transforms one or more emotion-scaled three-dimensional feature point coordinates 316 into an ear position by applying the ear relationship 206 to the emotion-scaled three-dimensional feature point coordinates 316 indicated in the ear relationship 206. However, in instances where the two-dimensional feature point coordinates 308 and the emotion-scaled three-dimensional feature point coordinates 316 include ear feature points, the emotion-aware tracking application 120 identifies the emotion-scaled three-dimensional feature point coordinates 316 corresponding to the ear feature points and uses these emotion-scaled three-dimensional feature point coordinates 316 as the emotion-specific ear position 318.

[0090] At step 614, the audio application 122 generates a processed audio signal based on the emotion-specific ear position 318. In some embodiments, the audio application 122 further generates a processed audio signal based on the head orientation 310. For example, the audio application 122 identifies one or more HRTFs based on the emotion-specific ear position 318, the head orientation 310, and the speaker configuration. The audio application 122 generates a processed audio signal based on the HRTFs to maintain a desired audio effect that dynamically adapts to the emotion-specific ear position 318. The processed audio signal is generated to create a sound field and / or provide various adaptive audio effects, such as noise cancellation, crosstalk cancellation, spatial / positional audio effects, etc., where the adaptive audio effects adapt to the user's emotion-specific ear position 318. In some embodiments, the computing system 100 includes one or more microphones, and the audio application 122 uses the microphone audio signals and / or other audio signals to generate the processed audio signal. In some embodiments, the audio application 122 further generates a processed audio signal based on the speaker configuration of the set of speakers 160.

[0091] At step 616, the audio application 122 provides the processed audio signal to the speaker 160. The speaker 160 generates a sound field based on the processed audio signal. As a result, the sound field includes one or more audio effects that dynamically adapt to the user in real time based on the emotion-specific ear position 318 and the head orientation 310. In embodiments where one or more registered head geometries 142 are used, the emotion-specific ear position 318 is also user-specific. The process returns to step 502 such that the sound field is dynamically adjusted based on the updated emotion-specific ear position 318 identified using the updated two-dimensional image data 152.

[0092] In summary, techniques for head and ear tracking using image scaling with emotion detection are disclosed. Some embodiments relate to a method that includes: obtaining one or more images of a user; determining the emotion of the user based on the one or more images; processing the one or more images to generate emotion-specific 3D ear positions of the user based on the user's three-dimensional (3D) head geometry and emotion; and processing one or more audio signals to generate one or more processed audio signals based on the 3D positions of the ears.

[0093] At least one technical advantage of the disclosed techniques over the prior art is that, using the disclosed techniques, the accuracy of camera-based head and / or ear tracking of personal and / or near-field audio systems (such as headrest audio systems, seat / chair audio systems, soundbars, vehicle audio systems, etc.) is improved. The improved camera-based tracking provides users with improved noise cancellation, improved crosstalk cancellation, and an improved three-dimensional audio listening experience. The described techniques or the disclosed techniques enable the tracking of user-specific scaled ear positions in three dimensions using a single standard two-dimensional imaging camera that does not have stereo or depth capabilities. These technical advantages represent one or more technical improvements over prior art methods.

[0094] Aspects of the present disclosure are also described in accordance with the following clauses.

[0095] 1. In some embodiments, a computer-implemented method includes: obtaining one or more images of a user; determining the emotion of the user based on the one or more images; processing the one or more images to generate emotion-specific 3D ear positions of the user based on the user's three-dimensional (3D) head geometry and emotion; and processing one or more audio signals to generate one or more processed audio signals based on the 3D positions of the ears.

[0096] 2. The computer-implemented method according to clause 1, further comprising performing a registration for generating the 3D head geometry based on the one or more images of the user.

[0097] 3. The computer-implemented method according to clause 1 or 2, further comprising performing a registration for generating a plurality of emotion-specific 3D head geometries based on the one or more images of the user.

[0098] 4. The computer-implemented method according to any one of clauses 1 to 3, wherein the 3D head geometry includes a selected emotion-specific 3D head geometry from a general 3D head geometry or a plurality of emotion-specific 3D head geometries.

[0099] 5. A computer-implemented method as described in any one of clauses 1 to 4, wherein processing the one or more images to generate the emotion-specific 3D ear position further includes: determining an initial feature point depth estimate of a plurality of feature points identified in the one or more images based on the 3D head geometry; and performing scaling on the initial feature point depth estimate based on the user's emotion to generate the emotion-specific 3D ear position.

[0100] 6. A computer-implemented method as described in any one of clauses 1 to 5, wherein performing the scaling includes modifying the initial feature point depth estimate based on a baseline scaling factor of a pair of feature points among the plurality of feature points and a real-time scaling factor of the pair of feature points among the plurality of feature points.

[0101] 7. A computer-implemented method as described in any one of clauses 1 to 6, wherein processing the one or more images to determine the emotion-specific 3D ear position includes: selecting an emotion-specific 3D head geometry based on the emotion; and using the emotion-specific 3D head geometry to determine an emotion-scaled feature point depth estimate, wherein the emotion-specific 3D ear position is generated using the emotion-scaled feature point depth estimate.

[0102] 8. A computer-implemented method as described in any one of clauses 1 to 7, wherein determining the 3D position of the user's ears includes: generating two-dimensional (2D) feature point coordinates for a plurality of feature points based on the one or more images using a face detection model; and generating the 3D feature point coordinates based on the emotion-scaled feature point depth estimate of the 2D feature point coordinates using the 3D head geometry, wherein the emotion-specific 3D ear position is based on the 3D feature point coordinates.

[0103] 9. A computer-implemented method as described in any one of clauses 1 to 8, wherein the 3D position of the user's ears is generated based on one or more ear relationships in the 3D head geometry, wherein the one or more ear relationships associate the 3D feature point coordinates with the emotion-specific 3D ear position.

[0104] 10. A computer-implemented method as described in any one of clauses 1 to 9, wherein the plurality of feature points includes one or more of an eye center feature point, an eye outer point feature point, an eye inner point feature point, an eyebrow outer point feature point, an eyebrow inner point, a nose bridge feature point, a nose tip feature point, a nose bottom feature point, a nose root feature point, a glabella feature point, a mouth tip feature point, an upper lip midpoint feature point, a lower lip midpoint feature point, a chin feature point, or a mandibular line feature point.

[0105] 11. A computer-implemented method as described in any one of clauses 1 to 10, wherein processing the one or more audio signals includes: determining one or more head-related transfer functions (HRTFs) based on the 3D position of the ear; and modifying the one or more audio signals based on the HRTFs to generate the one or more processed audio signals.

[0106] 12. A computer-implemented method as described in any one of clauses 1 to 11, further comprising generating, using one or more speakers, an acoustic field including one or more audio effects based on the one or more processed audio signals.

[0107] 13. A computer-implemented method as described in any one of clauses 1 to 12, wherein the audio effect includes one or more of a spatial audio effect, noise cancellation, or crosstalk cancellation.

[0108] 14. In some embodiments, one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: receiving two-dimensional (2D) image data of a user; determining the user's mood based on the 2D image data; processing the 2D image data to determine a mood-specific 3D ear position of the user based on three-dimensional (3D) head geometry and the mood; and generating one or more processed audio signals based on the 3D ear position.

[0109] 15. One or more non-transitory computer-readable media as described in clause 14, wherein the step further includes generating the 3D head geometry based on the 2D image data to register the user.

[0110] 16. One or more non-transitory computer-readable media as described in clause 14 or 15, wherein the step further includes generating a plurality of mood-specific 3D head geometries based on the 2D image data to register the user.

[0111] 17. One or more non-transitory computer-readable media as described in any one of clauses 14 to 16, wherein processing the 2D image data to determine the mood-specific 3D ear position further includes: identifying an initial feature point depth estimate of a plurality of feature points identified in the 2D image data based on the 3D head geometry; and scaling the initial feature point depth estimate based on the user's mood to generate the mood-specific 3D ear position.

[0112] 18. One or more non-transitory computer-readable media as described in any one of clauses 14 to 17, wherein processing the 2D image data to determine the emotion-specific 3D ear position further comprises: selecting an emotion-specific 3D head geometry based on the emotion; and using the emotion-specific 3D head geometry to determine an emotion-scaled feature point depth estimate, wherein the emotion-specific 3D ear position is generated using the emotion-scaled feature point depth estimate.

[0113] 19. One or more non-transitory computer-readable media as described in any one of clauses 14 to 18, wherein processing the 2D image data to determine the emotion-specific 3D ear position further comprises: using a face detection model to generate 2D feature point coordinates for a plurality of feature points based on the 2D image data; and using the 3D head geometry to generate the 3D feature point coordinates based on the emotion-scaled feature point depth estimate of the 2D feature point coordinates, wherein the emotion-specific 3D ear position is based on the 3D feature point coordinates.

[0114] 20. In some embodiments, a system includes: one or more speakers; a camera that captures two-dimensional (2D) image data of a user; a memory that stores instructions; and one or more processors that, when executing the instructions, are configured to perform the following steps: identify the emotion of the user based on the 2D image data; determine the emotion-specific 3D ear position of the user based on a three-dimensional (3D) head geometry and the emotion; and generate a sound field using the one or more speakers, wherein the sound field includes one or more audio effects based on the 3D ear position.

[0115] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0116] Aspects of the present embodiments may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software aspects with hardware aspects, all of which may generally be referred to herein as a "module" or "system". Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0117] Any combination of one or more computer-readable media can be utilized. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, by way of example and not limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following media: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing media. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0118] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array or the like.

[0119] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0120] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and the scope is determined by the claims that follow.

Claims

1. A computer-implemented method comprising: Obtain one or more images of the user; determining an emotion of the user based on the one or more images; processing the one or more images to generate an emotion-specific 3D ear position of the user based on a three-dimensional (3D) head geometry and an emotion of the user; as well as One or more audio signals are processed to generate one or more processed audio signals based on the 3D position of the ear.

2. The computer-implemented method of claim 1 , further comprising: Registration of generating the 3D head geometry based on the one or more images of the user is performed.

3. The computer-implemented method of claim 1 , further comprising: A registration is performed to generate a plurality of emotion specific 3D head geometries based on the one or more images of the user. 4 . The computer-implemented method of claim 1 , wherein the 3D head geometry comprises a generic 3D head geometry or a selected emotion-specific 3D head geometry among a plurality of emotion-specific 3D head geometries.

5. The computer-implemented method of claim 1 , wherein processing the one or more images to generate the emotion-specific 3D ear positions further comprises: determining initial feature point depth estimates for a plurality of feature points identified in the one or more images based on the 3D head geometry; as well as The initial feature point depth estimates are scaled based on the emotion of the user to generate the emotion-specific 3D ear positions.

6. The computer-implemented method of claim 5, wherein performing the scaling comprises modifying the initial feature point depth estimate based on a baseline scaling factor for a pair of feature points in the plurality of feature points and a real-time scaling factor for the pair of feature points in the plurality of feature points.

7. The computer-implemented method of claim 1 , wherein processing the one or more images to determine the emotion-specific 3D ear positions comprises: selecting an emotion-specific 3D head geometry based on the emotion; as well as Emotion-scaled feature point depth estimates are determined using the emotion-specific 3D head geometry, wherein the emotion-specific 3D ear positions are generated using the emotion-scaled feature point depth estimates.

8. The computer-implemented method of claim 1 , wherein determining the 3D position of the ear of the user comprises: generating two-dimensional (2D) feature point coordinates for a plurality of feature points based on the one or more images using a facial detection model; as well as The 3D feature point coordinates are generated using the 3D head geometry based on emotion-scaled feature point depth estimates of the 2D feature point coordinates, wherein the emotion-specific 3D ear positions are based on the 3D feature point coordinates.

9. A computer-implemented method as described in claim 8, wherein the 3D position of the ear of the user is generated based on one or more ear relationships in the 3D head geometry, wherein the one or more ear relationships associate the 3D feature point coordinates with the emotion-specific 3D ear position.

10. The computer-implemented method of claim 8, wherein the plurality of feature points comprises one or more of an eye center feature point, an outer eye point feature point, an inner eye point feature point, an outer eyebrow feature point and an inner eyebrow feature point, a nose bridge feature point, a nose tip feature point, a nose base feature point, a nasal root feature point, a glabella feature point, a mouth tip feature point, an upper lip midpoint feature point, a lower lip midpoint feature point, a chin feature point, or a jawline feature point.

11. The computer-implemented method of claim 1 , wherein processing the one or more audio signals comprises: determining one or more head-related transfer functions (HRTFs) based on the 3D position of the ear; as well as The one or more audio signals are modified based on the HRTF to generate the one or more processed audio signals.

12. The computer-implemented method of claim 1, further comprising: A sound field including one or more audio effects based on the one or more processed audio signals is generated using one or more speakers.

13. The computer-implemented method of claim 12, wherein the audio effect comprises one or more of a spatial audio effect, noise cancellation, or crosstalk cancellation.

14. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: Receiving two-dimensional (2D) image data of a user; determining an emotion of the user based on the 2D image data; processing the 2D image data to determine an emotion-specific 3D ear position of the user based on a three-dimensional (3D) head geometry and the emotion; as well as One or more processed audio signals are generated based on the 3D ear positions.

15. The one or more non-transitory computer readable media of claim 14, wherein the steps further comprise: The 3D head geometry is generated based on the 2D image data to register the user.

16. The one or more non-transitory computer readable media of claim 14, wherein the steps further comprise: A plurality of emotion-specific 3D head geometries are generated based on the 2D image data to enroll the user.

17. The one or more non-transitory computer-readable media of claim 14, wherein processing the 2D image data to determine the emotion-specific 3D ear position further comprises: identifying initial feature point depth estimates for a plurality of feature points identified in the 2D image data based on the 3D head geometry; as well as The initial feature point depth estimates are scaled based on the emotion of the user to generate the emotion-specific 3D ear positions.

18. The one or more non-transitory computer-readable media of claim 14, wherein processing the 2D image data to determine the emotion-specific 3D ear position further comprises: selecting an emotion-specific 3D head geometry based on the emotion; as well as Emotion-scaled feature point depth estimates are determined using the emotion-specific 3D head geometry, wherein the emotion-specific 3D ear positions are generated using the emotion-scaled feature point depth estimates.

19. The one or more non-transitory computer-readable media of claim 14, wherein processing the 2D image data to determine the emotion-specific 3D ear position further comprises: generating 2D feature point coordinates for a plurality of feature points based on the 2D image data using a facial detection model; as well as The 3D feature point coordinates are generated using the 3D head geometry based on emotion-scaled feature point depth estimates of the 2D feature point coordinates, wherein the emotion-specific 3D ear positions are based on the 3D feature point coordinates.

20. A system comprising: one or more speakers; a camera that captures two-dimensional (2D) image data of a user; a memory storing instructions; as well as One or more processors, wherein the one or more processors are configured to perform the following steps when executing the instructions: identifying an emotion of the user based on the 2D image data; determining an emotion-specific 3D ear position of the user based on a three-dimensional (3D) head geometry and the emotion; as well as A sound field is generated using the one or more speakers, wherein the sound field includes one or more audio effects based on the 3D ear positions.