Personalized hrtfs via optical capture
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2019-07-25
- Publication Date
- 2026-08-07
Smart Images

Figure CN116528141B_ABST
Abstract
Description
[0001] Divisional application information
[0002] This application is a divisional application of the invention patent application filed on July 25, 2019, with application number "201980049009.2" and invention title "Personalized HRTFS via Optical Capture".
[0003] Cross-reference of related applications
[0004] This application claims the benefit of U.S. Provisional Application No. 62 / 703,297, filed July 25, 2018, entitled “Method and Apparatus to Personalized HRTF via Optical Capture,” which is incorporated herein by reference. Technical Field
[0005] This invention relates to audio processing, and more particularly, to generating customized audio based on the anthropometric and demographic characteristics of the listener. Background Technology
[0006] Unless otherwise stated herein, the methods described in this section are not prior art to the claims of this application and are not recognized as prior art because they are included in this section.
[0007] By placing sound at various locations and recording it through a simulated head, the perception of sound from corresponding locations relative to the listener can be achieved by playing back such recordings through headphones. This method has the undesirable side effect of causing blurred sound when using separate speakers instead of headphones; therefore, this technique is typically used for selected tracks in multi-track recordings, rather than the entire recording. To improve this technique, the simulated material may incorporate the shape of an ear (auricle) and may be designed to match the acoustic reflectivity / absorption rate of a real head and ear.
[0008] Alternatively, the Head Relative Transfer Function (HRTF) can be applied to the sound source, making the sound appear spatially localized when listening through headphones. Generally, the HRTF corresponds to the acoustic transfer function between a point in three-dimensional (3D) space and the entrance to the ear canal. The HRTF originates from the passive filtering functions of the ear, head, and body, and is ultimately used by the brain to infer the location of the sound. The HRTF consists of amplitude and phase frequency responses that vary depending on the elevation and azimuth angles (rotation around the listener) applied to the audio signal. Besides recording sound at specific locations around the simulated head, sound can be recorded in numerous ways (including conventional methods) and then processed using the HRTF to make it appear at the desired location. Of course, superposition allows for the simultaneous creation of many sounds at various locations to replicate a real-world audio environment or a simple artistic intent. Additionally, sound can be digitally synthesized instead of being recorded.
[0009] The improvement to the HRTF concept involves recording from the ear canals of actual humans. Summary of the Invention
[0010] In refining HRTF through recording from actual human ear canals, it was recognized that significant variability exists between individuals in HRTF, attributed to individual anatomical differences such as shoulder width, head size, ear shape, and other facial features. Additionally, subtle differences exist between the left and right ears of individual listeners. Due to this individual behavior of HRTF, problems persist when using generic HRTFs, such as those designed from simulated heads, individual individuals, or averages of multiple individuals. The use of generic HRTFs often leads to positional accuracy issues, such as difficulty in placing sound in front of the face, reversing positions, transmitting sound a specific distance from the head, and angular accuracy. Furthermore, generic HRTFs have often been found to lack timbre or spectral naturalness and an overall perception of soundstage depth. With this increased understanding, various efforts are underway to obtain HRTFs tailored to specific listeners using various technologies.
[0011] This document describes a system for personalized binaural audio playback to significantly improve the accuracy of perceived sound source location. In addition to user-provided demographic information, the system uses anthropometry-based optical image capture of the user, such as shoulder, head, and ear shape. This data is used to derive a personalized HRTF for the user. This personalized HRTF is then used to process sound sources represented as localized sound objects, where the location can range from precise location to diffuse source (e.g., using...). Atmos TM(Audio objects in the system). In some embodiments, the sound source may be in a multi-channel format, or even a stereo source converted into a location-based sound object. The sound source can be used in video, music, dialogue enhancement, video games, virtual reality (VR), and augmented reality (AR) applications, etc.
[0012] According to an embodiment, a method generates a Head-Related Transfer Function (HRTF). The method includes: generating an HRTF calculation system; and using the HRTF to generate a personalized HRTF for a user. Generating the HRTF calculation system includes: measuring multiple 3D scans of multiple training subjects; generating multiple HRTFs for the multiple training subjects by performing acoustic scattering calculations on the multiple 3D scans; collecting generated data from the multiple training subjects; and performing training on the HRTF calculation system to transform the generated data into the multiple HRTFs. Generating the personalized HRTF includes: collecting generated data from the user; and inputting the user's generated data into the HRTF calculation system to obtain the personalized HRTF.
[0013] Performing the training may include using linear regression in conjunction with lasso regularization.
[0014] The user's generated data includes at least one of anthropometric and demographic data.
[0015] The anthropometry measurements can be obtained by collecting multiple images of the user and using the multiple images to determine the anthropometry measurements. Determining the anthropometry measurements using the multiple images can be performed using a convolutional neural network. The method may further include scaling the user's anthropometry measurements using a reference object in at least one of the user's multiple images.
[0016] The method may further include generating audio output by applying the personalized HRTF to the audio signal.
[0017] The method may further include: storing the personalized HRTF by a server device; and transmitting the personalized HRTF to a user device by the server device, wherein the user device generates audio output by applying the personalized HRTF to an audio signal.
[0018] The method may further include: generating audio output by a user device by applying the personalized HRTF to an audio signal, wherein the user device includes one of headphones, a pair of earbuds, and a pair of wearable devices.
[0019] The audio signal may include multiple audio objects containing location information, and the method may further include generating binaural audio output by applying the personalized HRTF to the multiple audio objects.
[0020] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls a device to perform processing including one or more of the methods discussed above.
[0021] According to another embodiment, a device generates a Head-Related Transfer Function (HRTF). The device includes at least one processor and at least one memory. The at least one processor is configured to control the device to generate an HRTF calculation system and to use the HRTF calculation system to generate a personalized HRTF for a user. Generating the HRTF calculation system includes: measuring multiple 3D scans of multiple training subjects; generating multiple HRTFs for the multiple training subjects by performing acoustic scattering calculations on the multiple 3D scans; collecting generated data from the multiple training subjects; and performing training on the HRTF calculation system to transform the generated data into the multiple HRTFs. Generating the personalized HRTF includes: collecting generated data from the user; and inputting the user's generated data into the HRTF calculation system to obtain the personalized HRTF.
[0022] The generated data of the user may include at least one of anthropometric and demographic data, and the device may further include a user input device configured to collect multiple images of the user and use the multiple images of the user to determine the anthropometric of the user, wherein the anthropometric of the user is scaled using a reference object in at least one of the multiple images of the user.
[0023] The device may further include a user output device configured to generate audio output by applying the personalized HRTF to an audio signal.
[0024] The device may further include a server device configured to generate the HRTF calculation system, generate the personalized HRTF, store the personalized HRTF, and transmit the personalized HRTF to a user device, wherein the user device is configured to generate audio output by applying the personalized HRTF to an audio signal.
[0025] The device may further include a user device configured to generate audio output by applying the personalized HRTF to an audio signal, wherein the user device includes one of a headset, a pair of earbuds, and a pair of wearable devices.
[0026] The audio signal includes multiple audio objects containing location information, wherein at least one processor is configured to control the device to generate binaural audio output by applying the personalized HRTF to the multiple audio objects.
[0027] The device may further include a server device configured to generate the personalized HRTF for the user using the HRTF computing system, wherein the server device performs a photogrammetry component, a context transformation component, a feature point detection component, and an anthropometry component. The photogrammetry component is configured to receive multiple structural images of the user and generate multiple camera transforms and structural image sets using motion-to-structure reconstruction techniques. The context transformation component is configured to receive the multiple camera transforms and the structural image sets and generate transformed multiple camera transforms by translating and rotating the multiple camera transforms using the structural image sets. The feature point detection component is configured to receive the structural image sets and the transformed multiple camera transforms and generate a 3D feature point set corresponding to anthropometry feature points of the user identified using the structural image sets and the transformed multiple camera transforms. The anthropometry component is configured to receive the 3D feature point set and generate anthropometry data from the 3D feature point set, wherein the anthropometry data corresponds to a set of distances and angles measured between individual feature points in the 3D feature point set. The server device is configured to generate the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system.
[0028] The device may further include a server device configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server device performs a scaling component. The scaling component is configured to receive a scaling image containing a scaling reference and generate a homologous metric. The server device is configured to use the homologous metric to scale the user's structural image.
[0029] The device may further include a server device configured to generate the personalized HRTF for the user using the HRTF calculation system, wherein the server device performs a feature point detection component, a 3D projection component, and an angle and distance measurement component. The feature point detection component is configured to receive a cropped image set of anthropometric feature points of the user and generate a set of two-dimensional coordinates of the user's anthropometric feature point set from the cropped image set. The 3D projection component is configured to receive the set of 2D coordinates and multiple camera transformations and use the camera transformations to generate a set of 3D coordinates corresponding to each of the set of 2D components in 3D space. The angle and distance measurement component is configured to receive the set of 3D coordinates and generate anthropometric data from the set of 3D coordinates, wherein the anthropometric data corresponds to the angle and distance of the anthropometric feature points in the set of 3D coordinates. The server device is configured to generate the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system.
[0030] The HRTF calculation system can be configured to train a model corresponding to one of the left ear HRTF and the right ear HRTF, wherein the personalized HRTF is generated by generating one of the left ear personalized HRTF and the right ear personalized HRTF using the model and generating the other of the left ear personalized HRTF and the right ear personalized HRTF using reflections of the model.
[0031] The device may further include a server unit configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server unit performs a data compilation component. The data compilation component is configured to perform a moderate downgrade of the generated data to fill in missing portions of the generated data using estimates determined from known portions of the generated data.
[0032] The device may further include a server unit configured to generate the HRTF computing system, wherein the server unit performs a dimensionality reduction component. The dimensionality reduction component is configured to reduce the computational complexity of training the HRTF computing system by performing principal component analysis on the plurality of HRTFs for the plurality of training subjects.
[0033] The device may further include a server unit configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server unit performs a photography component. The photogrammetry component is configured to receive multiple structural images of the user, perform constrained image feature search on the multiple structural images using a facial landmark detection process, and generate multiple camera transformation and structural image sets using motion-based structure reconstruction techniques and the results of the constrained image feature search.
[0034] The device may further include a server device configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server device performs a context transformation component. The context transformation component is configured to receive a first plurality of camera transformations, a plurality of facial feature points, and a scaling measure; use the plurality of facial feature points to translate and rotate the plurality of camera transformations to generate a second plurality of camera transformations; and use the scaling measure to scale the second plurality of camera transformations.
[0035] The device may further include a server unit configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server unit performs a scaling component. The scaling component is configured to receive range imaging information and use the range imaging information to generate a homologous measure. The server unit is configured to use the homologous measure to scale the user's structural image.
[0036] The device may further include a user input device and a server device. The user input device is associated with a speaker and a microphone. The server device is configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server device performs a scaling component. The scaling component is configured to receive arrival time information from the user input device and use the arrival time information to generate a homologous metric, wherein the arrival time information is related to sound output by the speaker at a first location and received by the microphone at a second location, wherein the first location is associated with the user and the second location is associated with the user input device. The server device is configured to use the homologous metric to scale the user's structural image.
[0037] The device may further include a server unit configured to use the HRTF calculation system to generate the personalized HRTF for the user, wherein the server unit performs a pruning component and a feature point detection component. The pruning component and the feature point detection component are coordinated to implement constrained and recursive feature point search by pruning and detecting multiple sets of different feature points.
[0038] The device may include details similar to those discussed above regarding the method.
[0039] The following detailed description and accompanying figures provide a further understanding of the nature and advantages of the various implementation schemes. Attached Figure Description
[0040] Figure 1 This is a diagram of the audio ecosystem 100.
[0041] Figure 2 This is a flowchart of method 200 for generating the Head-Related Transfer Function (HRTF).
[0042] Figure 3 This is a block diagram of the audio environment 300.
[0043] Figure 4 This is a block diagram of the Anthropometric System 400.
[0044] Figure 5 This is a block diagram of the HRTF Computing System 500. Detailed Implementation Plan
[0045] This document describes techniques for generating head-related transfer functions (HRTFs). In the following description, numerous examples and specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention as defined by the claims may include some or all of these features, alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0046] The following description details various methods, processes, and procedures. Although specific steps may be described in a particular order, this order is primarily for convenience and clarity. A particular step may be performed more than once, may occur before or after other steps (even if the steps are described in a different order), and may occur in parallel with other steps. A second step is required to follow the first step only if the first step must be completed before the second step begins. This will be specifically indicated when it is unclear from the context.
[0047] In this document, the terms “and,” “or,” and “and / or” are used. These terms should be understood to have inclusive meanings. For example, “A and B” can mean at least: “both A and B,” or “at least both A and B.” As another example, “A or B” can mean at least: “at least A,” “at least B,” “both A and B,” or “at least both A and B.” As yet another example, “A and / or B” can mean at least: “A and B,” or “A or B.” When XOR is intended to be used, this will be specifically indicated (e.g., “A or B,” “at most one of A and B”).
[0048] For the purposes of this document, several terms will be defined as follows: Acoustic anatomy will refer to a portion of the human body, including the upper trunk, head, and auricles, which acoustically filters sound and thus contributes to HRTF. Anthropometry will refer to a set of important geometric measurements used to describe the acoustic anatomy of a person. Demographic data will refer to demographic information provided by an individual, which may include their sex, age, ethnicity, height, and weight. Generated data will refer to a set of combined complete or partial anthropometry and demographic data that can be used together to estimate an individual's HRTF. HRTF calculation system will refer to a function or series of functions that takes any generated data as input and returns the estimated personalized HRTF as output.
[0049] As described in more detail herein, the general process for generating a personalized HRTF is as follows. First, an HRTF calculation system is prepared that expresses the relationship between any set of generated data and a unique approximate HRTF. Then, the system uses an input device (e.g., a mobile phone with a camera) and a processing device (e.g., a cloud personalization server) to efficiently derive a set of generated data. The prepared HRTF calculation system is then used with the newly generated data to estimate a personalized HRTF for the user.
[0050] To prepare the HRTF calculation system for use in the system, the following mathematical processes are performed in the training environment. A database of mesh data consisting of high-resolution 3D scans is created for multiple individuals. Demographic data for each individual is also included in the database. A set of corresponding target data consisting of HRTFs is generated from the mesh data. In one embodiment, the HRTF is obtained by numerical simulation of the sound field surrounding the mesh data. For example, this simulation can be performed using the boundary element method or the finite element method. Another known applicable method for obtaining HRTFs without meshes is acoustic measurement. However, acoustic measurement requires the human subject to sit or stand still in an echo-free recording environment for a long period of time, where measurements are prone to error due to human movement and microphone noise. Furthermore, acoustic measurements must be performed individually for each sound source location being measured, making increasing the sampling resolution of the acoustic sphere potentially very expensive. For these reasons, the use of numerical simulation of HRTFs in the training environment can be seen as an improvement over the collection of the HRTF database. In addition, anthropometric data is collected for each individual in the database and combined with demographic data to form a set of generated data. Then, the machine learning process calculates an approximate relationship between the generated data and the target data, which is the model that will be used as part of the HRTF calculation system.
[0051] Once ready in the training environment, the HRTF computing system can be used to generate a personalized HRTF for any user without the need for grid data or acoustic measurements. The system queries the user for demographic data and extracts anthropometric data from structural images using a range of photogrammetric, computer vision, image processing, and neural network techniques. For the purposes of this description, the term structural image refers to multiple images, which may be a series of images or derived from “burst” images or video clips, in which the user’s acoustic anatomy is visible. It may be necessary to scale objects in the structural images to their true physical scale. Scaled images, either independently of or as part of the structural images, can be used for this purpose, as further described herein. In one embodiment, a mobile device can be readily used to capture structural images along with corresponding demographic data and any necessary scaled images. The resulting anthropometric and demographic data are compiled into generated data, which is then used by the ready HRTF computing system to generate a personalized HRTF.
[0052] Figure 5 This is a block diagram of the HRTF computing system 500. The HRTF computing system 500 can be powered by a personalization server 120a (see...). Figure 1 For example, it can be implemented by executing one or more computer programs. The HRTF computing system 500 can be implemented as a sub-component of another component, such as the HRTF generation component 330 (see...). Figure 3 The HRTF computing system 500 includes a training environment 502, a database collection component 510, a numerical simulation component 520, a mesh annotation component 522, a dimensionality reduction component 524, a machine learning component 526, an estimation component 530, a dimensionality reconstruction component 532, and a phase reconstruction component 534.
[0053] The HRTF calculation system 500 can be prepared once in a training environment 502, as discussed below. The training environment 502 may involve computer programs for some components and manual data collection for others. Generally, the training environment 502 determines the relationship between the measured generated data matrix 523 and its corresponding HRTF values for several objects (a few hundred). (This generated data matrix 523 and HRTF may correspond to high-resolution grid data 511, as discussed below.) The system uses front-end generated data approximations to reduce the need for 3D modeling, mathematical simulations, or acoustic measurements. By providing the machine learning component 526 with the “training” HRTF and a relatively small generated data matrix 523, the system generates a model 533 that estimates the values needed to synthesize the entire HRTF set 543. The system can store and distribute the HRTF set 543 in an industry-standard spatial orientation format for acoustic (SOFA).
[0054] In the database collection component 510, demographic data 513 and 3D mesh data 511 are collected from a small number (100s) of training subjects 512. Capturing high-resolution mesh scans can be a time-consuming and resource-intensive task. (This is one reason why personalized HRTFs are not widely used, and also one reason that easier methods for generating personalized HRTFs are being developed, such as the method described in this paper.) For example, a high-resolution scan can be captured using an Artec 3D scanner with a 100,000-triangle mesh. This scan might require 1.5 hours of skilled post-editing work, followed by 24 hours of distributed server time for numerical simulation of the corresponding HRTF. If the HRTF and generated data are obtained directly from other sources for use in the training environment, then the database collection component 510 may be unnecessary, such as the HRTF database from the Center for Image Processing and Integrated Computing (CIPIC) at the UC Davis CIPIC Interface Laboratory.
[0055] At numerical simulation component 520, in one embodiment, these “trained” HRTFs can be computed using the boundary element method and can be expressed as an H matrix 527 and an ITD matrix 525. For any given location of the sound source, the H matrix 527 corresponds to a matrix of amplitude data for all training subjects 512, consisting of the frequency impulse responses of the HRTF. For any given location of the sound source, the ITD matrix 525 corresponds to a matrix of interaural time differences (ITDs, e.g., left ITD and right ITD) for all training subjects 512. The HRTF simulation technique used in one embodiment requires highly complex 3D image capture and very cumbersome mathematical computations. For this reason, training environment 502 is only intended to be prepared once using a limited amount of training data. Numerical simulation component 520 provides the H matrix 527 and ITD matrix 525 to dimension reduction component 524.
[0056] The mesh annotation component 522 outputs anthropometric data 521 corresponding to the anthropometric characteristics identified from the mesh data 511. For example, the mesh annotation component 522 can use manual annotation (e.g., operator-identified anthropometric characteristics). The mesh annotation component 522 can also use angle and distance measurement components (see...). Figure 4 (418) transforms the annotated features into measurements for anthropometric data 521. The union of anthropometric data 521 and demographic data 513 is the generated data matrix 523.
[0057] In one embodiment, dimensionality reduction component 524 may perform principal component analysis on H matrix 527 and ITD matrix 525 to reduce the computational complexity of the machine learning problem. For example, H matrix 527 may have frequency response amplitudes for 240 frequencies; dimensionality reduction component 524 may reduce these to 20 principal components. Similarly, ITD matrix 525 may have values for 2500 source directions; dimensionality reduction component 524 may reduce these to 10 principal components. Dimensionality reduction component 524 provides the set of principal component scores 529 of H matrix 527 and ITD matrix 525 to machine learning component 526. The coefficients 531 required to reconstruct the HRTF from the principal component space are fixed and reserved for use in dimensionality reduction component 532, as discussed later. Depending on the algorithm used in machine learning component 526, other embodiments may omit dimensionality reduction component 524.
[0058] Machine learning component 526 typically prepares model 533 for estimation component 530 for general computation of personalized HRTF. Machine learning component 526 performs training on model 533 to fit the generated data matrix 523 to the set of principal component scores 529 of H matrix 527 and ITD matrix 525. Machine learning component 526 can use approximately 50 predictors from the generated data matrix 523 and can perform known backward, forward, or best subset selection methods to determine the optimal predictor to be used in model 533.
[0059] Once the components of training environment 502 have been executed, a generalizable relationship has been established between the generated data and the HRTF. This relationship includes model 533, and if dimensionality reduction is performed via dimensionality reduction component 524, then the relationship may include coefficients 531. This relationship can be used in the generation steps described below to compute a personalized HRTF corresponding to any new set of generated data. The generation steps described below can be performed by personalization server 120a (see...). Figure 1 For example, this can be implemented by executing one or more computer programs. The generation steps described below can be implemented as a subcomponent of another component, such as HRTF generation component 330 (see below). Figure 3 ).
[0060] The estimation component 530 applies model 533 to a set of generated data 535 to produce the ITD and principal components 537 of the amplitude spectrum of the HRTF. The generated data 535 may correspond to demographic data 311 and anthropometric data 325 (see [link to relevant documentation]). Figure 3 A combination of ) . Generated data 535 can correspond to generated data 427 (see ) Figure 4Then, dimension reconstruction component 532 uses coefficients 531 from component 537 to reverse the dimension reduction process, thereby obtaining an H matrix 541 of the full amplitude spectrum and an ITD matrix 539 describing the ITD of the entire acoustic sphere. Phase reconstruction component 534 then uses the H matrix 541 and the ITD matrix 539 to reconstruct a set of impulse responses describing the HRTF set 543 with phase and amplitude information. In one embodiment, phase reconstruction component 534 may implement a minimum phase reconstruction algorithm. Finally, machine learning system 500 uses the set of impulse responses describing the HRTF set 543 to generate a personalized HRTF, which can be represented in SOFA format.
[0061] Further details of the HRTF Computing System 500 are as follows.
[0062] In one embodiment, the machine learning system 526 may perform linear regression to fit the model 533. For example, a set of linear regression weights may be computed to fit all generated data matrices 523 to each individual directional slice of the magnitude score matrix. As another example, a set of linear regression weights may be computed to fit all generated data matrices 523 to the entire vector of ITD scores. The machine learning system 526 may use z-score normalization to standardize each predictive vector of the generated data matrix 523.
[0063] As a regularization method, regression algorithms can use a minimum absolute shrinkage and selection operator (“lasso”). The lasso process operates to identify and ignore parameters irrelevant to the model at a given location (e.g., by translating the coefficients to zero). For example, the interauricular distance can provide a prediction for the generated data, but has little or no effect on the amplitude of the impulse response between the user’s right ear and a sound source placed directly to the user’s right. Similarly, finer details of the auricle described by the predictions in the generated data may have little effect on the interauricular time difference. By ignoring irrelevant parameters, overfitting can be significantly reduced, and thus the accuracy of the model can be improved. Lasso regression can be compared to ridge regression, where ridge regression scales the weights or proportions of all predictions and does not set any coefficients to zero.
[0064] In other embodiments, the machine learning system 526 may use other machine learning methods to generate the HRTF set. For example, the machine learning system 526 may train a neural network to predict the entire matrix of magnitude scores. As another example, the machine learning system 526 may train a neural network to predict the entire vector of ITD scores. The machine learning system 526 may normalize the HRTF values via z-score normalization before training the neural network.
[0065] In one embodiment, the training environment 502 can be optimized by performing machine learning and / or dimensionality reduction on the transfer function of only one ear. For example, a single HRTF containing the transfer function for the entire sphere around the head can be considered as two sets of left-ear HRTFs, one of which simply reflects across the sagittal plane. In this example, if a numerical simulation component 520 is performed on 100 subjects using both ears as receivers above the entire acoustic sphere, then the right-ear HRTF value for each subject can be converted to a left-ear value, thus resulting in a set of HRTF values containing 200 instances of the left-ear HRTF. The HRTF can be expressed as a source location-dependent impulse response, amplitude spectrum, or interaural delay, and each right-ear position can be directly mapped to a left-ear position by reflecting its coordinates across the sagittal plane. The conversion can be performed by assigning the right-ear HRTF values to matrix indices of the reflection positions.
[0066] Because the predictions of the generated data matrix 523 used to train model 533 are scalar values, these predictions can also be considered independent of the side of the body on which the predictions are measured. Therefore, model 533 can be trained to approximate, for example, the HRTF of the left ear. Creating the user's right ear HRTF is as simple as mapping the spherical coordinates of the HRTF set generated using the right ear data back to its original coordinates. Therefore, even if the generated data and the resulting HRTF may not be symmetric, model 533 and the dimensionality reduction can be said to be symmetric. Overall, this reflection process has the expected result of reducing the complexity of the target data by a factor of two and increasing the sample size of the H matrix 527 and the ITD matrix 525 by a factor of two. A significant additional advantage of using this process is that the reflection-reconstructed HRTF can be more balanced. This is because the reflection process results in symmetric behavior of any noise in the HRTF calculation system 500 caused by overfitting and errors in the dimensionality reduction component 524 and the machine learning component 526.
[0067] Figure 1 This is a block diagram of an audio ecosystem 100. The audio ecosystem 100 includes one or more user devices 110 (two shown: user input device 110a and user output device 110b) and one or more cloud devices 120 (two shown: personalization server 120a and content server 120b). Although multiple devices are shown, this is for ease of description. A single user device can implement the functions of user input device 110a and user output device 110b. Similarly, the functions of personalization server 120a and content server 120b can be implemented by a single server or by multiple computers in a distributed cloud system. The devices of the audio ecosystem 100 can be connected via wireless or wired networks (not shown). The general operation of the audio ecosystem 100 is as follows.
[0068] User input device 110a captures user-generated data 130. User input device 110a may be a mobile phone with a camera. Generated data 130 consists of structural images and / or demographic data, and may also include zoomed images. Further details of the capture process and generated data 130 are described below.
[0069] Personalization server 120a receives generated data 130 from user input device 110a, processes the generated data 130 to generate a personalized HRTF 132 for the user, and stores the personalized HRTF 132. For example, personalization server 120a may implement estimation component 530, dimension reconstruction component 532, and phase reconstruction component 534 (see...). Figure 5 Further details regarding the generation of the personalized HRTF 132 are provided below. The personalization server 120a also provides the personalized HRTF 132 to the user output device 110b.
[0070] Content server 120b provides content 134 to user output device 110b. Generally, content 134 contains audio content. The audio content may contain audio objects, such as those defined by… Atmos TM The system. Audio content may contain multi-channel signals, such as stereo signals converted to locate sound objects. Content 134 may also contain video content. For example, content server 120b may be a multimedia server providing audio and video content, a game server providing game content, etc. Content 134 may be continuously provided from content server 120b, or content server 120b may provide content 132 to user output device 110b for current storage and future output.
[0071] User output device 110b receives a personalized HRTF 132 from personalization server 120a, receives content 134 from content server 120b, and applies the personalized HRTF 132 to content 134 to produce audio output 136. Examples of user output device 110b include mobile phones (and associated earphones), headphones, headsets, earbuds, wearable devices, etc.
[0072] User output device 110b can be the same device as user input device 110a. For example, a mobile phone with a camera can capture and generate data 130 (as user input device 110a), receive personalized HRTF 132 (as user output device 110b), and can be associated with a pair of earbuds that generate audio output 136. User output device 110b can also be a different device from the user input device 110a. For example, a mobile phone with a camera can capture and generate data 130 (as user input device 110a), and a headset can receive personalized HRTF 132 and generate audio output 136 (as user output device 110b). User output device 110b can also be associated with other devices, such as a computer, an audio / video receiver (AVR), a television, etc.
[0073] The audio ecosystem 100 is referred to as an "ecosystem" because it is adaptable to any output device currently being used by the user. For example, a user can be associated with a user identifier, and the user can log in to the audio ecosystem 100. Personalization server 120a can use the user identifier to associate a personalized HRTF 132 with the user. Content server 120b can use the user identifier to manage the user's subscriptions, preferences, etc., for content 134. User output device 110b can use the user identifier to communicate to personalization server 120a that the user output device 110b should receive the user's personalized HRTF 132. For example, when a user purchases new headphones (as user output device 110b), the headphones can use the user identifier to obtain the user's personalized HRTF 132 from personalization server 120a.
[0074] Figure 2 This is a flowchart of method 200 for generating the head-related transfer function (HRTF). Method 200 can be provided by audio ecosystem 100 (see...). Figure 1 One or more devices, for example, by executing one or more computer programs.
[0075] At point 202, the HRTF calculation system is generated. Generally, the HRTF calculation system corresponds to the relationship between anatomical measurements and HRTF. The HRTF calculation system can be generated by a personalized server 120a (see...). Figure 1 For example, by implementing the HRTF computing system 500 (see...) Figure 5 The generation of the HRTF calculation system includes sub-steps 204, 206, 208, and 210.
[0076] At point 204, several 3D scans of several training subjects are measured. Generally, the 3D scans correspond to a database of high-resolution scans of the training subjects, and the measurements correspond to measurements of anatomical features captured in the 3D scans. The 3D scans may correspond to grid data 511 (see...). Figure 5 The Personalized Server 120a can store a database of high-resolution scans.
[0077] At position 206, several HRTFs for the training subject are generated by performing acoustic scattering calculations based on measurements from a 3D scan. The personalized server 120a can be implemented, for example, by implementing a numerical simulation component 520 (see...). Figure 5 To perform acoustic scattering calculations to generate HRTF.
[0078] At point 208, generated data for the training subject is collected. Generally, the generated data corresponds to the anthropometric and demographic data of the training subject, with the anthropometric data determined from 3D scan data. For example, the generated data may correspond to demographic data 513, anthropometric data 521, and generated data matrix 523 (see [link to documentation]). Figure 5 One or more of the following. Anthropometric data 521 may be generated by a mesh annotation component 522 based on mesh data 511 (see...). Figure 5 )produce.
[0079] At 210, training is performed on the HRTF calculation system to transform the generated data into multiple HRTFs. Generally, a machine learning process is performed to generate a model for the HRTF calculation system, using which the generated data (see 208) is used to estimate the values of the generated HRTFs (see 206). Training may include linear regression in conjunction with lasso regularization, as discussed in more detail below. The personalized server 120a may, for example, implement machine learning component 526 (see... Figure 5 ) to perform the training process.
[0080] At point 212, a personalized HRTF for the user is generated using an HRTF calculation system. The personalization server 120a can be implemented, for example, by implementing an HRTF calculation system 500 (see...). Figure 5 To generate a personalized HRTF. Generating a personalized HRTF includes sub-steps 214 and 216.
[0081] At point 214, generated data for the user is collected. Generally, the generated data corresponds to the anthropometrics and demographics of a specific user (where the anthropometrics are determined from 2D image data) in order to generate their personalized HRTF. For example, the generated data might correspond to generated data 535 (see...). Figure 5 The reference object can also be captured in the image along with the user for zooming purposes. User input device 110a (see...) Figure 1This can be used to collect user-generated data. For example, the user input device 110a can be a mobile phone that includes a camera.
[0082] At 216, the user-generated data is input into the HRTF calculation system to obtain a personalized HRTF. The personalization server 120a can obtain a personalized HRTF by inputting the user-generated data (see 214) into the results of training the HRTF calculation system (see 210), for example, by implementing the estimation component 530, the dimension reconstruction component 532, and the phase reconstruction component 534 (see...). Figure 5 ).
[0083] At point 218, once a personalized HRTF has been generated, it can be provided to the user output device and used when generating audio output. For example, user output device 110b (see...) Figure 1 The system can receive a personalized HRTF 132 from a personalization server 120a, and an audio signal from content 134 from a content server 120b. Audio output 136 can be generated by applying the personalized HRTF 132 to the audio signal. The audio signal may contain an audio object with location information, and the audio output may correspond to binaural audio output generated by rendering the audio object using the personalized HRTF. For example, the audio object may contain... Atmos TM Audio object.
[0084] Further details of this process are provided below.
[0085] Figure 3 This is a block diagram of audio environment 300. Audio environment 300 is similar to audio environment 100 (see [link]). Figure 1 ), and provides additional details. Similar to audio environment 100, audio environment 300 can use one or more devices to generate a personalized HRTF, for example, by performing method 200 (see... Figure 2 One or more steps. Audio environment 300 includes input device 302, processing device 304, and output device 306. The details of audio environment 300's functionality are described in contrast to audio environment 100. The functions of the devices in audio environment 300 may be implemented by one or more processors, for example, executing one or more computer programs.
[0086] Input device 302 typically captures user input data. (The input data is processed into user-generated data, such as structural images 313 and / or demographic data 311.) Input device 302 may also capture zoomed-in images 315. Input device 302 may be a mobile phone with a camera. Input device 302 includes a capture component 312 and a feedback and local processing component 314.
[0087] The capture component 312 typically captures demographic data 311 and structural images 313 of the user's acoustic anatomy. The structural images 313 (described further below) are then used to generate a set of anthropometric data 325. For ease of further processing, the structural image 313 capture can be performed against a static background.
[0088] One option for capturing structural images 313 is as follows. The user places the input device 302 on a stable surface directly below eye level and positions it so that its acoustic anatomy is visible in the captured frame. The input device 302 generates a tone or other indicator, and the user slowly rotates 360 degrees. The user can rotate from a standing or sitting position with their arms at their sides.
[0089] Another option for capturing structural images 313 is as follows. The user holds the input device 302 at arm's length, while the user's acoustic anatomy is in the video frame. Starting with the input device 302 facing the user's ear, the user swings their arm forward, causing the video to capture an image from the user's ear to the front of the user's face. The user then repeats the process on the other side of their body.
[0090] Another option for capturing structural images 313 is as follows. As in the previous embodiment, the user holds the input device 302 at arm's length, and the user's acoustic anatomy is in the video frame. However, in this embodiment, the user rotates their head as comfortably as possible to the left and right. This allows the user's head and ears to be captured in the structural image.
[0091] The above options allow a user to capture an image of their own structure without the assistance of another person. However, an additional effective embodiment would be to have a second person walk around a stationary, standing user, with the camera of input device 302 pointed at the user's acoustic anatomy.
[0092] The extent, order, and manner of recording structural images are irrelevant as long as structural images from multiple azimuth or horizontal angles relative to the face are available. In one embodiment, it is recommended to capture structural images at intervals of ten degrees or less and within at least 90 degrees to the left and right of the user's face.
[0093] The capture component 312 can provide guidance to the user during the capture process. For example, the capture component 312 can output a beep or voice command to tilt the input device 302 up or down; to vertically shift the input device 302 to achieve perpendicularity with the user's ear; to increase or decrease the speed of the sweep or rotation process; etc. The capture component 312 provides the structural image to the feedback and local processing component 314.
[0094] The feedback and local processing component 314 typically evaluates the structural image 313 captured by the capture component 312 and performs local processing on the capture. Regarding evaluation, the feedback and local processing component 314 can assess various criteria of the captured image, such as whether the user remains within the frame or does not rotate too quickly; if a criterion indicates failure, the feedback and local processing component 314 can return the operation of the input device 302 to the capture component 312 to perform another capture. Regarding local processing, the feedback and local processing component 314 can subtract the background from each image and perform other image processing functions, such as blur / sharpness evaluation, contrast evaluation, and brightness evaluation, to ensure image quality. The feedback and local processing component 314 can also perform key feature point recognition, such as the center of the face and ear position in the video, to ensure that the final structural image 313 adequately describes the user's acoustic anatomy.
[0095] The captured video then contains structural images of the user's acoustic anatomy from multiple angles. Input device 302 then sends the final structural image 313, along with any demographic data 311, to processing device 304. Input device 302 may also send a zoomed image 315 to processing device 304.
[0096] Processing device 304 typically processes structural image 313 to generate anthropometric data 325 and generates personalized HRTF based on generated data consisting of anthropometric data 325 and / or demographic data 311. Processing device 304 may be hosted on a cloud-based server. Alternatively, input device 302 may implement one or more functions of processing device 304 where local implementation of cloud processing capabilities is required. Processing device 304 includes photogrammetry component 322, context transformation component 324, feature point detection component 326, anthropometric component 328, and HRTF generation component 330.
[0097] Photogrammetry component 322 receives the final version of structural image 313 from feedback and local processing component 314 and performs photogrammetry using techniques such as Structure from Motion (SfM) to generate camera transform 317 and structural image set 319. Generally, structural image set 319 corresponds to frames of structural image 313 successfully located by photogrammetry component 322, and camera transform 317 corresponds to the 3D position and orientation components for each image in structural image set 319. Photogrammetry component 322 provides structural image set 319 to context transformation component 324 and feature point detection component 326. Photogrammetry component 322 also provides camera transform 317 to context transformation component 324.
[0098] The context transformation component 324 uses the structured image set 319 to translate and rotate the camera transformation 317 to generate the camera transformation 321. The context transformation component 324 may also receive a scaled image 315 from the feedback and local processing component 314; when the camera transformation 321 is generated, the context transformation component 324 may use the scaled image 315 to scale the camera transformation 317.
[0099] Feature point detection component 326 receives and processes structural image set 319 and camera transformation 321 to generate 3D feature point set 323. Generally, 3D feature point set 323 corresponds to anthropometric feature points identified by feature point detection component 326 from structural image set 319 and camera transformation 321. For example, these anthropometric feature points may include various feature points on the visible surfaces of each auricle, such as the fossa, outer ear, tragus, and helix. Other anthropometric feature points of the user's acoustic anatomy detected by feature point detection component 326 may include eyebrows, chin, and shoulders; and the head and torso may be measured in appropriate frames. Feature point detection component 326 provides 3D feature point set 323 to anthropometric component 328.
[0100] The human measurement component 328 receives a 3D feature point set 323 and generates human measurement data 325. Generally, the human measurement data 325 corresponds to a set of distances and angles geometrically measured between individual feature points in the 3D feature point set 323. The human measurement component 328 provides the human measurement data 325 to the HRTF generation component 330.
[0101] The HRTF generation component 330 receives anthropometric data 325 and generates a personalized HRTF 327. The HRTF component 330 may also receive demographic data 311 and use it when generating the personalized HRTF 327. The personalized HRTF 327 may be presented in an acoustically oriented spatial orientation format (SOFA) file format. As part of generating the personalized HRTF 327, the HRTF generation component 330 typically uses a previously determined HRTF calculation system, as discussed in more detail herein (e.g., by...). Figure 5 The HRTF calculation system 500 trains a model 533. The HRTF generation component 330 provides a personalized HRTF 327 to the output device 306.
[0102] Output device 306 typically receives personalized HRTF 327 from processing device 304, applies the personalized HRTF 327 to audio data, and generates audio output 329. Output device 306 may be a mobile phone and associated speakers (e.g., headphones, earphones, etc.). Output device 306 may be the same device as input device 302. Where cloud processing capabilities are desired to be implemented locally, output device 306 may implement one or more functions of processing device 304. Output device 306 includes rendering component 340.
[0103] The rendering component 340 receives a personalized HRTF 327 from the HRTF generation component 330, performs binaural rendering on the audio data using the personalized HRTF 327, and produces audio output 329.
[0104] Figure 4 This is a block diagram of an anthropometry system 400. The anthropometry system 400 may consist of an audio ecosystem 100 (e.g., Figure 1 Personalized servers 120a), audio ecosystem 300 (e.g., Figure 3 The human body measurement system 400 can be implemented using components such as the processing device 304. Method 200 can be implemented using the human body measurement system 400 (see...). Figure 2 One or more steps. The anthropometry system 400 may be similar to the processing device 304 (see...). Figure 3 The human body measurement system 400 operates through one or more components, such as a photogrammetry component 322, a context transformation component 324, a feature point detection component 326, and an anthropometry component 328. The anthropometry system 400 includes a data extraction component 402, a photogrammetry component 404, a proportioning component 406, a facial feature point detection component 408, a context transformation component 410, a cropping component 412, a feature point detection component 414, a 3D projection component 416, an angle and distance component 418, and a data compilation component 420. The functions of the components of the anthropometry system 400 may be implemented by one or more processors, for example, processors executing one or more computer programs.
[0105] The data extraction component 402 receives input data 401 and performs data extraction and selection to generate demographic data 403, structural image 405, and scale image 407. This can be obtained from the user input device 110a (see [link to user input device]). Figure 1 ) or from input device 302 (see Figure 3 ) Receive input data 401, for example, data contained by the mobile phone's camera (see... Figure 2 The image data captured by the data extraction component 403 is directly provided to the data compilation component 420, the structural image 405 is provided to the photogrammetry component 404, and the scale image 407 is provided to the scale measurement component 406.
[0106] Photogrammetry component 404 typically performs a photogrammetric process on structural image 405, such as structure-from-motion (SfM), to generate camera transforms 411 and image set 409. The photogrammetric process acquires structural image 405 and generates a set of camera transforms 411 (e.g., camera viewpoint position and viewpoint orientation) corresponding to each frame of image set 409, which may be a subset of structural image 405. Viewpoint orientation is typically expressed in quaternion or rotation matrix format, but for the purposes of this document, the mathematical example will be expressed in rotation matrix format. Image set 409 is passed to facial feature point detection component 408 and cropping component 412. Camera transforms 411 are passed to context transform component 410.
[0107] Optionally, the photogrammetry component 404 can perform image feature detection on the structural image 405 using constrained image feature search before executing the SfM process. Constrained image feature search can improve the results of the SfM process by overcoming user errors during the capture process.
[0108] The scaling measurement component 406 uses the scaling image 407 from the data extraction component 402 to generate information for later use during the scaling camera transformation 411. The scaling information, referred to as the homologity measure 413, is generated as follows: The scaling image contains a visible scaling reference, which the scaling measurement component 406 uses to measure the scaling homologs visible in the same frame of the scaling image and in one or more frames of the structure image. The resulting measure of the scaling homologs, as the homologity measure 413, is passed to the context transformation component 410.
[0109] The facial feature point detection component 408 searches for visible facial feature points in frames of the image set 409 received from the photogrammetry component 404. The detected feature points may include points on the user's nose and the location of the pupils, which can later be used as scale homologs visible in both the image set 409 and the scale image 407. The resulting facial feature points 415 are then passed to the context transformation component 410.
[0110] The context transformation component 410 receives a camera transform 411 from the photogrammetry component 404, a homology measure 413 from the scaling component 406, and a set of facial feature points 415 from the facial feature point detection component 408. The context transformation component 410 effectively transforms the camera transform 411 into a set of camera transforms 417, which are appropriately centered, oriented, and scaled to the context of the acoustic anatomy captured in the structural images of the image set 409. In summary, the context transformation is achieved by: scaling the positional information of the camera transform 411 using the facial feature points 415 and the homology measure 413; rotating the camera transform 411 in 3D space using the facial feature points 415; and translating the positional information of the camera transform 411 using the facial feature points 415 to move the origin of the 3D space to the center of the user's head. The resulting camera transform 417 is then passed to the cropping component 412 and the 3D projection component 416.
[0111] Cropping component 412 typically uses camera transform 417 to select and crop a subset of frames from image set 409. After proper centering, orientation, and scaling, cropping component 412 uses camera transform 417 to estimate which subset of the images from image set 409 contains structural images of specific features of the user's acoustic anatomy. Furthermore, cropping component 412 can use camera transform 417 to estimate which portion of each image contains structural images of specific features. Cropping component 412 can thus be used to crop individual frames from a subset of image set 409 to produce the resulting image data of cropping 419.
[0112] Feature point detection component 414 typically provides predicted locations of specified feature points of the user's acoustic anatomy visible in the 2D image data of cropped 419. Therefore, feature points visible in a given image frame are labeled with corresponding ordered set of 2D point locations. Cropping component 412 and feature point detection component 414 can be coordinated to implement constrained and recursive feature point search by cropping and detecting multiple sets of different feature points visible in different subsets of image set 409. Feature point detection component 414 transmits the obtained 2D coordinates 421 of the anatomical feature points to 3D projection component 416.
[0113] The 3D projection component 416 typically uses camera transformation 417 to convert a series of 2D coordinates 421 of each anatomical feature point into a single position in 3D space. A complete set of 3D feature point positions, as 3D coordinates 423, is then passed to the angle and distance measurement component 418.
[0114] The angle and distance measurement component 418 uses a set of predetermined instructions to measure the angles and distances between various points on the 3D coordinate 423. These measurements can be achieved by applying simple Euclidean geometry. The resulting measurements can be effectively used as anthropometric data 425 and passed to the data compilation component 420.
[0115] Data compilation component 420 typically combines demographic data 403 with anthropometric data 425 to form a complete set of generated data 427. This generated data 427 can then be processed in an HRTF computing system (e.g., as described above). Figure 5 The generated data (535) is used to export a personalized HRTF for the user.
[0116] Further details and examples of the Anthropometry System 400 are as follows.
[0117] While all frames of structural imagery 405 can be used in photogrammetry component 404, the system can achieve better performance by reducing the number of frames for computational efficiency. To select the optimal frames, data extraction component 402 can evaluate frame content and sharpness metrics. An example of frame content selection might be searching for similarity in consecutive images to avoid redundancy. Sharpness metrics can be selected from several options, one example being a 2D spatial frequency power spectrum radially collapsed into a 1D power spectrum. Data extraction component 402 provides the selected set of structural images 405 to photogrammetry component 404. Because photogrammetry component 404 may perform time-consuming processes, data extraction component 402 can pass structural images 405 before collecting scaled images 407 or demographic data 403 from input data 401. This order of operations is an ideally optimized example if the system can process in parallel.
[0118] The SfM process takes a series of frames as input (e.g., structural images 405 that do not require sequential ordering) and outputs the assumed rigid object imaged during the capture process (3D point cloud) along with estimates of the calculated viewpoint position (x, y, z) and rotation indices for each frame input. These viewpoint position and rotation matrices are referred to in this document as camera transformations (e.g., camera transformation 411) because they describe the positioning and orientation of each camera with respect to world space, a general term encompassing the 3D coordinate system of all camera viewpoints and structural images. It should be noted that for this application, further use of the 3D point cloud itself is not required; that is, it is not necessary to generate a 3D mesh object for any part determined by the user's anthropometry. It is not uncommon for the SfM process to fail to derive camera transformations for one or more of the images in a set of optimal structural images 405. Therefore, any failed frames can be omitted from subsequent processing. Because autofocus camera applications are not necessarily optimized for the capture conditions of this system, for the image capture component (see...), Figure 3 Fixing the camera focus during the image capture process can be useful.
[0119] The SfM process assumes that the auricle and head are inherently rigid objects. In determining the shape and location of these objects, the SfM process first implements one or more known image feature detection algorithms, such as SIFT (Shift Invariant Feature Transform), HA (Hessian Affine Feature Detector), or HOG (Histogram of Oriented Gradients). The resulting image features differ from facial and anatomical feature points because they are not individually pre-trained and are not specific in any way to the context of the acoustic anatomy. Because other parts of the user's body may not remain rigid throughout the capture process, it can be useful to program the photogrammetry process to infer geometry using only the image features detected in the head region. This selection of image features can be achieved by executing the facial feature point detection component 408 before executing the photogrammetry component 404. For example, known face detection techniques can be used to estimate feature points defined by the bounding boxes of the head or face in each image. The photogrammetry component 404 can then apply a mask to each image to include only the image features detected by computer vision within the corresponding bounding boxes. In another embodiment, the photogrammetry component 404 may directly use facial feature points instead of or in addition to the detected image features. By limiting the range of image features used during photogrammetry, the system can be made more efficient and optimized. If the facial feature point detection component 408 is performed before the photogrammetry component 404, then the facial feature point detection component 408 may directly receive the structural image 405 from the data extraction component 402 instead of the image set 409, and may pass facial feature points 415 to both the photogrammetry component 404 and the context transformation component 410.
[0120] Photogrammetry component 404 may also include a focal length compensation component. The camera's focal length can exaggerate the depth of the ear shape (e.g., due to barrel distortion caused by a short focal length) or reduce such depth (e.g., due to pincushion distortion caused by a long focal length). In many smartphone cameras, this focal length distortion is typically of the barrel distortion type, and the focal length compensation component can detect the barrel distortion type based on the focal length and the distance to the captured image. Known methods can be used to apply this focal length compensation process to ensure that the image set 409 is not distorted. This compensation can be particularly useful when processing structural images according to a handheld capture method.
[0121] A more detailed description of the scale measurement component 406 follows. The term "scale reference" refers to an imaged object of known size, such as a banknote or ID card. The term "scale homolog" refers to an imaged object or distance shared by two or more images and used to infer the relative size or scale of objects in each image. This scale homolog is also shared with the structural image and can therefore be used to scale the structural image and any measurements performed therein. Various scale homologs and scale references can be used, with examples provided below.
[0122] The following is an example embodiment of the scaling component 406. A user can capture an image of themselves holding a card of a known size (e.g., 85.60 mm by 53.98 mm) at their face (e.g., in front of their mouth and perpendicular to their face). This image can be captured in a manner similar to that of the structural image 405 (e.g., before or after the capture of the structural image 405, so that the card does not otherwise obstruct the capture process), and can be captured from a position perpendicular to their face. The card can then be used as a scaling reference to measure the physical interpupillary distance between the user's pupils in millimeters, which can later be used as a scaling homolog by the context-transformation component 410 to apply an absolute scale to the structural image. This is possible because the structural image capture includes one or more images in which the user's pupils are visible. The scaling component 406 can implement one or more neural networks to detect and measure the scaling reference and scaling homolog, and computer vision processes can be applied to refine the measurement.
[0123] In one embodiment, the face detection algorithm used in the face feature point detection component 408 can also be used to locate the pixel coordinates of the user's pupils in the scale image. The scale measurement component 406 can also use a pre-trained neural network and / or computer vision techniques to define the boundaries and / or feature points of the scale reference. In this example embodiment, the scale measurement component 406 can use a pre-trained neural network to estimate the corners of the card. Next, thick lines can be fitted to each pair of points describing the corners of the card detected by the neural network. The pixel distance between these thick lines can be divided by the known size of the scale reference in millimeters to derive the number of pixels per millimeter of the image at the distance from the card.
[0124] For accuracy, the scale measurement component 406 can perform the following computer vision techniques to fine-tune the measurement of the card. First, a Canny edge detection algorithm can be applied to a normalized version of the scale image. Then, a Hough transform can be used to infer the fine lines in the scale image. The area between each thick line and each thin line can be calculated. Then, a threshold number (e.g., 10 pixels of the image size) can be used to select only those fine lines that are separated into small regions by the coarse neural network prediction. Finally, the median of the selected fine lines can be selected as the final boundary of the card and used to derive the pixels per millimeter of the image as described above. Because the card and the user's pupil are at similar distances from the camera, the distance between the user's pupils in pixels can be divided by the pixels calculated per millimeter to measure the user's interpupillary distance in actual millimeters. The scale measurement component 406 then passes this interpupillary distance as a homologous measure 413 to the context transformation component 410.
[0125] The above embodiments of the scaling technique have been observed to be relatively accurate and usable on a wide range of input devices containing a standard camera. Where additional sensors are available for the input device, the following embodiments are proposed as alternatives and can be used to infer scale information without a scale reference.
[0126] The second process for measuring proportionate homologs is a multi-mode approach that utilizes not only a camera but also a microphone and speaker (e.g., a pair of earplugs), all of which can serve as capture devices (e.g., Figure 1 The user input device 110a (e.g., a mobile phone) is a component. Since sound is known to travel reliably at a speed of 343 m / s, it is expected that the signal sound played on the earpiece near the user's face will be recorded on the phone with exactly the amount of delay required for the sound to travel to the phone's microphone. This delay can be multiplied by the speed of sound to obtain the distance from the phone to the face. An image of the face can be captured simultaneously with the sound being emitted, and this image will contain proportional homologs, such as the user's eyes. The proportional measurement component can use some simple trigonometric functions to calculate the distance between the user's pupils or another pair of reference points, in metric units:
[0127] d = delay * sos
[0128] w_mm=2*tan(aov / 2)*d
[0129] ipd_mm = w_mm * ipd_pix / w_pix
[0130] (In the above equations, ipd_pix is the pixel distance between the user's pupils, ipd_mm is the millimeter distance between the user's pupils, sos is the speed of sound in millimeters per millisecond, delay is the delay in millimeters between the played signal and the recorded sound signal, w_mm is the horizontal dimension in millimeters of the imaging plane at a distance of d, w_pix is the horizontal dimension of the image in pixels, and aov is the horizontal angle of view of the imaging camera.)
[0131] Wireless earbuds can also be used, and it is conceivable that even over-ear headphones could be used; the process can begin simply by turning on both wireless earbuds and over-ear headphones. The volume of the recorded signal can be used as an indicator of the proximity between the earbuds and the microphone. The earbuds can typically be placed close to a proportionally similar source. The sound signal can be within, below, or above the threshold of human hearing, such as chirps, frequency sweeps, dolphin calls, or other such attractive sounds. The sound may be very short (e.g., less than one second), thus allowing for numerous measurements (e.g., over many seconds) for purposes of redundancy, averaging, and statistical analysis.
[0132] Another option for establishing the scale of a structural image (which can be used in an alternative capture process using a physical scanning camera around the head) is to use input from a user input device (e.g., Figure 1 The inertial measurement unit (IMU) data in 110a) establishes the absolute distance between camera positions (e.g., accelerometer, gyroscope, etc., with acceptable tolerances).
[0133] Another option for establishing a scaled image, which can be used in any capture process embodiment, is to use image range imaging, which can be supplied by the input device in various forms. For example, some input devices (e.g., modern mobile phones) are equipped with range cameras that utilize techniques such as structured light, pixel separation, or interferometry. When any of these techniques is used in combination with a standard camera, an estimate of the depth for a given pixel can be derived via known methods. Therefore, these techniques can be used to directly estimate the distance between the scaled source and the camera, and subsequent processes for measuring the scaled source can be implemented as described above.
[0134] The facial landmark detection component 408 can perform face detection as follows. The facial landmark detection component 408 can extract landmarks from sharp frames of image set 409 using histograms of oriented gradients. As an example, the facial landmark detection component 408 can implement the process described by Navneet Dalal and Bill Triggs, namely, Histograms of Oriented Gradients for Human Detection, Computer Vision & Pattern Recognition (CVPR'05), June 2005, San Diego, United States, pp. 886-893. The facial landmark detection component 408 can implement a support vector machine (SVM) using a sliding window method to classify the extracted landmarks. The facial landmark detection component 408 can use non-maximum suppression to reject multiple detections. The facial landmark detection component 408 can use a model that has been pre-trained on several faces (e.g., 3000 faces).
[0135] The facial landmark detection component 408 can perform 2D coordinate detection using an ensemble of regression trees. As an example, the facial landmark detection component 408 can implement the process described by Vahid Kazemi and Josephine Sullivan, namely, One Millisecond Face Alignment with an Ensemble of Regression Trees, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1867–1874. The facial landmark detection component 408 can identify several facial landmarks (e.g., five landmarks, such as the inner and outer edges of each eye and the columella of the nose). The facial landmark detection component 408 can use a model pre-trained on several faces (e.g., 7198 faces) to identify facial landmarks. Feature points can include several points corresponding to various facial features, such as points that define the bounding boxes or boundaries of the face (cheeks, outer chin, etc.), eyebrows, pupils, nose (bridge of nose, nostrils, etc.), mouth, etc.
[0136] In another embodiment, a convolutional neural network can be used for 2D facial feature point detection. The facial feature point detection component 408 can use a neural network model trained on a database of annotated faces to identify several facial feature points (e.g., five or more). The feature points can contain the 2D coordinates of various facial feature points.
[0137] Context transformation component 410 uses a series of coordinate and orientation transformations to normalize camera transformation 411 with respect to the user's head. The context of the transformation is similar to the localization and orientation of a person's head measured or calculated in the generation of HRTF. For example, in one embodiment, target data (see...) Figure 5 The HRTF used in the training environment 502) is calculated or measured using the x-axis extending between the user's ear canals and the y-axis extending through the user's nose.
[0138] The context transformation component 410 first uses a least squares algorithm to find the plane that best fits the position data of the camera transformation 411. The context transformation component 410 can then rotate the best-fit plane (along with the camera transformation 411) to the xy plane (so that the z-axis is accurately described as "up" and "down").
[0139] To perform absolute scaling and other subsequent transformations, the context transformation component 410 can estimate the positions of several key feature points of facial feature points 415 in world space, which may include the pupils and nose. The system can use a process similar to that of the 3D projection component 416 (which will be detailed below) to locate these feature points using a complete set of 2D facial feature points 415. Once the 3D positions of the facial feature points are determined, the following transformations can be performed.
[0140] First, the origin of world space can be centered by subtracting the arithmetic mean of the 3D coordinates of the eyes from the position information of camera transformation 411 and the 3D positions of facial feature points 415. Next, to apply an absolute scale to world space, the context transformation component 410 can multiply the position information from camera transformation 411 and the 3D positions of facial feature points 415 by a scaling factor. This scaling factor can be derived by dividing the homologous measure 413 by an estimate of the interpupillary distance calculated using the 3D positions of the left and right sides of each eye or the 3D position of the pupil itself. This scaling process allows the HRTF synthesis system 400 to use real-world (physical) distances, as physical distances are specifically related to sound waves and their subsequent reflection, diffraction, absorption, and resonance behavior.
[0141] At this point, the scaled, centered camera transform needs to be oriented about the vertical axis of world space (which may be called the z-axis). In photography, a photograph in which the subject's face and nose are directly pointed at the camera is usually called a full-face photograph. It may be useful to rotate the camera transform 411 about the z-axis of world space so that the full-face frame of image set 409 corresponds to a camera transform that is oriented at zero degrees relative to one of the other two axes of world space (e.g., the y-axis). The context transform component 410 can perform the following process to identify the full-face frame of image set 409. The context transform component 410 minimizes the point asymmetry of facial feature points 415 to find the “full-face” reference frame. For example, a full-face frame can be defined as the frame in which each pair of pupils is closest to being equidistant from the nose. In mathematical terms, the context transform component 410 can calculate the asymmetry according to the asymmetry function |L–R| / F, where L is the centroid of the left feature point, R is the centroid of the right feature point, and F is the centroid of all facial feature points 415. The full-face frame is then the frame in which the asymmetry is minimized. Once the full frame is selected, the 3D positions of the camera transformation 411 and facial feature points 415 can be rotated around the z-axis, so that the full frame of the image set 409 corresponds to a camera transformation positioned at zero degrees relative to the y-axis of world space.
[0142] Finally, the context transformation component 410 translates the camera transformation 411 into the head, so that the origin of world space does not correspond to the center of the face, but rather to an estimated point between the ears. This can be simply accomplished by translating the camera transformation along the y-axis by the average of the orthogonal distances between the face and the interauricular axis. This value can be calculated using human anthropometric data from a mesh database (see [link to relevant documentation]). Figure 5 (Grid data 511 in the dataset). For example, the authors have found that the average orthogonal distance between the eyebrows and ears in the dataset is 106.3 mm. In one embodiment, to illustrate the angular pitch of the head, the camera transformation can also be rotated about the interauricular axis (the axis between the ear canals). This rotation can be made such that the 3D position of the nose is along the y-axis.
[0143] As a result of the above process, the context transformation component 410 generates a centered, horizontal, and scaled camera transformation 417. These camera transformations 417 can be used by the cropping component 412 to estimate images in which points describing the user's acoustic anatomy are visible. For example, assuming the tip of the user's nose is aligned at 0 degrees with the y-axis in world space, the image in which the camera transformation 417 is positioned clockwise around the z-axis in world space between 30 and 100 degrees might contain structural images of the user's right ear. Once these images are selected from the image set 409, they can be cropped to include only portions of the image that are estimated to contain structural images with the feature points of interest.
[0144] To crop each image to the appropriate portion, the cropping component 412 calculates the approximate position and size of a 3D point cloud containing the anatomical features of interest. For example, the average distance between ear canals is approximately 160 mm, and the camera transformation 411 has been centered around the estimated bisection point of this line. Therefore, assuming the tip of the user's nose is aligned with the y-axis of the world, the position of the point cloud for each ear can be expected to be approximately 80 mm along the x-axis of world space in any direction. In this example, the size of each 3D point cloud might be approximately 65 mm in diameter, describing the average length of the ear.
[0145] The following techniques can now be used to achieve cropping. The orientation information of each camera's transformation describes how the camera's three axes are linearly related to the three axes of world space, and conventionally, the camera's principal axis is considered as a vector describing travel from the camera's position through the center of the image frame. The feature point line is considered as a line between the camera's position in world space and the estimated position of the feature point cloud. The rotation matrix of the camera transformation can be directly used to express the feature point line in camera space or a specific camera's 3D coordinate system. The camera's viewpoint is an inherent parameter that can be calculated using a 35mm equivalent focal length or the camera's focal length and sensor size, both of which can be derived from a camera lookup table, from EXIF data encoded for each image, or from the input device itself at capture time. The feature point line can be projected onto the image using the camera's viewpoint and the image's pixel dimensions. For example, the horizontal pixel distance ("x") between the image center and feature points in the image plane can be approximated as follows:
[0146] d_pix = (w_pix / 2) / tan(aov / 2)
[0147] x_pix = d_pix * tan(fax / 2)
[0148] (In the above equations, d_pix is the distance between the camera and the image plane in pixels, w_pix is the horizontal dimension of the image in pixels, aov is the horizontal angle of view of the imaging camera, and fax is the horizontal angular component of the feature point line.)
[0149] Once the pixel location is close to the center of the feature point cloud, a similar method can be used to calculate the appropriate width and height for cropping. For example, since almost all ears are less than 100 mm along their longest diagonal, a 100 mm crop is reasonable for locating the ear's feature points. As described in the following steps, many neural networks use square images as input, meaning the final cropped image will have the same height and width. For this ear example, a crop of + / - 50 mm vertically and horizontally from the center of the feature points is appropriate. The distance from the camera to the image plane in world space units can be calculated by calculating the magnitude of the orthogonal projection of the feature point line onto the camera's principal axis in world space. Since this distance has been calculated in pixels above, the pixel-per-millimeter ratio can be calculated and applied to the 50 mm crop size to determine the cropping boundaries in pixels. Once the cropping component 412 has completed this cropping process, the image can be rescaled and used in the feature point detection component 414 as described below.
[0150] To identify 2D coordinates 421, the feature point detection component 414 can use a neural network. For example, the personalization server 120a (see...) Figure 1 ) or user input device 110a (see Figure 1 Neural networks can be used as components for implementing neural networks 326 (see [link]). Figure 3 ), facial detection component (see Figure 4 (410) or facial landmark detection component (see 410) Figure 4 The system implements a convolutional neural network (CNN) to label anatomical feature points, as shown in section 408. According to one embodiment, the system can use a MobileNets architecture pre-trained on the ImageNet database to implement the CNN by running the Keras neural network library (written in Python) on top of the TensorFlow machine learning software library. The MobileNets architecture can use an alpha multiplier of 0.25 and a resolution multiplier of 1, and can be trained on a database of structural images to detect facial feature points used to construct anthropometric data 425.
[0151] For example, an image (e.g., one of the structural image frames 313) may be downsampled to a smaller resolution image and may be represented as a tensor (multidimensional data array) of size 224×224×3. The image may be processed by a MobileNets architecture trained to detect a given set of feature points, resulting in a tensor of size 1×(2*n) that recognizes the x and y coordinates of n feature points. For example, this process may be used to generate the x and y coordinates of 18 ear feature points and 9 torso feature points. In different embodiments, the Inception V3 architecture or different convolutional neural network architectures may be used. The cropping component 412 and the feature point detection component 414 may be used repeatedly or simultaneously for different images and / or different sets of feature points.
[0152] To estimate the singular values of the coordinates of each feature point, the 2D coordinates 421 from the feature point detection component 414 are passed to the 3D projection component 416 along with the camera transformation 417. The 3D projection component 416 projects the 2D coordinates 421 from each camera into world space and then performs a least-squares calculation to approximate the intersection of a set of projected rays for each feature point. For example, the description of the clipping component 412 above details how feature point lines in world space can be projected onto the image plane using a series of known photogrammetric methods. This process is reversible, such that each feature point in the image plane can be represented as a feature point line in world space. Several known methods exist for estimating the intersection of multiple lines in 3D space, such as least-squares solutions. When the 3D projection component 416 ends, multiple feature point positions, which can be collected from different angular ranges and / or using different neural networks, have been calculated in world space. This set of 3D coordinates 423 may also include 3D coordinates calculated by other methods, such as the above calculation of the position of each pupil.
[0153] In one embodiment, it may be useful to sequentially repeat the processing of context transformation component 410, cropping component 412, feature point detection component 414, and 3D projection component 416 as part of an iterative refinement process. For example, the initial iteration of context transformation component 410 may be considered as “coarse” localization and orientation of camera transformation 417, and the initial iteration of cropping component 412 may be considered as “coarse” selection and cropping from image set 409. 3D coordinates 423 may contain the estimated position of each ear, which can be used to repeat context transformation component 410 in “fine” iterations. In a preferred embodiment, the “coarse” crop may be significantly larger than the estimated size of the feature point cloud to allow for errors in feature point line estimation. As part of the refinement process, cropping component 412 may be repeated with a more compact, smaller image crop after fine iteration of context transformation component 410. This refinement process may be repeated multiple times as needed, but in one embodiment, at least one refinement iteration is recommended. This is recommended because the authors have found that the feature point detection component 414 is more accurate when the cropping component 412 uses a more compact cropping; however, the cropping must include the structural image of the entire feature point set, and therefore must be set using an accurate estimation of the feature point lines.
[0154] The actual anthropometric data used by the system to generate personalized HRTFs are scalar values representing the length of anatomical features and the angles between anatomical features. These calculations are specified and performed across a subset of 3D coordinates 423 by the angle and distance measurement component 418. For example, the angle and distance measurement component 418 can specify the calculation of “shoulder width” as the Euclidean distance between the “left shoulder” and “right shoulder” coordinates belonging to a set of 3D coordinates 423. As another example, the angle and distance measurement component 418 can specify the calculation of “auricular spread angle” as the angular representation of the horizontal component of the vector between the “anterior outer ear” coordinate and the “superior auricle” coordinate. A set of known anthropometric methods for HRTF calculations have been proposed and can be collected during this process. For example, the anthropometric measurements determined from 3D coordinates 423 may include, for each auricle of the user, auricle spread angle, auricle rotation angle, auricle cleft angle, auricle backward offset, auricle downward offset, auricle height, first auricle width, second auricle width, first interauricular width, second interauricular gap width, fossa height, outer ear width, outer ear height, and concha height. At this point, the data compilation component 420 can aggregate the obtained anthropometric data 425 and the previously mentioned demographic data 403 to form the generated data 427 required to produce a personalized HRTF.
[0155] When compiling generated data 427, data compilation component 420 may perform so-called moderate degradation. Moderate degradation may be used when one or more predictors of demographic data 403 are not provided, or when the identification of one or more predictors of generated data 427 fails or is uncertain. In this case, data compilation component 420 may generate an estimate of the missing predictor based on other known predictors of generated data 427, and the estimated predictor may then be used as part of generating a personalized HRTF. For example, if the system cannot determine the measurement of shoulder width, the system may use demographic data (e.g., age, sex, weight, height, etc.) to generate an estimate for shoulder width. As another example, data compilation component 420 may use calculations of some auricular characteristics with high confidence metrics (e.g., low error in the least squares solution) to estimate values of other auricular characteristics calculated with lower confidence. Predetermined relations may be used to implement estimations of subsets of generated data 427 using other subsets. For example, as part of the training environment (see...) Figure 5 In section 502), the system can use anthropometric data from a training database of high-resolution grid data to perform linear regression between sets of generated data. As another example, the publicly released U.S. Army Personnel Anthropometric Survey (ANSUR 2 or ANSUR II) contains certain characteristics that can be included as predictive expressions in the generated data and used in the linear regression methods described above. In summary, data compilation component 420 avoids the problem of missing data by estimating any missing values from available information in demographic data 403 and anthropometric data 425. The use of a complete set of generated data 427 further avoids the need to consider missing data in the HRTF calculation system.
[0156] Implementation details
[0157] The embodiments may be implemented in hardware, as executable modules stored on computer-readable media, or a combination of both (e.g., a programmable logic array). Unless otherwise stated, the steps performed by the embodiments need not inherently relate to any particular computer or other device, although they may be in some embodiments. Specifically, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized devices (e.g., integrated circuits) to perform the desired method steps. Therefore, the embodiments may be implemented in one or more computer programs that execute on one or more programmable computer systems, each comprising at least one processor, at least one data storage system (containing volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and produce output information. The output information is applied to one or more output devices in a known manner.
[0158] Each such computer program is preferably stored on or downloaded to a storage medium or device readable by a general-purpose or special-purpose programmable computer (e.g., solid-state memory or media, or magnetic or optical media) for configuring and operating the computer when the storage medium or device is read by a computer system to execute the program described herein. The inventive system may also be considered as a computer-readable storage medium configured with a computer program, wherein such a configuration causes the computer system to operate in a specific and predefined manner to perform the functions described herein. (Software itself and intangible or transient signals are excluded as they are non-patentable subject matter.)
[0159] The foregoing description illustrates various embodiments of the invention and examples of how aspects of the invention can be implemented. These examples and embodiments should not be considered as limited embodiments, but are presented to illustrate the flexibility and advantages of the invention as defined by the appended claims. Other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art based on the foregoing disclosure and the appended claims, and may be employed without departing from the spirit and scope of the invention as defined by the claims.
Claims
1. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: Capture video data of a user, wherein the video data includes multiple views of the user's head; Process the video data to extract multiple images of the user; Receive the multiple images from the user; Processing the plurality of images to generate the user's anthropometric data, wherein processing the plurality of images to generate the anthropometric data includes: Convert the multiple images into a 3D point cloud model; The 3D point cloud model is used to select key image frames from the plurality of images; and The human body measurement data is generated from the key image frames; and The anthropometric data is input into the HRTF calculation system to obtain the personalized HRTF.
2. The method of claim 1, wherein capturing the video data comprises capturing an image of an object having a known size; and Processing the plurality of images to generate the anthropometric data includes using the known size to convert the anthropometric data from pixel measurements to absolute distance measurements.
3. The method according to claim 1, wherein, The electronic device includes a camera, and the capture of the video data is performed using the camera.
4. The method of claim 1, wherein the means at a given distance captures the video data of the user, the method further comprising: The given distance is determined by measuring the time delay between sound output from headphones located close to the user and sound received at the microphone of the device. Processing the plurality of images to generate the anthropometric data includes using the given distance to convert the anthropometric data from pixel measurements to absolute distance measurements.
5. The method of claim 1, wherein identifying or selecting key image frames of the plurality of images is based on frame content and one or more sharpness indicators.
6. The method of claim 1, wherein generating the personalized HRTF comprises: Provides an HRTF model trained by performing a machine learning process on anthropometric data and a high-resolution database of measured amplitude / frequency responses; and The personalized HRTF is generated by applying the HRTF model to the user's anthropometric data.
7. The method of claim 6, wherein the machine learning process comprises lasso regression.
8. The method of claim 1, further comprising: The personalized HRTF is generated on the server device; and The personalized HRTF is transmitted from the server device to the user device.
9. The method of claim 1, wherein the electronic device is a user device; and the method further comprises: The personalized HRTF is generated on the user device.
10. The method of claim 1, wherein processing the plurality of images to generate the anthropometric data comprises using at least one of a photogrammetry component, a context transformation component, a feature point detection component, and an anthropometric component of the electronic device.
11. The method of claim 1, wherein processing the plurality of images to generate the anthropometric data includes using a feature point detection component, a 3D projection component, and an angle and distance measurement component. The feature point detection component receives a cropped image set of the user's anthropometric feature points and generates a set of 2D coordinates for the user's anthropometric feature points from the cropped image set. The 3D projection component receives the set of 2D coordinates and multiple camera transformations, and uses the camera transformations to generate a set of 3D coordinates corresponding to the set of 2D coordinates for each of the anthropometric feature points in 3D space. The angle and distance measurement component receives the set of 3D coordinates and generates human body measurement data from the set of 3D coordinates, wherein the human body measurement data corresponds to the angle and distance of the human body measurement feature points in the set of 3D coordinates. The electronic device generates the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system.
12. The method of claim 9, wherein the user device generates audio output by applying the personalized HRTF to an audio signal, wherein the user device comprises one of a headset, a pair of earpieces, and a pair of hearing devices.
13. The method of claim 1, wherein the anthropometric data includes at least one of the following: the user's shoulder width, the user's neck width, the user's neck height, the user's face height, the user's interpupillary distance, and the user's zygomatic distance.
14. The method according to claim 1, wherein the anthropometric data includes, for each auricle of the user, at least one of the following: auricle opening angle, auricle rotation angle, auricle splitting angle, auricle posterior displacement, auricle downward displacement, auricle height, auricle width, first interorbital width, second interorbital width, concha height, outer ear width, outer ear height, and concha height.
15. The method of claim 1, wherein the anthropometric data further comprises other data, wherein the other data includes the user's age, the user's weight, the user's gender, and the user's height, wherein the other data is obtained from a source other than processing the plurality of images.
16. The method of claim 12, wherein the audio signal comprises a plurality of audio objects containing location information, wherein generating the audio output corresponds to generating binaural audio output by applying the personalized HRTF to the plurality of audio objects.
17. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: Capture video data of a user, wherein the video data includes multiple views of the user's head; Process the video data to extract multiple images of the user; Receive the multiple images from the user; Processing the plurality of images to generate the user's anthropometric data, wherein processing the plurality of images to generate the user's anthropometric data includes: The first frame of the plurality of images is identified by minimizing the asymmetry of key points in the first frame, wherein the first frame is a view perpendicular to the user's face; Identify a second frame of the plurality of images, wherein the second frame is a view perpendicular to the user's first auricle, based on a view at a 90-degree angle to the first frame; and Identify a third frame in the plurality of images, wherein the third frame is a view perpendicular to the user's second auricle based on a view at a 180-degree angle to the second frame; and The anthropometric data is input into the HRTF calculation system to obtain the personalized HRTF.
18. The method of claim 17, wherein the second frame is one of a plurality of second frames selected from the plurality of images within +45 degrees and -45 degrees of the view perpendicular to the first auricle; and The third frame is one of a plurality of third frames selected from the plurality of images within +45 degrees and -45 degrees around the view perpendicular to the second auricle.
19. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: Capture video data of a user, wherein the video data includes multiple views of the user's head; Process the video data to extract multiple images of the user; Receive the multiple images from the user; Processing the plurality of images to generate the user's anthropometric data, wherein processing the plurality of images to generate the user's anthropometric data includes: Identify key image frames from the multiple images; The key image frames are used to identify the user's anthropometric features; and The anthropometric data is generated by determining the measurement of the anthropometric features; and The anthropometric data is input into the HRTF calculation system to obtain the personalized HRTF.
20. A non-transitory computer-readable medium storing one or more computer programs that, when executed by one or more processors, control a device to perform processing comprising the method according to any one of claims 1 to 19.
21. An apparatus for generating a personalized head-related transfer function (HRTF), the apparatus comprising: At least one processor; and At least one memory, The at least one processor is configured to perform processing comprising the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
HRTF personalization based on anthropometric features
US20150312694A1