Personalized HRTF with optical capture

The system addresses individual anatomical variability by using optical capture and anthropometric data to generate personalized HRTFs, improving sound localization and timbre accuracy for individual users.

JP7785824B2Active Publication Date: 2025-12-15DOLBY LABORATORIES LICENSING CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024023315
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-07-25
Filing Date
2024-02-20
Publication Date
2025-12-15
Estimated Expiration
2039-07-25

AI Technical Summary

Technical Problem

Existing HRTF technologies struggle with individual variability in human anatomy, leading to inaccuracies in sound localization, timbral naturalness, and depth perception, particularly when using generic HRTFs derived from dummy heads or averages.

Method used

A system that uses optical image capture and anthropometric measurements to derive personalized HRTFs by training a calculation system with 3D scans and demographic data, employing machine learning to generate customized HRTFs for individual users.

Benefits of technology

Improves sound source localization accuracy and overall audio perception by tailoring HRTFs to individual listeners, enhancing spatial awareness and timbre fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785824000001
    Figure 0007785824000001
  • Figure 0007785824000002
    Figure 0007785824000002
  • Figure 0007785824000003
    Figure 0007785824000003
Patent Text Reader

Abstract

To provide a device and a method, of generating a personalized head part transfer function (HRTF).SOLUTION: An audio eco system 100 contains: a user device 110 (a user input device 110a and a user output device 110b); and a cloud device 120 (a personalized server 120a and a content server 120b). The user input device captures a generation data 130 of a user, receives and processes the generation data from the user input device, and generates and stores a personalized HRTF 132 on the user. The content server provides a content 134 to a user output device. The user output device receives the personalized HRTF from the personalized server, receives the content from the content server, and generates an audio output 136 by applying the personalized HRTF to the content.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 703,297, filed July 25, 2018, for "Method and Apparatus for Personalized HRTFs via Optical Capture," which is incorporated herein by reference.

[0002] Field This disclosure relates to audio processing, and more particularly to generating custom audio according to the anthropometric and demographic characteristics of the listener. [Background technology]

[0003] Unless otherwise stated herein, the approaches described in this section are not prior art to the claims of this application and are not admitted to be prior art by their inclusion in this section.

[0004] By placing sounds at various positions and recording them through a dummy head, playback of such recordings through headphones can achieve the corresponding perception of sounds coming from various positions relative to the listener. Because this approach has the undesirable side effect of causing a muffled sound when using standalone loudspeakers instead of headphones, this technique is often used for selected tracks of a multi-track recording rather than the entire recording. To improve this technique, the dummy material can include the shape of an ear (pinna) and can be designed to match the acoustic reflectivity / absorption of a real head and ear.

[0005] Alternatively, a head-related transfer function (HRTF) can be applied to the sound source so that the sound is spatially localized when heard through headphones. Generally, an HRTF corresponds to the acoustic transfer function between a point in three-dimensional (3D) space and the entrance to the ear canal. HRTFs arise from the passive filtering functions of the ear, head, and body and are ultimately used by the brain to infer sound location. HRTFs consist of magnitude and phase frequency responses as a function of elevation and azimuth (rotation around the listener) added to the audio signal. Rather than recording sounds at specific locations around a dummy head, sounds can be recorded in many ways, including traditional approaches, and then processed using HRTFs to appear at desired locations. Of course, many sounds can be generated simultaneously at various locations through superposition, to recreate either real-world audio environments or simply artistic intent. Furthermore, sounds can be digitally synthesized instead of being recorded.

[0006] A refinement of the HRTF concept involves recording from an actual human ear canal. Summary of the Invention [Problem to be solved by the invention]

[0007] In the process of improving HRTFs by recording from real human ear canals, it has been discovered that HRTFs vary widely from person to person due to individual differences in anatomy, such as shoulder bulk, head size, pinna shape, and other facial features. Furthermore, even within a single individual, there are subtle differences between the left and right ears. This individualized behavior of HRTFs poses challenges when using generic HRTFs, such as those designed from dummy heads, single individuals, or averages across many individuals. The use of generic HRTFs typically results in problems with location accuracy, such as difficulty placing sounds in front of the face, front-to-back inversion, transmission at specific distances from the head, and angular accuracy. Furthermore, generic HRTFs are commonly found to lack timbral or spectral naturalness and overall depth perception in the soundstage. This increased understanding has led to ongoing efforts to tailor HRTFs for specific listeners through a variety of techniques. [Means for solving the problem]

[0008] This paper describes a system for personalizing binaural audio playback to substantially improve the accuracy of perceived sound source location. In addition to demographic information provided by the user, the system uses optical image capture of the user's anthropometric measurements, such as shoulder, head, and ear shape. These data are used to derive a personalized HRTF for the user. This personalized HRTF is then used to process sound sources represented as positioned sound objects, where the locations can range from pinpoint locations to diffuse sources (e.g., using audio objects, as in the Dolby® Atmos™ system). In some embodiments, the sound sources may be multichannel formats, or even stereo sound sources converted into positioned sound objects. The sound sources may be for video, music, dialogue enhancement, video games, virtual reality (VR), augmented reality (AR) applications, etc.

[0009] According to one embodiment, a method for generating head-related transfer functions (HRTFs) includes generating an HRTF calculation system and using the HRTF calculation system to generate personalized HRTFs for a user. Generating the HRTF calculation system includes generating HRTFs for a plurality of training subjects by measuring multiple 3D scans of the plurality of training subjects and performing acoustic scattering calculations on the multiple 3D scans, collecting generative data for the plurality of training subjects, and training the HRTF calculation system to convert the generative data into the plurality of HRTFs. Generating the personalized HRTF includes collecting user generative data and inputting the user generative data into the HRTF calculation system to obtain personalized HRTFs.

[0010] Performing the training may include using linear regression with Lasso regularization.

[0011] The user-generated data may include at least one of anthropometric measurements and demographic data.

[0012] The anthropometric measurements may be obtained by collecting multiple images of the user and determining the anthropometric measurements using the multiple images. Determining the anthropometric measurements using the multiple images may be performed using a convolutional neural network. The method may further include scaling the anthropometric measurements of the user using a reference object in at least one image of the multiple images of the user.

[0013] The method may further include generating an audio output by applying the personalized HRTF to the audio signal.

[0014] The method may further include storing, by the server device, the personalized HRTF and transmitting, by the server device, the personalized HRTF to the user device, wherein the user device generates the audio output by applying the personalized HRTF to the audio signal.

[0015] The method may include generating, by a user device, an audio output by applying the personalized HRTF to an audio signal, wherein the user device includes one of a headset, an earphone, and a hearable.

[0016] The audio signal may include a plurality of audio objects including position information, and the method may further include generating a binaural audio output by applying the personalized HRTFs to the plurality of audio objects.

[0017] According to another embodiment, a non-transitory computer-readable medium stores a computer program that, when executed by a processor, controls an apparatus to perform processes including one or more of the methods described above.

[0018] According to another embodiment, an apparatus generates head-related transfer functions (HRTFs). The apparatus includes at least one processor and at least one memory. The at least one processor is configured to generate an HRTF calculation system and control the apparatus to generate personalized HRTFs for a user using the HRTF calculation system. Generating the HRTF calculation system includes generating HRTFs for a plurality of training subjects by measuring multiple 3D scans of the plurality of training subjects and performing acoustic scattering calculations on the multiple 3D scans, collecting generative data for the plurality of training subjects, and training the HRTF calculation system to convert the generative data into the plurality of HRTFs. Generating the personalized HRTF includes collecting user generative data and inputting the user generative data into the HRTF calculation system to obtain personalized HRTFs.

[0019] The generated data of the user may include at least one of anthropometric measurements and demographic data. The apparatus may further include a user input device configured to collect a plurality of images of the user and determine anthropometric measurements of the user using the plurality of images of the user, and the anthropometric measurements of the user may be scaled using a reference object in at least one image of the plurality of images of the user.

[0020] The apparatus may further comprise a user output device configured to generate an audio output by applying the personalized HRTF to an audio signal.

[0021] The device may further include a server device configured to generate the HRTF calculation system, generate the personalized HRTF, store the personalized HRTF, and transmit the personalized HRTF to a user device, the user device being configured to generate an audio output by applying the personalized HRTF to an audio signal.

[0022] The apparatus may have a user device configured to generate an audio output by applying a personalized HRTF to an audio signal, the user device including one of a headset, earphones, and hearables.

[0023] The audio signal may include a plurality of audio objects including position information, and the at least one processor is configured to control the device to generate a binaural audio output by applying personalized HRTFs to the plurality of audio objects.

[0024] The apparatus may include a server device configured to generate a personalized HRTF for a user using an HRTF calculation system, the server device executing a photogrammetry component, a context transformation component, a landmark detection component, and an anthropometry component. The photogrammetry component is configured to receive a plurality of structural images of the user and generate a plurality of camera transformations and a structural image set using a structure-from-motion technique. The context transformation component is configured to receive the plurality of camera transformations and the structural image set and generate a plurality of transformed camera transformations by translating and rotating the plurality of camera transformations using the structural image set. The landmark detection component is configured to receive the structural image set and the transformed plurality of camera transformations and generate a set of 3D landmarks corresponding to the user's anthropometric landmarks identified using the structural image set and the transformed plurality of camera transformations. The anthropometry component is configured to receive the 3D landmark set and generate anthropometric data from the 3D landmark set, the anthropometric data corresponding to a set of distances and angles measured between individual landmarks in the 3D landmark set. The server device is configured to generate a personalized HRTF for the user by inputting anthropometric data into the HRTF calculation system.

[0025] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF calculation system, the server device executing a scale measurement component configured to receive a scale image including an image of a scale reference and generate a homologue measure, the server device configured to scale a structural image of the user using the homologue measure.

[0026] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF calculation system, the server device executing the landmark detection component, the 3D projection component, and the angle and distance measurement component. The landmark detection component is configured to receive a cropped image set of the user's anthropometric landmarks and generate a set of 2D coordinates for the set of anthropometric landmarks from the cropped image set. The 3D projection component is configured to receive the set of 2D coordinates and a plurality of camera transformations and generate a set of 3D coordinates corresponding to the set of 2D components of each anthropometric landmark in 3D space using the camera transformations. The angle and distance measurement component is configured to receive the set of 3D coordinates and generate anthropometric data from the set of 3D coordinates, the anthropometric data corresponding to the angles and distances of the anthropometric landmarks in the set of 3D coordinates. The server device is configured to generate a personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system.

[0027] The HRTF computation system may be configured to train a model corresponding to one of a left ear HRTF and a right ear HRTF, in which case the personalized HRTF is generated by using the model to generate one of a left ear personalized HRTF and a right ear personalized HRTF, and using a reflection of the model to generate the other of the left ear personalized HRTF and the right ear personalized HRTF.

[0028] The apparatus may further include a server device configured to generate personalized HRTFs for the user using the HRTF calculation system, the server device executing a data aggregation component configured to implement graceful degradation of the generated data that fills in missing portions of the generated data using estimates determined from known portions of the generated data.

[0029] The apparatus may further include a server device configured to generate the HRTF computation system, the server device executing a dimensionality reduction component configured to reduce the computational complexity of a training run of the HRTF computation system by performing a principal component analysis on the multiple HRTFs for the multiple training subjects.

[0030] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF calculation system, the server device executing a photogrammetry component configured to receive a plurality of structural images of the user, perform a constrained image feature search using a facial landmark detection process on the plurality of structural images, and generate a plurality of camera transforms and a set of structural images using a structure-from-motion technique and results of the constrained image feature search.

[0031] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF computation system, the server device executing a context transformation component configured to receive the first plurality of camera transformations, the plurality of facial landmarks, and the scale measure, translate and rotate the plurality of camera transformations using the plurality of facial landmarks to generate a second plurality of camera transformations, and scale the second plurality of camera transformations using the scale measure.

[0032] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF calculation system, the server device executing a scale measurement component configured to receive range imaging information and generate a homolog measure using the range imaging information, and the server device configured to scale a structural image of the user using the homolog measure.

[0033] The apparatus may further include a user input device and a server device. The user input device is associated with a speaker and a microphone. The server device is configured to generate a personalized HRTF for the user using an HRTF calculation system, and the server device executes a scale measurement component. The scale measurement component is configured to receive arrival time information from the user input device and generate a homology measure using the arrival time information, the arrival time information relating to sound output by a speaker at a first location and received by a microphone at a second location, the first location associated with the user and the second location associated with the user input device. The server device is configured to scale a structural image of the user using the homology measure.

[0034] The apparatus may further include a server device configured to generate a personalized HRTF for the user using the HRTF computation system, the server device executing a cropping component and a landmark detection component, the cropping component and the landmark detection component cooperating to perform a constrained recursive landmark search by cropping and detecting a plurality of different sets of landmarks.

[0035] The apparatus may include similar details as those described above with respect to the method.

[0036] The following detailed description and accompanying drawings provide a further understanding of the nature and advantages of various implementations. [Brief explanation of the drawings]

[0037] [Figure 1] 1 is a block diagram of an audio ecosystem 100. [Figure 2] 2 is a flowchart of a method 200 for generating a head-related transfer function (HRTF). [Figure 3] FIG. 3 is a block diagram of an audio environment 300. [Figure 4] FIG. 4 is a block diagram of an anthropometric measurement system 400. [Figure 5] FIG. 5 is a block diagram of an HRTF calculation system 500. DETAILED DESCRIPTION OF THE INVENTION

[0038] This specification describes a technique for generating head-related transfer functions (HRTFs). In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may also include modifications and equivalents of the features and concepts described herein.

[0039] In the following description, various methods, processes, and procedures are set forth in detail. While specific steps may be described in a certain order, such order is primarily for convenience. Certain steps may be repeated more than once, may occur before or after other steps (even if those steps are described in a different order), or may occur in parallel with other steps. A second step is required to follow a first step only if the first step must be completed before the second step can begin. Such situations will be specifically pointed out if they are not clear from the context.

[0040] The terms "and," "or," and "and / or" are used herein. Such terms should be read as having an inclusive meaning. For example, "A and B" may mean at least "both A and B," or "at least both A and B." As another example, "A or B" may mean at least "at least A," "at least B," "both A and B," or "at least both A and B." As another example, "A and / or B" may mean at least "A and B," or "A or B." If an exclusive disjunction is intended, this will be specifically stated (e.g., "only one of A or B," "at most one of A and B").

[0041] For the purposes of this document, several terms are defined as follows: Acoustic anatomy refers to the parts of the human body that acoustically filter sound and contribute to HRTFs, including the upper body, head, and pinnae. Anthropometric data refers to a set of salient geometric measurements that can be used to describe a person's acoustic anatomy. Demographic data refers to demographic information provided by a person, which may include gender, age, race, height, and weight. Generative data is a combined set of complete or partial anthropometric and demographic data that can be used together to estimate a person's HRTFs. HRTF computation system refers to a function or set of functions that accepts arbitrary generative data as input and returns estimated personalized HRTFs as output.

[0042] As described in more detail herein, the general process for generating personalized HRTFs is as follows: First, an HRTF computation system is provided that expresses the relationship between an arbitrary set of generated data and a unique approximated HRTF. The system then efficiently derives a set of generated data using an input device, such as a mobile phone with a camera, and a processing device, such as a cloud personalization server. The provided HRTF computation system is then used on the new generated data to estimate a personalized HRTF for the user.

[0043] To prepare an HRTF calculation system for use in the present system, the following mathematical process is performed in a training environment: A database of mesh data composed of high-resolution 3D scans is created for multiple individuals. Demographic data for each individual is also included in the database. From the mesh data, a corresponding set of target data composed of HRTFs is created. In one embodiment, the HRTFs are obtained through a numerical simulation of the sound field around the mesh data. For example, this simulation can be achieved using the boundary element method or the finite element method. Another applicable known method for obtaining HRTFs that does not require a mesh is acoustic measurements. However, acoustic measurements require the subject to sit or stand stationary in an anechoic recording environment for a very long period of time, making the measurements prone to errors due to human movement and microphone noise. Furthermore, because acoustic measurements must be performed individually for each source position to be measured, increasing the sampling resolution of the sound sphere can be incredibly expensive. For these reasons, the use of numerically simulated HRTFs in a training environment is considered an improvement over HRTF database collection. Additionally, anthropometric data is collected for each individual in a database and combined with demographic data to form a set of generative data. A machine learning process then calculates the approximate relationship between the generative data and the target data, i.e., the model used as part of the HRTF calculation system.

[0044] Once prepared in the training environment, the HRTF computation system can be used to generate personalized HRTFs for any user without the need for mesh data or acoustic measurements. The system queries the user for demographic data and extracts anthropometric data from structural images using a series of photogrammetry, computer vision, image processing, and neural network techniques. For purposes of this description, the term structural imagery refers to multiple images in which the user's acoustic anatomy is visible, which may be a series of images or derived from "burst" images or video footage. It may be necessary to scale objects in the structural image to their true physical scale. Scaling imagery, which may be separate from or part of the structural image, may be used for this purpose, as described further herein. In some embodiments, a mobile device may be readily used to capture the structural image, as well as the corresponding demographic data and any necessary scaling imagery. The resulting anthropometric and demographic data is compiled into generated data, which is then used by a provided HRTF calculation system to generate personalized HRTFs.

[0045] 5 is a block diagram of an HRTF calculation system 500. The HRTF calculation system 500 may be implemented by the personalization server 120a (see FIG. 1), for example, by executing one or more computer programs. The HRTF calculation system 500 may also be implemented as a subcomponent of another component, such as the HRTF generation component 330 (see FIG. 3). The HRTF calculation system 500 includes a training environment 502, a database collection component 510, a numerical simulation component 520, a mesh annotation component 522, a dimensionality reduction component 524, a machine learning component 526, an estimation component 530, a dimensionality reconstruction component 532, and a phase reconstruction component 534.

[0046] The HRTF calculation system 500 may be prepared once, as described below, in a training environment 502, which may include computer programs for some components and manual data collection for others. Typically, the training environment 502 determines the relationship between a measured generative data matrix 523 for several subjects (perhaps hundreds) and the values ​​of their respective HRTFs. (This generative data matrix 523 and HRTFs may correspond to high-resolution mesh data 511, as described below.) The system uses front-end generative data approximations, obviating the need for 3D modeling, mathematical simulation, or acoustic measurements. By providing the machine learning component 526 with the "training" HRTFs and the relatively small generative data matrix 523, the system generates a model 533 that estimates the values ​​needed to synthesize the entire HRTF ensemble 543, which the system can store and distribute in the industry-standard spatially oriented format for acoustics (SOFA) format.

[0047] In the database collection component 510, demographic data 513 and 3D mesh data 511 are collected from a small number (a few hundred) of training subjects 512. Capturing high-resolution mesh scans can be a time-consuming and resource-intensive task. (This is one reason why personalized HRTFs are not more widely used, and one reason that motivates simpler methods for generating personalized HRTFs, such as those described herein.) For example, a high-resolution scan can be captured using an Artec 3D scanner that uses a 100,000 triangle mesh. This scan may require 1.5 hours of skilled post-editing work, followed by 24 hours of distributed server time, to numerically simulate the corresponding HRTFs. The database collection component 510 may be unnecessary if the HRTFs and generated data are obtained for use in the training environment directly from other sources, such as the Center for Image Processing and Integrated Computing (CIPIC) HRTF database from the UC Davis CIPIC Interface Laboratory.

[0048] In the numerical simulation component 520, in one embodiment, these "training" HRTFs may be calculated using the boundary element method and may be represented as an ITD matrix 525 and an H matrix 527. The H matrix 527 corresponds to a matrix of magnitude data for all of the training objects 512, consisting of the frequency impulse responses of the HRTFs, for any given position of the sound source. The ITD matrix 525 corresponds to a matrix of inter-aural time differences (ITDs, e.g., left ITD and right ITD), for all of the training objects 512, for any given position of the sound source. The HRTF simulation technique used in one embodiment requires highly sophisticated 3D image capture and a very tedious amount of mathematical calculations. For this reason, the training environment 502 is intended to be prepared only once with a finite amount of training data. The numerical simulation component 520 provides the H matrix 527 and the ITD matrix 525 to the dimensionality reduction component 524.

[0049] The mesh annotation component 522 outputs anthropometric data 521 corresponding to the anthropometric characteristics identified from the mesh data 511. For example, the mesh annotation component 522 may use manual annotation (e.g., an operator identifies the anthropometric characteristics). The mesh annotation component 522 may also use an angle and distance measurement component (see 418 in FIG. 4 ) to convert the annotated characteristics into measurements for the anthropometric data 521. The union of the anthropometric data 521 and the demographic data 513 is the generated data matrix 523.

[0050] In one embodiment, the dimensionality reduction component 524 may perform principal component analysis on the H matrix 527 and the ITD matrix 525 to reduce the computational complexity of the machine learning problem. For example, the H matrix 527 may have absolute frequency response values ​​for 240 frequencies, and the dimensionality reduction component 524 may reduce these to 20 principal components. Similarly, the ITD matrix 525 may have values ​​for 2500 source directions, and the dimensionality reduction component 524 may reduce these to 10 principal components. The dimensionality reduction component 524 provides collective principal component scores 529 of the H matrix 527 and the ITD matrix 525 to the machine learning component 526. The coefficients 531 needed to reconstruct the HRTFs from the principal component space are fixed and retained for use in the dimension recovery component 532, described below. Depending on the algorithm used in the machine learning component 526, other embodiments may omit the dimensionality reduction component 524.

[0051] The machine learning component 526 generally prepares a model 533 for use by the estimation component 530 in the generalized calculation of personalized HRTFs. The machine learning component 526 performs training of the model 533 to fit the generated data matrix 523 to the collective principal component scores 529 of the H matrix 527 and ITD matrix 525. The machine learning component 526 may use approximately 50 predictors from the generated data matrix 523 and may implement known backward, forward, or best subset selection methods to determine the optimal predictors for use in the model 533.

[0052] Once the components of the training environment 502 have been executed, a generalizable relationship between the generated data and the HRTFs has been established. The relationship includes the model 533 and, if dimensionality reduction is performed via the dimensionality reduction component 524, may also include the coefficients 531. This relationship can be used in the production steps described below to calculate personalized HRTFs corresponding to any new set of generated data. The production steps described below may be implemented by the personalization server 120a (see FIG. 1), for example, by executing one or more computer programs. The production steps described below may also be implemented as subcomponents of another component, such as the HRTF generation component 330 (see FIG. 3).

[0053] The estimation component 530 applies the model 533 to a set of generated data 535 to generate principal components 537 of the magnitude spectra and ITDs of the HRTFs. The generated data 535 may correspond to a combination of demographic data 311 and anthropometric data 325 (see FIG. 3). The generated data 535 may correspond to the generated data 427 (see FIG. 4). A dimension reconstruction component 532 then inverts the dimension reduction process using coefficients 531 on the components 537, resulting in an H matrix 541 of the full magnitude spectra and an ITD matrix 539 of the ITDs that describe the entire acoustic sphere. A phase reconstruction component 534 then uses the ITD matrix 539 and H matrix 541 to reconstruct a set of impulse responses that describe the HRTF set 543 with phase and magnitude information. In one embodiment, the phase reconstruction component 534 may implement a minimum phase reconstruction algorithm. Finally, the machine learning system 500 uses the set of impulse responses describing the HRTF ensemble 543 to generate personalized HRTFs, which may be expressed in SOFA format.

[0054] Further details of the HRTF calculation system 500 follow.

[0055] In one embodiment, the machine learning system 526 may implement linear regression to fit the model 533. For example, a set of linear regression weights may be calculated to fit all of the generated data matrix 523 to each individual directional slice of the absolute value score matrix. As another example, a set of linear regression weights may be calculated to fit all of the generated data matrix 523 to the entire vector of ITD scores. The machine learning system 526 may standardize each predictor vector in the generated data matrix 523 using z-score normalization.

[0056] As a regularization method, regression algorithms may use the least absolute shrinkage and selection operator ("lasso"). The lasso process works by identifying and ignoring parameters that are not relevant to the model at given locations (e.g., by shifting their coefficients toward zero). For example, interaural distance may serve as a predictor in the generated data, but may have little or no effect on the magnitude of the impulse response between the user's right ear and a sound source located directly to the user's right. Similarly, finer details of the pinna described by predictors in the generated data may have little or no effect on the interaural time difference. Ignoring irrelevant parameters may significantly reduce overfitting and therefore improve model accuracy. Lasso regression contrasts with ridge regression in that ridge regression scales the weights or contributions of all predictors and does not set any coefficients to zero.

[0057] In other embodiments, the machine learning system 526 may use other methods of machine learning to generate the set of HRTFs. For example, the machine learning system 526 may train a neural network to predict an entire matrix of absolute value scores. As another example, the machine learning system 526 may train a neural network to predict an entire vector of ITD scores. The machine learning system 526 may standardize the values ​​of the HRTFs via z-score normalization before training the neural network.

[0058] In some embodiments, the training environment 502 may be optimized by performing machine learning and / or dimensionality reduction on the transfer functions of only one ear. For example, a single HRTF set containing the transfer functions for the entire sphere around the head may be considered simply as one of two left-ear HRTF sets reflected about the sagittal plane. In this example, if the numerical simulation component 520 were performed on 100 subjects across the entire acoustic sphere with two ears as receivers, it could convert each subject's right-ear HRTF values ​​to left-ear values, creating a set of HRTF values ​​containing 200 examples of left-ear HRTFs. HRTFs may be represented as impulse responses, magnitude spectra, or interaural delays as a function of source position, and each right-ear position can be directly mapped to a left-ear position by reflecting its coordinates about the sagittal plane. This conversion can be performed by assigning the right-ear HRTF value to a matrix index of the reflected position.

[0059] Because the predictors in the generative data matrix 523 used to train the model 533 are scalar-valued, these predictors may also be considered independent of the side of the body on which they were measured. Thus, the model 533 may be trained to approximate only the left-ear HRTF, for example. The process of generating a user's right-ear HRTF is as simple as mapping the spherical coordinates of the HRTF ensemble generated using the generative data for the right ear back to the original coordinates. Thus, although the generative data and resulting HRTFs may not be symmetric, the model 533 and dimensionality reduction are said to be symmetric. Overall, this reflection process has the desirable result of reducing the complexity of the target data by a factor of two and doubling the sample size of the H matrix 527 and the ITD matrix 525. An important additional benefit of using this process is that the reflected reconstructed HRTFs may be more balanced. This is because the reflection process leads to symmetric behavior of any noise in the HRTF computation system 500 caused by overfitting and errors in the dimensionality reduction component 524 and the machine learning component 526.

[0060] FIG. 1 is a block diagram of an audio ecosystem 100. The audio ecosystem 100 includes one or more user devices 110 (two are shown: a user input device 110a and a user output device 110b) and one or more cloud devices 120 (two are shown: a personalization server 120a and a content server 120b). Multiple devices are shown for ease of illustration. A single user device may implement the functionality of the user input device 110a and the user output device 110b. Similarly, the functionality of the personalization server 120a and the content server 120b may be implemented by a single server or by multiple computers in a distributed cloud system. The devices in the audio ecosystem 100 may be connected by a wireless or wired network (not shown). The general operation of the audio ecosystem 100 is as follows.

[0061] A user input device 110a captures user generated data 130. The user input device 110a may be a cell phone with a camera. The generated data 130 may consist of structural images and / or demographic data, and may also include scaled images. Further details of the capture process and generated data 130 are provided below.

[0062] The personalization server 120a receives the generated data 130 from the user input device 110a, processes the generated data 130 to generate personalized HRTFs 132 for the user, and stores the personalized HRTFs 132. For example, the personalization server 120a may implement an estimation component 530, a dimension reconstruction component 532, and a phase reconstruction component 534 (see FIG. 5). Further details regarding the generation of the personalized HRTFs 132 are provided below. The personalization server 120a also provides the personalized HRTFs 132 to the user output device 110b.

[0063] The content server 120b provides content 134 to the user output device 110b. Generally, the content 134 includes audio content. The audio content may include audio objects, for example, according to the Dolby® Atmos™ system. The audio content may include multi-channel signals, such as stereo signals, converted into positioned sound objects. The content 134 may also include video content. For example, the content server 120b may be a multimedia server providing audio and video content, a game server providing game content, or the like. The content 134 may be continuously provided by the content server 120b, or the content server 120b may provide the content 132 to the user output device 110b for current storage and future output.

[0064] The user output device 110b receives the personalized HRTF 132 from the personalization server 120a, receives the content 134 from the content server 120b, and applies the personalized HRTF 132 to the content 134 to generate the audio output 136. Examples of the user output device 110b include a mobile phone (and associated earbuds), headphones, a headset, earbuds, hearables, etc. The user output device 110b may be the same device as the user input device 110a. For example, a mobile phone equipped with a camera may capture the generated data 130 (as the user input device 110a), receive the personalized HRTF 132 (as the user output device 110b), and be associated with a pair of earbuds that generate the audio output 136. The user output device 110b may be a different device from the user input device 110a. For example, a mobile phone with a camera may capture generated data 130 (as user input device 110a), and a headset may receive personalized HRTFs 132 (as user output device 110b) and generate audio output 136. User output device 110b may also be associated with other devices such as a computer, an audio / video receiver (AVR), a television, etc.

[0065] The audio ecosystem 100 is referred to as an "ecosystem" because the system adapts to whatever output device a user is currently using. For example, a user may be associated with a user identifier, and the user may log in to the audio ecosystem 100. The personalization server 120a may use the user identifier to associate a personalized HRTF 132 with the user. The content server 120b may use the user identifier to manage the user's subscriptions, preferences, etc. for content 134. The user output device 110b may use the user identifier to communicate to the personalization server 120a that the user output device 110b should receive the user's personalized HRTF 132. For example, when a user purchases a new headset (as the user output device 110b), the headset may use the user identifier to obtain the user's personalized HRTF 132 from the personalization server 120a.

[0066] 2 is a flowchart of a method 200 for generating a head-related transfer function (HRTF). Method 200 may be performed by one or more devices of audio ecosystem 100 (see FIG. 1), for example, by executing one or more computer programs.

[0067] At 202, an HRTF calculation system is generated. Generally, an HRTF calculation system corresponds to a relationship between anatomical measurements and HRTFs. The HRTF calculation system may be generated by personalization server 120a (see FIG. 1), for example, by implementing HRTF calculation system 500 (see FIG. 5). The generation of the HRTF calculation system includes substeps 204, 206, 208, and 210.

[0068] At 204, 3D scans of several training subjects are measured. Generally, the 3D scans correspond to a database of high-resolution scans of the training subjects, and the measurements correspond to measurements of anatomical features captured in the 3D scans. The 3D scans may correspond to mesh data 511 (see FIG. 5). The personalization server 120a may store the database of high-resolution scans.

[0069] At 206, several HRTFs are generated for the training subject by performing acoustic scattering calculations on the 3D scan measurements. The personalization server 120a may perform the acoustic scattering calculations to generate the HRTFs, for example, by implementing the numerical simulation component 520 (see FIG. 5).

[0070] At 208, generated data is collected for the training subject. Generally, the generated data corresponds to anthropometric measurements and demographic data for the training subject, where the anthropometric measurements are determined from the 3D scan data. For example, the generated data may correspond to one or more of demographic data 513, anthropometric data 521, and generated data matrix 523 (see FIG. 5). The anthropometric data 521 may be generated by mesh annotating component 522 based on mesh data 511 (see FIG. 5).

[0071] At 210, training is performed on an HRTF calculation system that converts the generated data into a plurality of HRTFs. Generally, a machine learning process is performed to generate a model for use in the HRTF calculation system. The generated data (see 208) is used by the model to estimate values ​​of the generated HRTFs (see 206). Training may include using linear regression with lasso regularization, as described in more detail below. Personalization server 120a may perform the training process, for example, by implementing machine learning component 526 (see FIG. 5).

[0072] At 212, a personalized HRTF is generated for the user using an HRTF calculation system. Personalization server 120a may generate the personalized HRTF, for example, by implementing HRTF calculation system 500 (see FIG. 5). Generating the personalized HRTF includes substeps 214 and 216.

[0073] At 214, user-generated data is collected. Typically, the generated data corresponds to anthropometric measurements and demographic data of a particular user (the anthropometric measurements are determined from the 2D image data) for generating a personalized HRTF. For example, the generated data may correspond to generated data 535 (see FIG. 5). For scaling purposes, a reference object may be captured in the image along with the user. A user input device 110a (see FIG. 1) may be used to collect the user-generated data. For example, the user input device 110a may be a mobile phone that includes a camera.

[0074] At 216, the user-generated data is input to an HRTF calculation system to obtain personalized HRTFs. The personalization server 120a can obtain personalized HRTFs by inputting the user-generated data (see 214) into the results of training the HRTF calculation system (see 210), for example, by implementing an estimation component 530, a dimension reconstruction component 532, and a phase reconstruction component 534 (see FIG. 5).

[0075] At 218, once the personalized HRTF is generated, it may be provided to a user output device and used when generating audio output. For example, user output device 110b (see FIG. 1) may receive personalized HRTF 132 from personalization server 120a, may receive audio signals in content 134 from content server 120b, and may generate audio output 136 by applying personalized HRTF 132 to the audio signals. The audio signals may include audio objects that include positional information, and the audio output may correspond to binaural audio output generated by rendering the audio objects using the personalized HRTF. For example, the audio objects may include Dolby® Atmos™ audio objects.

[0076] Further details of this process are provided below.

[0077] FIG. 3 is a block diagram of audio environment 300. Audio environment 300 is similar to audio environment 100 (see FIG. 1) and provides additional details. Like audio environment 100, audio environment 300 can generate personalized HRTFs using one or more devices, for example, by performing one or more steps of method 200 (see FIG. 2). Audio environment 300 includes input device 302, processing device 304, and output device 306. Compared to audio environment 100, details of audio environment 300 are described functionally. The functionality of the devices in audio environment 300 may be implemented, for example, by one or more processors executing one or more computer programs.

[0078] The input device 302 typically captures user input data (which is processed into structural images 313 and / or user-generated data such as demographic data 311). The input device 302 may also capture scaled images 315. The input device 302 may be a mobile phone equipped with a camera. The input device 302 includes a capture component 312 and a feedback and local processing component 314.

[0079] The capture component 312 generally captures demographic data 311 and a structural image 313 of the user's acoustic anatomy. The structural image 313 is then used (as described further below) to generate a set of anthropometric data 325. For ease of further processing, the capture of the structural image 313 may be performed against a static background.

[0080] One option for capturing a structural image 313 is as follows: The user places the input device 302 on a stable surface just below eye level and positions themselves so that their acoustic anatomy is visible in the capture frame. The input device 302 generates a tone or other indicator, and the user slowly rotates 360 degrees. The user may rotate in a standing or sitting position with their arms at their sides.

[0081] Another option for capturing structural images 313 is as follows: The user holds the input device 302 at arm's length, placing the user's acoustic anatomy in the video frame. Starting with the input device 302 facing the user's ear, the user sweeps their arm forward so that the video captures images from the user's ear to the front of the user's face. The user then repeats this process on the other side of the body.

[0082] Another option for capturing the structural image 313 is as follows: As in the embodiment described above, the user holds the input device 302 at arm's length, placing the user's acoustic anatomy in the video frame. However, in this embodiment, the user turns their head left and right as far as possible without discomfort. This allows the user's head and pinnae to be captured in the structural image.

[0083] The above options allow a user to capture their own structural image without the assistance of another person, however, an additional useful embodiment would be to have a second person walk around the user while they are standing motionless, pointing the camera of the input device 302 at the user's acoustic anatomy.

[0084] The extent, order, and manner in which the structural images are recorded is not important, as long as there are structural images from multiple azimuthal or horizontal angles relative to the face. In one embodiment, it is recommended that structural images be captured at intervals of 10 degrees or less and over a range of at least 90 degrees to the left and 90 degrees to the right of the user's face.

[0085] The capture component 312 may provide guidance to the user during the capture process. For example, the capture component 312 may output beeps or audio prompts to tilt the input device 302 up or down, to vertically shift the input device 302 so that it is perpendicular to the user's ear, to increase or decrease the speed of the sweep or rotation process, etc. The capture component 312 provides a structural image to the feedback and local processing component 314.

[0086] The feedback and local processing component 314 generally evaluates the structural images 313 captured by the capture component 312 and performs local processing on the capture. With regard to evaluation, the feedback and local processing component 314 may evaluate various criteria of the captured image, such as whether the user stayed in frame, whether the user was not rotating too quickly, etc.; if the criteria indicate failure, the feedback and local processing component 314 may return operation of the input device 302 to the capture component 312 to perform another capture. With regard to local processing, the feedback and local processing component 314 may subtract the background from each image and perform other image processing functions, such as blur / sharpness evaluation, contrast evaluation, and brightness evaluation, to ensure photographic quality. The feedback and local processing component 314 may also perform identification of key landmarks, such as the center of the face and the location of the ears in the video, to ensure that the final structural image 313 adequately describes the user's acoustic anatomy.

[0087] The captured video then includes structural images of the user's acoustic anatomy from multiple perspectives. The input device 302 then transmits the final structural images 313, and optional demographic data 311, to the processing device 304. The input device 302 may also transmit a scaled image 315 to the processing device 304.

[0088] The processing unit 304 generally processes the structural image 313 to generate anthropometric data 325 and generates personalized HRTFs based on the generated data, which may consist of the anthropometric data 325 and / or the demographic data 311. The processing unit 304 may be hosted by a cloud-based server. Alternatively, the input device 302 may implement one or more functions of the processing unit 304, in cases where it is desirable to implement cloud processing functions locally. The processing unit 304 includes a photogrammetry component 322, a context transformation component 324, a landmark detection component 326, an anthropometry component 328, and an HRTF generation component 330.

[0089] Photogrammetry component 322 receives the final version of structural image 313 from feedback and local processing component 314 and performs photogrammetry using techniques such as structure-from-motion (SfM) to generate camera transform 317 and structural image set 319. Generally, structural image set 319 corresponds to frames of structural image 313 where photogrammetry component 322 has been successfully positioned, and camera transform 317 corresponds to the three-dimensional position and orientation components for each image in structural image set 319. Photogrammetry component 322 provides structural image set 319 to context transformation component 324 and landmark detection component 326. Photogrammetry component 322 also provides camera transform 317 to context transformation component 324.

[0090] The context transformation component 324 uses the structural image set 319 to translate and rotate the camera transform 317 to generate the camera transform 321. The context transformation component 324 may also receive a scaled image 315 from the feedback and local processing component 314; the context transformation component 324 may use the scaled image 315 to scale the camera transform 317 when generating the camera transform 321.

[0091] The landmark detection component 326 receives and processes the structural image set 319 and the camera transform 321 to generate a 3D landmark set 323. Generally, the 3D landmark set 323 corresponds to anthropometric landmarks that the landmark detection component 326 has identified from the structural image set 319 and the camera transform 321. For example, these anthropometric landmarks may include detection of various landmarks on visible surfaces, such as the fossa of each pinna, the concha, the tragus, and the helix. Other anthropometric landmarks of the user's acoustic anatomy detected by the landmark detection component 326 may include the brows, chin, and shoulders, as well as measurements of the head and torso in the appropriate frame. The landmark detection component 326 provides the 3D landmark set 323 to an anthropometry component 328.

[0092] Anthropometry component 328 receives 3D landmark set 323 and generates anthropometry data 325. Generally, anthropometry data 325 corresponds to a set of geometrically measured distances and angles between individual landmarks in 3D landmark set 323. Anthropometry component 328 provides anthropometry data 325 to HRTF generation component 330.

[0093] The HRTF generation component 330 receives anthropometric data 325 and generates a personalized HRTF 327. The HRTF component 330 may also receive demographic data 311 and use it in generating the personalized HRTF 327. The personalized HRTF 327 may be in a spatially oriented acoustic format (SOFA) file format. The HRTF generation component 330 typically uses a previously determined HRTF calculation system (e.g., model 533 trained by the HRTF calculation system 500 of FIG. 5 ), discussed in more detail herein, as part of generating the personalized HRTF 327. The HRTF generation component 330 provides the personalized HRTF 327 to the output device 306.

[0094] The output device 306 typically receives the personalized HRTF 327 from the processing device 304 and applies the personalized HRTF 327 to the audio data to generate audio output 329. The output device 306 may be a mobile phone and associated speakers (e.g., a headset, earbuds, etc.). The output device 306 may be the same device as the input device 302. In cases where cloud processing functionality is desired to be implemented locally, the output device 306 may implement one or more functions of the processing device 304. The output device 306 includes a rendering component 340.

[0095] The rendering component 340 receives the personalized HRTF 327 from the HRTF generation component 330 and performs binaural rendering on the audio data using the personalized HRTF 327 to generate an audio output 329.

[0096] FIG. 4 is a block diagram of an anthropometry system 400. The anthropometry system 400 may be implemented by a component of the audio ecosystem 100 (e.g., the personalization server 120a in FIG. 1 ), a component of the audio ecosystem 300 (e.g., the processing unit 304 in FIG. 3 ), or the like. The anthropometry system 400 may implement one or more steps of the method 200 (see FIG. 2 ). The anthropometry system 400 may operate similarly to one or more components of the processing unit 304 (see FIG. 3 ), such as the photogrammetry component 322, the context transformation component 324, the landmark detection component 326, and the anthropometry component 328. The anthropometry system 400 includes a data extraction component 402, a photogrammetry component 404, a scale measurement component 406, a facial landmark detection component 408, a context transformation component 410, a cropping component 412, a landmark detection component 414, a 3D projection component 416, an angle and distance component 418, and a data summarization component 420. The functionality of the components of anthropometry system 400 may be implemented, for example, by one or more processors executing one or more computer programs.

[0097] The data extraction component 402 receives input data 401 and performs data extraction and selection to generate demographic data 403, structural images 405, and scale images 407. The input data 401 may be received from user input device 110a (see FIG. 1) or input device 302 (see FIG. 3), and may include, for example, image data captured by a cell phone camera (see 214 in FIG. 2). The data extraction component provides the demographic data 403 directly to the data compilation component 420, the structural images 405 to the photogrammetry component 404, and the scale images 407 to the scale measurement component 406.

[0098] The photogrammetry component 404 typically performs a photogrammetry process, such as structure from motion (SfM), on the structural image 405 to generate a camera transform 411 and an image set 409. The photogrammetry process receives the structural image 405 and generates a set of camera transforms 411 (e.g., camera viewpoint positions and viewpoint orientations) corresponding to each frame of the image set 409, which may be a subset of the structural image 405. The viewpoint orientation is often expressed in the form of either a quaternion or a rotation matrix, but for the purposes of this paper, the mathematical examples will be expressed in the form of a rotation matrix. The image set 409 is passed to a facial landmark detection component 408 and a cropping component 412. The camera transform 411 is passed to a context transformation component 410.

[0099] The photogrammetry component 404 may optionally perform image feature detection on the structural image 405 using constrained image feature searching before performing the SfM process. Constrained image feature searching may improve the results of the SfM process by overcoming user errors in the capture process.

[0100] The scale measurement component 406 uses the scale image 407 from the data extraction component 402 to generate information for subsequent use in scaling the camera transform 411. The scaling information, called the homologue measure 413, is generated in summary as follows: the scale image contains a visible scale reference, and the scale measurement component 406 uses that scale reference to measure scale homologues that are visible in the same frame of the scale image and in one or more frames of the structural image. The resulting scale homologue measure is passed to the context transformation component 410 as the homologue measure 413.

[0101] The facial landmark detection component 408 searches for visible facial landmarks within the frames of the image set 409 received from the photogrammetry component 404. The detected landmarks may include points on the user's nose as well as the location of the pupils, which may later be used as scale homologs visible in both the image set 409 and the scale image 407. The resulting facial landmarks 415 are passed to the context transformation component 410.

[0102] The context transformation component 410 receives a camera transform 411 from the photogrammetry component 404, homology measures 413 from the scale measurement component 406, and a set of facial landmarks 415 from the facial landmark detection component 408. The context transformation component 410 effectively transforms the camera transform 411 into a set of camera transforms 417 that are properly centered, oriented, and scaled to fit the context of the acoustic-anatomical structures captured in the structural images of the image set 409. In summary, the context transformation is achieved by scaling the positional information of the camera transform 411 using the facial landmarks 415 and homology measures 413, rotating the camera transform 411 in 3D space using the facial landmarks 415, and translating the positional information of the camera transform 411 to move the origin of the 3D space to the center of the user's head. The resulting camera transform 417 is passed to the cropping component 412 and the 3D projection component 416.

[0103] The cropping component 412 generally selects and crops a subset of frames from the image set 409 using a camera transform 417. After being properly centered, oriented, and scaled, the cropping component 412 uses the camera transform 417 to estimate which subset of images from the image set 409 contains structural images of particular characteristics of the user's acoustic anatomy. Additionally, the cropping component 412 may use the camera transform 417 to estimate which portion of each image contains structural images of particular characteristics. The cropping component 412 can thus be used to crop individual frames of the subset of the image set 409 to generate resulting image data, referred to as crops 419.

[0104] The landmark detection component 414 generally provides predicted locations of specified landmarks of the user's acoustic anatomy that are visible within the 2D image data of the crop 419. Thus, landmarks that are visible within a given image frame are labeled as a corresponding set of ordered 2D point locations. The cropping component 412 and landmark detection component 414 may be coordinated to implement a constrained recursive landmark search by cropping and detecting multiple different sets of landmarks that may be visible within different subsets of the image set 409. The landmark detection component 414 passes the resulting 2D coordinates 421 of the anatomical landmarks to the 3D projection component 416.

[0105] The 3D projection component 416 generally transforms the set of 2D coordinates 421 of each anatomical landmark into a single position in 3D space using a camera transform 417. The full set of 3D landmark positions is passed as 3D coordinates 423 to the angle and distance measurement component 418.

[0106] The angle and distance measurement component 418 uses a predetermined set of instructions to measure angles and distances between various points in 3D coordinates 423. These measurements may be achieved by applying simple Euclidean geometry. The resulting measures can effectively be used as anthropometric data 425 and are passed to the data summarization component 420.

[0107] The data aggregation component 420 typically combines the demographic data 403 with the anthropometric data 425 to form a complete set of generated data 427. These generated data 427 may then be used in the HRTF calculation system (e.g., generated data 535 of FIG. 5 ) as described above to derive personalized HRTFs for the user.

[0108] Further details and examples of the anthropometry system 400 follow.

[0109] Although all frames of the structural images 405 can be used in the photogrammetry component 404, the system may achieve better performance by reducing them for computational efficiency. To select the best frames, the data extraction component 402 may evaluate frame content and a sharpness metric. An example of frame content selection may be searching for similarities in consecutive images to avoid redundancy. The sharpness metric may be selected from several sharpness metrics; one example is using a radial collapse of a 2D spatial frequency power spectrum into a 1D power spectrum. The data extraction component 402 provides the selected set of structural images 405 to the photogrammetry component 404. Because the photogrammetry component 404 may implement a time-intensive process, the data extraction component 402 may pass the structural images 405 before scaling images 407 from the input data 401 or collecting demographic data 403. If the system is capable of parallel processing, this order of operations may be a desirable optimization.

[0110] The SfM process inputs a series of frames (e.g., structural images 405, which need not be ordered sequentially) and outputs, for each input frame, an estimate (e.g., a 3D point cloud) of the assumed rigid object imaged in the capture process, as well as a calculated viewpoint position (x, y, z) and rotation matrix. These viewpoint positions and rotation matrices are referred to herein as camera transforms (e.g., camera transform 411) because they describe where each camera is located and oriented with respect to world space. World space is a general term for a 3D coordinate system that includes all camera viewpoints and structural images. Note that the 3D point cloud itself need not be used further in this application; i.e., it is not necessary to generate a 3D mesh object for any part of determining a user's anthropometric measurements. It is not uncommon for the SfM process to fail to derive a camera transform for one or more of the images in the optimal set of structural images 405. For this reason, the failed frames may be omitted from subsequent processing. Because autofocus camera applications are not necessarily optimized for the capture conditions of this system, it may be useful for the image capture component (see Figure 3) to fix the camera focal length during the image capture process.

[0111] The SfM process considers the pinnae and head to be essentially rigid objects. In determining the shape and location of these objects, the SfM process first implements one or more known image feature detection algorithms, such as SIFT (shift-invariant feature transform), HA (Hessian Affine feature point detector), or HOG (histogram of oriented gradients). The resulting image features differ from face and anatomical landmarks in that they are not separately pre-trained and are not specific to the context of acoustic-anatomical features. Because other parts of the user's body may not remain rigid throughout the acquisition process, it is useful to program the photogrammetry process to infer geometric configuration using only image features detected in the head region. This selection of image features can be achieved by running the facial landmark detection component 408 before running the photogrammetry component 404. For example, known techniques for face detection can be used to estimate landmarks defined by the bounding box of the head or face in each image. The photogrammetry component 404 can then apply a mask to each image to include only the image features detected by computer vision within the corresponding bounding box. In another embodiment, facial landmarks may be used directly by the photogrammetry component 404 instead of, or in addition to, the detected image features. Constraining the range of image features used in the photogrammetry process can also make the system more efficient and optimized. If the facial landmark detection component 408 runs before the photogrammetry component 404, the facial landmark detection component 408 can receive the structural image 405 directly from the data extraction component 402 instead of the image set 409 and pass the facial landmarks 415 to both the photogrammetry component 404 and the context transformation component 410.

[0112] The photogrammetry component 404 may also include a focal length compensation component. The focal length of a camera can exaggerate the depth of the ear shape (e.g., due to barrel distortion caused by a short focal length) or reduce such depth (e.g., due to pincushion distortion caused by a long focal length). In many smartphone cameras, such focal length distortion is often a type of barrel distortion that the focal length compensation component can detect depending on the focal length and the distance to the captured image. This focal length compensation process may be applied using known methods to remove distortions in the image set 409. This compensation may be particularly useful when processing structural images from handheld capture methods.

[0113] A more detailed description of the scale measurement component 406 follows. The term scale reference refers to an imaged object of known size, such as a banknote or an identification card. The term scale homologue refers to an imaged object or distance that is common to two or more images and can be used to infer the relative size or scale of the object in each image. This scale homologue is also shared with the structural image and can therefore be used to scale the structural image and any measurements made thereon. A variety of scale homologues and scale references can be used, and the following embodiment is an example.

[0114] The following is an example embodiment of the scale measurement component 406. A user may capture an image of the user holding a card having a known size (e.g., 85.60 mm x 53.98 mm) to their face (e.g., in front of their mouth, perpendicular to the front of their face). This image may be captured in a manner similar to the capture of the structural image 405 (e.g., before or after the capture of the structural image 405, so that the card does not otherwise interfere with the capture process) and from a position perpendicular to the front of their face. The card may then be used as a scale reference to measure the physical interpupillary distance between the user's pupils in millimeters, which may later be used by the context transformation component 410 as a scale analog to apply an absolute scale to the structural image. This is possible because the structural image capture includes one or more images in which the user's pupils are visible. The scale measurement component 406 may implement one or more neural networks to detect and measure the scale reference and scale analog, and may apply computer vision processes to refine the measurement.

[0115] In one embodiment, the face detection algorithm used by the facial landmark detection component 408 may also be used to locate the pixel coordinates of the user's pupils in the scale image. The scale measurement component 406 may use a pre-trained neural network and / or computer vision techniques to define the boundaries and / or landmarks of the scale reference. In this exemplary embodiment, the scale measurement component 406 may use a pre-trained neural network to estimate the corners of the card. Coarse lines may then be fitted to each pair of points describing the corners of the card detected by the neural network. The pixel distance between these coarse lines may be divided by the known dimension of the scale reference in millimeters to derive the number of pixels per millimeter of the image at the card's distance.

[0116] For accuracy, the scale measurement component 406 may perform the following computer vision techniques to fine-tune the card's measurements. First, a Canny edge detection algorithm may be applied to a normalized version of the scale image. Then, a Hough transform may be used to estimate fine lines in the scale image. The area between each coarse line and each fine line may be calculated. A threshold (e.g., the image dimensions multiplied by 10 pixels) may then be used to select only those fine lines with small areas that deviate from the coarse neural network prediction. Finally, the median of the selected fine lines may be selected as the final boundary of the card and used as described above to derive the number of pixels per millimeter for the image. Because the card and the user's pupils are at similar distances from the camera, the distance between the user's pupils in pixels may be divided by the number of pixels per millimeter calculation to measure the user's interpupillary distance in actual millimeters. Thus, the scale measurement component 406 passes this interpupillary distance to the context transformation component 410 as the homology measure 413.

[0117] The above-described embodiments of the scaling technique have been observed to be relatively accurate and accessible to a wide range of input devices, including standard cameras. The following embodiments are presented as alternatives for cases where additional sensors are available to the input device and can be used to infer scale information without the need for a scale reference. A second process for measuring scale analogs is a multimodal approach that utilizes not only a camera but also a microphone and a speaker (e.g., a pair of earbuds), all of which may be components of a capture device (e.g., user input device 110a in FIG. 1, such as a mobile phone). Since sound is known to travel reliably at 343 m / s, a tone played on an earbud next to a user's face can be expected to be recorded by the phone with a delay precisely the amount required for the sound to travel to the phone's microphone. This delay can then be multiplied by the speed of sound to determine the distance from the phone to the face. An image of the face can be taken at the same time the sound is emitted, and this image includes scale analogs such as the user's eyes. The scale measurement component may use some simple trigonometry to calculate the distance between the user's pupils, or the distance between another pair of reference points, in metric units: d=delay*sos w_mm=2*tan(aov / 2)*d ipd_mm=w_mm*ipd_pix / w_pix (In the above formula, ipd_pix is ​​the pixel distance between the user's pupils, ipd_mm is the millimeter distance between the user's pupils, sos is the speed of sound in milliseconds per millimeter, delay is the delay in milliseconds between the sound being played and recorded, w_mm is the horizontal dimension of the imaging plane at distance d in millimeters, w_pix is ​​the horizontal dimension of the image in pixels, and aov is the horizontal angle of view of the imaging camera.)

[0118] Wireless earbuds, or perhaps over-the-ear headphones, may be used, and the process may begin simply by turning them on. The volume of the recorded signal may be used as an indicator of proximity between the earbuds and the microphone. The earbuds may generally be placed in close proximity to a scale homologue. The sound signal may be within, below, or above the threshold of human hearing and may be a chirp, a frequency sweep, a dolphin call, or other such attractive or soothing sounds. The sound may be extremely short (e.g., less than one second), allowing many measurements to be taken (e.g., over many seconds) for purposes of redundancy, averaging, and statistical analysis.

[0119] Another option for establishing the scale of the structural image, which may be used in an alternative capture process that physically sweeps a camera around the head, is to use inertial measurement unit (IMU) data from a user input device (e.g., 110a in FIG. 1 ) to establish absolute distances between camera positions (e.g., accelerometers, gyroscopes with acceptable tolerances, etc.). Another option for establishing the scale of the structural image, which may be used in any of the capture process embodiments, is to use image range imaging, which may be made available by various forms of input devices. For example, some input devices, such as modern mobile phones, are equipped with range cameras that utilize techniques such as structured light, split pixel, or interferometry. When any of these techniques are used in combination with a standard camera, an estimate can be derived via known methods for the depth of a given pixel. Thus, the distance between the scale analog and the camera may be directly estimated using these techniques, and the subsequent process for measuring the scale analog may be performed as described above.

[0120] The facial landmark detection component 408 may perform face detection as follows: The facial landmark detection component 408 may extract landmarks from clear frames of the image set 409 using a histogram of orientation gradients. As an example, the facial landmark detection component 408 may implement the process described in Non-Patent Document 1. The facial landmark detection component 408 may implement a support vector machine (SVM) with a sliding window approach for classification of extracted landmarks. The facial landmark detection component 408 may use non-max suppression to reject multiple detections. The facial landmark detection component 408 may use a model pre-trained on a large number of faces (e.g., 3000 faces). [Non-Patent Document 1] Navneet Dalal and Bill Triggs, Histograms of Oriented Gradients for Human Detection, International Conference on Computer Vision & Pattern Recognition (CVPR '05), June 2005, San Diego, United States, pp.886-893

[0121] The facial landmark detection component 408 may perform 2D coordinate detection using an ensemble of regression trees. As an example, the facial landmark detection component 408 may implement the process described in Non-Patent Document 2. The facial landmark detection component 408 may identify several facial landmarks (e.g., five landmarks, such as the inner and outer edges of each eye and the ear columella). The facial landmark detection component 408 may identify facial landmarks using a model pre-trained on a large number of faces (e.g., 7,198 faces). The landmarks may include several points corresponding to various facial features, such as points defining the bounding box or boundary of the face (outside the cheeks, chin, etc.), eyebrows, pupils, nose (columella, nostrils, etc.), mouth, etc. [Non-patent document 2] Vahid Kazemi and Josephine Sullivan, One Millisecond Face Alignment with an Ensemble of Regression Trees, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp.1867-1874

[0122] In another embodiment, a convolutional neural network may be used for 2D facial landmark detection. The facial landmark detection component 408 may use a neural network model trained on a database of annotated faces to identify several facial landmarks (e.g., five or more). The landmarks may include the 2D coordinates of various facial landmarks.

[0123] The context transformation component 410 normalizes the camera transform 411 with respect to the user's head using a series of coordinate and orientation transformations. The context of the transformations is similar to the positioning and orientation of the human head that is measured or calculated in generating the HRTFs. For example, in one embodiment, the HRTFs used in the target data (see training environment 502 in FIG. 5) are calculated or measured with an x-axis running between the user's ear canals and a y-axis passing through the user's nose.

[0124] The context transformation component 410 first uses a least squares algorithm to find the plane that best fits the position data of the camera transformation 411. The context transformation component 410 can then rotate the best-fit plane (along with the camera transformation 411) onto the xy plane (so that the z-axis is accurately described as "up" and "down").

[0125] To complete the absolute scaling and other subsequent transformations, the context transformation component 410 can estimate the location in world space of several key landmarks in the facial landmarks 415, which may include the pupils and nose. The system can use a process similar to the 3D projection component 416, described in more detail below, to locate these landmarks using the full set of 2D facial landmarks 415. Once the 3D locations of the facial landmarks have been determined, the following transformations may be performed:

[0126] First, the origin of world space may be centered by subtracting the arithmetic mean of the eye's 3D coordinates from the position information in the camera transform 411 and the 3D positions of the facial landmarks 415. Next, to apply an absolute scale to world space, the context transformation component 410 may multiply the position information in the camera transform 411 and the 3D positions of the facial landmarks 415 by a scaling ratio. This scaling ratio may be derived by dividing the homology measure 413 by an estimate of the interpupillary distance calculated using the 3D positions of the left and right sides of each eye, or the 3D positions of the pupils themselves. This scaling process allows the HRTF synthesis system 400 to use real-world (physical) distances, as these relate to, among other things, sound waves and their subsequent reflection, diffraction, absorption, and resonance behavior.

[0127] At this point, the scaled and centered camera transform needs to be oriented around the vertical axis of world space, sometimes referred to as the z-axis. In the field of photography, a photograph in which the subject's face and nose are pointed directly at the camera is commonly referred to as a frontal photograph. It may be useful to rotate the camera transform 411 around the z-axis of world space so that a frontal frame of the image set 409 corresponds to a camera transform positioned at 0 degrees relative to one of the other two axes of world space, e.g., the y-axis. The context transformation component 410 may implement the following process to identify a frontal frame of the image set 409: The context transformation component 410 minimizes point asymmetry of the facial landmarks 415 to find a "frontal" reference frame. For example, a frontal frame may be defined as a frame in which each pair of pupils is closest to being equidistant from the nose. In mathematical terms, the context transformation component 410 may calculate the asymmetry according to the asymmetry function |LR| / F. where L is the centroid of the left landmark, R is the centroid of the right landmark, and F is the centroid of all facial landmarks 415. The frontal frame is then the frame in which the asymmetry is minimized. Once the frontal frame is selected, the camera transform 411 and the 3D positions of the facial landmarks 415 may be rotated around the z-axis so that the frontal frame of the image set 409 corresponds to a camera transform positioned at 0 degrees relative to the y-axis of world space.

[0128] Finally, the context transformation component 410 may translate the camera transform 411 inside the head so that the origin of world space corresponds to an estimated point between the ears rather than the center of the face. This can be done simply by translating the camera transform along the y-axis by the average value of the orthogonal distance between the face and the interaural axis. This value may be calculated using human anthropometric data in a mesh database (see mesh data 511 in FIG. 5). For example, the authors found that the average orthogonal distance between the eyebrows and ears in their dataset was 106.3 millimeters. In one embodiment, to account for the angular pitch of the head, the camera transform may also be rotated around the interaural axis, i.e., the axis between the ear canals. This rotation may be done so that the 3D position of the nose lies along the y-axis.

[0129] As a result of the above process, the context transformation component 410 generates a centered, leveled, and scaled camera transform 417. These camera transforms 417 may be used by the cropping component 412 to estimate images in which various points describing the user's acoustic anatomy are visible. For example, images in which the camera transform 417 is positioned between 30 and 100 degrees clockwise around the z-axis of world space are likely to contain a structural image of the user's right ear, given that the tip of the user's nose is aligned with the y-axis of world space at 0 degrees. Once these images are selected from the image set 409, they may be cropped to include only the portion of the image estimated to contain a structural image of the landmark(s) of interest.

[0130] To crop each image to an appropriate portion, the cropping component 412 may calculate the approximate location and size of a 3D point cloud containing the anatomical landmarks of interest. For example, the average distance between ear canals is approximately 160 millimeters, and the camera transform 411 is centered on the estimated bisector of this line. Thus, the location of each ear point cloud may be expected to be approximately 80 millimeters in each direction along the x-axis of world space, given that the tip of the user's nose is aligned with the y-axis of world space. In this example, the size of each 3D point cloud is likely to be approximately 65 mm in diameter, which describes the average length of the ear.

[0131] Here, cropping can be achieved using the following techniques: The orientation information of each camera's transform describes how the camera's three axes linearly relate to the three axes of world space, with the camera's principal axis typically considered to describe a vector passing through the center of the image frame from the camera's position. The landmark line is a line in world space between the camera's position and the estimated location of the landmark point cloud. The camera transformation's rotation matrix can be used to directly represent the landmark line in camera space, or the specific camera's 3D coordinate system. The camera's field of view is an intrinsic parameter that can be calculated using the camera's 35-millimeter equivalent focal length or focal length and sensor size, either of which can be derived from a camera lookup table, from EXIF ​​data encoded with each image, or from the input device itself at the time of capture. Landmark lines may be projected onto the image using the camera's field of view and the image's pixel dimensions. For example, the horizontal pixel distance ("x") between the image's center and the landmark in the image plane can be approximated as follows: d_pix=(w_pix / 2) / tan(aov / 2) x_pix=d_pix*tan(fax / 2) (In the above formula, d_pix is ​​the distance in pixels between the camera and the image plane, w_pix is ​​the horizontal dimension of the image, aov is the horizontal field of view of the capturing camera, and fax is the horizontal angular component of the landmark line.)

[0132] Once the pixel location of the landmark point cloud center is approximated, the appropriate width and height of the crop can be calculated using a similar method. For example, because almost all ears are less than 100 millimeters along their longest diagonal, a 100-millimeter crop is reasonable for locating ear landmarks. As described in the steps below, many neural networks use square images as input, which means that the final cropped image has the same height and width. Therefore, for this ear example, a crop of ±50 millimeters vertically and horizontally from the landmark center may be appropriate. The distance from the camera to the image plane in world space units may be calculated by calculating the magnitude, in world space, of the orthogonal projection of the landmark line onto the camera's major axis. Because this distance was calculated in pixels above, the ratio of pixels per millimeter can be calculated and applied to the 50-millimeter cropping dimension to determine the crop boundary in pixels. Once the cropping component 412 completes this cropping process, the image can be rescaled and used in the landmark detection component 414, as described below.

[0133] To identify the 2D coordinates 421, the landmark detection component 414 may use a neural network. For example, the personalization server 120a (see FIG. 1) or the user input device 110a (see FIG. 1) may use a neural network as part of implementing the neural network component 326 (see FIG. 3), the face detection component (see 410 in FIG. 4), or the facial landmark detection component (see 408 in FIG. 4). According to one embodiment, the system implements a convolutional neural network (CNN) to perform anatomical landmark labeling. The system may implement the CNN by running the Keras neural network library (written in Python) on top of the TensorFlow machine learning software library using the MobileNets architecture pre-trained on the ImageNet database. The MobileNets architecture may use an alpha multiplier of 0.25 and a resolution multiplier of 1 and may be trained on a database of structural images to detect facial landmarks for constructing the anthropometric data 425.

[0134] For example, an image (e.g., one of the structural image frames 313) may be downsampled to a smaller resolution image and represented as a tensor (multidimensional data array) having size 224 x 224 x 3. The image may be processed by a MobileNets architecture trained to detect a given set of landmarks, resulting in a tensor having size 1 x (2 * n) that identifies the x and y coordinates of n landmarks. For example, this process may be used to generate x and y coordinates for 18 ear landmarks and 9 torso landmarks. In different embodiments, the Inception V3 architecture or a different convolutional neural network architecture may be used. The cropping component 412 and landmark detection component 414 may be used iteratively or simultaneously for different images and / or different sets of landmarks.

[0135] To estimate singular values ​​for each landmark's coordinates, the 2D coordinates 421 from the landmark detection component 414 are passed to the 3D projection component 416 along with the camera transform 417. The 3D projection component 416 may project the 2D coordinates 421 from each camera into world space and then perform a least-squares calculation to approximate the intersection of the set of projected rays for each landmark. For example, the description of the cropping component 412 above details how landmark lines in world space can be projected onto the image plane through a series of known photogrammetry methods. This process is reversible, and each landmark in the image plane may be represented as a landmark line in world space. There are several known methods for estimating the intersection of multiple lines in 3D space, such as least-squares solutions. At the conclusion of the 3D projection component 416, multiple landmark positions have been calculated in world space, which may be collected from different field-of-view ranges and / or by using different neural networks. This set of 3D coordinates 423 may also include 3D coordinates calculated through other methods, such as the above calculation of each pupil's position.

[0136] In one embodiment, it may be useful to repeat the processes of the context transformation component 410, the cropping component 412, the landmark detection component 414, and the 3D projection component 416 as part of an iterative refinement process. For example, the initial iteration of the context transformation component 410 can be considered a “coarse” positioning and orientation of the camera transform 417, and the initial iteration of the cropping component 412 can be considered a “coarse” selection and cropping from the image set 409. The 3D coordinates 423 may include an estimated location of each ear, which may be used to iterate the context transformation component 410 in “fine” iterations. In a preferred embodiment, the “coarse” crop may be significantly larger than the estimated size of the landmark point cloud to account for errors in landmark line estimation. As part of the refinement process, the cropping component 412 may be iterated with a tighter, smaller crop of the image after the fine iteration of the context transformation component 410. This refinement process can be repeated as many times as desired, although in one embodiment, at least one refinement iteration is recommended. This is recommended because the authors have found that the accuracy of the landmark detection component 414 is higher when the cropping component 412 uses a tighter crop, but the crop must include a structural image of the entire landmark set and therefore must be set using accurate estimates of the landmark lines.

[0137] The actual anthropometric data used by the system to generate personalized HRTFs are scalar values ​​representing lengths of anatomical features and angles between anatomical features. These calculations across subsets of 3D coordinates 423 are specified and performed by angle and distance measurement component 418. For example, angle and distance measurement component 418 may specify the calculation of "shoulder width" as the Euclidean distance between a "left shoulder" coordinate and a "right shoulder" coordinate belonging to 3D coordinate set 423. As another example, angle and distance measurement component 418 may specify the calculation of "pinna flare angle" as the angular representation of the horizontal component of the vector between a "concha front" coordinate and a "superior helix" coordinate. A known set of anthropometric measures for use in HRTF calculations has been proposed and may be collected during this process. For example, the anthropometric measurements determined from the 3D coordinates 423 may include, for each of the user's pinna, the pinna flare angle, pinna rotation angle, pinna cleft angle, pinna offset back, pinna offset down, pinna height, first pinna width, second pinna width, first intertragus width, second intertragus width, fossa height, concha width, concha height, and concha navicular height. At this point, the data aggregation component 420 can assemble the resulting anthropometric data 425 and the aforementioned demographic data 403 to form the generation data 427 necessary to generate personalized HRTFs.

[0138] The data compilation component 420 may perform what is called graceful degradation when compiling the generated data 427. Graceful degradation may be used when one or more predictors in the demographic data 403 are not provided or when identification of one or more predictors in the generated data 427 fails or is inconclusive. In such cases, the data compilation component 420 may generate estimates of the missing predictors based on other known predictors in the generated data 427 and then use the estimated predictors as part of generating personalized HRTFs. For example, if the system cannot determine a measurement for shoulder width, the system may use demographic data (e.g., age, gender, weight, height, etc.) to generate an estimate for shoulder width. As another example, the data compilation component 420 may use calculations of some pinna features made with a high confidence metric (e.g., low error in a least-squares solution) to estimate values ​​for other pinna features calculated with a lower confidence metric. Estimating subsets of the generated data 427 using other subsets can be accomplished using predetermined relationships. For example, as part of the training environment (see 502 in FIG. 5 ), the system may perform linear regressions between various sets of generated data using anthropometric data from a training database of high-resolution mesh data. As another example, the published U.S. Army Soldier Anthropometric Survey (ANSUR 2 or ANSUR II) may be included as a predictor in the generated data, containing certain characteristics that can be used in the linear regression method described above. In essence, the data summarization component 420 avoids the problem of missing data by estimating missing values ​​from available information in the demographic data 403 and anthropometric data 425. Using the complete sets of generated data 427 avoids the need to account for missing data in the HRTF calculation system.

[0139] Implementation details The embodiments may be implemented in hardware, in an executable module stored on a computer-readable medium, or in a combination of both (e.g., a programmable logic array). Unless otherwise specified, the steps performed by the embodiments need not inherently relate to any particular computer or other apparatus, although in certain embodiments they may. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Thus, the embodiments may be implemented in one or more computer programs running on one or more programmable computer systems, each having at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.

[0140] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., solid-state memory or media, or magnetic or optical media) to configure and operate a computer when the storage medium or device is read by a computer system to perform the procedures described herein. The system of the present invention may also be considered to be implemented as a computer-readable storage medium configured with a computer program, the storage medium so configured causing a computer system to operate in a specific, predefined manner to perform the functions described herein. (Software per se and intangible or transitory signals are excluded, to the extent that they are non-patentable subject matter.)

[0141] The above description illustrates various embodiments of the present disclosure, along with examples of how various aspects of the disclosure may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of the present disclosure, as defined by the claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be used without departing from the spirit and scope of the disclosure, as defined by the claims.

[0142] Several aspects will be described. [Aspect 1] 1. A method for generating a head-related transfer function (HRTF), the method comprising: Generating an HRTF calculation system, generating the HRTF calculation system comprising: Take multiple 3D scans of multiple training subjects generating a plurality of HRTFs for the plurality of training subjects by performing an acoustic scattering calculation on the plurality of 3D scans; collecting generation data for the plurality of training subjects; performing training of the HRTF calculation system to convert the generated data into the plurality of HRTFs; generating a personalized HRTF for a user using the HRTF calculation system, the generating the personalized HRTF comprising: Collect user-generated data, and inputting the generated data of the user into the HRTF calculation system to obtain a personalized HRTF. method. [Aspect 2] 2. The method of embodiment 1, wherein said performing training comprises using linear regression with lasso regularization. Aspect 3 3. The method of any one of aspects 1 to 2, wherein the generated data of the user includes at least one of anthropometric measurements and demographic data. Aspect 4 The anthropometric measurements are: Collect multiple images of the user, by determining anthropometric measurements using the plurality of images; The method of embodiment 3. Aspect 5 5. The method of embodiment 4, wherein determining the anthropometric measurements using the plurality of images is performed using a convolutional neural network. Aspect 6 and scaling the anthropometric measurements of the user using a reference object in at least one image of the plurality of images of the user. The method according to embodiment 4. Aspect 7 generating an audio output by applying the personalized HRTF to an audio signal. 7. The method of any one of embodiments 1 to 6. Aspect 8 storing, by a server device, the personalized HRTF; transmitting, by the server device, the personalized HRTF to a user device; the user device generates an audio output by applying the personalized HRTF to an audio signal. 8. The method of any one of embodiments 1 to 7. Aspect 9 generating, by a user device, an audio output by applying the personalized HRTF to an audio signal, wherein the user device includes one of a headset, a pair of earbuds, and a pair of hearables. 9. The method of any one of embodiments 1 to 8. Aspect 10 the audio signal comprises a plurality of audio objects, each of which includes position information; The method further comprises: generating a binaural audio output by applying the personalized HRTFs to the plurality of audio objects. 10. The method of any one of embodiments 1 to 9. Aspect 11 A non-transitory computer-readable medium storing a computer program that, when executed by a processor, controls an apparatus to perform processes including the method of any one of aspects 1 to 10. Aspect 12 1. An apparatus for generating a head-related transfer function (HRTF), the apparatus comprising: At least one processor; and having at least one memory, The at least one processor is configured to control the apparatus to generate an HRTF calculation system, wherein generating the HRTF calculation system includes: Take multiple 3D scans of multiple training subjects generating a plurality of HRTFs for the plurality of training subjects by performing an acoustic scattering calculation on the plurality of 3D scans; collecting generation data for the plurality of training subjects; performing training of the HRTF calculation system to convert the generated data into the plurality of HRTFs; The at least one processor is configured to control the device to generate a personalized HRTF for a user using the HRTF calculation system, wherein generating the personalized HRTF includes: Collect user-generated data, inputting the generated data of the user into the HRTF calculation system to obtain a personalized HRTF; Device. Aspect 13 The generated data of a user includes at least one of anthropometric measurements and demographic data, and the device further comprises: a user input device configured to collect a plurality of images of a user and determine anthropometric measurements of the user using the plurality of images of the user; the anthropometric measurements of the user are scaled using a reference object in at least one image of the plurality of images of the user; 13. The apparatus of embodiment 12. Aspect 14 a user output device configured to generate an audio output by applying the personalized HRTF to an audio signal. 14. The apparatus of embodiment 12 or 13. Aspect 15 a server device configured to generate the HRTF calculation system, generate the personalized HRTF, store the personalized HRTF, and transmit the personalized HRTF to a user device; the user device is configured to generate an audio output by applying the personalized HRTF to an audio signal. 15. The apparatus of any one of aspects 12 to 14. Aspect 16 a user device configured to generate an audio output by applying the personalized HRTF to an audio signal; the user device includes one of a headset, a pair of earbuds, and a pair of hearables; 16. The apparatus of any one of aspects 12 to 15. Aspect 17 17. The device of any one of aspects 12 to 16, wherein the audio signal includes a plurality of audio objects including position information, and the at least one processor is configured to control the device to generate binaural audio output by applying the personalized HRTF to the plurality of audio objects. Aspect 18 a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a photogrammetry component, a context transformation component, a landmark detection component, and an anthropometry component; the photogrammetry component is configured to receive a plurality of structural images of a user and generate a plurality of camera transformations and structural image sets using structure-from-motion techniques; the context transformation component is configured to receive the plurality of camera transformations and the structural image set and generate a plurality of transformed camera transformations by translating and rotating the plurality of camera transformations with the structural image set; the landmark detection component is configured to receive the structural image set and the transformed plurality of camera transformations and generate a set of 3D landmarks corresponding to anthropometric features of the user identified using the structural image set and the transformed plurality of camera transformations; the anthropometry component is configured to receive the set of 3D landmarks and generate anthropometry data from the set of 3D landmarks, the anthropometry data corresponding to a set of distances and angles measured between individual landmarks of the set of 3D landmarks; the server device is configured to generate the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system. 18. The apparatus of any one of embodiments 12 to 17. Aspect 19 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a scale measurement component. the scale measurement component is configured to receive a scale image including an image of a scale reference and generate a homology measure; 19. The apparatus of any one of aspects 12-18, wherein the server device is configured to scale a structural image of a user using the homology measure. Aspect 20 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a landmark detection component, a 3D projection component, and an angle and distance measurement component; the landmark detection component is configured to receive a cropped set of images of a user's anthropometric landmarks and generate a set of 2D coordinates of the set of anthropometric landmarks of the user from the cropped set of images; the 3D projection component is configured to receive the set of 2D coordinates and a plurality of camera transformations and use the camera transformations to generate a set of 3D coordinates corresponding to a set of 2D components of each anthropometric landmark in 3D space; the angle and distance measurement component is configured to receive the set of 3D coordinates and generate anthropometric data from the set of 3D coordinates, the anthropometric data corresponding to angles and distances of the anthropometric landmarks in the set of 3D coordinates; the server device is configured to generate the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system. 20. The apparatus of any one of aspects 12 to 19. Aspect 21 21. The device of any one of aspects 12 to 20, wherein the HRTF calculation system may be configured to train a model corresponding to one of a left ear HRTF and a right ear HRTF, and the personalized HRTF is generated by using the model to generate one of a left ear personalized HRTF and a right ear personalized HRTF, and using a reflection of the model to generate the other of the left ear personalized HRTF and the right ear personalized HRTF. Aspect 22 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a data aggregation component. the data aggregation component is configured to implement a graceful degradation of the generated data by filling in missing portions of the generated data using estimates determined from known portions of the generated data. 22. The apparatus of any one of aspects 12 to 21. Aspect 23 a server device configured to generate the HRTF calculation system, the server device executing a dimensionality reduction component; the dimensionality reduction component is configured to reduce the computational complexity of performing training of the HRTF computation system by performing a principal component analysis on the plurality of HRTFs for the plurality of training subjects. 23. The apparatus of any one of aspects 12 to 22. Aspect 24 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a photogrammetry component. the photogrammetry component is configured to receive a plurality of structural images of a user, perform a constrained image feature search using a facial landmark detection process on the plurality of structural images, and generate a plurality of camera transforms and a set of structural images using a structure from motion technique and results of the constrained image feature search; 24. The apparatus of any one of aspects 12 to 23. Aspect 25 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a context transformation component. the context transformation component is configured to receive a first plurality of camera transformations, a plurality of facial landmarks, and a scale measure, translate and rotate the plurality of camera transformations using the plurality of facial landmarks to generate a second plurality of camera transformations, and scale the second plurality of camera transformations using the scale measure; 25. The apparatus of any one of embodiments 12 to 24. Aspect 26 and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a scale measurement component. the scale measurement component is configured to receive range imaging information and generate a homology measure using the range imaging information; the server device is configured to scale the structural image of the user using the homology measure; 26. The apparatus of any one of embodiments 12 to 25. Aspect 27 a user input device associated with a speaker and microphone; and a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a scale measurement component; the scale measurement component is configured to receive arrival time information from the user input device and generate a homology measure using the arrival time information, the arrival time information relating to sound output by the speaker at a first location and received by the microphone at a second location, the first location relating to a user and the second location relating to the user input device; the server device is configured to scale the structural image of the user using the homology measure; 27. The apparatus of any one of embodiments 12 to 26. Aspect 28 a server device configured to generate the personalized HRTF for a user using the HRTF calculation system, the server device executing a cropping component and a landmark detection component; the cropping component and the landmark detection component are coordinated to perform a constrained recursive landmark search by cropping and detecting a plurality of different sets of landmarks; 28. The apparatus of any one of embodiments 12 to 27. Aspect 29 1. A method for generating a personalized head-related transfer function (HRTF), the method comprising: Receive multiple images of the user; processing the plurality of images to generate anthropometric data of a user; inputting the anthropometric data into an HRTF calculation system to obtain personalized HRTFs. method. Aspect 30 capturing video data of a user, the video data including a plurality of views of the user's head; and processing the video data to extract the plurality of images. 30. The method of embodiment 29. Aspect 31 capturing the video data includes capturing an image of an object having a known size; processing the plurality of images to generate the anthropometric data includes converting the anthropometric data from pixel measurements to absolute distance measurements using the known sizes; 31. The method of embodiment 29 or 30. Aspect 32 a device at a given distance captures said video data of a user, and the method further comprises: determining the given distance by measuring the time delay between emitting sound from headphones positioned proximate to the user and receiving the sound at a microphone on the device; processing the plurality of images to generate the anthropometric data includes converting the anthropometric data from pixel measurements to absolute distance measurements using the given distance; 32. The method of any one of embodiments 29 to 31. Aspect 33 Processing the plurality of images to generate the anthropometric data includes: converting the plurality of images into a three-dimensional point cloud model; selecting key image frames from the plurality of images using the three-dimensional point cloud model; generating the anthropometric data from the key image frames; 33. The method of any one of embodiments 29 to 32. Aspect 34 Processing the plurality of images to generate anthropometric data of a user includes: identifying a first frame of the plurality of images that is a view perpendicular to the user's face by minimizing asymmetry of key points in the first frame; Identifying a second frame of the plurality of images, the second frame being a view perpendicular to a first pinna of the user, according to a 90 degree view from the first frame; identifying a third frame of the plurality of images, the third frame being a view perpendicular to a second pinna of the user, according to a 180 degree view from the second frame; 34. The method of any one of embodiments 29 to 33. Aspect 35 the second frame is one of a plurality of second frames selected from the plurality of images within a range of +45 degrees and −45 degrees around a view perpendicular to the first pinna; the third frame is one of a plurality of third frames selected from the plurality of images within a range of +45 degrees and −45 degrees around a view perpendicular to the second pinna; 35. The method of any one of embodiments 29 to 34. Aspect 36 Processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using a neural network on the key image frames to identify anthropometric characteristics of the user; generating the anthropometric data by determining measurements of the anthropometric characteristics; 36. The method of any one of embodiments 29 to 35. Aspect 37 To generate a personalized HRTF: providing an HRTF model trained by performing a machine learning process, optionally including lasso regression, on a high-resolution database of anthropometric data and measured magnitude / frequency responses; generating the personalized HRTFs by applying the HRTF model to the anthropometric data of the user. 37. The method of any one of embodiments 29 to 36. Aspect 38 generating the personalized HRTF on a server device; transmitting the personalized HRTF from the server device to a user device. 38. The method of any one of embodiments 29 to 37. Aspect 39 A method according to any one of aspects 29 to 38, wherein a user device generates an audio output by applying the personalized HRTF to an audio signal, and the user device comprises one of a headset, a pair of earbuds, and a pair of hearables. Aspect 40 A method according to any one of aspects 29 to 39, wherein the anthropometric data includes at least one of the user's shoulder width, the user's neck width, the user's neck height, the user's face height, the user's interpupillary distance, and the user's cheekbone width. Aspect 41 41. The method of any one of aspects 29 to 40, wherein the anthropometric data includes, for each auricle of the user, at least one of pinna flare angle, pinna rotation angle, pinna cleft angle, posterior pinna offset, inferior pinna offset, pinna height, pinna width, first intertragus width, second intertragus width, fossa height, concha width, concha height, and concha navicular height. Aspect 42 A method according to any one of aspects 29 to 41, wherein the anthropometric data further includes other data, the other data including at least one of the user's age, the user's weight, the user's gender, and the user's height, and the other data originates from a source other than processing the plurality of images. Aspect 43 A method according to any one of aspects 39 to 42, wherein the audio signal includes a plurality of audio objects including positional information, and generating the audio output corresponds to generating a binaural audio output by applying the personalized HRTF to the plurality of audio objects. Aspect 44 A non-transitory computer-readable medium storing a computer program that, when executed by a processor, controls an apparatus to perform processing including the method described in any one of aspects 29 to 43. Aspect 45 1. An apparatus for generating a personalized head-related transfer function (HRTF), the apparatus comprising: At least one processor; and having at least one memory, the at least one processor is configured to control the device to receive a plurality of images of a user; the at least one processor is configured to control the device to process the plurality of images to generate anthropometric data of a user; the at least one processor is configured to input the anthropometric data into an HRTF calculation system to control the device to obtain the personalized HRTF. Device.

Claims

1. 1. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: capturing user-generated data, the capturing including obtaining demographic data provided by the user and capturing video data of the user, the video data including multiple views of the user's head; processing the video data to extract a plurality of images of a user; receiving the plurality of images of a user; Processing the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; inputting the anthropometric data and the demographic data into an HRTF calculation system to obtain the personalized HRTF; Processing the plurality of images to generate the anthropometric data includes: converting the plurality of images into a three-dimensional point cloud model; selecting key image frames from the plurality of images using the three-dimensional point cloud model; generating the anthropometric data based on the key image frames; method.

2. capturing the video data includes capturing an image of an object having a known size; processing the plurality of images to generate the anthropometric data includes converting the anthropometric data from pixel measurements to absolute distance measurements using the known sizes; The method of claim 1.

3. The method of claim 1 , wherein the electronic device includes a camera, and the capturing of the video data is performed using the camera.

4. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: capturing user-generated data, the capturing including obtaining demographic data provided by the user and capturing video data of the user, the video data including multiple views of the user's head; processing the video data to extract a plurality of images of a user; receiving the plurality of images of a user; Processing the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; inputting the anthropometric data and the demographic data into an HRTF calculation system to obtain the personalized HRTF; A device at a given distance captures the video data of the user, the method further comprising: determining the given distance by measuring the time delay between emitting sound from headphones positioned proximate to the user and receiving the sound at a microphone on the device; processing the plurality of images to generate the anthropometric data includes converting the anthropometric data from pixel measurements to absolute distance measurements using the given distance; method.

5. 1. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: capturing video data of a user, the video data including multiple views of the user's head; processing the video data to extract a plurality of images of a user; receiving the plurality of images of a user; Processing the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; inputting the anthropometric data into an HRTF calculation system to obtain the personalized HRTF; Processing the plurality of images to generate anthropometric data of a user includes: identifying a first frame of the plurality of images that is a view perpendicular to the user's face by minimizing asymmetry of key points in the first frame; Identifying a second frame of the plurality of images, the second frame being a view perpendicular to a first pinna of the user, according to a 90 degree view from the first frame; identifying a third frame of the plurality of images, the third frame being a view perpendicular to a second pinna of the user, according to a 180 degree view from the second frame; method.

6. processing the plurality of images to generate anthropometric data of the user includes identifying at least a second frame and a third frame; the second frame is one of a plurality of second frames selected from the plurality of images within a range of +45 degrees and −45 degrees around a view perpendicular to a first pinna of the user; the third frame is one of a plurality of third frames selected from the plurality of images within a range of +45 degrees and −45 degrees around a view perpendicular to the user's second auricle; The method of claim 1.

7. The method of claim 1 , wherein identifying or selecting key image frames from the plurality of images is based on frame content and one or more sharpness metrics.

8. 4. Generating a personalized HRTF: providing an HRTF model trained by performing a machine learning process, optionally including lasso regression, on a high-resolution database of anthropometric data and measured magnitude / frequency responses; generating the personalized HRTFs by applying the HRTF model to the anthropometric data of the user. The method of claim 1.

9. generating the personalized HRTF on a server device; transmitting the personalized HRTF from the server device to a user device. The method of claim 1.

10. The method of claim 1 , wherein the electronic device is a user device, and further comprising generating the personalized HRTF on the user device.

11. 10. The method of claim 1, wherein processing the plurality of images to generate the anthropometric data comprises using at least one of a photogrammetry component, a context transformation component, a landmark detection component, and an anthropometry component of the electronic device.

12. 1. A method for generating a personalized head-related transfer function (HRTF) on an electronic device, the method comprising: capturing video data of a user, the video data including multiple views of the user's head; processing the video data to extract a plurality of images of a user; receiving the plurality of images of a user; Processing the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; inputting the anthropometric data into an HRTF calculation system to obtain the personalized HRTF; processing the plurality of images to generate the anthropometric data includes using a landmark detection component, a 3D projection component, and an angle and distance measurement component; the landmark detection component receives a cropped image set of the user's anthropometric landmarks and generates a set of 2D coordinates of the user's anthropometric landmarks from the cropped image set; the 3D projection component receives the set of 2D coordinates and a plurality of camera transformations and uses the camera transformations to generate a set of 3D coordinates corresponding to the set of 2D coordinates of each anthropometric landmark in 3D space; the angle and distance measurement component receives the set of 3D coordinates and generates anthropometric data from the set of 3D coordinates, the anthropometric data corresponding to angles and distances of the anthropometric landmarks at the set of 3D coordinates; the electronic device generates the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system. method.

13. 11. The method of claim 10, wherein the user device generates an audio output by applying the personalized HRTF to an audio signal, and the user device comprises one of a headset, a pair of earbuds, and a pair of hearables.

14. 2. The method of claim 1, wherein the anthropometric data includes at least one of a user's shoulder width, a user's neck width, a user's neck height, a user's face height, a user's interpupillary distance, and a user's cheekbone width.

15. 2. The method of claim 1, wherein the anthropometric data includes, for each auricle of the user, at least one of auricle flare angle, auricle rotation angle, auricle cleft angle, posterior auricle offset, inferior auricle offset, auricle height, auricle width, first intertragus width, second intertragus width, fossa height, concha width, concha height, and concha navicular height.

16. 10. The method of claim 1, wherein the anthropometric data further comprises other data, the other data comprising at least one of a user's age, a user's weight, a user's gender, and a user's height, the other data being obtained from a source other than processing the plurality of images.

17. 14. The method of claim 13, wherein the audio signal comprises a plurality of audio objects including position information, and wherein generating the audio output corresponds to generating a binaural audio output by applying the personalized HRTFs to the plurality of audio objects.

18. When executed by one or more processors, capturing user-generated data, the capturing including obtaining demographic data provided by the user and capturing video data of the user, the video data including multiple views of the user's head; processing the video data to extract a plurality of images of a user; receiving the plurality of images of a user; Processing the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; inputting the anthropometric data and the demographic data into an HRTF calculation system to obtain a personalized head-related transfer function (HRTF), processing the plurality of images to generate the anthropometric data includes using a landmark detection component, a 3D projection component, and an angle and distance measurement component; the landmark detection component receives a cropped image set of the user's anthropometric landmarks and generates a set of 2D coordinates of the user's anthropometric landmarks from the cropped image set; the 3D projection component receives the set of 2D coordinates and a plurality of camera transformations and uses the camera transformations to generate a set of 3D coordinates corresponding to the set of 2D coordinates of each anthropometric landmark in 3D space; the angle and distance measurement component receives the set of 3D coordinates and generates anthropometric data from the set of 3D coordinates, the anthropometric data corresponding to angles and distances of the anthropometric landmarks at the set of 3D coordinates; the one or more processors generating the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system. Non-transitory computer-readable medium.

19. 1. An apparatus for generating a personalized head-related transfer function (HRTF), the apparatus comprising: at least one processor; and having at least one memory, the at least one processor is configured to control the device to receive user-generated data, the receiving including receiving demographic data provided by the user and receiving a plurality of images of the user; The at least one processor is configured to control the apparatus to process the plurality of images to generate anthropometric data of a user, wherein processing the plurality of images to generate anthropometric data of a user includes: identifying a key image frame from the plurality of images; using the key image frames to identify anthropometric features of a user; generating the anthropometric data by determining measurements of the anthropometric features; the at least one processor is configured to input the anthropometric data and the demographic data into an HRTF calculation system to control the device to obtain a personalized head-related transfer function (HRTF); processing the plurality of images to generate the anthropometric data includes using a landmark detection component, a 3D projection component, and an angle and distance measurement component; the landmark detection component receives a cropped image set of the user's anthropometric landmarks and generates a set of 2D coordinates of the user's anthropometric landmarks from the cropped image set; the 3D projection component receives the set of 2D coordinates and a plurality of camera transformations and uses the camera transformations to generate a set of 3D coordinates corresponding to the set of 2D coordinates of each anthropometric landmark in 3D space; the angle and distance measurement component receives the set of 3D coordinates and generates anthropometric data from the set of 3D coordinates, the anthropometric data corresponding to angles and distances of the anthropometric landmarks at the set of 3D coordinates; the at least one processor is configured to control the device to generate the personalized HRTF for the user by inputting the anthropometric data into the HRTF calculation system. Device.

Citation Information

Patent Citations

  • On-vehicle speaker system and audio device

    JP2006279548A

  • Sound system for seat speaker

    JP2007053622A

  • Method and apparatus for individualizing HRTFs through modeling

    JP2008527821A

  • Stereophonic sound reproduction device and program

    JP2017085362A

  • Method of Modifying Audio Content

    US20070270988A1