How to determine a personalized head transfer function
An artificial neural network-based method predicts personalized HRTFs from ear images, addressing the inefficiency of experimental measurement and enabling tailored immersive audio experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HARMAN INT IND INC
- Filing Date
- 2021-12-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for determining personalized head-related transfer functions (HRTFs) are complex and time-consuming, requiring experimental measurements in users' ears, which is inefficient for widespread application in immersive audio environments.
A method using artificial neural networks trained with ear images and directional data to predict personalized HRTFs, allowing for the calculation of direction-dependent auditory system-related function values without direct experimental measurement on the user.
Enables efficient determination of personalized HRTFs based on ear images, reducing the need for complex experimental setups and enabling immersive audio experiences tailored to individual anatomical features.
Smart Images

Figure 0007862389000001 
Figure 0007862389000002 
Figure 0007862389000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the determination of personalized head-related transfer functions. In particular, the present disclosure relates to systems, methods, and devices for determining personalized frequency responses of a personalized head-related transfer function (HRTF) based on an image of an ear. Applications include audio processing.
Background Art
[0002] The human auditory system can not only perceive sounds but also determine the direction from which the sounds arrive. The human brain achieves this using three sound characteristics. First, sounds arriving from one side of the head are registered as louder by the ear closer to the sound source (detection of the interaural amplitude difference). Second, the sound arrives earlier at the ear closer to the head (detection of the interaural time difference). Third, as the sound propagates through the listener's body, including the shoulders, head, outer ear, and middle ear, until it reaches the inner ear where it is converted into electrical nerve signals, the spectrum of the sound is distorted. In particular, the pinna of the outer ear is highly asymmetric and is configured to modify the spectrum of the sound depending on the direction of the sound source relative to the head.
[0003] Therefore, to create an immersive acoustic environment, for example, for a virtual reality application, it is necessary to mimic the position of the sound source by generating the interaural amplitude and time differences and modifying the spectrum of the sound according to the direction of the sound source and the listener's anatomical structure. The modification of the spectrum is achieved using a head-related transfer function (HRTF). To achieve a natural acoustic experience for the listener, the HRTF must be individualized according to the listener's anatomical structure, particularly the shape of the listener's ear pinna.
[0004] Head-related transfer functions (HRTFs) can be generated experimentally, that is, by placing microphones in the user's ears and recording the sound produced by a loudspeaker near the user's head. However, this procedure is complex and time-consuming. Therefore, there is a need to predict individualized HRTFs or their frequency responses based on simpler measurements. [Overview of the project] [Means for solving the problem]
[0005] This specification discloses and claims methods and systems for determining a personalized head-related transfer function based on one or more images of the ear.
[0006] According to a first aspect, a computer implementation method for determining a personalized head-related transfer function is disclosed. The method includes receiving a first training data subset, which includes one or more training images of the ears of one or more first users, and training input vectors indicating one or more first directions relative to the heads of the first users; receiving a second training data subset, which includes one or more values of a personalized head-related direction function associated with the first user and the first directions of the first training data subset; supplying the first and second training data subsets to an artificial neural network as a training dataset; training the artificial neural network on the training dataset to predict one or more personalized values of the direction function; receiving an inference dataset, which includes an inference image of the ear of a second user or a pair of inference images of the ears of a second user, and inference input vectors indicating a second direction relative to the head of the second user; and processing the inference dataset by the artificial neural network to predict one or more personalized values of the direction function, wherein the values are associated with the second user and the second direction.
[0007] The method aims to determine a personalized head-related direction function. This function is individual-specific because it depends on anatomical features, particularly the shape of the auricle. The function allows for the calculation of direction-dependent auditory system-related function values, such as changes in volume, spectrum, or delay of sound waves as they reach the inner ear, depending on the direction of the incoming sound wave. As will be described in more detail below, the function may, in one embodiment, include a personalized head-related transfer function.
[0008] The method uses an artificial neural network, such as a convolutional neural network, and includes a training phase and an inference phase. In the training phase, a training dataset is received, which includes a first training data subset and a second training data subset. The first training data subset relates to possible input data and includes images of the auricle and methods that may be given, for example, in spherical or Cartesian coordinates. The second training data subset relates to output data corresponding to the input data and includes one or more values of a personalized head-related orientation function related to a first user and a first orientation of the first training data subset. The one or more values supplied for training may be determined by experimental measurement or simulation and are considered correct values. In one embodiment, if the personalized head-related orientation function relates to the frequency response of a head-transfer function, the value of the function may be a spectral response function that relates the acoustic spectrum reaching the inner ear to the spectrum of impact sound. The artificial neural network is typically trained to reproduce the second training data subset using a number of images of the auricle of different first users and a number of different orientations. The first user may be a test user participating in an experimental measurement session in a specialized laboratory. Training may involve the use of numerous training datasets from a large number of test users, the training datasets including images of both ears and recorded head-related directional functions for a certain number of directions. Thus, the method allows for training a neural network using data from a laboratory setting and then applying the trained neural network to users who have not participated in laboratory testing. During inference, the trained artificial neural network receives images of one or both of the second user's ears and one or more directional vectors. The artificial neural network then predicts a personalized function for the second user. In particular, a fixed set of directions given as pairs of azimuth and elevation angles may be used as both the first direction for training and the second direction for inference. Values may be interpolated for intermediate directions, as long as only separate sets of directions are used for inference.Therefore, a second user, who may be a consumer, does not need to undergo measurements to determine the head-related direction function; they only need to take a photograph of one or both of their ears to determine the function. For example, a pair of headphones may be personalized so that the end user can provide an image of their ears to the algorithm, which then calculates the direction function according to the method described above, and this function may be applied to any sound played by the headphones to generate personalized sound. For example, sound effects in a movie or computer game may be tailored to the user's auditory system to create an acoustic experience that precisely matches the user's auditory system. By using multiple functions for different directions, it may be possible to create an immersive acoustic environment for the user.
[0009] In one embodiment, the direction function includes a personalized head impulse response function. HRIR represents the changes in sound as it crosses the human body, modified by anatomical features such as refraction, diffraction, and attenuation by the head and shoulders.
[0010] In further embodiments, the direction function includes a personalized head-related transfer function. Similar to the head-related impulse response (HRIR), the head-related transfer function (HRTF) represents the changes in sound as it crosses the human body, and is modified by anatomical features, such as refraction, diffraction, and attenuation by the head and shoulders.
[0011] For the purposes of this disclosure, HRIR is defined as the ratio of sound pressure in the blocked external auditory canal to sound pressure at the center of the head. The center of the head is defined as the midpoint between the two ears. HRIR generally depends on the location of the sound source (i.e., distance, elevation, and azimuth), as well as anatomical features (e.g., head size and shape, auricle shape), and wavelength. The Fourier transform of HRIR is called HRTF. The amplitude of the phase-free HRTF is called the frequency response.
[0012] The HRTF, therefore, provides transmission values depending on the direction and wavelength of the colliding sound waves. Thus, the HRTF may be expressed as a set with a spectrum, each representing the spectral amplitude of the transmission as a function of frequency for one angle. The spectrum may be represented as a vector containing discrete values. Spectra for multiple directions may be required. For example, the entire sphere around the user's head may be covered in azimuth and elevation angles, for example, in 20-degree increments. To make the auditory experience more natural for the user, finer increment angles may be selected, while coarser increment angles allow for savings in computational cost and memory space. The personalized frequency response may then be applied to audio signals, for example, from movies, simulations, or computer games, to generate the impression that the sound is coming from a second direction.
[0013] In a further embodiment, the direction function includes the frequency response of a personalized head-related transfer function. This allows only the spectral dependence of the HRTF to be determined using the method. The total amplitude and the amplitude difference between the two ears are determined using known physical relationships regarding amplitude decay with respect to distance and direction, without relying on an artificial neural network. This improves the convergence of the artificial neural network and the reliability of the method.
[0014] In a further embodiment, the frequency response is the frequency response of a free sound field. This means that the frequency response is calculated based on the assumption that the sound source is located at a considerable distance from the user. This is an advantage in determining the frequency response. In a laboratory setting for determining the training dataset, for example, an anechoic chamber, an audio signal can be generated using a speaker at a relatively distant distance. The audio signal reaching the ear can then be measured using an intra-ear microphone. The head impulse response is then determined by subtracting the signal emitted by the speaker from the signal recorded by the microphone. In this setting, neither the size of the sound source nor any reflections affect the measurement. This makes it possible to create a training dataset that reflects only the anatomical features of the first user.
[0015] In a further embodiment, the training dataset and / or inference dataset includes at least one pair of images, one of which is a mirror image of the left ear and the other of the right ear of the same user; and at least one pair of first input vectors corresponding to a first direction and / or a second direction, and a second input vector determined by mirroring the first input vectors with respect to a plane passing through the center of the user's head and perpendicular to a line between the user's ears.
[0016] This allows for the simultaneous training of artificial neural networks for both the left and right ears, with data from both ears included in the same dataset. In the above embodiment, both one image and its corresponding direction vector are transformed to be similar to the other image and its direction vector. Mirroring one of the vectors can lead to a 270-degree azimuth angle for the right ear corresponding to 90 degrees for the left ear. The artificial neural network can compute separately personalized values for the left and right ears. Thus, the personalized values of the direction function predicted by the artificial neural network include pairs of values for the right and left ears. Both values in this pair generally correspond to sounds from the same direction, associated with the same external sound source. For example, if the sound source is directly in front of the left ear, i.e., at a 90-degree azimuth angle in spherical coordinates, the sound source is on the opposite side of the right ear, which leads to the filter function for the left ear being similar, but not identical, to the filter function for the right ear when the sound source is directly in front of the right ear, i.e., at a 270-degree azimuth angle.
[0017] In a further embodiment, the image is a photograph. The photograph may be taken with a digital camera and stored in memory on a mobile device or a network-accessible server. A complete set of frequency responses covering all directions may be obtained, along with multiple directions. No further data, particularly anthropometric data, is required. Furthermore, there is no need to convert the photograph into a set of ear-related parameters, such as the longitudinal and transverse directions of the auricle. Rather, the photograph is simply input into an artificial neural network. Thus, the second user may be a consumer using headphones for an immersive acoustic environment, and only two images need to be taken to enable sound adaptation.
[0018] In a further embodiment, the image is a depth map. A depth map is a two-dimensional grayscale image where the position of each pixel is related to a lateral position, as in the case of a photograph, and the value is related to the height of the skin surface at that lateral position. A depth map can be obtained by processing one or more photographs with an image processing algorithm to determine the three-dimensional contour of the ear. Alternatively, multiple visible or infrared markers may be projected onto the ear when recording the photograph, thereby making it possible to obtain information about the three-dimensional structure of the auricle. This technique is known as the use of a depth camera. While a depth map does not contain the complete three-dimensional structure of the ear, it provides information that can be used to train an artificial neural network and is also relatively easy to measure. Thus, depth maps offer a compromise between the accuracy of function determination for which a complete three-dimensional model of the auricle is more suitable and the use of photographs, which are easier to capture.
[0019] In further embodiments, the method further includes determining an input audio signal transmitted to the user's head from a first direction; recording the audio signal transmitted into the ear of the first user; determining a head impulse response based on the input audio signal and the transmitted audio signal; and generating a second training data subset by converting the head impulse response into frequency space.
[0020] Here, the first user may be a test user participating in an experimental measurement session in a specialized laboratory. The input audio signal may be transmitted using a loudspeaker or another sound source. The transmitted audio signal may be recorded using an intra-ear microphone. An anechoic chamber may be used to determine the free-field transfer function. Alternatively, the room may include objects such as reflective surfaces to determine the transfer function in the presence of reflections. Determining the head-impulse response involves subtracting the input audio signal from the transmitted audio signal. The head-impulse response is then transformed into frequency space, for example by applying a Fourier transform or wavelet transform, to generate a head-transfer function that serves as a second training data subset. This may be combined with further processing steps, such as normalizing the values so that only relative spectral differences are reflected. This is particularly advantageous when other techniques are used to account for amplitude differences related in different directions.
[0021] In further embodiments, processing a training dataset and / or an inference dataset with an artificial neural network includes, by the head block of the artificial neural network, extracting features from an image of the auricle to generate feature data; creating a copy of the feature data for each coordinate of the direction vector; multiplying each copy by the coordinate of the input vector to generate a number of weighted copies; and by the tail block of the artificial neural network, processing the weighted copies to predict a head-associated direction function related to a second user and a second direction.
[0022] The processing step of extracting feature data enables the generation of preprocessed data associated with the extracted features. Known techniques for feature extraction, particularly combinations of convolutional, pooling, and fully connected layers, or any other form of machine learning algorithm may be used. The feature data is then copied so that one copy is created for each coordinate of the direction vector. The direction vector may be given in Cartesian or spherical coordinates and thus may have three components. Three copies of the data are then created and multiplied by the corresponding coordinate values to generate weighted copies. The tail block then processes the weighted copies. The processing step may include combinations of convolutional, pooling, and fully connected layers, or any other form of machine learning algorithm. Alternatively, a six-component vector may be used, as detailed below. The technique using head and tail blocks improves reliability and convergence. While the head and tail blocks can be trained individually, it is possible to train the entire algorithm containing both blocks together. This reduces the complexity of the training process.
[0023] In a further embodiment, the input vector is specified in a positive-definite 6-component format. This ensures that only positive values are used, thereby improving the convergence and performance of the artificial neural network.
[0024] In a further embodiment, the method further includes specifying one or more input vectors in Cartesian coordinates and transforming the input vectors into a format having six components by defining a pair of components for each Cartesian coordinate, wherein the first component of each pair is identical to the Cartesian coordinate if the Cartesian coordinate is non-negative and 0 if the Cartesian coordinate is negative, and the second component of each pair is 0 if the Cartesian coordinate is non-negative and identical to the absolute value of the Cartesian coordinate if the Cartesian coordinate is negative.
[0025] As a result, a three-dimensional Cartesian direction vector (X, Y, Z) is converted into a six-dimensional vector (Xp, Yp, Zp, Xn, Yn, Zn). Here, when X ≥ 0, Xp = X, and when X < 0, Xp = 0. Further, when X ≥ 0, Xn = 0, and when X < 0, Xn = abs(X). The other coordinates are calculated from Y and Z in the same way. As a result, all components become positive and half of the components become zero. This enables faster convergence of the training.
[0026] In a further embodiment, the method further includes post-processing one or more personalized values of a direction function to generate a filter, and applying the filter to a second input signal.
[0027] This filter can be generated by, for example, applying an inverse Fourier transform to convert the frequency response from the frequency domain to the time domain. Further, different volume levels may be applied to take into account the interaural amplitude difference, and the signal may be time-shifted to take into account the interaural time difference. As a result, a personalized head impulse HRIR-based filter is generated. However, other steps may be performed to generate other filters. When the personalized head impulse filter is applied to the second input audio signal, an audio signal is generated that appears to the second user to have arrived from a second direction. This approach makes it possible to use an artificial neural network only to determine the difference in spectral amplitudes associated with different directions. As a result, the computational cost is reduced. For example, the second user may be a consumer using headphones of an immersive acoustic system. Then, the head impulse filter may be generated using an image of the user's auricle, whereby the second input audio signal can be processed to appear to have arrived from a predetermined direction.
[0028] In a further embodiment, the method further includes determining and / or storing a second personalized frequency response in a mobile device. For example, any step of the method according to the present disclosure may be performed on a mobile device. Alternatively, the artificial neural network may be trained on a computing server, and only the inference step may be performed on the mobile device. According to yet another alternative, the artificial neural network may be used on a network-accessible server to benefit from the computing resources of the server and from the centralized maintenance and updating of the artificial neural network. For example, the artificial neural network may be developed and trained on one or more computing servers and then stored on a network-accessible server for performing inference steps for a plurality of second users. Thereby, the artificial neural network exhibits consistent behavior for a plurality of second users.
[0029] In a further embodiment, the method further includes determining and / or storing one or more personalized values of a direction function on a network-accessible server. Thereby, a user profile, each including a personalized value for a particular second user for one or more devices, may be used to manage the configuration. Thereby, access to the second personalized frequency values by a plurality of mobile devices is enabled, whereby an immersive acoustic environment may be consistently created on different devices used by the same user and for devices shared by a plurality of users.
[0030] Other aspects include a pair of headphones, a data processing system, a computer program product, and a computer-readable storage medium including the computer program product, all configured to perform the method of the present disclosure. All features of the first aspect of the present disclosure apply also to the other aspects. This specification also provides, for example, the following: (Item 1) A computer implementation method for determining a personalized head transfer function, One or more training images of the ears of one or more first users, and Training input vectors indicating one or more first directions relative to the heads of one or more first users Receiving a first training data subset, which includes, The second training data subset receives one or more values of personalized head-related direction functions associated with one or more first users and one or more first directions of the first training data subset, The first training data subset and the second training data subset are supplied to the artificial neural network as a training dataset, To predict one or more personalized values of the personalized head-related direction function, the artificial neural network is trained on the first training data subset and the second training data subset. An inferred image of the second user's ear or a pair of inferred images of the second user's ear, and Inference input vector indicating the second direction relative to the head of the second user Receiving an inference dataset that includes, Processing the inference dataset by the artificial neural network in order to predict one or more personalized values of the personalized head-related direction function, wherein the personalized values are related to the second user and the second direction. The method, including the method described above. (Item 2) The method according to item 1, wherein the personalized head-related direction function includes a personalized head-impulse response. (Item 3) The method according to item 1, wherein the personalized head-related direction function includes a personalized head transfer function. (Item 4) The method according to item 1, wherein the personalized head-related direction function includes the frequency response of the personalized head transfer function. (Item 5) The method according to item 4, wherein the frequency response is a free-field frequency response. (Item 6) At least one of the training dataset or the inference dataset is Images of at least one pair of the same user's left and right ears, wherein one of the pair of images is a mirror image of the image, A first input vector corresponding to at least one of the first or second directions, and a second input vector determined by mirroring the first input vector with respect to a plane passing through the center of the user's head and perpendicular to the line between the user's ears, The method described in item 1, including the method described in item 1. (Item 7) The method according to item 6, wherein the image of the left ear is a photograph. (Item 8) The method according to item 6, wherein the image of the left ear is a depth map. (Item 9) Determining an input audio signal transmitted from a first direction to the head of one of the one or more first users, Recording the audio signal transmitted into the ear of one of the one or more first users, Based on the input audio signal and the transmitted audio signal, the head impulse response is determined. Converting the head impulse response into frequency space, The method according to item 1, further comprising determining the second training data subset by means of the method. (Item 10) Processing the inference dataset using the artificial neural network is The head block of the artificial neural network extracts features from an image of one of the auricles of the second user's ear to generate feature data, To create a copy of the feature data for each coordinate of the direction vector, Multiplying each copy by the coordinates of the inference input vector generates multiple weighted copies, The tail block of the artificial neural network processes the weighted replication to predict the personalized head-related direction function associated with the second user and the second direction, The method described in item 1, including the method described in item 1. (Item 11) The method according to item 1, wherein one or more of the training input vectors or the inference input vectors are specified in positive definite 6-component format. (Item 12) Specifying one or more of the training input vectors or the inference input vectors in Cartesian coordinates, By defining a pair of components for each Cartesian coordinate, one or more of the training input vectors or the inference input vectors are converted into a format containing six components, It further includes, The first component of each pair is identical to the Cartesian coordinate if the Cartesian coordinate is not negative, and is 0 if the Cartesian coordinate is negative. The method according to item 11, wherein the second component of each pair is 0 if the Cartesian coordinate is not negative, and is the same as the absolute value of the Cartesian coordinate if the Cartesian coordinate is negative. (Item 13) Post-processing one or more personalized values of the personalized head-related direction function to generate a filter, The method according to item 1, further comprising applying the filter to a second input audio signal. (Item 14) The method according to item 1, further comprising determining, storing, or determining and storing one or more personalized values of the personalized head-related direction function on a mobile device. (Item 15) The method according to item 1, further comprising determining, storing, or determining and storing one or more personalized values of the personalized head-related direction function on a network-accessible server. (Item 16) It is a system, Memory and One or more processors configured to perform any of the methods described in items 1 to 15, The system comprising the above. (Item 17) A computer-readable storage medium that includes instructions that cause one or more processors to perform any of the methods described in items 1 to 15 when executed by one or more processors.
[0031] The features, purposes, and advantages of this disclosure will become more apparent from the detailed description below when interpreted in conjunction with the drawings, in which similar reference numbers refer to similar elements. [Brief explanation of the drawing]
[0032] [Figure 1] A block diagram of a system according to one embodiment of this disclosure is shown. [Figure 2] A flowchart of a computer implementation method for training an artificial neural network to determine a personalized head-related directional function, according to one embodiment of the present disclosure, is shown. [Figure 3] A flowchart of a computer implementation method for determining a personalized head-related direction function using an artificial neural network, according to one embodiment of the present disclosure, is shown. [Figure 4] A flowchart of a computer implementation method for processing training and / or inference datasets using an artificial neural network, according to one embodiment of the present disclosure, is shown. [Modes for carrying out the invention]
[0033] Figure 1 shows a block diagram of a system 100 according to one embodiment of the present disclosure. The system 100 comprises a server system 102 and one or more client systems 120, which are connected to each other so as to be able to communicate via a network 118.
[0034] The server system 102 is configured to train and apply an artificial neural network. For training, data is determined from one or more first users. The first users may be subjects participating in a measurement campaign to determine the training dataset. Determining the training dataset involves taking images of the first user's ear with camera 104. Camera 104 may be a conventional photographic camera or a camera with some three-dimensional capabilities. For example, a depth camera may be used, in which case optical markers are projected onto the auricle and their positions are used to obtain more information about shape. One or more speakers 106 and one or more microphones 108 may be part of a setup for determining the head-related free-field impulse response in an anechoic chamber. The setup may include an intra-ear microphone placed in the first user's ear and a plurality of speakers spherically arranged around the first user's head, each at a distance of 1.2 meters from the center of the head, to generate free-field audio signals. Multiple input audio signals may be generated. An intra-ear microphone is configured to record transmitted sound and simultaneously block the ear. However, this disclosure is not limited to this type of camera, speaker, and microphone. Rather, different devices may be used. For example, multiple speakers may be placed near the head to generate a near-field signal. A reverberant environment may be used instead of an anechoic chamber. In yet another example, a sound source other than a speaker may be used. In that case, two signals associated with the sound waves may be measured at two points in space. The two signals may include a signal outside the head and a transmitted signal measured by an intra-ear microphone. The generated data may then be processed by a server computer 110 having a processor 112 and memory 114. Processing may include training an artificial neural network (ANN) 116 on the data, testing the ANN 116 using the generated data, and performing inference steps to predict a direction function. Processing may further include pre-processing and post-processing of the data, as detailed below.The server system 102 may be localized in one location, but alternatively, it may include devices 104-116 distributed to different locations and connected via a network 118, such as the Internet.
[0035] One or more client systems 120 may include a camera 122, headphones 124, and a client computer 126 including a processor 128 and a memory device 130. The client system 120 may be used by a second user, who may be a consumer using headphones. The camera 122 is configured to take one or more photographs of the second user's ear. The camera 122 may be a depth camera as detailed above. For example, the camera 122 and the client computer 126 may be included in a smartphone. The client computer 126 may then preprocess the image, for example, by generating a depth map based on data generated by the camera 122, and the image may include, for example, a photograph or some three-dimensional data. Further preprocessing may include mirroring one of the images. The image may then be transmitted to a server system 102 via a network 118. The server computer 110 may then generate a predetermined number of direction vectors. The artificial neural network 116 may then process the image and direction vectors to generate a direction function. The direction function may include, for example, the head impulse response for each direction. The direction function 132 is then transmitted to the client computer 126 via the network 118 and can be stored in memory 130. This means that data transfer over the network is only required for the calibration of a new user's client system. The client computer 126 can then apply the function 132 to the original audio signal emitted by the headphones. Further steps may be performed, such as correcting the amplitude of the left and right ears and inducing a phase shift to compensate for delay. Alternatively, these steps may also be performed by a second client computer (not shown) which may be included in the set of headphones. This makes it possible to generate the impression that the original audio signal is coming from a given direction. The given direction may be stored in metadata stored with the original audio signal.
[0036] It should be noted that System 100 in Figure 1 represents only an exemplary embodiment. Alternatively, only the training and testing of the artificial neural network may be performed by the server system, and the inference steps may be performed on one or more client systems. In yet another alternative, a single local system including a camera, speaker, microphone, and computer may be used, and all method steps may be performed by the system.
[0037] Figure 2 shows a flowchart of a computer implementation method for training an artificial neural network to determine a personalized head-related directional function, according to one embodiment. In step 202, one or more images are created, for example, by taking a photograph with a camera. Images containing more 3D data may be used, but a photograph is already sufficient to train the artificial neural network. In addition, optional preprocessing steps may be performed. In step 204, one image of the same person's ear is mirrored; that is, the image is flipped horizontally. This provides the artificial neural network with two images of ears, which are identical in orientation but represent differences in the shape of the ear present, particularly the shape of the auricle. In step 206, one or more images may be converted into a depth map, i.e., a 2D grayscale image, where the gray values represent the position of the skin surface in the vertical direction. These steps may improve the conversion behavior of the artificial neural network. In step 208, one or more input vectors corresponding to directions are specified. The input vectors may be given in Cartesian or spherical coordinates. For example, a set of input vectors indicating a predetermined number of directions may be specified. For example, the direction may extend across the entire angular range of the sphere in 20-degree increments for both azimuth and elevation to train the artificial neural network over the entire range of available angles. The size of the increment is a trade-off between experimental and computational effort on the one hand and the accuracy of the trained artificial neural network on the other. The first training data subset includes one or more images and one or more input vectors.
[0038] Steps 210-216 relate to generating a direction function that is fed into the artificial neural network as a second training data subset. In this exemplary embodiment, the direction function includes a head-related transfer function (HRTF). In step 210, an input audio signal is transmitted by, for example, one of the speakers 106. The input audio signal may include a sine sweep, a logarithmic sweep, or another signal that can cover the spectrum. In step 212, the transmitted signal is recorded by, for example, one of the microphones 108. Processing the signal involves determining the impulse response in step 214. This may involve subtracting the input audio signal from the transmitted audio signal to generate a head-related impulse response (HRIR). In step 216, the impulse response is transformed into frequency space by, for example, a Fourier transform or wavelet transform. The output of step 216 yields a second training data subset. The first and second training data subsets are then sent to the artificial neural network as training datasets. In step 218, the artificial neural network is trained on the training dataset. Training may involve one or more of the steps described with reference to Figure 4.
[0039] Figure 3 shows a flowchart of a computer implementation method for determining a personalized head-related direction function using an artificial neural network, according to one embodiment. This relates to the inference steps of the artificial neural network. In steps 302-306, one or more images of one or both of the user's ears are created and optionally mirrored and / or transformed into one or more depth maps. These steps are mainly identical to steps 202-206 above, except that the images represent the ears of a second user and may be taken with different hardware. In 308, one or more input vectors are defined. In 310, the artificial neural network processes the input dataset, i.e., one or more images and one or more input vectors, to predict the values of the direction function for one or more directions indicated by the vectors. This may include calculating a head-related transfer function (HRTF). In 312, filters are generated based on the direction function. For example, when an artificial neural network determines an HRTF, a time-domain transformation, for example, by an inverse Fourier transform, determines a filter function that can be applied to one or more original audio signals in step 314, creating the impression that the audio signal is coming from a direction indicated by a vector. For example, a direction vector may be generated so that for each given direction, the artificial neural network can process one or more images and vectors to generate a value of the direction function associated with that given direction. This makes it possible to determine the filter function with high accuracy. Alternatively, in step 308, a set of input vectors covering all angles of a sphere, for example in 20-degree increments, is generated. The artificial neural network then generates a value of the direction function for each direction. A filter is also generated for each direction. The set of filters may then be stored in memory 130 so that the filters can be applied to the audio signal. Filters may be generated by interpolating the filters in the set in order to filter the audio signal and create the impression that the audio signal is coming from a direction for which no filters exist in the set.
[0040] Figure 4 shows a flowchart of a computer implementation method 400 for processing training and / or inference datasets using an artificial neural network, according to one embodiment. The steps of method 400 may be substeps of training step 218 and inference step 310. In 402, one or more images are processed by the head block of the artificial neural network to extract features. Then in 404, the feature data is copied, and one copy may be created for each coordinate of the direction vector; for example, if the direction vector is determined by six components, six copies may be created. The direction vectors specified in steps 208 and 308 are optionally converted to Cartesian coordinates in 406. The use of Cartesian coordinates makes the conversion of the artificial neural network faster. In 408, the direction vector is converted to a six-component format. Such a format may be defined by converting each component to a pair of components, one of which is identical to the absolute value of the original component and the other is zero. This makes the convergence of the artificial neural network faster. In 410, a second vector is optionally determined. When images of both the left and right ears of the same user are processed together to train a neural network or to simultaneously predict the direction function values for both ears for the same sound source, the first vector may indicate the direction of the sound source for the first ear, and the second vector may indicate the direction of the sound source for the second ear. In this case, the second vector is mirrored in a plane perpendicular to the line passing through the center of the head and between the ears. For example, a 90-degree azimuth angle for the left ear (first vector) corresponds to a 270-degree azimuth angle for the right ear (second vector). Nevertheless, all coordinates may be specified in Cartesian coordinates. In 412, each copy is multiplied by its coordinate value to generate a weighted copy. In 414, the artificial neural network processes the weighted copies to predict the direction function values. [Explanation of symbols]
[0041] 100 Systems 102 Server System 104 Camera 106 speakers 108 Microphone 110 Server Computers 112 processors 114 memory 116 Artificial Neural Networks 118 Network 120 Client Systems 122 Cameras 124 headphones 126 client computers 128 processors 130 memory 132 Functions 200 Computer implementation method for determining personalized head-related directional functions Steps for a computer implementation method to determine personalized head-related directional functions (202-218) 300 Computer implementation method for determining personalized head-related directional functions using artificial neural networks 302-314 Steps of a computer implementation method for determining personalized head-related directional functions using an artificial neural network. 400 Computer implementations for processing training and / or inference datasets using artificial neural networks. 402-414 Steps of a computer implementation method for processing training and / or inference datasets using an artificial neural network.
Claims
1. A computer implementation method for determining a personalized head transfer function, One or more training images of the ears of one or more first users, and Training input vectors indicating one or more first directions relative to the heads of one or more first users Receiving a first training data subset including, The second training data subset receives one or more values of personalized head-related direction functions associated with one or more first directions of the one or more first users and the first training data subset, The first training data subset and the second training data subset are supplied to the artificial neural network as a training dataset, To predict one or more personalized values of the personalized head-related directional function, the artificial neural network is trained on the first training data subset and the second training data subset. An inference image of the second user's ear or a pair of inference images of the second user's ear, and The inference input vector indicating the second direction relative to the head of the second user. Receiving an inference dataset that includes, Processing the inference dataset by the artificial neural network in order to predict one or more personalized values of the personalized head-related direction function, wherein the personalized values are related to the second user and the second direction. Includes, At least one of the training dataset or the inference dataset is Images of at least one pair of the same user's left and right ears, wherein one of the pair of images is a mirror image, A first input vector corresponding to at least one of the first or second directions, and a second input vector determined by mirroring the first input vector with respect to a plane passing through the center of the user's head and perpendicular to the line between the user's ears, Methods that include...
2. The method according to claim 1, wherein the personalized head-related direction function includes a personalized head-impulse response.
3. The method according to claim 1, wherein the personalized head-related direction function includes a personalized head transfer function.
4. The method according to claim 1, wherein the personalized head-related direction function includes the frequency response of a personalized head transfer function.
5. The method according to claim 4, wherein the frequency response is a free-field frequency response.
6. The method according to claim 1, wherein the image of the left ear is a photograph.
7. The method according to claim 1, wherein the image of the left ear is a depth map.
8. Determining an input audio signal transmitted from a first direction to the head of one of the one or more first users, Recording the audio signal transmitted into the ear of one of the one or more first users, Based on the input audio signal and the transmitted audio signal, the head impulse response is determined. Converting the head impulse response into frequency space, The method according to claim 1, further comprising determining the second training data subset by means of the method.
9. A computer implementation method for determining a personalized head transfer function, One or more training images of the ears of one or more first users, and Training input vectors indicating one or more first directions relative to the heads of one or more first users Receiving a first training data subset including, The second training data subset receives one or more values of personalized head-related direction functions associated with one or more first directions of the one or more first users and the first training data subset, The first training data subset and the second training data subset are supplied to the artificial neural network as a training dataset, To predict one or more personalized values of the personalized head-related directional function, the artificial neural network is trained on the first training data subset and the second training data subset. An inference image of the second user's ear or a pair of inference images of the second user's ear, and The inference input vector indicating the second direction relative to the head of the second user. Receiving an inference dataset that includes, Processing the inference dataset by the artificial neural network in order to predict one or more personalized values of the personalized head-related direction function, wherein the personalized values are related to the second user and the second direction. Includes, Processing the inference dataset using the artificial neural network is The head block of the artificial neural network extracts features from an image of one of the ears of the second user to generate feature data, To create a copy of the feature data for each coordinate of the direction vector, Multiplying each copy by the coordinates of the inference input vector generates multiple weighted copies, The tail block of the artificial neural network processes the weighted replication to predict the personalized head-related direction function associated with the second user and the second direction, Methods that include...
10. The method according to claim 1, wherein one or more of the training input vector or the inference input vector are specified in positive definite 6-component format.
11. Specifying one or more of the training input vectors or the inference input vectors in Cartesian coordinates, By defining a pair of components for each Cartesian coordinate, one or more of the training input vectors or the inference input vectors are converted into a format containing six components. It further includes, The first component of each pair is identical to the Cartesian coordinate if the Cartesian coordinate is not negative, and is 0 if the Cartesian coordinate is negative. The method according to claim 10, wherein the second component of each pair is 0 if the Cartesian coordinate is not negative, and is the same as the absolute value of the Cartesian coordinate if the Cartesian coordinate is negative.
12. Post-processing one or more personalized values of the personalized head-related direction function to generate a filter, The method according to claim 1, further comprising applying the filter to a second input audio signal.
13. The method according to claim 1, further comprising determining, storing, or determining and storing one or more personalized values of the personalized head-related direction function on a mobile device.
14. The method according to claim 1, further comprising determining, storing, or determining and storing one or more personalized values of the personalized head-related direction function on a network-accessible server.
15. It is a system, Memory and One or more processors configured to perform the method described in any one of claims 1 to 14, A system that includes these features.
16. A computer-readable storage medium comprising, when executed by one or more processors, an instruction causing one or more processors to execute the method according to any one of claims 1 to 14.