Audio system and method for HRTF estimation
By estimating the audio spatial parameters of the target user in the audio system, the problem of difficulty in obtaining HRTF data in the prior art is solved, and high-quality spatial audio reproduction and simplified measurement process are achieved.
Patent Information
- Application Number
- CN202411940594.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art requires expensive equipment and long measurement processes when acquiring personalized head-related transfer function (HRTF) data, resulting in poor user experience and inability to achieve high-quality spatial audio reproduction.
By using a processor in the audio system, audio data is obtained using a microphone arranged near the target user's ear canal, and audio spatial parameters are estimated through machine learning models to provide a personalized HRTF.
More accurate and accurate HRTF estimation is achieved, the quality of spatial audio generation is improved, the modeling and measurement process of HRTF is simplified, and the need for expensive equipment and long-term measurements is avoided.
Smart Images

Figure CN120201362A_ABST
Abstract
Description
[0001] The present disclosure relates to an audio system and related methods, specifically for analyzing, monitoring, and / or estimating transfer functions and / or audio spatialization parameters related to a user of an audio device, such as head-related transfer functions. Specifically, an audio system and a method for estimating one or more audio spatialization parameters, as well as a computer-implemented method for training a machine learning model to provide audio spatialization parameters, are disclosed. Background Art
[0002] Within the technical field of spatial audio, access to a personalized head-related transfer function (HRTF) dataset is considered crucial. Such a dataset can be regarded as a unique acoustic fingerprint of a user's outer ears and the position of the outer ears relative to the user's head. When personalized HRTFs are carefully measured for a specific user, it becomes easier for that user to reproduce high-quality spatial audio that sounds almost the same as the sound from a real physical sound source. This is in contrast to generic HRTFs, which can be measured, for example, on a mannequin head and typically result in lower-quality spatial audio.
[0003] Therefore, an audio device that relies on a single HRTF requires the user to spend time and effort measuring their HRTF. The process of measuring a person's HRTF typically involves generating stimuli from different directions and measuring the response of microphones placed in the ears. A typical recording process requires very expensive equipment and takes at least 30 minutes.
[0004] A simplified alternative to applying personalized HRTFs is to apply the best average HRTF derived from a population. The drawback of the simplification is that individual spatial perception deteriorates, i.e., the sound is perceived as coming from other points than intended. Summary of the Invention
[0005] Therefore, there is a need to provide an audio system, an electronic device, and a method for improved spatial audio generation.
[0006] A method for estimating one or more audio spatialization parameters for a specific target user is disclosed. The method is performed, for example, in an audio system including one or more processors. The method includes obtaining audio data including first audio data and second audio data, for example, in a normal usage environment. The method further includes obtaining the first audio data from a first microphone, for example, arranged near, in, or at a first ear canal of the target user, and obtaining the second audio data from a second microphone, for example, arranged near, in, or at a second ear canal of the target user; and providing one or more audio spatialization parameters. Providing one or more audio spatialization parameters includes applying a model to the audio data, for example, using an audio device or an accessory device, for parameter estimation of the one or more audio spatialization parameters, and determining the one or more audio spatialization parameters based on the parameter estimation. The method includes outputting the one or more audio spatialization parameters.
[0007] In addition, an audio system is disclosed. The audio system includes one or more processors. The one or more processors are configured to obtain audio data including first audio data and second audio data, for example, in a normal usage environment, by obtaining the first audio data from a first microphone, for example, arranged near, in, or at a first ear canal of the target user, and obtaining the second audio data from a second microphone, for example, arranged near, in, or at a second ear canal of the target user; and providing one or more audio spatialization parameters. Providing one or more audio spatialization parameters includes applying a model to the audio data for parameter estimation of the one or more audio spatialization parameters and determining the one or more audio spatialization parameters based on the parameter estimation. The one or more processors are configured to output (such as transmit and / or store, for example, for later retrieval by the target user) the one or more audio spatialization parameters.
[0008] In addition, a computer-implemented method is provided for training a machine learning model to process audio data including, for example, first audio data indicative of a first audio signal from a first microphone and second audio data indicative of a second audio signal from a second microphone in a normal usage environment as input, and providing a parameter estimation of one or more audio spatialization parameters as output. The method includes: obtaining, by the computer, one or more audio spatialization parameters for each training user; obtaining, by the computer, a first sound input position near, in, or at a first ear canal of the training user and a second sound input position near, in, or at a second ear canal of the training user for each training user; obtaining, by the computer, a plurality of training input sets for each training user, where each training input set represents a specific sound environment and a specific time period and includes a first audio input signal and a second audio input signal, the first audio input signal and the second audio input signal each representing ambient sound from a specific sound environment at a specific time period at the first sound input position and the second sound input position, respectively; and training the machine learning model by performing multiple rounds of training across the plurality of training input sets obtained for the plurality of training users by the computer, where each round of training includes: applying the machine learning model to one of the plurality of training input sets obtained for the corresponding training user, and adjusting parameters of the machine learning model, such as weights or other parameters, using the one or more audio spatialization parameters obtained for the corresponding training user as the target output of the machine learning model.
[0009] An advantage of the present disclosure is that more accurate and precise HRTFs are provided, which in turn can lead to improved spatial audio generation. In addition, customized HRTFs can be provided without the need for highly specialized and time-consuming HRTF determination in an audio laboratory.
[0010] In addition, the present disclosure provides an improved machine learning / neural network model architecture for effectively processing and analyzing audio from ear-mounted microphones.
[0011] In addition, the present disclosure simplifies HRTF modeling / estimation / determination by leveraging audio data obtained during normal use of an audio device.
[0012] The training of one or more models disclosed herein relies on receiving real or simulated sounds generated at specific other locations in a real or virtual space at specific locations in the real or virtual space, where such locations are defined relative to a real or simulated training user. Similarly, the use of a model for estimating one or more audio spatialization parameters by a real target user relies on receiving real sounds generated at specific other locations in the real space at specific locations in the real space, where such locations are defined relative to the real target user. To achieve the desired results, all such locations should be defined relative to a common reference frame in the real and / or virtual space.
[0013] A preferred reference frame is a three-dimensional coordinate system fixed to the head of each respective training or target user, such that its origin is located at the intersection of the sagittal and coronal planes of the head, at the height of the respective entrance to the ear canal, the first axis "X" of the three-dimensional coordinate system is in the coronal plane and extends orthogonally to the sagittal plane in the backward direction of the head, the second axis "Y" of the three-dimensional coordinate system is in the sagittal plane and extends orthogonally to the coronal plane in the forward direction of the head, and the third axis "Z" of the three-dimensional coordinate system extends along the intersection of the sagittal and coronal planes in the upward direction of the head. Hereinafter, all spatial positions or directions mentioned are relative to this preferred reference frame. Obviously, other reference frames can be used, and transformations between different reference frames can be applied when needed. For example, the upward direction can be determined according to the characteristics of the pinna of each training or target user. Minor deviations between reference frames, such as those caused by measurement errors in the real space, can be ignored; however, this may reduce the accuracy of the estimation of one or more audio spatialization parameters.
[0014] When a person is subjected to sounds from sound sources in their environment (e.g., in a normal usage environment), the sounds received at the person's left and right ear canals respectively and thus perceived by the person depend on the size and shape of the person's head and ears, the size, shape, and relative position of the respective pinnae, and the dimensions of the respective ear canals. All these head- and ear-related characteristics contribute to modifying the received sounds to include spatial cues that enable a person with normal hearing to determine the location of the sound source by processing the received sounds. Such spatial cues are well known in the art and include various modifications to the received sounds, such as the shadow effect, interaural intensity difference (ILD), interaural time difference (ITD), and peaks, valleys, and phase shifts caused by constructive and destructive interference at specific frequencies. Roughly speaking, the shadow effect and interaural cues mainly enable the determination of the distance to the source and the angle of incidence relative to the sagittal plane, while other more refined spatial cues mainly achieve the discrimination of front-back positions and the determination of the elevation angle.
[0015] The fact that the features mentioned in relation to the head and ears can vary significantly between individuals suggests that directional hearing is to a large extent an acquired skill. Nevertheless, it has been established that the characteristics of the features mentioned also tend to be governed by a so-called "prototype", such that there are large groups of people, for example, whose auricle shapes are very similar. Thus, there will also be at least partial similarity between the HRTFs of the people in that group. Additionally, the shadow effect and the interaural differences depend to a large extent on the size and shape of the head and the relative position of the auricles.
[0016] When a person is in an environment where there are multiple sound sources spatially distributed around the person, such as in a large office with other people, and when these sound sources move relative to the person's head, then the sounds from the sound sources will reach the person and the person's ear canals from different directions over a period of time, possibly via reflections, and sometimes from multiple directions simultaneously. In such a case, the sounds received at the person's left and right ear canals during that period of time will exhibit various spatial cues indicative of the various relative positions of the sound sources.
[0017] The present invention provides or utilizes a machine learning model (the model), which, simply put, is trained to analyze sounds received at corresponding sound input locations near, in, or at the left and right ear canals of a person during one or more periods of time in one or more sound environments (such as in a normal usage environment), and to provide corresponding estimates of one or more parameters of the person's HRTF (audio spatialization parameters). These sound input locations are hereinafter referred to as "target sound input locations". The target user of a terminal user device implementing such a model can thus wear the terminal user device near, in, or at their ears during one or more periods of time when the model analyzes the sounds received by the corresponding microphones of the terminal user device and ultimately provides estimates of one or more audio spatialization parameters for the target user. Hereinafter, the term "target device" refers to such a terminal user device that includes a model as described herein, which enables the terminal user device to analyze the sounds received by its microphones when worn by the target user and to provide estimates of one or more audio spatialization parameters of the target user. The target device may optionally also include one or more processors for providing spatialized audio to the target user by providing an audio output based on one or more audio spatialization parameters.
[0018] In this context, the normal usage environment may not be an environment designed for testing. For example, the normal usage environment may be characterized as a workplace or office environment; a public place environment, such as an airport, railway station or bus station; a street environment; a home environment; a concert environment; a party environment; a sports environment; a school or teaching environment; a natural environment, such as by the sea or in a forest; or a transportation environment, such as on a bicycle, motorcycle, moped, or in a bus, train, car, airplane, ship, ferry, or other vehicle. The normal usage environment is not a test facility, such as an anechoic chamber or a virtual anechoic chamber. Additionally, the normal usage environment may be characterized as an environment where specialized test equipment (such as a dedicated microphone and / or a dedicated sound source) is not required or used. Thus, the normal usage environment may be characterized in that the audio for generating the audio data is not generated using test equipment, such as one or more sound sources arranged and controlled to output test audio.
[0019] Generally, in order to achieve the reproduction of high-quality spatial audio, it is preferably determined the HRTF for a specific sound source position and a specific ear, such that it defines the transfer function between the sound source position and the position of the tympanic membrane center of that ear. For the same purpose, it is preferably trained a model to estimate audio spatialization parameters, such that the audio spatialization parameters correspond to such a source-to-tympanic membrane transfer function. Alternatively, the model may be trained to estimate audio spatialization parameters such that the audio spatialization parameters correspond to other transfer functions (such as source-to-ear canal or source-to-outer ear transfer functions), however, this may reduce the quality of the audio spatialized based on the estimated audio spatialization parameters.
[0020] Ideally, and for similar reasons, the sound input position (hereinafter referred to as the "training sound input position") for training the model will be located at the tympanic membrane center of each ear. However, the shape and size of the ear canal generally have a relatively small impact on a person's directional hearing, and thus it is usually more important that the training sound input position corresponds to the target sound input position. When properly trained, the model will at least partially compensate for the sound input position offset from the corresponding tympanic membrane. That is, selecting a training sound input position far from the ear canal (such as on the lateral outside of an ear-mounted earcup) may result in a lower correlation between the spatial cues in the sound at the sound input position and the spatial cues in the sound at the corresponding tympanic membrane, and thus may also reduce the correlation between the estimated audio spatialization parameters and the true HRTF of the target user.
[0021] The model is trained by a training computer using one or more audio spatialization parameters obtained for each of a plurality of training users as the target output of the model. Any subset of the training users can be real persons for whom such audio spatialization parameters have been determined based on acoustic measurements of sounds received near, in, or at the ear canals of the respective training users and / or based on acoustic, optical, or mechanical measurements of the characteristics of the heads and / or ears of the training users. Alternatively or additionally, any subset of the training users can be simulated persons with artificially generated head and ear-related characteristics, and for such simulated persons, such audio spatialization parameters have been determined based on simulations of sound fields and / or sound propagation. The plurality of training users preferably includes real and / or simulated persons that together cover the variations of head and ear-related characteristics corresponding to the variations expected in the target users. For any one of the plurality of training users, the training computer can determine one or more audio spatialization parameters itself. Alternatively or additionally, the training computer can obtain all or some of the one or more audio spatialization parameters from one or more databases optionally stored on one or more other computers, and / or determine all or some of the one or more audio spatialization parameters by modifying other audio spatialization parameters obtained from such databases.
[0022] For each training user, the training computer obtains the respective sound input locations near, in, or at each of the ear canals of the training user (real or simulated). Each such sound input location is preferably selected within the respective ear canal, 3 mm within the entrance to the ear canal, 6 mm within the entrance to the ear canal, or 15 mm within the entrance to the ear canal. For real training users, the sound input location preferably corresponds to the location of the respective sound input to the microphone of the target device. The sound input location can be determined based on acoustic, optical, or mechanical measurements of the location of the respective sound input to the target device, which receives the sound to be analyzed by the model when the training user wears the target device. For simulated training users, the sound input location preferably corresponds to a location calculated based on the training user data (such as the head size and the size, shape, and location of each of the ears of the training user) and based on the target device data (such as the size, shape, and location of the simulated target device and / or the relative location of the respective target sound input location).
[0023] For both real and simulated training users, a smaller deviation between the training sound input location and the target sound input location can be at least partially compensated for by the training model, while a larger deviation may result in a reduced correlation between the estimated audio spatialization parameters and the true HRTF of the target user. Description of the Drawings
[0024] The above and other features and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of its exemplary embodiments with reference to the accompanying drawings, in which:
[0025] Figure 1 An exemplary audio system according to the present disclosure is schematically shown,
[0026] Figure 2 An exemplary part of the audio system according to the present disclosure is schematically shown,
[0027] Figure 3 An exemplary part of the audio system according to the present disclosure is schematically shown,
[0028] Figure 4 is a flowchart of an exemplary method according to the present disclosure, and
[0029] Figure 5 is a flowchart of an exemplary method according to the present disclosure. Detailed Embodiments
[0030] Various exemplary embodiments and details are described below with reference to the accompanying drawings (when relevant). It should be noted that the drawings may be drawn to scale or not to scale, and throughout the drawings, elements of similar structure or function are denoted by similar reference numerals. It should also be noted that the drawings are only intended to facilitate the description of the embodiments. The drawings are not intended as an exhaustive description of the present invention or a limitation on the scope of the present invention. Additionally, the illustrated embodiments need not have all the aspects or advantages shown. Aspects or advantages described in connection with a particular embodiment are not necessarily limited to that embodiment and may be practiced in any other embodiment even if not so shown or if not so explicitly described.
[0031] An audio system is disclosed. The audio system includes one or more processors optionally distributed among one or more devices of the audio system.
[0032] The audio system may include an audio device that includes a first earphone and a second earphone. The audio device may be a listening device, an audible part, a headphone, an ear protector, an earplug, a hearing aid, a group of hearing devices, or any combination thereof.
[0033] The present disclosure relates to a method, such as a method for estimating one or more audio spatialization parameters of a particular target user. The method may be performed in an audio system including one or more processors, and the method includes: obtaining (such as one or more of receiving and retrieving) audio data from a microphone disposed near, in, or at the ear canal of the target user. The audio data may have a duration of at least 10 minutes (such as at least 20 minutes). In one or more examples, the audio data has a duration of at least 30 minutes. In other words, the audio data may represent audio recorded within a time period of at least 10 minutes or at least 20 minutes (such as at least 30 minutes).
[0034] The method includes: obtaining audio data in one or more (such as a plurality of) environments, the one or more environments optionally including a normal usage environment. A normal usage environment is an environment in which an audio device is typically used, for example, in public transportation, a car, a home, or a school.
[0035] Obtaining audio data including first audio data and second audio data may optionally include: obtaining first audio data from a first microphone disposed, for example, near, in, or at the first ear canal of the target user using a target device (such as a listening device, an audible part, headphones, ear protectors, earplugs, a hearing aid, or a hearing device), and obtaining second audio data from a second microphone disposed, for example, near, in, or at the first ear canal or the second ear canal of the target user. The first microphone has a sound input at a first target sound input position, and the second microphone has a sound input at a second target sound input position. The method includes: providing one or more audio spatialization parameters and outputting one or more audio spatialization parameters. Providing one or more audio spatialization parameters may include: applying a model to the audio data for parameter estimation of the one or more audio spatialization parameters and determining the one or more audio spatialization parameters based on the parameter estimation.
[0036] In one or more examples, a method for estimating one or more audio spatialization parameters for a particular target user is disclosed, wherein the method includes: obtaining audio data including first audio data and second audio data by obtaining the first audio data from a first microphone disposed near, in, or at a first ear canal of the target user and obtaining the second audio data from a second microphone disposed near, in, or at a second ear canal of the target user; and providing one or more audio spatialization parameters, including: applying a model to the audio data for providing a parameter estimate of the one or more audio spatialization parameters and determining the one or more audio spatialization parameters based on the parameter estimate; and outputting the one or more audio spatialization parameters. The first microphone has a sound input at a first target sound input location, and the second microphone has a sound input at a second target sound input location.
[0037] In one or more examples, determining the one or more audio spatialization parameters based on the parameter estimate may include: using the parameter estimate as the one or more spatialization parameters.
[0038] In one or more examples, obtaining the audio data includes: filtering the first audio data, for example, using an echo canceller and / or a feedback suppression algorithm to remove a first speaker audio component from a first speaker disposed near, in, or at the first ear canal of the target user.
[0039] In one or more examples, obtaining the audio data includes: filtering the second audio data, for example, using an echo canceller and / or a feedback suppression algorithm to remove a second speaker audio component from a second speaker disposed near, in, or at the first ear canal or the second ear canal of the target user.
[0040] In one or more examples, the method includes: detecting, for example, the presence of the speech of the target user based on the audio data, and in response to detecting the presence of the speech, forgoing (such as stopping, pausing, and / or deactivating) obtaining the audio data. In other words, the audio including the speech of the target user is optionally omitted, not recorded, or not included in the audio data, thereby, for example, optimizing the memory and processing resources in the audio system. Thus, the speech of the target user himself / herself is excluded from the audio data, thereby improving the HRTF estimation.
[0041] In one or more examples, the method includes, for example, detecting a sound parameter (such as an intensity parameter) based on audio data, and abandoning (such as stopping, pausing, and / or deactivating) obtaining audio data according to the sound parameter meeting a sound criterion. In other words, audio or sound with insufficient quality and / or intensity to be input into the model can be omitted or not included in the audio data, thereby optimizing memory and processing resources in an audio system, for example. The method can include, for example, outputting an audio tone and / or message to a target user according to the sound parameter meeting the sound criterion. The audio tone and / or message can prompt the target user to move to a more suitable environment for recording audio data. The audio tone and / or message can indicate a poor sound environment.
[0042] In one or more examples, providing one or more audio spatialization parameters includes determining whether a parameter estimate meets a first criterion, and wherein determining one or more audio spatialization parameters is performed according to the determination of meeting the first criterion. For example, the first criterion can be designed to ensure that outliers in the parameter estimate do not form the basis for the audio spatialization parameters.
[0043] In one or more examples, determining whether a parameter estimate meets a first criterion includes, for example, determining a variance parameter based on the parameter estimate and / or historical (previous) audio spatialization parameters, and determining whether the variance parameter meets a variance criterion, such as whether the variance parameter is less than a variance threshold or whether the variance parameter is greater than a variance threshold. In one or more examples, if the variance criterion is met, the first criterion can be met or at least partially met.
[0044] In one or more examples, determining whether a parameter estimate meets a first criterion includes determining whether a quality parameter of the parameter estimate meets a threshold. In other words, if the quality parameter meets the threshold, the first criterion can be met or at least partially met.
[0045] In one or more examples, determining whether a parameter estimate meets a first criterion includes determining a time parameter and determining whether the time parameter meets a time criterion. In one or more examples, the time parameter can indicate the time since starting to record / obtain the first audio data and the second audio data. In other words, if the time criterion is met, the first criterion can be met or at least partially met. For example, the time parameter can indicate the duration of the audio data, such as the cumulative duration. The first time parameter combined with the quality parameter can allow optimizing (such as reducing) the time required to determine the audio spatialization parameters. The first time parameter combined with the quality parameter can ensure that the obtained audio data has a sufficient duration to achieve an accurate parameter estimate. For example, a smaller duration combined with high quality may be sufficient, while lower quality may require a longer duration.
[0046] In one or more examples, outputting one or more audio spatialization parameters includes: storing one or more audio spatialization parameters in a memory, and / or transmitting one or more audio spatialization parameters to an audio device such as a listening device, an audible part, headphones, ear protectors, earplugs, a hearing device.
[0047] In one or more examples, one or more audio spatialization parameters include one or more of the following: interaural intensity difference, interaural intensity difference map, compressed interaural intensity difference map, double logarithmic absolute HRTF, interaural time difference, interaural time difference map, and compressed interaural time difference map.
[0048] Parameter estimation may include one or more of the following: interaural intensity difference, interaural intensity difference map, compressed interaural intensity difference map, double logarithmic absolute HRTF, interaural time difference, interaural time difference map, and compressed interaural time difference map.
[0049] One or more audio spatialization parameters may, for example, indicate the time delay of the left ear / first filter and / or the right ear / second filter of an audio device. Alternatively or additionally, one or more audio spatialization parameters may indicate the signal gain (or attenuation) of the left ear / first filter and / or the right ear / second filter of an audio device. Either or both of the time delay and the signal gain may be specified as a scalar value to be applied to the target audio signal as a whole, a set of scalar values to be applied to different frequency ranges of the target audio signal, and / or a (discrete or continuous) transfer function to be applied to the target audio signal as a whole.
[0050] Alternatively or additionally, one or more audio spatialization parameters may indirectly indicate such values and / or functions in the form of geometric values related to the shape of the head and / or ears of the corresponding target user (such as, for example, the distance between the ears, the size of the head, the position of the ears relative to the head, the size of the outer ear, etc.).
[0051] In one or more examples, determining one or more audio spatialization parameters based on parameter estimation may include: transforming or mapping the output of a model to one or more spatialization parameters to be used in an audio device.
[0052] In one or more examples, the model is a neural network configured to receive a first complex spectrogram based on first audio data as a first input and a second complex spectrogram based on second audio data as a second input, the neural network being configured to provide an output that includes a parameter estimate based on the first complex spectrogram and the second complex spectrogram. In other words, the first audio data can be represented or transformed into the first complex spectrogram, and the second audio data can be represented or transformed into the second complex spectrogram. In one or more examples, a neural network, such as a dilated convolutional neural network, can be configured to take as input direct audio samples having a duration of, for example, at least 10 minutes (such as 30 minutes).
[0053] In one or more examples, the neural network is a dilated convolutional neural network that serves to reduce the computational effort required by the neural network.
[0054] In addition, a method performed by an electronic device for providing spatialized audio to a specific target user is disclosed, wherein the method includes: performing a method disclosed herein, such as a method for estimating one or more audio spatialization parameters of a specific target user; and providing an audio output based on the one or more audio spatialization parameters.
[0055] In addition, an electronic device, such as an audio device, is disclosed, the electronic device including one or more processors, wherein the electronic device / one or more processors are configured to perform any of the methods described herein. The electronic device can be the target device.
[0056] The audio device optionally includes one or more speakers, also referred to as loudspeakers or receivers, for outputting audio to the first ear canal and / or the second ear canal of the target user. The audio device can include a first earbud configured to be disposed in, near, or at the first ear canal of the target user. In one or more examples, the first earbud includes a first microphone and optionally a first speaker. The audio device can include a second earbud configured to be disposed in, near, or at the second ear canal of the target user. In one or more examples, the second earbud includes a second microphone and optionally a second speaker.
[0057] Disclosed is an audio system, which includes one or more processors. Wherein, the one or more processors are configured to: obtain audio data including first audio data and second audio data by obtaining the first audio data from a first microphone disposed in or at a first ear canal of a target user and obtaining the second audio data from a second microphone disposed in or at a second ear canal of the target user; provide one or more audio spatialization parameters by applying a model to the audio data for providing a parameter estimation of the one or more audio spatialization parameters and determining the one or more audio spatialization parameters based on the parameter estimation; and output the one or more audio spatialization parameters. The first microphone has a sound input at a first target sound input position, and the second microphone has a sound input at a second target sound input position.
[0058] In one or more exemplary audio systems, a model (such as a neural network) implemented by the one or more processors is configured to estimate or determine a parameter estimation based on the audio data.
[0059] The audio system may include an audio device and an accessory device. The one or more processors may be part of the audio device, part of the accessory device (such as a mobile phone, a tablet computer, a personal computer), or part of a server device. The one or more processors may be distributed between the audio device and one or more accessory devices and / or between the accessory devices. The one or more processors may be part of a target device.
[0060] In one or more exemplary audio systems, the one or more processors include or implement a pre-processor, which is configured to pre-process the audio data and provide a neural network input to the neural network based on the audio data (such as the first audio data and the second audio data). The pre-processor may be disposed in the audio device and / or the accessory device.
[0061] In one or more exemplary audio systems, the pre-processor is configured to determine a first real spectrogram of the first audio data, which is also denoted as P_R_1, and a first virtual spectrogram P_I_1 of the first audio data, and provide the first real spectrogram P_R_1 and the first virtual spectrogram P_I_1 in an input to the model (such as a neural network input). In one or more exemplary audio systems, the pre-processor is configured to determine a second real spectrogram of the second audio data, which is also denoted as P_R_2, and a second virtual spectrogram P_I_2 of the second audio data, and provide the second real spectrogram P_R_2 and the second virtual spectrogram P_I_2 in an input to the model (such as a neural network input).
[0062] The audio device preferably includes an A / D converter or a plurality of A / D converters, such as one A / D converter per microphone, which is used to digitize the audio signals from the respective microphones for providing audio data. In some exemplary audio systems, A / D conversion can be performed in the first earbud and the second earbud.
[0063] Determining the real and virtual spectrograms can include: sampling the audio data at a sampling rate greater than 8 kHz (such as at least 16 kHz, such as in the range from 6 kHz to 40 kHz, such as 32 kHz).
[0064] The real and virtual spectrograms can each include at least 128 values (such as 256 values for the respective 256 frequency bands or frequency ranges), and can be based on a Hahn window with a frame size of 512 samples, for example, a hop size with 256 samples and / or 50% overlap.
[0065] The neural network input can include K real and virtual spectrograms for the audio data, where K is selected such that the spectrograms represent at least 10 minutes (such as at least 20 minutes) of audio. In other words, the neural network input can include K_1 real spectrograms and K_1 virtual spectrograms of the first audio data and K_2 real spectrograms and K_2 virtual spectrograms of the second audio data, where K_1 and K_2 are preferably equal to K. In one or more examples, the audio data can be represented by the real and virtual spectrograms. Thus, the audio data can be represented in a substantially lossless representation form, which advantageously preserves both the phase and the amplitude, both of which are important parameters for determining the audio spatialization parameters.
[0066] In one or more exemplary audio systems, one or more processors include or implement an ASP controller, where the ASP controller is configured to determine one or more audio spatialization parameters (ASP) based on the parameter estimates (PE) output from the model. The ASP controller can be used as a post - processor, which is configured to post - process the neural network output (such as the parameter estimates from the neural network) and provide the ASP based on the neural network output. In one or more exemplary audio systems, the parameter estimates can be used as the ASP.
[0067] The ASP controller can be configured to determine whether a parameter estimate is a valid estimate, e.g., whether the parameter estimate meets a first criterion, such as that the parameter estimate is a reasonable estimate for a real person, and / or that the parameter estimate is a reasonable estimate for a target user based on other data obtained for the target user, and wherein determining one or more audio spatialization parameters and / or outputting one or more audio spatialization parameters ASP is performed based on determining the parameter estimate to be valid (such as based on determining that the first criterion is met). Thereby, improved ASP determination is provided by avoiding using inaccurate, incorrect, or otherwise invalid parameter estimates to determine the ASP. Optionally, determining whether the parameter estimate is valid (such as meeting the first criterion) includes: determining whether a quality parameter of the parameter estimate meets a threshold. In other words, the ASP controller can be configured to determine a quality parameter indicative of the quality of the parameter estimate, and determine the ASP based on the quality of the parameter estimate being good enough (e.g., whether the quality parameter meets the threshold).
[0068] Optionally, determining whether the parameter estimate is valid (such as meeting the first criterion) includes: determining a time parameter based on, e.g., audio data, and determining whether the time parameter meets a time criterion. In other words, the ASP controller can be configured to determine whether the audio data has a sufficient time span or duration. In one or more examples, the quality parameter and the time parameter can be combined to determine whether the parameter estimate is valid. The time parameter can indicate the time since the start of recording or the start of the audio data. The time parameter can indicate the duration of the audio data.
[0069] Optionally, determining whether the parameter estimate is valid (such as meeting the first criterion) includes: determining whether the parameter estimate corresponds to one or more ASPs that may represent a real person. For example, an interaural time difference of more than 1.5 ms only occurs when the distance between the two ears of the target user exceeds about 0.5 m, which is clearly not applicable to anyone.
[0070] In one or more exemplary audio systems, the neural network of the model is a deep neural network that includes multiple components, such as one or more convolutional neural networks (CNNs), one or more dense layers, one or more transformers, and / or one or more recurrent networks, such as long short-term memory (LSTM) recurrent networks. The convolutional neural network can include an input, J layers, and an output. The J layers can include: J_C convolutional layers represented as CONV_j (j = 1, 2, ……, J_C), where the J_C convolutional layers include a first convolutional layer CONV_1 and a second convolutional layer CONV_2; and J_O output layers, where the J_O output layers include a first output layer OUT_1 and optionally a second output layer OUT_2. The first output layer OUT_1 can be a fully connected output layer, and / or the second output layer OUT_2 can be a fully connected output layer. The J_O output layers are preferably located after the J_C convolutional layers.
[0071] The input to the neural network can include complex spectrograms, such as the real and virtual spectrograms of audio data AD_1 and AD_2. The output of the model (such as the neural network) includes parameter estimates PE. The parameter estimates optionally include one or more first parameter estimates associated with a first audio spatialization parameter. The first audio spatialization parameter can be or include a first HRTF or indicate a first HRTF associated with the first ear or first ear canal of the target user. In one or more examples, the first audio spatialization parameter ASP_1 includes one or more coefficients or filter settings of a filter for implementing the first HRTF. The output of the neural network can include an ILD map and / or a double-log absolute HRTF. The output of the neural network can include an ITD map or a compressed version thereof.
[0072] The parameter estimates optionally include one or more second parameter estimates associated with a second audio spatialization parameter. The second audio spatialization parameter can be or include a second HRTF or indicate a second HRTF associated with the second ear or second ear canal of the target user. In one or more examples, the second audio spatialization parameter ASP_2 includes one or more coefficients or filter settings of a filter for implementing the second HRTF.
[0073] In one or more exemplary convolutional neural networks, the number of convolutional layers ranges from 5 to 15 (such as from 10 to 12). One or more of the convolutional layers can have a kernel size of 3×3. One or more of the convolutional layers can have a stride of 2,1. One or more of the convolutional layers can have a stride of 1,1. One or more of the convolutional layers can have a dilation of 1,2. One or more of the convolutional layers can have a dilation of 1,4. One or more of the convolutional layers can have a dilation of 1,8.
[0074] The number of layers, such as the total number of layers in a neural network, the number of convolutional layers J_C, and / or the number of output layers J_O, can be varied to improve performance and / or reduce power consumption. This also applies to the characteristics of audio data (such as sampling rate, number of spectrogram segments, frame size, window overlap and length of the spectrogram buffer), and to the characteristics of individual network layers (such as kernel size, stride, and dilation). Generally, for a larger number of microphones (such as for an implementation with a plurality of first microphones disposed at a first ear canal and / or a plurality of second microphones disposed at a second ear canal), a larger neural network will be required.
[0075] Note that the description of the audio system herein is also applicable to a corresponding method for estimating one or more audio spatialization parameters of a particular user, and vice versa.
[0076] In addition, a computer-implemented method is provided, which is used to train a machine learning model (such as a neural network) to process audio data including first audio data indicating a first audio signal from a first microphone and second audio data indicating a second audio signal from a second microphone as input, and provide a parameter estimate of one or more audio spatialization parameters as output. Wherein, the method includes: obtaining, by a computer, one or more audio spatialization parameters for each of a plurality of training users; obtaining, by a computer, a first sound input position near, in, or at a first ear canal (such as a left ear canal) of a training user and a second sound input position near, in, or at a second ear canal (such as a right ear canal) of the training user for each training user; obtaining, by a computer, a plurality of training input sets for each training user, wherein each training input set represents a specific sound environment and a specific time period, and includes a first audio input signal and a second audio input signal, and the first audio input signal and the second audio input signal each represent ambient sounds from a specific sound environment at a specific time period at the first sound input position and the second sound input position respectively; and training the machine learning model by performing multiple rounds of training across the plurality of training input sets obtained for the plurality of training users by a computer, wherein each round of training includes: applying the machine learning model to one of the plurality of training input sets obtained for the corresponding training user, and adjusting the parameters of the machine learning model, such as weights or other parameters, using the one or more audio spatialization parameters obtained for the corresponding training user as the target output of the machine learning model.
[0077] The one or more audio spatialization parameters of the training user may include the HRTF of each training user. The HRTF of each training user can be carefully measured in an audio laboratory. The HRTF of each training user may include a diffuse-field calibrated HRTF.
[0078] Audio data can be input into a machine learning model, such as a neural network, in different formats. For example, an audio signal can be input into a machine learning model, such as a complex neural network, such as a real and virtual spectrogram or other suitable representation as described herein.
[0079] Figure 1 A block diagram of an exemplary audio system is shown. The audio system 2 includes an audio device 4, an accessory device 6, and a server device 8. The audio device 4 is configured to communicate wirelessly (as shown) or wired with the accessory device 6 and transmit audio data AD, which includes first audio data AD_1 and second audio data AD_2, to the accessory device. The audio device 4 is shown as a set of earbuds, and the set of earbuds includes a first earbud 4A and a second earbud 4B. The first earbud 4A includes a first microphone 10A configured to be disposed in or at a first ear canal of a target user and provide the first audio data AD_1, and the second earbud 4B includes a second microphone 10B configured to be disposed in or at a second ear canal of the target user and provide the second audio data AD_2. The accessory device 6 is configured to communicate wirelessly 7 with the server device 8 via a network 9.
[0080] One or more processors of the accessory device 6 and / or the server device 8 are configured to provide one or more audio spatialization parameters, where providing one or more audio spatialization parameters includes: applying a model (such as a neural network) to the audio data AD for providing a parameter estimate PE of one or more audio spatialization parameters ASP, and determining one or more audio spatialization parameters ASP, such as ASP_1 and ASP_2, based on the parameter estimate. One or more processors of the accessory device 6 are configured to output one or more audio spatialization parameters ASP, where outputting one or more audio spatialization parameters ASP includes: transmitting the ASP (such as ASP_1 and ASP_2) to the audio device 4. As Figure 1As shown, the ASP includes a first ASP (denoted as ASP_1) for the first earbud 4A and a second ASP (denoted as ASP_2) for the second earbud 4B. The earbuds 4A, 4B store and apply ASP_1 and ASP_2 by providing corresponding audio outputs based on ASP_1 and ASP_2 respectively. In other words, after determining or receiving the ASP, the first earbud 4A applies ASP_1 and provides a first audio output based on ASP_1, and the second earbud 4B applies ASP_2 and provides a second audio output based on ASP_2. For example, the first earbud 4A and the second earbud 4B may receive a first input audio signal and a second input audio signal respectively for playback to a target user, and apply ASP_1 and ASP_2 to the received input audio signals to provide a corresponding first spatialized audio output signal and a second spatialized audio output signal to the target user. In this way, the audio device 4 can spatialize an input audio signal received from, for example, the accessory device 6 based on one or more audio spatialization parameters determined according to the first audio data AD_1 and the second audio data AD_2 provided by the first microphone 10A and the second microphone 10B respectively at an earlier time, to provide a first spatialized audio output signal and a second spatialized audio output signal, whereby the first spatialized audio output signal and the second spatialized audio output signal can be specifically adapted to the target user.
[0081] Figure 2 Part of the exemplary audio system 2 is shown in more detail. One or more processors 12 of the audio system 2 obtain the audio data AD_1 and AD_2 from the corresponding microphones 10A and 10B, and include a neural network module 14 having a neural network that implements a model for providing a parameter estimation PE for one or more audio spatialization parameters. The parameter estimation PE is fed to an ASP controller 16 of the one or more processors 12, and the ASP controller 16 is configured to determine one or more audio spatialization parameters ASP based on the parameter estimation PE and output one or more audio spatialization parameters ASP.
[0082] Figure 3More particularly, portions of an exemplary audio system 2A are shown. One or more processors 12 of the audio system 2A include a preprocessor block 18 configured to preprocess audio data AD_1 and AD_2 for providing a neural network input 18A to a neural network module 14 based on the audio data AD_1 and AD_2. In the illustrated audio system 2A, the preprocessor block 18 is configured to determine a first real spectrogram, also denoted as P_R_1, and a first virtual spectrogram P_I_1 of the first audio data AD_1 of the first microphone 10A, and a second real spectrogram P_R_2 and a second virtual spectrogram P_I_2 of the second audio data AD_2 of the second microphone 10B, and provide the real spectrograms P_R_1, P_R_2 and the virtual spectrograms P_I_1, P_I_2 as the neural network input 18A.
[0083] The preprocessor block 18 is optionally configured to compensate for or remove audio components from corresponding speakers or loudspeakers near the microphones 10A, 10B. For example, the preprocessor block 18 may be configured to filter the first audio data AD_1 to remove, for example, a first speaker audio component from a first speaker (such as the first speaker of the first earbud 4A) disposed in or at a first ear canal of a target user, using, for example, an echo canceller / DFS algorithm. The preprocessor block 18 may be configured to filter the second audio data AD_2 to remove, for example, a second speaker audio component from a second speaker (such as the second speaker of the second earbud 4B) disposed in or at a second ear canal of the target user, using, for example, an echo canceller / DFS algorithm. The preprocessor block 18 may be configured to align or synchronize the first audio data AD_1 and the second audio data AD_2.
[0084] Figure 4A flowchart of an exemplary method for estimating one or more audio spatialization parameters for a specific target user is shown. Method 100 includes: obtaining S102 audio data including first audio data and second audio data having a duration of at least 30 minutes, for example. Obtaining S102 the audio data includes: obtaining S102A the first audio data from a first microphone disposed near, in, or at a first ear canal of the target user, and obtaining S102B the second audio data from a second microphone disposed near, in, or at a second ear canal of the target user. Method 100 continues by providing S104 one or more audio spatialization parameters, wherein providing S104 the one or more audio spatialization parameters includes: applying S104A a model to the audio data, for example using an audio device or an accessory device, for providing a parameter estimate of the one or more audio spatialization parameters; determining S104B the one or more audio spatialization parameters based on the parameter estimate; and outputting S104C the one or more audio spatialization parameters.
[0085] Figure 5 A flowchart of an exemplary method for training a machine learning model (such as a neural network, for example a neural network) for determining or providing a parameter estimate of one or more audio spatialization parameters is shown. Method 200 is a computer-implemented method for training a machine learning model (such as, a neural network, for example as a CNN) of an audio system (such as audio system 2, 2A) to process audio data including first audio data indicating a first audio signal from a first microphone and second audio data indicating a second audio signal from a second microphone as input, and providing a parameter estimate of one or more audio spatialization parameters (such as HRTF, other transfer functions, interaural differences (time and / or intensity), etc.) related to a specific target user of the audio system as output.
[0086] Method 200 includes: obtaining S204, by a computer, one or more audio spatialization parameters ASP for each of a plurality of training users; obtaining S206, by the computer, for each training user a first sound input position near, in, or at a first ear canal (such as a left ear canal) of the training user and a second sound input position near, in, or at a second ear canal (such as a right ear canal) of the training user; obtaining S208, by the computer, for each training user a plurality of training input sets, where each training input set represents a particular acoustic environment and a particular time period and includes a first audio input signal and a second audio input signal, the first audio input signal and the second audio input signal each representing ambient sound from the particular acoustic environment at the particular time period at the first sound input position and the second sound input position, respectively; and training S210, by the computer executing, a machine learning model over multiple rounds of training across the plurality of training input sets obtained for the plurality of training users, where each round of training includes: applying S210B the machine learning model to one of the plurality of training input sets obtained for the corresponding training user, and adjusting S210C parameters of the machine learning model, such as weights or other parameters, using the one or more audio spatialization parameters obtained for the corresponding training user as the target output of the machine learning model.
[0087] Method 200 optionally includes: applying S212 a machine learning (ML) model or neural network in an audio system, such as by storing model parameters of the ML model in an ML model module (such as neural network module 10 of audio systems 2, 2A).
[0088] In method 200, some of the method steps required to obtain S208 the training input sets (such as steps that produce the same result for all rounds of training) may be performed before training (i.e., before the first round of training is executed), while other steps for obtaining S208 the training input may be performed during training S210, such as interleaved with and / or during each round of training. Obviously, avoiding repeated calculations can save energy and time.
[0089] The use of terms such as "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. does not imply any particular order but is included to identify individual elements. Further, the use of terms such as "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. does not denote any order or importance, but rather the terms "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. are used to distinguish one element from another. Note that the words "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. are used herein and elsewhere for labeling purposes only and are not intended to denote any specific spatial or temporal ordering.
[0090] The memory can be one or more of the following: a buffer, flash memory, a hard disk drive, a removable medium, volatile memory, non-volatile memory, random access memory (RAM), or other suitable devices. In a typical arrangement, the memory can include non-volatile memory for long-term data storage and volatile memory that serves as the system memory for the processor. The memory can exchange data with the processor via a data bus. The memory can be considered a non-transitory computer-readable medium.
[0091] The memory can be configured to store information (such as information indicating a neural network, such as configuration and parameters, such as weights or other parameters thereof) in a portion of the memory.
[0092] Further, the labeling of a first element does not imply the existence of a second element, and vice versa.
[0093] It will be appreciated that the figures include some modules or operations shown in solid lines, and some modules or operations shown in dashed lines. The modules or operations included in the solid lines are the modules or operations included in the broadest exemplary embodiment. The modules or operations included in the dashed lines are exemplary embodiments that may be included in or a part of the modules or operations of the solid-line exemplary embodiment, or additional modules or operations that may be taken in addition to the modules or operations of the solid-line exemplary embodiment. It should be understood that these operations need not be performed in the order presented. Further, it should be understood that not all operations need to be performed. The example operations can be performed in any order and in any combination.
[0094] It should be noted that the word "comprising" does not necessarily exclude the existence of other elements or steps than those listed.
[0095] It should be noted that the word "a" or "an" before an element does not exclude the existence of a plurality of such elements.
[0096] It should also be noted that any reference signs do not limit the scope of the claims, and the exemplary embodiments can be implemented at least in part by both hardware and software, and several "tools", "units" or "devices" can be represented by the same piece of hardware.
[0097] The various exemplary methods, devices, and systems described herein are described in the general context of method steps processes. In one aspect, these method steps processes can be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions executed by a computer in a networked environment, such as program code. The computer-readable medium can include removable and non-removable storage devices, including but not limited to read-only memory (ROM), random access memory (RAM), compact discs (CDs), digital versatile discs (DVDs), etc. Generally, program modules can include routines, programs, objects, components, data structures, etc. that perform specified tasks or implement specific abstract data types. The computer-executable instructions, associated data structures, and program modules represent examples of program code for performing the steps of the methods disclosed herein. The specific sequences of such executable instructions or associated data structures represent examples of corresponding actions for implementing the functions described in such steps or processes.
[0098] Although the features have been shown and described, it should be understood that they are not intended to limit the claimed invention, and it will be apparent to those skilled in the art that various changes and modifications can be made without departing from the spirit and scope of the claimed invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive. The claimed invention is intended to cover all alternatives, modifications, and equivalents.
[0099] List of Reference Signs
[0100] 2 Audio System
[0101] 4 Audio Device
[0102] 4A First Earbud
[0103] 4B Second Earbud
[0104] 6 Accessory Device
[0105] 6A Smart Phone
[0106] 7 Wireless Communication
[0107] 8 Server Device
[0108] 9 Network
[0109] 10A First Microphone
[0110] 10B Second Microphone
[0111] 12 One or more processors
[0112] 14 Neural network module
[0113] 16 ASP controller
[0114] 18 Pre - processor block
[0115] 18A Neural network input
[0116] 100 Method for estimating one or more audio spatialization parameters for a specific target user
[0117] S102 Obtain audio data including first audio data and second audio data
[0118] S102A Obtain first audio data from a first microphone disposed near, in, or at a first ear canal of a target user
[0119] S102B Obtain second audio data from a second microphone disposed near, in, or at a second ear canal of a target user
[0120] S104 Provide one or more audio spatialization parameters
[0121] S104A Apply a model to the audio data for parameter estimation for providing one or more audio spatialization parameters;
[0122] S104B Determine one or more audio spatialization parameters based on the parameter estimation
[0123] S104C Output one or more audio spatialization parameters
[0124] 200 Method for training a machine - learning model (such as a neural network)
[0125] S204 Determine one or more audio spatialization parameters ASP for each training user
[0126] S206 Determine the sound input location
[0127] S208 Obtain multiple training input sets for each training user
[0128] S210 Train a machine - learning model
[0129] S210A Perform multiple rounds of training by a computer across multiple training input sets obtained for multiple training users
[0130] 210B Apply the machine - learning model to multiple training input sets obtained for multiple training users
[0131] S210C adjusts the parameters of a machine learning model, such as weights or other parameters, using one or more audio spatialization parameters determined for a corresponding training user as the target output of the machine learning model.
[0132] S212 Apply the machine learning model
[0133] AD Audio data
[0134] AD_1 First audio data
[0135] AD_2 Second audio data
[0136] ASP Audio spatialization parameter
[0137] ASP_1 First audio spatialization parameter
[0138] ASP_2 Second audio spatialization parameter
[0139] PE Parameter estimation.
Claims
1. A method for estimating one or more audio spatialization parameters for a specific target user, wherein: The method comprises: obtaining audio data including the first audio data and the second audio data by obtaining first audio data from a first microphone disposed near, in, or at a first ear canal of the target user and obtaining second audio data from a second microphone disposed near, in, or at a second ear canal of the target user; and Providing the one or more audio spatialization parameters comprises: applying a model to the audio data for providing parameter estimates of the one or more audio spatialization parameters; determining the one or more audio spatialization parameters based on the parameter estimate; and The one or more audio spatialization parameters are output.
2. The method according to claim 1, wherein: Obtaining audio data includes filtering the first audio data to remove a first speaker audio component from a first speaker disposed in or at the first ear canal of the target user.
3. The method according to any one of claims 1 to 2, wherein: Obtaining audio data includes filtering the second audio data to remove a second speaker audio component from a second speaker disposed in or at the second ear canal of the target user.
4. The method according to any one of claims 1 to 3, wherein: The method includes: detecting the presence of the target user's voice, and abandoning obtaining audio data based on detecting the presence of the voice.
5. The method according to any one of claims 1 to 4, wherein: Providing the one or more audio spatialization parameters comprises determining whether the parameter estimate satisfies a first criterion, and wherein determining the one or more audio spatialization parameters is performed in dependence on determining that the first criterion is satisfied.
6. The method according to claim 5, wherein: Determining whether the parameter estimate meets a first criterion includes determining whether a quality parameter of the parameter estimate meets a threshold.
7. The method according to any one of claims 5 to 6, wherein: Determining whether the parameter estimate meets a first criterion includes determining a time parameter, and determining whether the time parameter meets a time criterion.
8. The method according to any one of claims 1 to 7, wherein: Outputting the one or more audio spatialization parameters includes storing the one or more audio spatialization parameters and / or transmitting the one or more audio spatialization parameters.
9. The method according to any one of claims 1 to 8, wherein: The one or more audio spatialization parameters include one or more of: interaural intensity difference, interaural intensity difference map, compressed interaural intensity difference map, double log absolute HRTF, interaural time difference, interaural time difference map, and compressed interaural time difference map.
10. The method according to any one of claims 1 to 9, wherein: The model is a neural network configured to receive a first complex spectrogram based on the first audio data as a first input and a second complex spectrogram based on the second audio data as a second input, and the neural network is configured to provide an output, the output comprising the parameter estimates based on the first complex spectrogram and the second complex spectrogram.
11. The method according to claim 10, wherein: The neural network is a dilated convolutional neural network.
12. A method performed by an electronic device for providing spatialized audio to a specific target user, wherein: The method comprises: performing a method according to any one of claims 1 to 11; and providing an audio output based on one or more audio spatialization parameters.
13. An electronic device comprising one or more processors, wherein: The electronic device is configured to perform any of the methods according to any one of claims 1 to 12.
14. An audio system comprising one or more processors, wherein: The one or more processors are configured to: obtaining audio data including the first audio data and the second audio data by obtaining first audio data from a first microphone disposed in or at a first ear canal of a target user and obtaining second audio data from a second microphone disposed in or at a second ear canal of the target user; Provide one or more audio spatialization parameters by doing the following: applying a model to the audio data for providing parameter estimates of the one or more audio spatialization parameters; determining the one or more audio spatialization parameters based on the parameter estimates; and The one or more audio spatialization parameters are output.
15. A computer-implemented method for training a machine learning model to process as input audio data comprising first audio data indicative of a first audio signal from a first microphone and second audio data indicative of a second audio signal from a second microphone, and to provide as output parameter estimates of one or more audio spatialization parameters, wherein: The method comprises: obtaining, by a computer, one or more audio spatialization parameters for each training user of a plurality of training users; Obtaining, by a computer, for each training user, a first sound input position near, in, or at a first ear canal of the training user and a second sound input position near, in, or at a second ear canal of the training user; Obtaining a plurality of training input sets for each training user by means of a computer, wherein each training input set represents a specific sound environment and a specific time period, and includes a first audio input signal and a second audio input signal, wherein the first audio input signal and the second audio input signal each represent an ambient sound from the specific sound environment at the first sound input position and the second sound input position, respectively, within the specific time period; and The machine learning model is trained by performing, by a computer, multiple rounds of training across the multiple training input sets obtained for the multiple training users, wherein each round of training includes: applying the machine learning model to one of the multiple training input sets obtained for the corresponding training user, and adjusting parameters of the machine learning model, such as weights or other parameters, using the one or more audio spatialization parameters obtained for the corresponding training user as a target output of the machine learning model.