HRTF determination using head-mounted viewer and in-ear device

By using head-mounted devices and in-ear devices in virtual reality and augmented reality systems to collect audio signals in natural environments and estimate the location of sound sources, the time-consuming and expensive problems of traditional methods are solved, and the technical application of personalized HRTF to users is realized, which improves the rendering accuracy and immersiveness of spatial sound.

CN120693884APending Publication Date: 2025-09-23CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480012766.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2024-03-07
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies make it difficult to personalize head-related transfer functions (HRTFs) for different users in virtual reality and augmented reality systems, resulting in a degraded auditory experience. Traditional measurement methods are time-consuming and expensive, making them difficult to scale to a large number of users.

Method used

By using head-mounted devices and in-ear devices worn by users, audio signals are collected in a natural environment, the location of the sound source is estimated, and personalized HRTF is determined based on the signals collected by these devices. A user-specific HRTF data point cloud is gradually constructed, reducing user participation and dependence on specialized measurement systems.

Benefits of technology

It achieves efficient and low-cost determination of personalized HRTF in the user environment, improves the accuracy and immersion of spatial sound rendering, and is suitable for a large number of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120693884A_ABST
    Figure CN120693884A_ABST
Patent Text Reader

Abstract

Techniques for determining a personalized head-related transfer function (HRTF) using a head-mounted device and an in-ear device include receiving a first sound signal from a sensor array of the head-mounted device, the first sound signal associated with sound from a sound source in a local environment of a user of the head-mounted device; determining, based on the first sound signal, that reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determining that the sound source is stationary within a period of time; determining a relative position of the sound source relative to the user; receiving a second sound signal from an in-ear device in the ear of the user, the second sound signal being associated with sound from the sound source; and determining an HRTF or one or more parameters of the HRTF of the user associated with the relative position of the sound source based at least on the second sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 488,895, filed on March 7, 2023, entitled “HRTF DETERMINATION USING AHEADSET AND IN-EAR DEVICES.” Technical Field

[0003] The present disclosure generally relates to determining head-related transfer functions (HRTFs), and more particularly, to determining HRTFs or HRTF parameters using head-mounted devices (e.g., headsets) and in-ear devices. Various inventive embodiments, including devices, systems, methods, structures, and processes, are described herein. Background Art

[0004] Artificial reality systems such as head-mounted display (HMD) or heads-up display (HUD) systems typically include a near-eye display system in the form of a head-mounted viewer or a pair of glasses, and are configured to present content to a user through an electronic or optical display, for example, within about 10 mm to 20 mm in front of the user's eyes. As in virtual reality (VR) applications, augmented reality (AR) applications, or mixed reality (MR) applications, the near-eye display system can display virtual objects, or combine images of real objects with virtual objects. The near-eye display typically includes an optical system that is configured to form an image of a computer-generated image displayed by an image source (e.g., a display panel). For example, the optical system can pass an image generated by the image source to create a virtual image that appears to be more than a few centimeters away from the user's eyes. In addition to displaying virtual images at a target image plane, AR / VR systems may also require spatial sound or three-dimensional (3D) sound rendering, allowing users to perceive the sounds of virtual objects as if they were originating from the virtual objects' target locations. This enhances the immersive user experience and contributes to the successful implementation of VR / AR systems. Personalized transfer functions that describe how sound interacts with the user's head and torso before reaching their ear canals can be used to render high-fidelity spatial sound. Summary of the Invention

[0005] According to a first aspect, a method is provided, the method comprising: receiving a first sound signal from a sensor array of a head-mounted device, the first sound signal being associated with sound from a sound source in a local environment of a user of the head-mounted device; determining, based on the first sound signal, that reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determining that the sound source is stationary over a period of time; determining a relative position of the sound source relative to the user; receiving a second sound signal from an in-ear device in an ear of the user, the second sound signal being associated with sound from the sound source; and determining, based at least on the second sound signal, a head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF.

[0006] Determining the relative position of the sound source relative to the user may include determining an azimuth angle of the sound source relative to the user, an elevation angle of the sound source relative to the user, or a combination thereof.

[0007] Determining the relative position of the sound source relative to the user may include: determining the direction of arrival of the sound based on a first sound signal from the sensor array and the positions of two or more sensors in the sensor array; determining the relative position of the sound source relative to the user based on images captured by one or more cameras on the head-mounted device; or a combination thereof.

[0008] Determining the relative position of the sound source with respect to the user may include determining a confidence level for the determined relative position of the sound source with respect to the user.

[0009] Determining the HRTF of the user associated with the relative position of the sound source or one or more parameters of the HRTF may include: determining a reference sound signal based on the first sound signal and the determined relative position of the sound source; and determining the HRTF or one or more parameters of the HRTF based on a spectrum of the reference sound signal and a spectrum of the second sound signal;

[0010] Determining the reference sound signal may include performing beamforming in a direction of a relative position of the sound source based on the first sound signal.

[0011] The method may also include determining a relative position of the user's torso with respect to the user's head based on data from one or more position sensors of the head-mounted device.

[0012] The method may further include saving the HRTF or one or more parameters of the HRTF and the relative position of the sound source in a data repository, the data repository storing a plurality of HRTFs of the user.

[0013] The reverberation characteristics and spectral characteristics of sound may include signal-to-noise ratio, frequency range, reverberation level, reverberation time, or a combination thereof.

[0014] The method may further include generating a model or a lookup table for mapping the relative positions of the sound sources to one or more parameters of the HRTF.

[0015] The one or more parameters of the HRTF may include a frequency scaling factor or parameters of one or more filters used to implement the HRTF.

[0016] The method may further include repeatedly performing the operations of the method according to claim 1 to determine a plurality of HRTFs or parameters of the plurality of HRTFs associated with a plurality of sound source directions relative to the user.

[0017] This time period may be greater than 10 milliseconds.

[0018] According to a second aspect, a system is provided, comprising: an in-ear device configured to generate a first sound signal associated with a sound from a sound source in a user's local environment; and a head-mounted device comprising: a sensor array configured to generate a second sound signal associated with the sound; and an audio controller configured to: determine, based on the second sound signal, that the reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determine that the sound source is stationary over a period of time; determine the relative position of the sound source relative to the user; and determine, based at least on the first sound signal, a head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF.

[0019] The audio controller may be configured to determine an azimuth angle of the sound source relative to the user, an elevation angle of the sound source relative to the user, or a combination thereof.

[0020] The audio controller can be configured to determine the relative position of the sound source relative to the user by performing an operation including the following steps: determining the direction of arrival of the sound based on a first sound signal from the sensor array and the positions of two or more sensors in the sensor array; determining the relative position of the sound source relative to the user based on images captured by one or more cameras on the head-mounted device; or a combination thereof.

[0021] The audio controller can be configured to determine the HRTF or one or more parameters of the HRTF associated with the user's relative position to the sound source by performing an operation including the following steps: determining a reference sound signal based on the first sound signal and the determined relative position of the sound source; and determining the HRTF or one or more parameters of the HRTF based on the spectrum of the reference sound signal and the spectrum of the second sound signal.

[0022] The reference sound signal may be a sound signal at the center of the user's head determined by performing beamforming in a direction of a relative position of a sound source based on the second sound signal.

[0023] The one or more parameters of the HRTF may include a frequency scaling factor or parameters of one or more filters used to implement the HRTF.

[0024] According to a third aspect, a system is provided, which includes: one or more processors; and one or more processor-readable media, which store multiple instructions, which, when executed by the one or more processors, enable the one or more processors to: receive a first sound signal from a sensor array of a head-mounted device, the first sound signal being associated with sound from a sound source in a local environment of a user of the head-mounted device; determine based on the first sound signal that the reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determine that the sound source is stationary over a period of time; determine the relative position of the sound source relative to the user; receive a second sound signal from an in-ear device in the user's ear, the second sound signal being associated with sound from the sound source; and determine, based at least on the second sound signal, a head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF.

[0025] The foregoing features and examples, as well as other features and examples, are described in more detail below in the following specification, claims, and drawings. It will be appreciated that any feature described herein as suitable for incorporation into one or more aspects is intended to be universal across any and all aspects described herein. The foregoing general overview and the following detailed description are exemplary and illustrative only and are not limiting of the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Illustrative embodiments are described in detail below with reference to the following drawings.

[0027] Figure 1 is a perspective view of an example of a near-eye display in the form of a pair of glasses for implementing some of the multiple examples disclosed herein.

[0028] Figure 2 is a perspective view of an example of a near-eye display in the form of a head-mounted display (HMD) device for implementing some of the examples disclosed herein.

[0029] Figure 3 is a block diagram of an example of an audio system in a near-eye display according to certain embodiments.

[0030] Figure 4A and Figure 4B Shows the spatial coordinates of the sound source relative to the center of the user's head.

[0031] Figure 5A An example of measuring a user's head-related transfer function (HRTF) is shown.

[0032] Figure 5B An example of generating spatialized audio content based on a user's HRTF is shown.

[0033] Figure 6A An example of a system for measuring a user's HRTF is shown.

[0034] Figure 6B An example of a system for determining a personalized HRTF or parameters of a personalized HRTF using the techniques disclosed herein is shown in accordance with some embodiments.

[0035] Figure 7 Included is a flowchart illustrating an example of a process for determining an HRTF (or parameters of an HRTF) for a user using a head-mounted device and an in-ear device, according to some embodiments.

[0036] Figure 8 An example of a process for building a personalized set of HRTFs for a user using a system including a head-mounted device and in-ear devices is shown in accordance with some embodiments.

[0037] Figure 9 is a block diagram of an example of a sound filter subsystem in an audio system of a head mounted device according to some embodiments.

[0038] Figure 10 is a functional block diagram of an example of an audio time and level difference renderer (TLDR) for processing a single-channel input audio signal to generate spatialized audio content for multiple channels, according to some embodiments.

[0039] Figure 11 An example of an implementation of generating audio TLDR for spatialized audio content based on an approximation of a personalized HRTF in accordance with certain embodiments is shown.

[0040] Figure 12 Depicted is a block diagram of an example of a system including a head mounted device for implementing some examples disclosed herein, according to certain embodiments.

[0041] These figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown may be employed without departing from the principles of the present disclosure or the benefits claimed in the present disclosure.

[0042] In the drawings, similar components and / or features may have the same reference number. In addition, various components of the same type may be distinguished by following the reference number with a hyphen and a second reference number that distinguishes multiple similar components. If only the first reference number is used in the specification, the description applies to any similar component having the same first reference number, regardless of the second reference number. DETAILED DESCRIPTION

[0043] The present disclosure generally relates to determining head-related transfer functions (HRTFs), and more particularly, to determining HRTFs or HRTF parameters using head-mounted devices (e.g., headsets) and in-ear devices. Various inventive embodiments, including devices, systems, methods, structures, and processes, are described herein.

[0044] In virtual reality (VR) systems, augmented reality (AR) systems, or other near-eye display systems, in order to improve the immersive experience using the near-eye display system, in addition to foveated image rendering and improved image quality (e.g., high resolution, large color gamut, large field of view, etc.), it may be desirable to render multi-channel spatialized audio content (e.g., based on a single-channel input audio signal) so that the user can hear the sound of virtual objects as if the sound originated from the target location of the virtual objects. To generate such spatialized audio content, audio rendering can be performed, for example, in a binaural configuration or a transaural configuration. In a binaural configuration using headphones or in-ear devices (IEDs, such as earbuds), before generating binaural sound using, for example, transducers, it may be necessary to determine acoustic transfer functions that characterize the modifications to the sound along the path from a (virtual) sound source to the user's ears. The headphones or IED can then use these acoustic transfer functions to modify (e.g., filter) the sound signals to synthesize binaural sound that appears to originate from a specific point in space, and the user can infer the spatial location of the sound source based on localization cues in the binaural sound.

[0045] The modification of the sound on the path from the (virtual) sound source to the listener's ear can include attenuating signals of different frequencies differently and may depend on, for example, the size and shape of the user's outer ear (e.g., the pinna), the size, shape, and density of the user's head and / or torso, and the acoustic characteristics of the space in which the sound is played. In this regard, the acoustic transfer function between the sound source and the user's ear (e.g., at the outer end of the ear canal) can be referred to as the head-related transfer function (HRTF) or head shadow. The HRTF can include relevant acoustic cues for localizing the real sound source, such as interaural level difference (ILD), interaural time difference (ITD), and monaural spectral cues. The time domain representation of the HRTF is the head-related impulse response (HRIR). The HRTF depends on the direction of the sound source relative to the center of the head. For a sound source at a given position, the HRTF for the left ear and the HRTF for the right ear can be different. The relative position of the user's torso with respect to the user's head may also affect the HRTF. Due to the different anatomical structures of different users (for example, the size and shape of the pinnae, head, and torso), HRTFs are often different for different users. Therefore, an HRTF that works for one user may not work for another. Using non-personalized HRTFs to synthesize binaural sound may degrade the listening experience, such as impaired localization accuracy and perceptual externalization.

[0046] HRTFs are typically measured in an anechoic chamber (e.g., using a dummy) to minimize the effects of early reflections and reverberation on the measured response. HRTFs can be measured at small increments of azimuth and elevation, and interpolation can be used to synthesize an HRTF for any position. Using small increments of azimuth and elevation may require measuring HRTFs for many (e.g., more than 100, such as several hundred) spatial locations. Even with small increments, interpolation can lead to front-back confusion and can be difficult to optimize. Furthermore, as mentioned above, HRTFs vary from person to person because sound propagation varies depending on the size and shape of each person's head, torso, and pinnae. Due to differences in individual characteristics, applying an HRTF measured from a dummy or another person to a specific person may reduce the performance of immersive sound. Therefore, HRTFs need to be personalized to achieve the desired localization performance. However, the process of creating personalized HRTFs based on measurements in an anechoic chamber may require the use of specialized and expensive measurement systems and can be time-consuming and computationally intensive. Consequently, such a process may not be scalable to a large number of users.

[0047] According to certain embodiments disclosed herein, a user-worn head-mounted device and an in-ear device (IED) (which may or may not be part of the head-mounted device) can be used to determine a user's personalized HRTF or at least some parameters of a personalized HRTF by, for example, collecting audio signals in a natural environment suitable for HRTF measurement; estimating the locations (e.g., directions) of sound sources of the collected audio signals; and determining HRTFs for these locations based on the audio signals collected by the head-mounted device and the IED. In this regard, the head-mounted device and the IED can listen to random incidental sounds in the user's natural environment with minimal or no user involvement to gradually add data points associated with different spatial locations to the data point cloud of the user's personalized HRTF, so that a user-specific HRTF across all expected source directions can be constructed over time. The head-mounted device and the IED can be worn by the user for a period of time (e.g., days or weeks) for other purposes (e.g., AR / VR applications) to accumulate HRTFs or HRTF parameters for different directions. In this way, the user's HRTF or HRTF parameters can be determined with minimal or no user involvement and without the use of specialized measurement systems, such as sound attenuation chambers and speaker arrays. In some embodiments, when HRTFs for a sufficient number of directions have been collected, these HRTFs can be interpolated to generate a personalized HRTF for the user for any sound source direction.

[0048] Each sound signal collected by the head-mounted device and the IED can have a short duration (e.g., a few seconds, hundreds of milliseconds, or tens of milliseconds, such as a click or other very short duration pulse), and / or can have a frequency band that may be in at least a portion of the human hearing range (e.g., between approximately 20 Hz and 20 kHz). A HRTF for the entire human hearing range can be determined using a set of sound signals, in which each sound signal can cover a different corresponding frequency range. In some embodiments, the HRTF for a portion of the human hearing range can be determined by averaging the results determined using multiple sound signals to improve accuracy. In some embodiments where different filters can be used to implement different frequency bands of HRTFs, the filter can be selected based on the HRTFs for each portion of the human hearing range determined using different sound signals covering different portions of the human hearing range. In some embodiments, the HRTF or parameters of the HRTF can be used to personalize a non-personalized HRTF, for example, by personalizing the interaural time difference (ITD) and parameter scaling factors, such as factors for compressing or stretching the amplitude spectrum of the HRTF in the frequency domain (referred to herein as frequency scaling factors).

[0049] The head-mounted device for generating HRTF or parameters of HRTF can be an AR / VR system, which includes, for example, a microphone array and an audio controller. The head-mounted device can include one or more IEDs, or can communicate with one or more IEDs worn by a user of the head-mounted device. In some embodiments, the head-mounted device can include a camera system and / or another sensor system (e.g., a microphone array) or can communicate with the camera system and / or another sensor system, which can be used to determine the position (e.g., direction) of the object generating the sound. The microphone array and / or IED can be used to collect audio signals in a natural environment, and the audio controller can analyze the collected audio signals to determine whether they are suitable for HRTF measurement. For example, sounds suitable for personalized HRTF measurement can have high spatial staticity (at least when the sound is collected by a head-mounted device and an in-ear device, for example, within 1 second, within a few hundred milliseconds, or within tens of milliseconds), and can also have a high signal-to-noise ratio (SNR), low reverberation level, low reverberation time (for example, low RT60), and a wide frequency spectrum, etc.

[0050] The direction or position of a sound source suitable for HRTF measurement can be determined based on a direction of arrival (DOA) determined using, for example, two or more microphones (e.g., in a microphone array of a head-mounted device), or one or more cameras on the head-mounted device or in communication with the head-mounted device. In some embodiments, one or more cameras or one or more position sensors (e.g., an inertial measurement unit (IMU)) on the head-mounted device can be used to determine the relative position of the user's torso with respect to the user's head, as the HRTF may be affected by the relative position of the user's torso with respect to the user's head.

[0051] The audio signals collected by the microphone array can be used to determine a nearly anechoic reference signal for use in determining the head-related transfer function (e.g., by performing beamforming in the estimated direction of the sound source to determine a reference sound signal). The audio signals collected by the IED can be used to determine a HRTF or HRTF parameters specific to the direction of the sound source (and the relative position of the user's torso with respect to the user's head) by dividing the audio signals collected by the IED by a reference signal determined based on the audio signals collected by the microphone array.

[0052] In some embodiments, HRTFs for different users can be approximated by low-complexity signal processing using parameters in a lower-dimensional parameter space. For example, in some embodiments, the lower-dimensional parameters of the HRTF determined using the techniques disclosed herein can include an ITD for a sound source direction and lower-dimensional parameters of the HRTF, such as parameters of a filter used to implement the HRTF (e.g., the center frequency, gain, and Q value of the filter, or other parameters used to define the filter). In some examples, the lower-dimensional parameters of the HRTF determined using the techniques disclosed herein can include a personalized ITD and a personalized parameter scaling factor for personalizing a non-personalized HRTF. In some embodiments, to determine the parameters in the lower-dimensional parameter space for HRTF rendering, a set of parameters (e.g., filter parameters or frequency scaling factors) can be initialized and then optimized to match the HRTF measured for a sound source direction. In some embodiments, a machine learning model such as a neural network can be trained to fit the HRTF using lower-dimensional parameters (e.g., filter parameters or frequency scaling factors) in a way that these parameters can vary smoothly across space and exhibit similar behavior across different users.

[0053] Embodiments of the present invention may include an artificial reality system, or may be implemented in conjunction with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, and artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with collected (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that produces a three-dimensional effect to the viewer). In addition, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used to create content in artificial reality and / or be used in artificial reality in other ways. Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including wearable devices (e.g., head-mounted displays) connected to a host computer system, standalone wearable devices (e.g., head-mounted displays), mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0054] In the following description, for the purpose of explanation, various specific details are set forth in order to provide a thorough understanding of each example of the present disclosure. However, it will be apparent that various examples can be put into practice without these specific details. For example, devices, systems, structures, components, methods and other parts can be shown as parts in the form of block diagrams to avoid blurring these examples in terms of unnecessary details. In other examples, well-known devices, processes, systems, structures and techniques can be shown without the necessary details to avoid blurring each example. These figures and descriptions are not intended to be limiting. The terms and expressions adopted in this disclosure are used as descriptive terms rather than restrictive terms, and when using these terms and expressions, there is no intention to exclude any equivalents or parts thereof of the features shown and described. The word "example" is used herein to mean "used as an example, instance or illustration". Any embodiment or design described herein as an "example" is not necessarily to be interpreted as being preferred or advantageous over other embodiments or designs.

[0055] Figure 1 is a stereoscopic image of an example of a near-eye display (NED) 100 in the form of a pair of glasses for implementing some of the multiple examples disclosed herein. Typically, the NED 100 can be worn on the head (e.g., face) of a user such that content (e.g., media content) is presented to the user using a display assembly and / or an audio system. However, the NED 100 can also be used to present media content to the user in different manners. Examples of media content presented by the NED 100 include images, video, audio, or a combination thereof. In the example shown, the NED 100 includes a frame 110 and can include other components such as a display assembly including one or more display elements 120, a depth camera assembly (DCA), an audio system, and one or more position sensors 190. Although Figure 1 Components of the NED 100 are shown as being located at certain locations on the NED 100, but these components may be located elsewhere on the NED 100; on a peripheral device that is paired with the NED 100; or some combination thereof. Figure 1 More components or fewer components than shown may be used.

[0056] The frame 110 can hold other components of the NED 100. The frame 110 can include a front member that holds one or more display elements 120, and end members (e.g., temples) for attaching the NED 100 to the user's head. The front member of the frame 110 spans across the top of the user's nose. The length of the end members can be adjustable (e.g., adjustable temple length) to fit different users. The end members can also include portions that curve behind the user's ears (e.g., temple tips, earpieces).

[0057] One or more display elements 120 can provide light to a user wearing the NED 100. As shown, the NED 100 includes a display element 120 for each eye of the user. In some embodiments, the display element 120 generates image light, which is provided to the eyebox of the NED 100. The eyebox is the position in space occupied by the user's eyes when wearing the NED 100. For example, the display element 120 may include a waveguide display. A waveguide display includes a light source (e.g., a two-dimensional light source, one or more line light sources, one or more point light sources, etc.) and one or more waveguides. Display light from the light source can be coupled into one or more waveguides, which can replicate the display light and couple the display light out of an array of waveguide positions to replicate the pupil within the eyebox of the NED 100. The coupling of light into and / or the coupling of light out of the one or more waveguides can be accomplished using, for example, one or more diffraction gratings or mirrors. In some embodiments, the waveguide display includes a scanning element (e.g., a waveguide, a reflector, etc.) that scans light from a light source as it is coupled into one or more waveguides. In some embodiments, one or both of the display elements 120 may be opaque and may not transmit light from the environment surrounding the NED 100. For example, the surrounding environment may be a room in which a user wearing the NED 100 is located, or the user wearing the NED 100 may be outdoors and the surrounding environment may be an outdoor area. In this context, the NED 100 may generate and present VR content. Alternatively, in some embodiments, one or both of the display elements 120 are at least partially transparent so that light from the surrounding environment can be combined with light from the one or more display elements to generate and present AR content and / or MR content.

[0058] In some embodiments, the display element 120 may not generate image light, but rather the display element is a lens that transmits light from a local area to the eyepiece. For example, one or both of the display elements 120 may be an uncorrected (over-the-counter) lens or a prescription lens (e.g., a single vision lens, a bifocal and trifocal lens, or a gradient lens) to help correct the user's vision defects. In some embodiments, the display element 120 may be polarized and / or tinted to protect the user's eyes from the sun. In some embodiments, the display element 120 may include additional display optics (not shown). The display optics may include one or more optical elements (e.g., a lens, a Fresnel lens, etc.) that direct light from the display element 120 to the eyepiece. The display optics may, for example, correct aberrations in some or all of the image content, amplify some or all of the image, or some combination thereof.

[0059] A depth camera assembly (DCA) may determine depth information of a portion of a local area surrounding the NED 100. The DCA may include, for example, one or more imaging devices 130 and a DCA controller ( Figure 1 (not shown), and in some embodiments may further include an illuminator 140. For example, the illuminator 140 may be used to illuminate a portion of a local area with light. For example, the light may be a flash, structured light (e.g., dot pattern, stripes, etc.) in the visible or infrared (IR) band. In some embodiments, one or more imaging devices 130 (cameras) may capture an image of a portion of the local area that includes light from the illuminator 140. In some embodiments, the NED 100 may include one or more light sensors (e.g., IR sensors, Figure 1 (not shown), the one or more light sensors can detect light from the illuminator 140 and reflected by objects in the surrounding environment to determine the time of flight and the distance between the object and the NED 100. In the example shown, Figure 1 A single illuminator 140 and two imaging devices 130 are shown. In alternative embodiments, there may be zero to multiple illuminators 140, zero to multiple imaging devices 130, and zero to multiple light sensors. The DCA controller may use the captured images and one or more depth determination techniques to determine depth information for the portion of the local area, such as direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereo analysis, active stereo analysis (e.g., using texture added to the scene by light from illuminator 140), some other technique for determining scene depth, or a combination thereof.

[0060] The audio system of NED 100 can provide audio content. The audio system can include, for example, an array of transducers (e.g., speakers), an array of acoustic sensors (e.g., microphones), and an audio controller 150. In other embodiments, the audio system can include different components and / or additional components. In some examples, the functionality described with reference to the components of the audio system can be distributed among the components in a manner different from that described herein. For example, some or all of the various functions of the controller can be performed by a remote server or processor of NED 100.

[0061] The transducer array can include a plurality of transducers that present sound to the user. The transducer can be a speaker 160 or a tissue transducer 170 (e.g., a bone conduction transducer or a cartilage conduction transducer). Although the speaker 160 is shown as being located outside of the frame 110 in the illustrated example, in some examples the speaker 160 can be enclosed within the frame 110. In some embodiments, rather than a separate speaker for each ear, the NED 100 can include a speaker array comprising multiple speakers integrated into the frame 110 to improve the directionality of the presented audio content. The tissue transducer 170 can be coupled to the user's head and can directly vibrate the user's tissue (e.g., bone or cartilage) to produce sound. The number and / or position of the transducers in the transducer array can be the same as that in the embodiment of the present invention. Figure 1 Different as shown.

[0062] The acoustic sensor array can detect sounds within a local area of ​​the NED 100. The acoustic sensor array includes a plurality of acoustic sensors 180. The acoustic sensors 180 collect sounds emanating from one or more sound sources in a local area (e.g., a room). Each acoustic sensor is configured to detect sounds and convert the detected sounds into electronic (analog or digital) signals. The acoustic sensors 180 can be acoustic wave sensors, microphones, sound transducers, or similar sensors suitable for detecting sounds. In some embodiments, one or more acoustic sensors 180 can be placed in the ear canal of each ear (e.g., acting as binaural microphones). In some embodiments, the acoustic sensors 180 can be placed on an outer surface of the NED 100, on an inner surface of the NED 100, separate from the NED 100 (e.g., as part of some other device), or some combination thereof. In different embodiments, the number and / or location of the acoustic sensors 180 can vary depending on the device. Figure 1 The number and / or location of the microphones may be different from that shown in . For example, the number of acoustic detection locations may be increased to increase the amount of audio information collected and to improve the sensitivity and / or accuracy of that information. The acoustic detection locations may be oriented so that the microphones can detect sounds in a wide range of directions around the user wearing the NED 100.

[0063] The audio controller 150 can process information from the sensor array describing the sounds detected by the sensor array. The audio controller 150 can include a processor and one or more computer-readable storage media and can be configured to: generate direction of arrival (DOA) estimates; generate acoustic transfer functions (e.g., array transfer functions and / or head-related transfer functions); track the location of a sound source; form a beam in the direction of a sound source; classify the sound source; generate an acoustic filter for the speaker 160; or a combination thereof. In some embodiments, the audio controller 150 can select an audio time difference and intensity difference renderer (TLDR) that approximates a given HRTF with a specific level of accuracy. For example, the TLDR can be selected based on input parameters such as a target power consumption, a target computational load specification, a target memory footprint, a target accuracy level for the HRTF approximation, or a combination thereof. In these embodiments, the audio controller 150 can select an audio TLDR from a set of audio TLDRs based on the input target accuracy level and configure the selected audio TLDR based on input parameters such as a target sound source angle and a target fidelity for the audio rendering. The audio controller 150 may apply the selected and configured one or more audio TLDRs to an input audio signal received at a single channel to generate multi-channel spatialized audio content for provision to the speaker 160 .

[0064] The position sensor 190 can generate one or more measurement signals in response to movement of the NED 100. The position sensor 190 and the imaging device 130 can be used alone or in combination to determine, for example, the relative position of a user's torso relative to the user's head. The position sensor 190 can be located on a portion of the frame 110 of the NED 100. The position sensor 190 can include, for example, an inertial measurement unit (IMU). The position sensor 190 includes, for example, one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor that detects movement, a type of sensor used for error correction in the IMU, or a combination thereof. The position sensor 190 can be located externally to the IMU, internally to the IMU, or a combination thereof.

[0065] In some embodiments, the NED 100 may include simultaneous localization and mapping (SLAM) functionality for determining the position of the NED 100 and updating a model of the local area. For example, the NED 100 may include a passive camera assembly (PCA) that generates color image data. The PCA may include one or more RGB cameras that capture images of part or all of the local area. In some embodiments, some or all of the imaging devices 130 of the DCA may also function as PCAs. The images captured by the PCA and the depth information determined by the DCA may be used to determine parameters of the local area, generate a model of the local area, update the model of the local area, determine the position of a user, or a combination thereof. In some embodiments, the position sensor 190 may track the position (e.g., location and posture) of a user of the NED 100 within a room. Additional details regarding the components of the NED 100 are discussed below.

[0066] Figure 2 2 is a perspective view of an example of a near-eye display in the form of a head-mounted display (HMD) 200 for implementing some of the examples disclosed herein. HMD 200 can be, for example, part of a VR system, an AR system, an MR system, or any combination thereof. HMD 200 can include a body 220 and a headband 230. Figure 2 The bottom side 223, front side 225, and left side 227 of the main body 220 are shown in perspective. The headband 230 may have an adjustable or extendable length. There may be sufficient space between the main body 220 and the headband 230 of the HMD 200 to allow the user to wear the HMD 200 on the user's head. The HMD 200 may include at least some of the components of the NED 100 described above. In some embodiments, the HMD 200 may include additional, fewer, or different components.

[0067] The HMD 200 can present media to the user, which includes virtual views and / or augmented views of a physical real-world environment with computer-generated elements. Examples of media presented by the HMD 200 can include images (e.g., two-dimensional (2D) images or three-dimensional (3D) images), video (e.g., 2D video or 3D video), audio, or any combination thereof. The images and videos can be displayed via one or more display components (e.g., a display unit) enclosed in the body 220 of the HMD 200. Figure 2The HMD 200 may include two eyebox areas.

[0068] In some embodiments, the HMD 200 may include various sensors (not shown), such as depth sensors, motion sensors, position sensors, acoustic sensors, and eye-tracking sensors. Some of these sensors may sense using structured light patterns. In some embodiments, the HMD 200 may include an input / output interface for communicating with a console. In some embodiments, the HMD 200 may include a virtual reality engine (not shown) that may execute applications within the HMD 200 and may receive depth information, position information, acceleration information, velocity information, predicted future position, or any combination thereof of the HMD 200 from various sensors. In some embodiments, the information received by the virtual reality engine may be used to generate signals (e.g., display instructions) to one or more display components. In some embodiments, the HMD 200 may include multiple locators (not shown) located at fixed positions on the body 220 relative to each other and relative to a reference point. Each of the multiple locators may emit light that can be detected by an external imaging device. The HMD 200 may also include an audio system that may include, for example, a system described above with reference to a controller. Figure 1 The transducer (e.g., speaker) array, acoustic sensor (e.g., microphone) array, and audio controller described above may need to provide spatial sound or three-dimensional (3D) sound rendering so that the user can perceive the sound of a virtual object as if it originates from the target location of the virtual object, thereby obtaining an immersive user experience and successfully implementing a VR / AR system.

[0069] Figure 3 is a block diagram of an example of an audio system 300 in a near-eye display or head-mounted display according to some embodiments. The audio system 300 may be Figure 1 or Figure 2 An example of an implementation of an audio system in . The audio system 300 can generate one or more acoustic transfer functions of a user and can implement the one or more acoustic transfer functions to generate audio content for the user. Figure 3In the example shown, the audio system 300 includes a transducer array 310, a sensor array 320, and an audio controller 330. Some other embodiments of the audio system 300 may have the same Figure 3 In some embodiments, the functionality of audio system 300 may be distributed among components in a manner different from that described herein.

[0070] As described above, the transducer array 310 can be configured to present audio content to the user and can include multiple transducers positioned at different locations of a head-mounted display or near-eye display (generally referred to as a head-mounted view). A transducer is a device that provides audio content. A transducer can be, for example, a speaker (e.g., speaker 160), a tissue transducer (e.g., tissue transducer 170), or another device that can provide audio content. A tissue transducer can be configured to function as a bone conduction transducer or a cartilage conduction transducer. The transducer array 310 can present audio content by air conduction (e.g., through one or more speakers), by bone conduction (through one or more bone conduction transducers), by cartilage conduction (through one or more cartilage conduction transducers), or a combination thereof. In some embodiments, the transducer array 310 can include one or more transducers to cover different portions of a frequency range. For example, a piezoelectric transducer can be used to cover a first portion of a frequency range, and a dynamic coil transducer can be used to cover a second portion of the frequency range.

[0071] The bone conduction transducer can generate sound pressure waves by vibrating the bones / tissue of the user's head. The bone conduction transducer can be coupled to a portion of the headset and can be configured to be located behind the pinna, which is attached to a portion of the user's skull. The bone conduction transducer can receive vibration instructions from the audio controller 330 and vibrate the portion of the user's skull based on the received instructions. The vibrations from the bone conduction transducer can generate tissue-propagated sound pressure waves that bypass the eardrum and propagate toward the user's cochlea.

[0072] The bone conduction transducer can generate sound pressure waves by vibrating one or more parts of the auricular cartilage of the user's ear. The bone conduction transducer can be coupled to a portion of the headset and can be configured to be attached to one or more parts of the auricular cartilage of the ear. For example, the bone conduction transducer can be attached to the back of the auricle of the user's ear. The bone conduction transducer can be located anywhere along the auricular cartilage around the outer ear (e.g., the pinna, the tragus, some other part of the auricular cartilage, or a combination thereof). Vibrating one or more parts of the auricular cartilage can generate, for example: air-borne sound pressure waves outside the ear canal; sound pressure waves generated by tissue, which cause some parts of the ear canal to vibrate, thereby generating air-borne sound pressure waves within the ear canal; or a combination thereof. The generated air-borne sound pressure waves can propagate along the ear canal toward the eardrum.

[0073] The transducer array 310 can generate audio content according to instructions from the audio controller 330. In some embodiments, the audio content is spatialized. Spatialized audio content is audio content that appears to originate from a specific direction and / or target area (e.g., an object in a local area and / or a virtual object in a target location). For example, the spatialized audio content can cause a user of the audio system 300 to perceive the sound as originating from a virtual singer in a specific position or direction relative to the user (e.g., next door or on a stage). The transducer array 310 can be coupled to a wearable device (e.g., NED 100 or HMD 200) or can be part of the wearable device. In alternative embodiments, the transducer array 310 can be a plurality of speakers separate from the wearable device (e.g., coupled to an external console). In some embodiments, the transducer array 310 can include a pair of in-ear devices (e.g., in the form of earbuds).

[0074] The sensor array 320 can detect sounds within a local area surrounding the sensor array 320. For example, the sensor array 320 can include multiple acoustic sensors that can each detect air pressure changes associated with sound waves and convert the detected air pressure changes into electronic (analog or digital) signals. The multiple acoustic sensors can be located on a wearable device (e.g., NED 100 or HMD 200), on a user (e.g., in the form of earbuds in the user's ear canal or a headset on the user's ear), on a neckband, or a combination thereof. The acoustic sensors can include, for example, microphones, vibration sensors, accelerometers, or another sensor capable of detecting air pressure changes. Two or more acoustic sensors in the sensor array 320 can be used, for example, to determine the location of a sound source. In some embodiments, the sensor array 320 can be configured to use at least some of the multiple acoustic sensors to monitor audio content generated by the transducer array 310. Increasing the number of acoustic sensors can improve the accuracy of information describing the sound field generated by the transducer array 310 and / or sounds from the local area (e.g., directionality).

[0075] The audio controller 330 can control the operation of the audio system 300. The audio controller 330 can include, for example, a data repository 335, a direction of arrival (DOA) estimation subsystem 340, a transfer function subsystem 350, a tracking subsystem 360, a beamforming subsystem 370, and a sound filter subsystem 380. The audio controller 330 can be located inside the HMD or inside a console connected to the HMD. In different embodiments, the audio controller 330 can have Figure 3 The functions of the audio controller 330 may be distributed among the components in a manner different from that described here. For example, some functions of the audio controller 330 may be performed outside the HMD. The user may allow the audio controller 330 to transmit data collected by the HMD to a system external to the HMD (e.g., a console or server), and the user may select privacy settings that control access to any such data.

[0076] The data repository 335 may store data used or generated by the audio system 300. The data in the data repository 335 may include, for example, sounds recorded in a local area of ​​the audio system 300, audio content, a head-related transfer function (HRTF), some parameters of the HRTF, transfer functions of one or more sensors, array transfer functions (ATFs) of one or more acoustic sensors in each acoustic sensor, sound source locations, a virtual model of the local area, direction of arrival estimation results, sound filters, a model (e.g., a lookup table) for retrieving HRTFs or parameters of HRTFs based on the direction of a sound source, other data related to the use of the audio system 300 or generated by the audio system 300, or any combination thereof.

[0077] For example, the data repository 335 may store the collected sound signal, the sound source location information determined from the collected sound signal, the HRTF or at least some parameters of the HRTF determined based on the collected sound signal and the sound source location information. Some parameters of the HRTF may be parameters in a lower-dimensional parameter space, such as parameters of individual filters (e.g., notch filters, bandpass filters, high-shelf filters, and / or low-shelf filters). Some parameters of the HRTF may be used to modify the non-personalized HRTF to generate a personalized HRTF, such as a frequency scaling factor and a personalized interaural time difference (ITD).

[0078] The data repository 335 may also store data associated with the operation of the audio filter subsystem 380, which is associated with the selection and application of the audio time and intensity difference renderer (TLDR). The stored data may include static filter parameter values ​​and one-dimensional and / or two-dimensional interpolation lookup tables for finding frequency / gain / Q triplet parameter values, such as filter parameters (e.g., center frequency, notch depth, and slope), for a given azimuth and / or elevation angle of a target sound source. The data repository 335 may also store single-channel audio signals for processing by the audio TLDR and presenting them to the user as spatialized audio content via multiple channels via the HMD. In some embodiments, the data repository 335 may store default values ​​for input parameters, such as target fidelity of rendered audio content in the form of target frequency response values, target signal-to-noise ratio, target power consumption of the selected audio TLDR, target computational requirements of the selected audio TLDR, and target memory usage of the selected audio TLDR. The data repository 335 can store values ​​such as desired spectral profiles and equalization of spatialized audio content generated from an audio TLDR. In some embodiments, the data repository 335 can store a selection model for selecting an audio TLDR based on input parameter values. The stored selection model can take the form of a lookup table that maps a range of input parameter values ​​to one of a plurality of audio TLDRs. In some embodiments, the stored selection model can take the form of a specific weighted combination of input parameter values ​​that are mapped to one of a plurality of audio TLDRs. In some embodiments, the data repository 335 can store data for use by a parameterized filter fitting system. The stored data can include a set of measured HRTFs associated with a context vector, the spatial position of a sound source (e.g., azimuth and elevation values), and anthropometric features of one or more users. The data repository 335 can also store updated audio filter parameter values ​​determined by the parameterized filter fitting system.

[0079] The DOA estimation subsystem 340 can be configured, for example, to localize sound sources in a local area based at least in part on information from the sensor array 320. Localization is the process of determining the location of a sound source relative to the user of the audio system 300. The DOA estimation subsystem 340 can perform a DOA analysis to localize one or more sound sources in the local area. The DOA analysis can include analyzing the intensity, spectrum, and / or time of arrival of a sound at each acoustic sensor of the sensor array 320 to determine the direction from which the sound originates. The DOA analysis can include any suitable algorithm for analyzing the ambient acoustic environment in which the audio system 300 is located.

[0080] For example, DOA analysis can be designed to receive input signals from sensor array 320 and apply digital signal processing algorithms to these input signals to estimate the direction of arrival. These algorithms can include, for example, a delay-sum algorithm in which the input signal is sampled and weighted and delayed versions of the resulting sampled signals are averaged together to determine the DOA. In another example, a least mean squared (LMS) algorithm can be implemented to create an adaptive filter, which can then be used to identify, for example, differences in signal strength or differences in arrival time. These differences can then be used to estimate the DOA. In yet another example, the DOA can be determined by converting the input signal into the frequency domain and selecting specific bins within the time-frequency (TF) domain for processing. For example, each selected TF bin can be processed to determine whether the bin includes a portion of the audio spectrum containing a direct path audio signal. Those bins containing portions of a direct path signal can then be analyzed to identify the angle at which the sensor array 320 received the direct path audio signal. The determined angle can then be used to identify the DOA of the received input signal. Other algorithms not discussed above may also be used alone or in combination with the above algorithms to determine DOA.

[0081] In some embodiments, the DOA estimation subsystem 340 can determine the DOA relative to the absolute position of the audio system 300 within a local area. The position of the sensor array 320 can be received from an external system, such as some other component of the HMD, an artificial reality console, a map-building server, a position sensor (e.g., position sensor 190), and the like. The external system can create a virtual model of the local area, in which the local area and the position of the audio system 300 are mapped. The received position information can include the position and / or orientation of some or all components of the audio system 300 (e.g., sensor array 320). The DOA estimation subsystem 340 can update the estimated DOA based on the received position information. As described above, other components of the HMD (e.g., a camera and / or position sensor) can also be used alone or in combination with the DOA estimation subsystem 340 to determine the direction of a sound source or improve the accuracy of the direction of a sound source determined by the DOA estimation subsystem 340.

[0082] The transfer function subsystem 350 is configured to generate one or more acoustic transfer functions. A transfer function may include a mathematical function that gives a corresponding output value for each possible input value. The transfer function subsystem 350 may generate one or more acoustic transfer functions associated with the audio system based on parameters of the detected sound. The acoustic transfer functions may include, for example, an array transfer function (ATF), a head-related transfer function (HRTF), other types of acoustic transfer functions, or any combination thereof.

[0083] ATF can be used to characterize how a microphone receives sound from a point in space. The ATF can include multiple transfer functions that characterize the relationship between a sound source and the corresponding sound received by each acoustic sensor in the sensor array 320. Accordingly, for a sound source, each acoustic sensor in the sensor array 320 can have a corresponding transfer function, and a set of transfer functions of the acoustic sensors in the sensor array 320 can be referred to as the ATF of the sound source. The sound source can include, for example, someone or something that generates sound in a local area, a user, or one or more transducers of the transducer array 310. Since a person's anatomical structure (e.g., the size and shape of the ear, head, body, etc.) affects the sound when it propagates to the human ear, the ATF relative to a specific sound source position of the sensor array 320 may vary from user to user. Accordingly, each ATF of the sensor array 320 can be a personalized transfer function for each user of the audio system 300.

[0084] In some embodiments, the transfer function subsystem 350 can determine one or more HRTFs for a user of the audio system 300. HRTFs characterize how an ear receives sound from a sound source in space. Because a person's unique anatomy (e.g., the size and shape of the ears, head, body, etc.) can affect the sound as it is transmitted to their ears, the HRTF for a particular sound source location relative to a person can be unique for each ear of that person (and unique to that person). In some embodiments, the transfer function subsystem 350 can use a calibration process to determine the user's HRTFs. In some embodiments, the transfer function module 350 can provide information about the user to a remote system. The user can adjust privacy settings to allow or prevent the transfer function subsystem 350 from providing information about the user to any remote system. The remote system can use, for example, machine learning or other techniques to determine a set of HRTFs customized for the user and provide this customized set of HRTFs to the audio system 300. Further details on determining HRTFs using the audio controller 330 or another processing unit (e.g., a computer or remote server) are described below.

[0085] The tracking subsystem 260 can be configured to track the position of one or more sound sources. For example, the tracking subsystem 360 can compare a current DOA estimate with stored historical DOA estimates to determine, based on changes in the DOA estimate for a sound source during a time period, whether and how much the sound source has moved relative to the user during that time period. In some embodiments, the audio system 300 can recalculate the DOA estimate on a periodic schedule (e.g., once per second or every 100 milliseconds). In some embodiments, the tracking subsystem 360 can additionally or alternatively detect changes in position based on visual information received from the HMD or some other external source (e.g., one or more cameras or other image sensors). The tracking subsystem 360 can track the movement of one or more sound sources over time and can store values ​​for multiple sound sources and the position of each sound source at each point in time. In response to changes in the sound source position values, the tracking subsystem 360 can determine that the sound source has moved. In some embodiments, the tracking subsystem 360 can calculate an estimate of localization variance. The localization variance can be used as a confidence level for each change in movement determination. The results of tracking one or more sound sources can be used to determine whether the sound sources were moving during the time period when the sound signals were collected, and therefore whether the collected sound signals are suitable for determining a transfer function (e.g., ATF or HRTF) by the transfer function subsystem 350.

[0086] The beamforming subsystem 370 can be configured to analyze the sounds detected by the sensor array 320 and selectively emphasize (e.g., amplify) sounds originating from sound sources within a certain area (or from a certain direction) while de-emphasizing (e.g., attenuating) sounds originating from other areas (or directions). When analyzing the sounds detected by the sensor array 320, the beamforming subsystem 370 can combine information from different acoustic sensors to emphasize sounds originating from a specific area (or direction) within the local area while de-emphasizing sounds originating outside of that area (or direction). In one example, the beamforming subsystem 370 can use this technique to determine a reference sound signal based on sound signals measured by two or more acoustic sensors of the sensor array 320. In some embodiments, the beamforming subsystem 370 can isolate audio signals associated with sounds originating from a specific sound source in the local area from audio signals associated with other sound sources based on, for example, different DOA estimation results from the DOA estimation subsystem 340 and the tracking subsystem 360. Thus, the beamforming subsystem 370 can selectively analyze discrete sound sources within the local area. In some embodiments, beamforming subsystem 370 can enhance the signal from a sound source. For example, beamforming subsystem 370 can apply sound filters that eliminate signals above certain frequencies, below certain frequencies, or between certain frequencies. Signal enhancement can be used to enhance the sound associated with a given identified sound source relative to other sounds detected by sensor array 320.

[0087] The sound filter subsystem 380 can determine the sound filter used to generate audio data to drive the transducer array 310. The sound filter can positively or negatively amplify the sound according to the frequency change. The audio content presented by the transducer array can be multi-channel spatialized audio. For example, the sound filter can cause the audio content to be spatialized so that the audio content seems to originate from a specific direction and / or target area (for example, an object in a local area and / or a virtual object in a target position). The sound filter subsystem 380 can implement the sound filter using HRTF and / or acoustic parameters. Acoustic parameters can describe the acoustic properties of the local area and can include, for example, reverberation time, reverberation level, and room impulse response. In some embodiments, the sound filter subsystem 380 can calculate one or more of these acoustic parameters. In some embodiments, the sound filter subsystem 380 can obtain the acoustic parameters from (for example, as described below with respect to Figure 12 Acoustic parameters are requested from the map building server described in (described in).

[0088] In some embodiments, the sound filter subsystem 380 can select and configure an audio TLDR from a set of possible audio TLDRs based on received input parameters. The received input parameters may include, for example, target sound source angle, target fidelity of audio rendering, target power consumption, target computational load, target memory footprint, and the accuracy level of approximating the target of a given HRTF. The selected and configured audio TLDR can be used to generate spatialized audio content in multiple channels from an input single-channel audio signal. The input single-channel audio signal (also referred to as a mono audio signal (mono-audio signal, monaural audio signal, monophonic audio signal), etc.) is the audio content that arrives at a single channel, and when provided to a speaker, can be heard as a sound emitted from a single position. The selected and configured audio TLDR can be used to process the input single-channel audio to generate a multi-channel audio signal, such as stereo audio content through two separated audio channels (e.g., left channel and right channel). The selected audio TLDR can be configured to use static audio filters, dynamic audio filters, and delays so that it approximates a given HRTF with a specific level of accuracy. Filtering an input single-channel audio signal with the configured audio TLDR simulates applying one or more HRTFs of a user of the audio system to the single-channel audio signal, thereby generating multi-channel spatialized audio content. In some embodiments, the sound filter subsystem 380 can request data associated with a filter parameter model from a parametric filter fitting system for HRTF rendering.

[0089] The principles of spatial hearing can be based on binaural and monaural cues. Binaural cues can be related to the differences between the sound signals received by the two ears. These differences include the arrival time difference and intensity difference of the sound signals received by the two ears, which can be referred to as interaural time difference (ITD) and interaural intensity difference (ILD), respectively. These binaural cues can be related to the perceived horizontal direction (azimuthal localization) of the sound source. Monaural cues can include direction-related spectral cues caused by the head, body, and pinna. Monaural cues can modify the amplitude spectrum of the sound source and can be closely related to the perceived vertical direction of the sound source, and can therefore be used by the brain to estimate the height of the sound source. Another monaural cue is the reverberation factor, which is defined as the amount of reflection and reverberation relative to the direct sound and can be related to the perceived distance of the sound source. Although there is no simple relationship between direction and sound localization cues, the human brain can use these cues to accurately estimate the position of the sound source in space. To simulate an acoustic scene with sound sources in different directions, the audio content from these sound sources needs to be modified according to their direction. In binaural audio, this simulation can be achieved using directionally dependent acoustic filters, which can be called head-related impulse responses (HRIRs) in the time domain or head-related transfer functions (HRTFs) in the frequency domain. HRTFs are frequency responses that describe the modification (e.g., filtering) of sound along its path from the sound source to the ear canal. HRTFs can be measured in the form of linear time-invariant filters and synthesized using various models for real-time applications.

[0090] Figure 4A and Figure 4B The spatial coordinates of the sound source used to describe the HRTF relative to the center of the user's head are shown. Figure 4A and Figure 4B The spherical coordinate system and the cross section of the head used to specify the location of the sound source are shown. Figure 4A and Figure 4B , the origin of the coordinate system is located at the center of the user's head 400, between the two ear canal entrances. Starting from the origin, the x-axis, y-axis, and z-axis point to the right ear, the front, and the top of the head, respectively. The horizontal plane, the median plane, and the lateral plane can be defined by these three axes. The position of the sound source 410 is defined as (r, θ, φ) in the spherical coordinate system, where r is the distance from the sound source 410 to the origin. θ is the azimuth angle between the y-axis and the horizontal projection of the position vector of the sound source 410, and is defined as -180°<θ≤+180°, where -90°, 0°, +90°, and +180° represent the left, front, right, and back directions on the horizontal (e.g., xy) plane, respectively. φ is the elevation angle between the horizontal plane and the position vector of the sound source 410, defined as -90°≤φ≤+90°, where -90°, 0°, and +90° represent the bottom, front, and top directions in the median (e.g., yz) plane, respectively.

[0091] During the transmission process, the sound emitted from the sound source 410 reaches Figure 4A The two ears shown may be diffracted and reflected by the trunk, head, and auricle. The sound pressure change of the sound generated at the sound source 410 can be expressed as P S (r,θ,φ) represents the sound pressure change at the right ear entrance, which can be expressed as P R The sound pressure change at the left ear entrance can be expressed as P L express. Figure 4A It is also shown that the sound pressure change at the center of the head 400 (for example, when there is no user) can be represented by P0. The transfer function from the sound source 410 to the left ear of the user can be represented by H R It can be expressed as H R =P R / P S The transfer function from the sound source 410 to the user's right ear can be determined by H L It can be expressed as H L =P L / P S The transfer function from the sound source 410 to the center of the head 400 can be represented by H0, which can be determined according to H0=P0 / P S HRTF can be defined as the sound obtained from the center point (e.g., P0) when the listener is not present to the sound obtained at the listener's ear in the anechoic field (e.g., P L or P R ) is an acoustic transfer function. HRTF characterizes the sound transmission process and explains the overall sound filtering effect of the human anatomy.

[0092] Figure 5A An example of measuring the HRTF of a user is shown. R or H L ) is divided by the origin transfer function at the center of the head without the head (e.g., H0) to offset the influence of the measurement system characteristics to calculate the free-field HRTF. For each sound source direction, a pair of left and right HRTFs can be calculated by dividing the corresponding pair of binaural transfer functions by the origin transfer function using complex division. Figure 5A As shown, in order to determine the HRTF for the direction of a sound source, the microphone 520 can be positioned at the center of the user's head 500, while the subject is not present. In the frequency domain, the output of the speaker 510 can be expressed as P S (f), and the output of the microphone 520 at the center of the user's head 500 can be represented by P0(f). The microphone 520 can then be positioned at the user's ear (e.g., Figure 5A), and in the frequency domain, the output of microphone 520 at the user's right ear can be expressed as P R (f) represents. Then, for a sound source located in the direction of the speaker 510, the HRTF of the right ear can be determined as P R (f) / P0(f).

[0093] Figure 5B An example of generating spatialized audio content based on the user's HRTF is shown. Figure 5B In the example shown, an audio signal is provided as input to an audio controller 530. The audio controller 530 processes the input audio signal to generate a spatialized multi-channel audio signal for presentation to the user via a pair of speakers of a headset or IED. The audio controller 530 may include a set of left-ear filters 532 for implementing HRFT for the left ear, and a set of right-ear filters 534 for implementing HRFT for the right ear. As described above, the HRTFs may be different for sound sources from different directions. Therefore, the set of left-ear filters 532 used to synthesize audio content from a sound source from a first direction may be different from the set of left-ear filters 532 used to synthesize audio content from a sound source from a second direction. Similarly, the set of right-ear filters 534 used to synthesize audio content from a sound source from the first direction may be different from the set of right-ear filters 534 used to synthesize audio content from a sound source from the second direction. A set of left-ear filters 532 and a set of right-ear filters 534 for the direction of the sound source can be selected based on the direction of the sound source and / or the corresponding HRTFs for the left and right ears for the direction of the sound source. The output of the set of left-ear filters 532 can be provided to a transducer in or on the left ear, while the output of the set of right-ear filters 534 can be provided to a transducer in or on the right ear. The user can hear the sound using the left and right ears and perceive the sound as if it comes from the direction of the sound source.

[0094] As described above, the HRTF may depend on, for example, the size and shape of the user's outer ear (e.g., the pinna), the size, shape, and density of the user's head and torso, the acoustic properties of the space in which the sound is played, the direction of the sound source, and the like. HRTFs are typically measured in an anechoic chamber (e.g., using a dummy) to minimize the effects of reflections and reverberation on the measured response. HRTFs can be measured in small increments of azimuth and elevation, and interpolation can be used to synthesize HRTFs for arbitrary spatial positions. Using small increments, it may be necessary to measure HRTFs for many (e.g., more than 100, such as several hundred) spatial positions.

[0095] The HRTF measurement system may include, for example, a sound source (e.g., a loudspeaker) for generating stimuli, two in-ear microphones for recording binaural data, an audio interface for audio input and output, a head tracker for recording user orientation data, an optional display for visualizing current and previously measured orientations, and a computing system for signal processing. The head tracker can provide three Euler angles (tilt, pitch, and yaw), where yaw is the azimuth angle and pitch is the elevation angle. The head tracker can provide the head pose in the form of a quaternion. The sound source can be placed at a defined position relative to the user. The excitation signal can be provided by the computing system, processed by the audio interface (e.g., via a digital-to-analog converter (DAC) and a power amplifier), and reproduced by the loudspeaker. The sound signal reproduced by the loudspeaker can be collected by the in-ear microphones, amplified, digitized by an analog-to-digital converter (ADC) circuit, and transmitted to the computing system. The computing system can then use the measured sound signal and the excitation signal to calculate a pair of HRTFs for the sound source location. HRTFs for other sound source positions can be determined in a similar manner.

[0096] Figure 6A An example of a system for measuring a user's HRTF is shown. Figure 6A The measurement system shown in includes a speaker array 630 and a pair of binaural microphones 640 and 642 worn by a user 610. The measurement system can be placed in an acoustically treated room. In one example, the measurement system can be anechoic below about 500 Hz. It may be useful to collect audio test data from a large number of people of different ages, different body shapes, different genders, and different hair lengths. In some examples, the user 610 can be a human model, for example, which can have physical features representative of an average person (e.g., the size and shape of the auricle, head, and torso, etc.).

[0097] The speaker array 630 can generate test sounds according to instructions from the audio controller of the measurement system. The test sound can be an audible signal suitable for determining the HRTF, or at least some parameters of the HRTF or a certain frequency band. The test sound can have one or more specified characteristics, such as frequency range, volume and transmission length. The test sound can include, for example, a continuous sine wave of a constant frequency, a chirp, some other audio content (e.g., music), or a combination thereof. A chirp signal is a signal whose frequency sweeps up or down over a period of time. The speaker array 630 can include a plurality of speakers 632 positioned to project the sound into a target area. The target area is where the user 610 is located during the measurement. Each speaker of the plurality of speakers 632 can be located at a different respective position or direction relative to the user 610 in the target area. Although the speaker array 630 is in Figure 6A 630 is depicted in two dimensions, but it should be noted that the speaker array 630 may also include speakers in other positions and / or dimensions (e.g., in three-dimensional space). In one example, the speakers 632 in the speaker array 630 may be spaced apart at approximately 10° elevation and approximately 10° azimuth around the entire sphere, thereby producing a total of 612 (36×17) different spatial angles relative to the user 610. In some embodiments, one or more speakers 632 of the speaker array 630 may dynamically change their position (e.g., azimuth and / or elevation) relative to the target area. In the above description, the user 610 is stationary (e.g., the position of the ear within the target area remains substantially constant). In other embodiments, the user 610 may be on a rotatable stage that can position the user 610 at different azimuths relative to the speakers 632.

[0098] In the example shown, binaural microphone 640 is placed in the ear canal of the right ear of user 610, while binaural microphone 642 is placed in the ear canal of the left ear of user 610. In some embodiments, binaural microphones 640 and 642 can be embedded in foam earplugs worn by user 610. Binaural microphones 640 and 642 can collect test sounds emitted by speaker array 630. The collected test sounds can be referred to as audio test data. The audio test data can be used to determine a set of HRTFs. For example, the test sounds emitted by speaker 632 of speaker array 630 are collected as audio test data by binaural microphones 640 and 642. Speaker 632 can have a specific position relative to the head of test user 610. Accordingly, the associated audio test data can be used to determine the specific HRTF of each ear.

[0099] As described above, with small increments, it may be necessary to measure multiple HRTFs for many (e.g., more than 100, such as several hundred) spatial locations. Even if the increments are small, interpolation may cause front-to-back confusion and may be difficult to optimize. In addition, HRTFs vary from person to person because sound propagation varies due to the size and shape of each person's head, torso, and pinnae. Therefore, an HRTF that works for one user may not work for another user. Applying an HRTF measured from a dummy or another person to a specific person may reduce the performance of immersive sound due to differences in personal characteristics. Therefore, individualization of the HRTFs may be required to obtain the desired positioning performance. However, based on information about, for example, Figure 6A The described process of measuring to create a personalized HRTF can be time consuming and computationally intensive, and not scalable to a large number of users.

[0100] According to certain embodiments disclosed herein, a user-worn head-mounted device and an in-ear device (IED) (which may or may not be part of the head-mounted device) can be used to determine a user's personalized HRTF or at least some parameters of the personalized HRTF by, for example, the following steps: acquiring audio signals in a natural environment suitable for HRTF measurement; estimating the locations (e.g., directions) of sound sources of the acquired audio signals; and determining HRTFs for these locations based on the audio signals acquired by the head-mounted device and the IED. In this regard, the head-mounted device and the IED can randomly listen to incidental sounds in the user's natural environment with minimal or no user involvement to gradually add data points associated with different spatial locations to the data point cloud of the user's personalized HRTF, so that a user-specific HRTF across all expected source directions can be constructed over time. The head-mounted device and the IED can be worn by the user for a period of time (e.g., days, weeks, or months) for other purposes (e.g., AR / VR applications) to accumulate HRTFs or HRTF parameters for different directions. In this way, the user's HRTF or HRTF parameters can be determined with minimal or no user involvement and without the use of specialized measurement systems, such as sound attenuation chambers and speaker arrays. In some embodiments, when HRTFs for a sufficient number of directions have been collected, these HRTFs can be interpolated to generate a personalized HRTF for the user for any sound source direction.

[0101] Each sound signal collected by the headset and IED can have a short duration (e.g., a few seconds, a few hundred milliseconds, a few tens of milliseconds, or even less than a few milliseconds in the case of impulsive sounds such as clicks) and / or can have a frequency band that may be within at least a portion of the human hearing range (e.g., between approximately 20 Hz and 20 kHz). An HRTF for the entire human hearing range can be determined using a set of sound signals, each of which can cover a different corresponding frequency range. In some embodiments, an HRTF for a portion of the human hearing range can be determined by averaging the results determined using multiple sound signals to improve accuracy. In some embodiments, different filters can be used to implement different frequency bands of HRTFs, and the filter can be selected based on the HRTFs for each portion of the human hearing range determined using different sound signals covering different portions of the human hearing range. In some embodiments, the HRTF or parameters of the HRTF can be used to personalize the non-personalized HRTF, for example, by personalizing the interaural time difference (ITD) and frequency scaling factors.

[0102] Figure 6B An example of a system for determining a personalized HRTF or parameters of a personalized HRTF using the techniques disclosed herein according to certain embodiments is shown. As shown, in the system disclosed herein, a user 650 can be in the user's normal environment (e.g., living room, office, outdoors, etc.) and can wear a head mounted device 660 and a pair of in-ear devices 670 and 672. The head mounted device 660 can be an AR / VR system that the user uses for AR / VR applications, such as the NED 100 and HMD 200 described above. The head mounted device 660 can include an audio system, such as the one described above with respect to Figure 3 The audio system 300 is described. The audio system may include, for example, a microphone array and an audio controller. The IEDs 670 and 672 may be part of the head mounted device 660 or may be separate from the head mounted device 660 but in communication therewith.

[0103] In some embodiments, the head-mounted device 660 may include a camera system (e.g., a SLAM system) or another sensor or may communicate with the camera system or the other sensor, and the camera system and / or the other sensor may be used to determine the location (e.g., direction) of the object generating the sound. The microphone array and / or IED may be used to collect audio signals in a natural environment, and the audio controller may analyze the collected audio signals to determine whether they are suitable for HRTF measurement. For example, sounds suitable for personalized HRTF measurement may have high spatial stationarity (at least when the sound is collected by the head-mounted device 660 and the in-ear devices 670 and 672, for example, within about one second, within about a few hundred milliseconds, or within about tens of milliseconds), and may also have a high signal-to-noise ratio (SNR), a low reverberation level, a low reverberation time (e.g., a low RT60), and a wide frequency spectrum, etc.

[0104] The direction or position of a sound source suitable for HRTF measurement can be determined based on a direction of arrival (DOA) determined using, for example, two or more microphones (e.g., in a microphone array of the head-mounted device), or one or more cameras on or in communication with the head-mounted device 660. In some embodiments, one or more cameras or one or more position sensors (e.g., an inertial measurement unit (IMU)) on the head-mounted device 660 can be used to determine the relative position of the user's torso with respect to the user's head, as the HRTF may be affected by the relative position of the user's torso with respect to the user's head.

[0105] The audio signal collected by the microphone array can be used to determine a nearly anechoic reference signal for use in determining the head-related transfer function (e.g., by performing beamforming in the estimated direction of the sound source to determine a reference sound signal). The audio signal collected by the IED can be used to determine a HRTF or HRTF parameters for the direction of the sound source (and the relative position of the user's torso with respect to the user's head) by dividing the audio signal collected by the IED by a reference signal determined based on the audio signal collected by the microphone array.

[0106] The IED may include an IED 670 for the right ear and another IED 672 for the left ear, and may be configured to be worn in the user's respective ear canals, such that the IED may be configured to detect sounds arriving at the user's ear canals from a local area. Each IED may include an outward-facing microphone (e.g., toward the local area) for collecting sound signals arriving at the user's ear canals. The collected sound signals may be used to determine the user's HRTF. In some examples, each IED may include an inward-facing speaker (e.g., toward the eardrum) so that the IED can present audio content to the user during normal use of the IED (e.g., when the user is using an AR / VR application). In some embodiments, the IED may include other components, such as a transmitter, a receiver, a power supply, and one or more processors. For example, the IED may be connected to the head-mounted device 660 or another device (e.g., a console or portable device) wirelessly or using a wired connection for transmitting data between the IED and the head-mounted device 660 (or another device).

[0107] The head-mounted device 660 can be implemented as, for example, an eye-mounted device, such as smart glasses. In other embodiments, the head-mounted device 660 can be implemented as a head-mounted display, such as the HMD 200 described above. In some embodiments, the eye-mounted device is a near-eye display (NED), such as the NED 100 described above. In some embodiments, the head-mounted device 660 can be worn on the face of a user so that content (e.g., media content) can be presented to the user using a display component and / or an audio system. In some embodiments, the head-mounted device 660 can be another device that can be worn on the user's head. Examples of media content presented by the head-mounted device 660 include images, videos, audio, or a combination thereof. The head-mounted device 660 may include a frame, and may include other components such as a display component including one or more display elements, a depth camera component (DCA), an audio system, and one or more position sensors, such as those described above with reference to FIG. Figure 1 described.

[0108] The DCA for determining depth information of a local area around the head-mounted device 660 can be used to determine the position or direction of a sound source, which is used to determine the HRTF for the position or direction of the sound source. The DCA may include one or more imaging devices and a DCA controller, and may optionally include an illuminator. In some embodiments, the illuminator illuminates a portion of the local area with light. The one or more imaging devices can capture images of the portion of the local area. The DCA controller can use the captured images and one or more depth determination techniques to calculate the depth information of the portion of the local area. Depth determination techniques may include, for example, direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereo analysis, active stereo analysis (using texture added to the scene by light from the illuminator), another technique for determining the depth of the scene, or a combination thereof.

[0109] Position sensors on the head-mounted device 660 (e.g., position sensor 190) can be used to generate one or more measurement signals in response to movement of the user's head. These measurement signals can be used to determine, for example, the relative position of the user's torso with respect to the user's head. The position sensor can be located on a portion of the frame of the head-mounted view device. The position sensor can include an inertial measurement unit (IMU). Examples of position sensors include: one or more accelerometers; one or more gyroscopes; one or more magnetometers; another suitable type of sensor that detects motion; a type of sensor used for error correction in an IMU, or a combination thereof.

[0110] In some embodiments, the head mounted device 660 may be provided with a simultaneous localization and mapping (SLAM) function for tracking the position of the head mounted device 660 and updating a map of the local environment. For example, the head mounted device 660 may include a passive camera assembly (PCA) that generates color image data. The PCA may include one or more RGB cameras that capture images of part or all of the local area. In some embodiments, some or all of the imaging devices of the DCA may also be used as the PCA. The images captured by the PCA and the depth information determined by the DCA may be used to determine parameters of the local area, generate a model of the local area, update the model of the local area, determine the position of the user, or a combination thereof. For example, the images captured by the PCA and the depth information determined by the DCA may be used to determine the position or direction of a sound source relative to the user.

[0111] The audio system of the head-mounted device 660 (e.g., audio system 300) may include a transducer array (e.g., transducer array 310), a sensor array (e.g., sensor array 320), and an audio controller (e.g., audio controller 330). The transducer array may be used to present sounds to the user. The sensor array may be used to detect sounds that may be suitable for HRTF determination, and may be used to determine the direction of a sound source and / or determine a reference source signal for HRTF determination (e.g., by beamforming a sound signal detectable at a central location on the head). The sensor array may include a plurality of acoustic sensors that detect sounds within a local area of ​​the head-mounted device 660. Each acoustic sensor may be configured to detect sounds and convert the detected sounds into an electronic format (analog or digital). The acoustic sensors may include, for example, acoustic wave sensors, microphones, sound transducers, or similar sensors suitable for detecting sounds. The number and / or location of the acoustic sensors may be determined to optimize the amount of audio information collected and the sensitivity and / or accuracy of the information. The acoustic sensors can be oriented so that they can detect sounds in a wide range of directions around the user of the head mounted device 660.

[0112] The audio controller can process data from the sensor array describing sounds detected by the sensor array and data from the IED describing sounds at the user's ear canal. The audio controller can include a processor and a computer-readable storage medium. The audio controller can use the received data to generate direction of arrival (DOA) estimates, generate acoustic transfer functions (e.g., ATFs and / or HRTFs), track the location of sound sources, form beams in the direction of sound sources, classify sound sources, generate sound filters for speakers, and the like, or a combination thereof.

[0113] For example, the audio controller can be configured to determine whether the sound detected by the sensor array has characteristics within a predetermined range, so that the sound can be suitable for HRTF determination. These characteristics can include, for example, reverberation characteristics, bandwidth characteristics, and spatial stationarity. In one example, the audio system can measure the amount of reverberation present in the scene and the spectrum-time characteristics of the sound detected by the sensor array (for example, every 10ms to 300ms). If the sound is broadband without obvious temporal modulation or reverberation, the audio controller can determine that these characteristics are within a predetermined range and can be suitable for HRTF determination. If the characteristics are determined to be outside the predetermined range and therefore the sound may not be suitable for HRTF determination, the audio controller can continue to monitor the sound detected by the sensor array. The audio system can also determine whether the sound source was spatially stationary during the time period when the sound was recorded, and if the sound source was spatially stationary, the audio system can use the recorded sound for HRTF determination, or if the sound source was not spatially stationary during the time period when the sound was recorded, the audio system can continue to monitor the sound detected by the sensor array.

[0114] If the sound detected by the sensor array has reverberation characteristics and spectral characteristics within a predetermined range, and the sound source is stationary during the time period when the sound is recorded, the audio controller can determine the relative position of the sound source relative to the head-mounted device 660, and optionally determine a confidence value associated with the determined position. The audio controller can use, for example, DOA technology, images from DCA and / or PCA, information from position sensors, or a combination thereof to determine the relative position of the sound source. The relative position can also take into account the position of the user's torso relative to the head. The audio controller can determine the confidence value of the determined position based in part on the difference between the position determined by DOA and the position determined from images and / or information from the position sensors.

[0115] If the confidence value meets a threshold, the audio controller can use the sounds detected by the IED to determine one or more HRTFs associated with the relative position of the determined sound source (e.g., one HRTF for the right ear and one HRTF for the left ear). The multiple sounds detected simultaneously by the sensor array can be used to determine a reference transfer function or reference sound signal, which is used to determine the HRTFs. For example, microphones on the head-mounted device 660 used for DOA and other applications can be used to beamform the determined location using information from different acoustic sensors to emphasize sounds from a specific direction while de-emphasizing sounds from other directions, as described above with respect to the beamforming subsystem 370. In one example, the audio controller can determine a nearly anechoic sound signal based on sound signals measured by two or more acoustic sensors of the sensor array. The beamformed signal can be used as a reference signal, and the signal detected by the IED can be used as a measurement signal to determine the HRTF by, for example, dividing the measurement signal by the reference signal in the frequency domain or time domain. In some examples, the sound signal may be suitable only for determining the HRTF or HRTF parameters for a certain frequency band (e.g., a portion of the complete HRTF that is within the human hearing frequency range). In some examples, the sound signal can be used to determine some parameters of the HRTF in a lower dimensional parameter space, such as parameters of individual filters (e.g., notch filters, bandpass filters, high-shelf filters, and / or low-shelf filters) or parameters for modifying non-individualized HRTFs to generate personalized HRTFs (e.g., frequency scaling factors and personalized interaural time difference (ITD)). For example, an HRTF can be represented by a small set of spatial principal components combined with frequency and individual related weights. In one example, an HRTF can be a linear combination of some basic spectral shapes (or basis functions).

[0116] The determined HRTF for the determined sound source direction, part of the HRTF or some parameters associated with the HRTF can be saved in a data repository that stores a set of HRTFs for various sound source directions. The audio controller can continue to detect sounds in the user's local environment to determine sounds that may be suitable for HRTF determination, and can then use the sounds detected as described above to determine HRTFs for various sound source directions, parts of the HRTF or some parameters associated with the HRTF. In some embodiments, HRTFs, parts of HRTFs or parameters of HRTFs can be added, averaged, weighted averaged (e.g., based on confidence levels) or otherwise combined to generate more accurate and complete HRTFs. Over a sufficiently long period of time (e.g., days or weeks), personalized HRTFs or parameters associated with personalized HRTFs for desired sound source directions or desired resolutions (e.g., elevation intervals less than 10° or 5°, azimuth intervals less than 10° or 5°) can be accumulated. HRTFs or parameters of HRTFs for other sound source directions can be determined using linear or nonlinear interpolation to achieve higher spatial resolution. In some embodiments, as HRTFs are accumulated, they can be projected into a lower-dimensional HRTF parameter space (e.g., the coefficients of a cascade of biquad filters such as infinite impulse response (IIR) filters), thereby enabling spatial interpolation. In this way, over time, a complete set of HRTFs customized for the user can be generated.

[0117] The audio controller can implement a suitable HRTF for the target sound source direction based on the set of HRTFs or HRTF parameters to synthesize the audio content of the target sound source, and provide the synthesized audio content to the transducer array and / or IED to present spatialized audio content to the user. In some embodiments, the lower dimensional parameters of the HRTF and the corresponding sound source direction can be stored in a model or a lookup table so that the sound source direction and the model or the lookup table can be used to retrieve the lower dimensional parameters of the HRTF for the sound source direction. For example, in some embodiments, the audio controller or another processing unit can generate a model and / or a lookup table that maps ITD and filter parameters, which are used to approximate the real HRTF for various target positions (azimuth and / or elevation). In some embodiments, the audio controller or another processing unit can generate a model and / or a lookup table that maps personalized ITD and personalized frequency scaling factors, which are used to personalize the non-personalized HRTF for various target positions (azimuth and / or elevation). The model and / or the lookup table may then be used to retrieve parameters for implementing the HRTF for the target sound source direction based on the target sound source direction.

[0118] Figure 7 Included is a flowchart 700 illustrating an example of a process for determining a user's personalized HRTF using a head-mounted device and an in-ear device in accordance with certain embodiments. The operations in flowchart 700 may be performed using, for example, an audio controller, a head-mounted device, a wearable device, a personal electronic device, a server, or another computing system, or a combination thereof. Although flowchart 1400 may depict multiple operations as a sequential order, some of the multiple operations may be performed in parallel or simultaneously. Furthermore, the order of the operations may be rearranged. The process may have additional steps not included in the flowchart. Some operations may be optional or, in some embodiments, may be omitted. Some operations may be performed more than once.

[0119] The operation at block 710 of flowchart 700 may include receiving a first sound signal associated with a sound from a sound source in a local area of ​​a user. The user may be wearing a head-mounted device, such as NED 100, HMD 200, or head-mounted device 660. The user may also be wearing an in-ear device (e.g., IED 670 and 672) that includes an outward-facing microphone. The sound source may be located in a certain direction relative to the user. Each sound may last, for example, for tens of milliseconds, hundreds of milliseconds, several seconds, or longer. The first sound signal may be detected by, for example, a sensor array (e.g., multiple acoustic sensors 180 or sensor array 320) of an audio system of the head-mounted device. In one example, the sensor array may include two or more acoustic sensors on the frame (e.g., temple) of a pair of glasses and may detect sounds in the local area of ​​the user. In some embodiments, the first sound signal associated with the sound from the sound source may be detected by other acoustic sensors in a device worn or carried by the user. The microphone of the IED may also detect sounds from the sound source in the local area of ​​the user.

[0120] The operation at box 720 may include: determining, based on at least the first sound signal, whether the spectral characteristics of the sound meet predetermined criteria. Sounds suitable for personalized HRTF determination may have, for example, a high signal-to-noise ratio (SNR), a low reverberation level, a low reverberation time (e.g., a low RT60), and a wide frequency spectrum. Therefore, the criteria may include, for example, an SNR greater than a threshold, a reverberation level less than a threshold level, a reverberation time (e.g., RT60) shorter than a threshold length, a frequency greater than a certain threshold frequency (e.g., several kilohertz (kHz)), a frequency bandwidth wider than a threshold range, or a combination thereof. For example, the first sound signal may be analyzed by, for example, an audio controller to determine the SNR, reverberation level, reverberation time (e.g., RT60), and frequency band, thereby determining whether the first sound signal meets the criteria. If it is determined that the first sound signal meets the criteria, the first sound signal may be suitable for HRTF determination.

[0121] The operation at box 730 may include: determining that the sound source is stationary during the time period in which the first sound signal is collected. Sounds that are suitable for personalized HRTF determination may also have high spatial stationarity (at least when the head-mounted device and the in-ear device collect the sound, for example, within a few seconds, one second, hundreds of milliseconds, or tens of milliseconds). The spatial stationarity of the sound source can be determined based on, for example, the position of the sound source tracked by the audio controller, the position of the sound source tracked by the SLAM of the head-mounted device, the position of the sound source tracked using DCA, PCA, or images collected by one or more cameras or other image sensors, or a combination thereof. If the audio controller or another processor determines that the sound source is stationary during the time period in which the first sound signal is collected, the first sound signal may be suitable for HRTF determination.

[0122] The operation at box 740 may include: estimating the relative position of the sound source relative to the user. For example, the audio controller of the head-mounted device can use the sound signals collected by two or more acoustic sensors of the sensor array to estimate the direction of arrival of the sound based on, for example, the time difference between the sound arriving at different acoustic sensors. In some embodiments, alternatively or additionally, the relative position of the sound source relative to the user can be determined based on, for example, the sound source position determined by the SLAM of the head-mounted device, the sound source position determined by images collected using one or more cameras or other image sensors, or a combination thereof. The relative position of the sound source relative to the user can be described using, for example, the azimuth angle of the sound source relative to the user, the elevation angle of the sound source relative to the user, or a combination thereof. In some embodiments, the above description of, for example, Figure 6B In some embodiments, the relative position of the user's torso with respect to the user's head can be determined based on data from one or more position sensors of the head-mounted device.

[0123] The operations at block 750 may include receiving a second sound signal from an in-ear device in an ear of the user, the second sound signal being associated with a sound from a sound source. For example, an acoustic sensor (e.g., a microphone) on an in-ear device (IED) in each ear may sense changes in sound pressure at an ear entrance, and these changes in sound pressure are used to generate a sound signal associated with the sound from the sound source. The IED device and the sensor array may simultaneously acquire the sound signal associated with the sound from the sound source to generate the first sound signal and the second sound signal.

[0124] The operations at block 760 may include determining a HRTF or one or more parameters of the HRTF associated with the relative position of the user to the sound source based on the first sound signal and the second sound signal. For example, the first sound signal may be used to determine a reference sound signal, which may then be used as a reference to determine the HRTF or parameters of the HRTF. In one example, a frequency domain analysis (e.g., a Fourier transform or a z-transform) may be performed on the reference sound signal and the second sound signal to determine a spectrum of the reference sound signal and a spectrum of the second sound signal, and the HRTF may be determined by dividing the spectrum of the second sound signal by the spectrum of the reference sound signal. The reference sound signal may, for example, be a sound signal received by an acoustic sensor positioned at the center of the user's head and may be determined based on a first sound signal measured by two or more acoustic sensors of a sensor array. In one example, the audio controller may use the first sound signal collected by the sensor array (e.g., a microphone) and the determined direction of the sound source to beamform the determined location using information from different acoustic sensors, for example, to emphasize sounds from a particular direction while de-emphasizing sounds from other directions, as described above with respect to the beamforming subsystem 370. The beamformed signal can be used as a reference sound signal, and the signal detected by the IED can be used as a measurement signal to determine the HRTF (or HRIR) by, for example, dividing the measurement signal by the reference signal in the frequency domain (or time domain). One or more parameters of the HRTF may include parameters in a lower-dimensional parameter space, such as parameters of one or more filters for implementing a personalized HRTF, or parameter scaling factors for scaling parameters of the HRTF (e.g., frequency, gain, Q, or other parameters). The HRTF or one or more parameters of the HRTF and the relative position of the sound source (and the relative position of the torso with respect to the head) can be saved to a data repository storing multiple HRTFs for the user.

[0125] The operations in blocks 710 to 760 may be performed repeatedly to collect sounds that appear in the user's local environment and are suitable for HRTF determination, and to determine HRTFs or HRTF parameters for a plurality of different sound source directions to populate the user's personalized HRTF data points until the HRTFs or HRTF parameters for all desired sound source directions are determined, or until the spatial resolution of the data points is higher than the desired spatial resolution. In some embodiments, a model or lookup table may be generated for mapping the relative position of a sound source to one or more corresponding parameters of a corresponding HRFT or HRTF, and the model or lookup table may be used to retrieve the corresponding one or more parameters for the target sound source direction.

[0126] As described above, in some examples, the collected sound signal may be applicable only to determining the HRTF of a specific frequency band (e.g., a portion of the HRTF for the entire human auditory frequency range) or some HRTF parameters. In some examples, the collected sound signal can be used to determine some parameters of the HRTF in a lower dimensional parameter space, such as parameters of individual filters (e.g., notch filters, bandpass filters, high shelf filters and / or low shelf filters), parameters of individual audio time difference and intensity difference renderers (TLDRs), or parameters for modifying non-personalized HRTFs to generate personalized HRTFs (e.g., frequency scaling factors and personalized interaural time difference (ITD) etc.).

[0127] Figure 8 An example of a process for building a personalized HRTF set for a user using a system including a head-mounted device and an in-ear device is shown in accordance with certain embodiments. In the example shown, the process may include three stages. In the first stage, incident sound from a sound source in the user's local environment may be detected and analyzed, and other information about the sound source and the user may also be collected. For example, at 810, a microphone of the in-ear device may detect incident sound arriving at the in-ear device from the sound source, and at 820, a sensor array on the head-mounted device may detect incident sound arriving at the sensor array (e.g., microphone) on the head-mounted device from the sound source. At 850, the incident sound detected by the sensor array may be analyzed for reverberation and masking estimation to determine whether the level and duration of reverberation and the volume of reflections are low enough to avoid masking the direct sound. If the reverberation level / duration is low, the sound may be suitable for HRTF determination. In some embodiments, at 830, one or more cameras of the head-mounted device may capture images of the user's local environment, which are used to determine the location of the sound source. In some implementations, at 840 , one or more position sensors may generate data for torso position determination.

[0128] In the second stage, data collected by the IED, sensor array, camera, and position sensor can be processed to estimate a local HRTF or parameters of a local HRTF for the location of the sound source. For example, at 852, the incident sound detected by the sensor array can be used to perform a spectro-temporal stationarity test to determine whether the sound source is spectrally and temporally stationary. In one example, a set of spectro-temporal filters can be applied to the input signal, and the change in the filter response can be measured over a period of time. If the measured change is low, the signal can be determined to be stationary. Other methods for determining the amount of change in the sound spectrum over time can also be used for spectro-temporal stationarity detection. As described above, in some examples, the audio controller can determine a reference sound signal (e.g., a sound signal received by an acoustic sensor positioned at the center of the user's head) based on the sound signal detected by the sensor array (e.g., via beamforming). The incident sound detected by the sensor array and / or the images captured by one or more cameras can be used to determine the direction of arrival of the incident sound at 860, determine whether the sound source is spatially stationary at 862, and estimate the direction of the sound source if the sound source is determined to be spatially stationary. In one example, the direction of arrival of an incident sound can be determined based on the different arrival times of the incident sound at different acoustic sensors with known locations. In another example, the direction of arrival of a sound source relative to the user can be determined based on the different perspectives of the sound source viewed by two or more cameras at known locations, or other camera-based object localization techniques. In some examples, the relative position of a sound source can be tracked based on the incident sound detected by a sensor array and / or images captured by one or more cameras to determine whether the sound source is spatially stationary. In some examples, the position or direction of a sound source relative to the user can be determined based on the incident sound detected by the sensor array and images captured by one or more cameras. In some embodiments, the confidence level of the determined direction of the sound source can be determined based on the incident sound detected by the sensor array and / or images captured by one or more cameras. As described above, the relative position of the user's torso with respect to the user's head may also affect the HRTF. Therefore, the relative position of the user's torso with respect to the user's head can be determined, for example, using sensor data generated by one or more position sensors. The audio controller may estimate audio properties of a sound signal detected by a microphone on the in-ear device at 812, estimate audio properties of a reference sound signal, and determine an HRTF or parameters of an HRTF for a sound source direction and a torso position at 880, as described above and below, e.g., with respect to Figure 7 as described in block 760.

[0129] In the third stage, at 890, the determined HRTF (or parameters of the HRTF), the corresponding direction of the stationary sound source, and the corresponding relative position of the user's torso with respect to the user's head can be saved to a data repository. The system can continue to collect sounds in the user's local environment to detect sounds that may be suitable for HRTF determination, determine the direction of the sound source, determine the HRTF (or parameters of the HRTF) for the corresponding sound source direction, and save the HRTF (or parameters of the HRTF) for the corresponding sound source direction to the data repository. In this way, over a period of time (e.g., several days or weeks), a set of personalized HRTFs can be generated for the desired sound source direction and / or with a spatial angular resolution above a certain threshold. As described above, in some embodiments, the HRTF (or parameters of the HRTF) for the sound source direction can be determined using sound signals collected at different times, and / or the HRTF (or parameters of the HRTF) can be further processed (e.g., added, averaged, or weighted averaged) to determine the personalized HRTF (or parameters of the HRTF) for the sound source direction. The set of personalized HRTFs may be used to estimate HRTFs (or parameters of HRTFs) for other sound source directions, for example, by interpolation.

[0130] In some embodiments, an audio controller or another processing unit can utilize the user's response to the synthetically generated sounds (e.g., explicitly indicating the apparent direction of the sound source in space, or implicitly reacting to the generated spatial audio) to adjust parameters over time to more closely model the HRTF and provide the user with a more realistic spatial perception.

[0131] As described above, a set of personalized HRTFs can be multi-valued functions personalized for each user. A user's HRTF may include redundant information / patterns. In addition, the HRTFs of multiple users may have similar functional information across the multiple users. Therefore, the parameters in a lower-dimensional parameter space can be used to approximate the HRTFs of multiple users through low-complexity signal processing. For example, in some embodiments, the lower-dimensional parameters of the HRTF can be determined using the technology disclosed herein and may include the ITD and lower-dimensional parameters of the HRTF for the direction of the sound source, such as the parameters of the filter for implementing the HRTF (e.g., the center frequency, gain, and Q value of the filter, or other parameters for defining the filter). In some examples, the lower-dimensional parameters of the HRTF determined using the technology disclosed herein may include personalized ITD and personalized frequency scaling factors for personalizing non-personalized HRTFs. In some embodiments, in order to determine the parameters in the lower-dimensional parameter space for HRTF rendering, a set of parameters (e.g., filter parameters or frequency scaling factors) can be initialized and then optimized to match the HRTF measured for the direction of the sound source. In some embodiments, a machine learning model such as a neural network can be trained to fit the HRTF with lower spatial parameters (e.g., filter parameters or frequency scaling factors) in such a way that these parameters can vary smoothly across space and exhibit similar behavior across different users.

[0132] In some embodiments, the audio controller or another processing unit can generate a model and / or lookup table that maps the ITD and filter parameters, which is used to approximate the real HRTF for various target positions (azimuth and / or elevation). For example, lower dimensional parameters (e.g., parameters of the filter) and the corresponding sound source direction can be saved in a lookup table so that the target sound source direction and the lookup table can be used to retrieve the parameters of the filter for implementing the HRTF. In some embodiments, the model and / or lookup table can be later installed and downloaded from an external server to the audio system. In some embodiments, the model and / or lookup table can be located on an external server, and the audio system requests the filter parameters from the external server by providing the target sound source direction.

[0133] In some embodiments, the HRTF or HRTF parameters determined using the techniques disclosed herein can be used to select and configure appropriate parameters and apply the appropriate parameters to a time and intensity difference renderer (TLDR) to generate spatialized audio content, which can then be provided to the user via a head-mounted device (e.g., a headset) or an in-ear device. For example, the audio system can use information such as the target sound source direction and the target fidelity of the audio rendering to select audio TLDR parameters from a set of possible audio TLDR parameters for generating multi-channel spatialized audio content from a single-channel audio signal. The selected audio TLDR can use static audio filters, dynamic audio filters, delays, or a combination thereof to simulate applying one or more head-related transfer functions (HRTFs) of the user to the audio signal, thereby generating multi-channel spatialized audio content from the input single-channel audio signal. After configuration, the audio TLDR can be applied to an audio signal received on a single channel to generate spatialized audio content corresponding to multiple channels (e.g., a left channel audio signal and a right channel audio signal).

[0134] In one example, the audio TLDR selected and configured can include a series of infinite impulse response (IIR) filters and a pair of delays of cascade. The audio TLDR selected and configured can have a monaural static filter (0,1,2 or more monaural static filters are arranged in this group) of one group of configuration and a monaural dynamic filter (0,1,2 or more monaural dynamic filters are arranged in this group) of one group of configuration connected to this group of monaural static filters. In certain embodiments, the audio TLDR selected and configured can also include the binaural static filter that can perform, for example, individualized left / right loudspeaker equalization. In certain embodiments, the audio TLDR selected and configured can also include a binaural dynamic filter of one group of one or more configurations in each channel of multiple audio channels (for example, connected left channel and connected right channel). In addition, in certain embodiments, the audio TLDR selected and configured can have the delay of configuration between multiple audio channels.

[0135] In some embodiments, selecting and configuring a specific audio TLDR involves selecting and configuring filters in each of the following in the specific audio TLDR: a set of monaural static filters; a set of monaural dynamic filters; a set of binaural static filters; and a set of binaural dynamic filters. The selection and configuration of filters can be based on, for example, the desired target power consumption of the audio TLDR, the target computational load specification associated with the selected audio TLDR, the target memory footprint associated with the selected audio TLDR, the target sound source direction, the target sound source distance, the target audio fidelity of the audio rendering, or a combination thereof. As described above, the target sound source direction describes the angular position of the virtual sound source relative to the user and can be described by both an azimuth parameter value (e.g., azimuth) and an elevation parameter value (e.g., elevation).

[0136] Using a parametric audio TLDR approach to generate spatialized audio content can have several advantages. One advantage is computational and memory efficiency, as the computational complexity of using a cascaded series of infinite impulse response (IIR) filters can be much lower than the computational complexity of implementing the equivalent impulse response convolution of the HRTF in the time domain (e.g., when using finite impulse response filters), and memory usage can be one to two orders of magnitude smaller. This reduced complexity makes the embodiments described herein implementable even in hardware with low computational and memory resources. Another advantage of this approach is that by using IIR filters, the approximated HRTF can be interpolated, personalized, and manipulated in real time. For example, moving a notch in a time-domain impulse response can be complex, whereas in a parametric framework, the filter's center frequency can be easily adjusted (e.g., by modifying a few parameters in the model, such as values ​​in a lookup table). This provides greater flexibility for personalizing HRTFs and adjusting and correcting filter parameters for individual device equalization or hardware output curves. Another advantage of the parametric audio TLDR approach is its scalability, balancing computational and memory usage to achieve the desired accuracy. For example, in the audio TLDR method, more or fewer filters can be applied to more or less closely approximate the HRTF, as the number of filters used may affect the accuracy of the audio rendering. By increasing the number of filters employed, the method allows the rendering to be modified as needed between different devices or on the same device. For example, when a device has more computing power / battery, it can use an architecture that utilizes more filters to more closely approximate the HRTF. In low battery mode or on a device with limited computing resources, the parametric audio TLDR method can switch to an architecture that uses fewer filters to perform the audio spatialization that is possible for the allocated filter resources. In some embodiments, when considering room acoustics, the direct sound can be spatialized at the highest resolution, while early reflections and late reverberation can be rendered with gradually decreasing detail or accuracy.

[0137] Figure 9 FIG. 9 is a block diagram of an example of a sound filter subsystem 900 in an audio system of a head mounted device according to some embodiments. The sound filter subsystem 900 may be a block diagram of an example of a sound filter subsystem 900 in an audio system of a head mounted device according to some embodiments. Figure 3 1 and 2. The embodiment of the sound filter subsystem 380 described herein. In the example shown, the sound filter subsystem 900 includes an audio TLDR selection module 910, an audio TLDR configuration module 920, and an audio TLDR application module 930. In alternative configurations, the sound filter subsystem 900 may include different modules and / or additional modules, and the functionality of the sound filter subsystem 900 may be distributed among the modules in a manner different from that described herein.

[0138] The audio TLDR selection module 910 can select an audio TLDR from a set of possible audio TLDRs for generating multi-channel spatialized audio content from a single-channel input audio signal. The set of possible audio TLDRs can include a range of audio TLDRs, from audio TLDRs with fewer configured filters to audio TLDRs with more configured filters. Compared to audio TLDRs with an increasing number of cascaded static and dynamic filters (which have a corresponding increase in power consumption, computational load, and / or memory usage requirements), audio TLDRs with fewer filters can have lower power consumption, lower computational load, and / or lower memory usage requirements. As the number of static and dynamic audio filters in the audio TLDR increases, the accuracy of its amplitude spectrum approximation of a given HRTF can correspondingly improve. For example, an audio TLDR with several configured dynamic binaural filters may be able to closely approximate a given HRTF. Therefore, there is a trade-off when selecting an audio TLDR with additional filters: while such an audio TLDR can improve the approximation of a given HRTF when used to generate spatialized audio content, it may also result in a corresponding increase in power consumption, computational load, and memory requirements.

[0139] In some embodiments, the group of possible audio TLDRs may include three audio TLDRs that provide different levels of accuracy when approximating the amplitude spectrum of a given HRTF. In some embodiments, the group of possible audio TLDRs may include: (i) an audio TLDR that uses two biquad filters and a delay and a one-dimensional interpolation lookup table for configuring the filter to provide a first approximation of the given HRTF, (ii) a second audio TLDR that uses six biquad filters, two gain adjustment filters, and a one-dimensional and two-dimensional interpolation lookup table for configuring the filter to provide a second approximation of the given HRTF, and (iii) a third audio TLDR that uses twelve biquad filters and a one-dimensional and two-dimensional interpolation lookup table for configuring the filter to provide a third approximation of the given HRTF. In these embodiments, as the number of filters in the selected audio TLDR increases, the corresponding approximation of the given HRTF is closer to the given HRTF. In addition, each audio TLDR in the group of audio TLDRs can be associated with a specific range of memory usage, computational load, power consumption, etc. In alternative embodiments, the audio TLDRs in the set may have different numbers of static and dynamic filters, including more or less than a pair of binaural biquad filters, etc. In some embodiments, the filters in the audio TLDRs may be coupled differently than described herein.

[0140] The audio TLDR selection module 910 can select a specific audio TLDR from the set of possible audio TLDRs based on certain input parameters. In some embodiments, the input parameters may include target power consumption, target computational requirements, target memory usage, and a target accuracy level for approximating a given HRTF, or a combination thereof. The input parameters may also specify the target fidelity of the audio content rendering as a target frequency response and target signal-to-noise ratio for the rendered audio content. In some embodiments, a weighted combination of the received input parameters may be used to select the audio TLDR. In some embodiments, the audio TLDR selection module 910 may obtain default values ​​for these parameters from the data repository 335 and use these default values ​​when selecting the audio TLDR. For given input parameters (e.g., target memory usage and target computational load), the audio TLDR selection module 910 may use a selection model retrieved from the data repository 335 to select a specific audio TLDR from the set of possible audio TLDRs. The selection model may take the form of a lookup table that maps multiple ranges of input parameter values ​​to each audio TLDR in the set of possible audio TLDRs. In some embodiments, the selection model may map a specific weighted combination of input parameter values ​​to one of the multiple audio TLDRs. Other selection models are also possible. In some embodiments, the audio TLDR selection module 910 can receive an input parameter in the form of a specification for a target accuracy level when approximating a given HRTF. In these embodiments, the audio TLDR selection module 910 can select an audio TLDR from the group of audio TLDRs based on a model. The model can take the form of, for example, a lookup table that maps specific audio TLDRs in the group to achieve a specific accuracy level when approximating a given HRTF. In such embodiments, a virtual and / or physical input mechanism (e.g., a dial) that can be adjusted to specify a target approximate accuracy level can be used to specify the approximate target accuracy level for a given HRTF as an input parameter.

[0141] The audio TLDR configuration module 920 can configure the various filters of the selected audio TLDR to provide an approximation of a given HRTF. In some embodiments, the audio TLDR configuration module 920 can retrieve one or more models for configuring the various filters of the selected audio TLDR from the data repository 335. The audio TLDR configuration module 920 can receive and use input parameters (e.g., target sound source direction) and the retrieved model to configure the filters of the selected audio TLDR. As described above, the HRTFs for different sound source directions can be different. The audio TLDR configuration module 920 can configure the filters to approximate the corresponding HRTF for the target sound source direction, so that the configured audio TLDR can then receive and process a single-channel audio signal to generate spatialized audio content corresponding to a multi-channel audio signal (e.g., left channel audio signal and right channel audio signal) for presentation to the user.

[0142] In some embodiments, the audio TLDR configuration module 920 can configure the selected audio TLDR as a series of infinite impulse response (IIR) filters and fractional delays or non-fractional delays in cascade to generate spatialized audio content corresponding to a multi-channel audio signal (e.g., a left channel audio signal and a right channel audio signal) from an input single-channel audio signal. In some embodiments, the cascaded series of IIR filters can include a biquad filter, which can be a third-order recursive linear filter with two poles and two zeros. The biquad filters used in the embodiments herein can include "shelf" filters and "peak / notch" filters. The parameters of these biquad filters can be specified using filter type (shelf vs. peak / notch) and center frequency / gain / Q triplet parameter values ​​(or frequency band, gain / attenuation, and slope). The cascaded series of IIR filters can include one or more single-channel (i.e., monaural) static filters, one or more monaural dynamic filters, and one or more multi-channel (i.e., binaural) dynamic filters.

[0143] The audio TLDR configuration module 920 can be configured as a scalar value for the fixed (that is, constant relative to the target sound source direction) parameter of each monaural static filter in the selected audio TLDR. The static filter can be configured by the audio TLDR configuration module 920 to simulate those components that are substantially constant and independent of the position relative to the user (for example, the center frequency, gain and Q value configured for the static filter). For example, the static filter can be considered to be approximated to the shape of one or more HRTFs and allows adjustment of the overall color (for example, spectrum profile, equalization, etc.) of the spatialized audio content generated. In one example, the static filter can be adjusted to match the color of the real HRTF so that final binaural output feels more natural to the user from an aesthetic point of view. Therefore, the configuration of the static filter can involve adjusting the parameter value (for example, any one of the center frequency, gain and Q value) of the filter in a manner independent of the sound source position but aesthetically suitable for the user. The audio TLDR configuration module 920 can configure the static filter to be applied to the audio signal received at a single channel. In embodiments where the selected audio TLDR has multiple static filters, the multiple static filters can process the incoming single-channel audio signal in series, in parallel, or in a combination thereof. The static filters can include, for example, a static high-shelf filter, a static notch filter, another type of filter, or a combination thereof.

[0144] The dynamic filters in the selected audio TLDR can process the input audio signal to generate spatialized audio content that appears to originate from a specific spatial position relative to the user. The dynamic filters in the selected audio TLDR can include monaural dynamic filters and binaural dynamic filters. Compared to static filters, the filter parameters of the dynamic filters (either monaural or binaural) can be based in part on the target position relative to the user's position (e.g., specified by azimuth and elevation). The monaural dynamic filter can be coupled to the above-mentioned monaural static filter in a single channel. The binaural dynamic filter can be coupled to each individual channel of multiple audio channels (e.g., a connected left channel and a connected right channel). The binaural dynamic filter can be used to reproduce frequency-dependent interaural intensity differences (ILDs) across the ears, including the contralateral head shadow and the pinna shadow effect observed in the posterior hemifield of view. The binaural dynamic filters can include, for example, peak filters and shelf filters, and can be applied in series to each audio channel signal of multiple audio channels. Although the same type of dynamic filter (e.g., peak filter) can be configured for multiple audio channel signals, the specific shape of each filter can be different. A typical HRTF for a user may have a first peak at approximately 4kHz to 6kHz and a main notch at approximately 5kHz to 7kHz. In some embodiments, the monaural dynamic audio filter may be configured to produce such a main first peak (e.g., at approximately 4kHz to 6kHz) and such a main notch (e.g., at approximately 5kHz to 7kHz) found in the typical HRTF. In alternative embodiments, the binaural dynamic filter may be configured to produce such a main first peak and main notch.

[0145] The audio TLDR configuration module 920 can retrieve one or more models for configuring the selected audio TLDR from the data repository 335. These models may include lookup tables, functions, and models trained using machine learning techniques, or a combination thereof. The retrieved model can map various values ​​of the target sound source direction to corresponding filter parameter values, such as center frequency / gain / Q triplet values ​​(or other combinations of filter parameters that characterize the filter). In some embodiments, the model can be represented as one or more lookup tables that use the input azimuth and / or elevation parameter values ​​to output the linear interpolation value of the triplet value. For example, as described above, the technology disclosed herein can be used to determine the lower dimensional parameters of the HRTF for the sound source direction, such as the parameters of the filter for implementing the HRTF. The lower dimensional parameters (e.g., the parameters of the filter) and the corresponding sound source direction can be stored in a lookup table so that the parameters of the filter for implementing the HRTF can be retrieved using the sound source direction. In some embodiments, the model can map the received azimuth and / or elevation parameter input values ​​to dynamic filter parameters by interpolating the one-dimensional lookup table. In some embodiments, the model can map both received azimuth and elevation parameters to dynamic filter parameters by interpolating a one-dimensional lookup table. In some embodiments, the model can map both received azimuth and elevation parameter input values ​​to dynamic filter parameters by interpolating a two-dimensional lookup table.

[0146] The audio TLDR configuration module 920 can use the model retrieved based on the input target source direction to configure the dynamic filter of the selected audio TLDR with the target frequency / gain / Q triplet value (or other filter parameters). The audio TLDR configuration module 920 can use the retrieved one-dimensional interpolation lookup table to input one of the azimuth or elevation values ​​from the input target sound source direction to obtain filter parameters, such as center frequency / gain / Q triplet value (or other filter parameters). Alternatively, the audio TLDR configuration module 920 can use the retrieved one-dimensional interpolation lookup table to input both the azimuth and elevation values ​​of the input target sound source direction to obtain filter parameters, such as center frequency / gain / Q triplet value. Using a two-dimensional lookup table can achieve an approximation closer to a given HRTF.

[0147] In some embodiments, the audio TLDR configuration module 920 can configure a fractional delay between the left and right audio channels. For example, the audio TLDR configuration module 920 can use a model (e.g., a lookup table) retrieved from the data repository 335 to determine the amount of delay to apply based on an input target location. This delay can be determined based on, for example, the interaural time difference (ITD) between the sound signal received by the IED in the left ear and the sound signal received by the IED in the right ear during the HRTF determination described above. This delay can be saved in a lookup table along with the corresponding sound source direction so that the delay can be retrieved based on the target source direction. The configured delay can be a fractional delay or a non-fractional delay that simulates the delay between sounds incident on different ears based on the location of the sound source relative to the user, thereby reproducing the interaural time difference (ITD). For example, if the sound source is located to the right of the user, the sound from the sound source may be rendered at the right ear first, followed by the left ear. The audio TLDR configuration module 920 can determine the delay by, for example, inputting the target location (e.g., azimuth and / or elevation) into the model (e.g., a lookup table).

[0148] The audio TLDR application module 930 can apply the configured audio TLDR to an audio signal received at a single channel to generate spatialized audio content for multiple audio channels (e.g., a left audio channel and a right audio channel). The audio TLDR application module 930 can ensure that the (mono) audio signal received at a single channel is processed by any monaural static filters and monaural dynamic filters (if any) in the configured audio TLDR. The (possibly processed) audio signal can then be separated into multiple individual signals (e.g., a left signal and a right signal), which are then processed by any binaural filters in the configured audio TLDR. The audio TLDR application module 930 can also ensure that the spatialized audio content generated at each of the multiple channels is provided to a transducer array (e.g., in a headset or IED) for presentation to the user. Thus, the configured set of monaural static filters and the configured set of monaural dynamic filters can be connected via a single channel to receive and output a single-channel audio signal. In addition, the configured binaural dynamic filters can be connected via corresponding left and right channels to receive a single-channel audio signal and output corresponding left and right audio signals. In some embodiments, the audio TLDR application module 930 can also generate spatialized audio content for additional audio channels. The audio TLDR application module 930 can provide the generated spatialized audio content to the transducer array 310 to present the spatialized audio content to the user.

[0149] Figure 101 is a functional block diagram 1000 illustrating an example of an audio TLDR 1005 for processing a single-channel input audio signal and generating spatialized audio content for multiple channels, according to certain embodiments. The audio TLDR 1005 may be an audio TLDR that has been selected and configured by the sound filter subsystem 900. In some embodiments, there may be additional or different elements or elements arranged in a different order than described herein.

[0150] In some embodiments, the input parameters 1010 of the audio TLDR 1005 may include a target sound source direction, such as a target azimuth angle and a target elevation angle. Figure 10 The model 1020 in the audio TLDR 1005 may be a model for obtaining filter parameter values ​​and delays of static filters and dynamic filters, such as a lookup table and a function. In some embodiments, the model 1020 may be obtained from the data repository 335. The model 1020 may be about Figure 9 Thus, in some embodiments, the model 1020 may include one-dimensional and two-dimensional interpolation lookup tables, which may be used to obtain filter parameter values ​​and delay values ​​based on input sound source direction values ​​(e.g., azimuth parameter values ​​and / or elevation parameter values).

[0151] An audio signal may be provided as input to the audio TLDR 1005 at a single audio channel 1032 of the selected audio TLDR 1005. The input audio signal may be processed by the audio TLDR 1005 to generate a spatialized multi-channel audio signal for presentation to a user (e.g., via a head-mounted viewer or IED). The input audio signal may be provided as input to one or more static filters 1060. The static filters 1060 may be the same as described above with respect to Figure 9 Any of the multiple static filters described above, for example, a monaural static filter. The audio signal processed by the static filter 1060 may then be provided to one or more monaural dynamic filters 1070. The monaural dynamic filter 1070 may be one of the monaural dynamic filters described above, for example Figure 9 The monaural dynamic filter 1070 may receive an input audio signal via the single audio channel 1032 and / or the static filter 1060 and may provide a processed output audio signal to one or more binaural dynamic filters 1080 in the plurality of audio channels 1034.

[0152] The binaural dynamic filter 1080 may be the same as described above with respect to, for example Figure 9In some embodiments, the output audio signals from the monaural filters (e.g., one or more static filters 1060 and / or monaural dynamic filters 1070) can be split and provided as input to the binaural dynamic filter 1080 via the plurality of audio channels 1034. A plurality of audio signals can be generated at the plurality of audio channels as outputs of the binaural dynamic filter 1080. In some embodiments, the audio signals in the plurality of audio channels can be processed by the delay unit 1090 to introduce delays between the channels, for example, as described with respect to Figure 9 The spatialized audio content generated at the plurality of audio channels may include audio content output to the left channel 1036 and audio content output to the right channel 1038. Figure 10 An input mono audio signal is depicted flowing through the single audio channel 1032 and the plurality of audio channels 1034 in a particular order, but other embodiments may use a different order to process the mono audio channels through the audio TLDR 1005 to generate multi-channel spatialized audio content.

[0153] Figure 11 An example of an audio TLDR 1100 that can generate spatialized audio content based on an approximation of a personalized HRTF according to some embodiments is shown. In the example shown, the audio TLDR 1100 can be as described above with respect to Figure 9 and Figure 10 The audio TLDR 1100 may be configured based on the following: an input azimuth angle (θ) 1112 and elevation angle (ρ) 1114 specifying the direction of the target sound source; and a model (e.g., a lookup table 1126) that maps the target sound source direction to parameters of the filters and / or delays as described above. In the example shown, a mono audio signal received at a single audio channel 1132 may be processed by the audio TLDR 1100 to generate a multi-channel spatialized audio signal at a left channel 1136 and a right channel 1138.

[0154] The input audio signal received at a single audio channel 1132 can be processed by any static and / or dynamic monaural filter (not shown) and then split and provided as input to multiple audio channels 1134. Since the binaural properties of some of the multiple filters may vary with elevation values, in some embodiments, the inputs to the multiple audio channels 1134 can be scaled by a binaural scaling unit 1116, for example, using the cosine value of the elevation angle (ρ) 1114 of the target sound source. In the example shown, the configured audio TLDR 1100 includes a dynamic binaural filter 1186 and an associated fractional delay 1196. The dynamic binaural filter 1186 can be configured using two-dimensional interpolation lookup tables 1126A, 1126B, 1126C, 1126D, 1126E, and 1126F. These tables can be looked up using both the azimuth and elevation values ​​of the input target sound source directions. Using a two-dimensional lookup table can achieve a close approximation of a given HRTF, such as approximating the spectral shape of any given HRTF with an error of less than one decibel within the range of human hearing (e.g., up to about 20 kHz, such as from about 5 kHz to about 13 kHz).

[0155] In some examples, alternatively or additionally, the HRTF can be personalized by scaling parameters in the HRTF (e.g., frequency, gain, Q, slope, or other parameters) such that the spectral envelope of the HRTF is linearly translated on a logarithmic frequency scale, and / or by personalizing the frequency-dependent phase delay difference between the left HRTF and the right HRTF. In one example, the HRTF can be personalized by personalizing a factor (referred to herein as a frequency scaling factor) used to compress or stretch the amplitude spectrum of the HRTF in the frequency domain.

[0156] In some embodiments, an audio controller or another processing unit may generate a model and / or lookup table that maps personalized ITDs and personalized frequency scaling factors for personalizing non-personalized HRTFs for various target locations (azimuth and / or elevation). In an HRTF, spectral features such as notches and peaks can provide clues to the direction and color of a sound. The frequencies of these spectral features may systematically differ between users, depending on their physical characteristics. Scaling the non-personalized HRTF in the frequency domain can produce a personalized HRTF. Frequency scaling can include resampling the HRTF by changing the sampling frequency while keeping the sampling frequency of the audio signal unchanged, and then processing the audio signal with the resampled HRTF. The ratio of the two sampling frequencies can be referred to as a frequency scaling factor, or simply a scaling factor. Frequency scaling may introduce errors in interaural time difference (ITD). According to certain embodiments, errors in the ITD introduced by frequency scaling can be compensated so that the resulting ITD matches the ITD of the original non-personalized HRTF. In addition, the ITD and frequency scaling factors can be personalized by, for example, adjusting the ITD and frequency scaling factors based on the user's head width, interpupillary distance and / or other anatomical markers; adjusting the ITD and frequency scaling factors based on the user's feedback on the direction of the sound source perceived by the user based on the audio content rendered to the user's two ears using the ITD and frequency scaling factors; adjusting the ITD and frequency scaling factors based on the user's feedback on the signal strength of the audio content rendered to the user's two ears using the ITD and frequency scaling factors; or a combination thereof.

[0157] Figure 12 A block diagram of a system 1200 including a head mounted device 1205 for implementing some examples disclosed herein is depicted in accordance with some embodiments. In some embodiments, the head mounted device 1205 may be Figure 1 NED100, or Figure 2 The HMD 200 of the system 1200 can operate in an artificial reality environment (e.g., a virtual reality environment, an augmented reality environment, a mixed reality environment, or a combination thereof). Figure 12 The example of the system 1200 shown in FIG includes a head mounted device 1205, a console 1215, an input / output (I / O) interface 1210 coupled to the console 1215, a network 1220, a map building server 1225, and an HRTF rendering system 1270. Figure 12The example of system 1200 is shown to include one head mounted device 1205 and one I / O interface 1210, but in other embodiments, system 1200 may include any number of these components. For example, there may be multiple head mounted devices, each having an associated I / O interface 1210, wherein each head mounted device and I / O interface 1210 communicates with console 1215. In alternative configurations, different components and / or additional components may be included in system 1200. Additionally, in some embodiments, in combination with Figure 12 The functionality described in one or more of the components shown in FIG may be combined with Figure 12 The different ways of distributing the components are as follows. For example, some or all of the functions of the console 1215 may be provided by the head mounted device 1205.

[0158] The head mounted device 1205 may include a display assembly 1230, display optics 1232, one or more position sensors 1234 and one or more cameras 1236, an audio system 1238, a communication subsystem 1240, a memory 1242, and one or more other devices 1244, such as an eye tracking subsystem. Some embodiments of the head mounted device 1205 have a Figure 12 In addition, in other embodiments, the components described are different from the components described. Figure 12 The functionality provided by the various components described may be distributed variously among the components of the head mounted device 1205 , or embodied in separate components remote from the head mounted device 1205 .

[0159] Display assembly 1230 can display content to the user based on data received from console 1215. Display assembly 1230 can use one or more display elements (e.g., display element 120) to display content. The display element can be, for example, an electronic display. In various embodiments, display assembly 1230 can include a single display element or multiple display elements (e.g., one display for each eye of the user). Examples of electronic displays include: a liquid crystal display (LCD); an organic light emitting diode (OLED) display; an active-matrix organic light-emitting diode display (AMOLED); a micro-LED display; a light emitting polymer display (LPD); a waveguide display; another type of display; or a combination thereof. Note that in some embodiments, the display element can also include some or all of the functionality of display optics 1232.

[0160] The display optics 1232 can amplify image light received from an electronic display, correct optical errors associated with the image light, and present the corrected image light to one or two eye zones of the head-mounted device 1205. In various embodiments, the display optics 1232 can include one or more optical elements. Examples of optical elements included in the display optics 1232 include: an aperture; a Fresnel lens; a convex lens; a concave lens; a filter; a reflective surface; or any other suitable optical element that affects image light. In addition, the display optics 1232 can include a combination of different optical elements. In some embodiments, one or more of the multiple optical elements in the display optics 1232 can have one or more coatings, such as a partially reflective coating or an anti-reflective coating.

[0161] The magnification and focusing of image light by display optics 1232 allows the electronic display to be physically smaller, lighter, and consume less power than larger displays. Furthermore, magnification can increase the field of view of content presented by the electronic display. For example, the field of view of the displayed content can be such that the displayed content is presented using substantially all of the user's field of view (e.g., approximately 110 degrees diagonally), and in some cases, the displayed content is presented using the entire user's field of view. Furthermore, in some embodiments, the amount of magnification can be adjusted by adding or removing optical elements.

[0162] In some embodiments, display optics 1232 can be designed to correct for one or more types of optical errors. Examples of optical errors include barrel or pincushion distortion, longitudinal chromatic aberration, or lateral chromatic aberration. Other types of optical errors may include spherical aberration, chromatic aberration, or errors due to lens field curvature, astigmatism, or any other type of optical error. In some embodiments, content provided to an electronic display for display is pre-distorted, and display optics 1232 corrects for this distortion when it receives image light from the electronic display (the image light being generated based on the content).

[0163] Each position sensor 1234 is an electronic device that generates data indicating the position of the head-mounted device 1205. Position sensor 1234 can generate one or more measurement signals in response to movement of the head-mounted device 1205. Position sensor 190 can be an example of a position sensor 1234. Examples of position sensor 1234 include: one or more IMUs; one or more accelerometers; one or more gyroscopes; one or more magnetometers; another suitable type of sensor that detects motion; or a combination thereof. Position sensor 1234 can include multiple accelerometers for measuring translational motion (forward / backward, up / down, left / right) and multiple gyroscopes for measuring rotational motion (e.g., pitch, yaw, roll). In some embodiments, the IMU rapidly samples the measurement signals and calculates an estimated position of the head-mounted device 1205 based on the sampled data. For example, the IMU integrates the measurement signals received from the accelerometers over time to estimate a velocity vector, and integrates the velocity vector over time to determine the estimated position of a reference point on the head-mounted device 1205. A reference point is a point that can be used to describe the position of the head mounted device 1205. Although a reference point can generally be defined as a point in space, in practice, a reference point is defined as a point within the head mounted device 1205.

[0164] One or more cameras 1236 may form a depth camera assembly (DCA) for generating depth information of a portion of a local area. The DCA may include one or more cameras and a DCA controller, and optionally an illuminator. Figure 1 The operation and structure of the DCA and camera 1236 are described.

[0165] The audio system 1238 can provide audio content to the user of the head-mounted device 1205. The audio system 1238 can be substantially similar to the audio system 300 described above. The audio system 1238 may include one or more acoustic sensors, one or more transducers, and an audio controller. In some embodiments, the audio system 1238 may include multiple in-ear devices that include microphones and speakers (e.g., transducers). The audio system 1238 can provide spatialized audio content to the user. In some embodiments, the audio system 1238 can request acoustic parameters from the map building server 1225 via the network 1220. These acoustic parameters describe one or more acoustic properties of the local area (e.g., room impulse response, reverberation time, reverberation level, etc.). The audio system 1238 can provide information describing at least a portion of the local area from, for example, a DCA and / or position information of the head-mounted device 1205 from the position sensor 1234. The audio system 1238 can use one or more of the multiple acoustic parameters received from the map building server 1225 to generate one or more sound filters and use these sound filters to provide audio content to the user. In some embodiments, the audio system performs parameter selection for a suitable audio time difference and intensity difference renderer (TLDR) for generating spatialized audio content. The system can use input parameters to select an audio TLDR from a set of possible audio TLDRs for generating spatialized audio content from a single-channel input audio signal (e.g., mono). The selected audio TLDR can be configured using static monaural filters and dynamic monaural filters, and static binaural filters and dynamic binaural filters, and delays to simulate applying an approximation of a given HRTF to the input audio signal. The audio system uses the selected and configured audio TLDR to generate multi-channel spatialized audio content for presentation to the user through a head-mounted viewer. Various audio TLDRs can provide different levels of accuracy when approximating a given HRTF. In some embodiments, input parameters for selecting and configuring an audio TLDR may include target device metrics (eg, target power consumption, target computational load, etc.) and / or a target level of accuracy for approximating HRTFs.

[0166] The communication subsystem 1240 may include, for example, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (e.g., Devices, IEEE 802.11 devices, Wi-Fi devices, WiMax devices, cellular communication facilities, etc.), and / or similar communication interfaces. In some embodiments, system 1200 may include one or more antennas for wireless communication, either as part of the communication subsystem 1240 or as separate components coupled to any portion of the system. Depending on the desired functionality, the communication subsystem 1240 may include separate transceivers to communicate with base transceiver stations and other wireless devices and access points. This may include communication with different data networks and / or network types, such as wide-area networks (WANs), wireless wide-area networks (WWANs), local area networks (LANs), wireless local area networks (WLANs), personal area networks (PANs), or wireless personal area networks (WPANs). A WWAN may be, for example, a WiMax (IEEE 802.16) network. A WLAN may be, for example, an IEEE 802.11x network. A WPAN may be, for example, a Bluetooth network, IEEE 802.15x, or some other type of network. The techniques described herein may also be used for any combination of WAN, LAN, PAN, WWAN, WLAN, and / or WPAN. The communication subsystem 1240 may allow data to be exchanged with a network, other computer systems, and / or any other device described herein. The communication subsystem 1240 may include a device for sending or receiving data (e.g., text, photos, audio, or video). The communication subsystem 1240, one or more processors, and memory 1242 may together constitute at least a portion of one or more of the devices for performing some of the functions disclosed herein.

[0167] Memory 1242 can be coupled to one or more processors. In some embodiments, memory 1242 can provide both short-term storage and long-term storage and can be divided into several units. Memory 1242 can be volatile, such as static random access memory (SRAM) and / or dynamic random access memory (DRAM), and / or the memory can be non-volatile, such as read-only memory (ROM), flash memory, and solid-state drive. In addition, memory 1242 can include removable storage devices, such as secure digital (SD) cards. Memory 1242 can provide storage for computer-readable instructions, data structures, program modules, and other data for system 1200. In some embodiments, memory 1242 can be distributed among different hardware modules. A set of instructions and / or code can be stored on memory 1242. These instructions may take the form of executable code that can be executed by system 1200, and / or may take the form of source code and / or installable code that, once compiled and / or installed on system 1200 (e.g., using any of a variety of commonly used compilers, installers, compression / decompression utilities, etc.), may take the form of executable code. Memory 1242 may include an operating system loaded therein. The operating system may be operable to initiate execution of instructions provided by the plurality of application modules and / or manage other hardware, as well as interface with the communication subsystem 1240, which may include one or more wired and / or wireless transceivers. The operating system may be adapted to perform other operations across the various components of system 1200, including threading, virtualization, resource management, data storage control, and other similar functions. In some embodiments, memory 1242 may store a plurality of application modules, which may include any number of applications.

[0168] In some embodiments, the head-mounted device 1205 may include one or more other devices 1244. Each of the other devices 1244 can be a physical subsystem. Although each of the other devices 1244 can be permanently configured as a structure, some of the other devices 1244 can be temporarily configured to perform a specific function or be temporarily activated. Examples of other devices 1244 can include, for example, an eye tracking unit, a near field communication (NFC) device, a rechargeable battery, a battery management system, and a wired / wireless battery charging system. In some embodiments, one or more functions of the other devices 1244 can be implemented in software.

[0169] The eye tracking unit may include one or more eye tracking systems. Eye tracking may refer to determining the position, orientation, and location of an eye relative to the head-mounted device 1205. The eye tracking system may include an imaging system for imaging one or more eyes and may optionally include a light emitter that generates light directed toward the eye so that light reflected from the eye can be captured by the imaging system. For example, the eye tracking unit may include an incoherent or coherent light source (e.g., a laser diode) and a camera that emits light in the visible or infrared spectrum, and the camera captures light reflected from the user's eye. As another example, the eye tracking unit may capture reflected radio waves emitted by a miniature radar unit. The eye tracking unit may use low-power light emitters that emit light at a frequency and intensity that will not harm the eye or cause physical discomfort. The eye tracking unit may be arranged to increase the contrast of the image of the eye captured by the eye tracking unit while reducing the overall power consumed by the eye tracking unit (e.g., reducing the power consumed by the light emitter and imaging system included in the eye tracking unit). For example, in some embodiments, the eye tracking unit can consume less than 120 milliwatts of power. The head-mounted device 1205 can use the orientation of the eyes to, for example, determine the user's inter-pupillary distance (IPD), determine gaze direction, introduce depth cues (e.g., blurring images outside the user's primary line of sight), provide a concave display to reduce power consumption, collect heuristic information about user interaction in VR media (e.g., time spent on any particular object, object, or frame based on the stimulus experienced), perform some other function based in part on the orientation of at least one of the user's eyes, or any combination thereof. For example, because the orientation of the user's eyes can be determined, the eye tracking unit may be able to determine where the user is looking, and therefore the head-mounted device 1205 can generate images with high resolution / intensity for certain areas of the display and images with lower resolution / intensity for other areas of the display.

[0170] I / O interface 1210 is a device that allows a user to send action requests to console 1215 and receive responses from console 1215. An action request is a request to perform a specific action. For example, an action request may be an instruction to start or end the acquisition of image data or video data, or an instruction to perform a specific action within an application. I / O interface 1210 may include one or more input devices. Example input devices include: a keyboard; a mouse; a game controller; or any other suitable device for receiving action requests and transmitting these action requests to console 1215. Action requests received by I / O interface 1210 are transmitted to console 1215, which performs the action corresponding to the action request. In some embodiments, I / O interface 1210 includes an IMU that collects calibration data indicating an estimated position of I / O interface 1210 relative to an initial position of I / O interface 1210. In some embodiments, I / O interface 1210 may provide tactile feedback to the user based on the instructions received from console 1215. For example, tactile feedback is provided when an action request is received, or the console 1215 transmits an instruction to the I / O interface 1210 when the console 1215 performs an action, thereby causing the I / O interface 1210 to generate tactile feedback.

[0171] The console 1215 provides content to the head mounted device 1205 for processing based on information received from one or more of: the DCA; the head mounted device 1205; and the I / O interface 1210. Figure 12 In the example shown, the console 1215 includes an application repository 1255, a tracking module 1260, and an engine 1265. Some embodiments of the console 1215 have Figure 12 Similarly, the functions described further below may be combined with Figure 12 The described methods are distributed among the components of the console 1215 in different ways. In some embodiments, the functions discussed herein with respect to the console 1215 can be implemented in the head mounted device 1205 or in a remote system.

[0172] The application repository 1255 can store one or more applications for execution by the console 1215. An application is a set of instructions that, when executed by the processor, generates content for presentation to the user. The content generated by the application can be in response to user input received through movement of the head-mounted device 1205 or the I / O interface 1210. Examples of applications include gaming applications, conferencing applications, video playback applications, or other suitable applications.

[0173] The tracking module 1260 can use information from the DCA, one or more position sensors 1234, or a combination thereof to track the movement of the head mounted device 1205 or the I / O interface 1210. For example, the tracking module 1260 can determine the location of a reference point of the head mounted device 1205 in the constructed map of the local area based on the information from the head mounted device 1205. The tracking module 1260 can also determine the location of an object or virtual object. The tracking module 1260 can also determine the location of the object's torso relative to the object's head. In addition, in some embodiments, the tracking module 1260 can use a portion of the data from the position sensor 1234 indicating the location of the head mounted device 1205 and the representation of the local area from the DCA to predict the future location of the head mounted device 1205. The tracking module 1260 can provide the estimated or predicted future location of the head mounted device 1205 or the I / O interface 1210 to the engine 1265.

[0174] Engine 1265 can execute an application and receive position information, acceleration information, velocity information, predicted future position, or a combination thereof of head-mounted device 1205 from tracking module 1260. Based on the received information, engine 1265 can determine content to provide to head-mounted device 1205 for presentation to the user. For example, if the received information indicates that the user has looked to the left, engine 1265 can generate content for head-mounted device 1205 that reflects the user's movement within a virtual local area, or content that reflects the user's movement within a local area that is augmented with additional content. Furthermore, engine 1265 can perform an action within an application executing on console 1215 in response to an action request received from I / O interface 1210, and provide feedback to the user that the action has been performed. The feedback provided can be visual or auditory feedback via head-mounted device 1205, or tactile feedback via I / O interface 1210.

[0175] The network 1220 can couple the head mounted device 1205 and / or the console 1215 to the map building server 1225. The network 1220 can include any combination of local area networks and / or wide area networks using both wireless communication systems and / or wired communication systems. For example, the network 1220 can include the Internet and a mobile phone network. In one embodiment, the network 1220 uses standard communication technologies and / or protocols. Thus, the network 1220 can include links using technologies such as Ethernet, IEEE 802.11, worldwide interoperability for microwave access (WiMAX), 2G / 3G / 4G mobile communication protocols, digital subscriber line (DSL), asynchronous transfer mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, network protocols used on the network 1220 may include multiprotocol label switching (MPLS), transmission control protocol / Internet protocol (TCP / IP), user datagram protocol (UDP), hypertext transport protocol (HTTP), simple mail transfer protocol (SMTP), file transfer protocol (FTP), etc. Data exchanged through the network 1220 may be represented using technologies and / or formats including binary image data (e.g., Portable Network Graphics (PNG)), hypertext markup language (HTML), extensible markup language (XML), etc.In addition, all or some links can be encrypted using conventional encryption technologies, such as secure sockets layer (SSL), transport layer security (TLS), virtual private network (VPN), Internet Protocol security (IPsec), etc.

[0176] The map-building server 1225 may include a database storing virtual models describing multiple spaces, wherein a location in the virtual models corresponds to the current configuration of the local area of ​​the head-mounted device 1205. The map-building server 1225 receives information describing at least a portion of the local area and / or location information of the local area from the head-mounted device 1205 via the network 1220. The user may adjust privacy settings to allow or prevent the head-mounted device 1205 from sending information to the map-building server 1225. The map-building server 1225 may determine a location in the virtual model associated with the local area of ​​the head-mounted device 1205 based on the received information and / or location information. The map-building server 1225 may determine (e.g., retrieve) one or more acoustic parameters associated with the local area based in part on the determined location in the virtual model and any acoustic parameters associated with the determined location. The map-building server 1225 may send the location of the local area and any acoustic parameter values ​​associated with the local area to the head-mounted device 1205.

[0177] The HRTF rendering system 1270 can utilize a machine learning model (e.g., a neural network) to fit the measured HRTF with a parameterized filter. The filter is determined in such a way that the filter parameters vary smoothly across space and exhibit similar behavior across different users. The fitting method can use a neural network encoder and a differentiable decoder using a digital signal processing solution, and a loss function can be used to perform optimization of the weights of the neural network encoder to generate one or more filter parameter models that are fitted to the HRTF database. The HRTF rendering system 1270 can provide the filter parameter model to the audio system 1250 periodically or upon request for generating spatialized audio content for presentation to the user of the head-mounted device 1205. In some embodiments, the provided filter parameter model is stored in a data memory of the audio system 1238.

[0178] One or more components in the system 1200 may include a privacy module that stores one or more privacy settings for user data elements. The user data elements describe the user or the head-mounted device 1205. For example, the user data elements may describe a physical characteristic of the user, an action performed by the user, the location of the user of the head-mounted device 1205, the position of the head-mounted device 1205, the HRTF of the user, etc. The privacy settings (or "access settings") for the user data elements may be stored in any suitable manner, such as, for example, in association with the user data elements, in an index on an authorization server, in another suitable manner, or any suitable combination thereof.

[0179] The privacy settings for a user data element specify how the user data element (or specific information associated with the user data element) may be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, executed, displayed, or identified). In some embodiments, the privacy settings for a user data element may specify a "blacklist" of entities that may not access certain information associated with the user data element. The privacy settings associated with a user data element may specify any suitable granularity of allowing or denying access. For example, some entities may have the right to view the existence of a particular user data element, some entities may have the right to view the content of a particular user data element, and some entities may have the right to modify a particular user data element. Privacy settings may allow a user to allow other entities to access or store a user data element for a limited period of time.

[0180] Privacy settings may allow a user to specify one or more geographic locations from which user data elements may be accessed. Access to or denial of access to a user data element may depend on the geographic location of the entity attempting to access the user data element. For example, a user may allow access to a user data element and specify that the user data element is accessible to an entity only when the user is in a particular location. If the user leaves the particular location, the user data element may no longer be accessible to the entity. As another example, a user may specify that a user data element is accessible only to entities within a threshold distance from the user (e.g., another user of a head-mounted view device in the same local area as the user). If the user subsequently changes location, the entity that had access to the user data element may lose access, while a new set of entities may gain access when they come within the threshold distance of the user.

[0181] System 1200 may include one or more authorization / privacy servers for enforcing privacy settings. A request from an entity for a particular user data element may identify the entity associated with the request, and the user data element may only be sent to the entity if the authorization server determines that the entity is authorized to access the user data element based on the privacy settings associated with the user data element. If the requesting entity is not authorized to access the user data element, the authorization server may prevent the requested user data element from being retrieved, or may prevent the requested user data element from being sent to the entity. Although this disclosure describes enforcing privacy settings in a particular manner, this disclosure contemplates enforcing privacy settings in any suitable manner.

[0182] Embodiments of the present invention may include an artificial reality system, or may be implemented in conjunction with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, and artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with collected (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or a combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that produces a three-dimensional effect to the viewer). In addition, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or a combination thereof, which are used to create content in artificial reality and / or be used in artificial reality in other ways. Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including wearable devices (e.g., head-mounted displays) connected to a host computer system, standalone wearable devices (e.g., head-mounted displays), mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0183] The methods, systems, and devices discussed above are examples. Where appropriate, various embodiments may omit, replace, or add various procedures or components. For example, in alternative configurations, the described methods may be performed in an order different from that described, and / or various stages may be added, omitted, and / or combined. In addition, features described with reference to certain embodiments may be combined in various other embodiments. Different aspects and elements of the various embodiments may be combined in a similar manner. Furthermore, technology is evolving, and therefore many of the multiple elements are examples that do not limit the scope of this disclosure to those specific examples, as defined in the appended claims.

[0184] Specific details are given in the description to provide a thorough understanding of the various embodiments. However, the various embodiments can be practiced without these specific details. For example, well-known circuits, processes, systems, structures, and techniques have been shown without non-essential details to avoid obscuring the various embodiments. This description provides only example embodiments and is not intended to limit the scope, applicability, or configuration of the present invention as defined in the appended claims. Instead, the foregoing description of the embodiments will provide those skilled in the art with an enabling description for implementing the various embodiments. Various changes can be made to the function and arrangement of the elements without departing from the scope of the appended claims.

[0185] In addition, some embodiments are described as processes depicted as flow charts or block diagrams. Although each process can describe multiple operations as a sequential process, many of the multiple operations can be performed in parallel or simultaneously. In addition, the order of these operations can be rearranged. The process can have additional steps not included in the figure. In addition, various embodiments of the method can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the associated tasks can be stored in a computer-readable medium such as a storage medium. Each processor can perform the associated tasks.

[0186] It will be apparent to those skilled in the art that substantial variations can be made based on a variety of specific requirements. For example, customized or dedicated hardware can be used, and / or specific elements can be implemented in hardware, software (including portable software, such as applets), or both. In addition, connections to other computing devices, such as network input / output devices, can be employed.

[0187] Any of the various techniques, operations, methods, programs, algorithms, or codes described herein can be converted into or expressed as a programming language or computer program implemented on a computer, processor, or machine-readable medium. As used herein, the terms "programming language" and "computer program" include any language used to specify instructions to a computer or processor, and include, but are not limited to, the following languages ​​and their derivatives: assembler, Basic, batch files, BCPL, C, C+, C++, Delphi, Fortran, Java, JavaScript, machine code, operating system command languages, Pascal, Perl, PL1, Python, scripting languages, Visual Basic, meta-languages ​​that specify programs themselves, and all first, second, third, fourth, fifth, or later generation computer languages. Databases and other data models and any other meta-languages ​​are also included. There is no distinction between interpreted languages, compiled languages, or languages ​​that use both compiled and interpreted methods. There is no distinction between the compiled version and the source version of a program. Thus, a reference to a program (where a programming language may exist in more than one state (e.g., source code, compiled, object, or linked)) is a reference to any and all of these states. A reference to a program may include the actual instructions and / or the intent of those instructions.

[0188] With reference to the accompanying drawings, components that may include memory may include non-transitory machine-readable media. The terms "machine-readable medium" and "computer-readable medium" may refer to any storage medium that participates in providing data that causes a machine to operate in a particular manner. In the embodiments provided above, various machine-readable media may be involved in providing instructions / code to a processing unit and / or one or more other devices for execution. Additionally or alternatively, a machine-readable medium may be used to store and / or carry such instructions / code. In many embodiments, a computer-readable medium is a physical and / or tangible storage medium. Such media may take a variety of forms, including but not limited to non-volatile media, volatile media, and transmission media. Common forms of computer-readable media include, for example, magnetic and / or optical media such as compact disks (CDs) or digital versatile disks (DVDs), punch cards, paper tape, any other physical medium having a pattern of holes, RAM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), FLASH-EPROM, any other memory chip or cartridge, a carrier wave as described below, or any other medium from which a computer can read instructions and / or code. A computer program product may include code and / or machine-executable instructions that may represent a procedure, function, subroutine, program, routine, application program (App), subroutine, module, software package, class, or any combination of instructions, data structures, or program statements.

[0189] Those skilled in the art will recognize that the information and signals used to convey the messages described herein may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the foregoing description may be represented using voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0190] As used herein, the terms "and" and "or" can include multiple meanings, which are also expected based at least in part on the context in which the terms are used. Generally, "or", if used in connection with a list, such as A, B, or C, is intended to mean A, B, and C (used in an inclusive sense here), as well as A, B, or C (used in an exclusive sense here). In addition, as used herein, the term "one or more" can be used to describe any feature, structure, or characteristic in the singular, or can be used to describe some combination of features, structures, or characteristics. However, it should be noted that this is merely an illustrative example, and the claimed subject matter is not limited to this example. In addition, the term "at least one of", if used in connection with a list, such as A, B, or C, can be interpreted as A, B, C, or any combination of A, B, and / or C (e.g., AB, AC, BC, AA, ABC, AAB, AABBCCC, etc.).

[0191] In addition, although certain embodiments have been described using specific combinations of hardware and software, it will be appreciated that other combinations of hardware and software are possible. Certain embodiments may be implemented solely in hardware, solely in software, or using a combination thereof. In one example, the software may be implemented using a computer program product comprising computer program code or instructions that may be executed by one or more processors to perform any or all of the steps, operations, or processes described in this disclosure, wherein the computer program may be stored on a non-transitory computer-readable medium. The various processes described herein may be implemented on the same processor or on different processors in any combination.

[0192] Where a device, system, component, or module is described as being configured to perform certain operations or functions, such configuration may be accomplished, for example, by designing electronic circuitry to perform the operations; by programming programmable electronic circuitry (e.g., a microprocessor) to perform the operations, such as by executing computer instructions or code, or by a processor or core programmed to execute code or instructions stored on a non-transitory storage medium; or any combination thereof. Processes may communicate using a variety of technologies, including but not limited to conventional technologies for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.

[0193] The specification and drawings are therefore to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, subtractions, deletions, and other modifications and changes may be made thereto within the scope of the appended claims. Therefore, although specific embodiments have been described, these are not intended to be limiting. Various modifications and equivalents are intended to fall within the scope of the appended claims.

Claims

1. A method comprising: receiving a first sound signal from a sensor array of a head mounted device, the first sound signal associated with a sound from a sound source in a local environment of a user of the head mounted device; determining, based on the first sound signal, whether reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determining that the sound source is stationary over a period of time; determining a relative position of the sound source relative to the user; receiving a second sound signal from an in-ear device in an ear of the user, the second sound signal being associated with the sound from the sound source; as well as A head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF is determined based on at least the second sound signal.

2. The method according to claim 1, wherein Determining the relative position of the sound source relative to the user includes determining an azimuth angle of the sound source relative to the user, an elevation angle of the sound source relative to the user, or a combination thereof.

3. The method according to claim 1 or 2, wherein Determining the relative position of the sound source with respect to the user includes: determining a direction of arrival of the sound based on the first sound signal from the sensor array and positions of two or more sensors in the sensor array; determining the relative position of the sound source relative to the user based on images captured by one or more cameras on the head mounted device; or A combination of the above; Optionally, determining the relative position of the sound source relative to the user comprises: determining a confidence level of the determined relative position of the sound source relative to the user.

4. A method according to any preceding claim, wherein: Determining the HRTF of the user associated with the relative position of the sound source or one or more parameters of the HRTF includes: determining a reference sound signal based on the first sound signal and the determined relative position of the sound source; and determining the HRTF or the one or more parameters of the HRTF based on the spectrum of the reference sound signal and the spectrum of the second sound signal; Optionally, determining the reference sound signal includes: performing beamforming in the relative position direction of the sound source based on the first sound signal.

5. The method according to any preceding claim, further comprising: Based on data from one or more position sensors of the head mounted device, a relative position of the user's torso with respect to the user's head is determined.

6. The method according to any preceding claim, further comprising: The HRTF or the one or more parameters of the HRTF and the relative position of the sound source are saved in a data repository, wherein the data repository stores a plurality of HRTFs of the user.

7. A method according to any preceding claim, and one or more of the following: (i) wherein the reverberation characteristics and the spectral characteristics of the sound include signal-to-noise ratio, frequency range, reverberation level, reverberation time, or a combination thereof; (ii) of which: The one or more parameters of the HRTF include a frequency scaling factor or parameters of one or more filters for implementing the HRTF.

8. The method according to any preceding claim, further comprising: A model or lookup table is generated for mapping the relative positions of the sound sources to the one or more parameters of the HRTF.

9. The method according to any preceding claim, further comprising: The operations of the method according to claim 1 are repeatedly performed to determine a plurality of HRTFs or parameters of the plurality of HRTFs associated with a plurality of sound source directions relative to the user.

10. A method according to any preceding claim, wherein The time period is greater than 10 milliseconds.

11. A system comprising: an in-ear device configured to generate a first sound signal associated with a sound from a sound source in a local environment of a user; as well as A head-mounted device, comprising: a sensor array configured to generate a second sound signal associated with the sound; and An audio controller, the audio controller being configured to: determining, based on the second sound signal, whether reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determining that the sound source is stationary over a period of time; determining a relative position of the sound source relative to the user; and Based on at least the first sound signal, a head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF is determined.

12. The system of claim 11, and one or more of the following: (i) wherein the audio controller is configured to determine an azimuth angle of the sound source relative to the user, an elevation angle of the sound source relative to the user, or a combination thereof; (ii) of which: The audio controller is configured to determine the relative position of the sound source with respect to the user by performing an operation comprising the following steps: determining a direction of arrival of the sound based on the first sound signal from the sensor array and positions of two or more sensors in the sensor array; determining the relative position of the sound source relative to the user based on images captured by one or more cameras on the head mounted device; or A combination of the above.

13. The system according to claim 11 or 12, wherein: The audio controller is configured to determine the HRTF or one or more parameters of the HRTF of the user associated with the relative position of the sound source by performing an operation comprising the following steps: determining a reference sound signal based on the first sound signal and the determined relative position of the sound source; as well as determining the HRTF or one or more parameters of the HRTF based on the spectrum of the reference sound signal and the spectrum of the second sound signal; Optionally, the reference sound signal is a sound signal at the center of the user's head determined by performing beamforming in the direction of the relative position of the sound source based on the second sound signal.

14. The system according to any one of claims 11 to 13, wherein: The one or more parameters of the HRTF include a frequency scaling factor or parameters of one or more filters for implementing the HRTF.

15. A system comprising: one or more processors; as well as One or more processor-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: receiving a first sound signal from a sensor array of a head mounted device, the first sound signal associated with a sound from a sound source in a local environment of a user of the head mounted device; determining, based on the first sound signal, whether reverberation characteristics and spectral characteristics of the sound meet predetermined criteria; determining that the sound source is stationary over a period of time; determining a relative position of the sound source relative to the user; receiving a second sound signal from an in-ear device in an ear of the user, the second sound signal being associated with the sound from the sound source; as well as A head-related transfer function (HRTF) of the user associated with the relative position of the sound source or one or more parameters of the HRTF is determined based on at least the second sound signal.