Virtual auditory display filters and associated systems, methods, and non-transitory computer readable media
By using virtual auditory display filters and employing spectrum shaping and image processing techniques to generate and manipulate audio signals, the problem of inaccurate virtual 3D sound positioning in existing technologies is solved, achieving a high-quality personalized audio experience.
Patent Information
- Application Number
- CN202480023363.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2024-03-29
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to accurately provide a personalized audio experience for each user when generating virtual 3D sound, especially in terms of sound localization at near-zero azimuth and zero elevation angles, and existing methods may introduce complexity and inaccuracies.
A virtual auditory display filter is used. By generating a digital filter based on the center frequency distribution of the virtual auditory space location, and utilizing spectrum shaping technology and image processing algorithms, audio signals are generated and manipulated to accurately represent sound in the virtual auditory space.
It achieves accurate audio signal processing in virtual auditory space, providing a high-quality, clear sound experience, accurately representing the details of the original recording, and adapting to the user's personalized head orientation and acoustic environment.
Smart Images

Figure CN121002897A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present technology relates generally to virtual auditory display filters, and more specifically, to generating virtual auditory display filters, applying virtual auditory display filters to audio signals to generate virtual auditory display sounds in a virtual auditory space, and applications related to virtual auditory display filters. BACKGROUND
[0002] Three-dimensional (3D) sound systems can be implemented by arranging multiple loudspeakers in space, allowing sound to arrive from different directions. Headphones, headphones, and earbuds (collectively referred to as headphones) are commonly used to listen to music or other audio. Headphones can use a head-related transfer function (HRTF) to simulate 3D sound. The HRTF can be a compressed representation of how sound waves interact with a person's head and ears. More generally, the HRTF can be used to simulate the effects of sound waves propagating through 3D space. SUMMARY
[0003] In some aspects, the technology described herein relates to one or more non-transitory computer-readable media comprising executable instructions that, when executed by one or more processors of a system, cause the system to perform a method comprising: generating, for each of a plurality of virtual auditory spatial positions, one or more first digital filters, the one or more first digital filters comprising one or more first notch filters comprising one or more first center frequencies, the one or more first center frequencies based on a first generally S-shaped distribution of center frequencies as a function of virtual auditory spatial position, the one or more first notch filters configured to, when applied to a first audio signal, produce one or more first notches in a first frequency spectrum of the first audio signal based on the one or more first center frequencies; generating, for each of the plurality of virtual auditory spatial positions, one or more second digital filters, the one or more second digital filters comprising one or more second notch filters comprising one or more second center frequencies, the one or more second center frequencies based on a second generally S-shaped distribution of center frequencies as a function of virtual auditory spatial position, the one or more second notch filters configured to, when applied to a second audio signal, produce one or more second notches in a second frequency spectrum of the second audio signal based on the one or more second center frequencies; receiving an audio signal, the audio signal having one or more audio sub-signals, an audio sub-signal being associated with a virtual auditory spatial position; for each of the one or more audio sub-signals: based on the virtual auditory spatial position associated with the audio sub-signal, selecting a particular one or more first digital filters and a particular one or more second digital filters; applying the particular one or more first digital filters to the audio sub-signal to obtain a first processed audio sub-signal; and applying the particular one or more second digital filters to the audio sub-signal to obtain a second processed audio sub-signal; generating, based on the plurality of first processed audio sub-signals, a first output audio signal for a first device; generating, based on the plurality of second processed audio sub-signals, a second output audio signal for a second device; and providing the first output audio signal to the first device and the second output audio signal to the second device.
[0004] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the virtual auditory spatial position is a first virtual auditory spatial position, and the method further comprises: receiving a head orientation of the user; and for each of the one or more audio sub-signals, determining a second virtual auditory spatial position based on the first virtual auditory spatial position and the head orientation associated with the audio sub-signal, wherein selecting the particular one or more first digital filters and the particular one or more second digital filters based on the virtual auditory spatial position associated with the audio sub-signal comprises selecting the particular one or more first digital filters and the particular one or more second digital filters based on the second virtual auditory spatial position.
[0005] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the particular one or more first digital filters is a first particular one or more first digital filters, the particular one or more second digital filters is a first particular one or more second digital filters, the head orientation of the user is a first head orientation of the user, and the method further comprises: receiving a personalized audio signal associated with a third virtual auditory spatial position; selecting a second particular one or more first digital filters and a second particular one or more second digital filters based on the third virtual auditory spatial position; applying the second particular one or more first digital filters to the personalized audio signal to obtain a first processed personalized audio signal; applying the second particular one or more second digital filters to the personalized audio signal to obtain a second processed personalized audio signal; generating, based on the first processed personalized audio signal, a third output audio signal for the first device; generating, based on the second processed personalized audio signal, a fourth output audio signal for the second device; providing the third output audio signal to the first device and the fourth output audio signal to the second device; receiving a second head orientation of the user; determining a fourth virtual auditory spatial position based on the second head orientation; determining an amount of change between the third virtual auditory spatial position and the fourth virtual auditory spatial position; and modifying the one or more first digital filters and the one or more second digital filters based on the amount of change.
[0006] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein modifying the one or more first digital filters and the one or more second digital filters based on the amount of change comprises modifying one or more first center frequencies on which the one or more first notch filters are based and one or more second center frequencies on which the one or more second notch filters are based.
[0007] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, the method further comprising generating, using the one or more image processing algorithms, a first notch mask and a second notch mask, the first notch mask specifying a first gain modifier as a function of the virtual auditory spatial location, the second notch mask specifying a second gain modifier as a function of the virtual auditory spatial location, wherein: the one or more first notch filters comprise one or more first center frequencies and a first gain modified by the first gain modifier, and the one or more first notch filters are configured to, when applied to the first audio signal, produce one or more first notches in a first frequency spectrum of the first audio signal based on the one or more first center frequencies and the first gain, and the one or more second notch filters comprise one or more second center frequencies and a second gain modified by the second gain modifier, and the one or more second notch filters are configured to, when applied to the second audio signal, produce one or more second notches in a second frequency spectrum of the second audio signal based on the one or more second center frequencies and the second gain.
[0008] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the one or more image processing algorithms comprise one or more of a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a thresholding function, an edge detection function, and a segmentation function.
[0009] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, the method further comprising: receiving a selection of an acoustic environment; and determining, based on the acoustic environment, a first acoustic environment digital filter and a second acoustic environment digital filter, wherein for each audio sub-signal of the one or more audio sub-signals, applying the particular one or more first digital filters to the audio sub-signal to obtain a first processed audio sub-signal comprises applying the particular one or more first digital filters and the first acoustic environment digital filter to the audio sub-signal to obtain the first processed audio sub-signal, and applying the particular one or more second digital filters to the audio sub-signal to obtain a second processed audio sub-signal comprises applying the particular one or more second digital filters and the second acoustic environment digital filter to the audio sub-signal to obtain the second processed audio sub-signal.
[0010] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the acoustic environment is represented by one or more surround sound arrays, and determining, based on the acoustic environment, a first acoustic environment digital filter and a second acoustic environment digital filter comprises determining, based on the one or more surround sound arrays, the first acoustic environment digital filter and the second acoustic environment digital filter.
[0011] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the one or more first digital filters and the one or more second digital filters are infinite impulse response filters.
[0012] In some aspects, the techniques described herein relate to one or more non-transitory computer-readable media, wherein the first device comprises a first ear-wearable device and the second device comprises a second ear-wearable device.
[0013] In some aspects, the techniques described herein relate to a system comprising at least one processor and at least one memory including executable instructions that, when executed by the at least one processor, cause the system to: generate, for each of a plurality of virtual auditory spatial locations, one or more first digital filters, the one or more first digital filters comprising one or more first notch filters comprising one or more first center frequencies, the one or more first center frequencies based on a first generally S-shaped distribution of center frequencies as a function of the virtual auditory spatial location, the one or more first notch filters configured to, when applied to a first audio signal, produce one or more first notches in a first frequency spectrum of the first audio signal based on the one or more first center frequencies; generate, for each of the plurality of virtual auditory spatial locations, one or more second digital filters, the one or more second digital filters comprising one or more second notch filters comprising one or more second center frequencies, the one or more second center frequencies based on a second generally S-shaped distribution of center frequencies as a function of the virtual auditory spatial location, the one or more second notch filters configured to, when applied to a second audio signal, produce one or more second notches in a second frequency spectrum of the second audio signal based on the one or more second center frequencies; receive an audio signal, the audio signal having one or more audio sub-signals, an audio sub-signal being associated with a virtual auditory spatial location; for each of the one or more audio sub-signals: based on the virtual auditory spatial location associated with the audio sub-signal, select a particular one or more first digital filters and a particular one or more second digital filters; apply the particular one or more first digital filters to the audio sub-signal to obtain a first processed audio sub-signal; and apply the particular one or more second digital filters to the audio sub-signal to obtain a second processed audio sub-signal; generate, based on the plurality of first processed audio sub-signals, a first output audio signal for a first device; generate, based on the plurality of second processed audio sub-signals, a second output audio signal for a second device; and provide the first output audio signal to the first device and the second output audio signal to the second device.
[0014] In some aspects, the technology described herein relates to a system, wherein the virtual auditory spatial position is a first virtual auditory spatial position, and the executable instructions, when executed by the at least one processor, further cause the system to: receive a head orientation of the user; and for each of the one or more audio sub-signals, determine a second virtual auditory spatial position based on the first virtual auditory spatial position associated with the audio sub-signal and the head orientation, wherein the selecting the particular one or more first digital filters based on the virtual auditory spatial position associated with the audio sub-signal and the selecting the particular one or more second digital filters based on the virtual auditory spatial position associated with the audio sub-signal comprises selecting the particular one or more first digital filters and the particular one or more second digital filters based on the second virtual auditory spatial position.
[0015] In some aspects, the technology described herein relates to a system, wherein the one or more first digital filters are first one or more first digital filters, the one or more second digital filters are first one or more second digital filters, the particular one or more first digital filters are first particular one or more first digital filters, the particular one or more second digital filters are first particular one or more second digital filters, the head orientation is a first head orientation, the audio signal having the one or more audio sub-signals is a first audio signal having first one or more audio sub-signals, and the executable instructions, when executed by the at least one processor, further cause the system to: receive a personalized audio signal having a third virtual auditory spatial position; based on the third virtual auditory spatial position, select second particular one or more first digital filters and second particular one or more second digital filters; apply the second particular one or more first digital filters to the personalized audio signal to obtain a first processed personalized audio signal; apply the second particular one or more second digital filters to the personalized audio signal to obtain a second processed personalized audio signal; generate, based on the first processed personalized audio signal, a third output audio signal for the first device; generate, based on the second processed personalized audio signal, a fourth output audio signal for the second device; provide the third output audio signal to the first device, generate the fourth output audio signal to the second device; receive a second head orientation of the user; determine a fourth virtual auditory spatial position based on the second head orientation; determine an amount of change between the third virtual auditory spatial position and the fourth virtual auditory spatial position; and based on the amount of change, select second one or more first digital filters and second one or more second digital filters for use in receiving a second input audio signal having a second one or more audio sub-signals.
[0016] In some aspects, the technology described herein relates to a system, wherein the executable instructions, when executed by the at least one processor, further cause the system to generate, using one or more image processing algorithms, a first notch mask and a second notch mask, the first notch mask specifying a first gain modifier as a function of a virtual auditory spatial position, the second notch mask specifying a second gain modifier as a function of the virtual auditory spatial position, wherein: the one or more first notch filters are generated using one or more first center frequencies based on a first generally S-shaped distribution of center frequencies as a function of the virtual auditory spatial position and a first gain modified by the first gain modifier, and the one or more first notch filters are configured to, when applied to the first audio signal, produce one or more first notches in a first frequency spectrum of the first audio signal based on the one or more first center frequencies and the first gain, and the one or more second notch filters are generated using one or more second center frequencies based on a second generally S-shaped distribution of center frequencies as a function of the virtual auditory spatial position and a second gain modified by the second gain modifier, and the one or more second notch filters are configured to, when applied to the second audio signal, produce one or more second notches in a second frequency spectrum of the second audio signal based on the one or more second center frequencies and the second gain.
[0017] In some aspects, the technology described herein relates to a system, wherein the executable instructions, when executed by the at least one processor, further cause the system to: receive a selection of an acoustic environment; and determine, based on the acoustic environment, a first acoustic environment digital filter and a second acoustic environment digital filter, wherein for each audio sub-signal of the one or more audio sub-signals, applying the particular one or more first digital filters to the audio sub-signal to obtain a first processed audio sub-signal comprises applying the particular one or more first digital filters and the first acoustic environment digital filter to the audio sub-signal to obtain the first processed audio sub-signal, and applying the particular one or more second digital filters to the audio sub-signal to obtain a second processed audio sub-signal comprises applying the particular one or more second digital filters and the second acoustic environment digital filter to the audio sub-signal to obtain the second processed audio sub-signal.
[0018] In some aspects, the technology described herein relates to a system, wherein the one or more first digital filters and the one or more second digital filters are infinite impulse response filters.
[0019] In some aspects, the technology described herein relates to a system, wherein the first device comprises a first ear-wearable device, and the second device comprises a second ear-wearable device.
[0020] In some aspects, the technology described herein relates to a method comprising: generating a first virtual auditory display filter comprising a first set of first functions, one or more first functions, when applied to a first audio signal having a first location in a virtual auditory space, generate a first processed audio signal having a first frequency response with one or more first notches at one or more first center frequencies based on the first location, the one or more first notches having one or more first peak-to-trough depths of at most -10 dB; generating a second virtual auditory display filter comprising a second set of second functions, one or more second functions, when applied to the first audio signal, generate a second processed audio signal having a second frequency response with one or more second notches at one or more second center frequencies based on the first location, the one or more second notches having one or more second peak-to-trough depths of at most -10 dB; receiving a second audio signal having a second location in the virtual auditory space; applying the first virtual auditory display filter comprising a first subset of the first functions selected based on the second location to the second audio signal to generate a third processed audio signal having a third frequency response; applying the second virtual auditory display filter comprising a second subset of the second functions selected based on the second location to the second audio signal to generate a fourth processed audio signal having a fourth frequency response; providing the third processed audio signal to a first sound output device; and providing the fourth processed audio signal to a second sound output device.
[0021] In some aspects, the technology described herein relates to a method wherein the one or more first center frequencies are based on a first generally S-shaped distribution of center frequencies as a function of location in the virtual auditory space and the one or more second center frequencies are based on a second generally S-shaped distribution of center frequencies as a function of location in the virtual auditory space.
[0022] In some aspects, the technology described herein relates to a method, further comprising receiving a head orientation of the user, wherein: applying the first virtual auditory display filter comprising the first subset of first functions based on the second position selection to the second audio signal to generate the third processed audio signal having the third frequency response comprises applying the first virtual auditory display filter comprising the third subset of first functions based on the second position selection and the head orientation to the second audio signal to generate the third processed audio signal having the third frequency response, and applying the second virtual auditory display filter comprising the second subset of second functions based on the second position selection to the second audio signal to generate the fourth processed audio signal having the fourth frequency response comprises applying the second virtual auditory display filter comprising the fourth subset of second functions based on the second position selection and the head orientation to the second audio signal to generate the fourth processed audio signal having the fourth frequency response.
[0023] In some aspects, the technology described herein relates to a method, further comprising: generating, using one or more image processing algorithms, a first notch mask specifying a first depth modifier as a function of a position in a virtual auditory space and a second notch mask specifying a second depth modifier as a function of the position in the virtual auditory space; modifying one or more first peak-trough depths based on the first depth modifier; and modifying one or more second peak-trough depths based on the second depth modifier.
[0024] In some aspects, the technology described herein relates to a method, wherein the one or more image processing algorithms comprise one or more of a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a threshold function, an edge detection function, and a segmentation function.
[0025] In some aspects, the technology described herein relates to a method, wherein the first set of first functions comprises a first infinite impulse response digital filter and the second set of second functions comprises a second infinite impulse response digital filter.
[0026] In some aspects, the technology described herein relates to a method comprising: receiving a set of a plurality of first digital filters, one or more first digital filters being generated for each of a plurality of virtual auditory spatial locations, the one or more first digital filters comprising one or more first notch filters, the one or more first notch filters comprising one or more first center frequencies, the one or more first center frequencies being based on a first generally S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more first notch filters being configured to, when applied to a first audio signal, produce one or more first notches in a first frequency spectrum of the first audio signal based on the one or more first center frequencies; receiving a set of a plurality of second digital filters, one or more second digital filters being generated for each of the plurality of virtual auditory spatial locations, the one or more second digital filters comprising one or more second notch filters, the one or more second notch filters comprising one or more second center frequencies, the one or more second center frequencies being based on a second generally S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more second notch filters being configured to, when applied to a second audio signal, produce one or more second notches in a second frequency spectrum of the second audio signal based on the one or more second center frequencies; receiving a personalized audio signal having a virtual auditory spatial location; selecting a particular one or more first digital filters and a particular one or more second digital filters based on the virtual auditory spatial location; applying the particular one or more first digital filters to the personalized audio signal to obtain a first processed personalized audio signal; applying the particular one or more second digital filters to the personalized audio signal to obtain a second processed personalized audio signal; providing a first output audio signal to a first device based on the first processed personalized audio signal and a second output audio signal to a second device based on the second processed personalized audio signal; receiving a user perception of the first sound output by the first device and the second sound output by the second device; and modifying the set of a plurality of first digital filters and the set of a plurality of second digital filters based on the user perception.
[0027] In some aspects, the technology described herein relates to a method wherein the virtual auditory spatial location is a first virtual auditory spatial location, and wherein modifying the set of a plurality of first digital filters and the set of a plurality of second digital filters based on the user perception comprises: determining a second virtual auditory spatial location based on the user perception; determining an amount of change between the first virtual auditory spatial location and the second virtual auditory spatial location; and modifying the set of a plurality of first digital filters and the set of a plurality of second digital filters based on the amount of change.
[0028] In some aspects, the technology described herein relates to a method, wherein receiving a user perception comprises receiving a head orientation of the user, and wherein determining the second virtual auditory spatial position based on the user perception comprises determining the second virtual auditory spatial position based on the head orientation of the user.
[0029] In some aspects, the technology described herein relates to a method, wherein receiving a user perception comprises receiving one or more gestures of the user, and wherein determining the second virtual auditory spatial position based on the user perception comprises determining the second virtual auditory spatial position based on the one or more gestures of the user.
[0030] In some aspects, the technology described herein relates to a method, wherein the set of multiple first digital filters is a first set of multiple first digital filters, the set of multiple second digital filters is a first set of multiple second digital filters, wherein modifying the set of multiple first digital filters based on the user perception comprises selecting a second set of multiple first digital filters based on the user perception, and wherein modifying the set of multiple second digital filters based on the user perception comprises selecting a second set of multiple second digital filters based on the user perception.
[0031] In some aspects, the technology described herein relates to a method, wherein modifying the set of multiple first digital filters based on the user perception comprises modifying one or more first center frequencies, and wherein modifying the set of multiple second digital filters based on the user perception comprises modifying one or more second center frequencies.
[0032] In some aspects, the technology described herein relates to a method, further comprising: determining a spatialization precision estimate based on the user perception; and providing the spatialization precision estimate. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 FIG. 1 is an illustration of an environment in which virtual auditory display systems and virtual auditory display devices can operate in some embodiments.
[0034] Figure 2A FIG. 3 is a block diagram depicting components of a virtual auditory display system in some embodiments.
[0035] Figure 2B FIG. 5 is a block diagram depicting components of an ear-wearable device in some embodiments.
[0036] Figure 2C FIG. 7 is a block diagram depicting a process for generating an acoustic environment digital filter in some embodiments.
[0037] Figure 2D FIG. 9 is a block diagram depicting operations of a spatialization engine of a virtual auditory display system in some embodiments.
[0038] Figure 3Ais a block diagram of a method of generating and applying a digital filter in some embodiments.
[0039] Figure 3B is a block diagram depicting components of a filter generation system in some embodiments.
[0040] Figures 4A-4C is a plot of a frequency response of a digital audio signal in some embodiments.
[0041] Figure 5A depicts a distribution of center frequencies for the left ear as a function of azimuth (x-axis) and elevation (y-axis).
[0042] Figure 5B depicts a distribution of center frequencies for the right ear as a function of azimuth (x-axis) and elevation (y-axis).
[0043] Figure 6A is a plot of a center frequency of a digital filter as a function of elevation relative to head orientation in accordance with some embodiments.
[0044] Figure 6B is a plot of user experience data for a multiple trial in some embodiments with five different digital filters varying as a function of notch center frequency.
[0045] Figures 7A-7X depicts a parametric modifier mask that can be applied in some embodiments to modify the gain of a digital filter.
[0046] Figure 8A and Figure 8B depicts a head shadow gain resulting from a digital filter in some embodiments.
[0047] Figure 8C depicts an output of applying a digital filter to a digital audio signal in accordance with some embodiments.
[0048] Figure 8D depicts user experience data based on a transfer function of a digital filter and user experience data for a prior art transfer function in accordance with some embodiments.
[0049] Figure 8E depicts an example head-related transfer function (HRTF).
[0050] Figure 9A and Figure 9B depicts a method of generating a digital filter in accordance with some embodiments.
[0051] Figure 10A and Figure 10B depicts a method of applying a digital filter in accordance with some embodiments.
[0052] Figure 10C Methods of generating and applying virtual auditory display filters in some embodiments are depicted.
[0053] Figure 11A and Figure 11B Example user interfaces for displaying representations of virtual audio displays in some embodiments are depicted.
[0054] Figure 11C Example user interfaces for adjusting settings of virtual audio displays in some embodiments are depicted.
[0055] Figure 12 A plurality of images depicting example use cases of display filter techniques in some embodiments.
[0056] Figure 13A and Figure 13B A schematic of a method of personalizing digital filters in some embodiments.
[0057] Figure 14A and Figure 14B A method of personalizing digital filters in some embodiments is depicted.
[0058] Figures 15A-15C Example user interfaces for calibrating virtual auditory display devices in some embodiments are depicted.
[0059] Figures 15D-15F Example user interfaces for personalizing virtual auditory displays of virtual auditory display devices in some embodiments are depicted.
[0060] Figures 15G-15J Example user interfaces for providing information about calibration of virtual auditory display devices and personalization of virtual auditory displays of virtual auditory display devices in some embodiments are depicted.
[0061] Figure 16 A block diagram of an example digital device in some embodiments.
[0062] Throughout the drawings, like reference numerals will be understood to refer to like parts, components, and structures. DETAILED DESCRIPTION
[0063] HRTFs can be specific to a person. Generating individual HRTFs typically requires a highly specialized environment and acoustic testing equipment. A person must remain stationary in an anechoic chamber for about 30 minutes while audio signals are emitted from different known positions. Microphones are placed in each of the person's ears to capture the audio signals. However, this approach presents challenges because of factors such as the chamber, the audio signal sources, and the microphones, there can be stray responses that need to be eliminated in order to obtain accurate head-related impulse responses (HRIRs), which can then be converted to HRTFs. Furthermore, any movement of the person can affect the measurements, which can result in inaccurate HRTFs for the person. Another practical limitation of measuring HRIRs is that the time of direct collection is proportional to the number of discrete coordinates, and in practice limits the resolution of the resulting HRTFs.
[0064] So-called generic HRTFs have been used to overcome the shortcomings of individual HRTFs. Such generic HRTFs can be produced by averaging or otherwise combining measurements from multiple persons. However, such combinations typically result in a loss of individual characteristics of each person that are necessary for producing accurate virtual 3D sound for that person. As a result, such generic HRTFs can not accurately localize sound in a virtual 3D space for all users, especially sound that is directly in front of the user at approximately zero degrees azimuth and zero degrees elevation. Figure 8E An example HRTF 810 is depicted.
[0065] Another existing approach attempts to model individualized HRTFs via high-precision scanning of the head, torso, or pinna using photogrammetry, or using other methods, either time-of-flight or structured light. Physical acoustic models are then generated based on the resulting scanned forms. However, this approach can not produce a convincing virtual 3D space presentation because the physics-based simulation of sound interacting with the modeled surfaces can introduce complexities and inaccuracies in the resulting psychoacoustic cues after the physical scans are measured.
[0066] The technology described herein provides a technical solution to the technical problems of the existing approaches described above. The technology can utilize a virtual auditory display filter that can produce accurate presentations of sound at its location in a virtual auditory space. The virtual auditory display filter can utilize spectral shaping techniques, using equalizers, filters, and / or dynamic range compression to manipulate the spectrum of the audio signals. The virtual auditory display filter can be generated without resorting to direct physical measurements (e.g., measurements in an anechoic chamber, photogrammetry, etc.).
[0067] The virtual auditory display filter can be or include a function that manipulates the spectrum of the audio signal. The virtual auditory display filter can be or include a digital filter, such as a parametric equalization (EQ) filter, that allows for adjustment of parameters such as center frequency, gain, quality (Q or q), cutoff frequency, slope, bandwidth, and / or filter type. The parameters can be set as a function of the location of the sound in the virtual auditory space. The function or digital filter can affect the spectrum of the audio signal by producing notches and peaks in the audio signal. The notches, peaks, and other spectral shaping of the audio signal will produce sounds that are placed accurately in the virtual auditory space. In addition, the notches, peaks, and other spectral shaping of the audio signal produce a processed audio signal that can be used to output high quality, clear sounds that, in the example of a music recording, can accurately represent the original recorded performance and allow a listener to hear nuances and subtleties of the original recorded performance. As described herein, a digital filter can refer to a digital filter, a function, and / or some combination of one or more functions or one or more digital filters.
[0068] The virtual auditory space can be described as a virtual 3D sound environment for a person in which the person can perceive sounds as emanating from any location in the virtual 3D sound environment. In the described technology, each location in the virtual auditory space can have an associated function or digital filter applied to an audio signal having that location. Applying the function or digital filter to the audio signal having the location produces a sound that can be referred to as a virtual auditory display sound that is perceived by the person as coming from that location. Thus, a person who can wear headphones, earbuds, or other ear-worn devices can experience virtual auditory display sounds. Other advantages of the described technology will be apparent.
[0069] Figure 1 FIG. 1 is an illustration of an environment 150 in which virtual auditory display systems and virtual auditory display devices that interface with virtual auditory display systems can operate in some embodiments. As depicted, the environment 150 includes a virtual auditory display system 102 and a virtual auditory display device 100. The virtual auditory display system 102 and the virtual auditory display device 100 can together comprise a system. The virtual auditory display system 102 and the virtual auditory display device 100 can together present sounds in a virtual auditory space to a wearer of the virtual auditory display device 100.
[0070] The virtual auditory display system 102 can include a binaurizer 138. The binaurizer can include a system memory 118 that can include a left ear digital filter map 120a and a right ear digital filter map 120b. The binaurizer 138 can also include a left ear convolution engine 116a, a right ear convolution engine 116, and a spatialization engine 114. The virtual auditory display system 102 can include other components, modules, and / or engines, such as described with reference to, for example,Figure 2A Those described.
[0071] In some embodiments, the virtual auditory display system 102 can be or include a software application that can be executed on a digital device. A digital device is any device having at least one processor and memory. For example, reference will be made herein to a digital device that includes a processor and memory. It is to be understood that the virtual auditory display system 102 can be executed on any digital device having at least one processor and memory. Figure 16 Further discussion of digital devices is provided. For example, the virtual auditory display system 102 can be a software application that is executed on a general purpose computing device, such as a laptop or desktop computer. As another example, the virtual auditory display system 102 can be a software application that is executed on a mobile device, such as a phone or tablet computer. In other embodiments, the virtual auditory display system 102 can be or include a software application or firmware application that is executed on a special purpose computing device, such as the virtual auditory display device 100.
[0072] The virtual auditory display device 100 can include a first ear-worn device 102a and a second ear-worn device 102b. The first ear-worn device 102a and the second ear-worn device 102b can each be any ear-worn, ear-mounted, or ear-proximal device, such as an earphone in a pair of headphones, an earbud in a pair of earbuds, a headset, a speaker for a virtual reality headset, etc. In some embodiments, the virtual auditory display device 100 can be an embodiment of a virtual auditory display device as described in the above-referenced co-pending U.S. Patent Application No. __________, filed on the same day as this application and entitled “VIRTUAL AUDITORY DISPLAY DEVICES AND ASSOCIATED SYSTEMS, METHODS, AND DEVICES.” The first ear-worn device 102a and / or the second ear-worn device 102b can include components such as an inertial measurement unit (IMU), an accelerometer, a gyroscope, and / or a magnetometer that detect a head orientation of a wearer wearing the first ear-worn device 102a and the second ear-worn device 102b.
[0073] In some embodiments, a digital device (e.g., a laptop or desktop computer) can receive an encoded audio file 106 having one or more channels. Examples of encoded audio files 106 include 2.0 (two channels), 2.1 (three channels), 5.1 (six channels), 7.1.4 (twelve channels), and 9.1.6 (sixteen channels). The digital device can decode the encoded audio file 106 to obtain decoded audio objects 108 and an input audio signal 112 including one or more audio sub-signals (alternatively, channels). Each of the decoded audio objects 108 and / or audio sub-signals can have an associated coordinate identifying a location of the audio object in a virtual auditory space. The coordinate can be a Cartesian coordinate, a spherical coordinate, and / or a polar coordinate. Although specific examples of encoded audio files are described herein, the technology is not limited to such examples, and can be used with audio files having any number of channels.
[0074] The digital device can send the coordinates 110 to a spatialization engine 114 and send the input audio signal 112 to a left ear convolution engine 116a and a right ear convolution engine 116b. In some embodiments, the virtual auditory display system 102 receives the encoded audio file 106 and decodes the encoded audio file 106 to obtain the decoded audio objects 108 and the input audio signal 112.
[0075] As described with reference to, for example, Figure 11A and Figure 11B A user interface component of the virtual auditory display system 102 can provide a user interface that allows a user to select an acoustic environment, as described with reference to, for example,
[0076] As described with reference to, for example, Figures 15A-15J A user interface component of the virtual auditory display system 102 can provide a user interface 128 that allows a user to perform a calibration and / or personalization procedure 136 to calibrate and / or personalize the virtual auditory display system 102, as described with reference to, for example,
[0077] The spatialization engine 114 can determine the first acoustic environment digital filter and the second acoustic environment digital filter based on the acoustic environment 132. The acoustic environment digital filter can be or include a digital filter that is applied to an audio signal to manipulate the audio signal in order to produce an effect of audio that is played, generated, or produced in a particular acoustic environment. The spatialization engine 114 can provide the first acoustic environment digital filter to the left ear convolution engine 116a and the second acoustic environment digital filter to the right ear convolution engine 116b.
[0078] While the virtual auditory display system 102 is receiving the input audio signal 112, one or both of the first ear-wearable device 102a and the second ear-wearable device 102b can detect a head orientation of a wearer of the virtual auditory display device 100 and provide the head orientation and the audio source distance 126 (which can be specified by the wearer) to the virtual auditory display system 102.
[0079] For each of the one or more audio sub-signals, the binauralizer 138 can obtain the plurality of first processed audio sub-signals and the plurality of second processed audio sub-signals. The binauralizer 138 can do so by determining a particular first location of the audio sub-signal in the virtual auditory space based on the virtual auditory spatial location associated with the audio sub-signal and the head orientation. The left ear digital filter map 120a maps locations in the virtual auditory space to digital filters and / or functions of the first ear-wearable device 102a, and the right ear digital filter map 120b maps locations in the virtual auditory space to digital filters and / or functions of the second ear-wearable device 102b.
[0080] The virtual auditory display filter can be or include a function and / or a digital filter that the virtual auditory display system 102 applies to an audio signal to create a virtual auditory display sound. Reference is made to, for example, Figure 3A and Figure 3B The generation system, discussed in more detail below, can generate the virtual auditory display filter that the virtual auditory display system 102 applies to an audio signal.
[0081] The binauralizer 138 can select the particular first digital filter and / or function from the left ear digital filter map 120a in the system memory 118 and the particular second digital filter and / or function from the right ear digital filter map 120b. The binauralizer 138 can provide the particular first digital filter and / or function to the left ear convolution engine 116a and the particular second digital filter and / or function to the right ear convolution engine 116.
[0082] The left ear convolution engine 116a can apply the particular first digital filter and / or function and the first acoustic environment digital filter to the audio sub-signals to obtain first processed audio sub-signals. The left ear convolution engine 116a can then generate an output audio signal 122a for the first ear-wearable device 102a based on the plurality of first processed audio sub-signals. The right ear convolution engine 116b can apply the particular second digital filter and / or function and the second acoustic environment digital filter to the audio sub-signals to obtain second processed audio sub-signals. The right ear convolution engine 116b can then generate an output audio signal 122b for the second ear-wearable device 102b based on the plurality of second processed audio sub-signals. Plot 124a depicts an example impulse response of the output audio signal 122a, and plot 124b depicts an example impulse response of the output audio signal 122b.
[0083] Figure 2A is a block diagram depicting components of the virtual auditory display system 102 in some embodiments. The virtual auditory display system 102 can include a binauralizer 138, a communication module 202, an audio input module 204, an audio output module 206, a calibration and personalization module 208, a user interface module 210, and a data store 220.
[0084] The communication module 202 can send requests and / or data between components of the virtual auditory display system 102 and any other components or devices, such as the virtual auditory display device 100 and the generation system 380 (see, e.g., FIG. 1). Figure 3A and Figure 3B The communication module 202 can also receive requests and / or data between components of the virtual auditory display system 102 and any other components or devices.
[0085] The audio input module 204 can receive the input audio signal 112 from, e.g., a general-purpose computing device on which the virtual auditory display system 102 executes. The audio output module 206 can provide the output audio signal 122a to the first ear-wearable device 102a and the output audio signal 122b to the second ear-wearable device 102b.
[0086] The calibration and personalization module 208 can calibrate the IMUs and / or other sensors of the first ear-wearable device 102a and the second ear-wearable device 102b. The calibration and personalization module 208 can also generate a personalized audio signal and receive personalization information for personalizing filters. The user interface module 210 can provide a user interface that, among other things, allows a user to select an acoustic environment, select an audio visualization, control a volume, and request that a calibration and / or personalization procedure be performed by the virtual auditory display system 102.
[0087] The data storage 220 can include data stored, accessed, and / or modified by any of the engines, components, modules, etc. of the virtual auditory display system 102. The data storage 220 can include any number of data storage structures, such as tables, databases, lists, etc. The data storage 220 can include data stored in memory (e.g., random access memory (RAM)), on disk, or some combination of in memory and on disk.
[0088] Figure 2B is a block diagram depicting components of the first ear-wearable device 102a and the second ear-wearable device 102b in some embodiments. The first ear-wearable device 102a can include a memory 250, an IMU sensor system 252 (inertial measurement unit sensor system), a magnetometer 254, a microcontroller 256, a power management component 258, an audio DSP 260 (audio digital signal processor), a microphone 262, and a speaker 264. The second ear-wearable device 102b can include a memory 250, an IMU sensor system 252 (inertial measurement unit sensor system), a magnetometer 254, an audio DSP 260, a microphone 262, and a speaker 264.
[0089] The memory 250 can store software and / or firmware. The IMU sensor system 252 and / or the magnetometer 254 can detect a head orientation of a wearer of the virtual auditory display device 100 and / or user interactions with the virtual auditory display device 100. The microcontroller 256 can execute software and / or firmware stored in the memory 250 or a storage of the microcontroller 256.
[0090] The power management component 258 can provide power management. The audio DSP 260 can process audio signals to perform functions such as noise cancellation. The microphone 262 can capture audio, such as ambient audio and / or audio from a wearer of the first ear-wearable device 102a. The speaker 264 can output sound based on the output audio signal 122a and the output audio signal 122.
[0091] The first ear-wearable device 102a and / or the second ear-wearable device 102b can include components other than those depicted in Figure 2B FIG. 1, such as switches, interconnects, and oscillators. The first ear-wearable device 102a can be a primary device, and the second ear-wearable device 102b can be a secondary device. As such, the second ear-wearable device 102b can not include the microcontroller 256. In some embodiments, the second ear-wearable device 102b includes the microcontroller 256.
[0092] The virtual auditory display system 102, the first ear-wearable device 102a, the second ear-wearable device 102b, or the generation system 380 (see, e.g., FIG. 1) can include additional components, such as a user interface, a communication interface, and / or a power source. Figure 3BThe engines, components, modules, etc. described herein can be hardware, software, firmware, or any combination thereof. For example, each engine, component, module, etc. can include functionality executed by a special-purpose hardware (e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.), software, instructions maintained in memory, and / or any combination thereof. The software and / or firmware can be executed by one or more processors.
[0093] Although a limited number of engines, components, and modules are depicted in FIGS. 1-3, any number of engines, components, and modules, etc. can exist. Further, the various engines, components, and modules can perform any number of functions, including the functions of the various modules described herein. Further, although the virtual auditory display system 102, the first ear-wearable device 102a, the second ear-wearable device 102b, and the generation system 380 can be depicted as having a single one of several engines, components, or modules, the virtual auditory display system 102, the first ear-wearable device 102a, and the second ear-wearable device 102b, and the generation system 380 can have multiple engines, components, modules, etc. that perform specific functions. For example, the first ear-wearable device 102a is depicted as having a single audio DSP 260, but the first ear-wearable device 102a can include multiple audio DSPs 260. Figure 2A Figure 2B Figure 3B Although a limited number of engines, components, and modules are depicted in FIGS. 1-3, any number of engines, components, and modules, etc. can exist. Further, the various engines, components, and modules can perform any number of functions, including the functions of the various modules described herein. Further, although the virtual auditory display system 102, the first ear-wearable device 102a, the second ear-wearable device 102b, and the generation system 380 can be depicted as having a single one of several engines, components, or modules, the virtual auditory display system 102, the first ear-wearable device 102a, and the second ear-wearable device 102b, and the generation system 380 can have multiple engines, components, modules, etc. that perform specific functions. For example, the first ear-wearable device 102a is depicted as having a single audio DSP 260, but the first ear-wearable device 102a can include multiple audio DSPs 260.
[0094] Figure 3A is a block diagram of a method 300 of personalizing, generating, and applying a digital filter in some embodiments. The generation system 380 (see Figure 3B ) can perform the generation of the digital filter (steps 304-310), and the virtual auditory display system 102 can perform the personalization of the digital filter (step 302) and the application of the digital filter (steps 312-314).
[0095] The digital filter can be or include one or more parametric equalization (EQ) filters that allow for adjustment of parameters such as center frequency, gain, quality (Q or q), cutoff frequency, slope, bandwidth, and / or filter type. The parametric EQ filter can be or include a biquad filter. The biquad filter can be or include a peak filter, a low shelf filter, and a high shelf filter. In some embodiments, the digital filter can be or include one or more finite impulse response (FIR) filters. The FIR filter can be generated from or based on one or more infinite impulse response (IIR) filters. In some embodiments, the digital filter can be or include one or more IIR filters, or any other suitable type of digital filter.
[0096] The digital filters generated by the generation system 380 can be organized into a plurality of groups. The digital filter groups can include a notch filter group, a head shadow filter group, a shelf filter group, a peak filter group, a beam filter group, a stereo filter group, a post filter group, a top filter group, a top transition filter group, a bottom filter group, a bottom transition filter group, and a wide side filter group. Other groups are possible. Certain digital filters or digital filter groups can be used to set the position of a sound in a virtual auditory space (e.g., the notch filter group). Certain digital filters or digital filter groups can be used to ensure that a sound meets a desired threshold for tone quality, clarity, brightness, etc.
[0097] A digital filter can be or include an algorithm and one or more parameters for the algorithm. For example, the algorithm can be or include a high shelf, low shelf, and peak algorithm. The one or more parameters can be or include a center frequency, a quality (Q or q), a gain, and a sample frequency. For example, a notch digital filter can specify a peak algorithm, an initial center frequency of 6600 Hz, a Q of 15, and an initial gain of -85 decibels (dB). The one or more parameters can be modified. For example, the initial center frequency can be shifted to obtain a shifted center frequency, and the initial gain can be modified by a parameter modifier (see, e.g., the discussion of reference Figures 7A-7X The one or more parameters can be modified. For example, the initial center frequency can be shifted to obtain a shifted center frequency, and the initial gain can be modified by a parameter modifier (see, e.g., the discussion of reference
[0098] A digital filter can be generated and utilized based on how the digital filter represents the geometric shape of a person interacting with a sound wave. For example, a digital filter with a high shelf algorithm can produce a high shelf that can be a virtual representation of how the geometric shape of a person’s concha bowl interacts with a sound wave.
[0099] The method 300 can include a step 302 of calibration and / or personalization. The virtual auditory display system 102 (e.g., the calibration and personalization module 208) can perform calibration of the IMU and / or other sensors of the virtual auditory display device 100 using various devices and / or services 326, such as one or both of the first ear-wearable device 102a and the second ear-wearable device 102b, a cloud-based computing service, and / or a peripheral device of a computing device, such as a camera.
[0100] The virtual auditory display system 102 (e.g., the calibration and personalization module 208) can perform personalization of the virtual auditory display system using various methods and / or techniques, such as: 1) user-directed actions and / or perception of acoustic cues 316; 2) acoustic quality user feedback 318; 3) anatomical measurements 320; 4) demographic information 322; and 5) hearing measurements 324.
[0101] User-directed actions and / or perception of acoustic cues 316 can include capturing a user's response to a location of an acoustic cue. The response can be a user's vocal response captured using a microphone of a computing device, a user's gesture (e.g., head and / or arm movement) captured using a camera of the computing device and / or one or both of the first ear-wearable device 102a and the second ear-wearable device 102b, and user input via a graphical user interface (GUI) of the computing device.
[0102] Acoustic quality user feedback 318 can include user-directed feedback regarding acoustic quality (e.g., responses to questions such as regarding quality metrics of brightness, warmth, clarity, etc., responses to questions provided by a GUI or an audio user interface (AUI)), and observations of user behavior such as user song and / or notification preferences via, for example, the GUI or the AUI.
[0103] Anatomical measurements 320 can include measurements of user anatomical features such as a head, an auricle, and / or a concha via scanning or prediction. Anatomical measurements 320 of one or more users can also include direct measurements (e.g., via a silicone ear impression) and indirect measurements obtained via sensors and / or computer peripherals.
[0104] Demographic information 322 can include user-provided information such as a user's age or other demographics, and a digital fingerprint of the user generated according to one or more user features such as age, gender, and / or other user characteristics.
[0105] Hearing measurements 324 can include those provided by user input and / or obtained through acoustic measurements such as in an anechoic chamber while audio signals are emitted from different known locations.
[0106] The method 300 can include generating a system 380 (e.g., a model generation module 386 of the generating system 380, see Figure 3B Step 304 of generating, modifying, and / or receiving a plurality of models 370. The plurality of models 370 can include one or more outer ear models 354, which can include one or more auricle models 356 and one or more concha models 358. The plurality of models 370 can also include one or more head and torso models 360 and one or more ear canal models 362. The generating system 380 can generate the plurality of models 370 based on the calibration and / or personalization information obtained in step 302.
[0107] For each model of the plurality of models 370, for each location in the virtual auditory space, the generating system 380 can generate one or more first digital filters (for the left ear) and one or more second digital filters (for the right ear) that the model can generate. Thus, for the plurality of models 370, for each location in the virtual auditory space, the generating system 380 can generate a plurality of first digital filters and a plurality of second digital filters.
[0108] For example, for one or more head and torso models 360, the generating system 380 can generate one or more first digital filters and one or more second digital filters that take into account shoulder width and / or amplitude, head diameter, neck height, and other factors. For one or more concha models 358, the generating system 380 can generate one or more first digital filters and one or more second digital filters that represent the acoustic effects of the physical characteristics of the concha. These characteristics include, but are not limited to, concha depth, width, and angle.
[0109] As another example, for one or more pinna models 356, the generating system 380 can generate one or more first digital filters and one or more second digital filters that represent the acoustic effects of the physical characteristics of the pinna. These characteristics include, but are not limited to, pinna height, width, depth, location on the head, and angle of flaring relative to the head. For one or more ear canal models 362, the generating system 380 can generate one or more first digital filters and one or more second digital filters that take into account the physical proportions of the pinna, concha, and other ear components.
[0110] Also at step 304, for each location in the virtual auditory space, the generating system 380 can sum, aggregate, or otherwise combine the plurality of first digital filters into a combined first digital filter and can sum, aggregate, or otherwise combine the plurality of second digital filters into a combined second digital filter. The combined first digital filter can be or include one or more finite impulse response (FIR) filters. The combined second digital filter can be or include one or more FIR filters. Thus, at the end of step 304, for all locations in the virtual auditory space, there can be a set of combined first digital filters and a set of combined second digital filters.
[0111] At step 306, the generating system 380 can generate a mapping or association of the combined first digital filters to their corresponding locations in the virtual auditory space for the left ear. The generating system 380 can generate a mapping or association of the combined second digital filters to their corresponding locations in the virtual auditory space for the right ear. The generating system 380 can utilize Cartesian coordinates, polar coordinates, and / or spherical polar coordinates for the mapping or association.
[0112] At step 308, the generation system 380 can generate a file, database, or other data structure that includes the mapping or association of the combined first digital filters to their corresponding locations in the virtual auditory space, and the mapping and association of the combined second digital filters to their corresponding locations in the virtual auditory space.
[0113] At step 310, the generation system 380 can provide or store the file, database, or other data structure on one or more non-transitory computer-readable media of a device. The device can be the first ear-wearable device 102a and / or the second ear-wearable device 102b, a mobile device such as a phone or tablet, a laptop or desktop computer, another device, or any combination of the foregoing.
[0114] At step 312, the virtual auditory display system 102 can select the combined first digital filters and the combined second digital filters for use. After selection, at step 314, the virtual auditory display system 102 can utilize the combined first digital filters and the combined second digital filters in various applications, such as to present music. For example, with reference to Figure 12 Various applications of the disclosed technology are discussed.
[0115] In some embodiments, at step 304, for each model 370 in the plurality of models, the generation system 380 can generate one or more first digital filters and one or more second digital filters for each azimuth and elevation combination at a location in the virtual auditory space that is at a distance of one meter (1 m) from a center point representing a virtual listener within the virtual auditory space, with an increment of one degree for azimuth and elevation. The increment of one degree for azimuth is from about negative 180 degrees (inclusive) to about positive 180 degrees (inclusive). The increment of one degree for elevation is from about negative 90 degrees (inclusive) to about 90 degrees (inclusive). Thus, there are 65,160 combinations of azimuth and elevation, and thus 65,160 locations in the virtual auditory space, each 1 m from the center point. Thus, the generation system 380 can generate 65,160 sets of one or more first digital filters and 65,160 sets of one or more second digital filters.
[0116] In some embodiments, the method 300 can include a step in which the generation system 380 reduces the number of locations in the virtual auditory space at which digital filters are generated or stored. For example, after step 304, the generation system 380 can select an appropriate subset from the set of combined first digital filters, and an appropriate subset from the set of combined second digital filters.
[0117] In embodiments in which there are 65,160 positions in the virtual auditory space, the generation system 380 can select an appropriate subset from the set of combined first digital filters including approximately 7,000 (such as 7,220) combinations. Similarly, the generation system 380 can select an appropriate subset from the set of combined second digital filters including approximately 7,000 (such as 7,220) combinations.
[0118] The generation system 380 can select appropriate subsets that adequately represent the positions in the virtual auditory space while reducing the amount of storage required for the multiple sets of digital filters and reducing the amount of time to select and process the digital filters. The generation system 380 can achieve these goals in other ways, such as by generating mappings or associations for a reduced number of positions in the virtual auditory space, or storing mappings or associations for a reduced number of positions in the virtual auditory space.
[0119] In some embodiments, at step 304, the generation system 380 does not sum, aggregate, or otherwise combine the multiple first digital filters into combined first digital filters, and does not sum, aggregate, or otherwise combine the multiple second digital filters into combined second digital filters. Thus, at the end of step 304, there can be a set of multiple first digital filters and a set of multiple second digital filters for all of the positions in the virtual auditory space. As described herein, an appropriate subset of the set of multiple first digital filters and an appropriate subset of the set of multiple second digital filters can be utilized.
[0120] In such embodiments, at step 306, the generation system 380 can instead generate mappings or associations of the multiple first digital filters to their corresponding positions in the virtual auditory space of the left ear, and generate mappings or associations of the multiple second digital filters to their corresponding positions in the virtual auditory space of the right ear.
[0121] Further, in such embodiments, at step 308, the generation system 380 can instead generate a file, database, or other data structure including mappings or associations of the multiple first digital filters to their corresponding positions in the virtual auditory space, and mappings or associations of the multiple combined second digital filters to their corresponding positions in the virtual auditory space.
[0122] In some embodiments, the generation system 380 generates multiple sets of digital filters for locations in a virtual auditory space. As described herein, the generation system 380 can generate a first set of digital filters for a left ear and a first set of digital filters for a right ear. The generation system 380 can then generate one or more second sets of digital filters for the left ear and one or more second sets of digital filters for the right ear based on the first set of digital filters for the left ear and the first set of digital filters for the right ear. Each pair of sets can be used to represent a different prototype of a different user population or user grouping.
[0123] The generation system 380 can generate the one or more second sets of digital filters for the left ear and the one or more second sets of digital filters for the right ear by modifying one or more parameters of the digital filters for the left ear and the digital filters for the right ear. For example, the generation system 380 can modify a center frequency of a notch filter included in the first set of digital filters for the left ear and the first set of digital filters for the right ear. The generation system 380 can modify the center frequency of the notch filter to personalize the digital filters to a user, as described with reference to, for example, Figures 15A-15F The generation system 380 can do so to adjust an amount of variation between an actual location of a sound in the virtual auditory space and a location of the sound as perceived by a wearer.
[0124] In some embodiments, the generation system 380 can generate the first set of digital filters for the left ear and the first set of digital filters for the right ear at a distance of 1 m from a center point representing a virtual listener in the virtual auditory space, as described herein. The generation system 380 can generate one or more second sets of digital filters for the left ear and one or more second sets of digital filters for the right ear for other distances from the center point. The generation system 380 can generate the one or more second sets of digital filters for the left ear based on the first set of digital filters for the left ear and generate the one or more second sets of digital filters for the right ear based on the first set of digital filters for the right ear. For example, the generation system 380 can increase a gain of the digital filters for distances less than 1 m from the center point and can decrease a yield of the digital filters for distances greater than 1 m from the center point. Other approaches will be apparent.
[0125] Figure 3B is a block diagram depicting components of the generation system 380 in some embodiments. The generation system 380 can include a communication module 382, a filter generation module 384, a model generation module 386, a parameter generation module 388, a parameter mask module 390, a digital filter tuning module 392, a user interface module 394, and a data store 396.
[0126] The communication module 382 can transmit requests and / or data between the components of the generation system 380 and any other system, component, or device, such as the virtual auditory display system 102. The communication module 382 can also receive requests and / or data between the components of the generation system 380 and any other system, component, or device.
[0127] The filter generation module 384 can generate digital filters and acoustic environment digital filters. The filters can be or include one or more algorithms, and optionally, one or more parameters of the one or more algorithms.
[0128] The model generation module 386 can generate, modify, or access a plurality of models. The parameter generation module 388 can generate parameters for the digital filters.
[0129] The parameter mask module 390 can generate parameter modifier masks. The parameter mask module 390 can generate the parameter modifier masks using image processing techniques. The parameter mask module 390 can determine one or more parameter modifiers for one or more parameters of a filter using the parameter modifier masks. The parameter mask module 390 can modify the one or more parameters using the one or more parameter modifiers.
[0130] The digital filter tuning module 392 can receive parameters for a digital filter from a user and modify the digital filter based on the received parameters. The user interface module 394 can provide a user interface that allows a user to, among other things, listen to a sound output of an audio signal generated from the application of the digital filter and modify the parameters of the digital filter.
[0131] The data storage 396 can include data stored, accessed, and / or modified by any of the engines, components, modules, etc. of the generation system 380. The data storage 396 can include any number of data storage structures, such as tables, databases, lists, etc. The data storage 396 can include data stored in memory (e.g., random access memory (RAM)), on disk, or some combination of in memory and on disk.
[0132] Figures 4A-4C is a plot of the frequency response of a digital audio signal in some embodiments. Figure 4A is a plot 400 of the frequency responses of three audio signals. Each audio signal has a notch at a different center frequency. The center frequency of the notch is a factor in specifying the location of the sound corresponding to the audio signal in a virtual auditory space, which means the location of the sound as perceived by a user (e.g., a wearer of the first ear-wearable device 102a and the second ear-wearable device 102b).
[0133] The first audio signal is for a first sound that has a first position in the virtual auditory space at a distance of one (1) meter (m), zero degrees (0°) azimuth, and zero degrees (0°) elevation. The second audio signal is for a second sound that has a second position in the virtual auditory space at a distance of one (1) meter, five degrees (5°) azimuth, and zero degrees (0°) elevation. The third audio signal is for a third sound that has a third position in the virtual auditory space at a distance of one (1) meter, ten degrees (10°) azimuth, and zero degrees (0°) elevation. Figure 4A A plot 420 of frequency responses of audio signals to which digital filters have been applied is shown in accordance with some embodiments. The frequency responses have three notches at three different center frequencies. The audio signals are for sounds that have positions in the virtual auditory space at a distance of one (1) meter (m), zero degrees (0°) azimuth, and zero degrees (0°) elevation. The virtual auditory display system 102 has applied three notch filters to produce the three notches in the frequency responses of the audio signals. The notch filters can be or include parametric EQ filters with parameters of center frequency, gain, and bandwidth.
[0134] Figure 4B A plot 420 of frequency responses of audio signals to which digital filters have been applied is shown in accordance with some embodiments. The frequency responses have three notches at three different center frequencies. The audio signals are for sounds that have positions in the virtual auditory space at a distance of one (1) meter (m), zero degrees (0°) azimuth, and zero degrees (0°) elevation. The virtual auditory display system 102 has applied three notch filters to produce the three notches in the frequency responses of the audio signals. The notch filters can be or include parametric EQ filters with parameters of center frequency, gain, and bandwidth.
[0135] Figure 4C A plot 440 of frequency responses of two audio signals to which digital filters have been applied is shown in accordance with some embodiments. Each frequency response has a notch at a different center frequency. The virtual auditory display system 102 has applied three notch filters to produce the three notches in each frequency response. The notch filters can be or include parametric EQ filters with parameters of center frequency, gain, and bandwidth.
[0136] In Figures 4A-4C In the example shown, the peak-to-valley decibel values of the notches between the azimuth values of negative ten degrees (-10°) to ninety-five degrees (95°) and the elevation values of negative thirty degrees (-30°) to forty-five degrees (45°) reach less than negative thirty (-30) decibels (dB). Peak-to-valley decibel values of -30 dB or more can be beneficial to produce accurate sound source localization for a listener when a virtual sound source is present within the proposed azimuth-elevation range.
[0137] Figure 5A A plot 500 of a distribution of center frequencies for a left ear is depicted as a function of azimuth (x-axis) and elevation (y-axis). Figure 5BA plot 600 of a center frequency curve 602 for the right ear as a function of elevation angle (x-axis) is depicted, where the azimuth angle is zero degrees (0°). From a negative 90 degree elevation angle to a 90 degree elevation angle, the center frequency values range from approximately 4900 Hz to approximately 8700 Hz. The center frequency curve 602 has a generally S-shaped shape or distribution.
[0138] For example, Figure 6A A plot 600 of a center frequency curve 602 for the right ear as a function of elevation angle (x-axis) is depicted, where the azimuth angle is zero degrees (0°). From a negative 90 degree elevation angle to a 90 degree elevation angle, the center frequency values range from approximately 4900 Hz to approximately 8700 Hz. The center frequency curve 602 has a generally S-shaped shape or distribution.
[0139] Returning to Figure 5A and Figure 5B , the virtual auditory display system 102 can utilize the distribution 500 and / or the distribution 550 to determine the center frequency of one or more notches in the frequency spectrum of an audio signal. The virtual auditory display system 102 can determine the center frequency of the one or more notches based on a location in the virtual auditory space at which the audio signal will cause a sound to be produced by the first ear-wearable device 102a and the second ear-wearable device 102b.
[0140] That is, based on a location of the sound in the virtual auditory space (as specified by, for example, an azimuth angle and an elevation angle), the virtual auditory display system 102 can determine the center frequency of one or more notches in the frequency spectrum of an audio signal that will cause a sound to be produced by the first ear-wearable device 102a and the second ear-wearable device 102b. The virtual auditory display system 102 can determine the center frequency of a first notch in the frequency spectrum of the audio signal by accessing the distribution 500 and the distribution 550. The virtual auditory display system 102 can determine the center frequency of a second notch and subsequent notches in the frequency spectrum of the audio signal based on the distribution 500 and the distribution 550 and one or more offsets from the center frequencies obtained from the distribution 500 and the distribution 550.
[0141] In some embodiments, in addition to or as an alternative to utilizing the distribution 500 and / or the distribution 550, the virtual auditory display system 102 can utilize one or more center frequency curves, each of which can be for a different azimuth angle value, such as Figure 6A the center frequency curve 602 of FIG. 6B, or a different elevation angle value. The virtual auditory display system 102 determines the center frequency of one or more notches in the frequency spectrum of an audio signal based on a location of a resulting sound in the virtual auditory space.
[0142] Figure 6BIn some embodiments, a plot 650 of multiple trials of user experience data with five different digital filters as a function of the notch center frequency. Each of the point 652a, point 652b, point 652c, point 652d, and point 652e is an average of 15 user trials that collected real-time user feedback of perceived sound location in a virtual auditory space, for a total of 75 user trials. Bars 654a, 654b, 654c, 654d, and 654e each represent plus or minus one (1) standard deviation. Points 652b, 652c, 652d, and 652e indicate that for every 150 Hz increase in the notch center frequency, the amount of observed change is approximately 2.5°. Line 656 can be fit to the points 652. The virtual auditory display system 102 can utilize the linear function that produces line 656 to determine the center frequency for one or more notches based on an elevation angle of a sound to be produced. For example, for certain elevation angle ranges (e.g., between approximately 0 degrees and approximately 50 degrees, or between approximately 10 degrees and approximately 40 degrees), the virtual auditory display system 102 can utilize the linearity that produces line 656 to determine one or more offsets from the center frequency in the range. In addition to utilizing the distribution 500 and / or the distribution 550 depicted in FIGS. 5A and 5B, respectively, the virtual auditory display system 102 can do so in addition to or as an alternative to utilizing the distribution 500 and / or the distribution 550. Figure 5A and Figure 5B In addition to or as an alternative to utilizing the distribution 500 and / or the distribution 550, the virtual auditory display system 102 can do so.
[0143] Figures 7A-7X A parameter modifier mask is depicted that can be applied to modify parameters used in some embodiments to generate digital filters. Figure 7A A right ear parameter modifier mask 700a for a notch filter is depicted, and Figure 7B A left ear parameter modifier mask 700b for a notch filter is depicted. Figure 7C A right ear parameter modifier mask 705a for a head shadow filter is depicted, and Figure 7D A left ear parameter modifier mask 705b for a head shadow filter is depicted. Figure 7E A right ear parameter modifier mask 710a for a shelf filter is depicted, and Figure 7F A left ear parameter modifier mask 710b for a shelf filter is depicted. Figure 7G A right ear parameter modifier mask 715a for a peak filter is depicted, and Figure 7H A left ear parameter modifier mask 715b for a peak filter is depicted. Figure 7I A right ear parameter modifier mask 720a for a beam filter is depicted, and Figure 7J A left ear parameter modifier mask 720b for a beam filter is depicted. Figure 7KA right ear parametric modifier mask 725a for a stereo filter is depicted, and Figure 7L A left ear parametric modifier mask 725b for a stereo filter is depicted.
[0144] Figure 7M A right ear parametric modifier mask 730a for a post filter is depicted, and Figure 7N A left ear parametric modifier mask 730b for a post filter is depicted. Figure 7O A right ear parametric modifier mask 735a for a top filter is depicted, and Figure 7P A left ear parametric modifier mask 735b for a top filter is depicted. Figure 7Q A right ear parametric modifier mask 740a for a top transition filter is depicted, and Figure 7R A left ear parametric modifier mask 740b for a top transition filter is depicted. Figure 7S A right ear parametric modifier mask 745a for a bottom filter is depicted, and Figure 7T A left ear parametric modifier mask 745b for a bottom filter is depicted. Figure 7U A right ear parametric modifier mask 750a for a bottom transition filter is depicted, and Figure 7V A left ear parametric modifier mask 750b for a bottom transition filter is depicted. Figure 7W A right ear parametric modifier mask 755a for a wide side filter is depicted, and Figure 7X A left ear parametric modifier mask 755b for a wide side filter is depicted.
[0145] Figures 7A-7X The depicted parametric modifier masks specify parametric modifier values as a function of position in a virtual auditory space (e.g., specified by azimuth and elevation). The parametric modifier values can range between any two values. In some embodiments, the range of parametric modifier values is between zero (0) and another value (including 0), such as one (1), 0.2, 0.4, 0.8, including these values. The parametric mask module 390 can generate a parametric modifier mask by specifying a particular region in the virtual auditory space with a value of one (1) and by specifying a region other than the particular region with a value of zero (0). For example, for the right ear parametric modifier mask 700a, Figure 7A a particular region in the virtual auditory space can be from about negative 50 (-50) degrees azimuth to about 110 degrees azimuth, and from about negative 70 (-70) degrees elevation to about 30 degrees elevation. The parametric mask module 390 can use other particular regions for the parametric modifier mask 700a and other parametric modifier masks in Figures 7A-7X
[0146] The parameter mask module 390 can use image processing algorithms to create a continuous transition of values between the specific region and other regions to generate the parameter modifier mask with parameter modifier values. For example, the parameter mask module 390 can use image processing algorithms such as a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a threshold function, an edge detection function, and / or a segmentation function. In some embodiments, the parameter mask module 390 generates the parameter modifier values using a Gaussian blur mask. The parameter mask module 390 can generate a parameter modifier mask for the right ear and then reflect the parameter modifier mask for the right ear around a vertical axis at an azimuthal value of zero (0) to obtain a parameter modifier mask for the right ear.
[0147] The filter generation module 384 can utilize the parameter modifier mask depicted in FIG. 3B to select one or more parameter modifiers that the filter generation module 384 can use to modify one or more parameters to obtain one or more modified parameters based on the location of the sound in the virtual auditory space. For example, the filter generation module 384 can utilize one or more parameter modifiers to modify the gain of a digital filter. In some embodiments, the filter generation module 384 multiplies one or more parameters by one or more parameter modifiers to obtain one or more modified parameters. Other uses of parameter modifiers will be apparent. Figures 7A-7X
[0148] According to some embodiments, Figure 8A depicts a gain profile 870 for the head shadow for the left ear, and Figure 8B depicts a gain profile 880 for the head shadow for the right ear. The gain profile 870 depicts how the gain changes as the sound source transitions from a position 872 that is typically located to the right of the ear to a position 874 that is typically located in front of the wearer to a position 876 that is typically located to the left of the ear. The gain profile 880 depicts how the gain changes as the sound source transitions from a position 882 that is typically located to the left of the ear to a position 884 that is typically located in front of the wearer to a position 886 that is typically located to the left of the ear.
[0149] Figure 8C depicts a gain profile 860 of applying a digital filter to an audio signal according to some embodiments. The gain profile 890 shows several notches 864 across the head shadow 862. The center frequencies of the several notches 864 are between 10Λ3 Hz and 10Λ4 Hz.
[0150] Figure 8D depicts user experience data for a digital filter and user experience data for a prior art head-related transfer function (HRTF) according to some embodiments. The prior art HRTF is used as a standard model for many past and present HRTF applications. Figure 8D Panel 800 of FIG. 8 reports the difference between the user-perceived elevation of the virtual sound object and the actual elevation of the sound object for both the digital filter for 150 trials and the prior art HRTF for 150 trials. For the digital filter trials, point 804 is the average user-perceived elevation and band 802 is the standard deviation of the user-perceived elevation. For the HRTF trials, point 808 is the average user-perceived elevation and band 806 is the standard deviation of the user-perceived elevation.
[0151] Figure 8D Panel 850 of FIG. 8 reports the difference between the user-perceived azimuth of the virtual sound object and the actual azimuth of the sound object for both the digital filter for 150 trials and the prior art HRTF for 150 trials. For the digital filter trials, point 852 is the average user-perceived azimuth and band 854 is the standard deviation of the user-perceived azimuth. For the HRTF trials, point 858 is the average user-perceived azimuth and band 856 is the standard deviation of the user-perceived azimuth.
[0152] For the elevation trials, the closer the elevation change amount is to 0°, the more accurate the representation of the virtual sound object. Similarly, for the azimuth trials, the closer the azimuth change amount is to 0°, the more accurate the representation of the virtual sound object. The user experience data shows that the digital filter improved the elevation change amount from an average of about 30.19° (a standard deviation of about 12.54°) to an average of about -0.03° (a standard deviation of about 4.12°). The user experience data also shows that the digital filter improved the azimuth change amount from an average of about -0.64° (a standard deviation of about 7.76°) to an average of about 0.02° (a standard deviation of about 2.04°). The data shows that the digital filter according to some embodiments improved the accuracy and precision of the virtual sound object.
[0153] Figure 9A Method 900 of generating a digital filter according to some embodiments is depicted. Generating system 380 can perform method 900. Generating system 380 can perform method 900 to generate a first set of combined digital filters and a second set of combined digital filters. Method 900 begins at step 902, where generating system 380 (e.g., filter generation module 384) can generate an approximately S-shaped distribution of center frequencies for the right ear (e.g., see FIG. 9A) and an approximately S-shaped distribution of center frequencies for the left ear (e.g., see FIG. 9B). Figure 5A ) and an approximately S-shaped distribution of center frequencies for the left ear (e.g., see Figure 5B ).
[0154] At step 904, generating system 380 (e.g., parameter mask module 390) generates a parameter modifier mask for the right ear and a parameter modifier mask for the left ear (e.g., see Figures 7A-7X). The generation system 380 can generate the parameter modifier mask using one or more image processing algorithms. The one or more image processing algorithms can include one or more of a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a thresholding function, an edge detection function, and a segmentation function.
[0155] At step 906, for each of the plurality of positions in the virtual auditory space, the generation system 380 (e.g., the parameter generation module 388) can generate one or more first parameters of one or more first digital filters and one or more second parameters of one or more second digital filters. The one or more first parameters can include one or more first q, one or more first gain, and one or more first center frequency. The one or more second parameters can include one or more second q, one or more second gain, and one or more second center frequency.
[0156] The generation system 380 can utilize the parameter modifier mask for the right ear and the parameter modifier mask for the left ear to select one or more parameter modifiers based on the position in the virtual auditory space, which the generation system 380 can use to modify the one or more parameters to obtain one or more modified parameters. In some embodiments, the generation system 380 multiplies the one or more parameters by the one or more parameter modifiers to obtain the one or more modified parameters.
[0157] The generation system 380 can utilize the approximately S-shaped distribution of the center frequencies for the right ear and / or the approximately S-shaped distribution of the center frequencies for the left ear to determine one or more center frequencies of one or more notches in the frequency spectrum of the audio signal for the right ear and the audio signal for the left ear. The generation system 380 can determine the center frequencies of the one or more notches based on the position in the virtual auditory space.
[0158] At step 908, for each position, the generation system 380 (e.g., the filter generation module 384) can generate one or more first digital filters including one or more first notch filters including one or more first parameters. The generation system 380 can generate the one or more first notch filters utilizing the one or more first q, the one or more first gain, and the one or more first center frequencies. The one or more first notch filters are configured to produce one or more first notches in a first frequency spectrum of the first audio signal according to the one or more first q, the one or more first gain, and the one or more first center frequencies when the generation system 380 applies the one or more notch filters to the audio signal for the right ear.
[0159] At step 910, for each location, the generation system 380 (e.g., filter generation module 384) can generate one or more combined first digital filters for the location based on the one or more first digital filters. In some embodiments, the one or more first digital filters are IIR filters, and the one or more combined first digital filters are FIR filters.
[0160] At step 912, for each location, the generation system 380 can store (e.g., in the data storage 220) the one or more combined first digital filters in association with the location.
[0161] At step 914, for each location, the generation system 380 (e.g., filter generation module 384) can generate one or more second digital filters including one or more second notch filters, the second notch filters including one or more second parameters. The generation system 380 can generate the one or more second notch filters with the one or more second q, the one or more second gains, and the one or more second center frequencies. The one or more second notch filters are configured to produce one or more second notches in a second frequency spectrum of the second audio signal according to the one or more second q, the one or more second gains, and the one or more second center frequencies when the generation system 380 applies the one or more notch filters to the audio signal for the left ear.
[0162] At step 916, for each location, the generation system 380 (e.g., filter generation module 384) can generate one or more combined second digital filters for the location based on the one or more second digital filters. In some embodiments, the one or more second digital filters are IIR filters, and the one or more combined second digital filters are FIR filters.
[0163] At step 918, for each location, the generation system 380 can store (e.g., in the data storage 220) the one or more combined second digital filters in association with the location.
[0164] At step 920, the generation system 380 tests to see if there are more locations for which the generation system 380 is to generate digital filters. If so, the method 900 returns to step 906. The generation system 380 can perform the method 900 multiple times to generate multiple sets of combined first digital filters and multiple sets of combined second digital filters. Each pair of sets of digital filters can be used for a different prototype
[0165] In some embodiments, the generation system 380 can perform the method 900 several times to generate multiple sets of combined first digital filters and combined second digital filters. Each pair of sets can be used to represent a different archetype of a different user population or user grouping.
[0166] Figure 9B A method 950 of generating digital filters in some embodiments is depicted. The generation system 380 can perform the method 950. The method 950 includes certain steps that can be substantially similar to certain steps of the method 900. The generation system 380 (e.g., various components of the generation system 380) can perform the method 950. The generation system 380 can perform the method 950 to generate a set of first digital filters and a set of second digital filters.
[0167] The method 950 begins at step 952, where the generation system 380 (e.g., the parameter generation module 388) can generate a first substantially S-shaped distribution of center frequencies and a second substantially S-shaped distribution of center frequencies. At step 954, the generation system 380 (e.g., the parameter modifier mask module 390) can generate a first parameter modifier mask and a second parameter modifier mask.
[0168] At step 956, for each of a plurality of virtual auditory spatial positions, the generation system 380 (e.g., the parameter generation module 388) can generate one or more first parameters of one or more first digital filters and one or more second parameters of one or more second digital filters. The one or more first parameters can include one or more first q’s, one or more first gains, and one or more first center frequencies. The one or more second parameters can include one or more second q’s, one or more second gains, and one or more second center frequencies.
[0169] At step 958, for each virtual auditory spatial position, the generation system 380 can generate one or more first digital filters including one or more first notch filters, the first notch filters including the one or more first parameters. Step 958 is substantially similar to step 980 of the method 900.
[0170] At step 960, for each virtual auditory spatial position, the generation system 380 can store the one or more first digital filters in association with the virtual auditory spatial position. Step 960 is substantially similar to step 912 of the method 900.
[0171] At step 962, for each virtual auditory spatial position, the generation system 380 can generate one or more second digital filters including one or more second notch filters, the second notch filters including the one or more second parameters. Step 962 is substantially similar to step 914 of the method 900.
[0172] At step 964, for each virtual auditory spatial position, the generation system 380 can store the one or more second digital filters in association with the virtual auditory spatial position. Step 964 is generally similar to step 918 of the method 900.
[0173] At step 966, the generation system 380 performs a test to see if there are more virtual auditory spatial positions for which the generation system 380 is to generate digital filters. If so, the method 900 returns to step 956. The generation system 380 can perform the method 950 multiple times to generate multiple sets of one or more first digital filters and multiple sets of one or more second digital filters.
[0174] The method 900 and the method 950 can include additional steps. For example, the generation system 380 can provide a test digital filter. The generation system 380 (e.g., the user interface module 394) can provide a user interface that allows for playback of a sound generated from an audio signal to which the digital filter is applied. A user can listen to the sound and determine that one or more parameters of the digital filter should be modified. For example, the user can modify the parameters of the digital filter to ensure that the sound meets a desired threshold for tone quality, intelligibility, loudness, etc. The generation system 380 can provide another user interface that allows the user to modify the one or more parameters of the digital filter. The generation system 380 (e.g., the digital filter tuning module 392) can receive the one or more parameters of the digital filter from the user and modify the digital filter based on the received one or more parameters.
[0175] Figure 10A A method 1000 of applying digital filters is depicted in accordance with some embodiments. The virtual auditory display system 102 and the virtual auditory display device 100 can perform the method 1000. The method 1000 begins at step 1002, where the virtual auditory display system 102 (e.g., the binauralizer 138) receives a set of combined first digital filters and a set of combined second digital filters. At step 1004, the virtual auditory display system 102 (e.g., the binauralizer 138) receives an input audio signal that includes one or more audio sub-signals. Each of the one or more audio sub-signals has a position in a virtual auditory space.
[0176] The virtual auditory display system 102 performs steps 1006-1020 of the method 1000 while receiving the input audio signal. At step 1006, one or both of the first ear-wearable device 102a and the second ear-wearable device 102b detects a head orientation of a user wearing the first ear-wearable device 102a and the second ear-wearable device 102b. The first ear-wearable device 102a and / or the second ear-wearable device 102b provides the head orientation to the virtual auditory display system 102.
[0177] At step 1008, for each of the one or more audio sub-signals, the virtual auditory display system 102 determines a particular location in the virtual auditory space based on the location and the head orientation of the audio sub-signal. At step 1010, for each audio sub-signal, the virtual auditory display system 102 selects a particular one or more combined first digital filter and a particular one or more combined second digital filter based on the particular location.
[0178] At step 1012, for each audio sub-signal, the virtual auditory display system 102 applies the particular one or more combined first digital filter to the audio sub-signal to obtain a first processed audio sub-signal. At step 1014, for each audio sub-signal, the virtual auditory display system 102 applies the particular one or more combined second digital filter to the audio sub-signal to obtain a second processed audio sub-signal.
[0179] At step 1016, the virtual auditory display system 102 tests whether there are more audio sub-signals to process. If so, the method 1000 returns to step 1008. If not, the method 1000 continues to step 1018. After processing all of the audio sub-signals, the virtual auditory display system 102 obtains a plurality of first processed audio sub-signals and a plurality of second processed audio sub-signals.
[0180] At step 1018, the virtual auditory display system 102 generates a first output audio signal for the left ear-worn device based on the plurality of first processed audio sub-signals and a second output audio signal for the right ear-worn device based on the plurality of second processed audio sub-signals. The virtual auditory display system 102 provides the first output audio signal to the first ear-worn device 102a and the second output audio signal to the second ear-worn device 102b.
[0181] At step 1020, the first ear-worn device 102a outputs a first sound based on the first output audio signal and the second ear-worn device 102b outputs a second sound based on the second output audio signal. The virtual auditory display system 102 can thus utilize the method 1000 to provide a virtual auditory display sound based on an audio signal that can have a plurality of audio sub-signals (or channels) that would normally require a plurality of speakers to produce a surround sound effect. The virtual auditory display system 102 can provide the virtual auditory display sound to a user using only the first ear-worn device 102a and the second ear-worn device 102b.
[0182] Figure 10B A method 1050 of applying digital filters is depicted in accordance with some embodiments. The method 1050 includes certain steps that can be generally similar to certain steps of the method 1000. The virtual auditory display system 102 can perform the method 1050.
[0183] Method 1050 begins at step 1052, where virtual auditory display system 102 (e.g., binauralizer 138) receives a set of one or more first digital filters and a set of one or more second digital filters. At step 1054, virtual auditory display system 102 (e.g., binauralizer 138) receives an audio signal having one or more audio sub-signals. Each of the one or more audio sub-signals is associated with a virtual auditory spatial location.
[0184] At step 1056, virtual auditory display system 102 receives a head orientation of a user. At step 1058, for each of the one or more audio sub-signals, virtual auditory display system 102 determines a particular virtual auditory spatial location based on the virtual auditory spatial location and the head orientation. At step 1060, for each audio sub-signal, virtual auditory display system 102 selects a particular set of one or more first digital filters and a particular set of one or more second digital filters based on the virtual auditory spatial location or the particular virtual auditory spatial location.
[0185] At step 1062, for each audio sub-signal, virtual auditory display system 102 applies the particular set of one or more first digital filters to the audio sub-signal to obtain a first processed audio sub-signal. At step 1064, for each audio sub-signal, virtual auditory display system 102 applies the particular set of one or more second digital filters to the audio sub-signal to obtain a second processed audio sub-signal.
[0186] At step 1066, virtual auditory display system 102 tests whether there are more audio sub-signals to process. If so, method 1050 returns to step 1058. If not, method 1000 continues to step 1068. After processing all of the audio sub-signals, virtual auditory display system 102 obtains a plurality of first processed audio sub-signals and a plurality of second processed audio sub-signals.
[0187] At step 1068, virtual auditory display system 102 generates a first output audio signal for a first device based on the plurality of first processed audio sub-signals and a second output audio signal for a second device based on the plurality of second processed audio sub-signals. The first device can be or include, for example, first ear-wearable device 102a, and the second device can be or include, for example, second ear-wearable device 102. At step 1070, virtual auditory display system 102 provides the first output audio signal to the first device and the second output audio signal to the second device.
[0188] Figure 10CA method 1080 of generating and applying virtual auditory display filters in some embodiments is depicted. The generation system 380 and the virtual auditory display system 102 can perform the method 1080. The method 1080 begins at step 1082, where the generation system 380 (e.g., the parameter generation module 388) can generate a first substantially S-shaped distribution of center frequencies and a second substantially S-shaped distribution of center frequencies. At step 1084, the generation system 380 (e.g., the parameter mask module 390) can generate first and second parameter modifier masks, including first and second notch parameter modifier masks.
[0189] At step 1086, the generation system 380 (e.g., the filter generation module 384) can generate first and second virtual auditory display filters. The first virtual auditory display filter can include a first set of first functions. The one or more first functions, when applied to a first audio signal having a first location in a virtual auditory space, can generate a first processed audio signal having a first frequency response with one or more first notches at one or more first center frequencies based on the first location. The one or more first notches can have one or more first peak-to-trough depths of at most -10 dB (e.g., approximately -30 dB).
[0190] The second virtual auditory display filter can include a second set of second functions. The one or more second functions, when applied to the first audio signal, can generate a second processed audio signal having a second frequency response with one or more second notches at one or more second center frequencies based on the second location. The one or more second notches can have one or more second peak-to-trough depths of at most -10 dB (e.g., approximately -30 dB).
[0191] At step 1088, the virtual auditory display system 102 can receive an audio signal having a second location in a virtual auditory space. For example, the virtual auditory display system 102 can receive an audio signal from a digital device on which the virtual auditory display system 102 is executing. At step 1090, the virtual auditory display system 102 can receive a head orientation of a user, for example, from a virtual auditory display device 100 that the user is utilizing.
[0192] In step 1092, the virtual auditory display system 102 may apply a first virtual auditory display filter, including a first subset of a first function based on a second position selection, to the second audio signal to generate a third processed audio signal with a third frequency response. In step 1094, the virtual auditory display system 102 may apply a second virtual auditory display filter, including a second subset of a second function based on a second position selection, to the second audio signal to generate a fourth processed audio signal with a fourth frequency response. In step 1094, the virtual auditory display system 102 may provide the first processed audio signal to a first sound output device (e.g., a first ear-worn device 102a) and the second processed audio signal to a second sound output device (e.g., a second ear-worn device 102b).
[0193] When the virtual auditory display system 102 is receiving an input audio signal that may correspond to, for example, a song file, an audio stream, a podcast, or any other audio, the virtual auditory display system 102 may perform steps 1088 to 1094 of method 1080.
[0194] Methods 1000, 1050, and 1080 may include Figures 10A-10C Additional steps are not shown. For example, these methods may include the steps of the virtual auditory display system 102 receiving a selection of an acoustic environment and the virtual auditory display system 102 determining a first acoustic environment digital filter and a second acoustic environment digital filter based on the acoustic environment. The acoustic environment may be represented by one or more ambisonic arrays. The virtual auditory display system 102 may determine the acoustic environment digital filters based on one or more ambisonic arrays. The virtual auditory display system 102 may apply the digital filters and the acoustic environment digital filters to obtain a processed audio sub-signal. Other modifications to these methods will be apparent.
[0195] Figure 2C This is a block diagram depicting a process 290 for generating digital filters for an acoustic environment in some embodiments. A first speaker 292a and a second speaker 292b may be located in a specific acoustic environment, such as a concert hall, vehicle, nightclub, etc. The sound output from the first speaker 292a and the second speaker 292b is captured by a microphone 294 and converted into a signal. A virtual auditory display system 102 (e.g., a filter generation module 384) generates one or more surround sound digital filters 296 based on the signal. The one or more surround sound digital filters 296 are a set of one or more acoustic environment digital filters 298.
[0196] Figure 2Dis a block diagram depicting the operation of a spatialization engine 270 of the virtual auditory display system 102 in some embodiments. The spatialization engine 270 can be part of the binauralizer 138 or a separate component. The spatialization engine 270 receives a user interface (UI) selected acoustic environment at block 272 and determines an acoustic environment digital filter based on the selected acoustic environment at block 274. At block 276, the spatialization engine 270 receives coordinates of decoded audio objects and an input audio signal. At block 278, the spatialization engine 270 matches an index of the acoustic environment digital filter to an index of the input audio signal. At block 280, the spatialization engine 270 applies a convolution matrix to the output of block 278 and the coordinates and input audio signal.
[0197] At block 282, the spatialization engine 270 receives a user head orientation and audio source distance signal. At block 284, the spatialization engine performs ambisonic-to-binaural conversion based on the output of block 280 and the user head orientation and audio source distance signal by applying a digital filter to the audio signal received at block 276 based on the location of the audio signal in the virtual auditory space. At block 286, the spatialization engine 270 outputs an audio signal for the left ear and at block 288, the spatialization engine 270 outputs an audio signal for the right ear.
[0198] Figure 11A and Figure 11B depicts an example user interface 1100 for displaying a representation of a virtual audio display in some embodiments. The virtual auditory display system 102 (e.g., the user interface module 210) can provide the user interface 1100. The user interface 1100 is described with reference to the virtual auditory display device 100 Figure 11A and Figure 11B but the virtual auditory display system 102 can provide the user interface 1100 for other devices.
[0199] The user interface 1100 includes an icon 1104 labeled “VAD” indicating that the virtual auditory display device 100 is connected to the virtual auditory display system 102 and an icon 1102 labeled “IMU” indicating that the IMU-based sensor system of the virtual visual display device 100 is calibrated. The user interface 1100 also includes an encoded representation drop-down list 1114 that allows the wearer to select how the virtual auditory display system 102 should represent audio received by the virtual auditory display system 102. Example encoded representations are mono (single channel), stereo (two channels), 5.1 5.1 (six channels), 7.1 (eight channels), 7.1.4 (twelve channels), and 9.1.6 (sixteen channels).
[0200] The user interface 1100 also includes an acoustic environment drop-down menu 1116 that allows the wearer to select an acoustic environment in which the virtual auditory display system 102 should present the virtual auditory display. Example acoustic environments include a dry acoustic environment, a studio acoustic environment, a car acoustic environment, a telephone acoustic environment, a club acoustic environment, and a headphone acoustic environment. The virtual auditory display system 102 can select an acoustic environment digital filter based on the selected acoustic environment and apply the acoustic environment digital filter with the virtual auditory display filter. The virtual auditory display sound will sound different to the wearer based on the selected acoustic environment. The user interface 1100 also includes an output volume slider 1122 that allows the wearer to adjust the volume of the sound output by the virtual auditory display device 100.
[0201] The user interface 1100 also includes a representation 1108 of the virtual audio display. In Figure 11A the representation 1108 is depicted as a virtual sphere surrounding a head 1112, representing the wearer's head from an upper rear perspective. The user interface 1100 also displays sounds in the virtual auditory space at locations relative to the wearer's head on the representation 1108 at corresponding locations relative to the head 1112. For example, a sound 1110a is depicted to the left, below, and rear of the head 1112. This is because the actual sound corresponding to the sound 1110a has those locations in the virtual auditory space. A sound 1110b is depicted above, rear, and slightly to the left of the head 1112, a sound 1110c is depicted in front of and to the right of the head 1112, and a sound 1110d is depicted to the right, below, and rear of the head 1112.
[0202] While the sound is output, the virtual auditory display device 100 detects the wearer's head orientation and sends the head orientation to the virtual auditory display system 102. The virtual auditory display system 102 updates the representation 1108 based on the detected head orientation. The virtual auditory display system 102 can move the head 1112 and the sounds 1110 based on the detected head orientation.
[0203] The user interface 1100 also includes a virtual auditory display representation drop-down menu 1118 that allows the wearer to select how the virtual auditory display system 102 provides the virtual auditory display representation. Example virtual auditory display representations include a custom representation Figure 11A depicted in), which provides the wearer with a right upper rear perspective of the representation 1108, a top representation Figure 11B depicted in), which provides the wearer with a top perspective of the representation 1108 from Figure 11A the top of the sphere in), and a back representation, which provides the wearer with a rear perspective of the representation 1108 from Figure 11B the back of the sphere in).
[0204] The user interface 1100 also includes a position button 1120 that, if selected by the wearer, can cause the virtual auditory display system 102 to change the representation 1108 so that the location specified by some coordinate (e.g., zero degrees azimuth, zero degrees elevation) can be directly in front of the head 1112. The user interface 1100 also includes a settings icon 1106 that, if selected by the wearer, can cause the virtual auditory display system 102 to provide an example user interface for adjusting the virtual audio display settings.
[0205] Figure 11C An example user interface 1150 for adjusting the settings of the virtual audio display in some embodiments is depicted. The virtual auditory display system 102 (e.g., the user interface module 210) can provide the user interface 1150. The user interface 1150 includes an icon labeled “IMU” 1152 that indicates that the IMU-based sensor system of the virtual auditory display device 100 is calibrated. The user interface 1150 also includes a button labeled “Re-calibrate” 1154 that the wearer can select to cause the virtual auditory display system 102 to perform the calibration portion of the calibration and / or personalization process.
[0206] The user interface 1150 also includes an icon 1156 labeled “VAD” that indicates that the virtual auditory display device 100 is connected to the virtual auditory display system 102, a recommendation 1158 of a virtual auditory display filter, and a button labeled “Personalize” 1160 that the wearer can select to cause the virtual auditory display system 102 to perform the personalization portion of the calibration and / or personalization process. The user interface 1150 also indicates an estimate of the spatialization accuracy of the wearer’s virtual auditory display and a button labeled “Test” 1162 that the wearer can select to cause the virtual auditory display system 102 to provide a test procedure for the wearer to allow the wearer to test to see if the wearer can accurately localize the virtual auditory display sounds.
[0207] The user interface 1150 also includes a group of icons 1164 (labeled “A” through “G”) that indicate the set of virtual auditory display filters that are creating the virtual auditory display for the wearer. As depicted, the current set of virtual auditory display filters is “VAD C.” The wearer can select a different set of virtual auditory display filters by selecting a different icon in the group 1164. The wearer can perform the calibration portion of the calibration and / or personalization process by selecting the button 1154 and / or the personalization portion of the calibration and / or personalization process by selecting the button 1160.
[0208] The user interface 1150 also allows the wearer to select a set of custom digital filters for the virtual auditory display system 102 to use in generating the virtual auditory display. The wearer can do so by selecting the button 1168 labeled "Upload," which allows the wearer to upload a file containing a set of custom digital filters to the virtual auditory display system 102. The user interface element 1166 can then display the name of the file. This functionality can be desirable for a wearer who already has a custom HRTF and wishes for the virtual auditory display system 102 to utilize the custom HRTF.
[0209] Figure 12 Figures 1200 are a plurality of images that depict example use cases for the virtual auditory display filter techniques described in this application in some embodiments. One example use case is to improve virtual surround sound for a television or movie using only a pair of speakers, as depicted by image 1202. A group of example use cases relate to making or listening to music. Image 1210 depicts an example use case of listening to music on headphones. The virtual auditory display filter techniques can make the music played on headphones indistinguishable from music played using a physical surround sound setup.
[0210] Image 1218 depicts using the virtual auditory display filter techniques in a virtual monitor to provide noise isolation, sound quality, and virtualization for a musician. Image 1204 depicts mixing music in any acoustic environment using the virtual auditory display filter techniques. Image 1220 depicts how the virtual auditory display filter techniques can provide a listening experience that reinvigorates a listener's favorite music. Image 1212 depicts using the virtual auditory display filter techniques in a game to provide an immersive gaming experience. The virtual auditory display filter techniques can allow a user to hear sounds emanating from locations not shown on the user's display and thus improve the user's awareness / perception.
[0211] Another group of example use cases relates to military, non-military (e.g., first responders such as police and firemen), and / or other organizational applications. For example, military personnel can use military radio systems to communicate with soldiers, commanders, and other military personnel. The present techniques can be used in scenarios including military operations, emergency services, aviation, maritime, and other scenarios.
[0212] Image 1206 depicts the virtual auditory display filter techniques providing improved voice pickup and voice display for organizational communications. Image 1214 depicts the virtual auditory display filter techniques augmenting visual instrumentation with auditory signals in maritime operations. Image 1222 depicts the virtual auditory display filter techniques providing a hyper-realistic virtual audio environment that facilitates virtual training for military and / or non-military personnel.
[0213] Image 1208 depicts a virtual auditory display filter technique that provides audio augmentation for situational awareness of a combat infantry situation. Image 1216 depicts a virtual auditory display filter technique that provides audio augmentation for situational awareness of an air force directed control. For example, a pilot can use the location of virtual beacons to help with situational awareness. Image 1224 depicts a virtual auditory display filter technique that provides audio augmentation for super-situational awareness of a drone operation.
[0214] Another example use case for the virtual auditory display filter technique involves a phone call or a video conference. For example, multiple people can be talking simultaneously in a phone call or a video conference, making it difficult for a listener to focus on the speaker they want to hear, which can lead to confusion and miscommunication. Current technology allows a user to virtually select the speaker they want to hear by simply moving their head or other gestures. This attention selection mechanism can help avoid confusion and improve conference efficiency.
[0215] As another example, an air traffic controller communicates with pilots using radio messages. The air traffic controller monitors the position, speed, and altitude of aircraft within a designated airspace through vision and radar and issues instructions to the pilots through the radio. Often, the air traffic controller needs to communicate with multiple pilots simultaneously. Today, these multiple pilot communication situations are addressed by a physical switchboard that does not allow the user to direct / lead attention selection. The present technology can allow the air traffic controller to use gestures (e.g., head or hand movements) or other actions to position the radio communications so that there is seamless attention selection. In a simple example, multiple radio communication signals are statically arranged in unique virtual locations. The air traffic controller then looks at these predefined locations to hear the radio signals. In other examples, the radio communication signals are dynamically updated with the position, speed, and altitude of the aircraft.
[0216] Other example use cases include virtual auditory display notifications for positioning notifications such as voice, text-to-speed messages, email alerts, phone messages, or other audio notifications, spatial navigation for conveying direction and distance of virtual or real objects using virtual auditory display cues (which can also be used for wayfinding or directional awareness), and spatial ambience for providing a user with a virtual sound environment that can be mixed with local or virtual sounds (e.g., for experiencing music as if in a concert hall). Other example use cases for the virtual auditory display filter technique are possible.
[0217] Figure 13A and Figure 13Bis an illustration of a method of personalizing a digital filter in some embodiments that involves providing an action (e.g., playing a sound) and detecting a user's perception of the action. As described herein, a user can indicate a perception of an action in various ways, such as by pointing his or her head, making one or more gestures, indicating a location of a sound on a graphical user interface, and so on.
[0218] Figure 13A A view 1300 is depicted that indicates how a user can perceive a location of a sound in a virtual auditory space as indicated by an azimuth angle 1314 and an elevation angle 1316. Figure 13B A view 1350 is depicted that indicates how a user can perceive a location of a sound in a virtual auditory space as indicated by a distance 1318 and an elevation angle 1316. In both view 1300 and view 1350, the user 1302 is wearing a first ear-wearable device 102a and a second ear-wearable device 102b (not shown in FIG. 13). Figure 13A and Figure 13B The virtual auditory display system 102 can provide instructions to the user 1302 to follow the sound in the virtual auditory space with the user's head as the user hears the sound. The virtual auditory display system 102 can then generate audio signals that cause the first ear-wearable device 102a and the second ear-wearable device 102 to output one or more sounds with a location 1304 in the virtual auditory space. As the user 1302 hears the one or more sounds, the user 1302 can point his or her head to a perceived location 1306 from which the user perceives the sound, which can be different from the location 1304. In pointing his or her head to the perceived location 1306, the user 1302 can appear to look in the direction in which the user 1302 perceives the sound. Other gestures that the user can make with his or her head include nodding, shaking, tilting, and turning. Other head gestures will be apparent.
[0219] As the user 1302 moves his or her head to point to the perceived location 1306 of the one or more sounds, one or both of the first ear-wearable device 102a and the second ear-wearable device 102b can detect a head orientation of the user 1302 (e.g., using the IMU sensor system 252 and / or the magnetometer 254). The virtual auditory display system 102 can utilize the detected head orientation to determine the perceived location 1306. The virtual auditory display system 102 can then determine an amount of change 1308 between the location 1304 and the perceived location 1306. The virtual auditory display system 102 can calculate the amount of change 1308 based on a difference between the azimuth angle 1314 and / or the elevation angle 1316 of the location 1304 and the azimuth angle 1314 and / or the elevation angle 1316 of the perceived location 1306.
[0220] The user 1302 can use other gestures to indicate the distance, such as extending an arm a specified amount, and the virtual auditory display system 102 can determine the amount of change 1308 using such gestures based on the difference between the distance 1318 of the location 1304 and the distance 1315 of the perceived location 1306.
[0221] As indicated by solid line 1310, the virtual auditory display system 102 can generate an audio signal that causes the first ear-wearable device 102a and the second ear-wearable device 102 to output one or more subsequent sounds of which their locations in the virtual auditory space have changed. The user 1302 can move his or her head to follow the movement of the one or more subsequent sounds, and the perceived location of the one or more sounds can change, as indicated by dashed line 1312.
[0222] As the user 1302 moves his or her head, one or both of the first and second wearable devices 102a and 102b can detect the subsequent head orientation of the user 1302. The virtual auditory display system 102 can utilize the detected subsequent head orientation to determine a perceived location of the one or more subsequent sounds. The virtual auditory display system 102 can then determine one or more subsequent amounts of change between the location of the one or more subsequent sounds and the perceived location of the one or more subsequent sounds.
[0223] The virtual auditory display system 102 can store the amounts of change determined by the virtual auditory display system 102 and utilize the amounts of change to modify the digital filters in order to cause the user 1302 to perceive the location of the sounds in the virtual auditory space to be closer to the actual location in the virtual auditory space. In some embodiments, the virtual auditory display system 102 modifies the digital filters by selecting a different set of digital filters that the virtual auditory display system 102 determines will reduce or minimize the amounts of change for the user 1302. The virtual auditory display system 102 can then use the different set of digital filters for the user 1302.
[0224] In some embodiments, the virtual auditory display system 102 modifies parameters of the digital filters. For example, the virtual auditory display system 102 can modify parameters such as center frequency, gain, and / or q. For example, the user 1302 can have an amount of change in elevation of a few degrees. The virtual auditory display system 102 can modify the center frequency of a notch, a pair of notches, or a cluster of notches (see, e.g., FIG. 13B) to modify the elevation of the virtual sound object and thus reduce the amount of change in elevation. The virtual auditory display system 102 can then utilize the modified digital filters during real-time audio playback. Figure 6B
[0225] The virtual auditory display system 102 can repeat the personalization procedure one or more times until the virtual auditory display system 102 determines that the amounts of change are within a particular range or threshold.
[0226] In some embodiments, the virtual auditory display system 102 and / or other devices can capture other actions that a user can make in response to the user hearing a sound in the virtual auditory space to indicate the user's perception of the location of the sound. Example actions can include the user's vocal responses, gestures made by the user using parts of the user's body other than the user's head (e.g., pointing with the user's fingers or arms, waving, clapping, tapping, and hand signals). Such actions can be captured by devices connected to the virtual auditory display system 102, such as microphones, cameras, or motion sensing devices.
[0227] Other example actions include the user using a graphical user interface and / or user input devices of a digital device (such as a phone, tablet, or laptop or desktop computer) to indicate the perceived location of the sound. For example, the virtual auditory display system 102 can provide the user with a graphical user interface that graphically represents the virtual auditory space, and the user can utilize input devices (mouse, keyboard, touch screen, and / or voice commands) to indicate the perceived location of the sound on the graphical representation of the virtual auditory space. It will be understood that there are various methods of capturing user actions in response to the user's perception of the location of a sound, and the virtual auditory display system 102, optionally in cooperation with other devices, can use various methods.
[0228] Figure 14A A method 1400 of personalizing digital filters in some embodiments is depicted. The virtual auditory display system 102 can perform the method 1400. The method 1400 begins at step 1402, where the virtual auditory display system 102 (e.g., the binauralizer 138) receives a personalized audio signal having a first location in a virtual auditory space. At step 1404, the virtual auditory display system 102 (e.g., the binauralizer 138) determines a first particular first location in the virtual auditory space based on the first location.
[0229] At step 1406, the virtual auditory display system 102 selects, based on the first particular first location, a particular one or more of the first set of combined first digital filters from the first set of combined first digital filters and a particular one or more of the first set of combined second digital filters from the first set of combined second digital filters.
[0230] At step 1408, the virtual auditory display system 102 applies the particular one or more of the first set of combined first digital filters to the personalized audio signal to obtain a first processed personalized audio signal and applies the particular one or more of the first set of combined second digital filters to the personalized audio signal to obtain a second processed personalized audio signal.
[0231] At step 1410, the virtual auditory display system 102 generates a first output audio signal for the left-ear device based on the first processed personalized audio signal and a second output audio signal for the right-ear device based on the second processed personalized audio signal. At step 1412, the left-ear device outputs a first sound based on the first output audio signal and the right-ear device outputs a second sound based on the second output audio signal.
[0232] At step 1414, one or both of the left-ear device and the right-ear device detects a head orientation of a user wearing the left-ear device and the right-ear device. At step 1416, the virtual auditory display system 102 determines a second particular first location in the virtual auditory space based on the head orientation.
[0233] At step 1418, the virtual auditory display system 102 determines an amount of change between the first particular first location and the second particular first location. At step 1420, the virtual auditory display system 102 selects a second set of combined first digital filters and a second set of combined second digital filters based on the amount of change. The virtual auditory display system 102 can use the second set of combined first digital filters and the second set of combined second digital filters when receiving subsequent input audio signals.
[0234] Figure 14B A method 1450 of personalizing digital filters in some embodiments is depicted. The method 1450 includes certain steps that can be generally similar to certain steps of the method 1400. The virtual auditory display system 102 (e.g., various components of the virtual auditory display system 102) can perform the method 1450.
[0235] At step 1452, the virtual auditory display system 102 receives a set of multiple first digital filters. At step 1454, the virtual auditory display system 102 receives a set of multiple second digital filters. There is one or more first digital filters and one or more second digital filters generated for each of a plurality of virtual auditory space locations.
[0236] At step 1456, the virtual auditory display system 102 receives personalization information for the user. The personalization information can include user-directed actions or perceptions of acoustic cues, acoustic quality information, user anatomical measurements, user demographic information, and / or user hearing measurements.
[0237] At step 1458, the virtual auditory display system 102 modifies the set of multiple first digital filters based on the personalization information for the user. At step 1460, the virtual auditory display system 102 modifies the set of multiple second digital filters based on the personalization information for the user.
[0238] In some embodiments, modifying the set of first digital filters based on the individualization information includes modifying one or more first center frequencies of the first digital filters. Further, modifying the set of second digital filters based on the individualization information includes modifying one or more second center frequencies of the second digital filters.
[0239] In some embodiments, modifying the set of first digital filters based on the individualization information includes selecting a different set of first digital filters. Further, modifying the set of second digital filters based on the individualization information includes selecting a different set of second digital filters.
[0240] The virtual auditory display system 102 can provide a calibration and / or individualization process that allows a wearer of a virtual auditory display device to calibrate the virtual auditory display device and / or to individualize the virtual auditory display provided by the virtual auditory display device.
[0241] The calibration and / or individualization process can include a calibration portion and an individualization portion. The virtual auditory display device can include an inertial measurement unit (IMU). Calibrating the virtual auditory display device can refer to calibrating the IMU. Individualizing the virtual auditory display can refer to selecting a set of virtual auditory display filters for the wearer and / or modifying an existing set of virtual auditory display filters such that the virtual auditory display provided by the virtual auditory display device is customized for the wearer. The virtual auditory display system 102 can allow the wearer to perform both the calibration portion and the individualization portion of the calibration and / or individualization process, to perform only the calibration portion, or to perform only the individualization portion.
[0242] Figures 15A-15C An example user interface 1500 for calibrating a virtual auditory display device in some embodiments is depicted. The virtual auditory display device can be the virtual auditory display device 100 that includes the first ear-wearable device 102a and the second ear-wearable device 102. Figures 15A-15F The virtual auditory display device 100 is described with reference to, but other virtual auditory display devices can be calibrated and / or individualized.
[0243] The virtual auditory display system 102 (e.g., the user interface module 210) can provide the user interface 1500. The wearer can start the calibration and / or individualization process by selecting a button labeled “Start” (not shown in FIG. 15A) displayed by the virtual auditory display system 102. Figures 15A-15C Figure 15A The user interface 1500 is depicted providing a user interface element 1502 indicating a point in the calibration portion of the calibration and / or individualization process for the wearer along with instructions 1504 for the wearer.
[0244] Figure 15B A user interface 1500 providing first circle 1506a and second circle 1506b is depicted. The virtual auditory display system 102 can cause the first circle 1506a and / or the second circle 1506b to move up and down on the user interface 1500 and instruct the wearer to nod the head up and down to follow the first circle 1506a and the second circle 1506b.
[0245] Figure 15C A user interface 1500 providing circle 1508 is depicted. The virtual auditory display system 102 can cause the circle 1508 to move up and down on the user interface 1500 and instruct the wearer to nod the head up and down to follow the circle 1508. Additionally or alternatively, the virtual auditory display system 102 can cause the circle 1508 to move from side to side on the user interface 1500 and instruct the wearer to move their head from side to side to follow the circle 1508.
[0246] When the virtual auditory display system 102 performs the calibration portion of the calibration and / or personalization process, the virtual auditory display system 102 can receive a detection of a head orientation of the wearer’s head from the virtual auditory display device 100 based on data obtained from the IMU-based sensor system and / or other sensors of the first ear-wearable device 102a and / or the second ear-wearable device 102. The virtual auditory display system 102 can use the detection of the head orientation and other factors, such as a known or estimated distance from a display providing the user interface 1500, a width and height of the display, a position of the first circle 1506a, the second circle 1506b, and / or the circle 1508, and / or other data from the IMU-based sensor system, to calibrate the IMU-based sensor system.
[0247] Figures 15D-15F An example user interface 1550 for personalizing the virtual auditory display provided by the virtual auditory display device in some embodiments is depicted. The virtual auditory display system 102 (e.g., the user interface module 210) can provide the user interface 1550.
[0248] The wearer of the virtual auditory display device 100 can begin the personalization portion of the calibration and / or personalization process after completing the calibration portion. The virtual auditory display system 102 can cause the virtual auditory display device 100 to play sounds at several locations (e.g., five locations). The sounds can include, for example, objects that the wearer believes to be moving around his or her head, such as airplanes, helicopters, birds, and other flying creatures producing sounds. Figure 15D A user interface 1550 providing instructions 1554 instructing the wearer to point their nose toward the source of each sound as the virtual auditory display device 100 plays the sounds is depicted. The wearer can begin the personalization portion of the calibration and / or personalization process by selecting a button 1556 labeled “Continue.”
[0249] Figure 15E A user interface 1550 is depicted that provides a user interface element 1552 indicating a point in the individualization portion of the calibration and / or individualization process at which the wearer is positioned. Figure 15F A user interface 1550 is depicted in which the user interface element 1552 indicates that the wearer has positioned the virtual auditory display device 100 to play sound at a first location. The virtual auditory display system 102 can cause the virtual auditory display device 100 to play sound at a subsequent location and update the user interface 1550 accordingly.
[0250] As the virtual auditory display system 102 performs the individualization portion of the calibration and / or individualization process, the virtual auditory display system 102 can receive, from the virtual auditory display device 100, a detection of a head orientation of the wearer’s head based on data obtained from the IMU-based sensor systems and / or other sensors of the first ear-wearable device 102a and / or the second ear-wearable device 102. As described with reference to, for example Figure 13A and Figure 13B The virtual auditory display system 102 can use the detection of the head orientation and the location of the sound generated by the virtual auditory display system 102 to compute one or more amounts of change. The virtual auditory display system 102 can use the computed one or more amounts of change to select a set of virtual auditory display filters and estimate a spatialization accuracy of the virtual auditory display for the wearer.
[0251] Figures 15G-15J An example user interface 1570 is depicted that, in some embodiments, is used to provide information about the calibration of the virtual auditory display device and the individualization of the virtual auditory display of the virtual auditory display device. The user interface 1570 includes a recommendation 1572 of a set of virtual auditory display filters. In some embodiments, as described herein with reference to, for example Figure 13A and Figure 13B The virtual auditory display system 102 can select a set of virtual auditory display filters from among a plurality of sets of virtual auditory display filters based on the results of the calibration and / or individualization process, as described herein with reference to, for example
[0252] The virtual auditory display system 102 can base the estimate 1574, such as “very good” ( Figure 15G ), “medium” ( Figure 15H ), and “poor” ( Figure 15A), the spatialization accuracy of the virtual auditory display to the wearer is classified. The virtual auditory display system 102 can provide a recommendation to re-do the calibration and / or personalization process, and / or to use a custom filter. The user interface 1570 also includes a button 1576 labeled "Continue" that the wearer can select to return to the user interface 1100 depicted in FIGS. 11A and 11B. Figure 11A and Figure 11B depicted in FIGS. 11A and 11B.
[0253] Although the virtual auditory display system 102 is described as using circles, the virtual auditory display system 102 can use other visual user interface elements in the calibration and / or personalization process. Further, although the virtual auditory display system 102 is described as receiving detection of head orientation from the virtual auditory display device 100 in the calibration and / or personalization process, the virtual auditory display system 102 can receive detection of head orientation from other devices connected to the virtual auditory display system 102, such as cameras, motion sensing devices, virtual reality headsets, etc.
[0254] One advantage of the calibration and / or personalization process is that the virtual auditory display system 102 can personalize a set of virtual auditory display filters for a wide range of individuals. The virtual auditory display system 102 can personalize the set of virtual auditory display filters by modifying the set of virtual auditory display filters. The virtual auditory display system 102 can have multiple sets of virtual auditory display filters pre-configured, and can modify the set of virtual auditory display filters by selecting a different set of virtual auditory display filters based on the results of the user's calibration and / or personalization process.
[0255] Additionally or alternatively, the virtual auditory display system 102 can modify the set of virtual auditory display filters by modifying a digital filter or function included in the set of virtual auditory display filters. For example, in the case that the set of virtual auditory display filters includes a digital filter, the virtual auditory display system 102 can modify parameters of the digital filter, such as a center frequency, a gain, a q, an algorithm type, or other parameters, based on the results of the user's calibration and / or personalization process.
[0256] Personalization of the virtual auditory display filters allows a wide range of individuals to experience immersive, accurately rendered sound in a virtual auditory space. Moreover, these individuals do not have to have HRTFs generated for them using potentially difficult and / or unreliable physical measurement procedures. Such individuals can simply obtain a set of personalized virtual auditory display filters by having the virtual auditory display system 102 perform a calibration and / or personalization process for them. Modification of the virtual auditory display filters can be performed at an initial setup procedure for a person as well as at any subsequent point during the person's use of the virtual auditory display system 102 and / or virtual auditory display device 100.
[0257] One advantage of the virtual auditory display filters is that sound in many more locations in the virtual auditory space can be rendered compared to prior art. For example, a 9.1.6 configuration has 16 virtual speakers and can therefore be limited to accurately rendering sound for only those 16 virtual speaker locations. Such a configuration can render sound from other locations by blurring the sound from the virtual speaker locations to represent the other locations, but such artifacts can be apparent to the listener.
[0258] In contrast, the virtual auditory display filters are able to render sound at many more locations. For example, using locations with an increase in azimuth and elevation of one degree results in 65,160 locations. However, the described technology can generate virtual auditory display filters with smaller increases, resulting in even more locations at which the virtual auditory display filters can render sound. Moreover, typical approaches render sound at a modeled distance of 1 m from a center point representing the listener. The described technology can generate virtual auditory display filters for any number of distances from the center point. Thus, the described technology can accurately render sound at different distances.
[0259] One advantage of the described technology is that the described technology accurately renders virtual auditory display sound in a virtual auditory space, meaning that the sound is perceived by the listener as coming from the location at which the sound creator intended to emit the sound. Another advantage of the described technology is that the described technology can be utilized with any ear-worn device, such as headphones, headsets, and earbuds. Another advantage is that the virtual auditory display sound is high quality and clear. Another advantage is that the described technology can emphasize or de-emphasize sound in particular regions or locations of the virtual auditory space in order to focus the listener's attention on those particular regions or locations. Such an approach can improve the listening abilities of the listener and allow the listener to hear sounds that the listener would not otherwise hear.
[0260] Another advantage of the described technology is that any digital device with appropriate storage and processing capabilities can store and apply the virtual auditory display filter. As described herein, a general purpose computing device such as a laptop or desktop computer can store the virtual auditory display filter and apply it to an audio signal to generate a processed audio signal. The laptop or desktop computer can then send the processed audio signal to an ear-wearable device to generate sound based on the processed audio signal. Similarly, a digital device such as a phone, tablet, or virtual reality headset can store and apply the virtual auditory display filter to an audio signal to generate a processed audio signal and send the processed audio signal to an ear-wearable device.
[0261] Additionally or alternatively, an ear-wearable device such as the virtual auditory display device 100 described herein can store and apply the virtual auditory display filter. The ear-wearable device can receive an input audio signal from, for example, a digital device paired with the ear-wearable device such as a phone or tablet or from a cloud-based service. The ear-wearable device can apply the stored virtual auditory display filter to the input audio signal to generate a processed audio signal and output a virtual auditory display sound based on the processed audio signal. Another example is that a cloud-based service can store and apply the virtual auditory display filter to generate a processed audio signal and send the processed audio signal to an ear-wearable device. Other advantages will be apparent.
[0262] Figure 16 A block diagram of an example digital device 1600 in accordance with some embodiments is depicted. The digital device 1600 is shown in the form of a general purpose computing device. The digital device 1600 includes at least one processor 1602, RAM 1604, a communication interface 1606, input / output devices 1608, a storage device 1610, and a system bus 1612 that couples the various system components including the storage device 1610 to the at least one processor 1602. A system such as a computing system can be or include one or more of the digital device 1600.
[0263] The system bus 1612 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0264] Digital device 1600 generally includes various computer system readable media, such as computer system readable storage media. Such media can be any available media that is accessible by any of the systems described herein and includes both volatile and non-volatile media, removable and non-removable media.
[0265] In some embodiments, at least one processor 1602 is configured to execute executable instructions (e.g., programs). In some embodiments, at least one processor 1602 includes circuitry or any processor capable of processing executable instructions.
[0266] In some embodiments, RAM 1604 stores programs and / or data. In various embodiments, working data is stored within RAM 1604. Data within RAM 1604 can be cleared or eventually passed to storage 1610, such as prior to a digital device 1600 being reset and / or powered off.
[0267] In some embodiments, digital device 1600 is coupled to a network via communication interface 1606. Digital device 1600 can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet).
[0268] In some embodiments, input / output device 1608 is any device that inputs data (e.g., a mouse, a keyboard, a stylus, a sensor, etc.) or outputs data (e.g., a speaker, a display, a virtual reality headset).
[0269] In some embodiments, storage 1610 can include computer system readable media in the form of non-volatile memory, such as read only memory (ROM), programmable read only memory (PROM), solid state drives (SSDs), flash memory, and / or cache memory. Storage 1610 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage 1610 can be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a hard disk). Storage 1610 can include non-transitory computer readable media or a plurality of non-transitory computer readable media, which stores for execution by a computer system, such as digital device 1600, one or more sets of instructions Figure 2A 、 Figure 2B and Figure 3Bprograms or applications described. Although not shown, a disk drive can be provided to read from and write to a removable, non- volatile magnetic media (e.g., a "floppy drive" as represented by floppy disk 1625). Although not shown, an optical disk drive can be provided to read from and / or write to a removable, non- volatile optical media (e.g., a CD-ROM, DVD-ROM, etc.), as represented by optical disk 1626. In such instances, each can be connected to bus 1612 by one or more data media interfaces. As will be further depicted and described below, memory 1604 can include, among other things, a
[0270] One or more application programs 1622 can be loaded into the memory 1604 and run under the control of operating system 1621, which can be stored in the same or different memory locations within the memory 1604. By way of example, an application program implementing some or all of the aspects of virtual auditory display system 102 can be loaded into memory 1604. The system 1600, including memory 1604, can be used to implement memory 102 of FIG. 1.
[0271] It should be understood that, although not shown explicitly, other hardware and / or software components can be used in conjunction with the digital device 1600. These components include, but are not limited to, microcode, device drivers, redundant processing units, and external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0272] The exemplary embodiments are described herein with reference to the accompanying drawings. The disclosure can be implemented in various ways, and thus should not be construed as being limited to the embodiments disclosed herein. Rather, those embodiments are provided so that the disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0273] It should be understood that aspects of one or more embodiments can be implemented as a system, method, or computer program product. Accordingly, aspects can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects that can all generally be referred to herein as a "circuit," "module" or "system." Furthermore, aspects can take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.
[0274] Any combination of one or more computer readable medium can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a solid state drive (SSD), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0275] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof.
[0276] Computer readable program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0277] Computer program code for carrying out operations of aspects of the present application can be written in any combination of one or more programming languages, including one or more of a compiled or interpreted programming language such as Java, Smalltalk, C++, Python, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer program code can execute entirely on any of the systems described herein, or any combination of the systems described herein.
[0278] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0279] These computer program instructions can also be stored in a computer- readable medium that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0280] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0281] Although specific examples are described above for the purpose of illustration, various modifications can be made to the described examples. For example, although processes or blocks are presented in a given order, alternative implementations can perform routines having steps in a different order, or employ systems having blocks in a different order, and some processes or blocks can be deleted, moved, added, subdivided, combined, and / or modified to provide alternative or sub-combinations. Each of these processes or blocks can be implemented in a variety of different ways. Also, while processes or blocks are at times shown as being performed in series, these processes or blocks can instead be performed concurrently or in parallel, or at different times, and with desired results. Additionally, any specific numbers noted herein are merely meant to be illustrative: alternative implementations can employ different values or ranges.
[0282] Throughout this specification, plural instances can implement components, operations, or structures described as a single instance. The instances can be implemented as individual components or combinations of two or more components. Structures and functionality presented as separate components in example configurations can be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component can be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein. Additionally, any specific numbers noted herein are merely meant to be illustrative: alternative implementations can employ different values or ranges.
[0283] Components can be described or shown as being within other components or connected to other components. Such depictions are merely examples and other implementations can be employed. Components can be described or shown as being "coupled" with, "couplable" with, "operably coupled" with, "communicatively coupled" with, or otherwise "coupled with" other components. Such depictions can indicate that such components are in physical, electrical, or communicative contact with each other or otherwise cooperate or interact.
[0284] Components can be described or claimed as "configured" to perform a task or task(s) and include circuitry configured to perform the task(s) and / or one or more processors configured with executable instructions to perform the task(s), collectively configured to perform the task(s). Such components can include, in various examples, one or more storage mediums establishing one or more algorithms configurable to perform the task(s). Similarly, components can be described or claimed as "operable" to perform a task or task(s), even if the components are not configured to perform the task(s) at a given instance of time. In some examples, a component can be "operable" to perform a task at one instance of time and "not operable" to perform the task at another instance of time.
[0285] The use of "or" in the disclosure is intended to encompass "and / or" unless otherwise indicated. Rather, "or" should be interpreted as inclusive or, meaning one or the other and / or both, to the extent allowed by law. For example, a phrase such as "provide product or service" is intended to encompass the options of "provide product" and "provide service" and "provide product and service."
[0286] It can be apparent that various modifications can be made, and other embodiments can be used, without departing from the broader scope of what is discussed herein. For example, the virtual auditory display system 102 can utilize a bank of FIR filters for each of certain locations in the virtual auditory space, and a bank of IIR filters for each of other certain locations in the virtual auditory space. As another example, the virtual auditory display system 102 can provide audio signals to any device capable of directing sound toward a listener's ear. As another example, the virtual auditory display device can be any device or group of devices capable of producing sound based on the output audio signals generated by the virtual auditory display system 102, such as a pair of speakers.
[0287] Accordingly, the disclosure is intended to embrace all such alterations, modifications, and variations that fall within the scope of the example embodiments.
Claims
1. A non-transitory computer-readable medium comprising one or more executable instructions, said executable instructions causing the system to perform a method when executed by one or more processors of the system, said method comprising: One or more first digital filters are generated for each of a plurality of virtual auditory spatial locations, the one or more first digital filters including one or more first notch filters, the one or more first notch filters including one or more first center frequencies, the one or more first center frequencies being based on a first approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more first notch filters being configured to generate one or more first notches in a first spectrum of the first audio signal based on the one or more first center frequencies when applied to a first audio signal; One or more second digital filters are generated for each of the plurality of virtual auditory spatial locations, the one or more second digital filters including one or more second notch filters, the one or more second notch filters including one or more second center frequencies, the one or more second center frequencies being based on a second approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more second notch filters being configured to generate one or more second notches in a second spectrum of the second audio signal based on the one or more second center frequencies when applied to a second audio signal; Receive an audio signal, the audio signal having one or more audio sub-signals, the audio sub-signals being associated with a virtual auditory spatial location; For each of the one or more audio sub-signals: Based on the virtual auditory spatial location associated with the audio sub-signal, select one or more specific first digital filters and one or more specific second digital filters; One or more first digital filters are applied to the audio sub-signal to obtain a first processed audio sub-signal; as well as One or more of the specific second digital filters are applied to the audio sub-signal to obtain a second processed audio sub-signal; Based on multiple first processed audio sub-signals, a first output audio signal is generated for the first device; A second output audio signal is generated for the second device based on multiple second processed audio sub-signals; as well as The first output audio signal is provided to the first device, and the second output audio signal is provided to the second device.
2. The one or more non-transitory computer-readable media of claim 1, wherein the virtual auditory spatial location is a first virtual auditory spatial location, and the method further comprises: Receive user header orientation; as well as For each of the one or more audio sub-signals, a second virtual auditory spatial position is determined based on the first virtual auditory spatial position associated with the audio sub-signal and the head orientation. The selection of the specific one or more first digital filters and the specific one or more second digital filters based on the virtual auditory spatial location associated with the audio sub-signal includes the selection of the specific one or more first digital filters and the specific one or more second digital filters based on the second virtual auditory spatial location.
3. The one or more non-transitory computer-readable media of claim 2, wherein the particular one or more first digital filters are first-specific one or more first digital filters, the particular one or more second digital filters are first-specific one or more second digital filters, the user's head orientation is the user's first head orientation, and the method further comprises: Receive personalized audio signals associated with a third virtual auditory spatial location; Based on the third virtual auditory spatial location, select one or more first digital filters and one or more second digital filters of a second specific nature; One or more first digital filters of the second specific type are applied to the personalized audio signal to obtain a first processed personalized audio signal; One or more second digital filters of the second specific type are applied to the personalized audio signal to obtain a second processed personalized audio signal; Based on the first processed personalized audio signal, a third output audio signal is generated for the first device; Based on the second processed personalized audio signal, a fourth output audio signal is generated for the second device; The third output audio signal is provided to the first device, and the fourth output audio signal is provided to the second device; Receive the user's second head orientation; The location of the fourth virtual auditory space is determined based on the second head orientation. Determine the increment between the third virtual auditory spatial position and the fourth virtual auditory spatial position; as well as The one or more first digital filters and the one or more second digital filters are modified based on the change.
4. The one or more non-transitory computer-readable media of claim 3, wherein modifying the one or more first digital filters and the one or more second digital filters based on the change includes modifying the one or more first center frequencies on which the one or more first notch filters are based and the one or more second center frequencies on which the one or more second notch filters are based.
5. The method of one or more non-transitory computer-readable media according to claim 1, further comprising generating a first notch mask and a second notch mask using one or more image processing algorithms, wherein the first notch mask specifies a first gain modifier as a function of virtual auditory spatial location, and the second notch mask specifies a second gain modifier as a function of virtual auditory spatial location, wherein: The one or more first notch filters include the one or more first center frequencies and a first gain modified by the first gain modifier, and the one or more first notch filters are configured to generate one or more first notches in the first spectrum of the first audio signal based on the one or more first center frequencies and the first gain when applied to the first audio signal. The one or more second notch filters include the one or more second center frequencies and a second gain modified by the second gain modifier, and the one or more second notch filters are configured to generate one or more second notches in the second spectrum of the second audio signal based on the one or more second center frequencies and the second gain when applied to the second audio signal.
6. The one or more non-transitory computer-readable media of claim 5, wherein the one or more image processing algorithms include one or more of a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a thresholding function, an edge detection function, and a segmentation function.
7. The method further comprises: one or more non-transitory computer-readable media according to claim 1. The receiver's choice of acoustic environment; as well as Based on the acoustic environment, a first acoustic environment digital filter and a second acoustic environment digital filter are determined. Wherein, for each of the one or more audio sub-signals, applying the specific one or more first digital filters to the audio sub-signal to obtain the first processed audio sub-signal includes applying the specific one or more first digital filters and the first acoustic environment digital filter to the audio sub-signal to obtain the first processed audio sub-signal, and applying the specific one or more second digital filters to the audio sub-signal to obtain the second processed audio sub-signal includes applying the specific one or more second digital filters and the second acoustic environment digital filter to the audio sub-signal to obtain the second processed audio sub-signal.
8. The one or more non-transitory computer-readable media of claim 7, wherein the acoustic environment is represented by one or more surround sound arrays, and determining the first acoustic environment digital filter and the second acoustic environment digital filter based on the acoustic environment comprises determining the first acoustic environment digital filter and the second acoustic environment digital filter based on the one or more surround sound arrays.
9. The one or more non-transitory computer-readable media of claim 1, wherein the one or more first digital filters and the one or more second digital filters are infinite impulse response filters.
10. One or more non-transitory computer-readable media according to claim 1, wherein the first device comprises a first ear-worn device, and the second device comprises a second ear-worn device.
11. A system comprising at least one processor and at least one memory, said at least one memory including executable instructions that, when executed by said at least one processor, cause the system to: One or more first digital filters are generated for each of a plurality of virtual auditory spatial locations, the one or more first digital filters including one or more first notch filters, the one or more first notch filters including one or more first center frequencies, the one or more first center frequencies being based on a first approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more first notch filters being configured to generate one or more first notches in a first spectrum of the first audio signal based on the one or more first center frequencies when applied to a first audio signal; One or more second digital filters are generated for each of the plurality of virtual auditory spatial locations, the one or more second digital filters including one or more second notch filters, the one or more second notch filters including one or more second center frequencies, the one or more second center frequencies being based on a second approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location, the one or more second notch filters being configured to generate one or more second notches in a second spectrum of the second audio signal based on the one or more second center frequencies when applied to a second audio signal; Receive an audio signal, the audio signal having one or more audio sub-signals, the audio sub-signals being associated with a virtual auditory spatial location; For each of the one or more audio sub-signals: Based on the virtual auditory spatial location associated with the audio sub-signal, select one or more specific first digital filters and one or more specific second digital filters; One or more first digital filters are applied to the audio sub-signal to obtain a first processed audio sub-signal; as well as One or more of the specific second digital filters are applied to the audio sub-signal to obtain a second processed audio sub-signal; Based on multiple first processed audio sub-signals, a first output audio signal is generated for the first device; A second output audio signal is generated for the second device based on multiple second processed audio sub-signals; as well as The first output audio signal is provided to the first device, and the second output audio signal is provided to the second device.
12. The system of claim 11, wherein the virtual auditory spatial location is a first virtual auditory spatial location, and the executable instructions, when executed by the at least one processor, further cause the system to: Receive user header orientation; and For each of the one or more audio sub-signals, a second virtual auditory spatial position is determined based on the first virtual auditory spatial position associated with the audio sub-signal and the head orientation. Selecting the specific one or more first digital filters based on the virtual auditory spatial location associated with the audio sub-signal includes selecting the specific one or more first digital filters based on the second virtual auditory spatial location, and selecting the specific one or more second digital filters based on the virtual auditory spatial location associated with the audio sub-signal includes selecting the specific one or more second digital filters based on the second virtual auditory spatial location.
13. The system of claim 12, wherein the one or more first digital filters are first or more first digital filters, the one or more second digital filters are first or more second digital filters, the specific one or more first digital filters are first specific one or more first digital filters, the specific one or more second digital filters are first specific one or more second digital filters, the head orientation is a first head orientation, the audio signal having one or more audio sub-signals is a first audio signal having the first or more audio sub-signals, and the executable instructions, when executed by the at least one processor, further cause the system to: Receive personalized audio signals with a third virtual auditory spatial location; Based on the third virtual auditory spatial location, select one or more first digital filters and one or more second digital filters of a second specific nature; One or more first digital filters of the second specific type are applied to the personalized audio signal to obtain a first processed personalized audio signal; One or more second digital filters of the second specific type are applied to the personalized audio signal to obtain a second processed personalized audio signal; Based on the first processed personalized audio signal, a third output audio signal is generated for the first device; Based on the second processed personalized audio signal, a fourth output audio signal is generated for the second device; The third output audio signal is provided to the first device, and the fourth output audio signal is provided to the second device; Receive the user's second head orientation; The location of the fourth virtual auditory space is determined based on the second head orientation. Determine the amount of change between the third virtual auditory spatial position and the fourth virtual auditory spatial position; as well as Based on the change, a second or more first digital filters and a second or more second digital filters are selected for use when receiving a second input audio signal having a second or more audio sub-signals.
14. The system of claim 11, wherein the executable instructions, when executed by the at least one processor, further cause the system to use one or more image processing algorithms to generate a first notch mask and a second notch mask, the first notch mask specifying a first gain modifier based on the virtual auditory spatial position, and the second notch mask specifying a second gain modifier based on the virtual auditory spatial position, wherein: The one or more first notch filters are generated using the one or more first center frequencies based on a first approximately S-shaped distribution of the center frequency as a function of virtual auditory spatial location and a first gain modified by the first gain modifier, and the one or more first notch filters are configured to generate one or more first notches in the first spectrum of the first audio signal based on the one or more first center frequencies and the first gain when applied to the first audio signal. The one or more second notch filters are generated using the one or more second center frequencies based on the second approximately S-shaped distribution of the center frequency as a function of the virtual auditory spatial location and the second gain modified by the second gain modifier, and the one or more second notch filters are configured to generate one or more second notches in the second spectrum of the second audio signal based on the one or more second center frequencies and the second gain when applied to the second audio signal.
15. The system of claim 11, wherein the executable instructions, when executed by the at least one processor, further cause the system to: The selection of the acoustic environment for receiving; and Based on the acoustic environment, a first acoustic environment digital filter and a second acoustic environment digital filter are determined. Wherein, for each of the one or more audio sub-signals, applying the specific one or more first digital filters to the audio sub-signal to obtain the first processed audio sub-signal includes applying the specific one or more first digital filters and the first acoustic environment digital filter to the audio sub-signal to obtain the first processed audio sub-signal, and applying the specific one or more second digital filters to the audio sub-signal to obtain the second processed audio sub-signal includes applying the specific one or more second digital filters and the second acoustic environment digital filter to the audio sub-signal to obtain the second processed audio sub-signal.
16. The system of claim 11, wherein the one or more first digital filters and the one or more second digital filters are infinite impulse response filters.
17. The system of claim 11, wherein the first device includes a first ear-worn device, and the second device includes a second ear-worn device.
18. A method comprising: A first virtual auditory display filter is generated, the first virtual auditory display filter including a first set of first functions, one or more of the first functions generating a first processed audio signal having a first frequency response when applied to a first audio signal having a first position in a virtual auditory space, the first frequency response having one or more first notches at one or more first center frequencies based on the first position, the one or more first notches having one or more first peak-valley depths of up to -10dB. A second virtual auditory display filter is generated, the second virtual visual display filter including a second set of second functions, one or more of the second functions generating a second processed audio signal having a second frequency response when applied to the first audio signal, the second frequency response having one or more second notches at one or more second center frequencies based on the first position, the one or more second notches having one or more second peak-valley depths of up to -10dB; Receive a second audio signal having a second position in the virtual auditory space; The first virtual auditory display filter, comprising a first subset of a first function selected based on the second position, is applied to the second audio signal to generate a third processed audio signal with a third frequency response; The second virtual auditory display filter, comprising a second subset of a second function selected based on the second position, is applied to the second audio signal to generate a fourth processed audio signal with a fourth frequency response; The third processed audio signal is provided to the first sound output device; as well as The fourth processed audio signal is provided to the second sound output device.
19. The method of claim 18, wherein the one or more first center frequencies are based on a first approximately S-shaped distribution of center frequencies as a function of position in the virtual auditory space, and the one or more second center frequencies are based on a second approximately S-shaped distribution of center frequencies as a function of position in the virtual auditory space.
20. The method of claim 18, further comprising receiving the user's head orientation, wherein: Applying a first virtual auditory display filter, comprising a first subset of the first function selected based on the second position, to the second audio signal to generate a third processed audio signal having a third frequency response includes applying the first virtual auditory display filter, comprising a third subset of the first function selected based on the second position, and the head orientation to the second audio signal to generate the third processed audio signal having the third frequency response. Applying the second virtual auditory display filter, comprising a second subset of the second function selected based on the second position, to the second audio signal to generate a fourth processed audio signal having a fourth frequency response includes applying the second virtual auditory display filter, comprising a fourth subset of the second function selected based on the second position, and the head orientation to the second audio signal to generate the fourth processed audio signal having the fourth frequency response.
21. The method of claim 18, further comprising: A first notch mask and a second notch mask are generated using one or more image processing algorithms. The first notch mask specifies a first depth modifier as a function of the position in the virtual auditory space, and the second notch mask specifies a second depth modifier as a function of the position in the virtual auditory space. Modify the depth of one or more first peaks and valleys based on the first depth modifier; and The depth of the one or more second peaks and valleys is modified based on the second depth modifier.
22. The method of claim 21, wherein the one or more image processing algorithms include one or more of a Gaussian function, a sharpening function, a contrast adjustment function, a color correction function, a thresholding function, an edge detection function, and a segmentation function.
23. The method of claim 21, wherein the first set of first functions includes a first infinite impulse response digital filter, and the second set of second functions includes a second infinite impulse response digital filter.
24. A method comprising: A plurality of first digital filters are received, and one or more first digital filters are generated for each of a plurality of virtual auditory spatial locations. The one or more first digital filters include one or more first notch filters, and the one or more first notch filters include one or more first center frequencies. The one or more first center frequencies are based on a first approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location. The one or more first notch filters are configured to generate one or more first notches in a first spectrum of the first audio signal based on the one or more first center frequencies when applied to a first audio signal. A plurality of second digital filters are received, and one or more second digital filters are generated for each of a plurality of virtual auditory spatial locations. The one or more second digital filters include one or more second notch filters, and the one or more second notch filters include one or more second center frequencies. The one or more second center frequencies are based on a second approximately S-shaped distribution of center frequencies as a function of virtual auditory spatial location. The one or more second notch filters are configured to generate one or more second notches in a second spectrum of the second audio signal based on the one or more second center frequencies when applied to a second audio signal. Receive personalized audio signals with virtual auditory spatial location; Based on the virtual auditory spatial location, select one or more specific first digital filters and one or more specific second digital filters; One or more first digital filters are applied to the personalized audio signal to obtain a first processed personalized audio signal; One or more of the specific second digital filters are applied to the personalized audio signal to obtain a second processed personalized audio signal; Provide a first output audio signal based on the first processed personalized audio signal to a first device, and provide a second output audio signal based on the second processed personalized audio signal to a second device; Receive user perception of the first sound output by the first device and the second sound output by the second device; as well as Modify the set of multiple first digital filters and the set of multiple second digital filters based on the user's perception.
25. The method of claim 24, wherein the virtual auditory spatial location is a first virtual auditory spatial location, and wherein modifying the plurality of first digital filters and the plurality of second digital filters based on the user perception comprises: The location of the second virtual auditory space is determined based on the user's perception. Determine the amount of change between the first virtual auditory spatial position and the second virtual auditory spatial position; as well as Based on the change, modify the set of multiple first digital filters and the set of multiple second digital filters.
26. The method of claim 25, wherein receiving the user perception includes receiving the user's head orientation, and wherein determining the second virtual auditory spatial position based on the user perception includes determining the second virtual auditory spatial position based on the user's head orientation.
27. The method of claim 25, wherein receiving the user perception includes receiving one or more gestures of the user, and wherein determining the second virtual auditory spatial position based on the user perception includes determining the second virtual auditory spatial position based on the one or more gestures of the user.
28. The method of claim 24, wherein the plurality of first digital filters is a first plurality of first digital filters, the plurality of second digital filters is a first plurality of second digital filters, wherein modifying the plurality of first digital filters based on the user perception includes selecting a second plurality of first digital filters based on the user perception, and wherein modifying the plurality of second digital filters based on the user perception includes selecting a second plurality of second digital filters based on the user perception.
29. The method of claim 24, wherein modifying the set of multiple first digital filters based on the user perception includes modifying the one or more first center frequencies, and wherein modifying the set of multiple second digital filters based on the user perception includes modifying the one or more second center frequencies.
30. The method of claim 24, further comprising: The spatialization accuracy estimate is determined based on the user's perception. as well as Provide the spatialization accuracy estimate.