How to create linearly interpolated head transfer functions

By converting HRTFs to use full-pass filters with adjusted phase responses, the method enhances the efficiency of binaural audio processing, enabling effective linear interpolation and reduced complexity for immersive audio applications.

JP2026511605APending Publication Date: 2026-04-14DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing binaural audio processing techniques face inefficiencies in generating modified head-related transfer functions (HRTFs) that can be efficiently approximated by linear mixing while retaining psychoacoustic properties.

Method used

The method involves converting a first set of HRTFs to a second set by replacing delay components with full-pass filters, adjusting phase responses to be substantially linear below a threshold frequency and reducing interaural phase differences above the threshold, allowing for a more compact set of basis filters to be used for generating left and right audio data.

Benefits of technology

This approach enables efficient linear interpolation of HRTFs, maintaining audio quality and reducing computational complexity, suitable for immersive audio services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511605000001_ABST
    Figure 2026511605000001_ABST
Patent Text Reader

Abstract

Systems, devices, and methods are described for determining a "coupled" pair of left / right ear HRTFs adapted from the original pair of left / right ear HRTFs. The interaural delay of the coupled HRTF is formed using a whole-pass filter that provides the correct interaural delay at low frequencies. The whole-pass filter is adapted to limit the interaural phase difference at high frequencies. Furthermore, a low-complexity process for the rapid generation of a suitable whole-pass filter is described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 455,539 filed on 29 March 2023, U.S. Provisional Patent Application No. 63 / 595,752 filed on 2 November 2023, and U.S. Provisional Patent Application No. 63 / 567,376 filed on 19 March 2024, the entire contents of which are incorporated herein by reference.

[0002] Technical field This disclosure relates to the creation of a modified head-related transfer function (HRTF) from the original HRTF. [Background technology]

[0003] Unless otherwise specified herein, the methods described in this section are not prior art to the claims of this application, nor will they be deemed prior art by being included in this section.

[0004] Binaural audio signals contain two audio channels intended to be played back to the listener through each of their two (left and right) ears. Binaural playback can be achieved through loudspeakers placed near each ear, or through headphones (including over-ear and in-ear headphones).

[0005] Binaural signals can be generated by processing a source audio signal using a pair of head-related transfer function (HRTF) filter responses. HRTF responses can be defined in many ways, such as as time-domain impulse responses or frequency-domain responses. Typically, HRTF responses are grouped into pairs to provide a response for each ear transducer.

[0006] When used to process audio signals, HRTF filter pairs can be used to provide listeners with an experience that mimics the sound (in each ear) that would occur if the audio signal were presented from a particular direction of arrival. Different HRTF filter pairs produce different perceptions of sound source direction.

[0007] A pair of reference HRTF filters associated with a particular direction of arrival can be determined by measuring the acoustic transfer function from a sound source located at a certain distance in the same direction to each ear of the listener. Alternatively, the reference HRTF filters may be determined by other means, including numerical simulation or acoustic measurements of a mannequin.

[0008] A pair of modified HRTF filters may differ from a pair of acoustically measured HRTF filters, yet they still provide the listener with the desired impression of sound from the same direction. In particular, the phase difference between the high-frequency portions of the left and right modified HRTF filters may be substantially different from the phase difference between the high-frequency portions of the left and right reference HRTF filters without significant loss of perceived listener experience. This is possible because the interaural phase difference in the high-frequency range is of little importance to listener perception.

[0009] The HRTF set function is a function that determines the HRTF filters for the left and right ears, given the direction of arrival.

number

[0010] In Equation 1, the HRTF set function H(x,y,z) is given the direction of arrival in the form of a 3D unit vector (x,y,z), and the function returns a pair of HRTF filters for the left and right ears.

[0011] The disclosures made herein are presented with respect to these and other circumstances. [Overview of the project] [Problems that the invention aims to solve]

[0012] Techniques for processing audio signals are described. Various examples described herein provide systems, methods, and / or devices for creating and using modified HRTF filters with alternative high-frequency phase responses. [Means for solving the problem]

[0013] According to some exemplary embodiments, an audio processing method for a control system comprising one or more processors may involve the control system obtaining a first set of head-related transfer functions (HRTFs) and the control system converting the first set of HRTFs to a second set of HRTFs. In some exemplary embodiments, the conversion may involve replacing the delay components of the first set of HRTFs with full-pass filters in the second set of HRTFs. In some exemplary embodiments, the conversion may involve adjusting the phase response of each full-pass filter in the second set of HRTFs such that each interaural phase response is substantially linear at frequencies below the relevant threshold frequency of the corresponding full-pass filter, and each phase response has a reduced interaural phase difference at frequencies above the relevant threshold frequency of the corresponding full-pass filter.

[0014] In some exemplary embodiments, the method may involve outputting a second set of HRTFs. According to some exemplary embodiments, outputting a second set of HRTFs may involve storing a second set of HRTFs, transmitting a second set of HRTFs to a device configured to process audio data, providing a second set of HRTFs for further processing, or a combination thereof.

[0015] According to some exemplary embodiments, the method may involve a control system defining a set of basis filters based on a second set of HRTFs. The set of basis filters may have fewer elements than the second set of HRTFs. In some exemplary embodiments, the method may involve a control system obtaining a bitstream of input audio data in an input audio format, and the control system combining the input audio data with one or more basis filters from the set of basis filters to generate left audio data and right audio data.

[0016] In some exemplary embodiments, the method may involve a control system outputting left audio data and right audio data. According to some exemplary embodiments, outputting left audio data and right audio data may involve storing left audio data and right audio data, transmitting left audio data and right audio data, providing left audio data and right audio data to a set of loudspeakers for playback by a control system, providing left audio data and right audio data for further processing, or a combination thereof.

[0017] According to some exemplary embodiments, the transformation may also involve obtaining left-ear HRTFs and right-ear HRTFs from a first set of HRTFs, identifying the left-ear non-delayed impulse response and left-ear delay from each of the left-ear HRTFs, and identifying the right-ear non-delayed impulse response and right-ear delay from each of the right-ear HRTFs. In some exemplary embodiments, the transformation may also involve generating left-ear whole-pass filters, each of which is at least partially based on an instance of the left-ear delay. According to some exemplary embodiments, the transformation may also involve generating right-ear whole-pass filters, each of which is at least partially based on an instance of the right-ear delay. In some exemplary embodiments, the transformation may also involve combining instances of the left-ear and right-ear non-delayed impulse responses with corresponding instances of the left-ear and right-ear whole-pass filters to generate HRTF pairs of a second set of HRTFs.

[0018] In some exemplary embodiments, the method may also involve generating modified left-ear delay values ​​and modified right-ear delay values ​​based on one or more of the extracted left-ear delay and right-ear delay values. The left-ear full-pass filter and the right-ear full-pass filter may be based on the modified left-ear delay values ​​and modified right-ear delay values.

[0019] According to some exemplary embodiments, generating instances of modified left-ear delay values ​​and right-ear delay values ​​may involve determining the difference between the extracted left-ear delay and the extracted right-ear delay. In some such exemplary embodiments, generating instances of modified left-ear delay values ​​and right-ear delay values ​​may involve determining the maximum expected difference between the extracted left-ear delay and the extracted right-ear delay. According to some exemplary embodiments, the difference between the extracted left-ear delay and the extracted right-ear delay may be equal to the difference between the corresponding modified left-ear delay value and the modified right-ear delay value.

[0020] In some exemplary embodiments, the modified left ear delay value and the modified right ear delay value may correspond to a smooth function. According to some exemplary embodiments, each pair of modified left ear delay value and modified right ear delay value may include a lower delay value and a higher delay value. In some examples, the lower delay value may have less delay variation than the higher delay value. In some exemplary embodiments, the non-delayed impulse response may be the minimum phase filter response.

[0021] According to some exemplary embodiments, extracting the left ear non-delayed impulse response, the right ear non-delayed impulse response, the left ear delay, and the right ear delay from each of the left and right ear HRTFs may involve determining the frequency response of the original HRTF filter for a first set of HRTFs, determining the absolute response of the original HRTF filter, and determining the minimum phase frequency response of the new non-delayed minimum phase filter. In some such exemplary embodiments, extracting the left ear non-delayed impulse response, the right ear non-delayed impulse response, the left ear delay, and the right ear delay from each of the left and right ear HRTFs may involve determining the phase response of the original HRTF filter and the phase response of the new non-delayed minimum phase filter, and determining the delay associated with the original HRTF filter, at least in part, based on the phase response of the original HRTF filter and the phase response of the new non-delayed minimum phase filter. In some such exemplary embodiments, determining the minimum phase frequency response may involve implementing a Hilbert transform with respect to the absolute response of the original HRTF filter. According to some exemplary embodiments, determining the delay associated with the original HRTF filter may also be based at least in part on a delay measurement frequency in the range of 300 Hz to 1600 Hz.

[0022] In some exemplary embodiments, the set of base filters may have at least one order of magnitude fewer elements than the second set of HRTFs. According to some exemplary embodiments, the pass-through phase response may deviate from the linear ramp phase response and may smoothly approach zero phase at frequencies above the threshold frequency.

[0023] According to some exemplary embodiments, the control system may support at least a portion of the codecs for Immersive Voice and Audio Services (IVAS).

[0024] According to some further embodiments, one or more non-temporary computer-readable media may, when executed by one or more processors, store instructions causing the one or more processors to perform an operation of any one of the methods disclosed herein.

[0025] According to some additional exemplary embodiments, an audio processor device may be configured to process input audio data. In some exemplary embodiments, the audio processor device may include a receiver unit configured to receive input audio data and a computer unit. According to some exemplary embodiments, the computer unit may be configured to take a first set of head-related transfer functions (HRTFs) and convert the first set of HRTFs to a second set of HRTFs. In some exemplary embodiments, the conversion may involve replacing the delay component of the first set of HRTFs with a full-pass filter in the second set of HRTFs. In some exemplary embodiments, the conversion may involve adjusting the phase response of each full-pass filter in the second set of HRTFs such that each interaural phase response is substantially linear at frequencies below the relevant threshold frequency of the corresponding full-pass filter, and each phase response has a reduced interaural phase difference at frequencies above the relevant threshold frequency of the corresponding full-pass filter.

[0026] In some exemplary embodiments, a computer unit may be configured to output a second set of HRTFs. According to some exemplary embodiments, outputting a second set of HRTFs may involve storing the second set of HRTFs, sending the second set of HRTFs to a device configured to process audio data, providing the second set of HRTFs for further processing, or a combination thereof.

[0027] According to some exemplary embodiments, the computer unit may be further configured to define a set of basis filters based on a second set of HRTFs. The set of basis filters may have fewer elements than the second set of HRTFs. In some exemplary embodiments, the computer unit may be further configured to take a bitstream of input audio data in an input audio format and combine the input audio data with one or more basis filters from the set of basis filters to generate left audio data and right audio data.

[0028] In some exemplary embodiments, the computer unit may be further configured to output left audio data and right audio data. According to some exemplary embodiments, outputting left audio data and right audio data may involve storing left audio data and right audio data, transmitting left audio data and right audio data, providing left audio data and right audio data to a set of loudspeakers for playback by a control system, providing left audio data and right audio data for further processing, or a combination thereof.

[0029] According to some exemplary embodiments, the audio processor device may include a storage device configured to store a first HRTF, a second HRTF, left audio data, right audio data, input audio data, or a combination thereof. In some such exemplary embodiments, the storage device may include random-access memory, read-only memory, non-temporary computer-readable media, or a combination thereof.

[0030] In some exemplary embodiments, the audio processor device may support at least a portion of the codecs for immersive voice and audio services (IVAS).

[0031] Embodiments described herein may generally be described as "technical," and the term "technical" may refer to systems, devices, methods, computer-readable instructions, modules, components, hardware logic, and / or operations, as suggested by the context in which it applies herein.

[0032] Features and technical advantages other than those explicitly described above will become apparent from reading the detailed description below and examining the relevant drawings. This summary is provided in a simplified form to introduce the selection of the technology and is not intended to identify any significant or essential features of the claimed subject matter as defined by the attached claims. [Brief explanation of the drawing]

[0033] Embodiments of the present invention are described herein merely as examples with reference to the accompanying drawings.

[0034] [Figure 1A] This block diagram shows examples of components of a device capable of implementing various aspects of this disclosure.

[0035] [Figure 1B] A schematic block diagram of an exemplary device architecture that may be used to implement various aspects of this disclosure is shown.

[0036] [Figure 1C] Figure 1B shows a schematic block diagram of an exemplary CPU implemented in the device architecture, which may be used to implement various aspects of this disclosure.

[0037] [Figure 1D] This is a block diagram of one or more embodiments of an immersive speech and audio services (IVAS) coder / decoder ("codec") framework for encoding and decoding IVAS bitstreams.

[0038] [Figure 1E] This diagram shows a Cartesian coordinate system centered on the listener's head.

[0039] [Figure 2] This plot shows the HRTF referenced from the left ear.

[0040] [Figure 3] This plot shows the right ear reference HRTF.

[0041] [Figure 4] This plot shows the right ear HRTF filter with the delay removed.

[0042] [Figure 5] This plot shows the phase response of two alternative filters.

[0043] [Figure 6] This plot shows the phase response of two alternative filters.

[0044] [Figure 7] This plot shows the corrected HRTF in the left ear.

[0045] [Figure 8] This plot shows the corrected HRTF in the right ear.

[0046] [Figure 9] This plot shows the left ear HRTF with added delay.

[0047] [Figure 10] This plot shows the corrected HRTF in the right ear.

[0048] [Figure 11] This diagram shows the modification of the HRTF filter.

[0049] [Figure 12] This figure shows the formation of a full-pass impulse response with associated delays.

[0050] [Figure 13] This figure shows the conversion of the original HRTF library to a more compact HRTF basis set.

[0051] [Figure 14] This figure shows a compact HRTF basis set used to efficiently compute HRTFs.

[0052] [Figure 15] This figure shows a compact HRTF basis set used for processing scene-based audio signals.

[0053] [Figure 16] Figures 13-15 show additional details of the HRTF conversion block, based on several implementations.

[0054] [Figure 17] Figure 16 shows additional details of the HRTF transformation subblock, based on several implementations.

[0055] [Figure 18] Figure 17 shows an example of a function that can be implemented using the delayed processing block. [Figure 19] Figure 17 shows an example of a function that can be implemented using the delayed processing block. [Figure 20] Figure 17 shows an example of a function that can be implemented using the delayed processing block.

[0056] [Figure 21] This flowchart outlines an example of a method that can be performed by the apparatus or system disclosed herein. [Modes for carrying out the invention]

[0057] This disclosure relates to the creation of a modified HRTF from an original HRTF such that the modified HRTF can be more efficiently approximated by linear mixing while retaining the psychoacoustic properties of the original HRTF. This specification describes techniques for processing HRTF filters to generate modified HRTF filters suitable for use in a set of filters based on linear interpolation. The following description includes numerous examples and specific details for illustrative purposes to provide a full understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure, as defined by the claims, may include some or all of the features in these examples, either individually or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.

[0058] The following descriptions detail various systems, devices, methods, processes, and procedures. Certain steps may be described in a specific order, but such order is primarily for convenience and clarity. A particular step may be repeated more than once, may occur before or after other steps (even if those steps are described in a different order), or may occur in parallel with other steps. A second step is only required to follow a first step if the first step must be completed before the second step begins. Such situations will be specifically noted if they are not evident from the context.

[0059] In this paper, the terms “and,” “or,” and “and / or” are used. Such terms should be read as having an inclusive meaning. For example, “A and B” could mean at least: “both A and B,” or “at least both A and B.” Another example is “A or B” could mean at least: “at least A,” “at least B,” “both A and B,” or “at least both A and B.” Another example is “A and / or B” could mean at least: “A and B,” or “A or B.” When an exclusive OR is intended, such an OR is specified (e.g., “either A or B,” or “at most one of A and B”).

[0060] The term “including” and its variations should be read as open-ended terms meaning “including, but not limited to, …”. The terms “one exemplary implementation” and “a certain exemplary implementation” should be read as “at least one exemplary implementation”. The term “another implementation” should be read as “at least one other implementation”. The terms “determined,” “determine,” or “determine” should be read as observing, receiving, calculating, calculating, estimating, predicting, or deriving. Furthermore, in the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as generally understood by those skilled in the art to which this disclosure belongs.

[0061] This paper describes various processing functions related to structures such as blocks, elements, components, and circuits. Generally, these structures can be realized by a processor controlled by one or more computer programs.

[0062] Throughout this disclosure and in the relevant claims and / or drawings, various acronyms may appear, which are listed below. Other commonly used acronyms and technical terms may be omitted from this list for brevity. A short list of acronyms is provided below for the reader's reference. IVAS Immersive Voice and Audio Services HRTF (Head Related Transfer Function) LPC (Linear Predictive Coding) CLDFB Complex Low Delay Filter Bank SBA Scene Based Audio SPAR (Spatial Reconstruction)... a spatial audio encoding technique. DirAC Directional Audio Coding…another spatial audio coding technique MD Metadata BS Bitstream HOA Higher Order Ambisonics FOA First Order Ambisonics MDFT (Modified Discrete Fourier Transform) MDCT (Modified Discrete Cosine Transform)

[0063] Figure 1A is a block diagram showing examples of components of a device capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and number of elements shown in Figure 1A are provided for illustrative purposes only. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 101 may be or include a device configured to perform at least some of the methods disclosed herein, such as a smart audio device, laptop computer, cellular phone, tablet device, or smart home hub. In some such implementations, device 101 may be or include a server configured to perform at least some of the methods disclosed herein.

[0064] In this example, the device 101 includes an interface system 105 and a control system 110. In some implementations, the interface system 105 may be configured to provide the control system 110 with a first set of HRTFs. In some examples, the interface system 105 may be configured to output one or more of the results of the control system 110 processing the first set of HRTFs, such as a second set of HRTFs, a set of base filters based on the second set of HRTFs, and audio data processed by one or more of the base filters (such as left ear audio data and right ear audio data).

[0065] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementations, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as an optional memory system 115 shown in Figure 1A. However, in some examples, the control system 110 may include a memory system.

[0066] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0067] In some implementations, the control system 110 may reside in two or more devices. For example, part of the control system 110 may reside in a device within the environment (such as a laptop computer, tablet computer, or smart audio device), while another part of the control system 110 may reside in a device outside the environment, such as a server. In other examples, part of the control system 110 may reside in a device within the environment, while another part of the control system 110 may reside in one or more other devices within the environment.

[0068] In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. According to some examples, the control system 110 may be configured to receive a first set of HRTFs and to convert the first set of HRTFs to a second set of HRTFs. The second set of HRTFs may be approximated more efficiently by linear mixing than the first set of HRTFs, while retaining the psychoacoustic properties of the first set of HRTFs. In some such examples, the conversion may involve replacing the delay component of the first set of HRTFs with a whole-pass filter in the second set of HRTFs. According to some such examples, the conversion may involve tuning the phase response of each whole-pass filter in the second set of HRTFs such that each interaural phase response is substantially linear for frequencies below the relevant threshold frequency of the corresponding whole-pass filter, and each phase response has a reduced interaural phase difference for frequencies above the relevant threshold frequency of the corresponding whole-pass filter.

[0069] In some examples, the control system 110 may be configured to define a set of basis filters based on a second set of HRTFs. The set of basis filters may have fewer elements than the second set of HRTFs. In this context, an "element" of the second set of HRTFs is one of the HRTFs in the second set of HRTFs. Similarly, an "element" of the set of basis filters is one of the basis filters in the set of basis filters. According to some examples, the set of basis filters may have at least one order of magnitude fewer elements than the second set of HRTFs. For example, the second set of HRTFs may have hundreds or thousands of elements in some cases, while the set of basis filters may contain fewer than 100, fewer than 50, or even fewer than 20 elements.

[0070] In some examples, the control system 110 may be configured to receive a bitstream of input audio data in an input audio format via the interface system 105. The input audio format may be, for example, an ambisonic audio format, an audio object-based audio format (such as Dolby Atmos®), or a channel-based audio format. In some examples, the control system 110 may be configured to combine the input audio data with one or more base filters from a set of base filters to generate left audio data and right audio data, such as left ear audio data and right ear audio data. In some such examples, the control system 110 may be configured to output left audio data and right audio data via the interface system 105. Outputting left audio data and right audio data may involve storing left audio data and right audio data, transmitting left audio data and right audio data, providing left audio data and right audio data to a set of loudspeakers for playback, providing left audio data and right audio data for further processing, or a combination thereof.

[0071] In some examples, the control system 110 may be configured to implement at least a portion of the codec for Immersive Voice and Audio Services (IVAS). Several examples are described herein with reference to Figure 23.

[0072] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include memory devices such as those described herein, including but not limited to random-access memory (RAM) devices and read-only memory (ROM) devices. One or more non-temporary media may reside, for example, in an arbitrary memory system 115 and / or control system 110 shown in Figure 1A. Thus, various inventive aspects of the subject matter described herein may be implemented on one or more non-temporary media on which software is stored. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system, such as the control system 110 in Figure 1A.

[0073] In some examples, the device 101 may include an optional microphone system 120, as shown in Figure 1A. The optional microphone system 120 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker in a speaker system or a smart audio device.

[0074] In some implementations, the device 101 may include an arbitrary loudspeaker system 125 shown in Figure 1A. The arbitrary loudspeaker system 125 may include one or more loudspeakers. Loudspeakers are sometimes referred to as "speakers" in this paper. In some examples, at least some of the loudspeakers of the arbitrary loudspeaker system 125 may be positioned in any way. For example, at least some of the speakers of the arbitrary loudspeaker system 125 may be positioned in a location that does not correspond to any standard defined speaker layout such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the loudspeakers of the arbitrary loudspeaker system 125 may be positioned in a space-convenient location (e.g., a location where there is space to accommodate the loudspeakers), but not in any standard defined loudspeaker layout.

[0075] In some implementations, the device 101 may include an optional sensor system 130, as shown in Figure 1A. The optional sensor system 130 may include a touch sensor system, a gesture sensor system, one or more cameras, and so on.

[0076] In some implementations, the device 101 may include an optional display system 135 as shown in Figure 1A. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some examples, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples where the device 101 includes the display system 135, the sensor system 130 may include a touch sensor system and / or a gesture sensor system adjacent to one or more displays of the display system 135. According to some such implementations, the control system 110 may be configured to control the display system 135 to present a graphical user interface (GUI), such as a GUI, relating to implementing one of the methods disclosed herein.

[0077] Figure 1B shows a schematic block diagram of an exemplary device architecture 101 (in this example, Apparatus 101) that may be used to implement various aspects of the present disclosure. Apparatus 101 in Figure 1B is an example of Apparatus 101 in Figure 1A. Architecture 101 includes, but is not limited to, server and client devices, systems, etc., which may be configured to perform the methods described with reference to any or all of Figures 11-17 and Figure 21. As shown, architecture 101 includes a central processing unit (CPU) 141 that can perform various operations according to, for example, a program stored in read-only memory (ROM) 142, or a program loaded from a storage unit 148 into random access memory (RAM) 143. The CPU 141 may be, for example, an electronic processor 141. In these examples, the CPU 141 is an instance of the control system 110 in Figure 1A, and the ROM 142 and RAM 143 are instances of the memory system 115. RAM 143 also stores data as needed when the CPU 141 performs various operations. The CPU 141, ROM 142, and RAM 143 are interconnected via bus 144. The input / output (I / O) interface 145 is also connected to bus 144. Bus 144 and I / O interface 145 are instances of interface system 105 in Figure 1A.

[0078] The following components are connected to the I / O interface 145: an input unit 146 which may include a keyboard, mouse, etc.; an output unit 147 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 148 which includes a hard disk or another suitable storage device; and a communication unit 149 which includes a network interface card such as a network card (e.g., wired or wireless).

[0079] In some implementations, the input unit 146 includes one or more microphones located at different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other appropriate formats).

[0080] In some implementations, the output unit 147 includes a system with varying numbers of speakers. The output unit 147 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other appropriate formats) (depending on the capabilities of the host device).

[0081] In some embodiments, the communication unit 149 is configured to communicate with other devices (for example, via a network). The drive 150 is also connected to the I / O interface 145, if necessary. A removable medium 151, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is mounted on the drive 150, thereby installing computer programs read therefrom into the storage unit 148, if necessary. Those skilled in the art will understand that although the apparatus 101 is described as including the above-described components, in actual applications it is possible to add, remove, and / or replace some of these components, and all such modifications or changes will fall within the scope of this disclosure.

[0082] According to exemplary embodiments of the present disclosure, the processes described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded and mounted from a network via a communication unit 149 and / or installed from removable media 151, as shown in Figure 1B.

[0083] Figure 1C shows a schematic block diagram of an exemplary CPU 141 implemented in the device architecture 101 of Figure 1B, which may be used to implement various aspects of the present disclosure. The CPU 141 includes an electronic processor 160 and a memory 161. The electronic processor 160 is electrically and / or communicatively connected to the memory 161 for bidirectional communication. The memory 161 stores encoding software 162 and decoding software 163. The memory 161 may be, for example, ROM, RAM, or another non-temporary computer-readable medium. The electronic processor 160 can implement the encoding software 162 stored in the memory 161 to perform, among other things, method 2100 of Figure 21. Furthermore, the electronic processor 160 can implement the decoding software 163 stored in the memory 161 to perform, among other things, methods described with reference to any or all of Figures 11-17 and Figure 21.

[0084] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuitry (e.g., control circuits), software, logic, or any combination thereof. For example, the units described above may be executed by a control circuit (e.g., CPU 141 in combination with other components of Figure 1B), and thus the control circuit may perform the actions described in the present disclosure. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device (e.g., control circuits). Various aspects of the exemplary embodiments of the present disclosure are illustrated and described using block diagrams, flowcharts, or any other pictorial representation, but it will be understood that the blocks, apparatus, systems, techniques, or methods described herein may, in non-limiting examples, be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.

[0085] Furthermore, the various blocks shown in the flowchart can be viewed as method steps and / or as operations resulting from the operation of computer program code and / or as a group of coupled logic circuit elements constructed to perform related functions. For example, embodiments of the present disclosure include a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.

[0086] In the context of this disclosure, a machine-readable medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-temporary and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0087] Computer program code for performing the methods of this disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, a dedicated computer, or another programmable data processing device having a control circuit, so that when executed by the computer or other programmable data processing device processor, the program code performs the functions / operations specified in the flowcharts and / or block diagrams. The program code may run entirely on a computer, partially on a computer, as a standalone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0088] An example of an IVAS codec framework Figure 1D is a block diagram of an Immersive Speech and Audio Services (IVAS) coder / decoder ("codec") framework 170 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including, but not limited to, mono-to-stereo upmixing, as well as fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including, but not limited to, mobile and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices.

[0089] In this example, the IVAS codec 170 includes an IVAS encoder 171 and an IVAS decoder 174. In some examples, the IVAS encoder 171, the IVAS decoder 174, or both may be implemented by one or more instances of the control system 110 in Figure 1A, by the CPU 141 in Figures 1B and 1C, etc. In some examples, the IVAS encoder 171 may be implemented by the encoding software 162 in Figure 1C, and the IVAS decoder 174 may be implemented by the decoding software 163 in Figure 1C. According to some examples, a control system implementing the IVAS encoder 171, the IVAS decoder 174, or both may also be configured to perform some or all of the operations disclosed herein, such as those described with reference to one or more of Figures 11-17 and Figure 21.

[0090] In this example, the IVAS encoder 171 includes a spatial encoder 172 that receives N channels of input spatial audio (e.g., FOA, HOA). In some implementations, the spatial encoder 172 may be configured to implement Spatial Reconstruction (SPAR), Directional Audio Coding (DirAC), another spatial audio coding technique, or a combination thereof. In this example, the output of the spatial encoder 172 includes a spatial metadata (MD) bitstream (BS) and a spatial downmix of N_dmx channels. In this example, the spatial MD is quantized and entropy coded. In some implementations, the quantization can include fine, medium, coarse, and very coarse quantization strategies, and the entropy coding can include Huffman or arithmetic coding. In some implementations, the framework may allow quantization of three levels or less in a given mode of operation, but as the bitrate decreases, in some such implementations, the three levels become increasingly coarse as a whole to meet the bitrate requirements. In this example, the core audio encoder 173—which may, for example, be based on a monaural Enhanced Voice Services (EVS) encoding unit—is configured to encode N_dmx channels (N_dmx = 1 to 16 channels) of the spatial downmix into an audio bitstream, which is then combined with the spatial MD bitstream to form an IVAS encoded bitstream, which is then sent to the IVAS decoder 174.

[0091] In this example, the IVAS decoder 174 includes a core audio decoder 175 (e.g., an EVS decoder) that decodes the audio bitstream extracted from the IVAS bitstream to reconstruct N_dmx audio channels. According to this example, the spatial decoder / renderer 176 (e.g., SPAR / DirAC) decodes the spatial MD bitstream extracted from the IVAS bitstream to reconstruct the spatial MD and synthesizes / renders the output audio channels using the spatial MD and spatial upmix for playback on various audio systems with different speaker configurations and capabilities.

[0092] Figure 1E shows an example of a coordinate system based on the listener's head. A head-related transfer function (HRTF) filter can be used to process an audio signal to generate a binaural audio signal that provides the listener with the sensation of sound coming from a given direction of arrival. The direction of arrival can be defined using the (x,y,z) unit vector, where Cartesian coordinates can be defined as shown in Figure 1E. In the example shown in Figure 1E, the coordinate system is positioned such that its origin is approximately at the center of the listener's head 200, the X-axis 801 points forward (towards the listener's nose), the Y-axis 802 points to the listener's left, and the Z-axis 803 points upward through the top of the listener's head.

[0093] An audio signal s(t) can be processed using an HRTF filter to give the listener the sensation that the sound (of signal s(t)) is coming from a direction of arrival defined by a unit vector (x,y,z). This process involves applying a pair of HRTF filters (h) to the input audio signal. l (t) and h r By convolving with each of (t), we obtain the two ear signals e l (t) and e r Generate (t):

number

[0094] HRTF filter (h l (t) and h r (t)) is,

number

[0095] H(x,y,z) is referred to herein as the HRTF set function. This function is suitable for computing an HRTF filter for a set of (x,y,z) direction vectors. The set of (x,y,z) vectors for which the HRTF set function generates a valid HRTF filter is referred herein as the domain of the HRTF set function.

[0096] In the following description, a time-domain impulse response is used to represent the filter response. Those skilled in the art will understand that equivalent storage and manipulation of the filter response can be performed in other domains, including but not limited to the frequency domain.

[0097] The HRTF set function may be used to create an HRTF discrete library. The HRTF discrete library defines the left and right ear HRTF responses for N (x,y,z) unit vectors:

number

[0098] And when the HRTF set function is evaluated by Equation 4, the HRTF discrete library can be written as follows:

number

[0099] It is desirable to provide a means for defining an HRTF set function such that each output HRTF filter generated by the HRTF set function is formed from a linear combination of base filters. The linear HRTF set function may be defined according to Equation 6, where e l(t) and e r (t) The filter is calculated as follows.

Equation

[0100] According to Equation 6, the set of K left-ear basis filters b k l (t) and the set of K right-ear basis filters b k r (t) are linearly combined using weights defined by the gain functions g k l (x, y, z) and g k r (x, y, z).

[0101] In an alternative embodiment, a symmetric HRTF set function (where the left-ear HRTF filter for direction (x, y, z) is the same as the right-ear HRTF for direction (x, -y, z)) can be defined using a smaller set of basis filters and gain functions.

Equation

[0102] Without loss of generality, the first line of Equation 7 may be examined. It is understood that the following explanation also applies to the second line of Equation 7 and / or Equation 6.

[0103] For a set of N arrival directions ((x n , y n , z n ), n = 1…N), the first line of Equation 7 can be rewritten in matrix form (omitting the subscript l from e l (t) for simplicity):

Equation

[0104] Equation 8 can be rewritten in a simpler form as follows.

Equation

[0105] In Equation 9, the column vector E(t) is a set of N unit vectors ((x n ,y n ,z n A set of N left-ear HRTF filter responses is defined for n=1…N), and the column vector B(t) defines a set of K filter responses. In some embodiments, the goal is that the resulting HRTF filter E(t) is the same as the original set of HRTF filter responses E orig The goal is to determine these filter responses B(t) so that they are approximations of (t).

[0106] Various methods are known for determining an appropriate filter B(t). One example is as follows:

number

[0107] Other methods can also be used, and the goal of each method is the difference E(t)-E orig It can be understood that this can be achieved by minimizing the absolute value of (t).

[0108] A reasonable approximation (E(t)~E orig To provide (t), a very large number (K) of base filters may be required. The difficulty associated with using a linear mixing process (by equations 6, 7, or 8) is that the high-frequency components of the HRTF filter can generally be very difficult to define with respect to linear mixing.

[0109] In some embodiments, the original set of HRTF filters E orig (t) is modified, and the set of modified HRTF filters E mod(t) is generated. Here, the modified HRTF filter differs from the original filter in its phase response at high frequencies. For each of the N directions, the frequency responses of the original HRTF and the modified HRTF can be defined using the Fourier transform.

number

[0110] Frequency response function R orig,n (f) and R mod,n (f) is a complex value, and therefore we can say the following:

number

[0111] Figures 2, 3, and 4 show examples of impulse responses of HRTF filters. Figure 2 shows the impulse response 111 of the left ear HRTF filter for the direction of arrival: (x,y,z)=(1 / √2,1 / √2,0) (left-forward direction). Similarly, Figure 3 shows the impulse response 211 of the right ear HRTF for the same direction of arrival. Examining Figure 3, we can see that the impulse response 211 includes a delay of 0.4 ms.

[0112] Figure 4 shows the (delay-free) impulse response 311 created by removing a 0.4 ms delay from the impulse response 211 in Figure 3. Figure 5 shows an example graph of the phase response with respect to frequency. The 0.4 ms delay in Figure 3 may be defined as the linear phase response plot 411 in Figure 5. In addition, an alternative phase response 412 is plotted in Figure 5, and the alternative phase curve 412 closely matches the linear phase response 411 for frequencies from 0 to 1400 Hz.

[0113] In some embodiments, the original impulse response 211 in Figure 3 may be modified by removing a bulk delay of 0.4 ms to produce the delay-free impulse response 311 in Figure 4, and the phase response 412 in Figure 5 may be applied to the impulse response 311 to produce a new filtered impulse response with the correct phase response for frequencies below 1400 Hz. Unfortunately, in order to implement this filter in a real-time audio process, an additional 3 ms delay may be added to generate the impulse response, which may result in a new impulse response that is not causally related. See, for example, the impulse response 911 shown in Figure 10. To maintain compatibility with the right ear response, this exemplary impulse response 111 (left ear response) also requires the addition of a 3 ms delay, resulting in the impulse response 811 in Figure 9.

[0114] Some disclosed examples involve modifying the original HRTF filter for both the left and right ears to provide a biaural phase difference similar to that shown in phase response 412 in Figure 5, without the side effect of undesirable delays (e.g., the 3ms delay mentioned above with reference to Figures 9 and 10).

[0115] Figure 6 shows an example of a causal pass-through filter. Figures 7 and 8 show examples of modified HRTFs that can be generated by a causal pass-through filter. In some embodiments, the causal pass-through filters 511 and 512 in Figure 6 can be applied to the original left ear impulse response 111 and right ear impulse response 311, respectively, to generate the modified left ear HRTF 611 in Figure 7 and the modified right ear HRTF 711 in Figure 8, respectively.

[0116] Figure 11 shows an example of an HRTF processing block. In some examples, the block in Figure 11 can be implemented by the control system 110 in Figure 1A, for example, according to instructions stored on a computer-readable medium. In Figure 11, configuration 100 shows the original HRTF impulse response 211 h(t), which is received and processed by the HRTF processing block 151 to determine the bulk delay 140 d, which is the delay inherent in the impulse response 211. In this example, the HRTF processing block 151 also generates a delay-free HRTF 311 h'(t) such that h'(t) = h(t+d).

[0117] In the example shown in Figure 11, the pass-through generator 152 generates a pass-through filter impulse response 512 α(t) in response to a delay of 140 d, and the convolution process 153 combines the undelayed impulse response 311 and the pass-through response 512 to generate a modified HRTF 711 m(t).

[0118] The pass filter 512 α(t) can be defined as the following function:

number

number

[0119] We define Φ0(t) = arg(F{C(0,t)}(f)). This is the whole-pass phase response generated by the whole-pass generator 152 when the delay d = 0. This can be called zero-delay whole-pass. In some embodiments, it may be required that the whole-pass phase response satisfies the following equation.

number

[0120] The left-hand side of Equation 15 represents the phase difference between the zero-delay pass filter and the pass filter defined for delay d. This phase difference is equivalent to the phase response 412 in Figure 5. The right-hand side of Equation 15 represents the expected linear phase ramp for delay d. This is equivalent to the linear phase response 411 in Figure 5.

[0121] Therefore, in this example, equation 15 shows that the full-pass phase response 412 is F p This expresses the requirement that the linear ramp phase response 411 should match the frequencies up to . Furthermore, equation 15 gives an upper limit d for the range of delay values ​​over which the whole-pass generator function C(d,t) is expected to produce a valid result. max Define d max A typical value for this is d max =0.7ms (milliseconds), but in some applications, d max This could be any other value, such as a value between 0.6ms and 0.8ms, a value between 0.5ms and 0.8ms, a value between 0.6ms and 0.9ms, or a value between 0.5ms and 1.0ms.

[0122] In some embodiments, 0 to d max A finite set of M delay values ​​(d1, d2, ..., d) spanning the range M ) may be selected, and an appropriate full-pass response (α1(t), α2, ... (t), α M (t) may be pre-calculated according to an optimization process. In this case, the total pass generator function C(d,t) can be realized by a lookup table or interpolation function by utilizing M pre-stored total pass responses.

[0123] In some further embodiments, each of the total pass responses (e.g., the mth total pass filter α) m (t)) may be defined as an infinite impulse response (IIR) filter having T conjugate pole pairs and their corresponding conjugate zero pairs. 1,m ,p 2,m ,…,p T,m The basic set may be selected, in which case filter α m (t)

number

[0124] Therefore, all M total pass-through responses (α1(t), α2, ... (t), α M The basic set of (t) is defined as an IIR filter of order 2T, and the set of M total-pass responses is fully defined using the following [T × M] complex fundamental poles:

number

number

[0125] Given a matrix P of fundamental poles representing a set of M total-pass filters (each total-pass filter has T complex fundamental poles), and corresponding sets of delay values ​​(d1, d2, ..., d M Given ), a polynomial approximation may be formed such that the fundamental poles can be defined as polynomial functions of d. This polynomial approximation process can be carried out according to known methods, including but not limited to MATLAB's POLYFIT function.

[0126] In some embodiments, the number of complex fundamental poles is T=3, and a polynomial of degree 4 may be used to compute the fundamental extrema as a function of delay d. According to this embodiment, the process for computed the total-pass filter α(t) is performed by the following sequence of operations. 1. Given a delay d, calculate the fundamental poles (p1, p2, p3). 2.

number

[0127] The three steps above demonstrate the use of polynomial functions as a convenient method for calculating the poles of a filter. Of course, polynomials only provide an approximation to the "best" poles, and it is known that small errors in pole positions can lead to large changes in the resulting filter response. Several alternative methods involve using a nonlinear function (e.g., Equation 18) and equating it to a polynomial value (J l and J l+1 This includes applying it to ). This nonlinear function is a polynomial value (Jl and J l+1 The function can be defined such that a small error in the s-domain no longer results in a large error in the pole location. Furthermore, some such examples involve calculating pole locations in the "s-domain" and then mapping them to the "z-domain". Other choices of nonlinear mapping functions may be used, such nonlinear mapping functions may generate poles in the z-domain, s-domain, or other domains.

[0128] Figure 12 shows additional transformation processes that may be implemented by the pass-through generator 152 in Figure 11 in several embodiments. In the example shown in Figure 12, the pass-through generator 152 receives a selected delay d 140, which is processed by the delay processing block 172 to generate a set of intermediate values ​​180. In this example, the intermediate values ​​180 are then mapped by additional nonlinear processing to form a set of filter poles 182. The filter poles 182 are then processed by the pass-through calculation block 175 to form the pass-through impulse response α(t) 512 in Figures 11 and 12.

[0129] According to some examples corresponding to the block shown in Figure 12, the delayed processing block 172 applies the polynomial function described above to output a set of intermediate values ​​180. In some examples, this set of intermediate values ​​is a number J. l It may also be a set of intermediate values. In some such examples, the mapping block 173 applies a nonlinear mapping process to a set of intermediate values ​​180 by, for example, performing equation 18 to produce output 181, which in one example is the pole position in the s-domain. In some examples, the bilinear transform block 174 transforms the pole position in the s-domain to produce output 182, which in one example includes the pole position in the z-domain. In the example shown in Figure 12, the full-pass calculation block 175 calculates output 512, which in this example is alpha(t) (impulse response). In alternative examples, output 512 may be the output of another method that can be used to define the phase response, frequency response, or full-pass filter response.

[0130] When generating filter poles using a simple function such as a polynomial, as those skilled in the art will understand, when the poles are very close to the unit circle, small inaccuracies in the polynomial output can lead to large errors in the final pass-through response. A nonlinear processing, such as that applied in Figure 12 to convert the intermediate value 180 to the filter pole 182, may allow the processing 172 to be implemented more efficiently.

[0131] In some embodiments, the process 172 is implemented as a set of L polynomial functions that generate L intermediate values ​​180. For example, J l =Poly l (d) (l∈1…L). The intermediate value 180 can be used by the filter pole generation block 173 in this example to generate the s-plane filter pole 181. For example, the two intermediate values ​​are J l and J l+1 Using this, complex s-plane pole P n It is possible to define this.

number

[0132] Alternatively, a single intermediate value, for example, J l However, P n = -2πJ l According to this, a single real s-plane pole P n It may be used by the filter pole generation block 173 to define the

[0133] The s-plane pole 181 may then be transformed into the z-plane pole 182 (by transformation block 174 in this example). For example, MATLAB's BILINEAR function can perform this transformation: P(N) = bilinear(P(N), 1, 1, 48000) Or this conversion: P(N)=bilinear(P(N),1,1,48000,F P ) It may be used by conversion block 174 to apply the following, where Fp represents the upper frequency limit (as used in Equation 15) and 48000 is the sample rate according to this embodiment. It will be understood that alternative sample rates, including but not limited to 16000, 32000, 44100, or 96000, may be used.

[0134] Those skilled in the art will understand that other nonlinear processing methods can be used to facilitate the mapping of the selected delay d 140 to the set of full-range passing poles 182. In an alternative embodiment, the polynomial function applied by the delay processing block 172 may be used to define the frequency and the Q of the pole, and the nonlinear mapping process applied by the mapping block 173 may convert the frequency and Q value to the respective pole positions. In another embodiment, the nonlinear mapping process applied by the mapping block 173 may determine the z-domain pole positions, eliminating the need for the bilinear transformation of the transformation block 174.

[0135] Furthermore, by forming additional conjugate poles (for each of the complex poles in set 182) and by forming each filter zero as the reciprocal of each corresponding pole, the pass-through filter response may also be derived by the pass-through filter response block 175 in this example, and it will be understood that this pass-through filter is causal.

[0136] Figure 13 shows the process of generating a set of base filters from a set of HRTFs. In some examples, the blocks in Figure 13 may be implemented, at least partially, by the control system 110 in Figure 1A. Figure 13 shows a configuration 500 in which the original HRTF library 520 is processed by the HRTF conversion block 521 in this example to produce a modified HRTF library 541. In this example, the interaural delay component inherent in the HRTF filters of the original HRTF set is replaced by a pass-through filter satisfying equation 15, and the modified HRTF library is F pIt has reduced interaural phase at frequencies greater than . The modified HRTF 521 is processed by the base filter generation block 522 in this example and generates a set of base filters 523 according to a fitting process such as the fitting process of Equation 10.

[0137] Base filter set 523 has fewer elements than the modified HRTF set. In this context, an "element" of base filter set 523 is one of the base filters of base filter set 523, and an "element" of the modified HRTF set is one of the HRTFs of the modified HRTF set. In some examples, base filter set 523 can have at least an order of magnitude fewer elements than the modified HRTF set. For example, the modified HRTF set may have hundreds or thousands of elements in some cases, while base filter set 523 may contain fewer than 100, fewer than 50, or even fewer than 20 elements. Thus, base filter set 523 forms a compact representation of the original HRTF set 520.

[0138] Figure 14 shows the process of generating a set of basis filters from a set of HRTFs and the process of using the set of basis filters to form left and right HRTF filters. In some examples, the blocks in Figure 14 may be implemented, at least in part, by the control system 110 in Figure 1A. Figure 14 shows configuration 501 in which the original HRTF library 520 is processed by the HRTF conversion block 521 to produce a modified HRTF library 541, which is then processed by the basis filter generation block 522 to produce a set of basis filters 523. The direction of arrival 524 (which may be defined according to spherical coordinates (θ,φ), unit vectors (x,y,z), or by other forms known in the art) is processed by the weight coefficient generation block 525 to form weight coefficients 526. In some embodiments, the weight coefficients may be defined according to the spherical harmonic panning equations, and the basis filters may likewise be adapted to be compatible with the spherical harmonic panning equations. For example, g in equation 7 k (x, y, z).

[0139] In this example, the weight coefficient and base filter combination block 527 combines the weight coefficient 526 with the base filter 523 to form left-ear HRTF filters and right-ear HRTF filters (528 and 529, respectively) that represent modified HRTFs for a given direction of arrival. The weight coefficient and base filter combination block 527 can be implemented, for example, according to Equation 7 when the base filter represents a symmetric HRTF set. The weight coefficient and base filter combination block 527 can be implemented, for example, according to Equation 6 when the base filter represents an HRTF set that includes asymmetry.

[0140] Figure 15 illustrates the process of generating a set of base filters from a set of HRTFs and using the set of base filters to form left and right audio signals. In some examples, the blocks in Figure 15 may be implemented, at least partially, by the control system 110 in Figure 1A. Figure 15 shows configuration 502 in which the original HRTF library 520 is processed by the HRTF conversion block 521 to produce a modified HRTF library 541, which is then processed by the base filter generation block 522 to produce a set of base filters 523. According to this example, the audio generation block 530 generates an audio signal 531 in a form associated with a scene-based audio format such as ambisonics or higher-order ambisonics. The audio generation block 530 may also be, or include, an audio decoder adapted to generate a multi-channel audio bitstream from a transmitted or stored encoded bitstream. Alternatively, the audio generation block 530 may be, or include, an audio capture and / or processing device adapted to generate scene-based audio signals 531 representing a spatial audio scene.

[0141] In this example, the audio input and base filter coupling block 532 is adapted to combine the audio signal 531 with the base filter 523 to generate left and right ear audio signals (533 and 534, respectively). In some examples, the audio input and base filter coupling block 532 may also be configured to implement a convolution process, which may be implemented according to known time-domain or frequency-domain methods, as known in the art.

[0142] Figure 16 shows additional details of the HRTF conversion blocks in Figures 13-15, as in several implementations. In some examples, the blocks in Figure 16 may be implemented, at least partially, by the control system 110 in Figure 1A. Figure 16 shows a more detailed diagram of the process at the top of Figures 13-15 (conversion from the "original" HRTF library 520 to the "modified" HRTF library 521). In this example, each left / right HRTF pair is processed by the corresponding HRTF conversion subblock 150.

[0143] Figure 17 shows details of additional HRTF conversion subblocks from Figure 16, as in several implementations. In some examples, the blocks in Figure 17 can be implemented, at least partially, by the control system 110 in Figure 1A. Figure 17 shows an example of HRTF conversion subblock 150, where the L and R HRTFs (211L and 211R) are processed by the left HRTF processing block 151L and the right HRTF processing block 151R, respectively, to extract the undelayed impulse response 311L / R and the delayed 140L / R. The two delayed 140L / R are then processed by the delay processing block 138 to generate a new simplified delayed 141L / R. The simplified delayed is then processed by the pass-through filter generation blocks 152L and 152R, respectively, to form a pass-through filter 512L / R. In this example, the modified left HRTF generation blocks 153L and 153R are configured to combine the undelayed impulse response 311L / R with the pass-through filter 512L / R to form the modified HRTF pair 711L / R.

[0144] Exemplary delay definitions for each ear The delay processing block 138 in Figure 17 is configured to generate a new delay value 141L / R in response to the difference between delays 140L / R. Figures 18, 19, and 20 show examples of functions that may be implemented by the delay processing block 138 in Figure 17. For Figures 18, 19, and 20, the corresponding functions are as follows:

[0145]

number

[0146] In Equation 21, d max is, |d L -d R | represents the maximum expected value.

[0147] An important characteristic of the function implemented by the delayed processing block 138 is d' L -d' R =d L -d R d' that satisfies L and d' R The goal is to generate a HRTF, which preserves the interauricular delay difference between the left (L) HRTF and the right (R) HRTF.

[0148] The function in Figure 20 is the same function used in the exemplary MATLAB code shown below, and is defined to have the following properties: (a) the delay for both ears is a smooth function, and (b) ears with lower delays (which are typically also ears with larger amplitudes) have less delay variation (because the slope of the curve is lower when the delay is lower).

[0149] In some further embodiments, the left HRTF processing block 151L and the right HRTF processing block 151R may be adapted to produce non-delayed impulse responses 311L and 311R, respectively, which are minimum phase filter responses.

[0150] According to some embodiments, the left HRTF processing block 151L and the right HRTF processing block 151R may be implemented according to the following steps. 1. The frequency response of the original HRTF filter, for example, R defined by Equation 11. orig,n Determine (f). 2. Absolute response of the original HRTF filter M(f) = |R orig,n (f) To decide. 3. Determine the frequency response of the new (undelayed) minimum phase filter according to a method known in the art using the Hilbert transform.

number

number

number

[0151] In some examples, the delay measurement frequency f d The value is selected to be within the range of 300Hz to 1600Hz. For a detailed example, f d = 1200Hz. In some other examples, f d This can be a value within the range of 300Hz to 1200Hz, a value within the range of 600Hz to 1800Hz, a value within the range of 1000Hz to 1400Hz, or a value within another frequency range.

[0152] In some embodiments, the delay d and minimum phase response R'(f), as determined according to the above methods, are used to modify the HRTF response R according to the following steps. mod,n Used to determine (f). 1. For each of the original left-ear HRTF filters (with respect to a given direction of arrival), determine the delay d and minimum phase response R'(f) using the procedure described above, and label them as follows: left ear delay=dL Non-delayed filter for the left ear = R' L (f) Right ear delay = d R Non-delayed filter for the right ear = R' R (f) 2. Maximum interaural delay difference

Number

Number

[0153] Exemplary MATLAB implementation The following MATLAB implementation provides additional details according to the disclosed methods. An example of a MATLAB function that determines the HRTF basis filter set from an existing HRTF library is shown below. [Table 1-1] [Table 1-2] The get_allpass_IRs() function can be defined according to the following MATLAB code. [Table 2]

[0154] Figure 21 is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of Method 2100, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 2100 may be performed concurrently. Furthermore, some implementations of Method 2100 may include more or fewer blocks than those illustrated and / or described. The blocks of Method 2100 may be performed by one or more devices, which may be (or include) a control system, such as the control system 110 shown in Figure 1A and described above.

[0155] In this example, method 2100 is an audio processing method. According to this example, block 2105 involves the control system obtaining a first set of HRTFs. The first set of HRTFs could be, for example, the original HRTF library 520 shown in Figures 13-16.

[0156] In this example, block 2110 is involved in the control system converting a first set of HRTFs to a second set of HRTFs. The second set of HRTFs may be, for example, the modified HRTF library 541 shown in Figures 13-16. According to this example, the conversion process of block 2110 is involved in replacing the delay component of the first set of HRTFs with the full-pass filters in the second set of HRTFs. In this example, the conversion process of block 2110 is also involved in adjusting the phase response of each full-pass filter in the second set of HRTFs so that each interaural phase response is substantially linear at frequencies below the relevant threshold frequency of its corresponding full-pass filter, and each phase response has a reduced interaural phase difference at frequencies above the relevant threshold frequency of its corresponding full-pass filter. The threshold frequency may be, for example, the frequency at which the alternative phase curve 412 deviates from the linear phase response 411 in Figure 5. The threshold frequency may be, for example, a frequency in the range of 1300Hz to 1500Hz, a frequency in the range of 1000Hz to 1600Hz, or a frequency in the range of 1200Hz to 1600Hz. In some examples, the threshold frequency may be 1400Hz.

[0157] In this example, block 2115 is responsible for outputting the results of adjusting the phase response of each of the pass-through filters in the second set of HRTFs. Outputting the results could involve, for example, storing the results, transmitting the results, providing the results for further processing, or a combination of these.

[0158] In some examples, Method 2100 may involve additional processes, such as those described herein with reference to Figure 13. In some such examples, Method 2100 may also involve a control system defining a set of basis filters based on a second set of HRTFs. The set of basis filters may have fewer elements than the second set of HRTFs. In some examples, the set of basis filters may have at least one order of magnitude fewer elements than the second set of HRTFs.

[0159] In some examples, Method 2100 may also involve processes such as those described herein with reference to Figure 14 or Figure 15. In some such examples, Method 2100 may also involve a control system acquiring a bitstream of input audio data in an input audio format, and the control system combining the input audio data with one or more base filters from a set of base filters to generate left audio data and right audio data. According to some such examples, Method 2100 may also involve a control system outputting the left audio data and right audio data. Outputting the left audio data and right audio data may involve storing the left audio data and right audio data, transmitting the left audio data and right audio data, the control system providing the left audio data and right audio data to a set of loudspeakers for playback, providing the left audio data and right audio data for further processing—for example, to other modules implemented by the control system, to another control system, or a combination thereof.

[0160] In some examples, the transformation process of block 2110 may involve processes such as those described herein with reference to Figures 16 and 17. In some such examples, block 2110 may involve obtaining left-ear HRTFs and right-ear HRTFs from a first set of HRTFs, identifying the left-ear non-delayed impulse response and left-ear delay from each of the left-ear HRTFs, and identifying the right-ear non-delayed impulse response and right-ear delay from each of the right-ear HRTFs. In some such examples, block 2110 may involve generating left-ear full-pass filters, each of which is at least partially based on an instance of the left-ear delay, and generating right-ear full-pass filters, each of which is at least partially based on an instance of the right-ear delay. In some such examples, block 2110 may involve combining instances of the left-ear and right-ear non-delayed impulse responses with corresponding instances of the left-ear and right-ear full-pass filters to generate HRTF pairs of a second set of HRTFs.

[0161] In some such examples, block 2110 may be involved in generating modified left-ear delay values ​​and right-ear delay values ​​based on one or more of the extracted left-ear delays and right-ear delays. The left-ear full-pass filter and right-ear full-pass filter may be based on the modified left-ear delay values ​​and right-ear delay values, respectively. In some examples, generating instances of the modified left-ear delay values ​​and right-ear delay values ​​may be involved in determining the difference between the extracted left-ear delay and the extracted right-ear delay. According to some examples, generating instances of the modified left-ear delay values ​​and right-ear delay values ​​may be involved in determining the maximum expected difference between the extracted left-ear delay and the extracted right-ear delay.

[0162] In some examples, the difference between the extracted left ear delay and the extracted right ear delay may be equal to the difference between the corresponding modified left ear delay value and the modified right ear delay value. In some examples, the modified left ear delay value and the modified right ear delay value may correspond to a smooth function, such as that shown in Figure 20. In some examples, each pair of modified left ear delay value and modified right ear delay value may contain a lower delay value and a higher delay value. In some such examples, the lower delay value may have less delay variation than the higher delay value.

[0163] In some cases, the non-delayed impulse response can be the minimum phase filter response.

[0164] In some examples, extracting each left-ear non-delayed impulse response, each right-ear non-delayed impulse response, each left-ear delay, and each right-ear delay from the left-ear and right-ear HRTFs, respectively, may involve determining the frequency response of the original HRTF filter in the first set of HRTFs, determining the absolute value response of the original HRTF filter, determining the minimum phase frequency response of the new non-delayed minimum phase filter, determining the phase response of the original HRTF filter and the phase response of the new non-delayed minimum phase filter, and determining the delay associated with the original HRTF filter, at least partially based on the phase response of the original HRTF filter and the phase response of the new non-delayed minimum phase filter. In some such examples, determining the minimum phase frequency response may involve performing a Hilbert transform relating to the absolute value response of the original HRTF filter. In some examples, determining the delay associated with the original HRTF filter may also be at least partially based on a delay measurement frequency in the range of 300 Hz to 1600 Hz.

[0165] In some examples, the pass-through phase response may deviate from the linear ramp phase response and may smoothly approach zero phase at frequencies above the threshold frequency. The alternative phase curve 412 in Figure 5 provides one such example.

[0166] According to some examples, a control system configured to implement Method 2100 is also configured to implement at least a portion of the codecs for Immersive Voice and Audio Services (IVAS). Figure 1D shows one such example.

[0167] The above description illustrates various embodiments of the Disclosure, along with examples of how aspects of the Disclosure may be implemented. The above examples and embodiments should not be considered sole embodiments, but are presented to illustrate the flexibility and advantages of the Disclosure as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be adopted without departing from the spirit and scope of the Disclosure as defined by the claims.

Claims

1. An audio processing method for a control system including one or more processors, the method being: The control system includes the steps of: obtaining a first set of head-related transfer functions (HRTFs); The control system performs the step of converting the first set of HRTFs to a second set of HRTFs, wherein the conversion involves: Replace the delay component of the first set of HRTFs with the pass-through filter in the second set of HRTFs; Adjust the phase response of each pass-through filter in the second set of HRTFs, Each interaural phase response is substantially linear at frequencies below the corresponding threshold frequency of the full-pass filter. This includes ensuring that each phase response has a reduced interaural phase difference at frequencies above the relevant threshold frequency of the corresponding full-pass filter. Stages; The step of outputting the second set of HRTF and Methods that include...

2. The method according to claim 1, wherein outputting the second set of HRTFs relates to storing the second set of HRTFs, transmitting the second set of HRTFs to a device configured to process audio data, providing the second set of HRTFs for further processing, or a combination thereof.

3. The control system comprises the steps of defining a set of base filters based on the second set of HRTFs, wherein the set of base filters has fewer elements than the second set of HRTFs; The control system includes the steps of: acquiring a bitstream of input audio data in the input audio format; The control system comprises the steps of: generating left audio data and right audio data by combining the input audio data with one or more base filters from the set of base filters; The control system performs the steps of outputting the left audio data and the right audio data. The method according to claim 1, further comprising:

4. The method according to claim 3, wherein outputting the left audio data and the right audio data relates to storing the left audio data and the right audio data, transmitting the left audio data and the right audio data, providing the left audio data and the right audio data to a set of loudspeakers for playback by the control system, providing the left audio data and the right audio data for further processing, or a combination thereof.

5. The aforementioned conversion further: The steps include obtaining left ear HRTF and right ear HRTF from the aforementioned first set of HRTFs; The steps include identifying the left ear non-delayed impulse response and the left ear delay from each of the aforementioned left ear HRTFs; The steps include identifying the right ear non-delayed impulse response and the right ear delayed response from each of the aforementioned right ear HRTFs; A step of generating a left ear full-pass filter, wherein each of the left ear full-pass filters is at least partially based on an instance of the left ear delay; A step of generating a right ear full-pass filter, wherein each of the right ear full-pass filters is at least partially based on an instance of the right ear delay; The steps include: generating a second set of HRTF pairs of HRTFs by combining instances of the left ear non-delayed impulse response and the right ear non-delayed impulse response with corresponding instances of the left ear whole-pass filter and the right ear whole-pass filter; The method according to claim 1, further comprising:

6. The method according to claim 5, further comprising the step of generating a modified left ear delay value and a modified right ear delay value based on one or more of the extracted left ear delay and right ear delay, wherein the left ear full-pass filter and the right ear full-pass filter are based on the modified left ear delay value and the modified right ear delay value.

7. The method according to claim 6, wherein generating instances of the modified left ear delay and right ear delay values ​​relates to determining the difference between the extracted left ear delay and the extracted right ear delay.

8. The method according to claim 6, wherein generating instances of the modified left ear delay and right ear delay values ​​relates to determining the maximum expected difference between the extracted left ear delay and the extracted right ear delay.

9. The method according to claim 6, wherein the difference between the extracted left ear delay and the extracted right ear delay is equal to the difference between the corresponding modified left ear delay value and the modified right ear delay value.

10. The method according to claim 6, wherein the modified left ear delay value and the modified right ear delay value correspond to a smooth function.

11. The method according to claim 6, wherein each pair of the modified left ear delay value and the modified right ear delay value includes a lower delay value and a higher delay value, the lower delay value having less delay variation than the higher delay value.

12. The method according to claim 5, wherein the non-delayed impulse response is the minimum phase filter response.

13. Extracting each left ear non-delayed impulse response, each right ear non-delayed impulse response, each left ear delay, and each right ear delay from the respective left and right ear HRTFs is: Determine the frequency response of the original HRTF filter of the first set of HRTFs; Determine the absolute response of the original HRTF filter; Determine the minimum phase frequency response of a new non-delay minimum phase filter; The phase response of the original HRTF filter and the phase response of the new non-delay minimum phase filter are determined; Determining the delay associated with the original HRTF filter based at least partially on the phase response of the original HRTF filter and the phase response of the new non-delay minimum-phase filter. The method according to claim 5, relating to the present invention.

14. The method according to claim 13, wherein determining the minimum phase frequency response involves performing a Hilbert transform on the absolute value response of the original HRTF filter.

15. The method according to claim 13, wherein determining the delay associated with the original HRTF filter is at least partially based on a delay measurement frequency in the range of 300 Hz to 1600 Hz.

16. The method according to claim 3, wherein the set of basis filters has at least one order of magnitude fewer elements than the second set of HRTFs.

17. The method according to claim 1, wherein the pass-through phase response deviates from the linear ramp phase response and smoothly approaches the zero phase at frequencies above the threshold frequency.

18. The method according to claim 1, wherein the control system corresponds to at least a portion of a codec for immersive voice and audio services (IVAS).

19. One or more non-temporary computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform an operation according to any one of claims 1 to 18.

20. An audio processor device that processes input audio data, wherein the audio processor device: A receiver unit configured to receive the aforementioned input audio data; It has a computer unit, the computer unit is: The first step is to extract the first set of head-related transfer functions (HRTFs); The step of converting the first set of HRTFs to a second set of HRTFs, wherein the conversion is: Replace the delay component of the first set of HRTFs with the pass-through filter in the second set of HRTFs; Adjust the phase response of each of the full-pass filters in the second set of HRTFs, Each interaural phase response is substantially linear at frequencies below the corresponding threshold frequency of the full-pass filter. Each phase response should have a reduced interaural phase difference at frequencies above the relevant threshold frequency of the corresponding full-pass filter. including, Stages; The step of outputting the second set of HRTF and An audio processor device configured to perform the following actions.

21. The audio processor device according to claim 20, wherein outputting the second set of HRTFs involves storing the second set of HRTFs, transmitting the second set of HRTFs to a device configured to process audio data, providing the second set of HRTFs for further processing, or a combination thereof.

22. The aforementioned computer unit is: A step of defining a set of base filters based on the second set of HRTFs, wherein the set of base filters has fewer elements than the second set of HRTFs; The steps include: obtaining a bitstream of input audio data in the input audio format; The steps include: generating left audio data and right audio data by combining the input audio data with one or more base filters from the set of base filters; The steps of outputting the left audio data and the right audio data. The audio processor device according to claim 20, further configured to perform the following:

23. The audio processor device according to claim 22, wherein outputting the left audio data and the right audio data involves storing the left audio data and the right audio data, transmitting the left audio data and the right audio data, providing the left audio data and the right audio data to a set of loudspeakers for playback by the control system, providing the left audio data and the right audio data for further processing, or a combination thereof.

24. The audio processor device according to claim 22, further comprising a storage device configured to store the first HRTF, the second HRTF, the left audio data, the right audio data, the input audio data, or a combination thereof.

25. The audio processor device according to claim 24, wherein the storage device includes one or more of random access memory, read-only memory, and non-temporary computer-readable media.

26. The device is an audio processor device according to any one of claims 20 to 25, wherein the device corresponds to at least a portion of a codec for immersive voice and audio services (IVAS).