Apparatus and method for head-transfer function compression
Patent Information
- Application Number
- JP2026094568
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-18
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-01
Smart Images

Figure 2026139781000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio signal encoding, audio signal processing and audio signal decoding, in particular to an apparatus and a method for binaural rendering, and more particularly to an apparatus and a method for compression and expansion of head-related transfer functions (HRTF). [Background Art]
[0002] When sound waves are emitted from a speaker to a listener's ears, the sound is modified multiple times, for example, by reflection of the sound waves off walls. As a result, the sound reaching the pinna of the ear contains information about the listening environment in addition to, for example, music and speech.
[0003] In addition, sounds arriving from multiple directions are shaped differently by the listener's head and pinna. Using this information, the listener's brain can determine the approximate direction and distance of the sound source.
[0004] However, when headphones are used, audio is typically emitted almost directly to the listener's eardrum, so all such information is missing. This creates the impression that sound is generated inside the listener's head, which can be perceived as inconvenient, and spectral coloring may occur, for example, especially when earphones are used for an extended period of time.
[0005] It has been determined that the aforementioned modification of sound waves traveling toward a listener's pinna and eardrum can be measured and reproduced by digital filters, for example, by using head impulse responses, head-related transfer functions, binaural room impulse responses, and binaural room transfer functions. When such filters are applied to an audio signal to be reproduced by headphones or small earphones, spatial sound that creates a realistic sound impression is produced. Such audio signal processing is referred to as binaural processing or binaural rendering.
[0006] The head-related transfer function (HRTF) is an acoustic transfer function from a sound source to two ears. The HRTF contains positional information of the corresponding sound source. A virtual sound from a specific direction can be generated by convolution of the corresponding HRTF with the audio signal when listened to through headphones.
[0007] To render spatial sound binaurally, the HRTF (Head-Related Frequency Transform) at relevant locations around the listener is measured and stored. The magnitude of the HRTF is frequency-dependent and provides essential psychoacoustic cues for a reasonable binaural effect. However, these variations across frequency inevitably result in spectral distortion of the binauralized audio signal.
[0008] The degree to which a signal is spectrally distorted depends on several factors, such as the type of input signal (e.g., speech, music, atmosphere, special effects, etc.), the frequency spectrum of the signal, the frequency spectrum of the HRTF, whether dynamic head tracking is used during binaural playback, and the distribution of the signal around the head, and is more or less acceptable.
[0009] When the amplitude of an HRTF is flattened across frequency, signal distortion can be reduced. This flattening is hereafter referred to as HRTF compression. Similarly, the amplitude enhancement of the spectrum can be achieved by the inverse, called HRTF extension, which increases the spectral distortion of the input signal. To avoid redundancy, when referring to "HRTF compression" below, given the understanding that "extension" is simply "negative compression," the term "HRTF compression" includes both HRTF compression and HRTF extension. There is no algorithmic difference between compression and extension.
[0010] There is a trade-off between using an uncompressed HRTF, which provides a complete cue for reasonable binaural rendering but carries the risk of spectral distortion in the binaural audio signal, and using a compressed HRTF, which provides a less effective cue for reasonable binaural rendering but reduces spectral distortion in the binaural audio signal.
[0011] [1] and [2] describe modifications to the HRTF filter to reduce undesirable timbre effects. This technique also reduces undesirable timbre by reducing variations in the root-mean-square (RMS) spectrum of the HRTF.
[0012] [3] describes the effects of size compression or flattening on perceptual results (e.g., externalization).
[0013] [4] and [5] provide a concept for partially binaural virtualizing a single-channel audio signal by filtering. The control allows for a smooth transition between full binaural virtualization based on HRTF and non-binaural virtualization with panning.
[0014] [6], [7], and [8] offer concepts that generally attempt to reduce spectral distortion, but such concepts appear to affect the overall spatial impression and require complex operations or transformations such as principal component analysis (PCA).
[0015] Concepts are described in [8] and [9] that attempt to compress HRTFs to reduce redundancy and the amount of data stored. The HRTF representation is stored / transmitted, requiring less data space than the original dataset. The HRTF is restored before rendering, and the parameterization and reconstruction processes also affect the HRTF size spectrum, often flattening it. The concepts provided in [8] and [9] operate on the entire set of HRTFs and do not treat HRTFs from a particular angle differently, and therefore do not provide direct control over the perceived outcome.
[0016] Ideally, excellent localization and externalization should be achieved without distorting the spectrum of the input signal (and thus maintaining artistic intent). [Overview of the project] [Problems that the invention aims to solve]
[0017] The objective of this invention is to provide an improved concept for binaural rendering. [Means for solving the problem]
[0018] The object of the present invention is solved by the apparatus described in claim 1, the method described in claim 24, and the computer program described in claim 25.
[0019] An apparatus is provided. The apparatus comprises a rendering information processor configured to modify the original binaural rendering information in a manner dependent on directional information to obtain modified binaural rendering information in which spectral distortion has been adjusted.
[0020] Furthermore, a method is provided. The said method is Modify the original binaural rendering information to obtain corrected binaural rendering information with adjusted spectral distortion, depending on the directional information. Includes.
[0021] Further, a computer program for implementing the method described above when executed on a computer or a signal processor.
[0022] In contrast to some embodiments of the present invention, the concepts presented in [1] and [2] do not algorithmically adjust the amount of compression based on the azimuth and elevation angles of the HRTF, and therefore flatten the same amount of HRTF magnitude spectrum in all directions. Furthermore, while the concepts of [1] and [2] operate on HRTF pairs using a joint RMS spectrum of the original left-ear and right-ear filters, embodiments of the present invention use a single HRTF filter, and thus better preserve the interaural RMS difference. Further, embodiments provide the concept of parametrically adjusting compression by weighting the compression rate in multiple dimensions.
[0023] In contrast to some embodiments, [3] does not provide an algorithmic approach for how to adjust flattening to maintain good quality.
[0024] In contrast to some embodiments of the present invention, the concepts provided in [4] and [5] do not provide a parametric frequency-dependent compression approach, nor do they provide an approach for an entire database of HRTFs.
[0025] According to embodiments, spectral distortion in binaural rendering is adjusted.
[0026] In some embodiments, for example, spatial impression can be taken into account.
[0027] According to some embodiments, for example, one or more HRTFs can be processed.
[0028] In some embodiments, one or more HRTFs can, for example, be processed in the frequency domain.
[0029] According to some embodiments, the HRTF size spectrum can be processed in the frequency domain, for example.
[0030] In some embodiments, the HRTF size spectrum can be compressed or expanded, for example.
[0031] According to some embodiments, the HRTF size spectrum can be parametrically compressed or expanded by, for example, taking into account one or more parameters.
[0032] In some embodiments, one of the one or more parameters may be, for example, one or more HRTF elevation angles.
[0033] According to some embodiments, one of the one or more parameters may be, for example, one or more azimuth angles of the HRTF.
[0034] In some embodiments, one of the one or more parameters may be, for example, frequency.
[0035] According to some embodiments, for example, special processing of a front sound source can be provided.
[0036] In some embodiments, for example, an offset can be provided for special processing of a front sound source.
[0037] Some embodiments provide HRTF dynamic compression, which offers means for compressing the spectral magnitude of the HRTF in an azimuth, elevation, and frequency-dependent manner so that a cue of important regions can be preserved, but the magnitude of the HRTF can be controlled in different angular regions to control the spectral distortion and timbre of the resulting binaural output signal.
[0038] According to some embodiments, smooth HRTF compression ratios are provided across angles (az, ele) and frequencies. Applying these coefficients to the HRTF reduces spectral distortion while maintaining the best possible spatial impression.
[0039] In this embodiment, the tonal distortion of the HRTF can be reduced, for example, by reducing its spectral dynamics (in a direction-dependent manner) while maintaining the spatial impression.
[0040] Some embodiments can, for example, act only on the magnitude of the HRTF frequency without correcting the phase.
[0041] Spectral peaks and steps are essential components of any HRTF (especially elevated HRTFs). Peaks and steps are also specific to the direction (azimuth, elevation) of that HRTF. While peaks and steps are essential properties for synthesizing binaural audio, they introduce spectral distortion. Embodiments can provide, for example, HRTF compression means for controlling the trade-off between spectral distortion and spatial effects.
[0042] In one embodiment, the size of each frequency bin can be modified, for example, to be closer to (compressed) or further away from (expanded) the root mean square of the size of the HRTF frequency band.
[0043] According to some embodiments, the amount of compression applied to each HRTF is weighted in a azimuth-dependent manner, so that a front HRTF can obtain, for example, more compression and a rear HRTF can obtain less, and / or weighted in a elevation-dependent manner, so that a horizontal HRTF can obtain, for example, more compression and an elevated one can obtain less.
[0044] In one embodiment, the compression amount can be uniquely modified, for example, with respect to a selected angular focal region, for example, with respect to a central HRTF where the azimuth angle is 0.0 or close to it.
[0045] According to one embodiment, for example, different amounts of compression can be applied across different frequency domains.
[0046] In one embodiment, for example, by using HRTF dynamical compression according to one embodiment, excellent localization and externalization are provided without distorting the spectrum of the input signal (and thus maintaining artistic intent).
[0047] Embodiments of the present invention will be described in more detail below with reference to the drawings. [Brief explanation of the drawing]
[0048] [Figure 1] This is an apparatus according to an embodiment. [Figure 2] This is a device for binaural rendering according to one embodiment. [Figure 3] This figure shows compression for different vsp_angle_factor values without applying center_offset, according to one embodiment. [Figure 4] This shows the effect of center_offset on the overall compression value according to one embodiment. [Figure 5] This figure shows the overall HRTF compression process according to one embodiment. [Modes for carrying out the invention]
[0049] Figure 1 shows an apparatus according to one embodiment. The device includes a rendering information processor 110 configured to modify the original binaural rendering information in accordance with directional information to obtain modified binaural rendering information with adjusted spectral distortion.
[0050] According to one embodiment, the rendering information processor 110 may be configured to modify the original binaural rendering information, for example, such that the degree of adjustment of the spectral distortion depends on the directional information.
[0051] In one embodiment, the binaural rendering information may be suitable for use in processing one or more audio input signals to obtain a binaural signal containing two audio channels, for example.
[0052] Figure 2 shows an embodiment in which the apparatus further comprises a signal processor (120) configured to process one or more audio input signals in reliance on the modified binaural rendering information in order to acquire the binaural signal including the two audio channels.
[0053] According to one embodiment, the original binaural rendering information includes one or more original head-transfer function pairs, each of which includes a first original head-transfer function and a second original head-transfer function. Depending on the directional information, the rendering information processor 110 may be configured to modify the first original head-transfer function and / or the second original head-transfer function of each of the one or more original head-transfer function pairs, for example, to obtain a first modified head-transfer function and / or a second modified head-transfer function of each of the one or more original head-transfer function pairs.
[0054] In one embodiment, the signal processor 120 may be configured to process one or more audio input signals, for example, by relying on at least one modified head-to-transfer function pair from the one or more modified head-to-transfer function pairs, in order to acquire the binaural signal.
[0055] According to one embodiment, each of the first original head-transfer function and each of the second original head-transfer function of one or more original head-transfer function pairs, and each of the first modified head-transfer function and each of the second modified head-transfer function of one or more modified head-transfer function pairs, can be, for example, a head-transfer function, or for example a head-impulse response, or for example a binaural chamber transfer function, or for example a binaural chamber impulse response.
[0056] In one embodiment, the signal processor 120 may include, for example, one or more audio filters for applying at least one first modified head-transfer function and / or a second modified head-transfer function from one or more head-transfer function pairs on an audio input signal.
[0057] In one embodiment, the rendering information processor 110 may be configured to process, for example, the first original head-transfer function and / or the second original head-transfer function of each of one or more head-transfer function pairs in the frequency domain.
[0058] According to one embodiment, the rendering information processor 110 may be configured to modify the first and / or second original head transfer functions of each original head transfer function pair of one or more head transfer function pairs in a direction information-dependent manner, for example, so that the magnitude spectrum of the first original head transfer function and / or the magnitude spectrum of the second original head transfer function of the original head transfer function pair is modified.
[0059] In one embodiment, for example, the rendering information processor 110 may be configured to modify the first and / or second original head-transfer functions of each original head-transfer function pair of one or more head-transfer function pairs in a direction information-dependent manner, such that the magnitude difference between at least one of the two frequency bands of the first and / or second original head-transfer functions of the original head-transfer function pair is corrected.
[0060] In one embodiment, if the directional information indicates, for example, that spectral distortion should be reduced, the rendering information processor 110 may be configured to modify the first and / or second original head-transfer functions of each original head-transfer function pair of one or more head-transfer function pairs in a directional manner, such that the magnitude difference between at least one of the two frequency bands of the first and / or second original head-transfer functions of the original head-transfer function pair is reduced.
[0061] In one embodiment, the directional information may include, for example, the directional information of each of the one or more head transfer function pairs.
[0062] According to one embodiment, the directional information of one or more head-transfer function pairs includes the elevation angle and / or azimuth angle of the head-transfer function pair. The rendering information processor 110 can be configured, for example, to modify the first original head-transfer function and / or second original head-transfer function of the original head-transfer function pair depending on the elevation angle and / or azimuth angle of the head-transfer function pair.
[0063] In one embodiment, the rendering information processor 110 may be configured to determine one or more correction parameters depending on, for example, directional information. The rendering information processor 110 may also be configured to modify the first original head-transfer function and / or the second original head-transfer function of each of one or more original head-transfer function pairs depending on, for example, at least one of the one or more correction parameters. Each of the one or more correction parameters indicates the degree of adjustment of spectral distortion.
[0064] According to one embodiment, the rendering information processor 110 may be configured to determine one or more modification parameters by determining, for example, at least one modification parameter from one or more modification parameters for each of one or more head-transfer function pairs, depending on the elevation angle and / or azimuth angle of the head-transfer function pair. Each of the modification parameters for the head-transfer function pair indicates the degree of adjustment of spectral distortion in the first modified head-transfer function and / or the second modified head-transfer function of the modified head-transfer function pair, compared to the spectral distortion in the first original head-transfer function and / or the second original head-transfer function of the original head-transfer function pair.
[0065] In one embodiment, the rendering information processor 110 may be configured to determine, for example, the at least one modification parameter for each of the one or more head-transfer function pairs in a frequency-dependent manner.
[0066] According to one embodiment, the rendering information processor 110 may be configured to generate at least one modification parameter for each of the one or more head-related transfer function pairs such that, for example, when the azimuth angle of the head-related transfer function pair indicates the presence of a frontal sound source, the modification parameter can be generated in a different way for the same value of the elevation angle of the head-related transfer function pair compared to when the azimuth angle of the head-related transfer function pair does not indicate the presence of a frontal sound source.
[0067] In one embodiment, the rendering information processor 110 may be configured to generate at least one modification parameter for each of the one or more head-related transfer function pairs by generating an offset value corresponding to the elevation angle of the head-related transfer function pair when, for example, the azimuth angle of the head-related transfer function pair indicates the presence of the front sound source.
[0068] According to one embodiment, one or more original head-transfer function pairs are a plurality of original head-transfer function pairs, and the directional information includes different directional information for each of the plurality of head-transfer function pairs. The rendering information processor 110 may be configured to modify the first original head-transfer function and / or the second original head-transfer function of each original head-transfer function pair of the plurality of original head-transfer function pairs, depending on the directional information of the original head-transfer function pairs, in order to obtain the first modified head-transfer function and / or the second modified head-transfer function of each of the one or more modified head-transfer function pairs, which are a plurality of modified head-transfer function pairs. The signal processor 120 may be configured to process one or more audio input signals, depending on at least one of the plurality of modified head-transfer function pairs, in order to obtain the binaural signal.
[0069] In one embodiment, the signal processor 120 may be configured to process one or more audio input signals depending on one or more interpolated head-transfer function pairs to acquire, for example, a binaural signal. The rendering information processor 110 may be configured to determine an interpolated head-transfer function including, for example, a first interpolated head-transfer function and a second interpolated head-transfer function. The rendering information processor 110 may be configured to determine a first interpolated head-transfer function by, for example, interpolating between the first head-transfer functions of at least two head-transfer function pairs from a plurality of head-transfer function pairs, depending on directional information. Furthermore, the rendering information processor 110 may be configured to determine a second interpolated head-transfer function by, for example, interpolating between the second head-transfer functions of at least two head-transfer function pairs from a plurality of head-transfer function pairs, depending on directional information.
[0070] The following describes a specific embodiment. Where a head-related transfer function is referred to below, such reference should be understood as an example of a particular embodiment. The concepts provided are equally applicable to the time domain of head-related impulse responses and equally applicable to binaural-indoor transfer functions and binaural-indoor impulse responses. The term “head-related transfer function” includes head-related transfer functions, head-related impulse responses, binaural-indoor transfer functions, and binaural-indoor impulse responses.
[0071] For example, an HRTF measured in the quadrature mirror filter (QMF) region is considered "uncompressed." The most extreme version of HRTF compression yields a flat line across all frequencies. In such an extreme version of HRTF compression, the magnitude of all QMFs in all frequency bins is equal to the root mean square (RMS) of the uncompressed HRTF.
[0072] According to some embodiments, the HRTF compression stage takes into account the possible directions of sound arrival, i.e., the azimuth and elevation angles around the listener's head, and provides a way to smoothly transition between these two states, "uncompressed" and "fully compressed."
[0073] In some embodiments, compression ratios that depend on modification parameters, such as azimuth and elevation angles, can be calculated, for example, for each HRTF pair. In one embodiment, a single compression ratio for both the left and right ears for a specific azimuth and elevation angle can be calculated, for example. This coefficient may be further weighted, for example, for each frequency band, for example, for each QMF band, in one embodiment. This coefficient can be applied, for example, to the corresponding frequency bands of the corresponding HRTF. In another embodiment, two compression ratios for the left and right ears for a specific azimuth and elevation angle can be calculated, for example.
[0074] According to one embodiment, compression can be applied parametrically, depending on three parameters, for example, azimuth, elevation, and frequency. In one embodiment, compression may include the possibility of particularly reducing compression for a specific angular focal region (e.g., when the azimuth is equal to or close to 0°). This provides a flexible compression setting for the direction in which timbre-sensing signals, such as speech and vocals, are typically positioned.
[0075] The provided concept can, for example, take into account the HRTF angle, and keep the variation in the HRTF size spectrum relatively high at angles that are considered highly relevant to binaural rendering in order to maintain spatial cues. An advantage of some embodiments is that the concept can be highly customizable, for example, so that gradual changes in compression as a function of azimuth, elevation, and frequency can be adjusted to best suit, for example, the input content and / or the intent of binaural rendering.
[0076] Specific embodiments are provided below. A set of head-related transfer functions (HRTFs) can be used, for example, as input for calculating the compression ratio. The set of HRTFs can be defined, for example, as multiple HRTFs for different combinations of azimuth and elevation angles.
[0077] HRTF can be transferred to the frequency domain, for example, using a QMF filter bank.
[0078] Subsequently, the compression ratio can be calculated, for example, for each HRTF and each QMF region frequency band.
[0079] For example, in frequency band 1, the compression ratio can always be set to 1.0 for all HRTFs. For QMF bands 2-64, the loop is executed across all HRTFs, and one compression ratio can be calculated for each HRTF pair (left and right ears) for each QMF band, as defined in the pseudocode below. % compression factor hrtfComp_vspBaselineFactor = 1.0; % angle-dependence for compression factor hrtfComp_vspAzimuthFactor_anchor_horizontal = 0.25; hrtfComp_vspAzimuthFactor_anchor_elevated = 0.05; % centre offset hrtfComp_centerOffset_anchor_horizontal = 0.10; hrtfComp_centerOffset_anchor_elevated = 1.00; % weight the compression factor dependent on the angle and elevation compressionAngleFactor = cosd(abs(hrtf_elevations(i))); compressionAngleFactor = (compressionAngleFactor - cosd(hrtfComp_el2compress)) * ... ((hrtfComp_vspAzimuthFactor_anchor_horizontal - hrtfComp_vspAzimuthFactor_anchor_elevated) / ... (1 - cosd(hrtfComp_el2compress))) + hrtfComp_vspAzimuthFactor_anchor_elevated; compression2apply = hrtfComp_vspBaselineFactor * abs(sind(hrtf_azimuths(i) / 2)^compressionAngleFactor); % compute the weighting for the center (az=0) if (hrtf_azimuths(i)==0) compressionCenterOffset = sind(abs(hrtf_elevations(i))); compressionCenterOffset = compressionCenterOffset * ... ((hrtfComp_centerOffset_anchor_elevated - hrtfComp_centerOffset_anchor_horizontal) / ... sind(hrtfComp_el2compress)) + hrtfComp_centerOffset_anchor_horizontal; compression2apply = compression2apply + compressionCenterOffset; end % compute RMS before compression over all frequencies for e = 1:2 % L / R ears input_rms(e) = rms( squeeze(hrtf(:, e, i)) ); for b = 2:64 % QMF bands % apply compression hrtf(b, e, i) = input_rms(e) + ( (hrtf(b, e, i) - input_rms(e)) * compression2apply); end end The resulting compression ratio ranges from 0.0 (fully compressed) to 1.0 (uncompressed) for each HRTF and each bandwidth. This depends on the following three parameters:
[0080] Baseline compression is defined by hrtfComp_vspBaselineFactor, which defines the overall upper limit of compression. This can be set to 1.0 by default, for example, in the range between 0.0 (full compression) and 1.0 (uncompressed). Other values may be used, for example. For example, in another embodiment, 0 may indicate full compression, for example, and 100 may indicate uncompressed, for example. In one embodiment, baseline compression can set an upper limit on compression, for example, and the compression may typically be lower than this value, for example, since further azimuth and elevation weighting may be applied, for example.
[0081] The compressionAngleFactor defines an elevation-dependent weighting. This is limited by the compressionAngleFactor at elevation angle 0°, defined by hrtfComp_vspAzimuthFactor_anchor_horizontal, and the compressionAngleFactor at elevation angle 90°, defined by hrtfComp_vspAzimuthFactor_anchor_elevated. Between + / -90° and 0°, intermediate values can be calculated, for example, using a cosine function. % angle-dependence for compression factor hrtfComp_vspAzimuthFactor_anchor_horizontal = 0.25; hrtfComp_vspAzimuthFactor_anchor_elevated = 0.05; compressionAngleFactor = cosd(abs(hrtf_elevations(i))); compressionAngleFactor = (compressionAngleFactor - cosd(hrtfComp_el2compress)) * ... ((hrtfComp_vspAzimuthFactor_anchor_horizontal - hrtfComp_vspAzimuthFactor_anchor_elevated) / ... (1 - cosd(hrtfComp_el2compress))) + hrtfComp_vspAzimuthFactor_anchor_elevated; `compressionCenterOffset` can define an elevation-dependent offset that is applied, for example, when the azimuth angle of the current HRTF is 0°. `compressionCenterOffset` can be minimized, for example, when elevation = 0°. This ensures a compression ratio close to 0.0, which corresponds to a high level of compression and therefore results in a very clean timbre for important signals such as speech. As the elevation angle increases, `compressionCenterOffset` gradually increases, weighted, for example, by a sine function, up to a defined maximum value. If the azimuth angle is not equal to 0°, `compressionCenterOffset` is set to 0.0. % center offset hrtfComp_centerOffset_anchor_horizontal = 0.10; hrtfComp_centerOffset_anchor_elevated = 1.00; compressionCenterOffset = sind(abs(hrtf_elevations(i))); compressionCenterOffset = compressionCenterOffset * ... ((hrtfComp_centerOffset_anchor_elevated - hrtfComp_centerOffset_anchor_horizontal) / ... sind(hrtfComp_el2compress)) + hrtfComp_centerOffset_anchor_horizontal;
[0082] In one embodiment, for example, compressionCenterOffset is applied to avoid compression2apply becoming equal to 0.0 for an azimuth angle of 0°. Furthermore, compressionCenterOffset is sinusoidally weighted by the elevation angle to ensure that the total compression applied to the central source decreases as the elevation angle increases. A sine function can be used to calculate a value between 0.0 and 1.0 by applying, for example, a sine function sin(abs(elevation angle)). As a result, a very clean timbre is preserved for a horizontal central source, while the elevation angle and externalized cues are still preserved.
[0083] By applying the described algorithm, the angle-dependent bandwidth compression ratio is calculated. This is shown in Figure 3 for different compressionAngleFactor values before applying compressionCenterOffset.
[0084] In particular, Figure 3 shows compression for different compressionAngleFactor values without applying compressionCenterOffset according to one embodiment.
[0085] The plots represent different compressionAngleFactor values (and therefore different elevation angles). When compressionAngleFactor is greater than 0, the compression for a frontal sound source increases. Specifically, compressionAngleFactor 0 is indicated by line 2000, compressionAngleFactor 0.25 by line 2025, compressionAngleFactor 0.5 by line 2050, compressionAngleFactor 0.75 by line 2075, and compressionAngleFactor 1 by line 2100.
[0086] Increasing the compressionAngleFactor increases the compression of the front sound source (it becomes more strongly compressed).
[0087] Without `compressionCenterOffset`, azimuth = 0.0 will always result in `compression2apply=0.0`. This can be undesirable, for example, because 100% compression would destroy all elevation cues and externalizations.
[0088] In one embodiment, compressionCenterOffset is applied to overcome this problem. For example, compressionCenterOffset is minimized when the elevation angle is 0, for example, and a very clean timbre can be achieved for important signals such as speech. As the elevation increases, compressionCenterOffset can be weighted by a sine function, for example, and gradually increased up to a maximum value, for example, a maximum value defined by a chord.
[0089] Figure 4 shows the effect of applying compressionCenterOffset. In particular, Figure 4 shows the effect of compressionCenterOffset on the overall compression value in one embodiment.
[0090] According to some embodiments, the final compression ratio can be applied to, for example, HRTF.
[0091] For example, the final compression ratio can be applied to the HRTF by, for example, calculating the RMS (root mean square) of the uncompressed HRTF, or by calculating the RMS separately for the left / right ear for each individual HRTF. This RMS value can, for example, act as a pivot point across all frequencies, resulting in a flat line across frequencies. By calculating the RMS separately for the left and right ears, it is ensured that the overall ILD (inter-ear level difference) is preserved after compression.
[0092] In the most extreme case (100% compression), the size of all frequency bins is set to the same RMS value (a flat line across frequencies), resulting in passive downmixing. This produces an effect very similar to VSP 0 processing (VSP = virtual speaker position), except that the original energy of the HRTF is preserved.
[0093] Next, for each QMF domain frequency band, a weighted version of the difference between the RMS and the uncompressed HRTF (weighted by the corresponding compression ratio for each band) can be added to the RMS, for example. For each frequency, a weighted version of the difference between the RMS and the uncompressed HRTF can be added to the RMS, for example.
[0094] For example, regarding the frequency band b, b=0...63, For all azimuth angles az, For all elevation angles ele, the following formula can be used: HRTF compressed (az, ele, b) = rms (HRTF uncompressed ) + + (( HRTF uncompressed (az, ele, b) - rms (HRTF uncompressed )) · · compressionFactor (az, ele, b)) HRTF compressed This represents the compressed HRTF. HRTFin compressed This represents an uncompressed HRTF. RMS stands for Root Mean Square.
[0095] For example, the compression ratio (az,ele,0) may always be equal to 0.25 over all azimuth and elevation angles.
[0096] This means that if compressionFactor is equal to 0.0 (full compression), the output HRTF can be, for example, the RMS of the input HRTF. If compressionFactor is equal to, for example, 0.5 (50% compression), a portion of the uncompressed HRTF is converted back to RMS in a frequency-dependent manner.
[0097] For example, in one embodiment, for example, the rendering information processor 110 is, for example, HRTF compressed By using the above formula for, or by using another formula, an HRTF compression stage can be provided that provides a means for smoothly moving between these two states, "uncompressed" and "fully compressed".
[0098] Figure 5 shows the overall HRTF compression process according to one embodiment. In one embodiment, by adopting the above concept, means for individual (user-specific) HRTF personalization are provided. According to one embodiment, different compression presets may be provided to the user for selection, for example. In one embodiment, the user can individually adjust the HRTF compression to their liking.
[0099] Further embodiments are provided below. According to one embodiment, compression may be applied to a different frequency domain, for example, the FFT domain rather than the QMF domain.
[0100] In one embodiment, the processing of the lowest frequency band may be performed in a different way, for example. For example, the dynamic algorithm may start operating in band 0 instead of band 1. Alternatively, different predefined fixed compression values for the lowest band may be used.
[0101] According to one embodiment, the application of angle-dependent compression may be limited, for example, to different ranges of frequency bands (e.g., starting with band 2 instead of band 1, or only performing up to band 48).
[0102] In some embodiments, for example, a function other than the cosine or sine function can be used to define an angle-dependent parameter, such as a linear function of a quadratic function.
[0103] According to one embodiment, the orientation of the angle coefficient weighting pattern may be such that, for example, a high level of compression is applied to the rear sound source instead, while the front sound source is hardly compressed.
[0104] In some embodiments, for example, varying amounts of compression across different frequency domains can be used to preserve the elevation queuing of an HRTF that has risen at higher frequencies.
[0105] Some embodiments are described in the context of apparatus, but these embodiments also represent a description of the corresponding method, and it is clear that the block or device corresponds to a method step or a feature of a method step. Similarly, embodiments described in the context of a method step also represent a description of an item or feature of the corresponding block or corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0106] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or software, or at least partially in hardware or at least partially in software. Embodiments may be implemented using digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, which have electronically readable control signals stored therein and cooperate with (or are capable of cooperating with) a programmable computer system so that each method is performed. Thus, the digital storage media may be computer-readable.
[0107] Some embodiments of the present invention include a data carrier having an electronically readable control signal, such that one of the methods described herein is performed in cooperation with a programmable computer system.
[0108] Generally, embodiments of the present invention can be implemented as a computer program product having program code that operates to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0109] Other embodiments include a computer program for performing one of the methods described herein, which is stored on a machine-readable carrier.
[0110] Therefore, in other words, embodiments of the method of the present invention are computer programs having program code for performing one of the methods of the present invention when the computer program is executed on a computer.
[0111] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) recorded therein, which includes a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-temporary.
[0112] Therefore, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, such as the Internet.
[0113] Further embodiments include processing means configured to perform or applied to perform one of the methods described herein, such as a computer or a programmable logic device.
[0114] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0115] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0116] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0117] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0118] The methods described herein may be performed using hardware devices, or using a computer, or using a combination of hardware devices and a computer.
[0119] The embodiments described above are merely illustrative of the principles of the present invention. Modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended that the invention is limited only by the immediate claims and not by the specific details shown in the description and explanation of the embodiments herein.
[0120] literature [1] Merimaa, J.: “Modification of HRTF filters to reduce timbral effects in binaural synthesis”, In Audio Engineering Society Convention 127. Audio Engineering Society, 2009. [2] Merimaa, J.: “Modification of HRTF filters to reduce timbral effects in binaural synthesis, part 2: Individual HRTFs”, In Audio Engineering Society Convention 129, 2010. [3] Song Li: “On Externalization of Virtual Sound Images Presented via Headphones”, PhD Thesis University Hannover, 2021, (Experiment D). [4] DE 10 2019 135 690 A1, Pellegrini, “Verfahren und Vorrichtung zur Audiosignalverarbeitung fuer binaurale Virtualisierung“. [5] US 2021 0195361 A1, Pellegrini, “Method and device for audio signal processing for binaural virtualization”. [6] Marentakis, G. and Hoelzl, J.: “Compression Efficiency and Signal Distortion of Common PCA Bases for HRTF Modelling” 18th Sound and Music Computing Conference (SMC 2021), Virtual, 29 June - 01 July 2021. [7] Hoelzl, J.: “An initial Investigation into HRTF Adaptation using PCA”. [8] Lin Wang: “HRTF Compression via Principal Components Analysis and Vector Quantization”, 2008. [9] Jing Wang: “Compression of Head-Related Transfer Function Based on Tucker and Tensor Train Decomposition”, 2019.
Claims
1. Apparatus, the apparatus, A rendering information processor (110) configured to modify the original binaural rendering information based on directional information to obtain modified binaural rendering information with adjusted spectral distortion. A device equipped with the following features.
2. The rendering information processor (110) is configured to modify the original binaural rendering information such that the degree of adjustment of the spectral distortion depends on the directional information. The apparatus according to claim 1.
3. The aforementioned binaural rendering information is suitable for use in processing one or more audio input signals to obtain a binaural signal containing two audio channels. The apparatus according to claim 1 or 2.
4. The apparatus further comprises a signal processor (120) configured to process one or more audio input signals in reliance on the modified binaural rendering information in order to acquire the binaural signal including the two audio channels. The apparatus according to claim 3.
5. The original binaural rendering information includes one or more original head-transfer function pairs, each of which includes a first original head-transfer function and a second original head-transfer function. Depending on the directional information, the rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of each of the one or more original head transfer function pairs to obtain the first modified head transfer function and / or the second modified head transfer function of each of the one or more modified head transfer function pairs. The apparatus according to any one of claims 1 to 4.
6. The rendering information processor (110) determines coefficients depending on the direction, The rendering information processor (110) is configured to apply the coefficients to at least one of the first original head transfer function and / or the second original head transfer function from the one or more original head transfer function pairs to obtain at least one of the first modified head transfer function and / or the second modified head transfer function from the one or more modified head transfer function pairs. The apparatus according to claim 5.
7. The aforementioned coefficient is the compressibility ratio. The apparatus according to claim 6.
8. The signal processor (120) is configured to process one or more audio input signals in order to acquire the binaural signal, relying on at least one modified head-transfer function pair from the one or more modified head-transfer function pairs. The apparatus according to any one of claims 5 to 7, further dependent on claim 4.
9. Each of the first original head-transfer function and each of the second original head-transfer function of each of the one or more original head-transfer function pairs, and each of the first modified head-transfer function and each of the second modified head-transfer function of each of the one or more modified head-transfer function pairs, is either a head-transfer function, a head-impulse response, a binaural chamber-transfer function, or a binaural chamber-impulse response. The apparatus according to any one of claims 5 to 8.
10. The rendering information processor (110) is configured to process the first original head-transfer function and / or the second original head-transfer function of each of the one or more head-transfer function pairs in the frequency domain. The apparatus according to any one of claims 5 to 9.
11. The rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of each original head transfer function pair of the one or more head transfer function pairs in a direction-dependent manner, such that the magnitude spectrum of the first original head transfer function and / or the magnitude spectrum of the second original head transfer function of the original head transfer function pair is modified. The apparatus according to any one of claims 5 to 10.
12. The rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of each original head transfer function pair of the one or more head transfer function pairs, depending on the direction information, so as to correct the magnitude difference between at least one of the two frequency bands of the first original head transfer function and / or the second original head transfer function of the original head transfer function pair. The apparatus according to any one of claims 5 to 11.
13. If the directional information indicates that spectral distortion should be reduced, the rendering information processor (110) is configured to modify the first and / or second original head transfer functions of each original head transfer function pair of the one or more head transfer function pairs in a directional manner, such that the difference in magnitude of at least one between two frequency bands of the first and / or second original head transfer functions of the original head transfer function pair is reduced. The apparatus according to claim 12.
14. The directional information includes the directional information of each of the one or more head transfer function pairs. The apparatus according to any one of claims 5 to 13.
15. The directional information for each of the one or more head-transfer function pairs includes the elevation angle and / or azimuth angle of the head-transfer function pair. The rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of the original head transfer function pair depending on the elevation angle and / or azimuth angle of the head transfer function pair. The apparatus according to claim 14.
16. The rendering information processor (110) is configured to determine one or more modification parameters depending on the direction information, The rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of each of the one or more original head transfer function pairs, depending on at least one of the one or more modification parameters. Each of the one or more correction parameters indicates the degree of adjustment of the spectral distortion. The apparatus according to any one of claims 5 to 15, further dependent on claim 2.
17. The rendering information processor (110) is configured to determine one or more modification parameters by determining at least one of the one or more modification parameters for each of the one or more head-transfer function pairs, depending on the elevation angle and / or the azimuth angle of the head-transfer function pair. Each of the modification parameters of the head-to-head transfer function pair indicates the degree of adjustment of the spectral distortion in the first modified head-to-head transfer function and / or the second modified head-to-head transfer function of the modified head-to-head transfer function pair, compared to the spectral distortion in the first original head-to-head transfer function and / or the second original head-to-head transfer function of the original head-to-head transfer function pair. The apparatus according to claim 16, further dependent on claim 15.
18. The rendering information processor (110) is configured to determine, in a frequency-dependent manner, at least one modification parameter for each of the one or more head-related transfer function pairs. The apparatus according to claim 17.
19. The rendering information processor (110) is configured to generate at least one modification parameter for each of the one or more head-related transfer function pairs such that when the azimuth angle of the head-related transfer function pair indicates the presence of a frontal sound source, the modification parameter is generated in a different way for the same value of the elevation angle of the head-related transfer function pair compared to when the azimuth angle of the head-related transfer function pair does not indicate the presence of a frontal sound source. The apparatus according to claim 17 or 18.
20. The rendering information processor (110) is configured to generate at least one modification parameter for each of the one or more head-related transfer function pairs by generating an offset value corresponding to the elevation angle of the head-related transfer function pair when the azimuth angle of the head-related transfer function pair indicates the presence of the front sound source, The apparatus according to claim 19.
21. The one or more original head-transfer function pairs are a plurality of original head-transfer function pairs, and the directional information includes different directional information for each of the plurality of head-transfer function pairs. The rendering information processor (110) is configured to modify the first original head transfer function and / or the second original head transfer function of each of the plurality of original head transfer function pairs, depending on the direction information of the original head transfer function pairs, in order to obtain the first modified head transfer function and / or the second modified head transfer function of each of the one or more modified head transfer function pairs, which are a plurality of modified head transfer function pairs. The signal processor (120) is configured to process one or more audio input signals in order to acquire the binaural signal, depending on at least one modified head-transfer function pair among the plurality of modified head-transfer function pairs. The apparatus according to any one of claims 5 to 20, further dependent on claim 4.
22. The signal processor (120) is configured to process one or more audio input signals in order to acquire the binaural signal, relying on one or more interpolated head-transfer function pairs. The rendering information processor (110) determines the interpolated head transfer function which includes a first interpolated head transfer function and a second interpolated head transfer function. The rendering information processor (110) is configured to determine the first interpolated head transfer function by interpolating between the first head transfer functions of at least two pairs of head transfer functions among the plurality of head transfer function pairs, depending on the direction information. The rendering information processor (110) is configured to determine the second interpolated head transfer function by interpolating between the second head transfer functions of each of the at least two pairs of head transfer functions among the plurality of head transfer function pairs, depending on the direction information. The apparatus according to claim 21.
23. In order to modify the original binaural rendering information based on the directional information and obtain the modified binaural rendering information with adjusted spectral distortion, the rendering information processor (110) is configured to use the following formula: HRTF compressed (az, ele, b) = rms (HRTF uncompressed ) + + (( HRTF uncompressed (az, ele, b) - rms (HRTF uncompressed )) ・ ・ compressionFactor (az, ele, b)) In the formula, b represents the frequency band. az indicates the azimuth angle. ele indicates the angle of elevation. HRTF compressed This represents compressed HRTF, HRTF uncompressed This represents the uncompressed HRTF, rms represents the root mean square. The apparatus according to any one of claims 1 to 22.
24. A method, wherein the said method is Modify the original binaural rendering information to obtain corrected binaural rendering information with adjusted spectral distortion, depending on the directional information. Methods that include...
25. A computer program for carrying out the method according to claim 24, when executed on a computer or signal processor.