Methods, apparatus, and programs for fast-moving audio source rendering

The method addresses the challenge of rendering fast-moving audio sources in VR, AR, and XR by calculating a modeled position based on known audio object positions and sound speed, ensuring realistic auditory localization.

WO2026093068A1PCT designated stage Publication Date: 2026-05-07DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2025-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional audio renderers fail to provide perceptually realistic rendering of fast-moving audio sources in VR, AR, and XR environments due to assumptions of constant velocity and fixed direction, leading to inaccurate and potentially impossible auditory positions.

Method used

A method for determining a modeled rendering position of fast-moving audio objects by maintaining a list of known positions and calculating a time offset based on the speed of sound, allowing for perceptually realistic rendering without requiring knowledge of future movement.

Benefits of technology

The method ensures that the modeled auditory position corresponds to the actual audio object position, avoiding perceptually confusing or impossible positions, and is computationally efficient with minimal memory storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000004_0001
    Figure IMGF000004_0001
  • Figure IMGF000004_0002
    Figure IMGF000004_0002
  • Figure IMGF000004_0003
    Figure IMGF000004_0003
Patent Text Reader

Abstract

The disclosure relates to methods of rendering object-based audio content, comprising: determining a rendering position at each of a plurality of time instances and maintaining a list of one or more known rendering positions including a current rendering position and a number N (N ≥ 0) of previous rendering positions at a current and previous time instances, respectively; determining a measure of a time offset based on the current rendering position and a rendering position at the preceding time instance; determining a modeled time instance corresponding to one of the known rendering positions, closest in time to a timing that precedes the current time instance by the time offset; and determining a modeled rendering position based on a rendering position among the known rendering positions that corresponds to the determined modeled time instance, for modeling an offset between auditory and visual positions of an audio object. The disclosure further relates to corresponding apparatus, programs, and computer-readable storage media.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS, APPARATUS, AND PROGRAMS FOR FAST-MOVING AUDIO SOURCE RENDERING

[0002] Cross-Reference to Related Application

[0003] This application claims the priority of U. S. Provisional Patent Application No. 63 / 712,704, filed on 28 October 2024, and European Application No. 24209604.8 filed on 29 October 2024, each of which is incorporated by reference herein in its entirety.

[0004] Technical Field

[0005] The present disclosure relates to techniques for rendering object-based audio content. In particular, the present disclosure relates to rendering fast-moving audio objects (audio sources) in a perceptually realistic manner.

[0006] Background

[0007] For fast-moving audio sources in audio scenes such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), or Extended Reality (XR) audio scenes, the visual and acoustical perceptual localizations may significantly differ from each other. Since the speed of light is largely greater than the speed of sound, the acoustic position may significantly lag behind the visual perceptual localization for a fast-moving audio source. However, conventional audio renderers, such as the MPEG-I immersive audio Tenderer, do not support realistic modelling of delayed audio source positions.

[0008] Some feasible solutions for this problem modify the fast-moving audio source position and the corresponding distance. An example of one such solution 100A is schematically illustrated in Fig. 1A. Here, an audio source 20 at a current rendering position O(tr) isrendered to a listener (or observer) 10 at listener position L, applying distance attenuation depending on the distance dist = l|O(tr) - L\\, as well as applying a delay to the rendered audio signal.

[0009] Other feasible solutions are based on the assumption of constant velocity and fixed moving direction (for each calculation cycle) and do not take any psychoacoustical aspects into consideration. One example of program code 200 for such solution is shown in Fig. 2, which models an audio source trajectory as O(t) = xQ+ v • (t — t0) based on a position xQat a time t0and a velocity v estimated at time t0, and calculates a delayed audio source position based on the modeled trajectory and a time delay that is calculated from a distance and the speed of sound. However, the above assumptions are applicable only to close-to-linear audio object trajectories and / or low audio object velocities. They fail for situations in which the audio object trajectories are non-linear and / or not known in advance. Also, there may be no a-priori knowledge in which objects can be fast-moving. Thus, if the audio object velocity is high and / or its trajectory is nonlinear, applying conventional techniques as described above may result in situations where the delayed audio source position is not perceptually correct.

[0010] An example of where the conventional assumptions will fails is schematically shown in Fig. 3.

[0011] Here, an audio object (e.g., rotating object with high velocity) moves on a circular trajectory with previous audio object positions 320 and current audio object position 20. When extrapolating movement of the audio object based on the current audio object position 20, extrapolated velocity vector 360 would be obtained. When applying a certain delay, and correspondingly a delayed distance 340, this would result in a modeled auditory audio object position 330 (in a direction opposite to the extrapolated velocity vector 360), which is clearly different from and significantly removed from any of the actual previous audio object positions 320. Even worse, in some cases the modeled auditory audio object position 330 may be located in physically impossible areas, such as within occluder 370 in the example of Fig. 3. Each of these effects may result in perceptually unrealistic rendering of the audio object and may potentially lead to rendering artifacts.

[0012] There is thus need for improved techniques for rendering object-based audio content. There is particular need for such coding techniques that enable perceptually realistic rendering of fastmoving audio sources. Summary

[0013] In view of this need, the present disclosure provides methods of rendering object-based audio content, as well as corresponding apparatus, computer programs, and computer-readable storage media, having the features of respective independent claims.

[0014] One aspect of the present disclosure relates to a method of rendering object-based audio content. The method may include determining a rendering position of an audio object at each of a plurality of time instances. Rendering may be performed at each of a plurality of equi- spaced time instances, i.e., at a given rate or frequency. The method may further include maintaining (a list of) one or more known rendering positions of the audio object including a current rendering position at a current time instance and a number N of previous rendering positions at respective previous time instances. Therein, N may be a non-negative integer, N > 0. In total, N + 1 known time instances may be maintained. The audio object may be a moving audio object, in particular a fast-moving audio object. The method may further include, for the current time instance, determining a measure of a time offset based on the current rendering position and a rendering position of the audio object at the preceding time instance. The preceding time instance may relate to the immediately preceding time instance to the current time instance. The time offset may correspond to a time delay, for example. The method may further include, for the current time instance, determining a modeled time instance that corresponds to one of the known rendering positions and that is closest in time to a timing that precedes the current time instance by the time offset. The method may yet further include determining a modeled rendering position of the audio object based on a rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance, for modeling an offset between auditory and visual positions of the audio object. The modeled rendering position may account for different propagation speeds of light and sound in a (virtual) environment.

[0015] With these features, the proposed method allows to render fast-moving audio objects in a perceptually convincing way, by modelling the offset between visual and auditory positions for such audio objects. The proposed method particularly ensures that the modeled auditory position corresponds to an actual previous position of the audio object, which will be relevant especially for non-linear movement of the audio object, and thereby avoids rendering of the audio object from perceptually confusing or even physically impossible positions. All this can be done in computationally efficient manner, only by requiring memory storage of a certain number of known rendering positions, and without knowledge of future movement of the audio object. In some embodiments, determining the measure of the time offset may be based on a value representing the speed of sound and a measure of a distance derived from the current rendering position and the rendering position for the preceding time instance.

[0016] In some embodiments, the time offset toffsetmay be determined as toffset=

[0017] where t_{i*} is the current time instance, is the preceding time instance,

[0018]

[0019]

[0020] O(ti*) is the current rendering position, O(t_{i*-1}) is the rendering position of the audio object at the preceding time instance, s is a value for the speed of sound, and (-,-) is a distance function (e.g., distance metric).

[0021] In some embodiments, the modeled time instance t_{i*-k} may be determined via

[0022]

[0023] In some embodiments, the method may further include setting the time offset to zero if the modeled time instance is identical to the current time instance, and setting the modeled rendering position to the current rendering position. This may amount to disregarding the modeled rendering position in rendering the audio object at the current time instance.

[0024] In some embodiments, determining the modeled rendering position may involve at least one of low-pass filtering based on known rendering positions, applying curve fitting based on known rendering positions, and interpolation based on known rendering positions. In the simplest case however, the modeled rendering position may be identical to the rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance.

[0025] Having available a set of known rendering positions in memory and applying low-pass filtering, curve fitting, and / or interpolation allows for mitigating effects of discrete rendering and for achieving a more realistic listening experience, especially for non-linear movement of the audio object. In some embodiments, the method may further include determining a relative angle between the current rendering position and the modeled rendering position in relation to a user position. The user may be a listener and / or observer, located in the audio scene. The method may further include comparing the determined relative angle to a threshold for the relative angle. The method may yet further include using the modeled rendering position for rendering the audio object at the current time instance if (e.g., only if) the threshold for the relative angle does not exceed the relative angle. If the threshold exceeds the relative angle (or, possibly, is equal to the relative angle), the current rendering position may be used instead for rendering the audio object at the current time instance.

[0026] In some embodiments, the method may further include determining the threshold for the relative angle based on an orientation of the user. The orientation of the user may for example relate to a head orientation of the user. The threshold may be further based on a spatial localization listening sensitivity pattern.

[0027] In some embodiments, the method may further include determining a relative azimuth angle and a relative elevation angle between the current rendering position and the modeled rendering position in relation to a user position. The method may further include comparing the determined relative azimuth angle to a threshold for the relative azimuth angle and comparing the determined relative elevation angle to a threshold for the relative elevation angle. The method may yet further include using the modeled rendering position for rendering the audio object at the current time instance if (e.g., only if) the threshold for the relative azimuth angle does not exceed the relative azimuth angle and the threshold for the relative elevation angle does not exceed the relative elevation angle. The threshold for the relative elevation angle may be larger than the threshold for the relative azimuth angle, accounting for different positioning sensitivities of the human auditory system in azimuth and elevation.

[0028] In some embodiments, the method may further include determining a relative distance based on the current rendering position, the modeled rendering position, and the user position. The method may further include comparing the determined relative distance to a threshold for the relative distance. The method may yet further include using the modeled rendering position for rendering the audio object at the current time instance if (e.g., only if) the threshold for the relative distance does not exceed the relative distance. The relative distance may be defined for example as (the absolute value of) the difference between absolute values of distances between the current rendering position and the user position and between the modeled rendering position and the user position, respectively. If both an angular condition and a distance condition are applied, the modeled rendering position may only be used if both conditions are satisfied.

[0029] The foregoing embodiments allow to perform an efficient perceptual relevance check for deciding on whether the modeled rendering position shall be used for rendering, or disregarded. Results of this perceptual relevance check may be used, for example, for temporarily disabling determination of the modeled rendering position.

[0030] In some embodiments, a total number of IV + 1 known rendering positions including the current rendering position and N previous rendering positions may be maintained. The method may further include decreasing the number N if the modeled time instance corresponds to the current time instance. The method may yet further include increasing the number N if the modeled time instance corresponds to the least recent time instance among the N previous time instances. Increases / decreases of N may be by increments / decrements of 1, for example. It is understood that the steps of increasing / decreasing N and corresponding checks may be performed in an order to avoid N being frozen to zero.

[0031] Thereby, memory usage can be reduced to a minimum amount necessary for realistically modeling an audio object at a certain velocity.

[0032] In some embodiments, the method may further include receiving a bitstream comprising the object-based audio content. The method may yet further include determining, based on a bitstream parameter included in the bitstream, that a realistic rendering mode is to be used for rendering the object-based audio content that models the offset between auditory and visual positions of the audio object.

[0033] According to another aspect, an apparatus is provided. The apparatus may include one or more processors and a memory coupled thereto and storing instructions for the one or more processors. The one or more processors may be configured to perform the methods or method steps outlined throughout the present disclosure. This apparatus may relate to or be included in an audio decoder or an audio Tenderer, for example as part of computer-mediated reality, VR, AR, MR, or XR equipment, such as headset or goggles, as the case may be. According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., one or more processors).

[0034] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device (e.g., one or more processors) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.

[0035] It should be noted that the methods and apparatus including their preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.

[0036] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa.

[0037] Brief Description of the Drawings

[0038] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein

[0039] Fig. 1A schematically illustrates an example of conventional audio scene processing;

[0040] Fig. IB schematically illustrates an example of audio scene processing according to embodiments of the disclosure;

[0041] Fig. 2 shows an example of program code for conventional audio scene processing;

[0042] Fig. 3 schematically shows an example of determining a rendering position of a fast-moving audio object according to conventional techniques; Fig. 4 is a flowchart schematically illustrating an example of a method of rendering object-based audio content according to embodiments of the disclosure;

[0043] Fig. 5 schematically shows an example of determining a rendering position of a fast-moving audio object according to embodiments of the disclosure;

[0044] Fig. 6A, Fig. 6C, and Fig. 6C are flowcharts schematically illustrating examples of determining perceptual relevance of calculated rendering positions according to embodiments of the disclosure;

[0045] Fig. 7 is a flowchart schematically illustrating an example of memory management according to embodiments of the disclosure;

[0046] Fig. 8 is a flowchart schematically illustrating an example of bitstream processing according to embodiments of the disclosure; and

[0047] Fig. 9 schematically illustrates an apparatus suitable for implementing techniques according to embodiments of the disclosure.

[0048] Detailed Description

[0049] In the following, example embodiments of the disclosure will be described with reference to the appended figures. Identical elements in the figures may be indicated by identical reference numbers, and repeated description thereof may be omitted.

[0050] Overview

[0051] The present invention, broadly speaking, relates to rendering enhancement for auditory position modelling of fast-moving audio sources. In particular, it provides a solution for calculation of audio object positions for audio rendering, such as MPEG-I rendering, that results in physically consistent (e.g., delayed) auditory audio source localization.

[0052] An example 100B of determining a delayed auditory source location according to embodiments of the disclosure is schematically illustrated in Fig. IB. Here, an audio source 20 at a current rendering position O(tr) (which may be substantially identical to the visual position of the audio source 20) is rendered to the listener (or observer) 10 at listener position L from a delayed, or modeled, audio object position OdelayedCT'1)' possibly applying appropriate distance attenuation. In more detail, the present disclosure provides techniques for efficiently calculating the aforementioned modeled audio object position, using stored positions (e.g., coordinates) of audio sources, including fast-moving audio sources, and selection of corresponding positions for the audio rendering of delayed auditory audio source(s) from among the stored positions. Therein, the storing of positions may be velocity adaptive, for example with more previous positions being stored the faster the audio source moves.

[0053] The proposed techniques support runtime audio source position updates (e.g., do not require to have access to audio source position trajectory information in advance) as well as psycho-acoustically motivated means to control their application (e.g., the proposed processing may only be activated if its effect is perceptually significant).

[0054] Description of Example Embodiments

[0055] Techniques proposed by the present disclosure ensure that the modeled rendering positions for a fast-moving audio object are always located on the actual audio object trajectory (or its low-pass filtered approximation).

[0056] In some implementations, the proposed techniques may replace the “getLocation()” function of MPEG-I audio rendering by an alternative implementation.

[0057] Nomenclature and Notation

[0058] The following definitions and nomenclature may be used in the context of the present disclosure: t time

[0059] tj* time of the current function call (e.g., current time or current time instance,

[0060] “now”)

[0061] tt*_j time of previous function calls where j = 1,... (e.g., previous time instance) tt*-! preceding function call (e.g., preceding time instance)

[0062] O(t) “visual” position of the audio object (e.g., substantially equal to the actual position of the audio object)

[0063] Oauditory(f) physically accurate auditory position of audio object (e.g., “ideal” position for audio rendering) OdeiayedCO approximation of the auditory position of audio object (e.g., “modelled” position for audio rendering)

[0064] L(t) listener / observer position

[0065] t delay = d / s delay time (or time offset)

[0066] s speed of sound

[0067] d distance

[0068] DLPS(d function providing (potentially low-pass filtered) approximation of the distance data

[0069] OLPS-(O) function providing (potentially low-pass filtered) approximation of the positional data

[0070] r>(Oi, O2) function providing a measure of a (potentially low-pass filtered) distance between positions Oxand O2, such as a distance metric

[0071] N number of stored previous trajectory coordinate points

[0072] Problem Formulation

[0073] The present disclosure seeks to determine (e.g., calculate or estimate) an approximation of the auditory position

[0074]

[0075] where

[0076]

[0077] fdelay d / s.

[0078] (3) In the above, the delay time t^eiaycanbe approximated for example as

[0079]

[0080] When using low-pass filtering for the distance determination, this may amount to

[0081] fdelay

[0082]

[0083] or the simplest case

[0084]

[0085] Proposed Algorithm

[0086] Fig. 4 shows a flowchart that illustrates an example of a method 400 of rendering object-based audio content according to embodiments of the disclosure. Method 400 determines a modeled (or delayed) rendering position that approximates the aforementioned ideal auditory position.

[0087] At step S410, a rendering position of an audio object (e.g., moving audio object, in particular fast-moving audio object) is determined at each of a plurality of time instances (rendering time instances). Further, a list (or set) of one or more known rendering positions of the audio object including a current rendering position (e.g., O(t_{i*})) at a current time instance (e.g., t_{i*}) and a number N of previous rendering positions at respective previous time instances is maintained (e.g., stored). Here, N is a non-negative integer, i.e., N ∈ ℕ₀⁺. Accordingly, N + 1 known time instances may be maintained (stored) in total.

[0088] Audio rendering at step S410 may be performed at each of a plurality of equi-spaced time instances, i.e., at a given rate or frequency.

[0089] At step S420, for the current time instance, a measure of a time offset (or delay time) is determined based on the current rendering position and a rendering position of the audio object at the preceding time instance (e.g., O(t_{i*-1}) at preceding time instance t_{i*-1}). It isunderstood that the preceding time instance may relate to the immediately preceding time instance to the current time instance. At step S430, for the current time instance, a modeled time instance is determined that corresponds to one of the known rendering positions and that is closest in time to a timing that precedes the current time instance (e.g., t_{i*}) by the time offset (e.g., t_{delay}). At step S440, for the current time instance, a modeled rendering position (e.g.,

[0090]

[0091] of the audio object is determined based on a rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance, for modeling an offset between auditory and visual positions of the audio object.

[0092] As described above, the modeled rendering position determined at this step may account for different propagation speeds of light and sound in a (virtual) environment, and accordingly for an offset or delay between the visual object position and the auditory object position.

[0093] At step S450, which is an optional step, the time offset is set to zero if the modeled time instance is identical to the current time instance. Further, the modeled rendering position is set to the current rendering position. This may amount to disregarding the modeled rendering position in rendering the audio object at the current time instance.

[0094] In an implementation example, the following specific processing steps 1) to 4), of which processing steps 3) and 4) may be optional, may be performed in the context of method 400 for each function call time (or time instance) t_{i*}.

[0095] 1) Calculate delay time (or time offset) t_{delay} as defined above in Eq. (4), (5), or (6). This may correspond to or implement step S420 of method 400.

[0096] Further, calculate index k of the closest available time-stamp Akfrom among all N + 1 stored times

[0097]

[0098] example via

[0099]

[0100] This may correspond to or implement step S430 of method 400.

[0101] In line with the above, determining the measure of the time offset at step S420 may be based on a value representing the speed of sound and a measure of a distance derived from the current rendering position O(tf) and the rendering position O(tr-i) f°rthe preceding time instance. Further, the time offset t_{offset} = t_{delay} at step S420 may be determined as

[0102]

[0103] where t_{i*} is the current time instance,

[0104]

[0105] is the preceding time instance, O(t_{i*}) is the current rendering position, O(t_{i*-1}) is the rendering position of the audio object at the preceding time instance, s is (the value for) the speed of sound, and £)(•,•) is the distance function (e.g., distance metric).

[0106] Further in line with the above, the modeled time instance t_{i*-k} may be determined at step S430 via

[0107]

[0108] Once the modeled time instance t_{i*-k} is determined, it may be used for determining the modeled rendering position O_{delayed}(t_{i*}). As noted above, determining the modeled rendering position may involve at least one of low-pass filtering based on known rendering positions, applying curve fitting based on known rendering positions, and interpolation based on known rendering positions, for example via

[0109] O_{delayed}(

[0110]

[0111] In the simplest case however, the modeled rendering position may be identical to the rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance, i.e.,

[0112] O_{delayed}(

[0113]

[0114] 2) Check whether the audio object is moving fast and the auditory position modelling as per processing step 1) shall be applied. This may correspond to or implement optional step S450 of method 400.

[0115] Specifically, if k == 0, it may be concluded that the delay time t_{delay} is negligibly small and can be neglected,

[0116] f delay ~0,

[0117] (ID and auditory position modeling can be omitted,

[0118]

[0119] Otherwise, if k > 1, it may be concluded that the delay time t_{delay} is considerably large. Then, t_{delay} can be approximated as

[0120]

[0121] and the following approximation applies,

[0122]

[0123] 3) Optionally, perform a perceptual relevance check to determine whether the calculated auditory position approximation result is perceptually relevant for the listener (or observer) at listener position L, possibly in dependence on the listener’s orientation (e.g., head orientation). For this, any one or any combination of the implementation examples of the perceptual relevance check described below may be used.

[0124] In a first implementation example of the perceptual relevance check, the observer-relative angle between visual and auditory positions (Odeiayed‘^‘0 for fr, i.e., angle (Oc|e|ayec|(t(»)LO(t(»))j is considered. If the observer-relative angle is greater than a given fixed threshold (e.g., only if this is the case), then the auditory position approximation is calculated as per Eq. (9) or Eq. (10). In line with the above, method 400 may optionally include the steps of method 600A shown in the flowchart of Fig. 6A for performing the perceptual relevance check.

[0125] At step S610A, a relative angle between the current rendering position O(tr) and the modeled rendering position Odeiayed(T‘) inrelation to a user position L is determined. The user may be a listener and / or observer, located in the audio scene in question.

[0126] At step S620A, the determined relative angle is compared to a threshold (e.g., predetermined threshold) for the relative angle. At step S630A, the modeled rendering position O_{delayed}(t_{i*}) is used for rendering the audio object at the current time instance t_{i*} if (e.g., only if) the threshold for the relative angle does not exceed the relative angle. If the threshold exceeds the relative angle (or, possibly, is equal to the relative angle), the current rendering position O(t_{i*}) may be used for rendering the audio object at the current time instance

[0127]

[0128] .

[0129] In a second implementation example of the perceptual relevance check, the observer-relative azimuth and elevation angles between visual and auditory positions (Odeiayed‘^‘0 for ff)316considered. If (e.g., only if) the observer-relative azimuth and elevation angles are greater than a given fixed threshold (or greater than respective fixed thresholds), then the auditory position approximation is calculated as per Eq. (9) or Eq. (10).

[0130] In line with the above, method 400 may optionally include the steps of method 600B shown in the flowchart of Fig. 6B for performing the perceptual relevance check.

[0131] At step S610B, a relative azimuth angle and a relative elevation angle between the current rendering position O(tr) and the modeled rendering position O_{delayed}(t_{i*}) in relation to the user position are determined.

[0132] At step S620B, the determined relative azimuth angle is compared to a threshold for the relative azimuth angle and the determined relative elevation angle is compared to a threshold for the relative elevation angle.

[0133] At step S630B, the modeled rendering position O_{delayed}(t_{i*}) is used for rendering the audio object at the current time instance

[0134]

[0135] if (e.g., only if) the threshold for the relative azimuth angle does not exceed the relative azimuth angle and the threshold for the relative elevation angle does not exceed the relative elevation angle. If one or both thresholds exceed the respective relative angle (or, possibly, are equal to the respective relative angle), the current rendering position O(t_{i*}) may be used for rendering the audio object at the current time instance t_{i*}.

[0136] Here, the threshold for the relative elevation angle may be larger than the threshold for the relative azimuth angle, accounting for different positioning sensitivities of the human auditory system in azimuth and elevation. In some implementations however, the two thresholds may be identical, for reasons of simplicity. In a third implementation example of the perceptual relevance check, the observer-relative angle between visual and auditory positions (Odeiayed‘^‘0 for ff) 'sconsidered, as for the first implementation example of the perceptual relevance check. This time however, the threshold is implemented by a given threshold function that depends on at least one of a listener orientation (e.g., listener head orientation) and a spatial localization listening sensitivity pattern. That is, for a given listener orientation (e.g., listener head orientation) and / or spatial localization listening sensitivity pattern, the threshold for the observer-relative angle between visual and auditory positions may be given by the respective threshold function value of the threshold function.

[0137] In line with this implementation example, method 600A described above may further include determining the threshold for the relative angle based on an orientation (e.g., head orientation) of the user and / or based on a spatial localization listening sensitivity pattern.

[0138] In a fourth implementation example of the perceptual relevance check, an observer-relative distance difference (e.g.,

[0139]

[0140] for t_{i*}) considered. If (e.g., only if) the observer-relative distance difference is greater than a given fixed threshold, then the auditory position approximation is calculated as per Eq. (9) or Eq. (10).

[0141] In line with the above, method 400 may optionally include the steps of method 600C shown in the flowchart of Fig. 6C for performing the perceptual relevance check.

[0142] At step S610C, a relative distance is determined based on the current rendering position O(tj*), the modeled rendering position O_{delayed}(t_{i*}), and the user position L. The relative distance may be defined as (the absolute value of) the difference between absolute values of distances between the current rendering position O(tf) and the user position L and between the modeled rendering position O_{delayed}(t_{i*}) and the user position L, respectively, e.g. via

[0143]

[0144] At step S620C, the determined relative distance is compared to a threshold for the relative distance.

[0145] At step S630C, the modeled rendering position Odeiayed(T'‘) f°rrendering the audio object at the current time instance t_{i*} is used if (e.g., only if) the threshold for the relative distance does not exceed the relative distance. If the threshold exceeds the relative distance (or, possibly, is equal to the relative distance), the current rendering position O(tj*) may be used instead for rendering the audio object at the current time instance t_{i*}.

[0146] If more than one or a combination of the above implementation examples and methods is applied, the modeled rendering position O_{delayed}(t_{i*}) shall be used only if all relevant conditions are satisfied.

[0147] 4) Optionally, update the set of stored audio object trajectory coordinate points (or known rendering positions), for example as

[0148] {O(t_{i*-j})}, j = 0, ..., N.

[0149] (16) Here, if k == N (i.e., if the modeled rendering position corresponds to the earliest (oldest) stored / known rendering position), then the set of stored trajectory coordinate points may be extended, e.g.,

[0150] N = N + 1.

[0151] (17) On the other hand, if k == 0 (i.e., if the modeled rendering position corresponds to the current rendering position), then the set of stored trajectory coordinate points may be reduced, e.g.,

[0152] N = N - 1.

[0153] (18) In line with the above, method 400 may optionally include the steps of method 700 shown in the flowchart of Fig. 7. At any given time instance, a number of N + 1 known rendering positions including the current rendering position and N previous rendering positions are maintained. The number N may be adjusted as per steps S710 and S720, depending on a velocity of the audio object. Thereby, memory consumption for maintaining the known rendering positions can be optimized, while still being able to provide a perceptually accurate rendering of the audio object. For a fast-moving audio object, the number N will be (successively) increased by step S720, while it while be (successively) decreased for a slow-moving audio object by step S710. An initial value of N (specific to the audio object) can be set (and then limited) by a maximum modelled auditory position delay threshold (e.g., expressed in the time domain). Alternatively, N may be initialized to N = 0, so that memory of known rendering positions will be successively built up only for sufficiently fast-moving audio objects.

[0154] At step S710, the number N is decreased if the modeled time instance t_{i*-k} corresponds to the current time instance tt* (i.e., k == 0).

[0155] At step S720. the number N is increased if the modeled time instance

[0156]

[0157] corresponds to the least recent time instance t_{i*-N} among the N previous time instances (i.e., k == N).

[0158] In the above, increases / decreases of N may be by increments / decrements of 1, for example. It is understood that the steps of increasing / decreasing N and corresponding checks may be performed in an order to avoid N being frozen at zero.

[0159] Examples of Control Parameters

[0160] In some implementations, activation of the proposed algorithm may be performed based on a bitstream variable or / and renderer interface switch, e.g.:

[0161] • “Realistic mode” - proposed algorithm is applied

[0162] • “Artistic mode” - proposed algorithm is not applied

[0163] Accordingly, method 400 may optionally comprise the steps of method 800 shown in the flowchart of Fig. 8, for execution before step S410.

[0164] At step S810, a bitstream comprising the object-based audio content is received.

[0165] At step S820, it is determined, based on a bitstream parameter included in the bitstream, that a realistic rendering mode is to be used for rendering the object-based audio content that models the offset between auditory and visual positions of the audio object.

[0166] Alternatively, the decision at step S820 may be made based on a renderer interface input. Examples of further control parameters, apart from the bitstream variable and / or Tenderer interface switch may include:

[0167] • N (e.g., N = 100 - which corresponds to ca. 3 sec for 30 ms fixed update rate) • threshold(s) for observer relative angle and distance (e.g., 4° and 10m - which corresponds to ca. 30 ms) Apparatus, Programs, and Recording Media

[0168] While methods according to the present disclosure have been described above, it is understood that the present disclosure likewise relates to apparatus (e.g., computer apparatus or apparatus having processing capability in general) and systems for implementing these methods (or techniques in general).

[0169] An example of such apparatus 900 is schematically illustrated in Fig. 9. The apparatus 900 comprises a processor 910 (or multiple processors) and a memory 920 coupled to the processor 910. The memory 920 may store instructions for execution by the processor 910.

[0170] Processor 910 may be adapted to implement the apparatus described throughout the disclosure and / or to perform methods (e.g., methods of rendering object-based audio content) described throughout the disclosure. The apparatus 900 may receive inputs 930 (e.g., object-based audio content, a bitstream including object-based audio content, control parameters, etc.) and generate outputs 940 (e.g., modeled rendering positions, rendered audio content, etc.) as described throughout the disclosure. Accordingly, the apparatus 900 may relate to any of an audio renderer, an audio decoder, an apparatus included in an audio renderer or audio decoder, or an apparatus including an audio renderer or audio decoder (e.g., computer-mediated reality, VR, AR, MR, or XR equipment, such as headset or goggles, as the case may be.

[0171] The present disclosure further relates to programs (e.g., computer programs) comprising instructions that, when executed by a processor (or multiple processors), cause the processor (or multiple processors) to carry out any of the methods described throughout the disclosure, and to computer-readable storage media storing such programs.

[0172] Interpretation

[0173] Aspects of the techniques (e.g., methods, apparatus, systems) described herein may be implemented in an appropriate computer-based processing network environment (e.g., server or cloud environment) for processing digital or digitized files (e.g., digital or digitized audio files). Portions of such systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof. One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics.

[0174] Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.

[0175] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic -based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, computer-implemented neural networks described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.

[0176] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.

[0177] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.

[0178] Enumerated. Example Embodiments

[0179] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.

[0180] EEE1. A method of rendering object-based audio content, comprising:

[0181] determining a rendering position of an audio object at each of a plurality of time instances and maintaining a list of one or more known rendering positions of the audio object including a current rendering position at a current time instance and a number N of previous rendering positions at respective previous time instances, where N is a non-negative integer;

[0182] for the current time instance, determining a measure of a time offset based on the current rendering position and a rendering position of the audio object at the preceding time instance; determining a modeled time instance that corresponds to one of the known rendering positions and that is closest in time to a timing that precedes the current time instance by the time offset; and

[0183] determining a modeled rendering position of the audio object based on a rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance, for modeling an offset between auditory and visual positions of the audio object.

[0184] EEE2. The method according to EEE1, wherein determining the measure of the time offset is based on a value representing the speed of sound and a measure of a distance derived from the current rendering position and the rendering position for the preceding time instance.

[0185] EEE3. The method according to EEE1 or EEE2, wherein the time offset toffsetisdetermined as toffset=

[0186]

[0187] the current time instance, t,*-! is the preceding time instance, O(tf) is the current rendering position, O(ti*-i) is the rendering position of the audio object at the preceding time instance, s is a value for the speed of sound, and £>(•,•) is a distance function. EEE4. The method according to EEE3, wherein the modeled time instance t_{i*-k} is determined via

[0188]

[0189] EEE5. The method according to any one of the preceding EEEs, further comprising setting the time offset to zero if the modeled time instance is identical to the current time instance, and setting the modeled rendering position to the current rendering position.

[0190] EEE6. The method according to any one of the preceding EEEs, wherein determining the modeled rendering position involves at least one of low-pass filtering based on known rendering positions, applying curve fitting based on known rendering positions, and interpolation based on known rendering positions.

[0191] EEE7. The method according to any one of the preceding EEEs, further comprising: determining a relative angle between the current rendering position and the modeled rendering position in relation to a user position;

[0192] comparing the determined relative angle to a threshold for the relative angle; and

[0193] using the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative angle does not exceed the relative angle.

[0194] EEE8. The method according to EEE7, further comprising determining the threshold for the relative angle based on an orientation of the user.

[0195] EEE9. The method according to any one of the preceding EEEs, further comprising: determining a relative azimuth angle and a relative elevation angle between the current rendering position and the modeled rendering position in relation to a user position;

[0196] comparing the determined relative azimuth angle to a threshold for the relative azimuth angle and comparing the determined relative elevation angle to a threshold for the relative elevation angle; and

[0197] using the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative azimuth angle does not exceed the relative azimuth angle and the threshold for the relative elevation angle does not exceed the relative elevation angle. EEE10. The method according to any one of the preceding EEEs, further comprising: determining a relative distance based on the current rendering position, the modeled rendering position, and the user position; comparing the determined relative distance to a threshold for the relative distance; and using the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative distance does not exceed the relative distance.

[0198] EEE11. The method according to any one of the preceding EEEs, wherein a number of N + 1 known rendering positions including the current rendering position and N previous rendering positions are maintained; and

[0199] wherein the method further comprises:

[0200] decreasing the number N if the modeled time instance corresponds to the current time instance; and

[0201] increasing the number N if the modeled time instance corresponds to the least recent time instance among the N previous time instances.

[0202] EEE12. The method according to any one of the preceding EEEs, further comprising: receiving a bitstream comprising the object-based audio content; and

[0203] determining, based on a bitstream parameter included in the bitstream, that a realistic rendering mode is to be used for rendering the object-based audio content that models the offset between auditory and visual positions of the audio object.

[0204] EEE13. An apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of EEEl to EEE12.

[0205] EEE14. A computer program including instructions that when executed by one or more processors, cause the one or more processors to perform the method according to any one of EEE1 to EEE12.

[0206] EEE15. A computer-readable storage medium storing the computer program according to EEE14.

Claims

CLAIMS1. A method of rendering object-based audio content, comprising:determining a rendering position of an audio object at each of a plurality of time instances and maintaining a list of known rendering positions of the audio object including a current rendering position at a current time instance and a number N of previous rendering positions at respective previous time instances, where N is a non-negative integer;for the current time instance, determining a measure of a time offset based on the current rendering position and a rendering position of the audio object at the preceding time instance; determining a modeled time instance that corresponds to one of the known rendering positions and that is closest in time to a timing that precedes the current time instance by the time offset; anddetermining a modeled rendering position of the audio object based on a rendering position of the audio object among the known rendering positions that corresponds to the determined modeled time instance, for modeling an offset between auditory and visual positions of the audio object.

2. The method according to claim 1, wherein determining the measure of the time offset is based on a value representing the speed of sound and a measure of a distance derived from the current rendering position and the rendering position for the preceding time instance.

3. The method according to claim 1 or 2, wherein the time offset toffsetisdetermined as toffset=the current time instance, t,*-! is the preceding time instance, O(tf) is the current rendering position, O(tr-i) is the rendering position of the audio object at the preceding time instance, s is a value for the speed of sound, and £>(•,•) is a distance function.

4. The method according to claim 3, wherein the modeled time instance t_{i*-k} is determined via5. The method according to any one of the preceding claims, further comprising setting the time offset to zero if the modeled time instance is identical to the current time instance, and setting the modeled rendering position to the current rendering position.

6. The method according to any one of the preceding claims, wherein determining the modeled rendering position involves at least one of low-pass filtering based on known rendering positions, applying curve fitting based on known rendering positions, and interpolation based on known rendering positions.

7. The method according to any one of the preceding claims, further comprising: determining a relative angle between the current rendering position and the modeled rendering position in relation to a user position;comparing the determined relative angle to a threshold for the relative angle; and using the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative angle does not exceed the relative angle.

8. The method according to claim 7, further comprising determining the threshold for the relative angle based on an orientation of the user.

9. The method according to any one of the preceding claims, further comprising: determining a relative azimuth angle and a relative elevation angle between the current rendering position and the modeled rendering position in relation to a user position;comparing the determined relative azimuth angle to a threshold for the relative azimuth angle and comparing the determined relative elevation angle to a threshold for the relative elevation angle; andusing the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative azimuth angle does not exceed the relative azimuth angle and the threshold for the relative elevation angle does not exceed the relative elevation angle.

10. The method according to any one of the preceding claims, further comprising:determining a relative distance based on the current rendering position, the modeled rendering position, and the user position;comparing the determined relative distance to a threshold for the relative distance; and using the modeled rendering position for rendering the audio object at the current time instance only if the threshold for the relative distance does not exceed the relative distance.

11. The method according to any one of the preceding claims, wherein a number of N + 1 known rendering positions including the current rendering position and N previous rendering positions are maintained; andwherein the method further comprises:decreasing the number N if the modeled time instance corresponds to the current time instance; andincreasing the number N if the modeled time instance corresponds to the least recent time instance among the N previous time instances.

12. The method according to any one of the preceding claims, further comprising: receiving a bitstream comprising the object-based audio content; anddetermining, based on a bitstream parameter included in the bitstream, that a realistic rendering mode is to be used for rendering the object-based audio content that models the offset between auditory and visual positions of the audio object.

13. An apparatus comprising one or more processors and a memory coupled thereto, wherein the one or more processors are configured to perform the method according to any one of claims 1 to 12.

14. A computer program including instructions that when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 12.

15. A computer-readable storage medium storing the computer program according to claim 14.

Citation Information

Patent Citations

  • Active acoustics control for near- and far-field sounds

    EP4093058A1

  • Structural Modeling of the Head Related Impulse Response

    US20170094440A1