Method, apparatus, and system for controlling Doppler effect modeling

The method addresses the limitations of conventional Doppler effect modeling in 6DoF environments by using a predefined function to map relative velocities to pitch coefficient corrections, enhancing the listening experience through consideration of signal processing unit capabilities and content creator intent.

JP2026041729APending Publication Date: 2026-03-10DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Conventional methods for modeling the Doppler effect in six degrees of freedom (6DoF) environments, such as virtual reality (VR) and augmented reality (AR), fail to consider the capabilities of the underlying signal processing unit for pitch coefficient correction and do not allow control of the pitch coefficient correction value according to the content creator's intention, leading to an unsatisfactory listening experience.

Method used

A method and apparatus that utilize a predefined pitch coefficient correction function to map relative velocities to pitch coefficient correction values, taking into account the signal processing unit's capabilities and the content creator's intent, allowing flexible and efficient Doppler effect modeling.

Benefits of technology

Enables efficient and flexible Doppler effect modeling in 6DoF environments by accounting for signal processing unit capabilities and content creator intentions, improving the perceived listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041729000001_ABST
    Figure 2026041729000001_ABST
Patent Text Reader

Abstract

Methods, apparatus, and systems are provided for controlling Doppler effect modeling when rendering audio content for 6DoF environments, such as virtual reality and / or augmented reality environments. A method for modeling Doppler effects includes obtaining first parameter values ​​for one or more first parameters indicating an acceptable range of pitch coefficient correction values, obtaining second parameter values ​​for a second parameter indicating a desired strength of the Doppler effect to be modeled, determining a pitch coefficient correction value based on a relative velocity between a listener and a sound source in audio content and the first and second parameter values ​​using a predetermined pitch coefficient correction function, and rendering the sound source based on the pitch coefficient correction value. The predetermined pitch coefficient correction function has the first and second parameters and is a function for mapping the relative velocity to the pitch coefficient correction value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 273,185 (Reference No. D21092USP1), filed October 29, 2021.

[0002] The present disclosure relates to the general field of Doppler effect modeling, and more particularly to methods and apparatus for controlling Doppler effect modeling, for example, for use in virtual reality or augmented reality environments. [Background technology]

[0003] Generally speaking, the term Doppler effect (or Doppler shift) is typically used to refer to an audio effect experienced when there is a change in the frequency of a wave (e.g., a sound wave) relative to an observer (e.g., a listener) who is moving relative to the wave source (e.g., a sound source). More specifically, the Doppler effect may generally be perceived as a higher pitch (a commonly used measure of frequency or perceived frequency) as the wave source (e.g., an emergency vehicle siren) approaches the observer, and a lower pitch as the wave source passes and moves further away.

[0004] Today, the Doppler effect is beginning to be considered as an important aspect of audio rendering of dynamic scenes in six degrees of freedom (6DoF) environments, which are widely adopted, for example, in virtual reality (VR) and / or augmented reality (AR) scenarios (e.g., games). Broadly speaking, within the context of audio processing (e.g., rendering), the Doppler effect may generally be modeled using an audio pitch coefficient correction.

[0005] Some conventional implementations generally propose performing Doppler effect modeling based on a physical description or approximation of the Doppler effect. However, such approaches generally do not have the means or ability to consider the capabilities of the underlying signal processing unit for pitch coefficient correction (e.g., high magnitude of pitch coefficient correction value due to high relative velocity, singularity, etc.), nor the means or ability to control the pitch coefficient correction value (i.e., representing the strength of the Doppler effect) according to the content creator's intention (or in other words, the subjective listening experience).

[0006] More specifically, there is a need for techniques to perform (control) Doppler effect modeling when rendering audio content for 6DoF environments, such as virtual reality and / or augmented reality environments. Summary of the Invention

[0007] In view of the above, the present disclosure generally provides a method for modeling the Doppler effect when rendering audio content for a six degrees of freedom (6DoF) environment, a method for encoding parameters for use in modeling the Doppler effect when rendering audio content for a 6DoF environment, and corresponding audio renderers, encoders, programs and computer-readable storage media having the features of the respective independent claims.

[0008] According to a first aspect of the present invention, there is provided a method for modeling the Doppler effect when rendering audio content for a 6DoF environment, which method may be performed at the user side, or in other words, in the user (decoding) side environment.

[0009] In particular, the method may comprise obtaining first parameter values ​​for one or more first parameters indicating an acceptable range of pitch coefficient correction values. The acceptable range of pitch coefficient correction values ​​may be indicated, for example, by using an upper (e.g., boundary) and / or a lower (e.g., boundary). The method may further include obtaining a second parameter value for a second parameter indicating a desired strength (or sometimes referred to as "aggressiveness") of the Doppler effect to be modeled. The method may further include determining the pitch coefficient correction values ​​based on a relative velocity between the listener and a sound source in the audio content and the first and second parameter values ​​using a predetermined pitch coefficient correction function. In particular, the predetermined pitch coefficient correction function may have the first and second parameters (or, in other words, may take, inter alia, the first and second parameters as inputs) and may be a function for mapping the relative velocity to the pitch coefficient correction value. In particular, as will be understood and appreciated by those skilled in the art, the pitch coefficient correction value may be considered to be a value (perhaps expressed in any suitable format) commonly used to appropriately correct (e.g., shift) pitch, thereby enabling appropriate and proper modeling of Doppler effects and rendering of audio content in a 6DoF environment. Finally, the method may include rendering the sound source based on the (determined) pitch coefficient correction value.

[0010] In other words, in broad terms, the present disclosure generally proposes a method that utilizes a predefined (or predetermined / pre-implemented) pitch factor correction function to map relative velocities to corresponding pitch factor correction values ​​to model Doppler effects (e.g., when rendering audio content in a 6DoF environment). As will be described in more detail below, such a predefined pitch factor correction function may be implemented by any suitable means, subject to certain requirements (or characteristics) that generally need to be met. Specifically, the predefined pitch factor correction function may have multiple parameters (or, in other words, may receive multiple parameters as input), including (at least) a first parameter indicating an acceptable range of pitch factor correction values ​​and a second parameter indicating a desired strength (aggressiveness) of the Doppler effects to be modeled. Then, in practice, when the proposed method is executed, for example, by an audio rendering device (or simply referred to as an audio renderer) on which the predefined pitch factor correction function has been deployed, the audio rendering device may be configured to obtain first and second parameter values ​​corresponding to the first and second parameters, respectively. Obtaining the first and second parameter values ​​may be performed by any suitable means depending on various requirements and / or implementations. For example, in some possible cases, the first and second parameter values ​​may be derived (or simply extracted) from a bitstream received from an encoding device, or in some other possible cases, may be obtained (or simply read) from a file or a look-up table (LUT). In this manner, particularly by applying a pitch coefficient correction function, a pitch coefficient correction value for modeling the Doppler effect may be determined based on the relative velocity between the listener and the sound source and based on the first and second parameter values.

[0011] Configured as described above, the proposed method can provide an efficient and flexible mechanism for performing (e.g., controlling) Doppler effect modeling when rendering audio content for a 6DoF environment, while taking into account both the (acceptable or tolerable) capabilities of the underlying signal processing unit (e.g., of the audio renderer) for pitch coefficient modification (e.g., high magnitude of pitch coefficient modification values ​​for high relative velocities, singularities, etc.) and the possibility to control the (desired) pitch coefficient modification values ​​(e.g., representing the desired strength / aggressiveness of the Doppler effect to be modeled) according to the content creator's intention (in other words, the subjective listening experience), thereby improving the perceived listening experience (at the listener's side, e.g., a user playing a game in a VR environment).

[0012] Furthermore, because the pitch coefficient correction function is already pre-implemented as a predefined function that takes, among other things, values ​​of both the first and second parameters as inputs, it is generally not necessary to redesign (or re-implement) a new pitch coefficient correction function every time rendering conditions change (e.g., a different renderer with different processing capabilities is deployed, different audio content created by a different author and / or for a different scene, etc.). Rather, it is generally only necessary to communicate (e.g., using an encoded bitstream on the encoding side) different first and second parameter values ​​that respectively represent corresponding allowable ranges (e.g., limits) of pitch coefficient correction values ​​and corresponding desired strengths of the Doppler effect to be modeled. Thus, in some possible scenarios, the pre-defined pitch coefficient correction function may be implemented as simply as a plug-in or library (that takes the first parameter and the second parameter as inputs for modeling a desired Doppler effect), which may be deployed in various software and / or platforms according to different requirements and / or implementation forms and further customized if necessary. This avoids unnecessary redesign / reimplementation of pitch coefficient modification functions and further improves efficiency in the overall audio rendering process.

[0013] In some example implementations, the relative velocity may be calculated based on the positions (e.g., relative positions) of the listener and the sound source. For example, in some possible cases, the relative velocity may be determined from the rate of change of the relative distance between the sound source and the listener (e.g., by taking the first derivative) based on their respective positions. Of course, any other suitable means may be employed to obtain or calculate the relative velocity, as will be understood and appreciated by those skilled in the art.

[0014] In some example implementations, the one or more first parameters may comprise parameters indicating upper and / or lower bounds of an acceptable range of pitch coefficient correction values.

[0015] In some example implementations, the acceptable range of pitch coefficient correction values ​​may reflect the processing capabilities of an audio renderer that renders the audio content. That is, broadly speaking, the first parameter indicating the acceptable range (e.g., upper and / or lower bounds) of pitch coefficient correction values ​​can be viewed in one respect as representing the range of processing capabilities supported by a rendering device (e.g., an audio renderer), or more precisely, by an underlying processing unit of that rendering device for modeling the Doppler effect.

[0016] In some example implementations, if the thus obtained acceptable range of pitch coefficient correction values ​​cannot be supported by (the processing unit underlying) an audio renderer, a default range of pitch coefficient correction values ​​may be used by the audio renderer. An illustrative example of such a scenario may be a scenario of a mobile device (e.g., a mobile phone) with (relatively) limited processing (rendering) capabilities that obtains (e.g., receives) a range of pitch coefficient correction values ​​that was initially configured (e.g., by an encoding device) to target (relatively) more powerful rendering devices (e.g., game consoles or professional workstations). In such a scenario, to avoid unexpected or adversely affecting the rendering process, it may be considered more practical for the mobile device to apply a default range parameter setting (e.g., falling within the wider range initially obtained), which may be set by the mobile device's manufacturer to more accurately reflect the mobile device's actual processing (rendering) capabilities, instead of using the obtained unsupported parameter value.

[0017] In some example implementations, the second parameter may control the slope (or, in some possible cases, also referred to as "strength") of the pitch coefficient correction function, which may be considered to reflect the aggressiveness of the Doppler effect to be modeled.

[0018] In some example implementations, audio content may be extracted from a received bitstream. The bitstream may have been encoded by an encoding device, for example, in any suitable format by using any suitable means. Thus, according to various implementations, the first and second parameter values ​​may be derived (e.g., extracted, decoded, etc.) from instructions included in the bitstream. In some possible cases, the indications of the first and second parameter values ​​may be encoded as labels (or fields) in the bitstream, as will be understood and appreciated by those skilled in the art.

[0019] Of course, in some other possible implementations, it is also possible that the audio content and the first and second parameters may be obtained separately (eg, from two separate bitstreams).

[0020] In some exemplary implementations, the second parameter value may be set by a content creator of the audio content. In particular, the second parameter value may be set by the content creator of the audio content according to the intentions of the content creator. Thus, in a broad sense, the second parameter value may be considered to reflect a subjective listening experience targeted by (and controlled by) the content creator.

[0021] In some example implementations, the second parameter value may be set by modeling real-world standards and / or artistic expectations for a desired Doppler effect strength. Of course, any other suitable implementations for determining and setting the second parameter value may be possible, as will be understood and appreciated by those skilled in the art. In some example implementations, rendering the audio content based on the pitch coefficient modification value may comprise adjusting the pitch of a sound source in the audio content based on the pitch coefficient modification value.

[0022] In some example implementations, a positive pitch coefficient modification value may generally indicate an increase in the pitch of the sound source. Similarly, a negative pitch coefficient modification value may generally indicate a decrease in the pitch of the sound source.

[0023] In some example implementations, the pitch adjustment of the sound source may be performed in units of semitones. For example, a pitch coefficient modification value of 2 may simply mean increasing the pitch of the sound source by 2 semitones, and correspondingly, a pitch coefficient modification value of −2 may simply mean decreasing the pitch of the sound source by 2 semitones.

[0024] In some example implementations, the pitch coefficient correction function may be implemented based on a generalized logistic function. That is, implementing the pitch coefficient correction function may include, for example, appropriately modifying a logistic function, or specifically a generalized logistic function. However, as will become more apparent in view of the following description, it may be worth noting that any other suitable means (e.g., a formula or expression) may be used to implement such a pitch coefficient correction function, provided that the pitch coefficient correction function so implemented satisfies certain properties.

[0025] In some example implementations, the pitch coefficient correction function may have one or more of the following properties: being continuous and monotonic with respect to relative velocity, having asymptotic limits controlled by one or more first parameters, providing a zero pitch coefficient correction value at zero relative velocity, and / or having a slope near zero velocity controlled by a second parameter. As will be understood and appreciated by those skilled in the art, any other suitable properties may also be required in some possible implementations.

[0026] In some example implementations, the pitch coefficient modification function F is:

number

[0027] In some example implementations, the method may further include outputting the rendered audio (e.g., as part of the audio content) to speakers or headphones (or any other suitable playback device) for playback to the user, depending on various implementations or user environments (e.g., computer, game console, mobile, etc.).

[0028] According to a second aspect of the present invention, there is provided a method for encoding parameters for use in modeling the Doppler effect when rendering audio content for a six degrees of freedom (6DoF) environment. The parameters so encoded by this method (e.g., at the encoder side or in an encoding environment) may be used by any of the methods described in the above first aspect and its example implementations to model the Doppler effect when rendering audio content for a 6DoF environment (e.g., at the user side or in a user / decoding environment).

[0029] In particular, the method may comprise determining (e.g., calculating, setting, etc.) first parameter values ​​for one or more first parameters indicating an acceptable range of pitch coefficient correction values. The method may further comprise determining (e.g., calculating, setting, etc.) a second parameter value for a second parameter indicating a desired strength (or, in some cases, referred to as "aggressiveness") of the Doppler effect to be modeled. Finally, the method may further comprise encoding an indication of the first and second parameter values. Specifically, as indicated above, the first and second parameter values ​​may be used to map a relative velocity between a listener and a sound source of the audio content to a pitch coefficient correction value based on a predetermined pitch coefficient correction function, which may be used to render the sound source, and the predetermined pitch coefficient correction function may have first and second parameters and may be a function for mapping the relative velocity to the pitch coefficient correction value.

[0030] Configured as described above, the proposed method can provide an efficient and flexible mechanism for encoding parameters used for Doppler effect modeling when rendering audio content for a 6DoF environment, while taking into account both the (acceptable or tolerable) capabilities of the underlying signal processing unit (on the audio renderer side) for pitch coefficient modification (e.g., high magnitude of pitch coefficient modification value for high relative velocities, singularities, etc.) and the possibility to control the (desired) pitch coefficient modification value (i.e., representing the desired strength of the Doppler effect) according to the content creator's intention (in other words, the targeted subjective listening experience), thereby improving the perceived listening experience (on the listener side).

[0031] Furthermore, as described above, because the pitch coefficient modification function is already implemented and deployed as a predetermined function (on the renderer side) that takes the first and second parameters as inputs, it is generally not necessary to redesign (or reimplement) a new pitch coefficient modification function every time rendering conditions change (e.g., a different renderer with different processing capabilities is deployed, different audio content created by a different person and / or for a different scene, etc.). Instead, the encoding side may simply communicate (e.g., encoded in the bitstream) different first and second parameter values ​​that respectively represent the corresponding tolerance range (limit) of the pitch coefficient modification value and the corresponding desired strength of the Doppler effect to be modeled. Therefore, in some possible scenarios, the predefined pitch coefficient modification function may be implemented as simply as a (renderer-side) plug-in that can be deployed in various software and / or platforms according to different requirements and / or implementation forms and further customized if necessary. This avoids unnecessary redesign / reimplementation of the pitch coefficient modification function, further improving efficiency in the overall audio rendering process.

[0032] In some example implementations, the indications of the first and second parameter values ​​may be encoded as labels (or fields) in the bitstream. As will be understood and appreciated by those skilled in the art, such indications may also be implemented in any other suitable means, so long as a corresponding rendering device (on which the predetermined pitch coefficient correction function is deployed) can be enabled to derive the first and second parameter values ​​as needed. By way of example and not limitation of any kind, in some possible cases where the encoding method may be performed by a (game / control) engine (or sometimes also referred to as a game control logic engine) in, for example, an AR / VR gaming environment, it can be understood that the first and second parameter values ​​(or their respective indications) do not necessarily always need to be encoded into the bitstream (e.g., because in some cases the game engine may typically be located in the same environment as the rendering and / or listening components, e.g., in the form of a PC), but may be encoded (or encapsulated) in any other suitable format (or even as plain or clear parameter values ​​in some possible cases), together with or separately from the audio content.

[0033] In some example implementations, the indication of the first and second parameter values ​​may be encoded together with the audio content in a single bitstream or as separate bitstreams.

[0034] In some example implementations, the first and second parameter values ​​may be determined by a content creator or a game engine, as indicated above.

[0035] According to a third aspect of the present invention, there is provided an audio renderer (rendering apparatus) including a processor and a memory coupled to the processor, the processor may be adapted to cause the audio renderer to perform all steps according to any of the exemplary methods described in the first aspect.

[0036] According to a fourth aspect of the present invention, there is provided an encoder (encoder apparatus) including a processor and a memory coupled to the processor, wherein the processor may be adapted to cause the encoder to perform all steps according to any of the exemplary methods described in the second aspect.

[0037] According to a fifth aspect of the present invention, there is provided a computer program, which may include instructions that, when executed by a processor, cause the processor to perform all of the steps of the methods described throughout the present disclosure.

[0038] According to a sixth aspect of the present invention, there is provided a computer-readable storage medium, which may store the computer program described above.

[0039] It will be understood that apparatus features and method steps can be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding apparatus (or system), and vice versa, as will be understood by those skilled in the art. Furthermore, it will be understood that any statements made above with respect to a method(s) apply equally to a corresponding apparatus (or system), and vice versa. [Brief explanation of the drawings]

[0040] Exemplary embodiments of the present invention are described below with reference to the accompanying drawings.

[0041] [Figure 1] FIG. 10 is a schematic diagram illustrating an example function mapping between relative velocity and pitch correction value. [Figure 2] 10A-10C are schematic diagrams illustrating exemplary functional mappings between relative velocity and pitch correction values ​​for different settings of the Doppler effect modeling range in accordance with an embodiment of the present invention. [Figure 3]5A-5C are schematic diagrams illustrating exemplary functional mappings between relative velocity and pitch correction values ​​for different settings of Doppler effect modeling strength, in accordance with an embodiment of the present invention. [Figure 4] 1 is a schematic flow chart illustrating an example of a method according to an embodiment of the present invention. [Figure 5] 4 is a schematic flow chart illustrating another example of a method according to an embodiment of the present invention. [Figure 6A] 2A and 2B illustrate schematic diagrams of exemplary comparisons between audio signals processed by a conventional Doppler Effect modeling approach and audio signals processed in accordance with embodiments of the present invention; [Figure 6B] 2A and 2B illustrate schematic diagrams of exemplary comparisons between audio signals processed by a conventional Doppler Effect modeling approach and audio signals processed in accordance with embodiments of the present invention; [Figure 7A] 4A and 4B schematically illustrate another exemplary comparison between an audio signal processed by a conventional Doppler Effect modeling approach and an audio signal processed in accordance with an embodiment of the present invention. [Figure 7B] 4A and 4B schematically illustrate another exemplary comparison between an audio signal processed by a conventional Doppler Effect modeling approach and an audio signal processed in accordance with an embodiment of the present invention. [Figure 8A] FIG. 1 is a block diagram of an exemplary apparatus for performing methods according to embodiments of the present invention. [Figure 8B] FIG. 1 is a block diagram of an exemplary apparatus for performing methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] The drawings and the following description relate to preferred embodiments for purposes of illustration only. It should be noted from the following description that alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of what is claimed.

[0043] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that wherever practicable, like or similar reference numerals may be used in the figures and may indicate like or similar functionality. The drawings depict embodiments of the disclosed system (or method) for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be used without departing from the principles described herein.

[0044] Furthermore, when connecting elements such as solid or dashed lines or arrows are used in the figures to indicate a connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element is not meant to imply that the connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the invention. In addition, for ease of explanation, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such element represents one or more signal paths for affecting the communication, as appropriate.

[0045] As noted above, the terms "Doppler effect" or "Doppler shift" are generally used to refer to the audio effect experienced when there is a change in the frequency of a wave (e.g., a sound wave) to an observer (e.g., a listener) who is moving relative to the wave source (e.g., a sound source). Generally, the Doppler effect can be observed whenever the wave source is moving relative to the observer. The Doppler effect can be described as the effect produced by a moving wave source, where an observer seeing the wave source approaching will experience an apparent upward shift in frequency, and an observer seeing the wave source receding will experience an apparent downward shift in frequency. Nevertheless, it is important to note that this effect does not occur due to an actual change in the frequency of the source.

[0046] The Doppler effect can be observed for any type of wave, water wave, sound wave, light wave, etc., as can be understood and appreciated by those skilled in the art. An exemplary scenario in which the Doppler effect may be commonly perceived may be the case of a police or emergency vehicle moving toward a listener on a highway. As the vehicle approaches the siren, the pitch (a measure commonly used to indicate frequency) of the siren sound becomes higher (higher), and then, after the vehicle has passed and moved away, the pitch of the siren sound becomes lower (lower).

[0047] Also, as mentioned above, the Doppler effect has recently begun to be considered as an important aspect in the audio rendering of dynamic scenes in six degrees of freedom (6DoF) environments, which are widely adopted, for example, in virtual reality (VR) and / or augmented reality (AR) scenarios (e.g., games, immersive content, etc.).

[0048] Roughly speaking, a semitone, also called a half-step or half-tone, is generally considered the smallest interval commonly used in most tonal music and the most dissonant when sounded in harmony. A semitone is generally defined as the distance between two adjacent notes on a twelve-tone scale. That is, for many instruments designed based on the tempered scale, which divides the octave into 12 equal parts, the ratio of the frequencies of any two semitones (half-steps) is roughly the 12th root of two.

[0049] Within the context of audio processing (e.g., audio rendering), the Doppler effect may generally be modeled using an audio pitch coefficient (shift) correction value (sometimes simply referred to as p throughout this disclosure). Thus, following the general concept of semitones, in some possible implementations of the present invention, as described in more detail below, a positive pitch coefficient (shift) correction value may generally mean an increase in the pitch of a sound source, particularly in units of semitones. For example, a pitch coefficient correction value of 2 may simply translate into an increase of two semitones of the sound source. Correspondingly, a negative pitch coefficient (shift) correction value may generally mean a decrease in the pitch (in semitones) of the sound source. However, as those skilled in the art will understand, alternative frameworks and units for pitch coefficient correction values ​​may also be feasible in the context of the present invention.

[0050] Some approaches for modeling the Doppler effect may include modeling based on a physical description and / or approximation of the Doppler effect. Therefore, these techniques generally do not have a means for taking into account the capabilities of the underlying signal processing units for pitch coefficient correction (e.g., high relative velocities, high magnitude of pitch coefficient correction values ​​for singularities), nor a means for controlling the pitch coefficient correction value (i.e., representing the strength of the Doppler effect) according to the content creator's intention (in other words, the subjective listening experience).

[0051] In that regard, the present invention may generally attempt to address a problem given 1) a relative velocity (calculated based on listener and sound source positions, also referred to as v throughout this disclosure), 2) a range of pitch coefficient correction values ​​supported by the signal processing unit (also referred to as l throughout this disclosure), and 3) content creator settings (e.g., finding an appropriate pitch correction value p that can perceptually correspond to the input data in 1) through 3) based on (traditional) modeling equations, real-world standards, artistic expectations for Doppler effect strength, etc., also referred to as s throughout this disclosure). In particular, in some possible implementations, the range of pitch coefficient correction values ​​l may itself be bounded by a lower bound l. l and upper limit lh Therefore, the range l can be expressed as l={l l ,l h}

[0052] To address the above problems, in broad terms, the present invention proposes to generally consider implementing a pitch coefficient correction function F to map relative velocities v to pitch coefficient correction values ​​p, taking into account limitations l of the signal processing unit and user-adjustable settings s. In some possible implementations, the pitch coefficient correction function F may be implemented as a modified generalized logistic function.

[0053] A possible example (not intended to be limiting in any way) for implementing such a pitch coefficient modification function F may be as follows:

number

[0054] Any equivalent expression of equation (1) is intended to be encompassed by the present invention. For example, as will be appreciated by those skilled in the art, the supported range of pitch coefficient correction values ​​may be expressed by any suitable combination of two parameters that are contemplated to be encompassed by the present invention. For example, a lower bound l h and upper limit l h Instead of using the limit, we can use the tolerance size indication Δl=l h -l lIn this case, for example, l h =Δl+l l or l l =l h −Δl holds and can be used to modify equation (1) above.

[0055] Furthermore, the Euler number in equation (1) can also be replaced by an alternative constant greater than 1, or by an alternative constant less than 1 by swapping the sign of the exponent.

[0056] Additionally, as will be understood and appreciated by those skilled in the art, the pitch coefficient correction function F may be defined in any other suitable form / formula depending on various requirements and / or implementations, provided that the above-identified requirements (i.e., range / limit parameter l and user / content creator settings s) are taken into consideration. The characteristics of the pitch coefficient correction function F will become more apparent in consideration of the following description accompanying the drawings.

[0057] In particular, in a broad sense, by applying the approach proposed in the present invention, the values ​​of the parameters of the signal processing unit limit l and the user setting s may be allowed to be adjusted on the encoder side (and possibly, for example, encapsulated into the bitstream). Thus, the content creator may establish control over the Doppler effect modeling in order to adapt it to the capabilities of the audio renderer, and at the same time, be allowed to adjust the Doppler effect modeling according to the content creator's own preferences.

[0058] Turning now to the drawings, FIG. 1 is a schematic diagram illustrating exemplary functional mappings between relative velocity and pitch corrections based on different approaches to modeling the Doppler effect.

[0059] In particular, the x-axis schematically indicates the (input) relative velocity between the sound source and the observer / listener (e.g., a listener in a 6DoF environment). The relative velocity between the sound source and the observer / listener may be determined by any suitable means, for example, based on the positions of the listener and the sound source. For example, in some possible implementations, the relative velocity may be determined based on the rate of change (e.g., first derivative) of the positions of the listener and the sound source (e.g., the distance between the sound source and the observer / listener). As will be understood and appreciated by those skilled in the art, a negative value of the relative velocity generally means that the sound source and the observer / listener are approaching each other (getting closer to each other), and a positive value of the relative velocity generally means that the sound source and the observer / listener are moving further away from each other.

[0060] Meanwhile, the y-axis schematically shows the (output) pitch shift correction value (e.g., in semitones). As indicated above, in some possible implementations, a positive value of the pitch shift correction value may generally mean an increase in the pitch of the sound source, and a negative value of the pitch shift correction value may generally mean a decrease in the pitch of the sound source, e.g., both in semitones.

[0061] More specifically, diagram 101 in FIG. 1 generally illustrates an example of a (theoretical) basis for a Doppler effect model, e.g., based on a (theoretical) mathematical formula. Because diagram 101 generally represents a (theoretical) basis for modeling the Doppler effect, the diagram may not be compatible with, for example, an actual software implementation (or, in other words, an implementation for rendering audio content in a 6DoF environment), but may be considered to primarily serve as a purely mathematical illustration for modeling the “nature” of the Doppler effect. Thus, as will be understood and appreciated by those skilled in the art, the “cutoff” (approximately −343 m / s) as shown in diagram 101 may generally be attributed to the limitation of the speed of sound. Similarly, the nearly infinite pitch coefficient correction value as the relative velocity approaches −343 m / s (from 0) causes diagram 101 to appear “segmented” in some way.

[0062] On the other hand, the diagram 102 generally represents an example of a possible approach to modeling the Doppler effect, which may be performed, for example, based on estimations (approximations) of (theoretical) mathematical formulas.

[0063] Furthermore, diagram 103 (solid line) generally illustrates an example of a possible modeling of the Doppler effect according to an embodiment of the present invention. As mentioned above, this modeling of the Doppler effect, achieved through a (predetermined) pitch coefficient correction function F, includes multiple parameters (or, in other words, takes multiple parameters as inputs), among them the range of pitch coefficient correction values ​​supported by the signal processing unit (i.e., l) and content creator settings (i.e., s), which may be based on, for example, (traditional) modeling equations, real-world standards, artistic expectations of Doppler effect strength, etc.

[0064] In the particular example shown in FIG. 1 (as can be seen from exemplary diagram 103 and diagram 102), the range / limit parameter l is illustratively set to {-8,8}, i.e., l={l l ,l h Note that}={-8,8}. And the strength / aggressiveness parameter s is illustratively set to 0.015. However, as will be understood and appreciated by those skilled in the art, these values ​​of the parameters are set merely as possible examples (and not as limitations of any kind), and any other suitable values ​​may of course be used depending on various requirements and / or implementations.

[0065] As clearly shown in the example of FIG. 1, diagram 102 (i.e., representing a possible modeling of the Doppler effect) generally shows a "rough" / "intense" transition (or transition) from a lower velocity range (approximately 0 to ±170 m / s) to a higher velocity range (approximately above ±170 m / s). In contrast, the transitions in those regions of diagram 103 (representing a possible implementation according to an embodiment of the present disclosure) appear to be "softer." Therefore, the audio quality (of audio content being rendered according to the Doppler effect model of diagram 103) perceived by a listener (e.g., in a 6DoF environment) may be improved.

[0066] As noted above, any suitable form / formula (other than that illustrated in equation (1)) may be used to implement the pitch coefficient modification function F, provided that at least the range / limit parameter l and user / content creator settings s are taken into account. Nevertheless, it may be worth noting some properties that the pitch coefficient modification function F may need to satisfy in order to achieve more or less similar performance (e.g., in terms of perceived audio quality) comparable to that of the above-exemplified equation (1).

[0067] More specifically, according to embodiments of the present invention, the pitch coefficient modification function F may have one or more of the following properties: - Continuous and monotonic with respect to relative velocity, having asymptotic limits controlled by said one or more range / limit parameters (i.e., l); Providing a zero pitch coefficient correction value at zero relative velocity, and / or · Have a gradient near zero velocity that is controlled by the aggressiveness / strength parameter (i.e., s).

[0068] In some implementations, the pitch coefficient modification function F may have all of the above properties.

[0069] The above properties are also clearly reflected in diagram 103, which implements the pitch coefficient modification function F according to an embodiment of the present invention.

[0070] Reference is now made to Figures 2 and 3, which show in more detail a schematic of how different parameter settings affect modeling of the Doppler effect. As indicated above, it can be understood that, in a broad sense, the signal processing unit setting (e.g., l) generally affects the asymptotic limit of the function more in the high relative velocity region (e.g., as illustrated in Figure 2), while the content creator criteria setting (e.g., s) generally affects the slope of the function F more around the low relative velocity region (e.g., as illustrated in Figure 3).

[0071] In particular, Figure 2 is a schematic diagram illustrating an exemplary functional mapping between relative velocity and pitch correction value for different settings of the Doppler effect modeling range parameter 1, according to an embodiment of the present invention. In particular, diagram 201 of Figure 2, which is a (theoretical) mathematical representation of the Doppler effect, is the same as diagram 101 of Figure 1, and therefore a repeated description thereof may be omitted for the sake of brevity.

[0072] Specifically, in the example shown in FIG. 2, the range / limit parameter l={l l ,l h} are exemplarily set to l={-8,8}, l={-4,8}, and l={-8,4}, respectively. Meanwhile, the slope parameter s in all diagrams 202, 203, and 204 is the same (and may be set to any suitable value, e.g., 0.015). Thus, as can be seen from FIG. 2, diagrams 202, 203, and 204 appear to exhibit more or less similar slopes (especially in the low speed region), but only the respective upper and / or lower limits of the (output) pitch coefficient correction values ​​are set by the range / limit parameter l={l l ,l h} depending on the

[0073] As mentioned above, the range / limit parameter l is generally set to indicate the range of pitch coefficient correction values ​​supported by the (underlying) signal processing unit (e.g., on the renderer side) for performing Doppler modeling. In other words, such parameter l may be considered to generally represent the (processing) capabilities (e.g., in terms of hardware and / or software capabilities) of the signal processing unit (or, broadly speaking, the renderer). Furthermore, as mentioned above, the pitch coefficient correction function F may typically be implemented as a plug-in (or encapsulated as some kind of library) that simply receives the range parameter l (e.g., from the encoding side) along with other inputs (e.g., relative velocity, content creator settings s, etc.). Thus, in some implementations, it may be possible that the so-received range parameter l may unfortunately not be supported (fully or partially) by the (processing unit of) the renderer. An example of such a scenario may be when a mobile device (e.g., a mobile phone) with (relatively) limited processing (rendering) capabilities receives, along with audio content for rendering, a range parameter l (e.g., {-8,8}) that was originally set (e.g., by an encoder) to target a (relatively) more powerful rendering device (e.g., a game console or professional workstation). In such a case, it may be more practical for the mobile device to apply a default range parameter setting (e.g., {-4,8} or {-8,4}) rather than using the received, unsupported parameter l (e.g., {-8,8}). Here, the default range parameter setting may be set by the mobile device's manufacturer to more accurately reflect the mobile device's actual processing (rendering) capabilities, e.g., to avoid unexpectedly adversely affecting the rendering process.

[0074] 3 is a schematic diagram illustrating an exemplary functional mapping between relative velocity and pitch correction value for different settings of the Doppler effect modeling strength parameter s, according to an embodiment of the present invention. As above, diagram 301 of FIG. 3, which is a (theoretical) mathematical representation of the Doppler effect, is the same as diagram 101 of FIG. 1 (and the same as diagram 201 of FIG. 2), so a repeated description thereof may be omitted for the sake of brevity.

[0075] In the example shown in Figure 3, the slope / aggressiveness parameter s (e.g., primarily used to represent the content creator's intention for the subjective listening experience) in diagrams 302, 303, 304, and 305 is illustratively set to s = 0.015, s = 0.010, s = 0.005, and s = 0.0025, respectively. Meanwhile, the range parameter l in all diagrams 302, 303, 304, and 305 is the same (which may be set to any suitable value, e.g., {-8, 8}). Thus, as can be seen from Figure 3, although diagrams 302, 303, 304, and 305 appear to exhibit varying slopes (especially in the slow regions), the (theoretical) upper limits and / or bounds of the (output) pitch coefficient correction values ​​are more or less similar. Thereby, different "strengths" of the Doppler effect in the audio content being rendered may be perceived by the listener (eg, in a 6DoF environment), depending, for example, on the intention of the content creator.

[0076] In summary, by appropriately setting the slope / aggressiveness parameter s (possibly followed by encoding / encapsulating the parameter value, e.g., into a bitstream or other suitable format and transmitting or communicating it to a user / decoder-side device, or generally to a renderer), the content creator (or in some possible implementations the “game engine”) generally has the freedom to control the modeling behavior of the Doppler effect as desired, e.g., between no Doppler effect modeling at all and (nearly) “realistic” (theoretical) Doppler effect modeling, or even over-emphasized Doppler effect modeling.

[0077] Therefore, it can be said that the present invention provides content creators with additional degrees of freedom regarding Doppler effect modeling at the user side. In this way, the present invention allows content creators to selectively control or override the Doppler effect modeling by the decoder / renderer in an object-specific manner. This is achieved by providing a set of parameter values ​​in an appropriate format to the user / decoder-side device (which ultimately includes the actual renderer). These parameter values ​​may be encoded in the bitstream or provided to the renderer in any form suitable for or compatible with the renderer's data interface.

[0078] As an illustrative, non-limiting example, consider the use case of a VR scene with a supersonic flying jet. If the actual laws of physics for Doppler effect modeling (e.g., corresponding to diagram 301 in FIG. 3 ) were applied, the user / listener (e.g., a game player or other recipient of the VR scene content) would likely not perceive any sound from the jet at all, which would result in an unpleasant (but physically realistic) VR experience (e.g., a gaming experience). In that case, particularly by applying the methods proposed in this disclosure, the content creator (or a suitably configured game engine) has the freedom to control the modeling of the Doppler effect as desired. In other words, by appropriately setting the value of, for example, the gradient / aggressiveness parameter s, a content creator (or game engine) is given the freedom to override the renderer's Doppler modeling (e.g., corresponding to diagram 301 in FIG. 3 ) according to the laws of physics when deemed necessary or desirable by using other modeling settings (e.g., corresponding to diagrams 302, 303, 304, or 305, etc.) that are perhaps less accurate or less realistic for the user in the VR environment but would result in a more comfortable listening experience. This example assumes that the renderer used to render the VR scene is capable of applying “default” Doppler effect modeling according to or based on the laws of physics. In some example implementations, this default Doppler effect modeling may be realized by a specific set of parameter values ​​for the aforementioned equations.

[0079] As another illustrative, non-limiting example, weaker Doppler effect modeling may be applied to relatively slow objects or objects with speech (e.g., cartoon characters), or in some cases, even no Doppler effect modeling at all, while other objects may be considered for moderate or physically accurate Doppler effect modeling. For example, given a scene with very fast audio objects (e.g., flying superheroes) accompanied by speech, it may be desirable to apply little or no Doppler effect modeling to the speech, but at least some Doppler effect modeling to the remaining sounds.

[0080] 4 is a schematic flowchart illustrating an example of a method 400 for modeling the Doppler effect when rendering audio content for a 6DoF environment, according to an embodiment of the present invention. Depending on the implementation, the method may be performed in a decoder or a user-side environment (e.g., a VR / AR environment).

[0081] In particular, method 400 may begin in step S401 by obtaining (e.g., receiving) first parameter values ​​for one or more first parameters indicating an acceptable range of pitch coefficient correction values. Subsequently, in step S402, method 400 may include obtaining a second parameter value for a second parameter indicating a desired strength of the Doppler effect to be modeled. Method 400 may then proceed to step S403 by using a predetermined pitch coefficient correction function to determine a pitch coefficient correction value based on a relative velocity between a listener and a sound source in the audio content and the first and second parameter values. The pitch coefficient correction function may be predefined in any suitable form, for example, pre-implemented as a plug-in or library, in accordance with the above description with respect to FIGS. 1-3. More specifically, the pre-defined pitch coefficient correction function may have, among other things, first and second parameters (or, in other words, may take the first and second parameters as (additional) inputs) and may be a function for mapping relative velocity to a pitch coefficient correction value. Finally, the method 400 may include rendering the sound source based on the pitch coefficient correction value in step S404. Depending on the implementation, the method may optionally further include outputting the rendered sound source to, for example, an output (playback) device (e.g., corresponding to or including one or more speakers, headphones, etc.), and the rendered audio output (signal) with the modeled Doppler effect may be reproduced and perceived by a user.

[0082] As described above, the proposed method 400 may generally utilize a predefined (or predetermined / pre-implemented) pitch coefficient correction function to map relative velocities to corresponding pitch coefficient correction values ​​to model Doppler effects (e.g., when rendering audio content in a 6DoF environment). In operation, in practice, when the proposed method 400 is executed, for example, by an audio rendering device (e.g., an audio rendering device of an AR / VR device in a user-side environment) on which the predefined pitch coefficient correction function is deployed, the audio rendering device may be configured to obtain first and second parameter values ​​corresponding to the first and second parameters, respectively. For example, the audio rendering device may obtain first and second parameter values ​​for a sound source on a frame-by-frame basis, e.g., for each frame or for each keyframe. Thus, the present disclosure provides time-dependent and / or object-specific control of Doppler effect modeling.

[0083] The actual acquisition of the first and second parameter values ​​may be performed in any suitable manner depending on the requirements and / or implementation. For example, in some possible implementations, the first and second parameter values ​​may be derived (e.g., decoded or extracted) from a bitstream encoded and sent by an encoding device (e.g., as described in method 500 below in connection with FIG. 5). In some other possible implementations, the first and second parameter values ​​may be acquired (e.g., simply read) from a file or look-up table (LUT) stored in, for example, a memory of a user device based on instructions in the bitstream. In that case, the encoding environment / device may transmit, for example, an appropriate pointer, reference, or index, encoded in the bitstream, plain / clear, or in any other suitable format.

[0084] It should also be noted that, depending on how the actual audio content itself is encoded and / or transmitted, decoding of the audio content (e.g., audio signal) at the user side may be performed in any suitable manner and at any suitable time before final rendering (step S404) occurs, as will be understood and appreciated by those skilled in the art. Thus, decoding of the actual audio content (e.g., audio signal) is independent of the determination of the pitch coefficient modification values.

[0085] Configured as described above, the proposed method can provide an efficient and flexible mechanism for performing (e.g., controlling) Doppler effect modeling when rendering audio content for a 6DoF environment, while taking into account the (acceptable or tolerable) capabilities of the underlying signal processing unit (of the audio renderer) for pitch coefficient correction (e.g., high magnitude of pitch coefficient correction value for high relative velocities, singularities, etc.), and providing the possibility to control the (desired) pitch coefficient correction value (i.e., representing the desired strength / aggressiveness of the Doppler effect to be modeled) according to the content creator's intention (in other words, the subjective listening experience), thereby improving the perceived listening experience (on the listener's side, e.g., a gamer in a VR environment).

[0086] Furthermore, because the pitch coefficient correction function is already pre-implemented as a predetermined function that takes, among other things, values ​​of both the first and second parameters as inputs, it is generally not necessary to redesign (or re-implement) a new pitch coefficient correction function every time rendering conditions change (e.g., a different renderer with different processing capabilities is deployed, different audio content is created by a different author and / or for a different scene, etc.). Rather, it is generally only necessary to communicate (e.g., using an encoded bitstream on the encoding side) different first and second parameter values ​​that respectively represent corresponding allowable ranges (e.g., limits) of pitch coefficient correction values ​​and corresponding desired strengths of the Doppler effect to be modeled. Thus, in some possible scenarios, the pre-defined pitch coefficient correction function may be implemented as simply as a plug-in or library (that takes the first and second parameters as inputs for modeling a desired Doppler effect), which may be deployed in various software and / or platforms or further customized if necessary, depending on different requirements and / or implementation forms. Thereby, unnecessary redesign / reimplementation of pitch coefficient modification functions is avoided, further improving efficiency in the overall audio rendering process.

[0087] FIG. 5 is a schematic flowchart illustrating another example of a method 500 for encoding parameters for use in modeling the Doppler effect when rendering audio content for a 6DoF environment, according to an embodiment of the present invention. In other words, the parameters so encoded by this method 500 may be used to model the Doppler effect when rendering audio content for a 6DoF environment by the preceding method 400 as described with reference to FIG. 4. That is, in some possible implementations, the parameter values ​​encoded by the method 500 of FIG. 5 may be transmitted or communicated (in any suitable manner) to, for example, a user-side device (e.g., at the user side or in a decoding / rendering environment). The user-side device may be configured to suitably obtain (e.g., decode from a bitstream) the parameter values ​​and perform the method 400 for modeling the Doppler effect as described above with respect to FIG. 4. In particular, depending on the implementation, the encoding method 500 may be performed, for example, by an encoding device (or encoder for short) utilizing user input from a content creator, by a game engine, etc.

[0088] In particular, method 500 may begin in step S501 by determining first parameter values ​​for one or more first parameters indicating an acceptable range of pitch coefficient correction values. Subsequently, in step S502, method 500 may include determining a second parameter value for a second parameter indicating a desired strength of the Doppler effect to be modeled. Finally, in step S503, method 500 may include encoding an indication of the first and second parameter values. More specifically, the first and second parameter values ​​may be used to map a relative velocity between a listener and a sound source of the audio content to a pitch coefficient correction value based on a predetermined pitch coefficient correction function. As indicated above, the pitch coefficient correction value may be used to render the sound source, and the predetermined pitch coefficient correction function may have first and second parameters and may be a function for mapping the relative velocity to a pitch coefficient correction value.

[0089] As will be understood and appreciated by those skilled in the art, the first and second parameter values ​​may be encoded in any suitable manner. For example, in some implementations, the first and second parameters may be encoded into a single bitstream along with the audio content (e.g., an audio signal) or into separate bitstreams. Depending on the implementation and / or requirements, the first and second parameter values ​​(possibly along with the audio content / signal) may also be encoded, with or without compression, into any suitable format, such as a bitstream or data format compatible with an audio standard, such as the MPEG audio standard (e.g., the upcoming MPEG-I audio standard). In that case, the first and second parameter values ​​may be appropriately encoded, for example, as part of a header field, metadata, etc., as will be understood and appreciated by those skilled in the art. In some other possible cases, the first and second parameter values ​​may be inserted (encapsulated) as plain variables (e.g., floating-point numbers) in any suitable data format. The encoded bitstream may be transmitted or communicated to a user environment (eg, comprising a decoding or rendering device) by using any suitable means, for example in a wired or wireless manner.

[0090] Furthermore, as mentioned above, the encoding method 500 proposed in the present disclosure may be executed based on user input, for example, by a content creator. In that case, the determination of the first parameter value in step S501 and / or the determination of the second parameter value in step S502 may be based on the user input. Similarly, the encoding method 500 may be executed by a (software-based) game engine (game control logic engine) depending on the scenario and / or implementation. In this case, the determination in S501 and / or S502 is performed according to a decision routine of the game engine, for example, based on the type and / or speed of the sound source.

[0091] More specifically, in the case of user input from a content creator, in some possible implementations, the content creator may obtain information indicating the processing capabilities or profile of the target decoding / rendering device using any suitable means to appropriately determine and set a value for the first parameter indicating the acceptable range of pitch coefficient correction values. In some possible implementations, the content creator may input multiple parameter sets (i.e., with respective first and second parameter values) for encoding for respective multiple target devices, each parameter set comprising first and second parameter values ​​targeted to a respective decoding / rendering device. In that case, the decoding / rendering device may simply pick or choose from the received parameter sets to obtain respective first and second parameter values ​​that best suit the decoding / rendering device (e.g., that best suit the profile or capabilities of the decoding / rendering device). In addition, the content creator may also need to decide whether to apply Doppler effect modeling, and if so, to what extent (e.g., as shown above with respect to FIG. 3), depending on various scenarios and / or implementations. For example, in some possible implementations, a content creator may utilize and set a (global) flag (e.g., a specific bit field in the bitstream) to (globally) activate or deactivate Doppler effect modeling. This flag is then similarly encoded. In some other possible implementations, instead of using a (global) flag, a content creator may simply set the gradient (aggressiveness) of the Doppler effect to be modeled to 0 (e.g., by controlling the value of a second parameter indicating the desired strength). Thus, compared to global activation or deactivation, a content creator may have more freedom to control Doppler effect modeling in a more continuous manner (e.g., using frame-by-frame control of Doppler effect modeling).

[0092] On the other hand, in the case of a (software-based) game engine (e.g., for a VR / AR environment) performing the encoding task, the process is largely the same as that described above with respect to user input from a (human) content creator, except that the role of the content creator is replaced by the game engine. More specifically, here it is the game engine (or its developer) that may need to acquire knowledge of the corresponding capabilities / profiles of the rendering / decoding platform and, in addition, determine and control the gradient / aggressiveness of the Doppler effect modeling as appropriate, depending on the implementation and / or requirements (e.g., by using any suitable logic / algorithm, machine learning, hard coding, etc.). In some possible cases, mainly because in (real-time) rendering in a VR / AR environment, the game engine (performing the encoding task) is typically located together with the decoding / rendering device / component (e.g., in the same computer device or game console), the values ​​of the first and second parameters may not even need to be encoded into the bitstream but may be communicated / transmitted to the decoding / rendering device / component in other appropriate formats (e.g., as plain variables, etc.). Also, depending on the scenario and / or implementation, parameter values ​​may be communicated periodically (e.g., frame-based), on-demand, or in any other suitable manner.

[0093] Configured as described above, the proposed method can provide an efficient and flexible mechanism for encoding parameters used for Doppler effect modeling when rendering audio content for a 6DoF environment, while also taking into account the (acceptable or tolerable) capabilities of the underlying signal processing unit (at the audio renderer side) for pitch coefficient modification (e.g., high magnitude of pitch coefficient modification value for high relative velocities, singularities, etc.) and providing the possibility to control the (desired) pitch coefficient modification value (i.e., representing the desired intensity of the Doppler effect) according to the content creator's intention (in other words, subjective listening experience), thereby improving the perceived listening experience (at the listener / user side).

[0094] Furthermore, as described above, because the pitch coefficient modification function is already implemented and deployed as a predetermined function (on the renderer side) that takes the first and second parameters as inputs, it is generally not necessary to redesign (or reimplement) a new pitch coefficient modification function every time rendering conditions change (e.g., a different renderer with different processing capabilities is deployed, different audio content is created by a different person and / or for a different scene, etc.). Instead, the encoding side may simply communicate (e.g., encoded in the bitstream) different first and second parameter values ​​that respectively represent the corresponding tolerance range (limit) of the pitch coefficient modification value and the corresponding desired strength of the Doppler effect to be modeled. Therefore, in some possible scenarios, the predefined pitch coefficient modification function may be implemented as simply as a (renderer-side) plug-in that can be deployed in various software and / or platforms according to different requirements and / or implementation forms, or may even be further customized if necessary. This avoids unnecessary redesign / reimplementation of the pitch coefficient modification function, further improving efficiency in the overall audio rendering process.

[0095] Finally, note that the minimum requirement(s) for the pitch coefficient correction function F described above allow for a computationally simple implementation of this function while still achieving realistic modeling of Doppler effects. This may be particularly true for the explicit example of the pitch coefficient correction function F given in equation (1) and its equivalents.

[0096] 6A-6B and 7A-7B, a comparison is now shown schematically between an audio signal (e.g., at a user's environment) processed by a possible Doppler effect modeling approach and an audio signal (e.g., at a user's environment) processed in accordance with an embodiment of the present invention. In other words, FIGS. 6A-6B and 7A-7B generally illustrate and compare respective renderings (in the form of spectrograms) obtained by applying different modeling approaches to the Doppler effect. In particular, as will be understood and appreciated by those skilled in the art, in FIGS. 6A-6B and 7A-7B, the x-axis generally represents time and the y-axis generally represents frequency.

[0097] More specifically, as reflected in FIG. 6A (according to a possible conventional modeling approach) compared to FIG. 6B (according to a possible implementation of the present invention), in which the same exemplary audio signal "jet" is processed by each modeling approach, the pitch coefficient modification function F as proposed in the present invention (i.e., as exemplarily shown in FIG. 6B) generally exhibits a higher order of continuity (soft / smooth bends vs. hard / sharp bends), which results in better perceptual performance. Similar findings are observed in the comparisons shown in FIGS. 7A and 7B, in which the same exemplary audio signal "siren" is processed by each modeling approach. In both cases of FIGS. 6A-6B and 7A-7B, a constant acceleration of the sound source from -500 to +500 m / s is assumed for illustrative purposes. In particular, as will be understood by those skilled in the art, the lines or line structures shown in Figures 6A-6B and 7A-7B extending from the far left to the right of the figures may be considered to represent "equal energy" time / frequency slots, in the sense that these lines or line structures connect time / frequency slots having substantially equal or comparable energy densities. Thus, these lines or line structures show how energy density has moved across frequency as time progresses.

[0098] The present invention also relates to apparatuses for performing the methods and techniques described throughout the present invention. Figures 8A and 8B generally illustrate examples of such apparatuses 800 and 801, respectively. In particular, the apparatus 800 (or 801) comprises a processor 810 (or 811) and a memory 820 (or 821) coupled to the processor 810 (or 811). The memory 820 (or 821) may store instructions for the processor 810 (or 811). The processor 810 (or 811) may, among other things, receive input data (e.g., in the form of a bitstream or any other suitable format) 830 (or 831). The processor 810 (or 811) may be adapted to perform the methods / techniques described throughout the present invention and generate output data 840 (or 841) accordingly. For example, device 800 may, depending on the circumstances, implement an audio renderer configured to perform method 400 for modeling the Doppler effect when rendering audio content for a 6DoF environment such as that shown above with reference to FIG. 4, and device 801 may, depending on the circumstances, implement an encoder configured to perform method 500 for encoding parameters for use in modeling the Doppler effect when rendering audio content for a 6DoF environment such as that shown above with reference to FIG. 5 in accordance with an embodiment of the present invention.

[0099] interpretation A computing device implementing the above techniques may have the following exemplary architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the exemplary architecture includes one or more processors (e.g., a dual-core Intel® Xeon® processor), one or more output devices (e.g., an LCD), one or more network interfaces, one or more input devices (e.g., a mouse, a keyboard, a touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, a hard disk, an optical disk, flash memory, etc.). These components may communicate and exchange data over one or more communication channels (e.g., a bus), which may utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0100] The term "computer-readable medium" refers to any medium that participates in providing instructions to a processor for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media, including, but not limited to, coaxial cables, copper wire, and fiber optics.

[0101] The computer-readable medium may further include an operating system (e.g., a Linux operating system), a network communications module, an audio interface manager, an audio processing manager, and a live content distributor. The operating system performs basic tasks including, but not limited to, recognizing input from and providing output to network interfaces and / or devices, tracking and managing files and directories on the computer-readable medium (e.g., memory or storage devices), controlling peripheral devices, and managing traffic on one or more communications channels. The network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communications protocols such as TCP / IP, HTTP, etc.).

[0102] The architecture may be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device having one or more processors. The software may include multiple software components or may be a single body of code.

[0103] The described features may be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used directly or indirectly in a computer to perform a particular activity or bring about a particular result. Computer programs can be written in any form of programming language, including compiled or interpreted languages ​​(e.g., Objective-C, Java), and can be deployed in any form, including as a standalone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0104] Processors suitable for executing a program of instructions include, by way of example, both general-purpose and special-purpose microprocessors, and the sole processor or one of multiple processors or cores of any kind of computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer also includes, or is operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, application-specific integrated circuits (ASCs).

[0105] To provide for user interaction, features may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retinal display device, for displaying information to the user. The computer may have a touch surface input device (e.g., a touch screen) or keyboard, and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0106] Features may be implemented in a computer system that includes back-end components such as data servers, or includes middleware components such as application servers or Internet servers, or includes front-end components such as client computers with graphical user interfaces or Internet browsers, or includes any combination thereof. The components of the system may be connected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include, for example, LANs, WANs, and the computers and networks forming the Internet.

[0107] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server sends data (e.g., HTML pages) to client devices (e.g., for the purpose of displaying the data to and receiving user input from users interacting with the client devices). Data generated at the client devices (e.g., results of user interaction) may be received from the client devices at the server.

[0108] One or more computer systems may be configured to perform particular actions by having software, firmware, hardware, or a combination thereof installed on the system that, when in operation, causes the system to perform the actions. One or more computer programs may be configured to perform particular actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.

[0109] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a certain combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0110] Similarly, while operations are shown in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order shown, or in any sequential order, or that all of the shown operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0111] Unless otherwise stated, and as will be apparent from the following description, throughout the description of the present invention, descriptions utilizing terms such as "processing," "calculating," "computing," "determining," "analyzing," and the like will be understood to refer to the operations and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, into other data also represented as physical quantities.

[0112] Throughout this application, references to "one example embodiment," "example embodiments," or "example embodiments" mean that a particular feature, structure, or characteristic described in connection with an example embodiment is included in at least one example embodiment of the invention. Thus, the appearances of the phrases "in one example embodiment," "in example embodiments," or "in example embodiments" in various places throughout the application are not necessarily all referring to the same example embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this application, in one or more exemplary embodiments.

[0113] As used herein, unless otherwise specified, the use of ordinal adjectives such as "first," "second," "third," etc. to describe a common object merely indicates that different instances of a similar object are being referred to and is not intended to imply that the objects so described must be in a given order, temporally, spatially, in ranking, or in any other way.

[0114] It is also to be understood that the phraseology and terminology used herein is for purposes of description and should not be regarded as limiting. The use of "including," "comprising," or "having," and variations thereof, is meant to encompass the items listed thereafter and equivalents thereof, as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled," and variations thereof, are used broadly and encompass both direct and indirect mounting, connecting, supporting, and coupling.

[0115] In the following claims and in the description of this specification, any one of the terms "comprising," "comprised of," or "which comprises" is an open term meaning the inclusion of at least the element / feature that follows it, but not the exclusion of others. Therefore, when used in the claims, the term "comprising" should not be interpreted as being limited to the means or elements or steps listed thereafter. For example, the scope of the expression "device comprising A and B" should not be limited to a device consisting only of elements A and B. As used in this specification, any one of the terms "including," "which includes," or "that includes" is also an open term meaning the inclusion of at least the element / feature that follows it, but not the exclusion of others. Therefore, "including" is synonymous with "comprising" and means "comprising."

[0116] In the foregoing description of exemplary embodiments of the invention, it should be understood that various features of the invention may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of streamlining the invention and aiding in understanding one or more of the various inventive aspects. However, this method of the invention should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single, foregoing disclosed exemplary embodiment. Accordingly, the claims following this specification are expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the invention.

[0117] Furthermore, some exemplary embodiments described herein may include some features included in other exemplary embodiments but not other features, and combinations of features from different exemplary embodiments are meant to be within the scope of the present invention and form different exemplary embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.

[0118] In the description provided herein, numerous specific details are set forth. However, it will be understood that example embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0119] Thus, while what is believed to be the best mode of the invention has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as fall within the scope of the invention. For example, any formulas above are merely representative of procedures that may be used. Functions may be added to or deleted from block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure.

[0120] Various aspects of the invention can be understood from the following enumerated exemplary embodiments (EEE).

[0121] EEE1. A method for modeling the Doppler effect when rendering audio content for a six degree of freedom (6DoF) environment at a user site, comprising: obtaining first parameter values ​​for one or more first parameters indicative of an acceptable range of pitch coefficient correction values; obtaining a second parameter value of a second parameter indicative of a desired strength of the Doppler effect to be modeled; determining a pitch coefficient correction value based on a relative velocity between a listener and a sound source within the audio content and the first and second parameter values ​​using a predetermined pitch coefficient correction function; and rendering the sound source based on the pitch coefficient modification; The method of claim 1, wherein the predetermined pitch coefficient correction function has the first and second parameters and is a function for mapping relative velocity to a pitch coefficient correction value.

[0122] EEE2. The method of EEE1, wherein the relative velocity is calculated based on the positions of the listener and the sound source.

[0123] EEE3. The method of EEE1 or EEE2, wherein the one or more first parameters include parameters indicating upper and / or lower limits of an acceptable range of pitch factor correction values.

[0124] EEE4. The method of any one of the preceding EEE, wherein the allowable range of pitch coefficient modification values ​​reflects the processing capabilities of an audio renderer that renders the audio content.

[0125] EEE5. The method of EEE4, wherein if the allowed range of pitch coefficient modification values ​​is not supported by the audio renderer, a default range of pitch coefficient modification values ​​is used by the audio renderer.

[0126] EEE6. The method of any one of the preceding EEEs, wherein the second parameter controls the slope of the pitch coefficient correction function, which reflects the aggressiveness of the Doppler effect to be modeled.

[0127] EEE7. A method according to any one of the preceding EEE, wherein audio content is extracted from a received bitstream and wherein first and second parameter values ​​are derived from instructions included in said bitstream.

[0128] EEE8. The method of any one of EEE1 to EEE6, wherein said audio content and said first and second parameter values ​​are obtained from separate bitstreams.

[0129] EEE9. The method of any one of the preceding EEE, wherein the second parameter value is set by a content creator of the audio content.

[0130] EEE10. The method of any one of the preceding EEEs, wherein the second parameter value is set by modeling real-world standards and / or artistic expectations for the desired Doppler effect strength.

[0131] EEE11. Rendering the audio content based on the pitch factor modifications comprises: 10. The method of claim 1, further comprising adjusting the pitch of the sound source within the audio content based on the pitch coefficient modification value.

[0132] EEE12. The method of EEE11, wherein a positive pitch coefficient modification value indicates an increase in the pitch of the sound source.

[0133] EEE13. The method of EEE11 or 12, wherein the pitch adjustment of the sound source is performed in semitone increments.

[0134] EEE14. The method of any one of the preceding EEEs, wherein the pitch coefficient correction function is based on a generalized logistic function.

[0135] EEE15. The pitch coefficient correction function - continuous and monotonic with respect to relative velocity; - having an asymptotic limit controlled by said one or more first parameters; - providing a zero pitch coefficient correction value at zero relative velocity; and / or - having a slope in the vicinity of zero speed that is controlled by said second parameter; A method as set forth in any one of the preceding EEEs, having one or more of the following characteristics:

[0136] EEE16. The pitch coefficient correction function F is

number

[0137] EEE17. The method of any one of the preceding EEE, further comprising outputting the rendered audio source to speakers or headphones for playback to the user.

[0138] EEE18. Method of encoding parameters for use in modeling the Doppler effect when rendering audio content for a six degree of freedom (6DoF) environment determining first parameter values ​​for one or more first parameters indicative of an acceptable range of pitch coefficient correction values; determining a second parameter value of a second parameter indicative of a desired strength of the Doppler effect to be modeled; encoding an indication of the first and second parameter values; the first and second parameter values ​​can be used to map a relative velocity between a listener and a sound source of the audio content to a pitch coefficient modification value based on a predetermined pitch coefficient modification function, the pitch coefficient modification value being used to render the sound source, the predetermined pitch coefficient modification function having the first and second parameters and being a function for mapping a relative velocity to a pitch coefficient modification value.

[0139] EEE19. The method of EEE18, wherein the indication of the first and second parameter values ​​is encoded as a label in a bitstream.

[0140] EEE20. The method of EEE18 or 19, wherein the indication of the first and second parameter values ​​is encoded together with the audio content in a single bitstream or as a separate bitstream.

[0141] EEE21. The method of any one of EEE18 to 20, wherein the first and second parameter values ​​are determined by a content creator or a game engine.

[0142] EEE22. An audio renderer comprising a processor and a memory coupled to the processor, the processor adapted to cause the audio renderer to perform a method as set forth in any one of EEE1 to EEE17.

[0143] EEE23. An encoder comprising a processor and a memory coupled to the processor, the processor adapted to cause the encoder to perform a method according to any one of EEE18 to EEE21.

[0144] EEE24. A program comprising instructions which, when executed by a processor, cause the processor to perform the method of any one of EEE1 to EEE21.

[0145] A computer-readable storage medium storing a program according to EEE25.EEE24.

Claims

1. 1. A method for modeling the Doppler effect when rendering audio content at a user site for a six degree of freedom (6DoF) environment, comprising: obtaining a second parameter value of a second parameter indicative of a desired strength of the Doppler effect to be modeled; determining a pitch correction value based on a relative velocity between a listener and a sound source in the audio content and first and second parameters, the first parameter indicating an acceptable range of pitch correction values ​​and the second parameter indicating a desired strength of the Doppler effect to be modeled, the pitch correction value being determined based on a pitch correction function; and rendering the sound source based on the pitch modification value; the predetermined pitch correction function is a function for mapping relative velocity to a pitch correction value; The method of claim 1, wherein the second parameter controls the slope of the pitch correction function, which reflects the aggressiveness of the Doppler effect to be modeled.

2. The method of claim 1 , wherein the relative velocity is calculated based on the positions of the listener and the sound source.

3. The method of claim 1 , wherein if the allowed range of pitch correction values ​​is not supported by an audio renderer, a default range of pitch correction values ​​is used by the audio renderer instead.

4. The method of claim 1 , wherein the audio content is extracted from a received bitstream, and the first and second parameters are derived from instructions included in the bitstream.

5. An audio renderer comprising a processor and a memory coupled to the processor, the processor adapted to cause the audio renderer to perform the method of any one of claims 1 to 4.

6. A program comprising instructions which, when executed by a processor, cause the processor to carry out the method of any one of claims 1 to 4.

7. A computer-readable storage medium storing the program according to claim 6.