Method, apparatus and system for pre-rendering signals for audio rendering
By pre-rendering audio elements and using a simple distance attenuation model, the problem of insufficient sound field reproduction in virtual/augmented/mixed reality spaces by 6DoF audio renderers is solved, achieving high-quality audio reproduction and an immersive experience.
Patent Information
- Application Number
- JP2023179225
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-05
- Filing Date
- 2023-10-18
- Publication Date
- 2025-12-15
- Estimated Expiration
- 2039-04-08
AI Technical Summary
Existing 6DoF audio renderers cannot accurately reproduce the sound field desired by content creators in virtual/augmented/mixed reality spaces, mainly due to insufficient descriptive information about sound sources and environments, renderer functional limitations, and data and resource constraints.
By using encoding and decoding methods, pre-rendered audio elements are used to replace or supplement the original audio signal. Combined with position and direction information, effective audio elements are generated. A simple distance attenuation model is applied for rendering, and a predetermined rendering mode is used in the decoder to control the influence of the acoustic environment, ensuring high-quality audio reproduction.
It achieves high-quality reproduction of sound fields in virtual/augmented/mixed reality under limited resources and bandwidth conditions, preserving artistic intent and providing an immersive auditory experience.
Smart Images

Figure 0007785733000008 
Figure 0007785733000009 
Figure 0007785733000010
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 62 / 656,163 (Docket No. D18040USP1), filed April 11, 2018, and U.S. Provisional Application No. 62 / 755,957 (Docket No. D18040USP2), filed November 5, 2018, which are incorporated herein by reference.
[0002] Technical Field The present disclosure relates to providing apparatus, systems and methods for audio rendering. [Background technology]
[0003] FIG. 1 shows an example encoder configured to process metadata and audio renderer extensions.
[0004] In some cases, 6DoF renderers are unable to reproduce the sound field desired by the content creator at a certain position(s) (area, trajectory) in the virtual / augmented / mixed reality (VR / AR / MR) space. This is due to the following reasons: 1. Insufficient metadata describing the sound source and the VR / AR / MR environment Limited functionality of the 2.6DoF renderer and resources.
[0005] Some 6DoF renderers (which generate a sound field based only on the original audio source signal and the VR / AR / MR environment description) may not be able to reproduce the intended signal at the desired position due to the following reasons: 1.1) Bitrate limits for parameterized information (metadata) describing the VR / AR / MR environment and the corresponding audio signal; 1.2) Data for inverse 6DoF rendering is not available (e.g., a reference recording at one or more points of interest is available, but it is unclear how this signal is to be reproduced by a 6DoF renderer and what data inputs are required for this); 2.1) Artistic intent that may differ from the default (e.g., physics-consistent) output of a 6DoF renderer (e.g., similar to the concept of "artistic downmix"); 2.2) Functional limitations of the decoder (6DoF renderer) implementation (e.g., constraints on bitrate, complexity, latency, etc.).
[0006] At the same time, for a given position(s) in the VR / AR / MR space, one may require high audio quality (and / or fidelity to a predefined reference signal) of audio playback (i.e., 6DoF renderer output). For example, this may be required due to 3DoF / 3DoF+ compatibility constraints or compatibility requirements for different processing modes of 6DoF rendering (e.g., between a "baseline" mode and a "low power" mode that does not consider the effects of VR / AR / MR geometry). Summary of the Invention [Problem to be solved by the invention]
[0007] Thus, there is a need for an encoding / decoding method and corresponding encoder / decoder that improves the reproduction of a content creator's desired sound field in a VR / AR / MR space. [Means for solving the problem]
[0008] An aspect of the present disclosure relates to a method for decoding audio scene content from a bitstream by a decoder including an audio renderer with one or more rendering tools. The method may include receiving a bitstream. The method may further include decoding an audio scene description from the bitstream. The audio scene may include an acoustic environment, such as a VR / AR / MR acoustic environment. The method may further include determining one or more active audio elements from the audio scene description. The method may further include determining active audio element information indicating active audio element positions of the one or more active audio elements from the audio scene description. The method may further include decoding a rendering mode indication from the bitstream. The rendering mode indication may indicate whether the one or more active audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode. The method may further include, in response to the rendering mode instruction indicating that the one or more enabled audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode, rendering the one or more enabled audio elements using the predetermined rendering mode. Rendering the one or more enabled audio elements using the predetermined rendering mode may take into account the enabled audio element information. The predetermined rendering mode may define a predetermined configuration of rendering tools for controlling the influence of the acoustic environment of the audio scene on the rendered output. The enabled audio elements may, for example, be rendered at a reference position. The predetermined rendering mode may enable or disable certain rendering tools.The predetermined rendering mode may also enhance the sound effects (eg, add artificial sounds) for the one or more enabled audio elements.
[0009] The one or more effective audio elements encapsulate, so to speak, the effects of the acoustic environment, such as echo, reverberation, and acoustic masking. This allows the use of a particularly simple rendering mode (i.e., the predetermined rendering mode) in the decoder. At the same time, artistic intent can be preserved, and the user (listener) can be provided with a rich, immersive acoustic experience, even with a low-power decoder. Furthermore, the decoder's rendering tools can be individually configured based on the rendering mode instructions, which provide additional control over the acoustic effects. Ultimately, encapsulating the effects of the acoustic environment allows for efficient compression of metadata describing the acoustic environment.
[0010] In some embodiments, the method may further include obtaining listener position information indicating a position of a listener's head in the acoustic environment and / or listener orientation information indicating an orientation of the listener's head in the acoustic environment. A corresponding decoder may include an interface for receiving the listener position information and / or listener orientation information. Rendering the one or more enabled audio elements using a predetermined rendering mode may then further take the listener position information and / or listener orientation information into account. By referring to this additional information, the user's acoustic experience may be more immersive and meaningful.
[0011] In some embodiments, the effective audio element information may include information indicative of a sound radiation pattern of each of the one or more effective audio elements. In that case, rendering the one or more effective audio elements using the predetermined rendering mode may further take into account the information indicative of the sound radiation pattern of each of the one or more effective audio elements. For example, an attenuation factor may be calculated based on the sound radiation pattern of each effective audio element and the relative position between each effective audio element and a listener position. By taking radiation patterns into account, the user's acoustic experience may be more immersive and meaningful.
[0012] In some embodiments, rendering the one or more active audio elements using the predetermined rendering mode may apply sound attenuation modeling according to the respective distances between the listener position and the active audio element positions of the one or more active audio elements. That is, the predetermined rendering mode may (only) apply sound attenuation modeling (in empty space) without considering any acoustic elements in the acoustic environment. This defines a simple rendering mode that can be applied even by low-power decoders. Additionally, sound directivity modeling may be applied, for example, based on the sound radiation pattern of the one or more active audio elements.
[0013] In some embodiments, at least two active audio elements may be determined from the audio scene description. The rendering mode instruction may indicate a respective predetermined rendering mode for each of the at least two active audio elements. The method may further include rendering the at least two active audio elements using the respective predetermined rendering modes. Rendering each active audio element using the respective predetermined rendering mode may take into account active audio element information for that active audio element. The predetermined rendering modes for the active audio elements may further define respective predetermined configurations of rendering tools for controlling the influence of the audio scene's acoustic environment on the rendering output for that active audio element. This may provide additional control over the sound effects applied to individual active audio elements, thereby enabling a closer match to a content creator's artistic intent.
[0014] In some embodiments, the method may further include determining one or more original audio elements from the audio scene description. The method may further include determining audio element information indicating audio element positions of the one or more audio elements from the audio scene description. The method may further include rendering the one or more audio elements using a rendering mode for the one or more audio elements that differs from a predetermined rendering mode used for the one or more active audio elements. Rendering the one or more audio elements using the rendering mode for the one or more audio elements may take into account the audio element information. The rendering may further take into account the influence of the acoustic environment on the rendered output. Thus, active audio elements encapsulating the influence of the acoustic environment can be rendered using, for example, a simple rendering mode, while the (original) audio elements can be rendered using a more sophisticated, for example, reference, rendering mode.
[0015] In some embodiments, the method may further include obtaining listener position region information indicating listener position regions for which the predetermined rendering mode is to be used. The listener position region information may, for example, be encoded in the bitstream, thereby ensuring that the predetermined rendering mode is used only for listener position regions for which enabled audio elements provide a meaningful representation of (e.g., of) the original audio scene.
[0016] In some embodiments, the predetermined rendering mode indicated by the rendering mode indication may depend on a listener position. Further, the method may include rendering the one or more enabled audio elements using the predetermined rendering mode indicated by the rendering mode indication for the listener position region indicated by the listener position region information. That is, the rendering mode indication may indicate different (predetermined) rendering modes for different listener position regions.
[0017] Another aspect of the present disclosure relates to a method for generating audio scene content. The method may include obtaining one or more audio elements representing a captured signal from an audio scene. The method may further include obtaining effective audio element information indicating effective audio element positions of one or more effective audio elements to be generated. The method may further include determining the one or more effective audio elements from the one or more audio elements representing the captured signal by applying sound attenuation modeling according to a distance between a position at which the captured signal was captured and the effective audio element positions of the one or more effective audio elements.
[0018] This method allows the generation of audio scene content that, when rendered at a reference or capture position, gives a close perceptual approximation of the sound field that would emanate from the original audio scene. However, this audio scene content can be rendered at a listener position that is different from the reference or capture position, thus allowing for an immersive acoustic experience.
[0019] Another aspect of the present disclosure relates to a method for encoding audio scene content into a bitstream. The method may include receiving a description of an audio scene. The audio scene may include an acoustic environment and one or more audio elements at respective audio element positions. The method may further include determining one or more effective audio elements at respective effective audio element positions from the one or more audio elements. This determination may be performed by rendering the one or more effective audio elements at each effective audio element position to a reference position using a rendering mode that does not take into account the effect of the acoustic environment on the rendering output (e.g., applies distance attenuation modeling in empty space), thereby providing a psychoacoustic approximation of a reference sound field at the reference position that would result from rendering the one or more audio elements at each audio element position to the reference position using a reference rendering mode that takes into account the effect of the acoustic environment on the rendering output. The method may further include generating effective audio element information indicating the effective audio element positions of the one or more effective audio elements. The method may further include generating a rendering mode instruction indicating that the one or more active audio elements should be rendered using a predetermined rendering mode representing a sound field obtained from pre-rendered audio elements and specifying a predetermined configuration of a decoder's rendering tool for controlling an effect of an acoustic environment on a rendered output at the decoder. The method may further include encoding the one or more audio elements, the audio element positions, the one or more active audio elements, the active audio element information, and the rendering mode instruction into a bitstream.
[0020] The one or more effective audio elements encapsulate the effects of the acoustic environment, such as echo, reverberation, and acoustic hiding. This allows the use of a particularly simple rendering mode (i.e., the predetermined rendering mode) in the decoder. At the same time, artistic intent can be preserved, and the user (listener) can be provided with a rich, immersive acoustic experience, even with a low-power decoder. Furthermore, the decoder's rendering tools can be individually configured based on the rendering mode instructions, which provide additional control over the acoustic effects. Ultimately, encapsulating the effects of the acoustic environment enables efficient compression of metadata describing the acoustic environment.
[0021] In some embodiments, the method may further include obtaining listener position information indicative of a position of a listener's head in the acoustic environment and / or listener orientation information indicative of an orientation of the listener's head in the acoustic environment, and may further include encoding the listener position information and / or listener orientation information into a bitstream.
[0022] In some embodiments, the active audio element information may be generated to include information indicative of a sound radiation pattern for each of the one or more active audio elements. In some embodiments, at least two active audio elements may be generated and encoded into the bitstream, and the rendering mode indication may indicate a respective predetermined rendering mode for each of the at least two active audio elements.
[0023] In some embodiments, the method may further include obtaining listener position region information indicating a listener position region for which the predetermined rendering mode is to be used. The method may further include encoding the listener position region information into a bitstream.
[0024] In some embodiments, the predetermined rendering mode indicated by the rendering mode indication may depend on the listener position, such that the rendering mode indication indicates a respective predetermined rendering mode for each of a plurality of listener positions.
[0025] Another aspect of the present disclosure relates to an audio decoder including a processor coupled to a memory storing instructions for the processor, the processor being adapted to perform a method according to each of the above aspects or embodiments.
[0026] Another aspect of the present disclosure relates to an audio encoder including a processor coupled to a memory storing instructions for the processor, the processor may be configured to perform a method according to each of the above aspects or embodiments.
[0027] Further aspects of the present disclosure relate to corresponding computer programs and computer-readable storage media.
[0028] It will be understood that method steps and apparatus features may be interchanged in many ways. In particular, details of a disclosed method may be implemented as an apparatus adapted to perform some or all of the steps of the method, and vice versa, as will be understood by those skilled in the art. In particular, it will be understood that each statement made with respect to a method applies equally to the corresponding apparatus, and vice versa. [Brief explanation of the drawings]
[0029] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numbers indicate like or similar elements. [Figure 1] FIG. 1 is a schematic diagram illustrating an example of an encoder / decoder system. [Figure 2] An example of an audio scene is shown schematically. [Figure 3]Schematic representation of an example of the location of an audio scene in an acoustic environment. [Figure 4] 1 illustrates a schematic diagram of an example encoder / decoder system according to an embodiment of the present disclosure. [Figure 5] 10 illustrates a schematic diagram of another example of an encoder / decoder system according to an embodiment of the present disclosure. [Figure 6] 1 is a flowchart that schematically illustrates an example of a method for encoding audio scene content according to an embodiment of the present disclosure. [Figure 7] 1 is a flowchart that schematically illustrates an example of a method for decoding audio scene content according to an embodiment of the present disclosure. [Figure 8] 1 is a flowchart that schematically illustrates an example of a method for generating audio scene content according to an embodiment of the present disclosure. [Figure 9] 9 illustrates a schematic example of an environment in which the method of FIG. 8 can be implemented. [Figure 10] 1 illustrates a schematic diagram of an example environment for testing the output of a decoder according to an embodiment of the present disclosure. [Figure 11] 3A-3C illustrate schematic diagrams of examples of data elements transferred within a bitstream according to embodiments of the present disclosure. [Figure 12] Schematic examples of different rendering modes are shown with reference to an audio scene. [Figure 13] 10A and 10B illustrate schematic examples of encoder and decoder processing according to embodiments of the present disclosure with reference to an audio scene. [Figure 14] 10A and 10B illustrate schematic examples of rendering effective audio elements to different listener positions according to an embodiment of the present disclosure. [Figure 15] 1A and 1B illustrate schematic examples of audio elements, active audio elements, and listener positions in an acoustic environment, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0030] As mentioned above, the same or similar reference numerals in this disclosure indicate the same or similar elements, and repeated descriptions thereof may be omitted for reasons of brevity.
[0031] This disclosure relates to a VR / AR / MR renderer or audio renderer (e.g., an audio renderer whose rendering is compatible with the MPEG audio standard). This disclosure further relates to an artistic pre-rendering concept in which an encoder provides a quality and bitrate-efficient representation of the sound field within a pre-defined 3DoF+ region(s).
[0032] In one example, a 6DoF audio renderer may output a match to a reference signal (sound field) at a specific position(s). The 6DoF audio renderer may also extend to convert VR / AR / MR related metadata into a native format such as the MPEG-H 3D audio renderer input format.
[0033] The aim is to provide a standards-compliant (e.g., MPEG-compliant or future MPEG-compliant) audio renderer to generate audio output as predefined reference signals at 3DoF position(s).
[0034] A straightforward approach to support such requirements would be to forward predefined (pre-rendered) signal(s) directly to the decoder / renderer side. This approach has the following obvious drawbacks: 1. Increased bitrate (i.e., pre-rendered signal(s) are sent in addition to the original audio source signal); 2. Limited validity (i.e., the pre-rendered signal(s) are only valid for the 3DoF position(s)).
[0035] Broadly speaking, the present disclosure relates to efficiently generating, encoding, decoding, and rendering such a signal(s) to provide 6DoF rendering capabilities. Accordingly, the present disclosure describes methods to overcome the aforementioned drawbacks, including: 1. Using a pre-rendered signal in place of (or as a complementary addition to) the original audio source signal; 2. Increased applicability range (using 6DoF rendering) from 3DoF position(s) to 3DoF+ region for pre-rendered signal(s) by preserving a high level of sound field approximation.
[0036] An exemplary scenario in which the present disclosure is applicable is shown in Figure 2. Figure 2 shows an exemplary space, e.g., an elevator, and a listener. In one example, the listener may be standing in front of an elevator that is opening and closing its doors. Inside the elevator car, there are several people talking and ambient music. The listener can move around, but cannot enter the elevator cabin. Figure 2 shows a plan view and a front view of the elevator system.
[0037] Thus, the elevator and sound sources in Figure 2 (people talking, ambient music) can be said to define an audio scene.
[0038] In general, an audio scene in the context of this disclosure is understood to mean all audio elements, acoustic elements, and acoustic environments necessary to render the sounds in a scene, i.e., the input data required by an audio renderer (e.g., an MPEG-I audio renderer). In the context of this disclosure, an audio element is understood to mean one or more audio signals and associated metadata. An audio element may be, for example, an audio object, a channel, or an HOA signal. An audio object is understood to mean an audio signal with associated static / dynamic metadata containing information necessary to reproduce the sound of an audio source. An acoustic element is understood to mean a physical object in space that interacts with the audio element and affects the rendering of the audio element based on the user's position and orientation. An acoustic element may share metadata with the audio object (e.g., position and orientation). An acoustic environment is understood to mean metadata describing the acoustic characteristics of the virtual scene to be rendered, e.g., a room or local area.
[0039] For such scenarios (or indeed any other audio scene), it is desirable for an audio renderer to be able to render a sound field representation of the audio scene that is a faithful representation of the original sound field at least at the reference position, that satisfies the artistic intent, and / or that rendering is feasible within the (limited) rendering capabilities of the audio renderer, and that furthermore satisfies any bitrate limitations on the transmission of the audio content from encoder to decoder.
[0040] FIG. 3 shows a schematic overview of an audio scene associated with a listening environment. The audio scene includes an acoustic environment 100. The acoustic environment 100 includes one or more audio elements 102 at respective positions. The one or more audio elements may be used to generate one or more effective audio elements 101 at respective positions that are not necessarily equal to the positions of the one or more audio elements. For example, for a given set of audio elements, the position of the effective audio element may be set to the center (e.g., center of gravity) of the positions of the audio elements. The generated effective audio elements may have the property that rendering the effective audio elements at a reference position 111 within a listener position region 110 using a predetermined rendering function (e.g., a simple rendering function that simply applies distance attenuation in empty space) results in a sound field at the reference position 111 that is (substantially) perceptually equivalent to the sound field at the reference position 111 that would result from rendering the audio elements 102 using a reference rendering function (e.g., a rendering function that takes into account the characteristics (e.g., impact) of the acoustic environment, including acoustic elements (e.g., echo, reverberation, occlusion, etc.)). Of course, once generated, the useful audio element 101 may also be rendered using a predetermined rendering function to a listener position 112 within the listener position region 110 that is different from the reference position 111. The listener position may be at a distance 103 from the position of the useful audio element 101. An example for generating the useful audio element 101 from the audio element 102 is described in more detail below.
[0041] In some embodiments, the effective audio element 102 may alternatively be determined based on one or more captured signals 120 captured at a capture position within the listener position area 110. For example, a user in the audience of a musical performance may capture sound emanating from an audio element (e.g., a musician) on stage. Then, given a desired position of the effective audio element (e.g., its position relative to the capture position, such as by specifying a distance 121 between the effective audio element 101 and the capture position, possibly in relation to an angle indicating the direction of a distance vector between the effective audio element 101 and the capture position), the effective audio element 101 can be generated based on the captured signals 120. The generated effective audio element 101 may have properties such that rendering the effective audio element 101 at a reference position 111 (not necessarily equal to the capture position) using a predetermined rendering function (e.g., a simple rendering function that simply applies distance attenuation in empty space) provides a sound field that is (substantially) perceptually equivalent to the sound field at the reference position 111 emanating from the original audio element 102 (e.g., a musician). Examples of such use cases are described in more detail below.
[0042] In particular, the reference location 111 may in some cases be the same as the capture location, and the reference signal (i.e., the signal at the reference location 111) may be equal to the capture signal 120. This is a valid assumption for VR / AR / MR applications where a user may have an avatar in-head recording option. In real-world applications, this assumption may not be valid because the reference receiver is the user's ear, and the signal capture device (e.g., a cell phone or microphone) may be quite far away from the user's ear. Next, methods and apparatus for addressing the first-mentioned need are described.
[0043] 4 illustrates an example of an encoder / decoder system according to an embodiment of the present disclosure. An encoder 210 (e.g., an MPEG-I encoder) outputs a bitstream 220 that can be used by a decoder 230 (e.g., an MPEG-I decoder) to generate an audio output 240. The decoder 230 can further receive listener information 233. The listener information 233 is not necessarily included in the bitstream 220 and can come from any source. For example, the listener information may be generated and output by a head tracker and input to a (dedicated) interface of the decoder 230.
[0044] The decoder 230 includes an audio renderer 250, which includes one or more rendering tools 251. In the context of this disclosure, an audio renderer is understood to mean a canonical audio rendering module, e.g., MPEG-I, including the rendering tool and interfaces to external rendering tools and interfaces to system layers for external resources. A rendering tool is understood to mean a component of the audio renderer that performs aspects of rendering, e.g., room model parameterization, occlusion, reverberation, binaural rendering, etc.
[0045] The renderer 250 receives as input one or more enabled audio elements, enabled audio element information 231, and rendering mode instructions 232. The enabled audio elements, enabled audio element information, and rendering mode instructions 232 are described in more detail below. The enabled audio element information 231 and rendering mode instructions 232 may be derived (e.g., determined / decoded) from the bitstream 220. The renderer 250 renders a representation of the audio scene based on the enabled audio elements and enabled audio element information using one or more rendering tools 251. Here, the rendering mode instructions 232 indicate the rendering modes in which the one or more rendering tools 251 operate. For example, certain rendering tools 251 may be activated or deactivated according to the rendering mode instructions 232. Furthermore, certain rendering tools 251 may be configured according to the rendering mode instructions 232. For example, control parameters of certain rendering tools 251 may be selected (e.g., set) according to the rendering mode instructions 232.
[0046] In the context of this disclosure, an encoder (e.g., an MPEG-I encoder) has the tasks of determining 6DoF metadata and control data, determining enabled audio elements (e.g., including a mono audio signal for each enabled audio element), determining positions (e.g., x, y, z) for the enabled audio elements, and determining data for controlling rendering tools (e.g., enable / disable flags and configuration data). The data for controlling rendering tools can correspond to, include, or be included in the aforementioned rendering mode instructions.
[0047] In addition to the above, an encoder according to an embodiment of the present disclosure may minimize the perceptual difference of the output signal 240 relative to the reference signal R (if present) for the reference position 111. That is, for the rendering tool / rendering function F() used by the decoder, the processed signal A, and the position (x,y,z) of the active audio element, the encoder may perform the following optimization: {x,y,z;F}:||Output (参照位置) (F(x,y,z)(A))-R|| 知覚的 ->min
[0048] Furthermore, an encoder according to embodiments of the present disclosure may assign the "direct" portions of the processed signal A to the estimated locations of the original objects 102. For a decoder, this means, for example, that the decoder can recreate several useful audio elements 101 from a single captured signal 120.
[0049] In some embodiments, an MPEG-H 3D audio renderer extended with simple distance modeling for 6DoF may be used, where the effective audio element positions are expressed using azimuth, elevation, radius, and the rendering tool F(). The audio element positions and gains can be obtained manually (e.g., by encoder tuning) or automatically (e.g., by brute-force optimization).
[0050] FIG. 5 illustrates a schematic diagram of another example of an encoder / decoder system according to an embodiment of the present disclosure.
[0051] The encoder 210 receives the audio scene A instruction (processed signal), which is then encoded (e.g., MPEG-H encoding) in a manner described in this disclosure. In addition, the encoder 210 may generate metadata (e.g., 6DoF metadata) containing information about the acoustic environment. The encoder may also generate, possibly as part of the metadata, rendering mode instructions for configuring the rendering tools of the audio renderer 250 of the decoder 230. The rendering tools may include, for example, signal modification tools for the enabled audio elements. Depending on the rendering mode instruction, individual rendering tools of the audio renderer may be activated or deactivated. For example, if the rendering mode instruction indicates that the enabled audio element should be rendered, the signal modification tool may be activated while all other rendering tools are deactivated. The decoder 230 outputs an audio output 240, which can be compared to a reference signal R resulting from rendering the original audio element to the reference position 111 using a reference rendering function. An example of an arrangement for comparing the audio output 240 with a reference signal R is shown schematically in FIG.
[0052] FIG. 6 is a flowchart illustrating an example method 600 for encoding audio scene content into a bitstream, according to an embodiment of the present disclosure.
[0053] Step S610 In the audio scene description, a description of an audio scene is received, the audio scene including an acoustic environment and one or more audio elements at respective audio element positions.
[0054] Step S620In the method, one or more effective audio elements at each effective audio element position are determined from the one or more audio elements. The one or more effective audio elements are determined so as to provide a psychoacoustic approximation of a reference sound field at the reference position, which would result from rendering the one or more effective audio elements at each effective audio element position to a reference position using a rendering mode that does not consider the effect of the acoustic environment on the rendering output, and then rendering the one or more (original) audio elements at each audio element position to the reference position using a reference rendering mode that does consider the effect of the acoustic environment on the rendering output. The effect of the acoustic environment may include echo, reverberation, reflection, etc. The rendering mode that does not consider the effect of the acoustic environment on the rendering output may apply distance attenuation modeling (in empty space). Non-limiting examples of methods for determining such effective audio elements are described further below.
[0055] Step S630 Then, valid audio element information indicating valid audio element positions of the one or more valid audio elements is generated.
[0056] Step S640 a rendering mode instruction is generated indicating that the one or more enabled audio elements should be rendered using a predetermined rendering mode representing a sound field obtained from pre-rendered audio elements and defining a predetermined configuration of a decoder's rendering tools for controlling the effect of the acoustic environment on the rendering output at the decoder.
[0057] Step S650 In the bitstream, the one or more audio elements, the audio element positions, the one or more active audio elements, the active audio element information, and the rendering mode indication are encoded into a bitstream.
[0058] In the simplest case, a rendering mode instruction may be a flag indicating that all acoustics (i.e., the effects of the acoustic environment) are included (i.e., encapsulated) in the one or more enabled audio elements. Thus, a rendering mode instruction may instruct a decoder (or a decoder's audio renderer) to use a simple rendering mode in which only distance attenuation is applied (e.g., by multiplication with a distance-dependent gain) and all other rendering tools are deactivated. In more sophisticated cases, a rendering mode instruction may include one or more control values for configuring rendering tools. This may include activation and deactivation of individual rendering tools, but may also include more granular control of rendering tools. For example, rendering tools may be configured by the rendering mode instruction to enhance acoustic effects when rendering the one or more enabled audio elements. This may be used, for example, to add (artificial) acoustic effects such as echo, reverberation, reflection, etc., according to artistic intent (e.g., of the content creator).
[0059] In other words, method 600 may relate to a method for encoding audio data, the audio data representing one or more audio elements (e.g., representations of physical objects) at respective audio element positions within an acoustic environment that includes one or more acoustic elements. The method may include determining an effective audio element at an effective audio element position in the acoustic environment such that rendering the effective audio element to a reference position approximates a reference sound field at the reference position that would result from reference rendering of the one or more audio elements at the respective audio element positions to the reference position when using a rendering function that takes into account distance attenuation between the effective audio element position and a reference position but does not take into account the acoustic elements in the acoustic environment. The effective audio element and the effective audio element position may then be encoded into a bitstream.
[0060] In the above situation, determining the effective audio element at the effective audio element position may involve: rendering the one or more audio elements at a reference position in the acoustic environment using a first rendering function, thereby obtaining a reference sound field at the reference position, wherein the first rendering function takes into account the acoustic elements in the acoustic environment as well as the distance attenuation between the acoustic element position and the reference position; and determining the effective audio element at the effective audio element position in the acoustic environment based on the reference sound field at the reference position, such that rendering the effective audio element at the reference position using a second rendering function approximates the reference sound field, wherein the second rendering function takes into account the distance attenuation between the effective audio element position and the reference position but does not take into account the acoustic elements in the acoustic environment.
[0061] The method 600 described above may relate to 0DoF use cases where there is no listener data. In general, the method 600 supports the concept of a "smart" encoder and a "simple" decoder.
[0062] With respect to listener data, method 600 in some implementations may include obtaining listener position information indicating the position of the listener's head in the acoustic environment (e.g., in a listener position region). Additionally or alternatively, method 600 may include obtaining listener orientation information indicating the orientation of the listener's head in the acoustic environment (e.g., in a listener position region). The listener position information and / or listener orientation information may then be encoded into the bitstream. The listener position information and / or listener orientation information may be used by a decoder to render the one or more effective audio elements accordingly. For example, the decoder may render the one or more effective audio elements at the listener's actual position (rather than a reference position). Similarly, particularly for headphone applications, the decoder may perform a rotation of the rendered sound field depending on the orientation of the listener's head.
[0063] In some implementations, the method 600 may generate the active audio element information to include information indicating a sound radiation pattern of each of the one or more active audio elements. This information may then be used by a decoder to render the one or more active audio elements accordingly. For example, when rendering the one or more active audio elements, the decoder may apply a respective gain to each of the one or more active audio elements. These gains may be determined based on the respective radiation patterns. Each gain may be determined based on the angle between a distance vector between the respective active audio element and the listener position (or a reference position, if rendering to a reference position is performed) and a radiation direction vector indicating the radiation direction of the respective audio element. For more complex radiation patterns having multiple radiation direction vectors and corresponding weighting factors, the gain may be determined based on a weighted sum of each gain determined based on the angle between the distance vector and each radiation direction vector. The weight in the sum may correspond to the weighting factor. The gain determined based on the radiation pattern may be added to a distance attenuation gain applied by a given rendering mode.
[0064] In some implementations, at least two active audio elements may be generated and encoded into the bitstream. The rendering mode indication may then indicate a respective predetermined rendering mode for each of the at least two active audio elements. The at least two predetermined rendering modes may be different, thereby allowing different amounts of sound effects to be shown for different active audio elements, for example, according to the artistic intent of a content creator.
[0065] In some implementations, the method 600 may further include obtaining listener location region information indicating a listener location region for which a predetermined rendering mode is to be used. This listener location region information may then be encoded into the bitstream. At the decoder, if the listener location for which rendering is desired is within the listener location region indicated by the listener location information, the predetermined rendering mode should be used. Otherwise, the decoder may apply a rendering mode of its choice, such as a default rendering mode.
[0066] Furthermore, different predetermined rendering modes may be foreseen depending on the listener position for which rendering is desired. Thus, the predetermined rendering mode indicated by the rendering mode instruction may depend on the listener position, and the rendering mode instruction indicates a respective predetermined rendering mode for each of a plurality of listener positions. Similarly, different predetermined rendering modes may be foreseen depending on the listener position region for which rendering is desired. In particular, there may be different available audio elements for different listener positions (or listener position regions). Providing such rendering mode instructions enables control of (artificial) acoustics, such as (artificial) echoes, reverberations, and reflections, applied to each listener position (or listener position region).
[0067] 7 is a flowchart illustrating an example of a corresponding method 700 for decoding audio scene content from a bitstream by a decoder, according to an embodiment of the present disclosure. The decoder may include an audio renderer having one or more rendering tools.
[0068] Step S710 In, a bitstream is received. Step S720 In , the audio scene description is decoded from the bitstream. Step S730 In , one or more enabled audio elements are determined from the audio scene description.
[0069] Step S740 In the audio scene description, active audio element information indicating active audio element positions of one or more active audio elements is determined from the audio scene description.
[0070] Step S750 In the bitstream, a rendering mode indication is decoded, the rendering mode indication indicating whether the one or more enabled audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode.
[0071] Step S760 In response to the rendering mode instruction indicating that the one or more enabled audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using the predetermined rendering mode, the one or more enabled audio elements are rendered using the predetermined rendering mode. Rendering the one or more enabled audio elements using the predetermined rendering mode takes into account the enabled audio element information. Furthermore, the predetermined rendering mode defines a predetermined configuration of the rendering tool for controlling the impact of the acoustic environment of the audio scene on the rendered output.
[0072] In some implementations, method 700 may include obtaining listener position information indicative of a position of a listener's head in an acoustic environment (e.g., in a listener position region) and / or listener orientation information indicative of an orientation of a listener's head in an acoustic environment (e.g., in a listener position region). Rendering the one or more enabled audio elements using the predetermined rendering mode may then further take the listener position information and / or listener orientation information into account, e.g., in the manner described above with reference to method 600. A corresponding decoder may include an interface for receiving the listener position information and / or listener orientation information.
[0073] In some implementations of method 700, the active audio element information may include information indicative of a sound radiation pattern for each of the one or more active audio elements. Rendering the one or more active audio elements using the predetermined rendering mode may further take into account information indicative of a sound radiation pattern for each of the one or more active audio elements, e.g., in the manner described above with reference to method 600.
[0074] In some implementations of method 700, rendering the one or more active audio elements using the predetermined rendering mode may apply sound attenuation modeling (in empty space) depending on the respective distance between the listener position and the active audio element position of the one or more active audio elements. Such a predetermined rendering mode is referred to as a simple rendering mode. The simple rendering mode (i.e., distance attenuation in empty space only) can be applied because the effects of the acoustic environment are "encapsulated" in the one or more active audio elements. By doing so, part of the decoder's processing load can be delegated to the encoder, enabling even low-power decoders to render an immersive sound field consistent with the artistic intent.
[0075] In some implementations of method 700, at least two active audio elements may be determined from the audio scene description. The rendering mode instruction may then indicate a respective predetermined rendering mode for each of the at least two active audio elements. In such a situation, method 700 may further include rendering the at least two active audio elements using the respective predetermined rendering modes. Rendering each active audio element using the respective predetermined rendering modes may take into account active audio element information for that active audio element, and the rendering modes for that active audio element may define respective predetermined configurations of a rendering tool for controlling the influence of the acoustic environment of the audio scene on the rendering output for that active audio element. The at least two predetermined rendering modes may be distinguishable, thereby indicating different amounts of sound effects for different active audio elements, for example, according to the artistic intent of a content creator.
[0076] In some implementations, both the active audio elements and the (actual / original) audio elements may be encoded in the bitstream to be decoded. In this case, method 700 may include determining one or more audio elements from an audio scene description and determining audio element information indicating audio element positions of the one or more audio elements from the audio scene description. Rendering the one or more audio elements is performed using a rendering mode for the one or more audio elements that differs from the predetermined rendering mode used for the one or more active audio elements. Rendering the one or more audio elements using the rendering mode for the one or more audio elements may take audio element information into account. This allows the active audio elements to be rendered, for example, in the simple rendering mode while the (actual / original) audio elements are rendered, for example, in the reference rendering mode. Furthermore, the predetermined rendering mode can be configured separately from the rendering mode used for the audio elements. More generally, the rendering modes for audio elements and active audio elements may imply different configurations of associated rendering tools. Acoustic rendering (taking into account the effects of the acoustic environment) may be applied to the audio elements, while distance attenuation modeling (in empty space) may be applied to the active audio elements, possibly together with artificial acoustics (not necessarily determined by the acoustic environment assumed for encoding).
[0077] In some implementations, the method 700 may further include obtaining listener location region information indicating a listener location region for which a predetermined rendering mode is to be used. For rendering to listener positions indicated by the listener location region information within the listener location region, the predetermined rendering mode should be used. Otherwise, the decoder may apply a rendering mode of its choice (which may be implementation-dependent), such as a default rendering mode.
[0078] In some implementations of method 700, the predetermined rendering mode indicated by the rendering mode indication may depend on the listener position (or listener position region), and the decoder may then perform rendering of the one or more enabled audio elements using the predetermined rendering mode indicated by the rendering mode indication for the listener position region indicated by the listener position region information.
[0079] FIG. 8 is a flow chart illustrating an example of a method 800 for generating audio scene content.
[0080] Step S810 In the audio scene, one or more audio elements are obtained that represent captured signals from the audio scene. This can be done, for example, by sound capture using a microphone or a mobile device with recording capabilities.
[0081] Step S820 In the step S100, valid audio element information indicating valid audio element positions of one or more valid audio elements to be generated is obtained. The valid audio element positions may be estimated or may be received as user input.
[0082] Step S830wherein the one or more effective audio elements are determined from the one or more audio elements representing a captured signal by applying sound attenuation modeling according to the distance between the position at which the captured signal was captured and the effective audio element position of the one or more effective audio elements.
[0083] Method 800 enables real-world A( / V) recording of captured audio signals 120 representing audio elements 102 from discrete capture locations (see FIG. 3). Methods and apparatus according to the present disclosure enable consumption of this material (e.g., with as meaningful a user experience as possible using 3DoF+, 3DoF, 0DoF platforms) from a reference location 111 or other locations 112 and orientations (i.e., orientations in a 6DoF framework) within a listener location region 110. This is shown schematically in FIG. 9.
[0084] One non-limiting example for determining the effective audio elements from the (actual / original) audio elements in an audio scene follows.
[0085] As mentioned above, embodiments of the present disclosure relate to recreating a sound field at a "3DoF position" in a manner corresponding to a predefined reference signal (which may or may not be consistent with the physical laws of sound wave propagation). This sound field should be based on all original "audio sources" (audio elements) and reflect the effects of the complex (and potentially dynamically changing) geometric configuration of the corresponding acoustic environment (e.g., VR / AR / MR environment, i.e., "doors," "walls," etc.). For example, referring to the example of FIG. 2, the sound field may relate to all sound sources (audio elements) in an elevator.
[0086] Furthermore, to provide a high level of VR / AR / MR immersion for the "6DoF space", the corresponding renderer (e.g., 6DoF renderer) output sound field should be reproduced sufficiently well.
[0087] Thus, instead of rendering several original audio objects (audio elements) and considering the effects of a complex acoustic environment, embodiments of the present disclosure relate to introducing virtual audio objects (effective audio elements) that are pre-rendered in the encoder and represent the entire audio scene (i.e., take into account the effects of the acoustic environment on the audio scene). All effects of the acoustic environment (e.g., acoustic hiding, reverberation, direct reflections, echoes, etc.) are captured directly in the virtual object (effective audio element) waveforms that are encoded and transmitted to a renderer (e.g., a 6DoF renderer).
[0088] A corresponding decoder-side renderer (e.g., a 6DoF renderer) may operate in a "simple rendering mode" (without considering the VR / AR / MR environment) for such object types (element types) throughout the 6DoF space. The simple rendering mode (as an example of the above-mentioned predetermined rendering mode) may only take into account distance attenuation (in empty space) and may not consider effects of the acoustic environment (e.g., the effects of acoustic elements in the acoustic environment) such as reverberation, echoes, direct reflections, acoustic occlusion, etc.
[0089] To extend the applicability range of a predefined reference signal, the virtual object(s) (active audio elements) may be placed at a specific position in the acoustic environment (VR / AR / MR space) (e.g., at the center of the sound intensity of the original audio scene or original audio element). This position can be determined automatically by inverse audio rendering in the encoder or manually specified by the content provider. In this case, the encoder only transmits: 1.b) A flag signaling the "pre-render type" (or generally the rendering mode indication) of the virtual audio object; 2.b) a virtual audio object signal (enabled audio element) derived from at least a pre-rendered reference (e.g., a mono object); and 3.b) Coordinates of the "3DoF position" and a description of the "6DoF space" (e.g., valid audio element information including valid audio element positions).
[0090] The predefined reference signal for the conventional approach is not the same as the virtual audio object signal (2.b) for the proposed approach, i.e., a "simple" 6DoF rendering of the virtual audio object signal (2.b) should approximate the predefined reference signal as well as possible for a given "3DoF position(s)."
[0091] In one example, the following encoding method may be performed by an audio encoder: 1. Determining the desired "3DoF location(s)" and corresponding "3DoF+ region(s)" (e.g., listener location and / or listener location region(s) for which rendering is desired) 2. Reference rendering (or direct recording) of the "3DoF position(s)" 3. Inverse Audio Rendering: Determination of the signal(s) and position(s) of the virtual audio object(s) (enabled audio elements) that result in the best possible approximation of the obtained reference signal at the "3DoF position(s)" 4. Encoding the resulting virtual audio object(s) (enabled audio elements) and their position(s) together with "pre-rendered object" attributes (rendering mode instructions) that enable the signaling of the corresponding 6DoF space (acoustic environment) and the "simple rendering mode" of the 6DoF renderer.
[0092] The complexity of the inverse audio rendering (see item 3 above) directly correlates to the complexity of the 6DoF processing in the 6DoF renderer's "simple rendering mode." Furthermore, this processing occurs on the encoder side, which is assumed to be less limited in terms of computational power.
[0093] Examples of data elements that need to be transferred in the bitstream are shown schematically in Figure 11A. Figure 11B shows a schematic of the data elements that are transferred in the bitstream in a conventional encoding / decoding system.
[0094] Figure 12 shows the use cases for the straightforward "simple" and "reference" rendering modes: the left side of Figure 12 shows the operation of the aforementioned rendering modes, while the right side shows a schematic representation of rendering an audio object to the listener position using either rendering mode (based on the example in Figure 2). The "simple rendering mode" may not consider the acoustic environment (e.g., acoustic VR / AR / MR environment). That is, the simple rendering mode may only consider distance attenuation (e.g., in empty space). For example, as shown in the top left panel of Figure 12, in the simple rendering mode, F simple only considers distance attenuation and does not consider the effects of VR / AR / MR environments such as opening and closing doors (see Figure 2, for example). The "Reference Rendering Mode" (bottom panel on the left side of Figure 12) may take into account some or all of the VR / AR / MR environmental effects.
[0095] Figure 13 shows an example of encoder / decoder side processing for the simple rendering mode. The upper panel on the left shows the encoder processing, and the lower panel on the left shows the decoder processing. The right side shows a schematic of the inverse rendering of the audio signal at the listener position to the position of the active audio element.
[0096] The renderer (e.g., a 6DoF renderer) output may approximate the reference audio signal at the 3DoF position(s). This approximation may include the effects of the audio core coder and audio object aggregation (i.e., the representation of several spatially distinct audio sources (audio elements) by a smaller number of virtual objects (active audio elements). For example, the approximated reference signal may take into account changes in the listener's position in the 6DoF space and may also represent several audio sources (audio elements) based on a smaller number of virtual objects (active audio elements). This is shown schematically in FIG. 14.
[0097] In one example, FIG. 15 shows a sound source / object signal (audio element) 101, a virtual object signal (active audio element) 100, a desired rendering output in 3DoF 102, and a virtual object signal (active audio element) 103.
number
number
[0098] Further terms include: 3DoF given reference compatible position(s) ∈ 6DoF space 6DoF any allowed position(s) ∈ VR / AR / MR scene F reference (x) Encoder-determined reference rendering F simple (x) 6DoF "simple mode rendering" specified by the decoder x (NDoF) Sound field representation in 3DoF position / 6DoF space x reference (3DoF) Reference signal(s) for 3DoF position(s), as determined by the encoder: x reference(3DoF) :=F for 3DoF reference (x) x reference (6DoF) General Reference Rendering Output x reference (6DoF) :=F for 6DoF reference (x) Given (on the encoder side): Audio source signal(s) x Reference signal(s) for 3DoF position(s) x reference (3DoF) Available (in renderers): Virtual object signal(s) x virtual Decoder 6DoF "Simple Rendering Mode" 6DoF F simple , ∃F -1 simple problem: Provide the following x virtual and x (6DoF) Define Desired rendering output in 3DoF x (3DoF) →x reference (3DoF) Approximation of the desired rendering
number
number
[0099] The following main advantages of the proposed method can be identified: · Artistic Rendering Feature Support: The output of the 6DoF renderer can correspond to any (known on the encoder side) artistic pre-rendered reference signal. · amount of calculation:6DoF audio renderers (e.g., MPEG-I audio renderers) can operate in "simple rendering mode" for complex acoustic VR / AR / MR environments. · Coding efficiency : With this approach, the audio bitrate for the pre-rendered signal is proportional to the number of 3DoF positions (more precisely, the number of corresponding virtual objects) rather than to the number of original audio sources, which can be very beneficial when the number of objects is large and the 6DoF degrees of freedom of movement are limited. at a predetermined location or locations Audio Quality Control : The best perceptual audio quality can be explicitly guaranteed by the encoder for any position in the VR / AR / MR space and the corresponding 3DoF+ region(s).
[0100] The present invention supports the concept of reference rendering / recording (i.e., "artistic intent"), i.e., any complex acoustic environment (or artistic rendering effect) can be encoded by (transmitted in) a pre-rendered audio signal(s).
[0101] The following information can be signaled in the bitstream to allow reference rendering / recording: Pre-rendered signal type flag(s), which enables a "simple rendering mode" that ignores the effects of the acoustic VR / AR / MR environment for the corresponding virtual object. A parameter expression describing the applicability region (i.e., 6DoF space) for virtual object signal rendering.
[0102] During 6DoF audio processing (e.g., MPEG-I audio processing), the following may be specified: How a 6DoF renderer mixes such pre-rendered signals with each other and with regular signals.
[0103] Thus, the present invention provides: "Simple mode rendering" capabilities specified by the decoder (i.e., F simple ) is general; it may be of arbitrary complexity, but there should be a corresponding approximation at the decoder side (i.e., ∃F simple -1 Ideally, this approximation should be mathematically well-defined (e.g., algorithmically stable). General sound field and source representations (and their combinations): extensible and applicable to objects, channels, FOAs, and HOAs Aspects of audio source directionality can be considered (in addition to distance attenuation modeling) Applicable to multiple (possibly overlapping) 3DoF positions for pre-rendered signals Applicable to scenarios where pre-rendered signals are mixed with normal signals (ambient sounds, objects, FOA, HOA, etc.) ·Reference signal x for 3DoF position reference (3DoF) of -The output of any "production renderer" (of any complexity) applied by the content creator - Real audio signals / field recordings (and their artistic modifications) and allows it to be obtained.
[0104] Some embodiments of the present disclosure include:
number
[0105] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on a digital signal processor or microprocessor. Other components may be implemented as hardware and / or as application-specific integrated circuits. Signals encountered in the methods and systems described above may be stored on media such as random access memory or optical storage media. The signals may be transmitted over a network, such as an airwave network, a satellite network, a wireless network, or a wired network such as the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer devices used to store and / or render audio signals.
[0106] Exemplary implementations of methods and apparatus according to this disclosure will become apparent from the following enumerated example embodiments (EEE), which are not claims.
[0107] EEE1 relates to a method for encoding audio data, comprising the steps of: encoding a virtual audio object signal derived from at least a pre-rendered reference signal; encoding metadata indicative of a 3DoF position and a 6DoF spatial description; and transmitting the encoded virtual audio signal and the metadata indicative of the 3DoF position and the 6DoF spatial description. EEE2 relates to the method of EEE1, further comprising transmitting a signal indicating the presence of a pre-rendered type of said virtual audio object. EEE3 relates to the method of EEE1 or EEE2, wherein at least the pre-rendered reference is determined based on a reference rendering of the 3DoF position and the corresponding 3DoF+ region. EEE4 relates to the methods of any one of EEE1 to EEE3, further comprising determining a position of the virtual audio object relative to the 6DoF space. EEE5 relates to any one of the methods EEE1 to EEE4, wherein the position of the virtual audio object is determined based on at least one of reverse audio rendering or manual specification by a content provider. EEE6 relates to any one of the methods EEE1 to EEE5, wherein the virtual audio object approximates a predefined reference signal for 3DoF position. EEE7 relates to any one of the methods EEE1 to EEE6, wherein the virtual object is
number
[0108] EEE8 relates to a method for rendering a virtual audio object, the method comprising rendering a 6DoF audio scene based on said virtual audio object. EEE9 relates to the method of EEE8, wherein the rendering of the virtual object comprises:
number
[0109] Several aspects will be described. [Aspect 1] 1. A method for decoding audio scene content from a bitstream by a decoder including an audio renderer having one or more rendering tools, the method comprising: receiving the bitstream; decoding an audio scene description comprising an acoustic ambience from said bitstream; determining one or more effective audio elements from the audio scene description, the one or more effective audio elements encapsulating the effects of the acoustic environment and corresponding to one or more virtual audio objects representing the audio scene; determining effective audio element information indicating effective audio element positions of the one or more effective audio elements from the audio scene description, the effective audio element information including information indicating a sound radiation pattern of each of the one or more effective audio elements; decoding a rendering mode indication from the bitstream, the rendering mode indication indicating whether the one or more enabled audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode; in response to the rendering mode instruction indicating that the one or more enabled audio elements represent a sound field obtained from pre-rendered audio elements and should be rendered using a predetermined rendering mode, rendering the one or more enabled audio elements using the predetermined rendering mode; rendering the one or more active audio elements using the predetermined rendering mode takes into account the active audio element information and the information indicative of a sound radiation pattern of each of the one or more active audio elements, and the predetermined rendering mode defines a predetermined configuration of the rendering tool for controlling an influence of the acoustic environment of the audio scene on a rendered output; method. [Aspect 2] obtaining listener position information indicative of a position of the listener's head in the acoustic environment and / or listener orientation information indicative of an orientation of the listener's head in the acoustic environment; Rendering the one or more enabled audio elements using a predetermined rendering mode further takes into account the listener position information and / or listener orientation information. 2. The method of embodiment 1. Aspect 3 rendering the one or more active audio elements using the predetermined rendering mode includes applying sound attenuation modeling according to respective distances between a listener position and an active audio element position of the one or more active audio elements; 3. The method of embodiment 1 or 2. Aspect 4 At least two valid audio elements are determined from said audio scene description; the rendering mode indication indicates a respective predetermined rendering mode for each of the at least two enabled audio elements; The method includes rendering the at least two enabled audio elements using respective predetermined rendering modes; Rendering each enabled audio element using a respective predetermined rendering mode takes into account enabled audio element information for that enabled audio element, the rendering mode for that enabled audio element defining a respective predetermined configuration of rendering tools for controlling the influence of the acoustic environment of the audio scene on the rendered output for that enabled audio element; 4. The method of any one of embodiments 1 to 3. Aspect 5 determining one or more original audio elements from said audio scene description; determining audio element information indicating audio element positions of the one or more audio elements from the audio scene description; rendering the one or more audio elements using a rendering mode for the one or more audio elements that is different from a predetermined rendering mode used for the one or more enabled audio elements; rendering the one or more audio elements using a rendering mode for the one or more audio elements takes into account the audio element information; 5. The method of any one of embodiments 1 to 4. Aspect 6 obtaining listener position region information indicating a listener position region for which the predetermined rendering mode is to be used. 6. The method of any one of embodiments 1 to 5. Aspect 7 the predetermined rendering mode indicated by the rendering mode indication is dependent on the listener position; the method includes rendering the one or more enabled audio elements for the listener location region indicated by the listener location region information using a predetermined rendering mode indicated by the rendering mode indication; The method of embodiment 6. Aspect 8 1. A method for generating audio scene content, the method comprising: obtaining one or more audio elements representing a captured signal from an audio scene including an acoustic environment; obtaining effective audio element information indicative of effective audio element positions of one or more effective audio elements to be generated, the one or more effective audio elements encapsulating the effect of the acoustic environment and corresponding to one or more virtual audio objects representing the audio scene, the effective audio element information including information indicative of a sound radiation pattern of each of the one or more effective audio elements; determining the one or more effective audio elements from the one or more audio elements representing the captured signal by applying sound attenuation modeling according to a distance between a location where the captured signal was captured and an effective audio element location of the one or more effective audio elements; method. Aspect 9 1. A method for encoding audio scene content into a bitstream, the method comprising: receiving a description of an audio scene, the audio scene including an acoustic environment and one or more audio elements at respective audio element positions; determining one or more effective audio elements at each effective audio element position from the one or more audio elements, the one or more effective audio elements corresponding to one or more original audio objects, and the one or more effective audio elements corresponding to one or more virtual audio objects encapsulating the influence of the acoustic environment and representing the audio scene; generating effective audio element information indicating effective audio element positions of the one or more effective audio elements, the effective audio element information being generated to include information indicating a sound radiation pattern of each of the one or more effective audio elements; generating a rendering mode instruction indicating that the one or more enabled audio elements should be rendered using a predetermined rendering mode representing a sound field obtained from pre-rendered audio elements and defining a predetermined configuration of a decoder's rendering tools for controlling the effect of the acoustic environment on the rendered output at the decoder; encoding the one or more audio elements, the audio element position, the one or more active audio elements, the active audio element information, and the rendering mode indication into a bitstream; method. Aspect 10 obtaining listener position information indicative of a position of the listener's head in the acoustic environment and / or listener orientation information indicative of an orientation of the listener's head in the acoustic environment; encoding the listener position information and / or listener orientation information into the bitstream. 10. The method according to embodiment 9. Aspect 11 In some embodiments, at least two significant audio elements are generated and encoded into the bitstream; the rendering mode indication indicates a respective predetermined rendering mode for each of said at least two enabled audio elements; 11. The method of embodiment 9 or 10. Aspect 12 obtaining listener position area information indicating a listener position area in which the predetermined rendering mode is to be used; and encoding the listener location region information into the bitstream. 12. The method of any one of embodiments 9 to 11. Aspect 13 13. The method of claim 12, wherein the predetermined rendering mode indicated by the rendering mode indication depends on the listener position, and the rendering mode indication indicates a respective predetermined rendering mode for each of a plurality of listener positions. Aspect 14 8. An audio decoder comprising: a processor coupled to a memory storing instructions for the processor, the processor adapted to perform a method according to any one of aspects 1 to 7. Aspect 15 8. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the method of any one of aspects 1 to 7. Aspect 16 A computer-readable storage medium storing the computer program according to aspect 15. Aspect 17 14. An audio encoder comprising: a processor coupled to a memory storing instructions for the processor, the processor adapted to perform the method of any one of aspects 8 to 13. Aspect 18 14. A computer program comprising instructions that, when executed by a processor, cause the computer program to perform the method of any one of aspects 8 to 13. Aspect 19 A computer-readable storage medium storing the computer program according to aspect 18.
Claims
1. 1. A method for decoding audio scene content by a decoder, the method comprising: receiving, by the decoder, a bitstream including active audio elements of the audio scene, active audio element information, and listener position area information, wherein the active audio element information indicates active audio element positions of the active audio elements, and the listener position area information indicates listener position areas in an acoustic environment; and rendering the effective audio element based on the effective audio element information and the listener position area information. method.
2. receiving a rendering mode; determining, based on the rendering mode, that the active audio elements represent a sound field obtained from pre-rendered audio elements; determining that the valid audio element is to be rendered using a predetermined rendering mode; and rendering the active audio element within the listener position area using the predetermined rendering mode. The method of claim 1.
3. and rendering the effective audio elements using the predetermined rendering mode includes applying sound attenuation modeling according to respective distances between a listener position and an effective audio element position of the effective audio elements. The method of claim 2.
4. The method of claim 2 , wherein the predetermined rendering mode depends on the listener position area.
5. The method of claim 1 , wherein the acoustic environment is a virtual reality / augmented reality / mixed reality (VR / AR / MR) acoustic environment.
6. 10. A non-transitory computer-readable storage medium containing instructions for causing a processor executing the instructions to perform the method of claim 1.
7. 1. An apparatus for audio decoding, comprising: an audio decoder for receiving a bitstream including active audio elements of the audio scene, active audio element information, and listener position area information, wherein the active audio element information indicates active audio element positions of the active audio elements, and the listener position area information indicates listener position areas in an acoustic environment; a renderer for rendering the effective audio elements based on the effective audio element information and the listener position area information. Device.
Citation Information
Patent Citations
Acoustic processing device
JP2006074589A