Method and apparatus for immersive audio scene manipulation - Patents.com
The method for processing 3D audio scenes addresses distraction issues by modifying and integrating secondary audio elements based on environmental and content context, enhancing user focus and clarity in immersive audio environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2026-03-10
AI Technical Summary
In immersive audio environments, the placement of audio elements can negatively impact viewer attention and focus, leading to confusion and distraction, necessitating controlled modification and intelligent positioning to enhance user experience.
A method for processing 3D audio scenes involves receiving original and secondary audio data, extracting rendering parameters, and modifying the original scene to jointly render secondary audio elements, allowing for controlled integration based on environmental conditions and content context, with options for temporary or permanent manipulation of spatial volumes and audio elements.
This approach enables controlled integration of secondary audio elements, reducing distractions and enhancing user focus by intelligently positioning audio elements, improving clarity and attention guidance in immersive audio environments.
Smart Images

Figure 2026508387000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure generally relates to a method for processing a 3D audio scene. In particular, an original 3D audio scene including a first audio element in a 3D audio space may be modified. A second audio element may be rendered jointly with the modified audio scene or the original 3D audio scene. The present disclosure further relates to corresponding apparatus and computer program products.
[0002] Although some embodiments are described herein with particular reference to that disclosure, it will be understood that the disclosure is not limited to such fields of use but is applicable in a broader context. [Background technology]
[0003] Any discussion of background art throughout this disclosure should in no way be taken as an admission that such art is widely known or forms part of common general knowledge in the art.
[0004] In virtual or augmented viewing environments (augmented reality (AR), mixed reality (MR), extended reality (XR)), AR glasses being a prominent example, and in other contexts, the placement of audio elements for personalized advertising or informational purposes is becoming increasingly important. However, immersive audio can negatively impact the viewer / listener's attention and ability to focus on tasks and details. In virtual or augmented viewing environments, the placement of audio elements in complex rendered scenes can even lead to temporary confusion.
[0005] Therefore, there is a need to allow for the original immersive rendered audio scene to be modified in a controlled way, and furthermore, there is a need to allow for intelligent positioning of audio elements in the immersive rendered audio scene. Summary of the Invention [Means for solving the problem]
[0006] According to a first aspect of the present disclosure, there is provided a method for processing a 3D audio scene. The method may include receiving an original 3D audio scene. The original 3D audio scene may include first audio data and first metadata for a plurality of first audio elements in a 3D audio space. The method may further include receiving second audio data and second metadata for one or more second audio elements. The method may further include extracting one or more rendering parameters from the second metadata to render the one or more second audio elements relative to the original 3D audio scene. The method may further include modifying the original 3D audio scene to obtain a modified 3D audio scene. The method may then include jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters.
[0007] Thus, the original audio scene can be modified in a controlled way to provide room for the second audio element (guest audio element), which allows for a congruent rendering of the original audio scene and the guest audio element in a meaningful way based on environmental conditions and / or content context, and avoids harmful or even annoying effects for the listener.
[0008] In some embodiments, modifying the original 3D audio scene may be based on metadata.
[0009] In some embodiments, the original 3D audio scene may be associated with a first spatial volume that encloses a plurality of first audio elements. Modifying the original 3D audio scene may include mapping the first spatial volume to a second spatial volume associated with the modified 3D audio scene. The second spatial volume may be a modified spatial volume compared to the first spatial volume.
[0010] In some embodiments, the modified spatial volume may be a reduced spatial volume. Optionally, the volume fraction of the reduced spatial volume associated with the modified 3D audio scene may be between 1% and 99% of the first spatial volume.
[0011] In some embodiments, the modified spatial volume may be an extended spatial volume. Optionally, the volume fraction of the extended spatial volume associated with the modified 3D audio scene may be between 101% and 300% of the first spatial volume.
[0012] In some embodiments, the first metadata may include first tagging information for a plurality of first audio elements indicating whether each first audio element should be moved in the 3D audio space when mapping the first spatial volume to the second spatial volume, wherein only first audio elements indicated by the first tagging information as not to be moved may be rendered in their original positions.
[0013] In some embodiments, mapping the first volume of space to the second volume of space may be performed instantaneously.
[0014] In some embodiments, mapping the first spatial volume to the second spatial volume may be performed gradually or in stages over time.
[0015] In some embodiments, modifying the original 3D audio scene may further include selecting a third spatial volume within the 3D audio space to allocate the one or more second audio elements.
[0016] In some embodiments, the third spatial volume may at least partially overlap the first spatial volume.
[0017] In some embodiments, mapping the first spatial volume to the second spatial volume may include removing each first audio element from the third spatial volume, wherein only first audio elements indicated by the first tagging information as not to be moved may not be removed from the third spatial volume.
[0018] In some embodiments, removing the first audio element from the third volume of space may be performed instantaneously.
[0019] In some embodiments, removing the first audio element from the third spatial volume may be performed gradually or in stages over time.
[0020] In some embodiments, jointly rendering the one or more second audio elements and the modified 3D audio scene may include rendering the one or more second audio elements in a third spatial volume and rendering the modified 3D audio scene in a second spatial volume.
[0021] In some embodiments, the shape of the first, second, and / or third spatial volumes may include one or more of a sphere, a quadrant, an octant, and a point.
[0022] In some embodiments, the shape of the first, second, and / or third spatial volumes may be based on metadata.
[0023] In some embodiments, the shapes of the first, second, and / or third spatial volumes may be predefined.
[0024] In some embodiments, modifying the original 3D audio scene and / or the collective rendering may be based on first timing information, the first timing information indicating a point in time or a time period for applying said modification and / or said collective rendering.
[0025] In some embodiments, the method may further include, following the joint rendering, removing the one or more second audio elements and returning to rendering the original 3D audio scene.
[0026] In some embodiments, the removing of the one or more second audio elements and / or the returning to rendering the original 3D audio scene may be based on second timing information, the second timing information indicating a point in time or a time period for applying the removing and / or the returning.
[0027] In some embodiments, the second metadata may include second tagging information for one or more secondary audio elements indicating whether the respective secondary audio elements should be removed when returning to rendering the original 3D audio scene, wherein only secondary audio elements indicated by the second tagging information as not to be removed may not be removed prior to rendering the original 3D audio scene.
[0028] In some embodiments, the one or more rendering parameters may include an indication of a loudness of the one or more secondary audio elements, wherein the joint rendering may further include adjusting the loudness of the one or more secondary audio elements relative to the loudness of the modified 3D audio scene.
[0029] In some embodiments, the joint rendering may further include adapting a gain of the modified 3D audio scene and / or the one or more secondary audio elements.
[0030] In some embodiments, the method may include receiving different versions of second audio data for the one or more second audio elements, each version having a different audio configuration.
[0031] In some embodiments, the method may further include selecting a version of second audio data for the one or more second audio elements that best matches the audio composition of the original 3D audio scene.
[0032] In some embodiments, the congruent rendering may further include aligning the one or more second audio elements with an audio composition of the modified 3D audio scene.
[0033] In some embodiments, the one or more rendering parameters may further include an indication of a perceived complexity of the one or more second audio elements.
[0034] In some embodiments, the one or more rendering parameters may further include an indication of sensitivity to perceived interference of the one or more second audio elements by other audio elements.
[0035] In some embodiments, the one or more rendering parameters may further include an indication of an audio quality of the one or more secondary audio elements.
[0036] In some embodiments, the method may further include, prior to joint rendering, comparing the audio quality of the one or more second audio elements with the audio quality of the original 3D audio scene, and aligning the audio quality of the one or more second audio elements with the audio quality of the original 3D audio scene.
[0037] In some embodiments, the one or more rendering parameters may further include an indication of a priority level for each of the one or more secondary audio elements, the priority level indicating a priority for rendering the respective secondary audio element.
[0038] In some embodiments, the collaborative rendering may further include modifying some or all of the first audio elements based on their respective priority levels.
[0039] In some embodiments, the one or more rendering parameters may further include an indication of a category for each of the one or more second audio elements, the category including one or more of optional, required, supplementary, urgent, subordinate, dependent, within the context of the original 3D audio scene and outside the context of the original 3D audio scene.
[0040] In some embodiments, the joint rendering may be based on the perceived complexity and / or sensitivity of the original 3D audio scene, and optionally based on an indication of the perceived complexity and / or sensitivity of the one or more second audio elements.
[0041] In some embodiments, the joint rendering may be based on a priority level indication and a category indication of each of the one or more secondary audio elements, ignoring the perceived complexity and sensitivity of the original 3D audio scene.
[0042] In some embodiments, the prioritization of selecting a third spatial volume or allocating the one or more second audio elements to a third spatial volume may be based on one or more of a perceived complexity, sensitivity, audio quality, priority level, and category indication of the one or more second audio elements.
[0043] In some embodiments, a first audio element of the original 3D audio scene associated with ambience in the third spatial volume may be left unmodified, and the joint rendering may include attenuating or amplifying the one or more second audio elements relative to the first audio element associated with ambience.
[0044] According to a second aspect of the present disclosure, there is provided a method for processing a 3D audio scene. The method may include receiving an original 3D audio scene, the original 3D audio scene including first audio data and first metadata for multiple first audio elements in a 3D audio space. The method may further include receiving second audio data and second metadata for one or more second audio elements. The method may further include extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene. The method may then include jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters.
[0045] In some embodiments, the collaborative rendering may include embedding the one or more secondary audio elements into the original 3D audio scene such that the original 3D audio scene may be augmented by the one or more secondary audio elements.
[0046] In some embodiments, the collaborative rendering may include modifying some or all of the first audio element based on the one or more rendering parameters.
[0047] According to a third aspect of the present disclosure, there is provided a method for processing a 3D audio scene. The method may include receiving an original 3D audio scene, the original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space. The method may further include receiving information indicating modification of the original 3D audio scene. The method may further include obtaining one or more rendering parameters from the information. The method may then include rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene.
[0048] In some embodiments, the information may correspond to default metadata associated with one or more default audio elements.
[0049] In some embodiments, the information may be generated in real time.
[0050] In some embodiments, the information may be based on user preferences and / or location data.
[0051] According to a fourth aspect of the present disclosure, there is provided an apparatus for processing an audio scene. The apparatus may include one or more processors configured to perform the methods described herein. In some implementations, the one or more processors may be configured to perform a method including receiving an original 3D audio scene including first audio data and first metadata for multiple first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the second audio elements relative to the original 3D audio scene; modifying the original 3D audio scene to obtain a modified 3D audio scene; and jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters.
[0052] In some embodiments, the one or more processors may be further configured to, after joint rendering, remove the one or more second audio elements and return to rendering the original 3D audio scene.
[0053] According to a fifth aspect of the present disclosure, there is provided an apparatus for processing an audio scene, which may include one or more processors configured to perform a method including receiving an original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space, receiving second audio data and second metadata for one or more second audio elements, extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene, and jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters.
[0054] According to a sixth aspect of the present disclosure, there is provided an apparatus for processing an audio scene, which may include one or more processors configured to perform a method including receiving an original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space; receiving information indicating to modify the original 3D audio scene; obtaining one or more rendering parameters from the information; and rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene.
[0055] According to a seventh aspect of the present disclosure, there is provided a program comprising instructions that, when executed by a processor, cause the processor to perform the method described herein. The program may be stored on a computer-readable storage medium.
[0056] It will be understood that device (system) features and method steps can be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding device (system), and vice versa. This will be understood by those skilled in the art. Furthermore, it will be understood that any of the above statements made with respect to a method equally apply to a corresponding device (system), and vice versa. [Brief explanation of the drawings]
[0057] Exemplary embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0058] [Figure 1] 1 illustrates a first example of a method for processing a 3D audio scene according to an embodiment of the present disclosure. [Figure 2] 1 shows a schematic diagram of an example of modifying an original 3D audio scene to obtain a modified 3D audio scene and jointly rendering a second audio element and the modified 3D audio scene according to an embodiment of the present disclosure. [Figure 3] FIG. 10 shows a schematic diagram of a further example of modifying an original 3D audio scene to obtain a modified 3D audio scene and jointly rendering a second audio element and the modified 3D audio scene according to an embodiment of the present disclosure. [Figure 4] FIG. 10 shows a schematic diagram of a further example of modifying an original 3D audio scene to obtain a modified 3D audio scene and jointly rendering a second audio element and the modified 3D audio scene according to an embodiment of the present disclosure. [Figure 5] 10 illustrates a second example of a method for processing a 3D audio scene, according to an embodiment of the present disclosure. [Figure 6] 1 shows a schematic diagram of an example of jointly rendering a secondary audio element and an original 3D audio scene according to an embodiment of the present disclosure. [Figure 7]10 illustrates a third example of a method for processing a 3D audio scene, according to an embodiment of the present disclosure. [Figure 8] 1 shows a schematic diagram of an example of modifying an original 3D audio scene to obtain a modified 3D audio scene according to an embodiment of the present disclosure. [Figure 9] 10 shows a schematic diagram of a further example of modifying an original 3D audio scene to obtain a modified 3D audio scene according to an embodiment of the present disclosure; [Figure 10] 1 illustrates a schematic diagram of an example of an apparatus for implementing a method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0059] Overview The methods and apparatus described herein enable immersive audio renderers to intentionally modify, shape, or collapse the rendered sound field to achieve different goals based on environmental conditions, user preferences, and / or content context. These goals may be: Reducing distractions caused by spatial audio rendering, which can have a detrimental effect on the observer in certain scenarios, for example, when listening to an augmented reality (AR) sound field while walking through traffic. "Making space" for other content, such as audio with a different context, that is added to an existing immersive scene. For example, increasing the presence of advertising that is added to augment the existing rendered audio environment. Manipulation (modification) of an immersively rendered scene to direct the listener's attention to scene highlights, e.g. for training purposes. Improve the clarity of immersive scenes for users with certain sensory impairments or sensory preferences.
[0060] The methods and apparatus described herein provide a means to permanently or temporarily affect and manipulate the original immersively rendered audio scene, such that a user's attention or focus can be manipulated by intelligent placement of audio sources (e.g., audio elements or audio objects). An audio scene may be understood as a group of audio elements in 3D space, each having specific audio characteristics, e.g., loudness and direction, and an associated audio signal.
[0061] The methods and apparatus described herein further provide the ability to scale the modification, e.g., squashing, of the original audio scene from a minimum (unmanipulated / unmodified) to a selected maximum. For example, the entire original audio scene may be collapsed (mapped and squashed) to a single (mono) point in 3D audio space or downmix, or the entire original audio scene may be maximally expanded in 3D audio space. Furthermore, certain elements in the 3D audio space may be modified to become less audible, more audible, or to create certain spatial effects, such as focusing audio elements in the user's viewing direction.
[0062] Method and apparatus for processing 3D audio scenes Referring to Figure 1, a first example of a method for processing a 3D audio scene 100 is shown.
[0063] In step S101, an original 3D audio scene is received. The original 3D audio scene includes first audio data and first metadata for a plurality of first audio elements in a 3D audio space. The original 3D audio scene may be an original immersively rendered 3D audio scene, or in other words, the original 3D audio scene may be said to be an original rendered sound field in a 3D audio space.
[0064] In step S102, second audio data and second metadata for one or more second audio elements are received, where the second audio data / second audio elements may represent content other than the first audio data / first audio elements having a different context, such as an advertisement or emergency context.
[0065] In step S103, one or more rendering parameters are extracted from the second metadata for rendering the one or more second audio elements relative to the original 3D audio scene. The one or more rendering parameters may, for example, allow steering a user's / listener's attention or focus. Example properties / characteristics of the one or more rendering parameters are described in more detail below.
[0066] In step S104, the original 3D audio scene is modified (or manipulated) to obtain a modified 3D audio scene. Modifying the original 3D audio scene may allow for temporary or permanent shaping of the original 3D audio scene. The modification may allow for controlling / steering the placement / rendering of secondary audio elements in relation to the modified 3D audio scene. In an embodiment, modifying the original 3D audio scene may thus be based on the metadata.
[0067] In step S105, the one or more secondary audio elements and the modified 3D audio scene are then jointly rendered based on the one or more rendering parameters.
[0068] Modifying the original 3D audio scene is sometimes referred to as making room for rendering the one or more secondary audio elements. This can be achieved by collapsing or expanding the original 3D audio scene, as shown in the examples of FIGS. 2 and 3. Both examples show schematic diagrams of modifying the original 3D audio scene 200, 300 to obtain a modified 3D audio scene and jointly rendering the secondary audio element and the modified 3D audio scene. Both examples show the original 3D audio scene in the 3D audio space 201, 301 as the "starting point." In the examples of FIGS. 2 and 3, the original 3D audio scene is shown associated with a respective (first) spatial volume 202, 302. This first spatial volume 202, 302 may enclose a respective plurality of primary audio elements, which are not shown for simplicity.
[0069] In an embodiment, modifying the original 3D audio scene may include mapping a first spatial volume to a second spatial volume associated with the modified 3D audio scene, where the second spatial volume is a modified spatial volume compared to the first spatial volume. As already mentioned above, modifying the original 3D audio scene may be achieved by collapsing or expanding the original 3D audio scene, depending on the circumstances.
[0070] 2 illustrates collapsing an original 3D audio scene by mapping a first spatial volume 202 onto a second spatial volume 203 associated with the modified 3D audio scene, where the resulting modified spatial volume 203 is a reduced spatial volume compared to the first spatial volume 202. The volume fraction of the reduced spatial volume associated with the modified 3D audio scene is generally not limited, but in certain embodiments the volume fraction of the reduced spatial volume associated with the modified 3D audio scene may be between 1% and 99%, preferably between 10% and 80%, and more preferably between 25% and 60% of the first spatial volume.
[0071] Turning to the example of Figure 3, a further method for modifying the original 3D audio scene is shown to be expanding the original 3D audio scene. The expansion is performed by mapping a first spatial volume 302 onto a second spatial volume 303 associated with the modified 3D audio scene, with the resulting modified spatial volume 303 being an expanded spatial volume compared to the first spatial volume 302. The volume fraction of the expanded spatial volume associated with the modified 3D audio scene is also generally not limited, but in one embodiment, the volume fraction of the expanded spatial volume associated with the modified 3D audio scene may be between 101% and 300%, preferably between 150% and 250%, and more preferably between 175% and 225% of the first spatial volume.
[0072] The area for shrinking or expanding the original scene may be represented, for example, by a set of coordinates relative to the coordinate system of the original 3D audio scene and / or the shape of the modified 3D audio scene, which indicate the shrunken or expanded area / volume in 3D audio space compared to the original 3D audio scene. After shrinking or expanding by a given amount, the respective renderer can render the modified 3D audio scene using the shrunken or expanded space.
[0073] In general, the modification of the original 3D audio scene to obtain the modified 3D audio scene is not limited. Some examples for performing the modification are given below.
[0074] 1. Divide the 3D audio space (e.g., represented as a cube) into octants, redirect the center of gravity of the first (original) spatial volume for rendering the original 3D audio scene to the center of one octant, and provide an indication of the size (expanded or contracted) of the resulting second (modified) spatial volume.
[0075] 2. Shift the first spatial volume associated with the original 3D audio scene in the opposite direction in 3D audio space compared to the position where the second audio element should be inserted. The reference point can be the center of the first spatial volume associated with the original 3D audio scene / the center of the 3D audio space represented as a cube. Include an indication of the distance in the coordinate range of the original 3D audio scene.
[0076] 3. Shift the first spatial volume associated with the original 3D audio scene in the opposite direction relative to the (2D) center of the first spatial volume associated with the original 3D audio scene / center of the 3D audio space represented as a cube, compared to the position where the second audio element should be inserted. Include an indication of distance.
[0077] 4. Maintain a first spatial volume associated with the original 3D audio scene in the center of the 3D audio space represented as a cube, and remove / attenuate the first audio element that interferes with the second audio element that is placed within a radius x around the second audio element position (e.g., a third spatial volume).
[0078] 5. Maintain a first spatial volume associated with the original 3D audio scene in the center of the 3D audio space represented as a cube, and laterally warp the first audio element from the original 3D audio scene around a volume having the shape of a cylinder / cone oriented toward the center of the first spatial volume (e.g., a sphere) with a radius of x in the coordinate range of the original 3D audio scene, and release the space / volume within the cylinder / cone.
[0079] Referring to the example of Figure 4, a further example is shown of modifying an original 3D audio scene 400 to obtain a modified 3D audio scene and jointly rendering a second audio element and the modified 3D audio scene. The example of Figure 4 shows the modification of the original 3D audio scene in a 3D audio space 401 by mapping a first spatial volume 402 to a second spatial volume 405 associated with the modified 3D audio scene in a manner similar to that of Figure 2. The resulting modified spatial volume 405 is a reduced spatial volume compared to the first spatial volume 402. Additionally, the example of Figure 4 further shows the first spatial volume 402 enclosing a plurality of first audio elements 403.
[0080] In an embodiment, the first metadata may include first tagging information for a plurality of first audio elements indicating whether each first audio element should be moved in the 3D audio space when mapping the first spatial volume to the second spatial volume, and only first audio elements indicated by the first tagging information as not to be moved are rendered in their original positions.
[0081] 4, the filled circles indicate first audio elements that are indicated by the first tagging information as not being moved and that are rendered in their original positions 403b after modification, whereas the open circles indicate first audio elements 403a that are indicated by the first tagging information as being moved and that are not rendered in their original positions 403a after modification but are rendered in the modified / reduced spatial volume 405.
[0082] In particular, while the example of Figure 4 illustrates shrinking / collapse of the first spatial volume, similar considerations apply to expanding the first spatial volume during modification of the original 3D audio scene. That is, the methods described herein provide the ability to remove tagged audio elements (which may, for example, be tagged with accompanying metadata as being less relevant or of a certain category) and omit from the operation audio elements that are appropriately marked with accompanying metadata.
[0083] In general, the mapping of the first spatial volume to the second spatial volume is not limited to being performed, but in some embodiments, the mapping of the first spatial volume to the second spatial volume may be performed instantaneously. Alternatively, the mapping of the first spatial volume to the second spatial volume may be performed gradually or in stages over time. In other words, the mapping may be performed directly, or alternatively, in a slowly scaled / graceful manner.
[0084] 2 and 3, modifying the original 3D audio scene may further include selecting a third spatial volume 204, 304 within the 3D audio space 201, 301 for allocating the one or more second audio elements. The third spatial volume 204, 304 may at least partially overlap with the first spatial volume, for example, as shown in FIG. 3. Referring to the example of FIGS. 1-3, jointly rendering the one or more second audio elements and the modified 3D audio scene may include rendering the one or more second audio elements within the third spatial volume 204, 304 and rendering the modified 3D audio scene within the second spatial volume 203, 303.
[0085] 4, mapping the first spatial volume 402 to the second spatial volume 405 may include removing each first audio element 403 from the third spatial volume 404. As mentioned above, only first audio elements 403b that are indicated not to be moved by the first tagging information are not removed from the third spatial volume 404. This ensures that audio elements of higher relevance or audio elements of a certain category are not moved.
[0086] In general, the removal of the first audio element from the third spatial volume may be performed without limitation, but in some embodiments, the removal of the first audio element from the third spatial volume may be performed instantaneously, or alternatively, the removal of the first audio element from the third spatial volume may be performed gradually or in stages over time.
[0087] 2-4, while the shape of each spatial volume is shown as being elliptical or spherical, the shapes of the first, second, and / or third spatial volumes are generally not limited. However, in some embodiments, the shapes of the first, second, and / or third spatial volumes may include one or more of a sphere, a quadrant, an octant, and a point. The shapes of the first, second, and / or third spatial volumes may be based on metadata or, alternatively, may be predefined.
[0088] 1-4, modifying and / or congruent rendering of an original 3D audio scene may be based on first timing information, which may indicate a point in time or a time period for applying the modification and / or the congruent rendering.
[0089] As described above, the methods and apparatus described herein enable temporarily influencing and manipulating the original immersively rendered audio scene. In some embodiments, the method may further include, following the joint rendering, removing the one or more second audio elements and returning to rendering the original 3D audio scene. The removing of the one or more second audio elements and / or the returning to rendering the original 3D audio scene may be based on second timing information. The second timing information may indicate a point in time or a time period for applying the removing and / or the returning.
[0090] Similar to the first audio elements, the second metadata may include second tagging information for one or more second audio elements indicating whether the respective second audio elements should be removed when returning to rendering the original 3D audio scene. Then, only the second audio elements indicated by the second tagging information as not to be removed are not removed before rendering the original 3D audio scene.
[0091] 5, there is shown a second example of a method for processing a 3D audio scene 500. In contrast to the first example, in this case the original 3D audio scene remains unchanged.
[0092] In step S501, an original 3D audio scene is received, the original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space.
[0093] In step S502, second audio data and second metadata for one or more second audio elements are received.
[0094] In step S503, one or more rendering parameters are extracted from the second metadata for rendering the one or more second audio elements within the original 3D audio scene.
[0095] Then, in step S504, the one or more secondary audio elements and the original 3D audio scene are jointly rendered based on the one or more rendering parameters.
[0096] Referring to FIG. 6, a schematic diagram of an example of jointly rendering a second audio element with an original 3D audio scene 600 is shown. Similar to the examples of FIGS. 2 and 3, the original 3D audio scene is shown in a 3D audio space 601 associated with a respective (first) spatial volume 602. This first spatial volume 602 may enclose a respective plurality of first audio elements, which in this example are not shown for simplicity. In the present example, joint rendering may include embedding 603 one or more second audio elements into the original 3D audio scene 602 such that the original 3D audio scene is augmented by the one or more second audio elements. Joint rendering may further include modifying some or all of the first audio elements based on the one or more rendering parameters, as described herein.
[0097] 7, there is shown a third example of a method for processing a 3D audio scene 700. In this case, only the original 3D audio scene is modified.
[0098] In step S701, an original 3D audio scene is received, the original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space.
[0099] The metadata may define specific characteristics for each of the plurality of first audio elements. An example of the specific characteristics may be tagging information. The tagging information may indicate, for each of the plurality of first audio elements, whether the audio elements are modifiable when the original 3D audio scene is to be modified. In other words, the tagging information may protect certain audio elements of the plurality of first audio elements from being modified based on some or any received instructions to modify the original 3D audio scene.
[0100] In step S702, information indicating that the original 3D audio scene is to be modified is received. In one embodiment, the information may correspond to default metadata associated with one or more default audio elements. The default audio elements may be empty, and rendering thereof results in only modifications. Alternatively, the information may be generated in real time. Alternatively, the information may be based on user preferences and / or location data.
[0101] In some embodiments, the information may be received from a user, for example, via an interface. Alternatively, the information may be part of a bitstream received over a network.
[0102] The user preferences may include user-specific sensory impairments and sensory preferences to enhance the original 3D audio scene for the specific sensory impairment or preference. The sensory impairment may specifically relate to hearing, such as partial hearing loss in both ears or complete hearing loss in one ear. The partial hearing loss may relate to hearing loss in a specific frequency range. For example, the information may indicate an age-dependent hearing loss, which may correspond to hearing loss for frequencies above 2 kHz.
[0103] In step S703, one or more rendering parameters are obtained from the information.
[0104] Rendering parameters may include parameters for attenuating early reflections, parameters for reverberation decay, and / or parameters for distance-dependent gain stretching / compression.
[0105] Additionally, the rendering parameters may include one or more parameters for rendering the original 3D audio scene with a directional focus. Directional focus may be understood as focusing the user's attention in a certain direction by amplifying audio elements in that direction or attenuating elements outside that direction. The amplification or attenuation of audio elements may gradually increase / decrease for a certain distance or angle deviating from that direction. The direction may be the user's viewing direction. In other words, the default direction is steering toward the user's front viewing direction, but can also be reoriented in other directions to allow control through other services or modalities (e.g., an eye tracker). Optionally, audio elements authored to be in the user's coordinate system are associated with the listener and are not processed with directional focus.
[0106] The one or more parameters for rendering the original 3D audio scene with directional focus may include a flag to indicate whether directional focus should be applied to the original 3D audio scene, an angle range corresponding to a viewing direction, a transition angle for gradual attenuation, a maximum attenuation, a flag to indicate a default direction for the directional focus, a yaw angle to indicate the viewing direction, and a pitch angle to indicate the viewing direction.
[0107] As an example, if the received information indicates partial hearing loss, the obtained rendering parameters may be parameters for reverberation attenuation. By rendering the original audio scene with attenuated reverberation, intelligibility for a user with partial hearing loss may be significantly increased. Furthermore, if reverberation is attenuated or early reflections are attenuated, sensory overload for the user may be reduced.
[0108] Then, in step S704, the original 3D audio scene is rendered based on the one or more rendering parameters to obtain a modified 3D audio scene.
[0109] In other words, the original 3D audio scene is rendered using modified rendering parameters instead of the default rendering parameters, resulting in a modified 3D audio scene. As an example, audio elements outside the directional focus may be rendered with an attenuated gain, such that the modified 3D audio scene has a directional focus but corresponds to the original 3D audio scene. In other words, the user may perceive audio elements in the viewing direction as amplified.
[0110] In some embodiments, where the metadata includes tagging information for each of the plurality of first audio elements, if the tagging information indicates that a particular audio element should be left unchanged, the particular audio element is rendered using default rendering parameters instead of the rendering parameters obtained from the received information. Optionally, the tagging information may prohibit any modification or only certain modifications, such as modifications with directional focus, rendering with distance-dependent gain, or rendering with new coordinates in 3D space.
[0111] Additionally, when the received information provides that the user suffers from spectral hearing loss, the rendering of the original 3D audio may include dynamic equalization to compensate for the spectral hearing loss individually for both ears.
[0112] By taking user preferences into account directly within the rendering system, the ability to optimize audio rendition on the fly can improve the performance and quality of modifying 3D audio scenes.
[0113] In some embodiments, the original 3D audio scene may be rendered by any one of speakers, binaural headphones, a VR headset, or any other suitable device for rendering 3D audio.
[0114] 8 and 9, schematic diagrams 800, 900 of an example of modifying an original 3D audio scene to obtain a modified 3D audio scene are shown. Similar to the examples of FIGS. 2 and 3, the original 3D audio scene is shown in a 3D audio space 801, 901 as being associated with a respective (first) spatial volume 802, 902. This first spatial volume 802, 902 may enclose a respective plurality of first audio elements, which in this example are not shown for simplicity. Similarly to what has been described above, rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene may include mapping the first spatial volume to a second spatial volume associated with the modified 3D audio scene, which may be a modified spatial volume compared to the first spatial volume.
[0115] 8, an original 3D audio scene is shown collapsed by mapping a first spatial volume 802 onto a second spatial volume 803 associated with the modified 3D audio scene, the resulting modified spatial volume 803 being a reduced spatial volume compared to the first spatial volume 802. The volume fraction of the reduced spatial volume associated with the modified 3D audio scene is generally not limited, but in one embodiment the volume fraction of the reduced spatial volume associated with the modified 3D audio scene may be between 1% and 99%, preferably between 10% and 80%, and more preferably between 25% and 60% of the first spatial volume.
[0116] 9, an expansion of an original 3D audio scene is shown. The expansion is performed by mapping a first spatial volume 902 to a second spatial volume 903 associated with the modified 3D audio scene, with the resulting modified spatial volume 903 being an expanded spatial volume compared to the first spatial volume 902. The volume fraction of the expanded spatial volume associated with the modified 3D audio scene is also generally not limited, but in one embodiment, the volume fraction of the expanded spatial volume associated with the modified 3D audio scene may be between 101% and 300%, preferably between 150% and 250%, and more preferably between 175% and 225% of the first spatial volume.
[0117] Parameters In addition to the above, the manipulation / modification of the original 3D audio scene and / or the congruent rendering can be further steered based on the following parameters, applied alone or in combination depending on the respective use case:
[0118] In some embodiments, the one or more rendering parameters may include an indication of a loudness of the one or more secondary audio elements, and the joint rendering may further include adjusting the loudness of the one or more secondary audio elements relative to the loudness of the modified 3D audio scene.
[0119] In some embodiments, the joint rendering may further include adapting a gain of the modified 3D audio scene and / or the one or more secondary audio elements.
[0120] In some embodiments, the method may further include receiving different versions of second audio data for the one or more second audio elements, each version having a different audio configuration, and then selecting a version of second audio data for the one or more second audio elements that most closely matches the audio configuration of the original 3D audio scene.
[0121] In some embodiments, the joint rendering may further include aligning an audio composition of the one or more second audio elements and the modified 3D audio scene.
[0122] In some embodiments, the one or more rendering parameters may further include an indication of a perceived complexity of the one or more second audio elements.
[0123] In some embodiments, the one or more rendering parameters may further include an indication of sensitivity to perceived interference of the one or more second audio elements by other audio elements.
[0124] In some embodiments, the one or more rendering parameters may further include an indication of audio quality of the one or more secondary audio elements. The method may then further include, prior to joint rendering, comparing the audio quality of the one or more secondary audio elements with the audio quality of the original 3D audio scene and aligning the audio quality of the one or more secondary audio elements with the audio quality of the original 3D audio scene. An example of an audio quality indicator may be the bitrate of the stream, especially if there is a quality plateau. If the original 3D audio scene is from a bitrate plateau = 3, the decision logic may accordingly decide to fetch the secondary audio elements from the same bitrate plateau.
[0125] In some embodiments, the one or more rendering parameters may further include an indication of a priority level for each of the one or more secondary audio elements, the priority level indicating a priority for rendering each secondary audio element. The joint rendering may then further include modifying some or all of the primary audio elements based on the respective priority levels.
[0126] In some embodiments, the one or more rendering parameters may further include an indication of a category of each of the one or more second audio elements, the category including one or more of optional, required, supplementary, urgent, dependent, subordinate, within the context of the original 3D audio scene, and outside the context of the original 3D audio scene.
[0127] In one embodiment, the joint rendering may be based on the perceived complexity and / or sensitivity of the original 3D audio scene, and optionally based on an indication of the perceived complexity and / or sensitivity of the one or more second audio elements.
[0128] In one embodiment, the joint rendering may be based on a priority level indication and a category indication of each of the one or more second audio elements, ignoring the perceived complexity and sensitivity of the original 3D audio scene.
[0129] In an embodiment, selecting a third spatial volume or prioritizing the allocation of the one or more second audio elements to the third spatial volume may be based on one or more of the perceived complexity, sensitivity, audio quality, priority level, and category indication of the one or more second audio elements.
[0130] In some embodiments, a first audio element of the original 3D audio scene related to the ambience in the third spatial volume may be left unchanged, and the congruent rendering may include attenuating or amplifying the one or more second audio elements relative to the first audio element related to the ambience.
[0131] In other words, the parameters available to steer the manipulation / modification of the original 3D audio scene and the joint rendering of said one or more secondary audio elements may be: 1. [XYZ] 3D coordinates in the original scene to be crushed 2. [INT] Size of target shape in %, or other measure of fraction of original size, or area of collapse (e.g., one or more octants, quadrants) 3. [XYZ] 3D coordinates in the source scene that should be cleared from the audio element (first audio element) belonging to the host scene. 4. [INT] Size of the shape to be freed in %, or other indication of a percentage of the host scene size, or other indication of the area to free (e.g., octant, quadrant) 5. [INT] Loudness instructions for host scene and guest elements (secondary audio elements) 6. [INT] Indication of experiential complexity of audio elements (host scene, guest elements) 7. [INT] Audio element (host scene, guest element) sensitivity indication 8. [INT] Audio quality indication for audio elements (host scene, guest elements) 9. [INT] Priority level of audio elements (host scene, guest elements) 10. [Table] Guest Audio Element Categories (e.g., optional, supplementary, required, urgent, dependent, subordinate, in context of host scene, outside context of host scene) 11. [INT] Gain adaptation for host scene or guest audio elements 12. [TIME] Timing information about when to apply changes (absolute / UTC, relative to the host timeline) 13. [INT] An indication that the host audio scene or any guest element may, should, must, or must not be omitted from operation. 14. [FLOAT] Early reflection level reduction in decibels 15. [FLOAT] Reverberation level reduction in decibels 16. [BIN] Flag indicating whether distance index is enabled 17. [FLOAT] Parameter to exponentially modify (i.e., shrink or expand) the distance value r of the rendering item that is fed into the distance-dependent gain calculation. 18. [BIN] Flag indicating whether directional focus is enabled. 19. [INT] Radius of main lobe in degrees of directional focus 20. [INT] Width of the transition region between the main lobe and the stopband in degrees 21. [FLOAT] Directivity gain reduction in the stopband in dB 22. [BIN] Flag indicating whether the directional focus has a non-default direction 23. [INT] Yaw angle of the main direction of the directional focus relative to the frontal head orientation 24. [INT] Pitch angle of the principal direction of the directional focus relative to the forward head orientation 25. [BIN] If this authoring parameter is set, that particular element will never be processed with the directional focus effect. A. Adaptation of the rendering of the host scene based on the host scene's complexity [6] and sensitivity [7] dictates, taking into account the complexity [6] and sensitivity [7] of the guest elements being added. B. Selecting or prioritizing the placement of guest elements based on the rendering complexity [6] and sensitivity [7] and quality [8] of the host scene. C. Selecting or prioritizing the placement of guest elements based on guest element priority [9] D. Selection or prioritization of guest element placement based on element category
[10] E. Allowing adaptation of the rendering of the original host scene, ignoring its own indications of complexity [6], sensitivity [7] and quality [8], but focusing on the category
[10] and priority [9] of the guest elements. F. Selection of an appropriate guest element from an alternative selection based on the technical configuration of the host scene, i.e., its audio configuration (bit depth, sampling rate, channel / object configuration, loudness [5]). G. Processing or alignment (e.g., resampling, loudness adjustment) of the rendering of host scenes and guest elements, each with their own / different audio configurations (bit depth, sampling rate, channel / object configuration, loudness [5]); H. Gradual adaptation to the host scene's predefined or metadata-controlled geometry [2] and coordinates [1] (+ inversive) a. Instantaneous collapse from full to smaller shape b. Instantaneous collapse from full to single point (mono) c. A gradual or stepwise collapse from full to smaller shapes d. A gradual or stepwise collapse from full to a single point (mono) e. Instantaneous change from a predefined or metadata-controlled shape to another predefined or metadata-controlled shape f. A gradual change from a predefined or metadata-controlled shape to another predefined or metadata-controlled shape I. Gradual adaptation of guest elements to predefined or metadata-controlled shapes [4] and coordinates [3] (+ inversion) a. Instantaneous expansion from null or mono to full shape b. Gradual or stepwise expansion from null or mono to full shape c. Instantaneous change from a predefined or metadata-controlled shape to another predefined or metadata-controlled shape d. Gradual change from a predefined or metadata-controlled shape to another predefined or metadata-controlled shape J. Host Scene Timed Rendering Operations (Squash / Zoom)
[12] K. Embedding and Removal of Guest Element Timed Expressions
[12] L. Omitting from the operation certain audio elements that are already embedded in the host scene (e.g., when there are multiple / concurrent elements to be embedded) marked accordingly
[13] M. Embedding multiple (concurrent) guest elements to augment a host scene - Collapse adjustments of the host scene may be controlled by metadata or may be predefined (e.g., apply none, select one, or apply multiple concurrent collapse directives) N. Adjust the host scene by moving host scene elements away from the space to be freed up, but leave the general host scene ambience (channel bed, HOA ambience) unchanged or attenuated
[11] , controlled either by metadata or a predefined value. O. Adjust the host scene by moving host scene elements away from the space to be freed, but leaving the general host scene ambience (channel bed, HOA ambience) unchanged. Guest elements can be amplified or attenuated, controlled either by metadata or predefined values
[11] . P. The host scene is left unmodified. Guest elements can be amplified or attenuated, controlled either by metadata or predefined values.
[11]
[0132] Apparatus for implementing the method according to the present disclosure Finally, the present disclosure also relates to respective apparatuses (e.g., computer-implemented apparatuses) for performing the methods and techniques described throughout the present disclosure. Each apparatus may be implemented as a renderer. The renderer may be implemented on a server or on an end device. The end device may be a wearable device. FIG. 10 shows an example of such an apparatus 1000. In particular, the apparatus 1000 comprises a processor 1010 and a memory 1020 coupled to the processor 1010. The memory 1020 may store instructions for the processor 1010. The processor 1010 may also receive appropriate input data 1030, among other things, depending on the use case and / or implementation. The processor 1010 may be adapted to perform the methods / techniques described throughout the present disclosure and generate corresponding output data 1040 depending on the use case and / or implementation.
[0133] interpretation Aspects of the apparatus described herein may be implemented in a computer-based sound processing network environment suitable for processing digital or digitized audio files. Portions of an adaptive audio system may include one or more networks with any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0134] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or in terms of their behavior, register transfers, logical components, and / or other characteristics as data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media on which such formatted data and / or instructions may be embodied include various forms of physical (non-transitory) non-volatile storage media, such as, but not limited to, optical, magnetic, or semiconductor storage media.
[0135] While one or more implementations have been described in connection with specific embodiments by way of example, it is to be understood that the one or more implementations are not limited to the disclosed embodiments. On the contrary, the intention is to cover various modifications and similar arrangements, as will be apparent to those skilled in the art.
[0136] Itemized Exemplary Embodiments Aspects and implementations of the present disclosure can be understood from the following enumerated example embodiments, which are not claims. [EEE1] 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements relative to the original 3D audio scene; modifying the original 3D audio scene to obtain a modified 3D audio scene; jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters; A method comprising: [EEE2] The method according to EEE1, wherein modifying the original 3D audio scene is based on said metadata. [EEE3] The method of any one of EEE1 and EEE2, wherein the original 3D audio scene is associated with a first spatial volume that surrounds the plurality of first audio elements, and modifying the original 3D audio scene comprises mapping the first spatial volume to a second spatial volume associated with the modified 3D audio scene, the second spatial volume being a modified spatial volume compared to the first spatial volume. [EEE4] The method according to EEE3, wherein the modified spatial volume is a reduced spatial volume, and optionally a volume fraction of the reduced spatial volume associated with the modified 3D audio scene is between 1% and 99% of the first spatial volume. [EEE5] The method according to EEE3, wherein the modified spatial volume is an extended spatial volume, and optionally a volume fraction of the extended spatial volume associated with the modified 3D audio scene is between 101% and 300% of the first spatial volume. [EEE6] 6. The method of any one of claims 3 to 5, wherein the first metadata includes first tagging information for the plurality of first audio elements indicating whether each first audio element should be moved in 3D audio space when mapping the first spatial volume to the second spatial volume, and only first audio elements indicated by the first tagging information as not to be moved are rendered in their original positions. [EEE7] The method of any one of EEE3 to 6, wherein mapping the first spatial volume to the second spatial volume is performed instantaneously. [EEE8] The method of any one of EEE3 to 6, wherein mapping the first spatial volume onto the second spatial volume is performed gradually or in stages over time. [EEE9] 9. The method of any one of claims 3 to 8, wherein modifying the original 3D audio scene further comprises selecting a third spatial volume in the 3D audio space for allocating the one or more second audio elements. [EEE10] The method of EEE9, wherein the third spatial volume at least partially overlaps with the first spatial volume. [EEE11] The method of claim 8, wherein mapping the first spatial volume to the second spatial volume includes removing each first audio element from the third spatial volume, and wherein only first audio elements that are indicated as not to be moved by the first tagging information are not removed from the third spatial volume. [EEE12] The method of claim 3, wherein removing the first audio element from the third volume of space is performed instantaneously. [EEE13] The method of claim 8, wherein removing the first audio element from the third spatial volume is performed gradually or in stages over time. [EEE14] 36. The method of any one of EEE9 to 13, wherein jointly rendering the one or more second audio elements and the modified 3D audio scene comprises rendering the one or more second audio elements in the third spatial volume and rendering the modified 3D audio scene in the second spatial volume. [EEE15] The method of any one of EEE3 to 14, wherein the shape of the first spatial volume, the second spatial volume, and / or the third spatial volume comprises one or more of a sphere, a quadrant, an octant, and a point. [EEE16] The method of claim 3, wherein the shape of the first spatial volume, the second spatial volume, and / or the third spatial volume is based on the metadata. [EEE17] The method of EEE15, wherein the shapes of the first spatial volume, the second spatial volume, and / or the third spatial volume are predefined. [EEE18] 18. The method of any one of EEE1 to 17, wherein modifying the original 3D audio scene and / or the collective rendering may be based on first timing information, the first timing information indicating a point in time or a time period for applying the modification and / or the collective rendering. [EEE19] 19. The method of any one of EEE1 to 18, further comprising, following said joint rendering, removing said one or more secondary audio elements and returning to rendering the original 3D audio scene. [EEE20] The method according to EEE19, wherein removing the one or more second audio elements and / or returning to rendering the original 3D audio scene is based on second timing information, the second timing information indicating a point in time or a time period for applying the removing and / or returning. [EEE21] The method of any one of EEE19 and EEE20, wherein the second metadata includes second tagging information for one or more second audio elements indicating whether each second audio element should be removed when returning to rendering the original 3D audio scene, and wherein only second audio elements indicated as not to be removed by the second tagging information are not removed before rendering the original 3D audio scene. [EEE22] 22. The method of any one of EEE1 to 21, wherein the one or more rendering parameters include an indication of the loudness of the one or more secondary audio elements, and wherein the joint rendering further comprises adjusting the loudness of the one or more secondary audio elements relative to the loudness of the modified 3D audio scene. [EEE23] 23. The method of any one of EEE1 to EEE22, wherein the joint rendering further comprises adapting gains of the modified 3D audio scene and / or the one or more secondary audio elements. [EEE24] 24. The method of any one of EEE1 to EEE23, wherein the method comprises receiving different versions of the second audio data for the one or more second audio elements, each version having a different audio configuration. [EEE25] The method according to EEE24, further comprising selecting a version of the second audio data for the one or more second audio elements that best matches an audio composition of the original 3D audio scene. [EEE26] 26. The method of claim 24 or 25, wherein the joint rendering further comprises aligning the one or more second audio elements with an audio composition of the modified 3D audio scene. [EEE27] 27. The method of any one of EEE1 to 26, wherein the one or more rendering parameters further comprise an indication of a perceived complexity of the one or more second audio elements. [EEE28] 30. The method of any one of EEE1 to 27, wherein the one or more rendering parameters further comprise an indication of sensitivity to perceived interference of the one or more second audio elements by other audio elements. [EEE29] 29. The method of any one of EEE1 to 28, wherein the one or more rendering parameters further comprise an indication of an audio quality of the one or more secondary audio elements. [EEE30] The method according to EEE29, further comprising, prior to said joint rendering, comparing an audio quality of said one or more secondary audio elements with an audio quality of the original 3D audio scene, and aligning the audio quality of said one or more secondary audio elements with the audio quality of the original 3D audio scene. [EEE31] 31. The method of any one of EEE1 to EEE30, wherein the one or more rendering parameters further include an indication of a priority level for each of the one or more secondary audio elements, the priority level indicating a priority for rendering each secondary audio element. [EEE32] The method of EEE31, wherein the collaborative rendering further comprises modifying some or all of the first audio elements based on their respective priority levels. [EEE33] 33. The method of any one of EEE1 to 32, wherein the one or more rendering parameters further comprise an indication of a category for each of the one or more second audio elements, the category comprising one or more of optional, mandatory, supplementary, urgent, subordinate, dependent, within the context of the original 3D audio scene and outside the context of the original 3D audio scene. [EEE34] The method of any one of claims 8 to 10, wherein the joint rendering is based on the perceived complexity and / or sensitivity of the original 3D audio scene, and optionally based on an indication of the perceived complexity and / or an indication of the sensitivity of the one or more second audio elements. [EEE35] The method according to EEE34 when citing EEE31, wherein the joint rendering is based on an indication of a priority level and an indication of a category of each of the one or more second audio elements, ignoring the perceived complexity and sensitivity of the original 3D audio scene. [EEE36] The method of any one of EEE27, 28, 29, 31 and / or 33 when citing EEE9 or 10, wherein selecting the third spatial volume or prioritizing the allocation of the one or more second audio elements to the third spatial volume is based on one or more of the perceived complexity, sensitivity, audio quality, priority level and category indication of the one or more second audio elements. [EEE37] The method of any one of EEE9 to 36, wherein a first audio element of the original 3D audio scene related to an ambient environment in the third spatial volume is left unmodified, and the joint rendering comprises attenuating or amplifying the one or more second audio elements relative to the first audio element related to an ambient environment. [EEE38] 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene; jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters; A method comprising: [EEE39] The method according to EEE38, wherein the joint rendering comprises embedding the one or more secondary audio elements into the original 3D audio scene such that the original 3D audio scene can be augmented by the one or more secondary audio elements. [EEE40] 39. The method of claim 38, wherein the collaborative rendering includes modifying some or all of the first audio element based on the one or more rendering parameters. [EEE41] 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space; receiving information indicating that the original 3D audio scene is to be modified; obtaining one or more rendering parameters from said information; rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene; A method comprising: [EEE42] The method of claim 8, wherein the information corresponds to default metadata associated with one or more default audio elements. [EEE43] The method according to EEE41, wherein the information is generated in real time. [EEE44] The method of claim 31, wherein the information is based on user preferences and / or location data. [EEE45] 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the second audio element relative to the original 3D audio scene; modifying the original 3D audio scene to obtain a modified 3D audio scene; and jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters. Device. [EEE46] The apparatus of EEE45, wherein the one or more processors are further configured to, after the joint rendering, remove the one or more second audio elements and return to rendering the original 3D audio scene. [EEE47] 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene; and jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters. Device. [EEE48] 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space; receiving information indicating that the original 3D audio scene is to be modified; obtaining one or more rendering parameters from said information; rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene; 1. An apparatus comprising: [EEE49] A program comprising instructions which, when executed by a processor, cause the processor to perform a method according to any one of EEE1 to EEE44. [EEE50] A computer-readable storage medium storing a program according to EEE49.
[0137] .
Claims
1. 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space; receiving information indicating that the original 3D audio scene is to be modified; obtaining one or more rendering parameters from said information; rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene; A method comprising:
2. The method of claim 1 , wherein rendering the original 3D audio scene is further based on the metadata.
3. 3. The method of claim 1, wherein the metadata includes tagging information for the plurality of first audio elements indicating whether each first audio element should be modified when rendering the original 3D audio scene.
4. The method of claim 1 , wherein the information is based on user preferences and / or location data.
5. The method of claim 4 , wherein the user preferences include a user-specific sensory impairment and / or a user's sensory preferences.
6. The method of claim 5 , wherein the user-specific sensory impairment includes hearing.
7. 7. The method of claim 6, wherein the hearing ability includes a spectral hearing loss, and wherein rendering the original 3D audio scene based on the one or more rendering parameters to obtain the modified 3D audio scene includes dynamic equalization to compensate for the spectral hearing loss separately for both ears.
8. The method of claim 5 , wherein the sensory preferences of the user include improved clarity and / or reduced sensory overload.
9. 9. The method according to claim 1, wherein the rendering parameters comprise parameters for attenuating early reflections, parameters for reverberation decay, and / or parameters for stretching / compressing distance-dependent gain.
10. The method of claim 1 , wherein the rendering step further comprises adapting a gain of the original 3D audio scene.
11. The method of claim 10 when dependent on claim 9, wherein the gain of the original 3D audio scene is the distance dependent gain.
12. The method of claim 11 when dependent on claim 3, wherein the tagging information indicates whether each first audio element should be rendered with the distance-dependent gain.
13. 13. The method of any one of claims 1 to 12, wherein the rendering parameters include one or more parameters for rendering the original 3D audio scene with directional focus.
14. The method of claim 13 , wherein the directional focus attenuates sound outside a user's viewing direction.
15. The method of claim 14 , wherein the attenuation of sounds outside the viewing direction of the user is a gradual attenuation.
16. 16. The method of claim 15, wherein the one or more parameters for rendering the original 3D audio scene with the directional focus include a flag for indicating whether the directional focus should be applied to the original 3D audio scene, an angular range corresponding to the viewing direction, a transition angle for the gradual attenuation, a maximum attenuation, a flag for indicating a default direction for the directional focus, a yaw angle for indicating the viewing direction, and / or a pitch angle for indicating the viewing direction.
17. 17. The method of any one of claims 13 to 16 when dependent on claim 3, wherein the tagging information indicates whether each first audio element should be rendered with the directional focus.
18. 18. The method of any one of claims 1 to 17, wherein the information corresponds to default metadata associated with one or more default audio elements.
19. 19. The method of any one of claims 1 to 18, wherein the information is information generated in real time.
20. 20. The method of claim 1, wherein the original 3D audio scene is associated with a first spatial volume surrounding the plurality of first audio elements, and rendering the original 3D audio scene comprises mapping the first spatial volume to a second spatial volume associated with the modified 3D audio scene, the second spatial volume being a modified spatial volume compared to the first spatial volume.
21. 21. The method of claim 20, wherein the modified spatial volume is a reduced spatial volume, and optionally a volume fraction of the reduced spatial volume associated with the modified 3D audio scene is between 1% and 99% of the first spatial volume.
22. 21. The method of claim 20, wherein the modified spatial volume is an extended spatial volume, and optionally a volume fraction of the extended spatial volume associated with the modified 3D audio scene is between 101% and 300% of the first spatial volume.
23. 23. A method according to any one of claims 20 to 22 when relying on claim 3, wherein the tagging information indicates whether each first audio element should be moved in 3D audio space when mapping the first spatial volume onto the second spatial volume, and only first audio elements indicated by the first tagging information as not to be moved are rendered in their original positions.
24. 24. The method of any one of claims 20 to 23, wherein mapping the first volume of space onto the second volume of space is performed instantaneously.
25. 24. The method of any one of claims 20 to 23, wherein mapping the first spatial volume onto the second spatial volume is performed gradually or stepwise over time.
26. 26. The method of any one of claims 20 to 25, wherein the shape of the first spatial volume and / or the second spatial volume comprises one or more of a sphere, a quadrant, an octant, and a point.
27. The method of claim 26 , wherein the shape of the first spatial volume and / or the second spatial volume is based on the metadata.
28. 27. The method of claim 26, wherein the shape of the first spatial volume and / or the second spatial volume is predefined.
29. 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements relative to the original 3D audio scene; modifying the original 3D audio scene to obtain a modified 3D audio scene; jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters; A method comprising:
30. 30. The method of claim 29, wherein modifying the original 3D audio scene is based on the metadata.
31. 40. The method of claim 29 or 39, wherein the original 3D audio scene is associated with a first spatial volume surrounding the plurality of first audio elements, and modifying the original 3D audio scene comprises mapping the first spatial volume to a second spatial volume associated with the modified 3D audio scene, the second spatial volume being a modified spatial volume compared to the first spatial volume.
32. 32. The method of claim 31 , wherein the modified spatial volume is a reduced spatial volume, and optionally a volume fraction of the reduced spatial volume associated with the modified 3D audio scene is between 1% and 99% of the first spatial volume.
33. 32. The method of claim 31 , wherein the modified spatial volume is an extended spatial volume, and optionally a volume fraction of the extended spatial volume associated with the modified 3D audio scene is between 101% and 300% of the first spatial volume.
34. 34. The method of claim 31, wherein the first metadata includes first tagging information for the plurality of first audio elements indicating whether each first audio element should be moved in 3D audio space when mapping the first spatial volume to the second spatial volume, and wherein only first audio elements indicated by the first tagging information as not to be moved are rendered in their original positions.
35. 35. The method of any one of claims 31 to 34, wherein mapping the first volume of space to the second volume of space is performed instantaneously.
36. 35. The method of any one of claims 31 to 34, wherein mapping the first volume of space onto the second volume of space is performed gradually or stepwise over time.
37. 37. The method of any one of claims 31 to 36, wherein modifying the original 3D audio scene further comprises selecting a third spatial volume in the 3D audio space for allocating the one or more second audio elements.
38. 38. The method of claim 37, wherein the third spatial volume at least partially overlaps the first spatial volume.
39. 39. The method of claim 37 or 38 when relying on claim 34, wherein mapping the first spatial volume to the second spatial volume includes removing each first audio element from the third spatial volume, and only first audio elements indicated by the first tagging information as not to be moved are not removed from the third spatial volume.
40. 40. The method of claim 39, wherein removing the first audio element from the third volume of space is performed instantaneously.
41. 40. The method of claim 39, wherein removing the first audio element from the third volume of space is performed gradually or in stages over time.
42. 42. The method of claim 38, wherein jointly rendering the one or more second audio elements and the modified 3D audio scene comprises rendering the one or more second audio elements in the third spatial volume and rendering the modified 3D audio scene in the second spatial volume.
43. 43. The method of any one of claims 31 to 42, wherein the shape of the first spatial volume, the second spatial volume, and / or the third spatial volume comprises one or more of a sphere, a quadrant, an octant, and a point.
44. 44. The method of claim 43, wherein the shape of the first spatial volume, the second spatial volume, and / or the third spatial volume is based on the metadata.
45. 44. The method of claim 43, wherein the shapes of the first spatial volume, the second spatial volume, and / or the third spatial volume are predefined.
46. 46. The method of any one of claims 29 to 45, wherein modifying the original 3D audio scene and / or the collective rendering may be based on first timing information, the first timing information indicating a point in time or a time period for applying the modification and / or the collective rendering.
47. 47. The method of any one of claims 29 to 46, further comprising, following the joint rendering, removing the one or more secondary audio elements and returning to rendering the original 3D audio scene.
48. 48. The method of claim 47, wherein removing the one or more second audio elements and / or returning to rendering the original 3D audio scene is based on second timing information, the second timing information indicating a point in time or a time period for applying the removing and / or returning.
49. 49. The method of claim 47 or 48, wherein the second metadata includes second tagging information for one or more second audio elements indicating whether each second audio element should be removed when returning to rendering the original 3D audio scene, and wherein only second audio elements indicated by the second tagging information as not to be removed are not removed before rendering the original 3D audio scene.
50. 50. The method of claim 29, wherein the one or more rendering parameters include an indication of a loudness of the one or more secondary audio elements, and wherein the joint rendering further includes adjusting a loudness of the one or more secondary audio elements relative to a loudness of the modified 3D audio scene.
51. 51. The method of any one of claims 29 to 50, wherein the joint rendering further comprises adapting gains of the modified 3D audio scene and / or the one or more secondary audio elements.
52. 52. The method of any one of claims 29 to 51, wherein the method comprises receiving different versions of the second audio data for the one or more second audio elements, each version having a different audio configuration.
53. 53. The method of claim 52, further comprising selecting a version of the second audio data for the one or more second audio elements that most closely matches an audio composition of an original 3D audio scene.
54. 54. The method of claim 52 or 53, wherein the joint rendering further comprises aligning the one or more second audio elements with an audio composition of the modified 3D audio scene.
55. 55. The method of any one of claims 29 to 54, wherein the one or more rendering parameters further comprise an indication of a perceived complexity of the one or more second audio elements.
56. 56. The method of any one of claims 29 to 55, wherein the one or more rendering parameters further comprise an indication of sensitivity to perceived interference of the one or more second audio elements by other audio elements.
57. 57. The method of any one of claims 29 to 56, wherein the one or more rendering parameters further comprise an indication of an audio quality of the one or more second audio elements.
58. 58. The method of claim 57, further comprising, prior to the joint rendering, comparing an audio quality of the one or more secondary audio elements with an audio quality of the original 3D audio scene, and aligning the audio quality of the one or more secondary audio elements with the audio quality of the original 3D audio scene.
59. 59. The method of claim 29, wherein the one or more rendering parameters further comprise an indication of a priority level for each of the one or more secondary audio elements, the priority level indicating a priority for rendering each secondary audio element.
60. 60. The method of claim 59, wherein the collaborative rendering further comprises modifying some or all of the first audio elements based on their respective priority levels.
61. 61. The method of any one of claims 29 to 60, wherein the one or more rendering parameters further comprise an indication of a category for each of the one or more second audio elements, the category comprising one or more of optional, mandatory, supplementary, urgent, subordinate, dependent, within the context of the original 3D audio scene and outside the context of the original 3D audio scene.
62. 57. A method according to claim 55 or 56, wherein the joint rendering is based on the perceived complexity and / or sensitivity of the original 3D audio scene, and optionally on an indication of the perceived complexity and / or sensitivity of the one or more second audio elements.
63. 63. The method of claim 62 when citing claim 59, wherein the joint rendering is based on an indication of a priority level and an indication of a category of each of the one or more second audio elements, ignoring the perceived complexity and sensitivity of the original 3D audio scene.
64. 62. The method of any one of claims 55, 56, 57, 59 and / or 61 when citing claim 37 or 38, wherein selecting the third spatial volume or prioritizing the allocation of the one or more second audio elements to the third spatial volume is based on one or more of the perceived complexity, sensitivity, audio quality, priority level and category indication of the one or more second audio elements.
65. 65. The method of any one of claims 37 to 64, wherein a first audio element of the original 3D audio scene related to an ambient environment in the third spatial volume is left unmodified, and the joint rendering comprises attenuating or amplifying the one or more second audio elements relative to the first audio element related to the ambient environment.
66. 1. A method for processing a 3D audio scene, the method comprising: receiving an original 3D audio scene, the original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene; jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters; A method comprising:
67. 67. The method of claim 66, wherein the collaborative rendering includes embedding the one or more secondary audio elements into the original 3D audio scene such that the original 3D audio scene can be augmented by the one or more secondary audio elements.
68. 68. The method of claim 66 or 67, wherein the collaborative rendering includes modifying some or all of the first audio elements based on the one or more rendering parameters.
69. 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the second audio element relative to the original 3D audio scene; modifying the original 3D audio scene to obtain a modified 3D audio scene; jointly rendering the one or more second audio elements and the modified 3D audio scene based on the one or more rendering parameters; Including, Device.
70. 70. The apparatus of claim 69, wherein the one or more processors are configured to, after the joint rendering, remove the one or more second audio elements and return to rendering the original 3D audio scene.
71. 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including first audio data and first metadata for a plurality of first audio elements in a 3D audio space; receiving second audio data and second metadata for one or more second audio elements; extracting one or more rendering parameters from the second metadata for rendering the one or more second audio elements within the original 3D audio scene; jointly rendering the one or more second audio elements and the original 3D audio scene based on the one or more rendering parameters. Device.
72. 1. An apparatus for processing an audio scene, the apparatus comprising one or more processors configured to perform a method, the method comprising: receiving an original 3D audio scene including audio data and metadata for a plurality of first audio elements in a 3D audio space; receiving information indicating that the original 3D audio scene is to be modified; obtaining one or more rendering parameters from said information; rendering the original 3D audio scene based on the one or more rendering parameters to obtain a modified 3D audio scene; 1. An apparatus comprising:
73. 69. A program comprising instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 1 to 68.
74. 74. A computer readable storage medium storing the program of claim 73.