Apparatus and method for reproducing a spatially extended sound source or apparatus and method for generating a description of a spatially extended sound source using anchor information

By calculating the listener's position and rendering the sound source using shell projection, the problem of reproducing spatially extended sound sources in virtual reality and augmented reality is solved, achieving high-quality spatial range perception and stable timbre, while reducing the number of rendering point sources.

CN115280800BActive Publication Date: 2026-03-17FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-13
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively reproduce spatially extended sound sources with complex geometries, especially in virtual reality and augmented reality. Traditional methods rely on the listener's position and are perceptually unstable, often leading to changes in timbre and signal artifacts.

Method used

By calculating the relative position of the listener and the spatially extended sound source, at least two sound sources are rendered using shell projection on the projection plane, reducing the number of rendered point sources. A description of the spatially extended sound source is generated through decorrelation processing, which is suitable for two-dimensional or three-dimensional audio reproduction.

Benefits of technology

It achieves high-quality spatial extension sound source reproduction, reduces computation and rendering overhead, ensures a stable spatial range perceived by listeners in virtual reality and augmented reality, and avoids timbre alteration and signal artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280800B_ABST
    Figure CN115280800B_ABST
Patent Text Reader

Abstract

Apparatus for reproducing a spatially extended sound source having a defined position or orientation and geometry in a space, the apparatus comprising an interface (100) for receiving a listener position; a projector (120) for calculating a projection of a two- or three-dimensional hull associated with the spatially extended sound source onto a projection plane using the listener position, geometry information about the spatially extended sound source and position information about the spatially extended sound source; a sound position calculator (140) for calculating positions of at least two sound sources of the spatially extended sound source using the projection plane; and a renderer (160) for rendering the at least two sound sources at the positions to obtain a reproduction of the spatially extended sound source having two or more output signals, wherein the renderer (160) is configured to use different sound signals for different positions, wherein the different sound signals are associated with the spatially extended sound source, and wherein the renderer (160) is configured to render the at least two sound sources in response to a specific information received relative to a fixed position and / or orientation of the spatially extended sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to audio signal processing, and more particularly to the encoding, decoding, or reproduction of spatially extended sound sources. Background Technology

[0002] For a long time, people have been studying the reproduction of sound sources using multiple speakers or headphones. The simplest way to reproduce sound sources in this setup is to render them as point sources, i.e., very (ideally: infinitely) small sound sources. However, this theoretical concept is difficult to realistically simulate existing physical sound sources. For example, a grand piano has a large vibrating wooden enclosure with many spatially distributed strings inside, so it appears much larger to the ear than a point source (especially when the listener (and microphone) is close to the grand piano). Many real-world sound sources have a considerable size (“spatial range”), such as musical instruments, machines, orchestras or choirs, or ambient sounds (the sound of a waterfall).

[0003] Accurate / realistic reproduction of such sound sources has become the goal of many sound reproduction methods, whether using binaural headphones (i.e., using the so-called head-related transfer function HRTF or binaural room impulse response BRIR) or traditionally using speaker setups ranging from 2 speakers (“stereo”) to multiple speakers arranged on a horizontal plane (“surround sound”) and multiple speakers surrounding the listener in all three dimensions (“3D audio”). Summary of the Invention

[0004] The purpose of this invention is to provide a concept for encoding or reproducing spatially extended sound sources with potentially complex geometries.

[0005] Two-dimensional source width

[0006] This section describes methods related to rendering extended sound sources on a 2D surface facing the listener, such as within a certain azimuth range at zero elevation (as in the case of traditional stereo / surround sound) or a range of azimuth and elevation angles (as in the case of 3D audio or virtual reality, where the user's movement has 3 degrees of freedom ["3DoF"], i.e., the rotation of the head on the pitch / yaw / roll axes).

[0007] The apparent width of an audio object translated between two or more speakers can be increased (generating a so-called phantom or phantom source) by reducing the correlation of the signals involved in the audio channels (Blauert, 2001, pp. 241-257). As the correlation decreases, the propagation of the phantom source increases until the correlation value approaches zero (and the aperture angle is not too wide), covering the entire range between the speakers.

[0008] Decorcorrelated versions of the source signal are obtained by deriving and applying appropriate decorrelation filters. Lauridsen (Lauridsen, 1954) proposed adding / subtracting time-delayed and scaled versions of the source signal to obtain two decorrelation versions of the signal. For example, a more complex method was proposed by Kendall (Kendall, 1995). He derived pairwise decorrelation all-pass filters based on iterative combinations of random number sequences. Faller et al. (Baumgarte & Faller, 2003) proposed appropriate decorrelation filters (“diffusers”). Zotter et al. also derived filter pairs where frequency-dependent phase or amplitude differences are used to broaden the phantom source (Zotter & Frank, 2013). Furthermore, (Alary, Politis, & (2017) proposed a decorrelation filter based on velvet noise, and it was developed by (Schlecht, Alary, (&Habets, 2018) Further optimization.

[0009] Besides reducing the correlation of the corresponding channel signals of phantom sources, source width can also be increased by increasing the number of phantom sources attributed to an audio object. In (Pulkki, 1999), source width is controlled by translating the same source signal to a (slightly) different direction. This method was originally proposed to stabilize the perceptual phantom source propagation when VBAP-translated (Pulkki, 1997) source signals move in a sound scene. This is advantageous because, depending on the direction of the source, the rendered source is reproduced by two or more speakers, which can lead to undesirable changes in the perceptual source width.

[0010] Virtual World DirAC (Pulkki, Laitinen, & Erkut, 2009) is an extension of the traditional Directional Audio Coding (DirAC) (Pulkki, 2007) method for sound synthesis in virtual worlds. To render spatial extent, the directional sound components of the source are randomly translated within a certain range around the source's original direction, where the translation direction varies with time and frequency.

[0011] ( A similar approach was used in Santala & Pulkki (2014), where spatial range was achieved by randomly distributing the frequency bands of the source signal across different spatial directions. This is a method aimed at producing sound with the same spatial distribution and envelope from all directions, rather than controlling the degree of precision.

[0012] Verron et al. achieved spatial extenting of the source without using translation-dependent signals. Instead, they synthesized multiple disjoint versions of the source signals, evenly distributed them in a circle around the listener, and mixed them together (Verron, Aramaki, Kronland-Martinet & Pallone, 2010). The number of activated sources and their gain determined the intensity of the broadening effect. This method was implemented for spatial extension in ambient sound synthesizers.

[0013] 3D source width

[0014] This section describes methods related to rendering extended sound sources in 3D space, specifically in a volumetric manner, as this is necessary for virtual reality with 6 degrees of freedom (“6DoF”). This means 6 degrees of freedom for user movement (i.e., rotation of the head on the pitch / yaw / roll axes) plus 3 translational motion directions x / y / z.

[0015] Potard et al. extended the concept of source range to a one-dimensional parameter of the source (i.e., the width between two loudspeakers) by studying the perception of source shape (Potard, 2003). They generated multiple incoherent point sources by applying (time-varying) decorrelation techniques to the original source signal and then placing incoherent sources in different spatial locations, thus giving them a three-dimensional range (Potard & Burnett, 2004).

[0016] In MPEG-4Advanced AudioBIFS (Schmidt & In 2004, volume objects / shapes (shells, boxes, ellipses, and cylinders) can be filled with several uniformly distributed and decorrelated sound sources to evoke a three-dimensional sound source range.

[0017] To increase and control the source range using Ambisonics, Schmele et al. (Schmele & Sayin, 2018) proposed mixing that reduces the Ambisonics order of the input signal, which essentially increases the apparent source width and distributes a decorrelation copy of the source signal around the listening space.

[0018] Zotter et al. introduced another approach, which adopted the principle proposed in (Zotter & Frank, 2013) (i.e., deriving filter pairs, introducing frequency-dependent phase and amplitude differences to achieve source range in stereo reproduction settings) for Ambisonics (Zotter F., Frank, Kronlachner & Choi, 2014).

[0019] A common drawback of translation-based methods (e.g., (Pulkki, 1997)(Pulkki, 1999)(Pulkki, 2007)(Pulkki, Laitinen, & Erkut, 2009)) is their dependence on the listener's position. Even a small deviation from the most effective point can cause the spatial image to collapse into the speaker closest to the listener. This severely limits their application in virtual reality and augmented reality scenarios with 6 degrees of freedom (6DoF), where the listener should be able to move freely. Furthermore, allocating time-frequency bits in DirAC-based methods (e.g., (Pulkki, 2007)(Pulkki, Laitinen, & Erkut, 2009)) does not always guarantee the correct rendering of the spatial extent of the phantom source. Moreover, it often significantly degrades the timbre of the source signal.

[0020] Decorrelation of the source signal is typically achieved through one of the following methods: i) deriving a filter pair with complementary amplitudes (e.g., Lauridsen, 1954); ii) using an all-pass filter with constant amplitude but (randomly) scrambled phase (e.g., Kendall, 1995, Potard & Burnett, 2004); or iii) spatially distributing the time-frequency bits of the source signal randomly (e.g., ...). Santala, & Pulkki, 2014).

[0021] Each method has its own implications: according to i), complementary filtering of the source signal generally leads to an alteration in the perceived timbre of the decorrelated signal. While all-pass filtering in ii) preserves the timbre of the source signal, scrambling the phase disrupts the original phase relationship and, especially for transient signals, causes severe temporal dispersion and trailing artifacts. Spatially distributed time-frequency bits have proven effective for some signals but also alter the perceived timbre. Furthermore, it exhibits high signal dependence and introduces severe artifacts for impulsive signals.

[0022] Using Advanced AudioBIFS ((Schmidt& The multiple decorrelation versions of the source signal proposed in Potard (2004) and Potard & Burnett (2003) to fill a volume shape assume that there are a large number of mutually generating filters available for decorrelating the output signal (typically, each volume shape uses more than a dozen point sources). However, finding such filters is not a simple task, and the more filters required, the more difficult it becomes. Furthermore, if the source signal is not fully decorrelated and the listener moves around such a shape, for example in a (virtual reality) scene, the individual source distances to the listener correspond to different delays in the source signal and their superposition at the listener's ear, which can lead to position-dependent comb filtering and potentially introduce annoying source signal instability coloring.

[0023] By reducing the Ambisonics order, which only has an audible effect for transitions from the 2nd to the 1st or 0th order, the source width is controlled using Ambisonics-based techniques (Schmele & Sayin, 2018). Furthermore, these transitions are not only perceived as source expansion but are often thought of as the movement of the phantom source. While adding a decorrelated version of the source signal helps stabilize the perception of a significant source width, it also introduces a comb filter effect that alters the timbre of the phantom source.

[0024] The purpose of this invention is to provide an improved concept for describing the reproduction of spatially extended sound sources or the generation of spatially extended sound sources.

[0025] This objective is achieved by the apparatus for reproducing a spatially extended sound source, the apparatus for generating a bitstream, the method for reproducing a spatially extended sound source, the method for generating a description of a spatially extended sound source, the description of a spatially extended sound source, or a computer program of the present invention.

[0026] This invention is based on the discovery that the reproduction of a spatially extended sound source can be achieved, and in particular, rendering can be realized, by calculating the projection of a two-dimensional or three-dimensional shell associated with the spatially extended sound source onto a projection plane using the listener's position. This projection is used to calculate the positions of at least two sound sources of the spatially extended sound source, and at least two sound sources are rendered at those positions to obtain the reproduction of the spatially extended sound source. The rendering results in two or more output signals, and different sound signals are used at different positions, but the different sound signals are all associated with the same spatially extended sound source.

[0027] High-quality two-dimensional or three-dimensional audio reproduction is achieved because, on the one hand, the time-varying relative position between the spatially extended sound source and the (virtual) listener position is considered. The listener position may include only the user's geometric position, or it may be the user's orientation in space, or it may be both the user's geometric position and orientation. On the other hand, the spatially extended sound source is effectively represented by geometric information of the perceived sound source range and a number of at least two sound sources, such as peripheral point sources, which can be easily processed by well-known renderers. The geometric information is preferably acoustically valid geometric information. For example, a curtain is acoustically transparent but optically opaque. The situation is different for a thick glass wall. This wall is optically transparent but acoustically opaque. In particular, direct renderers in the art are always positioned to render sound sources at specific locations relative to a particular output format or speaker setting. For example, two sound sources calculated by a sound position calculator at certain locations can be rendered at those locations by, for example, amplitude translation.

[0028] For example, when a sound source is positioned between left-left surround channels in a 5.1 output format, and other sound sources are positioned between right-right surround channels in the output format, the amplitude shifting process performed by the renderer will result in very similar signals for the left and left-left surround channels of one sound source and correspondingly very similar signals for the right and right surround channels of the other sound source, so that the user perceives the sound source as originating from a location calculated by the sound location calculator. However, since all four signals ultimately relate to a spatially extended sound source, the user does not simply perceive two phantom sources associated with locations calculated by the sound location calculator; instead, the listener perceives a single spatially extended sound source.

[0029] An apparatus for reproducing a spatially extended sound source with a defined geometrical position and / or orientation in space includes an interface, a projector, a sound position calculator, and a renderer. This invention allows for consideration of enhanced sound conditions, such as those occurring within a piano. A piano is a large device, and its sound has traditionally been rendered as originating from a single point source. However, this does not fully represent the true sonic characteristics of a piano. According to the invention, a piano, as an example of a spatially extended sound source, is reflected by at least two sound signals. One sound signal can be recorded by a microphone located near the left side of the piano, i.e., near the bass strings, while the other sound source can be recorded by a different second microphone located near the right side of the piano, i.e., near the treble strings that produce high notes. Of course, due to the reflection conditions within the piano, both microphones will record different sounds from each other, also because the bass strings are closer to the left microphone than the right microphone, and vice versa. However, on the other hand, both microphone signals will have a large number of similar sonic components, ultimately constituting the unique sound of the piano. Specifically, the renderer is configured to render at least two sound sources relative to a fixed position and / or orientation of the spatially extended sound source in response to specific information received, i.e., in response to anchoring information.

[0030] According to the invention, a bitstream representing a spatially extended sound source, such as a piano, is generated by recording signals, by also recording geometric information about the spatially extended sound source and, optionally, by recording location information associated with different microphone locations (or two different locations typically associated with two different sound sources) or providing a description of the perceptual geometry of the sound (of a piano). To reflect the listener's position relative to the sound source, i.e., the listener may be "moving around" in virtual reality, augmented reality, or any other sound scene, the encapsulated projection of the sound source, such as a piano, associated with the spatially extended sound is calculated using the listener's position, and the positions of at least two sound sources are calculated using the projection plane, wherein, in particular, the preferred embodiment relates to the localization of the sound source at a point on the periphery of the projection plane. An output data formulator is configured to incorporate anchoring information or bitstream / description elements or flags into the description of the spatially extended sound source, which indicates the absolute anchoring of one or more different sound signals for the spatially extended sound source to the location or orientation of the spatially extended sound source. It can be used to describe spatially extended sound sources, such as as an XML description, a bitstream or compressed bitstream, or any other computer-readable format.

[0031] Example piano sounds in two-dimensional or three-dimensional cases can be practically represented by reducing computational and rendering overhead, so that, for example, when the listener is closer to the left side of a sound source such as a piano, the sound perceived by the listener is different from the sound that occurs when the user is closer to the right side of a sound source such as a piano or even behind a sound source such as a piano.

[0032] In view of the above, the unique aspect of this invention lies in providing a method for characterizing spatially extended sound sources on the encoder side, which allows the use of spatially extended sound sources in sound reproduction to achieve a true two-dimensional or three-dimensional setup. Furthermore, by calculating the projection of the two-dimensional or three-dimensional enclosure onto the projection plane using the listener's position, it is possible to use the listener's position in a highly flexible description of the spatially extended sound source in an efficient manner. The sound positions of at least two sound sources of the spatially extended sound source are calculated using the projection plane, and the positions of at least two sound sources calculated by the sound position calculator are rendered to obtain the reproduction of the spatially extended sound source, which has two or more headphone output signals or two or more multi-channel output signals in a stereo reproduction setup or a reproduction setup with more than two channels, such as five, seven, or even more channels.

[0033] Compared to existing methods that fill a 3D volume with sound by placing numerous different point sources throughout all parts of the volume to be filled, this projection avoids modeling numerous sound sources and significantly reduces the number of point sources used by only needing to fill the projection shell, i.e., a two-dimensional space. Furthermore, by optimizing the modeling of sound sources on the projection shell, the required number of point sound sources is further reduced; in extreme cases, it may be as simple as one sound source at the left boundary of the spatially extended sound source and one sound source at the right boundary of the spatially extended sound source. Both reduction steps are based on two psychoacoustic observations:

[0034] 1. The distance to a sound source cannot be reliably perceived compared to its azimuth (and elevation). Therefore, projecting the original volume onto a plane perpendicular to the listener does not significantly alter perception (but helps reduce the number of point sources required for rendering).

[0035] 2. Two decorrelated sounds are distributed as point sources on the left and right sides, respectively, and tend to fill the space between them with sound in perception.

[0036] Furthermore, the encoder end not only allows for the characterization of a single spatially extended sound source, but is also flexible, as descriptions representing the generated data, such as bitstreams, can include all data about two or more spatially extended sound sources, preferably related to a single coordinate system, including their geometric information and positions. At the decoder end, not only can a single spatially extended sound source be reproduced, but multiple spatially extended sound sources can also be reproduced, where a projector calculates the projection of each sound source using the (virtual) listener position. Additionally, a sound position calculator calculates the positions of at least two sound sources for each spatially extended sound source, and a renderer renders all calculated sound sources for each spatially extended sound source, for example, by adding two or more output signals from each spatially extended sound source in a signal-by-signal or channel-by-channel manner, and by providing the added channels to corresponding headphones for binaural reproduction or to corresponding speakers in a speaker-associated reproduction setup, or alternatively, to memory for storing (combined) two or more output signals for later use or transmission.

[0037] On the generator or encoder side, a device for generating a description of a spatially extended sound source is used to generate the description. This device includes a sound provider for providing one or more different sound signals to the spatially extended sound source, and an output data shaper for generating a description of the sound scene. This description includes one or more different sound signals, preferably compressed, such as by a bit-rate compression encoder, for example, an MP3, AAC, USAC, or MPEG-H encoder. The output data shaper is further configured to incorporate optional individual position information for each of the two or more different sound signals into the description, in the case of two or more different sound signals. This position information indicates the position of the corresponding sound signal and preferably has geometric information regarding the spatially extended sound source, i.e., the first signal is the signal recorded on the left side of the piano in the example above, and the signal recorded on the right side of the piano.

[0038] However, alternatively, the location information does not necessarily have to be related to the geometry of the spatially extended sound source, but can be related to the general coordinate origin, although being related to the geometry of the spatially extended sound source is preferred.

[0039] Furthermore, the apparatus for generating the description also includes a geometry provider for calculating information about the geometry of the spatially extended sound sources, and an output data formulator configured to incorporate information about the geometry, information about the individual location of each sound signal, and at least two sound signals, such as those recorded by a microphone, into the description. However, the sound provider does not necessarily have to actually pick up the microphone signal, but may also generate the sound signal on the encoder side using decorrelation processing, depending on the specific circumstances. Simultaneously, for spatially extended sound signals, only a small amount of sound signal, or even a single sound signal, may be transmitted, and decorrelation processing may be used on the reproduction side to generate the remaining sound signal. This is preferably represented by a description or bitstream element in the bitstream, so that the sound reproducer always knows how many sound signals are included in each spatially extended sound source, thus allowing the reproducer to determine, particularly within the sound location calculator, how many sound signals are available and how many sound signals should be derived on the decoder side, such as through signal synthesis or correlation processing.

[0040] In this embodiment, the output data former writes bitstream elements into a description or bitstream indicating the number of sound signals included for spatially extended sound sources. On the decoder side, the sound reproducer obtains bitstream elements from the transmitted description or bitstream, reads the bitstream elements, and determines, based on the bitstream elements, how much signal is used for preferred peripheral point sources or auxiliary sources located among peripheral sound sources. This calculation must be based on at least one received sound signal in the bitstream. The description of the spatially extended sound sources can be implemented, for example, as an XML description, a bitstream or compressed bitstream, or any other computer-readable format. Attached Figure Description

[0041] The preferred embodiments of the present invention will then be discussed in conjunction with the accompanying drawings, wherein:

[0042] Figure 1 This is an overview of a block diagram of a preferred embodiment of the reproduction side;

[0043] Figure 2 A spherical spatially extended sound source with a different number of peripheral point sources is shown;

[0044] Figure 3 An elliptical spatially extended sound source with multiple peripheral point sources is shown;

[0045] Figure 4 The diagram illustrates the distribution of peripheral point source locations using different methods for extending the linear spatial sound source.

[0046] Figure 5 A cuboid spatially extended sound source with different processes is shown to distribute peripheral point sources;

[0047] Figure 6 This illustrates spherical spatially extended sound sources at different distances;

[0048] Figure 7 This illustrates a piano-shaped spatially extended sound source within an approximate parametric elliptic shape;

[0049] Figure 8 A piano-shaped spatially extended sound source with three peripheral point sources distributed at the extreme points of the projected convex hull is shown;

[0050] Figure 9 A preferred embodiment of an apparatus or method for reproducing a spatially extended sound source is shown;

[0051] Figure 10 A preferred embodiment of an apparatus or method for generating a spatially extended sound source is shown;

[0052] Figure 11 It shows the result of Figure 10 The preferred embodiment described in the diagram is generated by the apparatus or method shown.

[0053] Figure 12a The image shows an object source with a cylindrical range and "user" alignment, observed in the right anterior hemisphere of the listener;

[0054] Figure 12b The image shows an object source with a cylindrical range and "user" alignment, observed in the left anterior hemisphere of the listener;

[0055] Figure 13 The relative signal channel positions are shown;

[0056] Figure 14a The image shows an object source (piano) with a box-shaped range and "object" alignment, i.e., a piano with orientation (front), range geometry, and label plane;

[0057] Figure 14b This shows the object source (piano) with a box-shaped range and "object" alignment, as viewed from the front of the piano; and

[0058] Figure 14c The image shows the object source (piano) with a box-shaped range and "object" alignment as viewed from the side of the piano. Detailed Implementation

[0059] Figure 9A preferred embodiment of an apparatus for reproducing a spatially extended sound source having a defined location or orientation and geometry in space is shown. The apparatus includes an interface 100, a projector 120, a sound location calculator 140, and a renderer 160. The interface is configured to receive a listener's location. Furthermore, the projector 120 is configured to use the listener's location received by the interface 100, along with information about the geometry of the spatially extended sound source, and information about the location of the spatially extended sound source in space, to calculate the projection of a two-dimensional or three-dimensional shell associated with the spatially extended sound source onto a projection plane. Preferably, the defined location or orientation of the spatially extended sound source in space, and the additional geometry of the spatially extended sound source in space, are received for reproducing the spatially extended sound source via a bitstream or a description reaching a demultiplexer or scene or description resolver 180. The demultiplexer 180 extracts geometric information about the spatially extended sound source from the description and provides this information to the projector. Additionally, the demultiplexer also extracts the location of the spatially extended sound source from the description or bitstream and forwards this information to the projector. Preferably, the description also includes location information for at least two different sound sources, and preferably, the demultiplexer also extracts compressed representations of at least two sound sources from the description, and the at least two sound sources are decompressed / decoded by the decoder as audio decoder 190. The decoded at least two sound sources are ultimately forwarded to renderer 160, and the renderer renders the at least two sound sources at the locations provided to renderer 160 by sound location calculator 140. Specifically, renderer 160 is configured to render at least two sound sources at fixed positions and / or orientations relative to spatially extended sound sources in response to specific information received, i.e., in response to anchoring information. The description of the spatially extended sound sources can be implemented as, for example, an XML description, a bitstream or compressed bitstream, or any other computer-readable format.

[0060] The use of anchoring information is particularly suitable for spatially extended sound sources defined by a multi-channel signal. In this scenario, each individual channel has associated alignment information. This alignment information can be, for example, left alignment of the left channel and right alignment of the right channel. Depending on the anchoring mode used—whether it is a "user-aligned" or "object-aligned" mode—and depending on the location information mapping a channel of the multi-channel signal to an external sound source, channels or waveforms are mapped to external sound sources based on the observer's position and orientation, i.e., the listening position, and based on the anchoring mode, and used by the renderer. Therefore, in this embodiment, the anchoring mode is used to interpret the location information as user-related or object-related. Thus, at least two sound sources determined by the sound location calculator are rendered by the renderer in response to the anchoring information.

[0061] although Figure 9A bitstream-related reproduction apparatus with a bitstream demultiplexer 180 and an audio decoder 190 is shown, but reproduction can also occur in situations different from the encoder / decoder scenario. For example, the location or orientation and geometry defined in space may already exist in the reproduction apparatus, such as in a virtual reality or augmented reality scenario, where data is generated on-site and used in the same location. The bitstream demultiplexer 180 and the audio decoder 190 are not actually necessary, and information about the geometry and location of the spatially extended sound sources is available without any extraction from the bitstream. Furthermore, the location information that associates the locations of at least two sound sources with information about the geometry of the spatially extended sound sources can also be pre-negotiated and therefore does not need to be transmitted from the encoder to the decoder, or alternatively, this data is generated on-site again.

[0062] Therefore, it should be noted that the location information is provided only in the embodiments and is not required to be sent even in the case of two or more sound source signals. For example, a decoder or reproducer can always treat the first sound source signal in the bitstream or description as a sound source placed on a projection further to the left. Similarly, the second sound source signal in the bitstream can be regarded as a sound source on a projection, with its projection position further to the right.

[0063] Furthermore, while the sound location calculator uses a projection plane to calculate the positions of at least two sound sources of spatially expanded sound sources, these at least two sound sources do not necessarily have to be received from the description or bitstream. Instead, only a single sound source of the at least two sound sources can be received via the bitstream and other sound sources, and therefore, other positions or location information can be actually generated only on the reproduction side without sending such information from the description generator to the reproducer. However, in other embodiments, all this information can be transmitted, and furthermore, when bit rate requirements are not stringent, a higher number of sound signals than one or two can be transmitted in the bitstream, and the audio decoder 190 will decode two, three, or even more sound signals representing the positions of the at least two sound sources calculated by the sound location calculator 140.

[0064] Figure 10 The encoder side of this scene is shown when reproduction is applied within the encoder / decoder application. Figure 10 An apparatus for generating a description of a spatially extended sound source is shown. Specifically, a sound provider 200 and an output data generator 240 are provided. In this embodiment, the spatially extended sound source is represented by a compressed description having one or more different sound signals, and the output data generator generates a description representing a preferably compressed sound scene, wherein the description includes at least one or more different sound signals and geometric information associated with the spatially extended sound source. This represents regarding... Figure 9 The situation shown, where all other information, such as the location of the spatially extended sound source (see...), Figure 9 The dashed arrow in the block (corresponding to projector 120) can be freely selected by the user on the reproduction side. Therefore, a unique description is provided of a spatially extended sound source having at least one or more different sound signals for this spatially extended sound source, wherein these sound signals are merely point source signals.

[0065] The apparatus for generation also includes a geometry provider 220 for providing information such as calculating the geometry of the spatially extended sound source. Other ways of providing geometry information other than calculation include receiving user input, such as graphics drawn manually by the user or any other information provided by the user, for example, through speech, tone, gestures, or any other user action. In addition to one or more different sound signals, information about the geometry is also incorporated into the description or bitstream.

[0066] Optionally, information regarding the individual location of each of one or more different sound signals is also incorporated into the bitstream, and / or location information regarding the spatially extended sound source is also incorporated into the bitstream or description. The location information for the sound source can be separate from or included within the geometric information. In the first case, the geometric information can be given relative to the location information. In the second case, the geometric information can include, for example, for a sphere, the center point and radius or diameter in coordinates. For a box-shaped spatially extended sound source, eight or at least one corner point can be given in absolute coordinates.

[0067] The location information of each of one or more different sound signals is preferably associated with geometric information about the spatially extended sound source. However, alternatively, absolute location information associated with the same coordinate system, where giving the location or geometry of the spatially extended sound source is useful, or alternatively, the geometry information can also be given using absolute coordinates within an absolute coordinate system rather than in a relative manner. However, providing this data in a relative manner, independent of a general coordinate system, allows the user to locate the spatially extended sound source within its own reproduction setting or itself, such as pointing to... Figure 9 The projector 120 is shown by the dashed line.

[0068] In a further embodiment, Figure 10 The sound provider 200 is configured to provide at least two different sound signals to a spatially extended sound source, and the output data former is configured to generate a bit stream such that the bit stream includes at least two different sound signals, preferably in an encoded format, and optionally, the individual position information of each of the at least two different sound signals in absolute coordinates or relative to the geometry of the spatially extended sound source.

[0069] In embodiments, the sound provider is configured to perform recording of natural sound sources at individual multiple microphone locations or orientations, or, for example, regarding... Figure 1 The one or more decorrelation filters discussed in blocks 164 and 166 perform the decorrelation of an audio signal from a single or multiple fundamental signals. The fundamental signal used in the generator may be the same as or different from the fundamental signal provided at the reproduction station or transmitted from the generator to the reproducer.

[0070] In a further embodiment, geometry provider 220 is configured to derive a parametric description or polygonal description from the geometry of the spatially extended sound source, and output data former is configured to introduce this parametric description or polygonal description into a bitstream.

[0071] Furthermore, in a preferred embodiment, the output data former is configured to incorporate a descriptive element into the bitstream or description, wherein this bitstream element indicates the number of at least one distinct audio signal for a spatially extended sound source, which is included in the bitstream or in an encoded audio signal associated with the bitstream, wherein the number is 1 or greater than 1. The bitstream generated by the output data former is not necessarily a complete description of both audio waveform data and metadata. Rather, the description or bitstream may simply be a separate metadata bitstream, including, for example, a descriptive field for the number of audio signals for each spatially extended sound source, geometric information about the spatially extended sound source, and, in embodiments, location information about the spatially extended sound source, and optionally, location information for each audio signal and each spatially extended sound source, geometric information about the spatially extended sound source, and, in embodiments, location information about the spatially extended sound source. The waveform audio signal, typically available in compressed form, is transmitted to the reproducer via a separate data stream or a separate transmission channel so that the reproducer receives encoded metadata from one source and (encoded) waveform signals from different sources.

[0072] The output data shaper (240) is further configured to introduce flags, bitstreams, or bitstream elements into the description or... Figure 10 The information shown at point 322 indicates the absolute anchoring of different sound signals from one or more spatially extended sound sources to the location or orientation of the spatially extended sound sources. Anchoring information 322 can be generated automatically or manually by the creator of the sound scene or spatially extended sound sources. Individual channels can be actually recorded in certain locations (such as in the piano example, via a first microphone located on the left side of the piano and a second microphone located on the right side of the piano) or can be synthesized or used with virtual microphones. In object anchoring mode, the location information of the sound signal or waveform will be derived from the microphone location or will be derived from the microphone location itself.

[0073] Furthermore, embodiments describing the generator include a controller 250. The controller 250 is configured to control the sound provider 200 regarding the amount of audio signal to be provided by the sound provider. According to this process, the controller 250 also provides bitstream element information to the output data formulator 240, indicated by shaded lines representing optional features. The output data formulator incorporates specific information about the amount of audio signal as bitstream elements by the controlled controller 250, and these elements are provided by the sound provider 200. Preferably, the amount of audio signal is controlled such that the output bitstream, including the encoded audio audio signal, meets external bitrate requirements. When the allowed bitrate is high, the sound provider will provide more audio signal compared to a lower allowed bitrate. In extreme cases, when bitrate requirements are stringent, the sound provider will only provide a single audio signal for a spatially extended sound source.

[0074] The reproducer will read the bitstream elements set accordingly and will continue to synthesize a corresponding number of other audio signals on the decoder side within the renderer 160 using the transmitted audio signals, so as to generate the final required number of peripheral point sources and optional auxiliary sources.

[0075] However, when bit rate requirements are less stringent, controller 250 will control the sound provider to supply a large number of different sound signals, for example, signals recorded by a corresponding number of microphones or microphone orientations. Then, on the reproduction side, any decorrelation processing is not needed at all or only necessary to a very small extent; therefore, ultimately, due to reduced or unnecessary decorrelation processing, the reproducer achieves better reproduction quality on the reproduction side. This trade-off between bit rate and quality is preferably achieved via a function of the bitstream element that indicates the amount of sound signal from each spatially extended sound source.

[0076] Figure 11 It shows the result of Figure 10 The description shown is a preferred embodiment of the description generated by the description generation device. The description includes, for example, a second spatial extension sound source 401 denoted as SESS2 and corresponding data, and another first spatial extension sound source denoted as SESS1 and data 301 to 322.

[0077] therefore, Figure 11 Detailed data for each spatially extended sound source associated with spatially extended sound signal 1 are shown. Figure 11In the example, for a spatially extended sound source generated in the generator, for example, microphone output data picked up from microphones placed at two different locations in the spatially extended sound source, there are two sound signals. The first sound signal is sound signal 1 indicated at 301 and the second sound signal is sound signal 2 indicated at 302, and both sound signals are preferably encoded via an audio encoder for bit-rate compression. Furthermore, item 311 represents a descriptive element indicating the number of sound signals for spatially extended sound source 1, for example by... Figure 10 The controller is 250.

[0078] As shown in block 331, geometric information about the spatially extended sound source is introduced. Item 301 indicates optional location information for the sound signals, preferably related to geometric information, such as, for the piano example, indicating "near the bass string" for sound signal 1 and "near the treble string" for sound signal 2, indicated at 302. Thus, item 302 represents location information. This location information is interpreted by anchoring information element 322 when reproducing the sound source. The geometric information may, for example, be a parametric or polygonal representation of the piano model, and this piano model will be, for example, different for a grand piano or a (small) piano. Reference numeral 341 additionally illustrates optional location information about the spatially extended sound source within space. As previously mentioned, when the user provides, such as Figure 9 When the dashed line indicates the position information pointing to the projector, this position information 341 is not required. However, even if the position information 341 is included in the bit stream, the user can still replace or modify the position information through user interaction.

[0079] Preferred embodiments of the invention are then discussed. These embodiments relate to rendering spatially extended sound sources in 6DoF VR / AR (Virtual Reality / Augmented Reality).

[0080] Preferred embodiments of the present invention pertain to methods, apparatus, or computer programs designed to enhance the reproduction of Spatial Extended Sound Sources (SESS). Specifically, embodiments of the methods or apparatus of the present invention take into account the time-varying relative position between the spatial extended sound source and the virtual listener's location. In other words, embodiments of the methods or apparatus of the present invention allow the width of the sound source to match the spatial extent of the represented sound object at any relative position to the listener. Therefore, embodiments of the methods or apparatus of the present invention are particularly suitable for 6-DOF virtual, mixed, and augmented reality applications where the spatial extended sound source complements conventionally employed point sources.

[0081] Embodiments of the method or apparatus of the present invention render a spatially extended sound source by using several peripheral point sources fed with (preferably significantly) decorrelated signals. Compared to other methods, the positions of these peripheral point sources depend on the listener's position relative to the spatially extended sound source. Figure 1A schematic block diagram of a spatially extended sound source renderer according to an embodiment of the method or apparatus of the present invention is depicted.

[0082] The key components of the block diagram are:

[0083] 1. Listener Position: This block provides the listener's instantaneous position, for example, as measured by a virtual reality tracking system. This block can be implemented as a detector for detection or as an interface 100 for receiving the listener's position.

[0084] 2. Location and geometry of spatially extended sound sources: This block provides the location and geometry data of the spatially extended sound sources to be rendered, for example, as part of a virtual reality scene representation.

[0085] 3. Projection and Convex Hull Calculation: This block or projector 120 calculates the convex hull of the spatially extended sound source geometry and then projects it toward the listener's position (e.g., the "image plane," see below). Alternatively, the same functionality can be achieved by first projecting the geometry toward the listener's position and then calculating its convex hull.

[0086] 4. Location of peripheral point sources: This block 140 calculates the locations of the peripheral point sources used based on the convex hull projection data calculated from the previous block. In this calculation, it can also consider the listener's location, thus taking into account the listener's proximity / distance (see below). The output is the locations of n peripheral point sources.

[0087] 5. Renderer Core: Renderer core 162 auralizes n peripheral point sources by positioning them at specific target locations. This can be, for example, a binaural renderer using a head-related transfer function or a renderer for speaker reproduction (e.g., vector-based amplitude translation). The renderer core generates l speaker or headphone output signals from k input audio fundamental signals (e.g., decorrelated signals from instrument recordings) and m ≥ (nk) additional decorrelated audio signals.

[0088] 6. Source Fundamental Signals: This block 164 is the input of k fundamental audio signals, which are (fully) decorrelated with each other and represent the sound source to be rendered (e.g., a mono - k = 1 - or stereo - k = 2 - a recording of an instrument). The k fundamental audio signals are, for example, taken from the bitstream received from the decoder-side generator (see, for example...). Figure 11 Elements 301, 302), or may be provided from an external source at the reproduction site. The mapping of the base audio signal to the location of the peripheral sound source, or the generation or waveform of the peripheral sound source, may be influenced by the location information along with anchoring information, exemplarily indicating the anchoring of the user or listener or the object.

[0089] 7. Decorcorrelator: This optional block 166 generates additional decorrelated audio signals as needed to render n peripheral point sources.

[0090] 8. Signal Output: The renderer provides one output signal for speaker (e.g., n=5.1) or binaural (typically n=2) rendering.

[0091] Figure 1 An overview of a block diagram illustrating an embodiment of the method or apparatus of the present invention is shown. Dashed lines represent the transmission of metadata such as geometry and location. Solid lines represent the transmission of audio, where k, l, and m represent the number of audio channels. Renderer core 162 may receive k+m audio signals and n (<=k+m) location data. Renderer core 162, block 164, and block 166 together form an embodiment of a general-purpose renderer 160. The renderer additionally receives anchoring information for interpreting geometry information, as well as positioning information, particularly in the case of describing several channel signals of a spatially extended sound source.

[0092] The location of the peripheral point source depends on the geometry of the spatially extended sound source, particularly its spatial extent and the listener's relative position to the spatially extended sound source. Specifically, the peripheral point source can lie on the projection of the convex hull of the spatially extended sound source onto a projection plane. The projection plane can be the image plane, i.e., a plane perpendicular to the line of sight from the listener to the spatially extended sound source, or a sphere surrounding the listener's head. The projection plane is located at an arbitrarily small distance from the center of the listener's head. Alternatively, the projected convex hull of the spatially extended sound source can be calculated from the azimuth and elevation angles, which are subsets of spherical coordinates relative to the listener's head perspective. In the illustrative example below, the projection plane is preferred due to its more intuitive nature. In the implementation of the projected convex hull calculation, angular representation is preferred because it is simpler to formalize and has lower computational complexity. Note that the projection of the spatially extended sound source's convex hull is the same as the convex hull projected onto the spatially extended sound source's geometry; that is, the convex hull calculation and the projection onto the image plane can be used in either order.

[0093] The locations of peripheral point sources can be distributed on the convex hull projection of the spatially extended sound source in various ways, including:

[0094] They can be uniformly disturbed around the projection of the casing.

[0095] They can be distributed at the extreme points of the cladding projection.

[0096] They can be located at the horizontal and / or vertical extreme points of the cladding projection (see the figure in the Actual Examples section).

[0097] In addition to peripheral point sources, other auxiliary point sources can be used to generate enhanced acoustic fill, but this increases computational complexity. Furthermore, the projected convex hull can be corrected before locating the peripheral point sources. For example, the projected convex hull can shrink towards its centroid. This shrinking of the projected convex hull can result in an additional spatial expansion of the individual peripheral point sources introduced by the rendering method. Hull correction can further differentiate between horizontal and vertical scaling.

[0098] When the listener's position relative to the spatially extended sound source changes, the projection of the spatially extended sound source onto the projection plane also changes. Conversely, the positions of the peripheral point sources also change accordingly. The positions of the peripheral point sources should preferably be chosen such that they change smoothly to accommodate continuous movement of both the spatially extended sound source and the listener. Furthermore, when the geometry of the spatially extended sound source changes, the projection convex hull also changes. This includes rotation of the spatially extended sound source geometry in 3D space, which alters the projection convex hull. The rotation of the geometry is equal to the angular displacement of the listener's position relative to the spatially extended sound source, and is, for example, inclusively referred to as the relative position of the listener and the spatially extended sound source. For example, the listener's circular motion around a spherical spatially extended sound source is represented by rotating the peripheral point sources around the center of gravity. Similarly, the rotation of the spatially extended sound source for a stationary listener results in the same change in the positions of the peripheral point sources.

[0099] The spatial extent generated by embodiments of the method or apparatus of the present invention inherently and accurately reproduces any distance between the spatially extended sound source and the listener. Naturally, the angle between peripheral point sound sources increases as the user moves closer to the spatially extended sound source, because it is suitable for modeling physical reality.

[0100] Although the angular position of the peripheral point sources is uniquely determined by their position on the projected convex hull of the projection plane, the distance to the peripheral point sources can be further selected in various ways, including...

[0101] • All peripheral point sources have the same distance, equal to the distance of the entire spatially extended sound source, for example, defined by the center of gravity of the spatially extended sound source relative to the listener's head.

[0102] • The distance to each peripheral point source is determined by back-projecting its position on the projected convex hull onto the geometry of the spatially extended sound source, such that the peripheral point sources are projected onto the same point on the projection plane. The back-projection from the projected convex hull to the peripheral point sources of the spatially extended sound source may not always be uniquely determined, therefore additional projection rules must be applied (see the Practical Examples section).

[0103] • If the rendering of peripheral point sources does not require distance attributes, but only the relative angular positions of azimuth and elevation angles, then it may be impossible to determine the distance of peripheral point sources at all.

[0104] To specify the geometry / convex hull of a spatially extended sound source, approximations are used (and may be passed to the renderer or renderer kernel), including simplified 1D shapes such as lines and curves; 2D shapes such as ellipses, rectangles, and polygons; or 3D shapes such as ellipsoids, cuboids, and polyhedra. The geometry or corresponding approximate shape of a spatially extended sound source can be described in various ways, including:

[0105] • Parametric description, which formalizes geometry through mathematical expressions that accept additional parameters. For example, an ellipsoid shape in 3D can be described using implicit functions in a Cartesian coordinate system, with the additional parameters being the extensions of the principal axes in all three directions. Other parameters may include 3D rotations and deformation functions of the ellipsoid's surface.

[0106] • Polygon description, which is a collection of primitive geometric shapes, such as lines, triangles, squares, tetrahedrons, and cuboids.

[0107] Primate polygons and polyhedra can be connected to larger and more complex geometries.

[0108] Peripheral point source signals are derived from the fundamental signal of spatially extended sound sources. The fundamental signal can be acquired in various ways, such as: 1) recording natural sound sources at the location and orientation of a single or multiple microphones (e.g., recording the sound of a piano as seen in a practical example); 2) synthesizing artificial sound sources (e.g., sound synthesis with different parameters); 3) combining any audio signals (e.g., various mechanical sounds from a car, such as the engine, tires, doors, etc.). Furthermore, additional peripheral point source signals can be artificially generated from the fundamental signal using multiple decorrelation filters (see previous sections).

[0109] In some application scenarios, the focus is on compact and interoperable storage / transmission of 6DoF VR / AR content. In this case, the entire chain involves three steps:

[0110] 1. Create / encode the required spatial expansion sound source into a description, such as a bitstream.

[0111] 2. Transmission / Storage of the Generated Bitstream. According to the invention, among other elements, the bitstream also includes a description of the spatially extended geometry (parameters or polygons) of the sound source and the associated source fundamental signal, such as a mono or stereo piano recording. The waveform can be compressed using perceptual audio coding algorithms such as mp3 or MPEG-2 / 4 Advanced Audio Coding (AAC) (see...). Figure 10 (Item 260 in the middle).

[0112] 3. As mentioned above, spatially extended sound sources are decoded / rendered based on the transmitted bitstream.

[0113] In addition to the core methods described above, there are several other options for further processing:

[0114] Option 1 – Dynamic selection of peripheral point source number and location

[0115] The number of peripheral point sources can vary depending on the distance from the listener to the spatially extended sound source. For example, when the spatially extended sound source is far from the listener, the angle (aperture) of the projection convex hull becomes smaller, allowing for the advantageous selection of fewer peripheral point sources, thus saving computational and storage complexity. In extreme cases, all peripheral point sources are reduced to a single remaining point source. Appropriate downmixing techniques can be applied to ensure that interference between the fundamental and derived signals does not degrade the audio quality of the resulting peripheral point source signal. Similar techniques can be applied to spatially extended sound sources at close distances from the listener's position if the geometry of the spatially extended sound source is irregular according to the listener's relative viewpoint height. For example, the geometry of a spatially extended sound source as a finite-length line may degenerate into a single point on the projection plane. Generally, if the angular range of peripheral point sources on the projection convex hull is small, the spatially extended sound source can be represented by fewer peripheral point sources. In extreme cases, all peripheral point sources are reduced to a single remaining point source.

[0116] Option 2 – Propagation Compensation

[0117] Since each peripheral point source also exhibits spatial propagation projected outwards from the convex hull, the perceived auditory image width of the rendered spatially extended sound source is slightly larger than the convex hull used for rendering. To align it with the desired target geometry, there are two possibilities:

[0118] 1. Compensation during creation: Consider additional propagation during the rendering process during content creation. Specifically, choose a slightly smaller spatial expansion geometry for the sound source during content creation so that the actual rendered size matches the requirements. This can be checked by monitoring the effect of the renderer or renderer core in the creation environment (e.g., a production studio). In this case, the transported description or bitstream and the renderer or renderer core use a reduced target geometry compared to the target size.

[0119] 2. Compensation during rendering: Spatial extension sound source renderers or renderer kernels can compensate for the additional perceptual extension during the rendering process by being aware of it. As a simple example, the geometry used for rendering could be...

[0120] o Reduce the constant factor a < 1.0 (e.g., a = 0.9), or

[0121] o Reduce the constant angular subtendance alpha = 5 degrees

[0122] Before being applied to the placement of peripheral point sources. In this case, the transmitted bitstream contains the final target size of the spatially expanded sound source geometry.

[0123] Furthermore, combinations of these methods are feasible.

[0124] Option 3 – Generation of peripheral point source waveforms

[0125] Furthermore, by taking into account the user's position relative to the spatially extended sound source, the actual signal for feeding the peripheral point source can be generated from the recorded audio signal in order to model spatially extended sound sources with geometrically related sound contributions, such as a piano with bass emanating from the left, and vice versa.

[0126] Example: The sound of an upright piano is characterized by its acoustic behavior. This is modeled by (at least) two fundamental audio signals, one near the lower end of the piano keyboard ("bass") and one near the upper end ("treble"). These fundamental signals can be acquired via an appropriate microphone when recording the piano sound and transmitted to a 6DoF renderer or renderer core, ensuring sufficient decorrelation between them.

[0127] Then, by considering the user's position relative to the spatially extended sound source, the peripheral point source signal is derived from these basic signals:

[0128] When the user faces the piano from the side (front of the keyboard), the two peripheral point sources are relatively far apart, located near the left and right ends of the piano keyboard, respectively. In this situation, the fundamental signal for bass can be directly fed into the left peripheral point source, and the fundamental signal for treble can be directly used to drive the right peripheral point source.

[0129] When the listener moves approximately 90 degrees to the right around the piano, the two peripheral point sources are translated to positions very close to each other because the projection of the piano's volumetric model (e.g., an ellipse) is small when viewed from the side. If the peripheral point source signals were to continue driving them directly with the fundamental signals, one peripheral point source would primarily contain treble, while the other would primarily carry bass. Since this is physically undesirable, the rendering can be improved by rotating the two fundamental signals to form the peripheral point source signals through a Givens rotation of the same angle as the user's movement relative to the piano's center of gravity. This way, both signals contain signals with similar spectral content while still being decorrelated (assuming the fundamental signals have already been decorrelated).

[0130] Option 4 – Post-processing of spatially extended sound sources in rendering

[0131] The actual signal can be pre-processed or post-processed to account for positional and directional effects, such as the directional patterns of spatially extended sound sources. In other words, the entire sound emanating from a spatially extended sound source, as previously described, can be modified to exhibit, for example, a direction-dependent sound radiation pattern. In the case of a piano signal, this might mean that radiation toward the back of the piano has fewer high-frequency components than radiation toward the front. Furthermore, the pre-processing and post-processing of peripheral point source signals can be adjusted individually for each peripheral point source. For example, different directional patterns can be selected for each peripheral point source. In a given example representing a spatially extended sound source of a piano, the directional patterns of the bass and treble ranges might be similar to those described above, but additional signals such as pedal noises have more omnidirectional directional patterns.

[0132] Subsequently, several advantages of the preferred embodiments are summarized.

[0133] Compared to using point sources to completely fill the space to expand the interior of the sound source (e.g., as used in Advanced AudioBIFS), the computational complexity is lower.

[0134] The possibility of destructive interference between point source signals is relatively small.

[0135] • Compact size of bitstream information (geometric approximation, one or more waveforms)

[0136] • Allows the use of traditional recordings made for music consumption (e.g., stereo recordings of pianos) for VR / AR rendering.

[0137] Then, various practical implementation examples are given:

[0138] ·Spherical spatial extended sound source

[0139] Elliptic spatial extension sound source

[0140] Linear spatial extension sound source

[0141] · Rectangular space extended sound source

[0142] • Distance-related peripheral point sources

[0143] Piano-shaped spatial sound source

[0144] As described in the above embodiments of the method or apparatus of the present invention, various methods can be applied to determine the location of peripheral point sources. The following practical examples demonstrate some isolated methods in specific cases. In a complete implementation of embodiments of the method or apparatus of the present invention, various methods can be appropriately combined, taking into account computational complexity, application purpose, audio quality, and ease of implementation.

[0145] The spatially extended sound source geometry is represented by a green surface mesh. Note that the mesh visualization does not imply that the spatially extended sound source geometry is described using a polygonal method, as it may actually be generated from a parametric specification. The listener's location is represented by a blue triangle. In the following example, the image plane is chosen as the projection plane and depicted as a transparent gray plane representing a finite subset of the projection plane. The projected geometry of the spatially extended sound source on the projection plane is depicted using the same green surface mesh. The peripheral point sources on the projection convex hull are depicted as red crosses on the projection plane. The back-projected peripheral point sources on the spatially extended sound source geometry are depicted as red dots. Corresponding peripheral point sources on the projection convex hull and back-projected peripheral point sources on the spatially extended sound source geometry are connected by red lines to aid in identifying visual correspondences. The positions of all involved objects are depicted in Cartesian coordinates in meters. The choice of the depicted coordinate system does not imply that the calculations involved were performed in Cartesian coordinates.

[0146] Figure 2 The first example considers a spherical spatially extended sound source. The spherical spatially extended sound source has a fixed size and fixed position relative to the listener. Three, five, and eight different sets of peripheral point sources are selected on the projected convex hull. All three sets of peripheral point sources are selected at uniform distances along the convex hull curve. The offset positions of the peripheral point sources on the convex hull curve are intentionally chosen to well represent the horizontal range of the spatially extended sound source geometry.

[0147] Figure 2 A spherical spatially extended sound source with different numbers (i.e., 3 (top), 5 (middle) and 8 (bottom)) of peripheral point sources uniformly distributed on a convex hull is shown.

[0148] Figure 3 The next example considers an elliptical spatially extended sound source. An elliptical spatially extended sound source has a fixed shape, position, and rotation in 3D space. Four peripheral point sources are selected in this example. Three different methods for determining the positions of the peripheral point sources are illustrated:

[0149] a) Two peripheral point sources are placed at two horizontal extrema, and two peripheral point sources are placed at two vertical extrema. However, extrema location is straightforward and often appropriate. This example demonstrates that this method may produce peripheral point source locations that are relatively close to each other.

[0150] (b) All four peripheral point sources are uniformly distributed on the projected convex hull. The offset of the peripheral point source positions is chosen such that the position of the topmost peripheral point source coincides with the position of the topmost peripheral point source in (a). It can be seen that the choice of the peripheral point source position offset has a considerable impact on the representation of the geometry through the peripheral point sources.

[0151] c) All four peripheral point sources are uniformly distributed on the contracted convex hull. The offset positions of the peripheral point sources are equal to the offset positions selected in b). The contraction operation of the projected convex hull is performed towards the centroid of the projected convex hull and has a stretching factor independent of direction.

[0152] Figure 3 An ellipsoidal spatially extended sound source with four peripheral point sources is shown under three different methods for determining the location of peripheral point sources: a / top) horizontal and vertical extrema, b / middle) points uniformly distributed on the convex hull, and c / bottom) points uniformly distributed on the contracting convex hull.

[0153] Figure 4 The next example considers a linear spatially extended sound source. While the previous examples considered volumetric spatially extended sound source geometries, this example shows that spatially extended sound source geometries can be well chosen as single-dimensional objects within 3D space. Subfigure a) depicts two peripheral point sources placed at the extrema of a finite linear spatially extended sound source geometry. b) Two peripheral point sources are placed at the extrema of a finite linear spatially extended sound source geometry, with an additional point source placed in the middle of the line. As described in embodiments of the method or apparatus of the present invention, placing additional point sources within the spatially extended sound source geometry can help fill large gaps in large spatially extended sound source geometries. c) Considers the same linear spatially extended sound source geometry as in a) and b), but with a changed relative angle toward the listener, resulting in a relatively small projected length of the linear geometry. As described in embodiments of the method or apparatus of the present invention above, the reduced size of the projected convex hull can be represented by a reduced number of peripheral point sources, in this particular example by a single peripheral point source located at the center of the linear geometry.

[0154] Figure 4 The diagram illustrates a spatially extended sound source that uses three different methods to allocate the location of peripheral point sources: a / top) two extrema on the projected convex hull; b / middle) two extrema on the projected convex hull with an additional point source at the center of the line; c / bottom) a peripheral point source at the center of the convex surface, because the projected convex hull of the rotating line is too small to allow more than one peripheral point source.

[0155] Figure 5 The next example considers a cuboid spatially extended sound source. The cuboid spatially extended sound source has a fixed size and a fixed position, but the listener's relative position changes. Subfigures a) and b) illustrate different methods of placing four peripheral point sources on the projected convex hull. The positions of the back-projected peripheral point sources are uniquely determined by the choice on the projected convex hull. c) depicts four peripheral point sources whose back-projected positions are not well separated. Instead, the distance between the peripheral point source positions is chosen to be equal to the centroid distance of the spatially extended sound source geometry.

[0156] Figure 5 A cuboid spatially extended sound source is shown, which has three different peripheral point source distribution methods: a / bottom) two peripheral point sources on the horizontal axis and two peripheral point sources on the vertical axis; b / middle) two peripheral point sources on the horizontal extremum of the projected convex hull and two peripheral point sources on the vertical extremum of the projected convex hull; c / bottom) the distance between the back-projected peripheral point sources is chosen to be equal to the centroid distance of the spatially extended sound source geometry.

[0157] Figure 6 The next example considers a spherical spatially extended sound source of fixed size and shape, but located at three different distances relative to the listener's position. The peripheral point sources are uniformly distributed along the convex hull curve. The number of peripheral point sources is dynamically determined by the length of the convex hull curve and the minimum distance between the possible locations of the peripheral point sources. a) The spherical spatially extended sound sources are very close, such that four peripheral point sources are selected on the projected convex hull. b) The spherical spatially extended sound sources are at a moderate distance, such that three peripheral point sources are selected on the projected convex hull. a) The spherical spatially extended sound sources are far away, such that only two peripheral point sources are selected on the projected convex hull. As described in the embodiments of the method or apparatus of the present invention above, the number of peripheral point sources can also be determined based on a range expressed in spherical angular coordinates.

[0158] Figure 6 The diagram shows spherical spatially extended sound sources of equal size but different distances: a / top) four peripheral point sources evenly distributed on the projected convex hull at close distance; b / middle) the middle distance of three peripheral point sources evenly distributed on the projected convex hull; c / bottom) two peripheral point sources evenly distributed on the projected convex hull at far distance.

[0159] Figure 7 and 8 The last example considers a spatially expanded sound source in the shape of a piano placed in a virtual world. The user wears a head-mounted display (HMD) and headphones. A virtual reality scene is presented to the user, including an open text canvas and a 3D upright piano model standing on the floor within a freely movable area (see [link to example]). Figure 7 An open-world canvas is a spherical, static image projected onto a sphere surrounding the user. In this particular case, the open-world canvas depicts a blue sky and white clouds. The user can move around and view and listen to the piano from various angles. In this scene, the piano is rendered as a single-point sound source placed at the center of gravity, or as a spatially extended sound source with three peripheral point sound sources on the projected convex hull (see...). Figure 8 Rendering experiments show that the peripheral point source rendering method has a significantly superior sense of realism compared to rendering with a single point source.

[0160] To simplify the calculation of the location of the peripheral point sources, the piano geometry is abstracted into an elliptical shape with similar dimensions, see... Figure 7Furthermore, two alternative point sources are placed at the left and right extreme points along the equator, while the third alternative point remains at the North Pole; see [link to relevant documentation]. Figure 8 This arrangement ensures that an appropriate horizontal source width is obtained from all angles, while significantly reducing computational costs.

[0161] Figure 7 A piano-shaped spatially extended sound source (shown in green) is shown, with an approximate parametric ellipsoidal shape (shown as a red grid).

[0162] Figure 8 A piano-shaped spatially extended sound source is shown, with three peripheral point sources distributed at the vertical extrema and the vertical top position of the projected convex hull. Note that for better visualization, the peripheral point sources are placed on the stretched projected convex hull.

[0163] Specific features of embodiments of the present invention are then provided. The features of the provided embodiments are as follows:

[0164] • To fill the perceived acoustic space of a spatially extended sound source, it is best not to fill its entire interior with the associated point source (peripheral point source), but only with its periphery facing the listener (e.g., "the projection of the spatially extended sound source's convex hull toward the listener"). Specifically, this means that the location of the peripheral point source is not attached to the geometry of the spatially extended sound source, but is dynamically calculated, taking into account the relative position of the spatially extended sound source with respect to the listener's position.

[0165] Dynamic calculation of peripheral point sources (number and location)

[0166] • Use an approximation of the spatially extended sound source shape (for scenarios using compressed representation: transmitted as part of a bitstream).

[0167] The application of this technology can be part of the 6DoF audio VR / AR standard. In this case, there is a classic encoder / bitstream / decoder (+renderer) scenario:

[0168] • In the encoder, the shape of the spatially extended sound source is encoded along with the "basic" waveform of the spatially extended sound source as side information, which may be

[0169] o Mono signal, or

[0170] o Stereo signal (preferably fully decorrelated), or

[0171] o represents more recorded signals from the spatially extended sound source (preferably also sufficient for decorrelation). These waveforms can be low-bit-rate encoded.

[0172] • In the decoder / renderer, the shape of the spatially extended sound source and the corresponding waveform are retrieved from the bitstream and used to render the spatially extended sound source, as described above.

[0173] Depending on the embodiment used and as an alternative to the described embodiment, it should be noted that the interface can be implemented as an actual tracker or detector for detecting the listener's location. However, the listening location is typically received from an external tracker device and fed into the playback device via the interface. However, the interface can represent only data input from the output data of an external tracker, or it can represent the tracker itself.

[0174] In addition, as outlined, additional auxiliary audio sources may be required between peripheral sound sources.

[0175] Furthermore, it has been found that left / right peripheral sound sources and optional horizontal (relative to the listener) spaced auxiliary sound sources are more important for the perceived impression than vertically spaced peripheral sound sources, i.e., the top peripheral sound source and the bottom of the spatially extended sound source. For example, when resources are scarce, it is preferable to use at least horizontally spaced peripheral (and optional auxiliary) sound sources, while vertically spaced peripheral sound sources can be omitted to save processing resources.

[0176] Furthermore, as outlined, the bitstream generator can be implemented to generate a bitstream with only one sound signal for a spatially extended sound source, and the remaining sound signals are generated on the decoder or reproduction side through decorrelation. When only a single signal exists, and when the entire space is to be uniformly filled with this signal, any location information is unnecessary. However, in this case, at least additional information about the geometry of the spatially extended sound source calculated by a geometry information calculator is useful, such as in... Figure 10 The one shown at position 220.

[0177] Further embodiments are discussed below:

[0178] ObjectSourceInputLayout

[0179] An ObjectSource with spatial extent can have a multi-channel AudioStream, allowing the renderer to render the ObjectSource in a more realistic way than a mono AudioStream. This is useful, for example, when rendering diffuse audio sources such as fountains, waterfalls, rivers, and crashing waves.

[0180] An object source with a range is always perceived by the listener within the elevation-azimuth region from the listener's perspective. This region is determined by the relative position of the object source to the listener and the range of the object source, all in an acoustically perceptible sense. This is in... Figure 12aThe example illustrates an object source with a cylindrical extent, where the ObjectSource is located in the listener's right anterior hemisphere. The intersection of the plane orthogonal to the observation vector at the center of the elevation-azimuth region and the elevation-azimuth region defines a rectangle. This rectangle represents the horizontal and vertical extent of the ObjectSource as perceived acoustically by the listener from their position. As the listener moves around, closer to, or away from the ObjectSource, this rectangle will translate, rotate, and resize in the world space coordinate system. Figure 12b This illustrates the case when the cylindrical ObjectSource is located in the listener's left front hemisphere. However, in the xy coordinate system with the center of these receptive field rectangles as the origin, these rectangles are always centered at the source's (0, 0) point.

[0181] The InputLayout child nodes described by ObjectSource consist of alignment markers and strings, including space-separated positioning mnemonics:

[0182]

[0183] The alignment attribute defines how the waveforms (channels) of the associated audio stream are positioned / anchored relative to the source. The positioning attribute is a string containing space-separated mnemonic labels, where a mnemonic label must be provided for each waveform. For example... Figure 13 As shown, it supports nine relative position mnemonics for referencing the previously described xy coordinate system.

[0184] Therefore, the supported channel specifications are the nine relative positions in this xy coordinate system, such as... Figure 13 As described in [the text].

[0185] In addition, ObjectSourceInputLayout can be a string containing space-separated position mnemonics. Figure 13 Nine possible locations are listed.

[0186] Alternatively, relative channel position can be used to indicate the use of waveforms for rendering ObjectSources with absolute 3D coordinate space dimensions (e.g., the sound of a grand piano with one channel primarily containing lower notes and another channel primarily containing higher notes). In this case, the label is applied to a rectangle on a plane perpendicular to the frontal direction of the ObjectSource when facing the ObjectSource's position (and the ObjectSource's "orientation" attribute must exist). This is indicated by the initial "A" mnemonic in the ObjectSourceInputLayout string.

[0187] Example:

[0188] inputLayout="LR"

[0189] This indicates that two waveforms are used to render the horizontal width of the source.

[0190] inputLayout="BT"

[0191] This indicates that two waveforms are used to render the vertical width of the source.

[0192] inputLayout="BL TL TR BR"

[0193] This indicates that four waveforms are used to render the horizontal and vertical width of the source.

[0194] inputLayout="AL R"

[0195] This indicates that two waveforms are used to render the horizontal width of a source with absolute left and right distribution.

[0196] In other words, the above embodiments involve an objectSource with two related waveforms (from a stereo recording, where the left channel ideally carries more bass and the right channel carries higher frequencies).

[0197] To accommodate this, refer to ObjectSourceInputLayout. Currently defined labels (such as L, C, R) are always defined for a projection plane perpendicular to the view direction. Therefore, this is unsuitable for static objects, such as (grand) pianos.

[0198] Therefore, according to the embodiment, the additional bitstream element is implemented as, for example, a small mark or additional mark (or letter) in the EIF specification, which allows the use of *absolute* tag anchors added to the current EIF specification. This allows for the resolution of the grand piano case and the description of the expected waveform using *fixed (absolute) position and orientation relative to the instrument*—as well as the use of size attributes. The orientation of the object will serve as a reference for the new projection plane. The additional bitstream element can also differ from the additional letter, provided the decoder is configured to parse the element for correct rendering.

[0199] In the example above, the letter "A" represents a marker, bitstream element, or anchoring information. This information is used by the renderer on the playback side in response to specific received information to render at least two sound sources relative to the fixed position and / or orientation of the spatially extended sound sources. Preferably, when the information is not in the encoded signal syntax, rendering occurs consistent with the transmitted information (e.g., left or right) but is related to the user's position. However, when the information is present, rendering is performed relative to the sound source position, not the user's or listener's position. In other words, when the information is present, for example, the piano is rendered as is regardless of whether the user is standing in front of or behind the piano. The first channel always comes from the bass side of the piano, and the second channel always comes from the treble side. However, when this information is absent, the channel positions are correct only when the user is standing in front of the piano, and incorrect when the user is standing behind the piano.

[0200] In other words, the embodiments involve anchoring the label relative to the listener's viewing direction (e.g., Figure 12a and 12b As described in the initial example (with the attribute alignment="user"), the signal channel relative position label can be used to indicate that an ObjectSource of a certain size is rendered using a waveform, anchoring it to an object in the scene (with the attribute alignment="object"). The example is the sound of a piano, where one signal channel primarily contains lower notes and the other primarily contains higher notes. In this case, the position label is applied to a rectangle on a plane passing through the object's position (center), which is perpendicular to the ObjectSource's orientation when viewing the front of the ObjectSource (the ObjectSource's "orientation" property must exist). During rendering, the position indicated by the label is then projected onto the user's view plane (a plane orthogonal to the view vector), as shown... Figures 14a to 14c As shown. This could also mean placing the sound sources (which may have a range) "behind" each other (see when viewing the piano from the side). Figure 14c Even swap them (when looking at the piano from behind).

[0201] Therefore, further examples are as follows:

[0202] <InputLayout alignment=”user”positioning=”L R” / >

[0203] This indicates that two waveforms are used to render the horizontal width of the source.

[0204] <InputLayout alignment=”user”positioning=”B T” / >

[0205] This indicates that two waveforms are used to render the vertical width of the source.

[0206] <InputLayout alignment=”user”positioning=”BL TL TR BR” / >

[0207] This indicates that four waveforms are used to render the horizontal and vertical width of the source.

[0208] <ObjectSource id=”src:piano”

[0209] position="1.2 1.0 -0.3"

[0210] orientation="38 0 0"

[0211] extent="geo:piano_extent"

[0212] signal = "signal:piano"

[0213] <InputLayout alignment=”object”positioning=”L R” / >

[0214]

[0215] This indicates that two waveforms are used to render the left and right sides of the piano object.

[0216] exist Figure 14a The example shown gives the following: Figure 13 The table shows a "map" of the location of a certain channel. In this example, the multichannel signal is a two-channel signal with a left channel or first channel for the left part, which has more of the sound recorded or synthesized from the bass or left part of the piano, and a right channel or second channel with more of the sound recorded or synthesized from the higher notes of the right part of the piano.

[0217] exist Figure 14b In the embodiments, Figure 1 or Figure 9 The sound location calculator 140 calculates the location of peripheral sound sources based on the listening position, i.e., the observer's projection plane, such as the four corners of a piano. Figure 14b As shown. Alternatively, the sound position calculator only calculates the left position, for example, the position at the middle left of the piano rectangle and the right position at the middle right of the piano rectangle.

[0218] For rendering, renderer 160 uses the first channel for [specific purposes] based on the anchoring mode and positioning information. Figure 14bA single peripheral sound source on the left side, or for the upper and lower positions on the left side. Furthermore, based on the anchoring mode and positioning information, renderer 160 uses the second channel for... Figure 14b A single peripheral sound source on the right side, or for upper and lower positions on the right. For example, this selection can be made by... Figure 1 Block 164 of the renderer example is executed.

[0219] In different Figure 14b In the case where the observer is in relation to Figure 14b Positioned at the same angle and distance behind the piano, Figure 1 or Figure 9 The sound location calculator 140 calculates the location of peripheral sound sources based on the listening position, i.e., the observer's projection plane, such as... Figure 14b The relevant four rear corners of the piano or only the left side position, such as the middle of the left side of the piano rectangle and the right side position of the middle of the right side of the piano rectangle.

[0220] Now, compared to the previous situation, renderer 160, based on the anchoring mode and positioning information, uses the first channel for... Figure 14b A single peripheral sound source on the right side or for the upper and lower positions on the right (compared to the left side outlined above). Furthermore, based on the anchoring pattern and positioning information, renderer 160 uses the second channel for... Figure 14b A single peripheral sound source on the left side of the center, or the upper and lower positions on the left side (compared to the right side as outlined above). For example, this selection can be made by... Figure 1 Block 164 of the renderer example is executed.

[0221] Specific details are as follows: Figure 14b As shown, the user is positioned to one side of the piano. In this embodiment, the waveforms used for all peripheral sound sources can be identical, and this waveform is calculated by adding the left or first channel to the right or second channel. This addition can include a weighted addition, such that... Figure 14c In one embodiment, where the user is more likely to stand on the left side of the piano, the weighting factor for the left channel is greater than that for the right channel because the right channel is slightly lower due to the greater distance from the user compared to the left channel, and also due to attenuation caused by the object itself, i.e., the piano. For example, this could be achieved by... Figure 1 Block 164 of the renderer example performs this calculation from the transmitted audio channels.

[0222] If the user is on the right side of the piano, the situation is similar to that outlined above, but the weighting factors are swapped in the case of weighted addition. This calculation and the determination of the weighting factors can, for example, be derived from... Figure 1 Block 164 of the renderer example is executed.

[0223] Note that the piano is just an example. An arbitrary spatially extended sound source can be represented as using... Figure 13 Exemplary agreements, such as Figures 14a to 14c The rectangle or any other shape such as an elliptic or block is schematically shown in the diagram.

[0224] For comparative purposes, which are not part of this invention, for the user mode and the above examples, consider the waveform of the audio channel mapped to the peripheral sound source. Figure 14b In the example where the listener is behind the object, the mapping will be the same as in the object pattern because the left and right sides will not change when the observer is in front of the object. However, in the example where the listener is behind the object, the situation will be reversed, meaning the user pattern will differ from the object pattern. Figure 14c The same applies to the implementation. In user mode instead of object mode, no (e.g., weighted) addition occurs, but the left channel will be used for the left peripheral sound source position, and the right channel will be used for the right peripheral sound source position.

[0225] When the listener is placed diagonally opposite the piano, such as in Figure 14b and Figure 14c The waveform of the sound source, positioned "between" the object, can be calculated in object-anchored mode by some mixture of the left and right channels. The waveform of the left peripheral sound source will be the first or left channel, weighted by a larger weight added to the right or second channel, which is weighted by a lower weight. The weights can be adjusted based on the observer's angle relative to the object, thus representing the waveform from... Figure 14b The situation to Figure 14c The weights change continuously in cases where both channels typically have the same weight. This calculation, and the determination of the weights, can be, for example, by... Figure 1 Block 164 of the renderer example is executed.

[0226] Furthermore, when the number of channels is less than the number of sound sources, a decorrelation device such as... Figure 9 Block 166, as shown, is used to generate additional sound sources. Figure 14c In some embodiments, decorrelation can be performed on the summation waveform derived from the sum of the left and right sides. For example, for Figure 14c Four peripheral sound sources at the four corners of the projection plane are used to obtain slightly different waveforms.

[0227] In such an embodiment, the spatially extended sound source is associated with a multi-channel signal having a first channel and a second channel, the first channel being associated with a first portion of the spatially extended object and the second channel being associated with a second portion of the spatially extended object, wherein the first portion is different from the second portion, and wherein specific information (320) indicates that at least two sound sources are rendered relative to a fixed position and / or orientation of the spatially extended sound source. The renderer (160) is then configured to determine different sound signals at different positions using mappings of the first and second channels to different positions or by adding the first and second channels to obtain different sound signals at different positions based on the listener's position and the first and second portions of the spatially extended sound source.

[0228] In such an embodiment, the first part is the left part of the spatially extended sound source, and the second part is the right part.

[0229] When the listener is positioned in front of the spatially expanded sound source ( Figure 14b The renderer is configured to use the first channel for the sound source location on the user's left and the second channel for the sound source location on the user's right.

[0230] Alternatively or additionally, when the listener is positioned behind the spatially extended sound source (with Figure 14b (In contrast), the renderer is configured to use the second channel for the sound source location on the user's left and the first channel for the location on the user's right.

[0231] Alternatively or additionally, when the listener is positioned on one side of the spatially extended sound source ( Figure 14c The renderer is configured to use the sum of the first and second channels for the sound source location to the left of the user, and the sum of the first and second channels for the location to the right of the user.

[0232] Alternatively or additionally, when the listener's position is on one side of the spatially extended sound source, the renderer is configured to use a weighted sum of the first and second channels for the sound source position to the user's left, and a weighted sum of the first and second channels for the position to the user's right, wherein the weighting factor of the weighted sum is determined such that the weighting factor of the channel associated with a portion of the spatially extended sound source closer to the listener's position is greater than the weighting factor of the other channel associated with another portion of the spatially extended sound source farther from the listener's position. Figure 14b The weight of L is greater than the weight of R; and Figure 14b Relatively speaking, the weight of R is greater than the weight of L.

[0233] Alternatively or additionally, when the listener's position is tilted relative to the spatially extended sound source, the renderer is configured to use a first weighted sum of the first and second channels for the sound source position to the user's left, and a second weighted sum of the first and second channels for the position to the user's right, wherein the weighting factor of the weighted sum is determined such that the weighting factor of the channel associated with a portion of the spatially extended sound source closer to the sound source position is greater than the weighting factor of the other channel associated with another portion of the spatially extended sound source farther from the sound source position (position in...). Figure 14b and Figure 14c "Between"; for the left sound source in the projection, the weight of the left channel is greater than the weight of the right channel, and for the right sound source in the projection, the weight of the left channel is lower than the weight of the right channel.

[0234] It should be noted that all the alternatives or aspects discussed above, as well as all aspects defined by the independent claims in the following claims, can be used individually; that is, there are no other alternatives or objectives other than the contemplated alternatives, objectives, or independent claims. However, in other embodiments, two or more of the alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.

[0235] The creatively encoded sound field description can be stored on a digital storage medium or a non-transitory storage medium, or it can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.

[0236] Although some aspects have been described in the context of the apparatus, these aspects clearly also represent descriptions of the corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus.

[0237] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, storing electronically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to perform the corresponding methods.

[0238] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0239] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.

[0240] Other embodiments include a computer program stored on a machine-readable carrier or non-transitory storage medium for performing one of the methods described herein.

[0241] In other words, embodiments of the method of the present invention are therefore computer programs having program code that, when run on a computer, performs one of the methods described herein.

[0242] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) having a computer program recorded thereon for performing one of the methods described herein.

[0243] Therefore, a further embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.

[0244] Further embodiments include processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0245] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.

[0246] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions that form the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.

[0247] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent is to be limited only by the scope of the appended claims and not by the specific details presented in the description and explanation of the embodiments herein.

[0248] References

[0249] Alary, B., Politis, A., & V.(2017).Velvet Noise Decorrelator.

[0250] Baumgarte,F.,&Faller,C.(2003).Binaural Cue Coding-Part I:Psychoacoustic Fundamentals and Design Principles.Speech and AudioProcessing,IEEE Transactions on,11(6),S.509–519.

[0251] Blauert,J.(2001).Spatial hearing(3Ausg.).Cambridge;Mass:MIT Press.

[0252] Faller,C.,&Baumgarte,F.(2003).Binaural Cue Coding-Part II:Schemes andApplications.Speech and Audio Processing,IEEE Transactions on,11(6),S.520–531.

[0253] Kendall,G.S.(1995).The Decorrelation of Audio Signals and Its Impacton Spatial Imagery.Computer Music Journal,19(4),S.p 71-87.

[0254] Lauridsen,H.(1954).Experiments Concerning Different Kinds of Room-Acoustics Recording.Ingenioren,47.

[0255] T.,Santala,O.,&Pulkki,V.(2014).Synthesis of SpatiallyExtended Virtual Source with Time-Frequency Decomposition of MonoSignals.Journal of the Audio Engineering Society,62(7 / 8),S.467–484.

[0256] Potard,G.(2003).A study on sound source apparent shape and wideness.

[0257] Potard,G.,&Burnett,I.(2004).Decorrelation Techniques for theRendering of Apparent Sound Source Width in 3D Audio Displays.

[0258] Pulkki,V.(1997).Virtual Sound Source Positioning Using Vector BaseAmplitude Panning.Journal of the Audio Engineering Society,45(6),S.456–466.

[0259] Pulkki,V.(1999).Uniform spreading of amplitude panned virtualsources.

[0260] Pulkki,V.(2007).Spatial Sound Reproduction with Directional AudioCoding.J.Audio Eng.Soc,55(6),S.503–516.

[0261] Pulkki,V.,Laitinen,M.-V.,&Erkut,C.(2009).Efficient Spatial SoundSynthesis for Virtual Worlds.

[0262] Schlecht,S.J.,Alary,B., V.,&Habets,E.A.(2018).OptimizedVelvet-Noise Decorrelator.

[0263] Schmele,T.,&Sayin,U.(2018).Controlling the Apparent Source Size inAmbisonics Unisng Decorrelation Filters.

[0264] Schmidt,J.,& E.F.(2004).New and Advanced Features for AudioPresentation in the MPEG-4Standard.

[0265] Verron,C.,Aramaki,M.,Kronland-Martinet,R.,&Pallone,G.(2010).A 3-DImmersive Synthesizer for Environmental Sounds.Audio,Speech,and LanguageProcessing,IEEE Transactions on,title=A Backward-Compatible MultichannelAudio Codec,18(6),S.1550–1561.

[0266] Zotter,F.,&Frank,M.(2013).Efficient Phantom Source Widening.Archivesof Acoustics,38(1),S.27–37.

[0267] Zotter,F.,Frank,M.,Kronlachner,M.,&Choi,J.-W.(2014).Efficient PhantomSource Widening and Diffuseness in Ambisonics.

Claims

1. Apparatus for a reproduction of a spatially extended sound source having a defined position or location and a geometry in space, the apparatus comprising: an interface for receiving a listener position; a projector for calculating a projection of a two- or three-dimensional hull associated with the spatially extended sound source on a projection plane using the listener position, information about the geometry of the spatially extended sound source and information about the defined position or location of the spatially extended sound source; a sound position calculator for calculating different positions of at least two sound sources of the spatially extended sound source using the projection plane; and a renderer for rendering the at least two sound sources at the different positions of the at least two sound sources to obtain a reproduction of the spatially extended sound source having two or more output signals, wherein the renderer is configured to use different sound signals for the different positions of the at least two sound sources, wherein the different sound signals are associated with the spatially extended sound source, wherein the renderer is configured to render the at least two sound sources relative to a fixed position and / or location of the spatially extended sound source in response to a received specific information.

2. Apparatus according to claim 1, further comprising a detector configured to detect an instantaneous listener position in space using a tracking system, or wherein the interface is configured to use position data entered via the interface.

3. Apparatus according to claim 1, configured for receiving a scene description comprising information about the defined position or location and the defined geometry of the spatially extended sound source and at least one basic sound signal associated with the spatially extended sound source, wherein the apparatus further comprises a scene description parser for parsing the scene description to retrieve the information about the defined position or location, the information about the defined geometry and the at least one basic sound signal, or wherein for the spatially extended sound source, the scene description comprises at least two basic sound signals of at least two basic sound signals and position information for each basic sound signal relative to the information about the geometry of the spatially extended sound source, and wherein the sound position calculator is configured to use the position information of the at least two basic sound signals when calculating the different positions of the at least two sound sources using the projection plane.

4. Apparatus according to claim 1, wherein the projector is configured to calculate a hull of the spatially extended sound source using the information about the geometry of the spatially extended sound source and to project the hull in a direction towards the listener position to obtain the projection of the two- or three-dimensional hull on the projection plane using the listener position or location, or wherein the projector is configured to project the geometry of the spatially extended sound source defined by the information about the geometry of the spatially extended sound source in a direction towards the listener position and to calculate a hull of the projected geometry to obtain the projection of the two- or three-dimensional hull on the projection plane.

5. Apparatus according to claim 1, wherein the sound position calculator is configured to calculate the different positions of the at least two sound sources in space from the hull projection data and the listener position.

6. Apparatus according to claim 1, wherein the sound position calculator is configured to calculate different positions of the at least two sound sources such that the at least two sound sources are peripheral sound sources and are located on a projection plane, or wherein the sound position calculator is configured to calculate different positions of the at least two sound sources such that the at least two sound sources are peripheral sound sources and are located on a projection plane, or 7. The apparatus of claim 1, wherein the renderer is configured to render the at least two sound sources using a panning operation according to the different positions of the at least two sound sources to obtain loudspeaker signals for a predefined loudspeaker setup, or a binaural rendering operation using head-related transfer functions according to the different positions of the at least two sound sources to obtain headphone signals.

8. The apparatus of claim 1, wherein a first number of basic source signals is associated with a spatially extended sound source, the first number being one or larger than one, wherein the first number of basic source signals is associated with the same spatially extended sound source, wherein the sound position calculator determines a second number of sound sources for rendering the spatially extended sound source, the second number being larger than one, and wherein the renderer comprises one or more decorrelators for generating decorrelated signals from the first number of one or more basic source signals, wherein the second number is larger than the first number.

9. The apparatus of claim 1, wherein the interface is configured to receive a time-varying position of a listener in a space, wherein the projector is configured to calculate a time-varying projection in the space, wherein the sound position calculator is configured to calculate a time-varying number of sound sources or time-varying positions of the at least two sound sources in the space, and wherein the renderer is configured to render the time-varying number of sound sources or the at least two sound sources as different sound signals at the time-varying positions of the at least two sound sources in the space.

10. The apparatus of claim 1, wherein the interface is configured to receive a listener position in six degrees of freedom, and wherein the projector is configured to calculate a projection according to six degrees of freedom.

11. The apparatus of claim 1, wherein, the projector is configured to calculate the projection as an image plane, such as a plane perpendicular to a line of sight of the listener, the image plane being the projection plane, or calculate the projection as a sphere around a head of the listener, the sphere being the projection plane, or calculate the projection as the projection plane located at a predetermined distance from a center of the head of the listener, or calculate a projection of a hull of the spatially extended sound source as the projection plane from an azimuth angle and an elevation angle, the azimuth angle and the elevation angle being derived from a perspective spherical coordinate with respect to the head of the listener, the hull being a convex hull.

12. The apparatus of claim 1, wherein the sound position calculator is configured to calculate different positions of the at least two sound sources such that the different positions are uniformly distributed around a projection of a hull, or such that the different positions of the at least two sound sources are located at extreme or peripheral points of the projection of the hull, or such that the different positions of the at least two sound sources are located at horizontal or vertical extreme or peripheral points of the projection of the hull.

13. The apparatus of claim 1, wherein The sound position calculator is configured to determine the position of an additional auxiliary sound source located above or before or after or inside the hull projection relative to the listener in addition to the position of the peripheral sound sources.

14. The apparatus of claim 1, wherein The projector is configured to additionally shrink the projection of the hull in different directions, such as in horizontal and vertical directions, such as towards the center of gravity of the hull or the projection, by a variable or predetermined amount or by different variable or predetermined amounts.

15. The apparatus of claim 1, wherein, The sound position calculator is configured to calculate such that at least one additional auxiliary sound source is located on the projection plane between the left and right peripheral sound sources relative to the listener position, or wherein the sound position calculator is configured for calculating such that at least one additional auxiliary sound source is located on the projection plane between the left and right peripheral sound sources relative to the listener position, wherein a single additional auxiliary sound source is located in the middle between the left and right peripheral sound sources, or two or more additional auxiliary sound sources are placed equidistant between the left and right peripheral sound sources.

16. The apparatus of claim 1, wherein, The sound position calculator is configured to perform a rotation of the different positions of the at least two sound sources around the center of gravity of the projection in case of a circular motion of the listener around the spatially extended sound source via the interface or in case of a rotation of the spatially extended sound source relative to a stationary listener via the interface.

17. The apparatus of claim 1, wherein, The renderer is configured to receive an opening angle for each sound source depending on the distance between the listener and the sound source and to render the sound source according to the opening angle.

18. The apparatus of claim 1, wherein, The renderer is configured to receive distance information for each sound source and wherein the renderer is configured to render the sound sources according to the distances such that a sound source placed closer to the listener is rendered with a greater volume than a sound source placed less close to the listener and having the same volume.

19. The apparatus of claim 1, wherein, The sound position calculator is configured to determine for each sound source a distance equal to the distance of the spatially extended sound source relative to the listener, or determine the distance of each sound source by back-projecting the position of the sound source on the projection of the geometry of the spatially extended sound source, and wherein the renderer is configured to use the information about the distances to render the at least two sound sources.

20. The apparatus of claim 1, wherein, the information about the geometry is defined as a one-dimensional straight or curved line, a two-dimensional area, such as an ellipse, a rectangle or a polygon, or a set of polygons, or a three-dimensional body, such as an ellipsoid, a cuboid or a polyhedron, and / or wherein the information is defined as a parametric description or a polygon description or a parametric representation of a polygon description.

21. The apparatus of claim 1, wherein the sound position calculator is configured to determine the number of sound sources depending on the distance of the listener to the spatially extended sound source, wherein the number of sound sources is higher for smaller distances compared to a smaller number for larger distances between the listener and the spatially extended sound source.

22. The apparatus of claim 1, configured for receiving information about the propagation introduced by the spatially extended sound source, and wherein The projector is configured to apply a contraction operation to the envelope or the projection using the information about the propagation to at least partially compensate for the propagation.

23. The apparatus of claim 1, wherein The renderer is configured to render the sound source by combining the at least two elementary sound signals associated with the spatially extended sound source in case the positions of the sound source are identical to each other within the defined tolerance range, to obtain rotated elementary sound signals using Givens rotations and to render the rotated elementary sound signals at the positions as the different sound signals.

24. The apparatus of claim 1, wherein, The spatially extended sound source is associated with a multi-channel signal having a first channel and a second channel, the first channel being associated with a first part of the spatially extended sound source and the second channel being associated with a second part of the spatially extended sound source, wherein the first part is different from the second part and wherein the received specific information indicates that at least two sound sources are to be rendered by the renderer relative to a fixed position and / or orientation of the spatially extended sound source, and wherein the renderer is configured to determine the different sound signals at the different positions using a mapping of the first channel and the second channel to the different positions or using an additive combination of the first channel and the second channel to obtain the different sound signals at the different positions depending on the listener position and the first part and the second part of the spatially extended sound source.

25. The apparatus of claim 24, wherein, the first part is a left part of the spatially extended sound source and the second part is a right part of the spatially extended sound source, wherein the renderer is configured to use the first channel for the sound source position to the left of the user and the second channel for the position to the right of the user when the listener position is in front of the spatially extended sound source, or wherein the renderer is configured to use the second channel for the sound source position to the left of the user and the first channel for the position to the right of the user when the listener position is behind the spatially extended sound source, or wherein the renderer is configured to use an additive combination of the first channel and the second channel for the sound source position to the left of the user and to use an additive combination of the first channel and the second channel for the position to the right of the user when the listener position is to one side of the spatially extended sound source, or wherein the renderer is configured to use a weighted additive combination of the first channel and the second channel for the sound source position to the left of the user and to use a weighted additive combination of the first channel and the second channel for the position to the right of the user when the listener position is to one side of the spatially extended sound source, wherein weighting factors for the weighted additive combination are determined such that a weighting factor of a channel associated with a part of the spatially extended sound source closer to the listener position is greater than a weighting factor of another channel associated with another part of the spatially extended sound source further away from the listener position, or wherein, when the listener position is tilted with respect to the spatially extended sound source, the renderer is configured to use a first weighted addition of the first sound channel and the second sound channel for sound source positions to the left of the user and a second weighted addition of the first sound channel and the second sound channel for positions to the right of the user, wherein the weighting factors of the weighted additions are determined such that the weighting factor of a sound channel associated with a part of the spatially extended sound source closer to the sound source position is larger than the weighting factor of another sound channel associated with another part of the spatially extended sound source further away from the sound source position.

26. The apparatus according to claim 1, configured for receiving a description of a spatially extended sound source, the description comprising a description element indicating a first number of different elementary sound signals of the spatially extended sound source comprised in the description received by the apparatus or in an encoded audio signal, the number being one or larger than one, reading the description element and retrieving the first number of different elementary sound signals of the spatially extended sound source comprised in the description or in the encoded audio signal, and wherein the sound position calculator determines a second number of sound sources for rendering the spatially extended sound source, the second number being larger than one, and wherein the renderer is configured to generate a third number of one or more decorrelated signals from the first number extracted from the description, the third number resulting from a difference between the second number and the first number, or receiving a flag or a bitstream or a description element or information indicating an absolute anchoring of one or more different elementary sound signals of the spatially extended sound source to a fixed position and / or orientation of the spatially extended sound source as a certain information received, and wherein the renderer is configured to render at least two sound sources with respect to the fixed position and / or orientation of the spatially extended sound source in response to the bitstream or the description element or the flag or the information, or receiving a flag or a bitstream or a description element or information indicating an absolute anchoring of one or more different elementary sound signals of the spatially extended sound source to a fixed position or orientation of the spatially extended sound source in one state and indicating a different processing than in said one state in another state as a certain information received, and wherein the renderer is configured to render at least two sound sources with respect to the fixed position and / or orientation of the spatially extended sound source in response to the flag or the bitstream or the description element or the information indicating said one state and to render said at least two sound sources in a different mode in the other state comprising said different processing.

27. An apparatus for generating a description of a spatially extended sound source, the apparatus comprising: a sound provider for providing one or more different elementary sound signals for a spatially extended sound source; a geometry provider for calculating information about a geometry of the spatially extended sound source; and an output data former for generating a description comprising the one or more different elementary sound signals and the information about the geometry, wherein the output data former is configured to introduce a certain information or a description element or a flag into the description indicating an absolute anchoring of the one or more different elementary sound signals of the spatially extended sound source to a fixed position and / or orientation of the spatially extended sound source. ​ 28. The apparatus of claim 27, wherein, The information about the geometry comprises position information indicative of a position of the spatially extended sound source in space.

29. The apparatus of claim 27, comprising: wherein The output data former is configured to introduce into the description position information about an individual position of each of the one or more different basic sound signals, such that the position information about the individual position represents a position of the respective basic sound signal.

30. The apparatus of claim 27, wherein, The sound provider is configured to provide at least two different basic sound signals for the spatially extended sound source, and wherein the output data former is configured to generate the description such that the description comprises the at least two different basic sound signals and position information of each of the at least two different basic sound signals relative to the information about the geometry of the spatially extended sound source.

31. The apparatus of claim 27, wherein, The sound provider is configured to perform a recording of a natural sound source at a single or multiple microphone positions or locations, or derive the sound signals from a single or multiple basic signals by one or more decorrelation filters.

32. The apparatus of claim 27, wherein the sound provider is configured to bit rate compress the one or more basic sound signals using an audio signal encoder to obtain bit rate compressed one or more basic sound signals for the spatially extended sound source, and wherein, the output data former is configured to use the bit rate compressed one or more basic sound signals for the spatially extended sound source.

33. The apparatus of claim 27, wherein the geometry provider is configured to derive a parametric description or a polygon description or a parametric representation of a polygon description from the geometry of the spatially extended sound source, and wherein the output data former is configured to introduce the parametric description or the polygon description or the parametric representation of a polygon description into the description as the information about the geometry.

34. The apparatus of claim 27, wherein, The output data former is configured to introduce into the description a description element indicative of a number of the one or more different basic sound signals on the spatially extended sound source for inclusion in the description or in an encoded audio signal associated with the description, the number being one or more than one.

35. The apparatus of claim 27, wherein, The specific information or the flag or the description element indicative of the absolute anchoring of the one or more different basic sound signals of the spatially extended sound source refers to an absolute position or an absolute location of the spatially extended sound source, or wherein the syntax element comprises a relative channel position, and wherein the flag or the description element or the information comprises a flag or a prefix or a specific letter indicative of the anchoring, or wherein the sound provider is configured to provide at least two different basic sound signals for the spatially extended sound source, and wherein the flag or the description element or the information is associated with the at least two different basic sound signals, or wherein the at least two different sound signals are associated with a first channel associated with a left side portion of a piano and a second channel associated with a right side portion of the piano.

36. A method for rendering a spatially extended sound source having a defined position or location and geometry in space, the method comprising: receiving a listener position; calculating a projection of a two- or three-dimensional hull associated with the spatially extended sound source onto a projection plane using the listener position, the information on the geometry of the spatially extended sound source, and the information on the defined position or orientation of the spatially extended sound source; calculating different positions of at least two sound sources of the spatially extended sound source using the projection plane; and rendering the at least two sound sources at the different positions of the at least two sound sources to obtain a reproduction of the spatially extended sound source with two or more output signals, wherein the rendering comprises using different sound signals for the different positions of the at least two sound sources, wherein the different sound signals are associated with the spatially extended sound source, wherein the rendering comprises rendering the at least two sound sources relative to a fixed position and / or orientation of the spatially extended sound source in response to the received specific information.

37. A method for generating a description of a spatially extended sound source, the method comprising: providing one or more different basic sound signals for the spatially extended sound source; providing information on a geometry of the spatially extended sound source; and generating a description comprising the one or more different basic sound signals and the information on the geometry of the spatially extended sound source, wherein the generating comprises introducing into the description a flag or a description element or specific information indicative of an absolute anchoring of the one or more different basic sound signals of the spatially extended sound source to a fixed position or orientation of the spatially extended sound source.

38. The method according to claim 37, wherein the information on the geometry of the spatially extended sound source comprises position information of the spatially extended sound source in space.

39. The method according to claim 37, wherein, the generating the description comprises introducing into the description position information on an individual position of each of the one or more different basic sound signals.

40. The method according to claim 37, wherein the providing comprises providing at least two different basic sound signals for the spatially extended sound source, and wherein the generating the description is performed such that the description comprises the at least two different basic sound signals and position information of each of the at least two different basic sound signals such that the information is indicative of a position of the corresponding basic sound signal relative to the information on the geometry of the spatially extended sound source.

41. The method of claim 37, wherein, the generating the description comprises introducing into the description a description element indicative of a number of the one or more different basic sound signals on the spatially extended sound source for inclusion in the description or in an encoded audio signal associated with the description, the number being one or more than one.

42. A storage medium having stored thereon a computer program for performing, when running on a computer or a processor, the method according to claim 36 or 37.

Citation Information

Patent Citations

  • Extracting decomposed representations of a sound field based on a second configuration mode

    US20160366530A1