Apparatus and method for synthesizing spatially extended sound source using cue information items

By utilizing spatial range indication information and inter-channel correlation values, a spatially extended sound source is synthesized, solving the problems of high computational complexity and poor sound quality in existing technologies. This achieves efficient and stable spatially extended sound source reproduction, suitable for headphone and speaker reproduction.

CN115668985BActive Publication Date: 2026-07-24FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2021-03-12
Publication Date
2026-07-24

Smart Images

  • Figure CN115668985B_ABST
    Figure CN115668985B_ABST
Patent Text Reader

Abstract

An apparatus for synthesizing a spatially extended sound source, comprising a spatial information interface (100) for receiving a spatial extent indication, the spatial extent indication indicating a limited spatial extent within a maximum spatial extent (600) for a spatially extended sound source; a cue information provider (200) for providing one or more cue information items in response to the limited spatial extent; and an audio processor (300) for processing an audio signal representing the spatially extended sound source using the one or more cue information items.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] manual

[0002] This invention relates to audio signal processing, and more particularly to the reproduction of one or more spatially extended sound sources.

[0003] For various applications, it is necessary to reproduce sound sources through several speakers or headphones. These applications include 6-degrees-of-freedom (6DoF) virtual, mixed, or augmented reality applications. The simplest way to reproduce a sound source with such a setup is to render it as a point sound source. However, this model is insufficient when the intention is to reproduce a physical sound source with a non-negligible auditory spatial range. Examples of such sound sources are a grand piano, a choir, or a waterfall, all of which have a certain "size".

[0004] The realistic reproduction of sound sources with spatial extent has become the goal of many sound reproduction methods. This includes binaural reproduction using headphones, as well as conventional reproduction using speaker setups ranging from two speakers (“stereo”) to many speakers arranged horizontally (“surround sound”), and many speakers surrounding the listener in all three dimensions (“3D audio”). Descriptions of existing methods are given below. Therefore, considering the source width in 2D or 3D space, the different methods are grouped into multiple categories.

[0005] Methods for rendering spatially extended sound sources (SESS) on a 2D surface viewed from the listener's perspective are described. For example, this could be within a certain azimuth range at zero elevation (as is the case in traditional stereo / surround sound), or within a certain azimuth and elevation range (as is the case in 3D audio or virtual reality (VR) with 3 degrees of freedom for user movement, i.e., head rotation on the pitch / yaw / roll axes).

[0006] Increasing the apparent width of an audio object that is translated between two or more speakers (generating a so-called phantom or phantom source) can be achieved by reducing the correlation of the signals involved in the channel, see reference [1, pp. 241-257].

[0007] As the correlation decreases, the phantom source widens until the correlation value approaches zero, covering the entire range between speakers. Decorrelated versions of the source signal are obtained by deriving and applying appropriate decorrelation filters. Lauridsen's [2] proposal adds / subtracts its time-delayed and scaled version from the source signal itself to obtain two decorrelation versions of the signal. Kendall's [3] proposal, for example, presents a more complex approach. He iteratively derives paired decorrelation all-pass filters based on combinations of random number sequences. Faller et al. proposed suitable decorrelation filters ("diffusers") in [4, 5]. Additionally, Zotter et al.'s [6] derived filter pairs where frequency-correlated phase or amplitude differences are used to widen the phantom source. Alary et al.'s [7] proposed a decorrelation filter based on velvet noise, which was further optimized by Schlecht et al. in [8].

[0008] In addition to reducing the correlation of the corresponding channel signals of phantom sources, the source width can be increased by increasing the number of phantom sources attributed to audio objects. In reference [9], the source width is controlled by translating the same source signal to (slightly) different directions. The initially proposed method is to stabilize the perceptual phantom source propagation of the VBAP-panned source signal as the VBAP-panned source signal (see reference

[10] ) moves in the sound scene. This is advantageous because, depending on the direction of the source, the rendered source is reproduced by two or more speakers, which may lead to undesirable changes in the perceptual source width.

[0009] Virtual world DirAC (see reference

[11] ) is an extension of the traditional directional audio coding (DirAC) (see reference

[12] ) method for sound synthesis in a virtual world. In order to render the spatial extent, the directional sound components of the source are randomly translated within a certain range around the original direction of the source, and the translation direction changes with time and frequency.

[0010] A similar approach was used in reference

[13] , where the spatial range was achieved by randomly assigning the frequency band of the source signal to different spatial directions. This is a method intended to produce the same spatial distribution and envelope sound from all directions rather than to control the degree of precision.

[0011] Verron et al. achieved spatial extent of the source by synthesizing multiple incoherent versions of the source signal instead of using translation-correlated signals, uniformly distributing them on a circle around the listener, and mixing them together, see reference

[14] . The number of active sources and the gain determine the intensity of the broadening effect. This method is implemented as a spatial extension of an ambient sound synthesizer.

[0012] Methods related to rendering extended sound sources in 3D space are described, specifically in a volumetric manner required for VR with 6DoF of user movement. These 6 degrees of freedom include head rotation on the pitch / yaw / roll axes plus three translational movement directions x / y / z.

[0013] Potard et al. extended the concept of source region to a one-dimensional parameter of the source (i.e., its width between two speakers) by studying the perception of source shape, see reference

[15] . They generated multiple incoherent point sources by applying (time-varying) decorrelation techniques to the original source signal and then placing incoherent sources at different spatial locations to give them a three-dimensional range, see reference

[16] .

[0014] In the MPEG-4 Advanced Audio BIFS document

[17] , a volumetric object / shape (shell, box, ellipsoid and cylinder) can be filled with several uniformly distributed and decoupled sound sources to evoke a three-dimensional source range.

[0015] Recently, Schlecht et al.

[18] proposed a method to project a convex hull of the SESS geometry toward the listener's position, which allows the SESS to be rendered toward the listener at any relative position. Similar to MPEG-4 Advanced Audio BIFS, several decorrelated point sources are then placed within this projection.

[0016] In order to increase and control the source region using high-fidelity stereo reproduction (Ambisonics), Schmele et al.

[19] proposed a hybrid approach, namely reducing the Ambisonics order of the input signal (which inherently increases the apparent source width) and distributing decorrelation copies of the source signal around the listening space.

[0017] Zotter et al. introduced another method, in which they used the principle proposed in reference [6] for Ambisonics (i.e., to derive a filter pair that introduces frequency-dependent phase and amplitude differences to achieve source extension in a stereo reproduction setting), see reference

[20] .

[0018] A common drawback of translation-based methods (e.g., [10, 9, 12, 11]) is their dependence on the listener's position. Even small deviations from the optimal position can cause the spatial image to collapse into the speaker closest to the listener. This severely limits their application in VR and Augmented Reality (AR) environments where the listener should be able to move freely. Furthermore, distributing time-frequency points in DirAC-based methods (e.g., [12, 11]) does not always guarantee accurate rendering of the spatial extent of the phantom source. Moreover, it often significantly degrades the audio quality of the source signal.

[0019] Decorcorrelation of the source signal is typically achieved by one of the following methods: i) deriving a filter pair with complementary amplitudes (e.g., reference [2]), or ii) using an all-pass filter with a constant amplitude but (randomly) scrambled phase (e.g., references [3, 16]). Alternatively, broadening of the source signal can be achieved by randomly distributing the time-frequency points of the source signal in space (e.g., reference

[13] ).

[0020] Each method has its own implications: according to i), complementary filtering of the source signal generally leads to changes in the perceived sound quality of the decorrelated signal. Although all-pass filtering in ii) can preserve the sound quality of the source signal, the scrambling phase disrupts the original phase relationship, especially for transient signals, resulting in severe diffusion and trailing artifacts. Spatially distributing time-frequency points has proven effective for some signals, but it also alters the perceived sound quality of the signal. It exhibits high signal dependence and introduces severe artifacts into impulse signals.

[0021] As proposed in Advanced Audio BIFS (references [17, 15, 16]), filling a volumetric shape with multiple decorrelated versions of the source signal assumes the use of a large number of filters that generate mutually decorrelated output signals (typically more than ten point sources per volumetric shape). However, finding such filters is not an easy task, and it becomes more difficult the more such filters are needed. If the source signals are not perfectly decorrelated, and the listener moves around such shapes, such as in a VR scene, the distances to the individual sources from the listener correspond to different delays in the source signals. Therefore, their superposition at the listener's ear will result in position-dependent comb filtering, which may introduce annoying unstable coloring of the source signals. Furthermore, the application of many decorrelated filters implies significant computational complexity.

[0022] Similar considerations apply to the method described in reference

[18] , where many decorrelational point sources are placed on the convex shell projection of the SESS geometry. Although the authors do not mention anything about the required number of decorrelational auxiliary sources, it may be necessary to have a large number to achieve a convincing range of sources. This leads to the disadvantages already discussed in the preceding paragraphs.

[0023] Using the Ambisonics-based technique described in reference

[19] to control the source width by reducing the Ambisonics order, it appears to have an audible effect only on transitions from the 2nd to the 1st or to the 0th order. These transitions are not only perceived as source widening, but are often perceived as a shift in the phantom source. While adding a decorrelated version of the source signal can help stabilize the perception of the apparent source width, it also introduces a comb filtering effect, thereby altering the sonic quality of the phantom source.

[0024] The purpose of this invention is to provide an improved concept for synthesizing spatially extended sound sources.

[0025] The objective is achieved by the apparatus for synthesizing a spatially extended sound source as claimed in claim 1, the method for synthesizing a spatially extended sound source as claimed in claim 23, or the computer program as claimed in claim 24.

[0026] This invention is based on the discovery that the reproduction of a spatially extended sound source can be effectively achieved by using a spatial range indicator, which indicates a finite spatial target range for the spatially extended sound source within a maximum spatial range. Based on the spatial range indicator, and particularly based on the finite spatial range, one or more prompt information items are provided, and the processor uses these one or more prompt items to process the audio signal representing the spatially extended sound source.

[0027] This process enables efficient processing of spatially extended sound sources. For headphone reproduction, for example, only two binaural channels are needed, either a left binaural channel or a right binaural channel. For stereo reproduction, only two channels are also needed. Therefore, instead of using a large number of peripheral sound sources to synthesize spatially extended sound sources, which fill the actual volume or area of ​​the spatially extended sound source, or typically fill a limited spatial range due to their individual placement, this is not necessary according to the invention, because the spatially extended sound source is not rendered using a large number of individual sound sources placed within the volume, but rather using two or possibly three channels that are mutually cued when a large number of peripheral sound sources are received at two or three locations.

[0028] Therefore, unlike existing methods that aim to realistically reproduce Spatial Extended Sound Sources (SESS), which typically require large amounts of decorrelated input signals, this invention takes a different approach. Generating such decorrelated input signals can be relatively expensive in terms of computational complexity. Earlier existing methods may also impair the perceived quality of sound through phonological differences or tailing. Moreover, finding a large number of mutually orthogonal decorrelationalizers is generally not an easy problem to solve. Therefore, in addition to the large amount of computational resources required, this earlier process always results in a trade-off between the degree of mutual decorrelation and the introduced signal degradation.

[0029] In contrast, this invention uses only two decorrelated input signals to synthesize a small number of channels for spatially extended sound sources, such as a left channel and a right channel. Preferably, the synthesis result is a left-ear signal and a right-ear signal for headphone reproduction. However, this invention can also be applied to other types of reproduction scenarios, such as speaker rendering or active crosstalk reduction speaker rendering. Instead of placing many different decorrelated sound signals at different locations within the volume of the spatially extended sound source, in response to a limited spatial range indication received from a spatial information interface, one or more cue information items derived from a cue information provider are used to process the audio signal for the spatially extended sound source consisting of one or more channels.

[0030] The preferred embodiment aims to efficiently synthesize the SESS for headphone reproduction. Therefore, the synthesis is based on an underlying model describing the SESS using an infinite number of densely spaced, decorrelated point sources distributed across the entire source extent range (ideally). The desired source extent range can be expressed as a function of azimuth and elevation, which makes the method applicable to 3DoF applications. However, it can be extended to 6DoF applications by continuously projecting the SESS geometry in a direction toward the current listener's position, as described in reference

[18] . As a specific example, the desired source extent range is described below based on the azimuth and elevation range.

[0031] Further preferred embodiments rely on using inter-channel correlation values ​​as cue information, or additionally using inter-channel phase difference, inter-channel time difference, level difference, and gain factor or paired first and second gain factor information items. Thus, the absolute level of a channel can be set by two gain factors or by a single gain factor and inter-channel level difference. Instead of actual cue items, or in addition to actual cue items, any audio filtering function can be provided as cue information items from the cue information provider to the audio processor, so that the audio processor can synthesize (e.g.) two output channels, such as two binaural output channels or paired left and right output channels, by applying the actual cue item and optionally, by using a head-related transfer function for each channel as a cue item, or by using a head-related impulse response function as a cue item, or by using a binaural or (non-binaural) room impulse response function as a cue item. Typically, setting only a single cue item may be sufficient, but in more detailed embodiments, the audio processor may apply more than one cue item with or without a filter to the audio signal.

[0032] Therefore, in embodiments, when inter-channel correlation values ​​are provided as cue information items, and where the audio signal includes a first audio channel and a second audio channel for spatially expanding the sound source, or where the audio signal includes a first audio channel and a second audio channel derived from the first audio channel by a second channel processor, which implements, for example, decorrelation processing or neural network processing or any other processing for deriving a signal that can be considered decorrelated, the audio processor is configured to apply correlation between the first and second audio channels using the inter-channel correlation values, and, in addition to this processing or before or after this processing, an audio filtering function may also be applied to ultimately obtain two output channels having a target inter-channel correlation indicated by the inter-channel correlation values ​​and also having other relationships indicated by the respective filtering functions or other actual cue items.

[0033] The prompt information provider may be implemented as a lookup table including memory, or as a Gaussian mixture model, or as a support vector machine, or as a vector codebook, a multidimensional function fit, or some other device that effectively responds to spatial range indications to provide the desired prompts.

[0034] For example, in examples of lookup tables, or in examples of vector codebooks or multidimensional function fitting, or in examples of Gaussian mixture models (GMMs) or support vector machines (SVMs), prior knowledge may already be provided, so the primary task of the spatial information interface is actually to find a matching candidate spatial range that best matches the input spatial range indication information among all available candidate spatial ranges. This information can be provided directly by the user, or it can be calculated using some kind of projection calculation using information about the spatially extended sound source and the listener's position or orientation (e.g., determined by a head tracker or such device). The geometry or size of the object and the distance between the listener and the object may be sufficient to derive the opening angle, and thus the finite spatial range for rendering the sound source. In other embodiments, when the data received by the interface is already in a format that the cue information provider can use, the spatial information interface is merely used to receive the finite spatial range as input and forward that data to the cue information provider.

[0035] The preferred embodiments of the invention will then be discussed with reference to the accompanying drawings, in which:

[0036] Figure 1a A preferred embodiment of a device for synthesizing spatially extended sound sources is shown;

[0037] Figure 1b Another embodiment of the audio processor and prompt information provider is shown;

[0038] Figure 2 Show Figure 1a A preferred embodiment of the second channel processor included within the audio processor;

[0039] Figure 3 A preferred embodiment of the apparatus for performing ICC regulation is shown;

[0040] Figure 4 A preferred embodiment of the invention is shown, wherein the prompt information item depends on the actual prompt item and the filter;

[0041] Figure 5 This illustrates another embodiment that relies on filter and inter-channel correlation terms;

[0042] Figure 6 A schematic sector diagram is shown, which illustrates the maximum spatial extent in two-dimensional or three-dimensional cases, as well as the finite spatial extent of each sector or that can be used as, for example, a candidate sector.

[0043] Figure 7 This illustrates an implementation method for the spatial information interface;

[0044] Figure 8 This demonstrates another implementation of a spatial information interface that relies on the projection calculation process;

[0045] Figure 9a and 9b An embodiment for performing projection calculations and spatial extent determination is shown;

[0046] Figure 10 This illustrates another preferred implementation of the spatial information interface;

[0047] Figure 11 This illustrates another embodiment of the spatial information interface related to the decoder implementation;

[0048] Figure 12 The calculation of the finite spatial range of a spherical spatially extended sound source is shown;

[0049] Figure 13 Further calculations are shown for the finite spatial extent of the sound source in the ellipsoidal space.

[0050] Figure 14 Further calculations are shown for the finite spatial extent of the linear spatial extension sound source;

[0051] Figure 15 Further explanation of the calculation of the finite spatial range of a cuboid-shaped extended sound source;

[0052] Figure 16 This illustrates another example of using a sphere to calculate the finite spatial extent of a spatially extended sound source;

[0053] Figure 17 A spatially extended sound source with a piano shape and an approximate parametric ellipsoidal shape is shown;

[0054] Figure 18 Points are shown to define a finite spatial range, which is used to render the spatially extended sound source of the piano shape.

[0055] Figure 1aA preferred embodiment of an apparatus for synthesizing a spatially extended sound source (SESS) is shown. The apparatus includes a spatial information interface 10 that receives spatial range indication information input, which indicates a limited spatial range within a maximum spatial range for the SESS. The limited spatial range is input to a cue information provider 200, which is configured to provide one or more cue information items in response to the limited spatial range given by the spatial information interface 10. The cue information items, or a plurality of cue information items, are provided to an audio processor 300, which is configured to process an audio signal representing the SESS using the one or more cue information items provided by the cue information provider 200. The audio signal for the SESS can be a single channel, or it can be a first audio channel and a second audio channel, or it can be two or more audio channels. However, for the purpose of low processing load, a small number of channels for the SESS or the audio signal representing the SESS is preferred. An audio signal is input to the audio signal interface 305 of the audio processor 300, and the audio processor 300 processes the input audio signal received by the audio signal interface, or when the number of input audio channels is less than the required number (e.g., only one), the audio processor includes... Figure 2 The second channel processor 310 shown includes a decorrelation unit for generating a second audio channel S2 decorrelated to the first audio channel S. Figure 2 S1 is also shown in the diagram. The information items can be actual information items, such as inter-channel correlation items, inter-channel phase difference items, inter-channel level difference and gain items, gain factor items G1, G2, together representing inter-channel level difference and / or absolute amplitude or power or energy level, for example. Alternatively, the information items can also be actual filter functions, such as head-related transfer functions, the number of which is required as the actual number of output channels to be synthesized in the synthesized signal. Therefore, when the synthesized signal is to have two channels, such as two binaural channels or two speaker channels, each channel requires a head-related transfer function. Instead of head-related transfer functions, head-related impulse response functions (HRIR) or binaural or non-binaural room impulse response functions (B)RIR are required. Figure 1a As shown, each channel requires such a transfer function, and Figure 1a An implementation with two channels is shown, hence the indices "1" and "2".

[0056] In this embodiment, the prompt information provider 200 is configured to provide inter-channel correlation values ​​as prompt information items. The audio processor 300 is configured to actually receive the first and second audio channels via the audio signal interface 305. However, when the audio signal interface 305 receives only a single channel, the optional second channel processor may be provided, for example, by means of… Figure 2 The process in the audio processor generates the second audio channel. The audio processor performs correlation processing to apply a correlation between the first and second audio channels using inter-channel correlation values.

[0057] Additionally or alternatively, other information items may be provided, such as inter-channel phase difference, inter-channel time difference, inter-channel level difference and gain, or first and second gain factor information items. Items may also be inter-aural (IACC) related values, i.e., more specific inter-channel related values, or inter-aural phase difference (IAPD) items, i.e., more specific inter-channel phase difference values.

[0058] In a preferred embodiment, the audio processor 300 applies a correlation in response to the relevant cue information item before performing ICPD, ICTD, or ICLD adjustments, or before performing HRTF or other transfer filtering function processing. However, the order may be set differently depending on the circumstances.

[0059] In a preferred embodiment, the audio processor includes memory for storing information about different cue information items related to different spatial range indications. In this case, the cue information provider also includes an output interface for retrieving from memory one or more cue information items associated with spatial range indications input to the corresponding memory. For example, in Figure 1b , Figure 4 or Figure 5 The image shows such a lookup table 210, which includes memory and an output interface for displaying the corresponding prompt information item. Specifically, the memory can store not only... Figure 1b The IACC, IAPD, or G shown l and G r Values, but the memory within the lookup table can also store, for example... Figure 4 and Figure 5 The filtering function shown in box 220 is indicated as "Select HRTF". In this embodiment, although in Figure 4 and Figure 5 The figures are shown separately, but boxes 210 and 220 may include the same memory, wherein corresponding prompt information items, such as IACC and optional IAPD, and transfer functions for filters (such as HRTF1 for the left output channel and HRTFr for the right output channel) are stored, associated with corresponding spatial range indicators indicating azimuth and elevation angles. Figure 4 or Figure 5 or Figure 1b In the middle, the left and right output channels are designated as S1 and S2 respectively. r .

[0060] The memory used by lookup table 210 or selection function box 220 may also utilize storage devices, where corresponding parameters are available based on certain sector codes, sector angles, or sector angle ranges. Optionally, the memory may, as appropriate, store vector codebooks or multidimensional function fitting routines, or Gaussian mixture models (GMMs) or support vector machines (SVMs).

[0061] Given a desired source region range, a SESS can be synthesized using two decorrelated input signals. These input signals are processed in a manner that accurately reproduces perceptually important auditory cues. These include the following interaural cues: interaural cross-correlation (IACC), interaural phase difference (IAPD), and interaural level difference (IALD). In addition, mono-channel spectral cues are also reproduced. These are crucial for source localization in the vertical plane. Although IAPD and IALD are also important for localization purposes, IACC is known to be a key cue for source width perception in the horizontal plane. During runtime, target values ​​for these cues are retrieved from a pre-computed store. In the following sections, lookup tables are used for this purpose. However, other methods of storing multidimensional data, such as vector codebooks or multidimensional function fitting, can also be used. Apart from the source region range considered, all cues depend only on the Head Relational Transfer Function (HRTF) dataset used. The derivation of different auditory cues is given later.

[0062] exist Figure 1b The overall block diagram of the proposed method is shown in the figure. [Φ1, Φ2] describes the desired source region with respect to the azimuth angle range. [θ1, θ2] is the desired source region with respect to the elevation angle range. S1(ω) and S2(ω) represent two decorrelated input signals, where ω describes the frequency index. Therefore, for S1(ω) and S2(ω), the following equation holds:

[0063]

[0064] Additionally, both input signals need to have the same power spectral density. Alternatively, only one input signal S(ω) can be given. The second input signal is generated internally using a decorrelation, such as... Figure 2 As shown. Given S l (ω) and S r (ω), an extended sound source is synthesized by sequentially adjusting inter-channel coherence (ICC), inter-channel phase difference (ICPD), and inter-channel level difference (ICLD) to match the corresponding interaural cue. The quantities required for these processing steps are read from a pre-computed lookup table. The resulting left and right channel signals S l (ω) and S r(ω) can be played through headphones and is similar to SESS. It should be noted that ICC adjustment must be performed first, but the ICPD and ICLD adjustment boxes are interchangeable. Instead of IAPD, the corresponding interaural time differences (IATDs) can also be reproduced. However, in the following text, only IAPD will be considered.

[0065] In the ICC adjustment box, the cross-correlation between the two input signals is adjusted to the desired value |IACC(ω)| using the following formula (see document

[21] ):

[0066]

[0067]

[0068]

[0069]

[0070] These formulas can lead to the desired cross-correlation if the input signals S1(ω) and S2(ω) are completely decorrelated. Furthermore, their power spectral densities must be identical. The corresponding block diagram is shown below. Figure 3 As shown.

[0071] The ICPD adjustment box is described by the following formula:

[0072]

[0073]

[0074] Ultimately, the ICLD adjustment was performed as follows:

[0075]

[0076]

[0077] Among them, G l (ω) describes the gain of the left ear, and G r (ω) describes the right ear gain. As long as... and They do indeed have the same power spectral density, which leads to the expected ICLD. Because the left and right ear gains are used directly, the mono spectral cue can be reproduced except for IALD.

[0078] To further simplify the previously discussed method, two options are described for simplification. As previously mentioned, the primary interaural cue affecting the perceived spatial range (in the horizontal plane) is the IACC. Therefore, it is conceivable to adjust directly via the HRTF instead of using pre-calculated IAPD and / or IALD values. For this, the HRTF corresponding to the location representing the desired source area range is used. At this location, the average of the desired azimuth / elevation range can be selected without loss of generality. Descriptions of the two options are given below.

[0079] The first option involves using pre-calculated IACC and IAPD values. However, the ICLD is adjusted using the HRTF corresponding to the center of the source region extent.

[0080] The first option's diagram is as follows Figure 4 As shown. Now calculate S using the following formula. l (ω) and D r (ω):

[0081]

[0082]

[0083] in and Describe the location of the HRTF, which represents the average of the desired azimuth / elevation range. The main advantages of the first option include:

[0084] • Compared to point sources at the center of the source region, there is no spectral shaping / coloring as the source region increases.

[0085] • Compared to the low storage requirements of a fully-blown system, because G l (ω) and G r (ω) does not need to be stored in the lookup table.

[0086] Compared to methods with all features, this approach offers greater flexibility in modifying the HRTF dataset during runtime because only the resulting ICC and ICPD, rather than ICLD, depend on the HRTF dataset used during pre-computation.

[0087] The main drawback of this simplified version compared to the non-extended source is that it will fail whenever the IALD changes significantly. In such cases, the IALD will not be reproduced with sufficient accuracy. This is the case, for example, when the source is not centered around a 0° azimuth and the source region becomes too large in the horizontal direction.

[0088] The second option involves using only pre-calculated IACC values. Adjust ICPD and ICLD using the HRTF corresponding to the center of the source region range.

[0089] The second option's flowchart is as follows Figure 5 As shown in the figure. Now calculate S using the following formula. l (ω) and S r (ω):

[0090]

[0091]

[0092] In contrast to the first option, now the phase and amplitude of the HRTF are used, not just the amplitude. This allows for adjustment not only of the ICLD but also of the ICPD. The main advantages of the second option include:

[0093] • For the first option, increasing the source region does not result in spectral shaping / coloring compared to a point source at the center of the source region.

[0094] • Even lower storage requirements than the first option, because G l (ω) and G r (ω) and IAPD do not need to be stored in the lookup table.

[0095] • Compared to the first option, this option offers greater flexibility in changing the HRTF dataset during runtime. Only the generated ICC depends on the HRTF dataset used during pre-computation.

[0096] • It can be effectively integrated into existing binaural rendering systems; simply put, two different inputs, and It needs to be used to generate left and right ear signals.

[0097] For the first option, this simplified version will fail as long as there is a significant change to the IALD compared to the non-extended source. Furthermore, the change to the IAP should not be too large compared to the non-extended source. However, since the IAP of the extended source is quite close to that of a point source at the center of the source region, the latter is not expected to be a major issue.

[0098] Figure 6An exemplary schematic sector diagram is shown. Specifically, the schematic sector diagram is shown at 600, and schematic sector diagram 600 shows the maximum spatial range. When the schematic sector diagram is considered as a two-dimensional schematic diagram of the three-dimensional surface of a sphere, this is achieved by showing the azimuth and elevation ranges (from 0° to 360° for azimuth and from -90° to +90° for elevation). It is evident that when the schematic sector diagram is wrapped around a sphere and the listener's position is placed within the center of the sphere, all the individual sectors shown exemplarily in some examples, namely S1 to S24, can subdivide the entire surface of the sphere into sectors. Therefore, for example, when applying... Figure 1b , Figure 4 , Figure 5 When the symbol is used, sector S3 extends with respect to the azimuth range from Φ1 = 60° to Φ2 = 90°. Sector S3 extends, exemplarily, within the elevation range between -30° and 0°.

[0099] However, schematic sector diagram 600 can also be used when the listener is not placed inside the center of the sphere, but rather with respect to the sphere being placed at some location. In this case, only certain sectors of the sphere are visible, but it is not necessary for certain informational items to be available for all sectors of the sphere. Only for some (required) sectors are certain informational items required to be available, and these informational items are preferably pre-calculated as discussed later, or alternatively obtained by measurement.

[0100] Optionally, the schematic sector diagram can be viewed as a two-dimensional maximum range where spatially extended sound sources can be located. In this case, the horizontal distance extends between 0% and 100%, and the vertical distance extends between 0% and 100%. The actual vertical distance or extension and the actual horizontal distance or extension can be mapped to absolute distance or extension by a certain absolute scaling factor. For example, when the scaling factor is 10 meters, 25% corresponds to 2.5 meters in the horizontal direction. In the vertical direction, the scaling factor can be the same as or different from the scaling factor in the horizontal direction. Thus, for the horizontal / vertical distance / extension example, sector S5 will extend between 33% and 42% of the (maximum) scaling factor with respect to the horizontal dimension, and sector S5 will extend between 33% and 50% of the vertical scaling factor within the vertical range. Thus, for example, the maximum spatial range of a sphere or non-sphere can be subdivided into finite spatial ranges or sectors S1 to S24.

[0101] To effectively adapt rasterization to human auditory perception, it is preferable to have low resolution in the vertical or height direction and higher resolution in the horizontal or azimuth direction. Exemplarily, sectors covering only the entire height range of a sphere can be used, meaning that a single-line sector extending only from, for example, S1 to S12 can be used as different sectors or limited spatial ranges, where the horizontal dimension is given by certain angular values, while the vertical dimension extends from -90° to +90° for each sector. Naturally, other sectorization techniques can also be used, such as in... Figure 6 There are 24 sectors, of which sectors S1 to S12 cover the entire height or vertical range between -90° and 0° or between 0% and 50%, while other sectors S13 to S24 cover the upper hemisphere between elevation angles of 0° and 90°, or the upper half of the "horizon" extending between 50% and 100%.

[0102] Figure 7 Show Figure 1a A preferred embodiment of the spatial information interface 10. Specifically, the spatial information interface includes an actual (user) receiving interface for receiving a spatial range indication. The spatial range indication can be input by the user, or, in the case of virtual reality, derived from head tracker information, or the enhancement matcher 30 can match the actually received finite spatial range with available candidate spatial ranges known from the prompt information provider 200 to find a matching candidate spatial range that best approximates the actual input finite spatial range. Based on the matching candidate spatial range, from Figure 1a The cue information provider 200 transmits one or more cue information items, such as inter-channel data or filter functions. Matching candidate spatial ranges or finite spatial ranges can include paired azimuth angles or paired elevation angles, or both, for example... Figure 1b As shown, it illustrates the azimuth and height ranges of the sector.

[0103] Alternatively, such as Figure 6 As shown, a finite spatial extent can be limited by information about horizontal distance, information about vertical distance, or information about both vertical and horizontal distance. When the maximum spatial extent is rasterized in two dimensions, not only is a single vertical or horizontal distance sufficient, but pairs of vertical and horizontal distances are also necessary, as shown with respect to sector S5. Alternatively again, the finite spatial extent information may include a code for a specific sector that identifies the finite spatial extent as the maximum spatial extent, where the maximum spatial extent comprises multiple distinct sectors. Such codes are given, for example, by labels S1 to S24, since each code is uniquely associated with a specific geometric two-dimensional or three-dimensional sector at the schematic sector diagram 600.

[0104] Figure 8Another embodiment of the spatial information interface is shown, which again comprises a user receiving interface 100, but now also includes a projection calculator 120 and a subsequently connected spatial range determiner 140. The user receiving interface 100 exemplarily receives the listener's position, which includes the user's actual location in an environment and / or the user's orientation at that location. Therefore, the listener's position may refer to the actual location, or the actual orientation, or both (actual listener location and actual listener orientation). Based on this data, the projection calculator 120 uses information about the spatially extended sound source to calculate so-called shell projection data. The SESS information may include the geometry of the spatially extended sound source and / or the location and / or orientation of the spatially extended sound source, etc. Based on the shell projection data, the spatial range determiner 140... Figure 6 One of the alternatives shown defines a finite spatial extent, or as per [the following] Figure 10 , 11 or Figures 12 to 18 The discussion focuses on the finite spatial extent, which is determined by... Figure 12 and Figure 18 The examples shown in the examples are given by two or more feature points, where the set of feature points is always limited to a certain finite spatial range from the entire spatial range.

[0105] Figure 9a and Figure 9b The calculation is shown by Figure 8 Different ways of outputting the shell projection data of frame 120. Figure 9a In one embodiment, the spatial information interface is configured to use the geometry of the spatially extended sound source, which serves as information about the spatially extended sound source, to calculate the shell of the spatially extended sound source, as shown in box 121. Using the listener's position, the shell of the spatially extended sound source is projected toward the listener 122 to obtain a projection of the two-dimensional or three-dimensional shell onto the projection plane. Alternatively, as... Figure 9b As shown, the spatially extended sound source, and particularly the geometry of the spatially extended sound source defined by information about its geometry, is projected in a direction toward the listener's position, as shown in box 123, and as shown in box 124, the shell of the projected geometry is calculated to obtain the projection of the two-dimensional or three-dimensional shell onto the projection plane. Finite spatial extent representation. Figure 9a The projected shell or in the embodiment Figure 9b The vertical / horizontal or azimuth / height extension of the shell of the projected geometry obtained by the implementation method.

[0106] Figure 10 A preferred embodiment of the spatial information interface 10 is shown. It includes a listener location interface 100, which in Figure 8 It is also shown as the user receiving interface. Additionally, as... Figure 8As shown, the input is the location and geometry of the spatially extended sound source, and a projector 120 and a calculator 140 for calculating the finite spatial range are also provided.

[0107] Figure 11 A preferred embodiment of a spatial information interface is shown, including interface 100, projector 120, and a limited spatial range location calculator 140. Interface 100 is configured to receive the listener's location. Projector 120 is configured to use the listener's location received by interface 100, along with additional information about the geometry of the spatially extended sound source and additional information about the location of the spatially extended sound source in space, to calculate the projection of a two-dimensional or three-dimensional shell associated with the spatially extended sound source onto a projection plane. Preferably, the defined location of the spatially extended sound source in space and additional information about the geometry of the spatially extended sound source in space are received for reproducing the spatially extended sound source via a bitstream arriving at a bitstream demultiplexer or scene resolver 180. Bitstream demultiplexer 180 extracts information about the geometry of the spatially extended sound source from the bitstream and provides this information to the projector. Bitstream demultiplexer also extracts the location of the spatially extended sound source from the bitstream and forwards this information to the projector.

[0108] Preferably, the bitstream also includes audio signals for SESS with one or two different audio signals, and preferably, the bitstream demultiplexer further extracts compressed representations of one or more audio signals from the bitstream, and one or more signals are decompressed / decoded by the decoder of the audio decoder 190. The decoded one or more signals are ultimately forwarded, for example, to... Figure 1a The audio processor 300, and the processor and Figure 1a The prompt information provider 200 renders at least two sound sources in accordance with the prompts provided.

[0109] although Figure 11 A bitstream-related reproduction device is shown, which has a bitstream demultiplexer 180 and an audio decoder 190. However, reproduction can also be performed in scenarios different from the encoder / decoder scenario. For example, the defined location and geometry in space may already exist at the reproduction device in a scene such as virtual reality or augmented reality, where data is generated on-site and used at the same location. The bitstream demultiplexer 180 and audio decoder 190 are not actually required, and information about the geometry and location of the spatially extended sound source is available without needing to be extracted from the bitstream.

[0110] Preferred embodiments of the invention are then discussed. These embodiments relate to rendering spatially extended sound sources in 6DoF VR / AR (Virtual Reality / Augmented Reality).

[0111] Preferred embodiments of the present invention pertain to a method, apparatus, or computer program designed to enhance the reproduction of a Spatial Extended Sound Source (SESS). Specifically, embodiments of the method or apparatus of the present invention take into account the time-varying relative position between the spatial extended sound source and the virtual listener's location. In other words, embodiments of the method or apparatus of the present invention allow the width of the auditory source to match the spatial extent of the represented sound object at any location relative to the listener. Thus, embodiments of the method or apparatus of the present invention are particularly suitable for 6-DOF virtual, mixed, and augmented reality applications, where the spatial extended sound source complements conventionally employed point sources.

[0112] Embodiments of the method or apparatus of the present invention render a spatially extended sound source by using a limited spatial extent. The limited spatial extent depends on the listener's position relative to the spatially extended sound source.

[0113] Figure 1a A general block diagram depicting a spatially extended sound source renderer according to an embodiment of the method or apparatus of the present invention. The key components of the block diagram are:

[0114] 1. Listener Position: The box provides the listener's instantaneous position, such as the position measured by a virtual reality tracking system. The box can be implemented as a detector 100 for detecting the listener's position or as an interface 100 for receiving the listener's position.

[0115] 2. Location and geometry of spatially extended sound sources: The box provides the location and geometry data of the spatially extended sound sources to be rendered, for example, as part of a virtual reality scene representation.

[0116] 3. Projection and Convex Shell Calculation: Box 120 calculates the convex shell of the spatially extended sound source geometry and then projects it in a direction toward the listener's position (e.g., the "image plane," see below). Alternatively, the same functionality can be achieved by first projecting the geometry toward the listener's position and then calculating its convex shell.

[0117] 4. Location Determination of the Limited Spatial Area: Box 140 calculates the location of the limited spatial area based on the convex hull projection data calculated from the previous box. The listener's location, and therefore proximity / distance (see below), may also be considered in this calculation. The output is, for example, the location of the points that collectively define the limited spatial area.

[0118] Figure 10 An overview of a block diagram illustrating an embodiment of the method or apparatus of the present invention. Dashed lines indicate the transmission of metadata, such as geometry and location.

[0119] The location of the point that jointly defines the finite spatial extent depends on the geometry of the spatially extended sound source, especially the spatial extent, and the listener's relative position to the spatially extended sound source. Specifically, the point defining the finite spatial extent can be the projection of the convex shell of the spatially extended sound source onto a projection plane. The projection plane can be a picture plane, i.e., a plane perpendicular to the line of sight from the listener to the spatially extended sound source, or it can be a spherical surface surrounding the listener's head. The projection plane is located at an arbitrarily small distance from the center of the listener's head. Alternatively, the projected convex shell of the spatially extended sound source can be calculated from the azimuth and elevation angles, which are subsets of the spherical coordinates relative to the listener's head. In the illustrative example below, a projection plane is preferred because it has more intuitive characteristics. When implementing the calculation of the projected convex shell, an angle representation is preferred due to its simpler formalization and lower computational complexity. The projection of the convex shell of the spatially extended sound source is the same as the convex shell of the projected spatially extended sound source geometry; that is, the convex shell calculation and the projection onto the picture plane can be used in any order.

[0120] When the listener's position relative to the spatially expanded sound source changes, the projection of the spatially expanded sound source onto the projection plane changes accordingly. In turn, the positions of the points defining the finite spatial extent change accordingly. These points are preferably chosen such that they change smoothly with continuous movement of both the spatially expanded sound source and the listener. Changing the geometry of the spatially expanded sound source alters the projected convex shell. This involves rotating the geometry of the spatially expanded sound source in 3D space, thereby changing the projected convex shell. The rotation of the geometry is equal to the angular displacement of the listener's position relative to the spatially expanded sound source, and is referred to, for example, in an inclusive manner, as the relative position of the listener and the spatially expanded sound source. For example, rotating the points defining the change in the finite spatial extent around the center of gravity represents the listener's circular motion around a spherical spatially expanded sound source. Similarly, rotating the spatially expanded sound source with a fixed listener results in the same change in the points defining the finite spatial extent.

[0121] For any distance between the spatially extended sound source and the listener, the spatial extent generated by embodiments of the method or device of the present invention is inherently and accurately reproduced. Naturally, as the user approaches the spatially extended sound source, the opening angle between points defining a finite spatial extent will increase, as it is suitable for modeling physical reality.

[0122] Therefore, the angular arrangement of points within a limited spatial range is uniquely determined by their positions on the projected convex shell on the projection plane.

[0123] To specify the geometry / convex hull of a spatially extended sound source, approximations are used (and possibly transferred to the renderer or renderer kernel), including simplified one-dimensional shapes such as straight lines and curves; 2D shapes such as ellipses, rectangles, and polygons; or 3D shapes such as ellipsoids, cuboids, and polyhedra. The geometry or corresponding approximate shape of a spatially extended sound source can be described in various ways, including:

[0124] • Parametric description, which formalizes geometry through mathematical expressions that accept additional parameters. For example, a 3D ellipsoid shape can be described by implicit functions in Cartesian coordinates, with the additional parameters being the extensions of the principal axes in all three directions. Other parameters may include 3D rotations and deformation functions of the ellipsoid's surface.

[0125] • A polygon description is a collection of primitive geometric shapes, such as lines, triangles, squares, tetrahedrons, and cuboids. Primitive polygons and polyhedra can be connected to form larger and more complex geometric shapes.

[0126] In some application scenarios, the focus is on compact and interoperable storage / transmission of 6DoF VR / AR content. In this case, the entire chain consists of three steps:

[0127] 1. Create / encode the desired spatial expansion sound source into a bitstream.

[0128] 2. Transmission / storage of the generated bitstream. According to the invention, among other elements, the bitstream also includes a description (parameterized or polygonal) of the spatially extended sound source geometry and the associated source underlying signal, such as a mono or stereo piano recording. The waveform can be compressed using perceptual audio coding algorithms such as mp3 or MPEG-2 / 4 Advanced Audio Coding (AAC).

[0129] 3. As mentioned above, the spatially extended sound source is decoded / rendered based on the transmitted bitstream.

[0130] Subsequently, various practical implementation examples are presented. These include spherical spatial extension sound sources, ellipsoidal spatial extension sound sources, linear spatial extension sound sources, cuboid spatial extension sound sources, distance-dependent finite spatial ranges, and / or piano-shaped spatial extension sound sources or spatial extension sound source shapes like any other musical instrument.

[0131] As described above in embodiments of the method or apparatus of the present invention, various methods can be applied to determine the location of points defining a finite spatial range. The practical examples below illustrate some isolated methods in specific situations. In a full implementation of embodiments of the method or apparatus of the present invention, various methods can be appropriately combined considering computational complexity, application purpose, audio quality, and ease of implementation.

[0132] The spatially extended sound source geometry is represented as a surface mesh. It's important to note that the mesh visualization does not imply that the spatially extended sound source geometry is described using a polygonal approach, as the geometry can actually be generated from a parametric specification. The listener's position is represented by a blue triangle. In the following example, the image plane is chosen as the projection plane and depicted as a transparent gray plane indicating a finite subset of the projection plane. The projected geometry of the spatially extended sound source onto the projection plane is depicted using the same surface mesh. Points defining a finite spatial extent on the projected convex hull are represented as crosses on the projection plane. Back-projection points defining a finite spatial extent on the spatially extended sound source geometry are depicted as points. Lines connecting corresponding points defining a finite spatial extent on the projected convex hull and back-projection points defining a finite spatial extent on the spatially extended sound source geometry aid in identifying visual correspondences. The positions of all objects involved are depicted in meters in a Cartesian coordinate system. The choice of the depicted coordinate system does not imply that the calculations involved are performed using Cartesian coordinates.

[0133] Figure 12 The first example considers a spherical spatially extended sound source. The spherical spatially extended sound source has a fixed size and fixed position relative to the listener. Three different sets of three, five, and eight points are selected on the projected convex shell to define a finite spatial extent. All three sets of points defining the finite spatial extent are selected at uniform distances on the curve of the convex shell. The offset positions of the points defining the finite spatial extent on the curve of the convex shell are deliberately chosen to well represent the horizontal extent of the spatially extended sound source geometry. Figure 12 A spherical spatially extended sound source is shown, which has a finite spatial range defined by a different number of points (i.e., 3 (top), 5 (middle), and 8 (bottom)) uniformly distributed on a convex shell.

[0134] Figure 13 The next example considers an ellipsoidally spatially extended sound source. An ellipsoidally spatially extended sound source has a fixed shape, position, and rotation in 3D space. In this example, four points are selected to define a finite spatial extent. Three different methods for determining the positions of points defining a finite spatial extent are illustrated:

[0135] a) Place two points with limited spatial ranges at two horizontal extreme points, and place two points with limited spatial ranges at two vertical extreme points. However, extreme point positioning is simple and often appropriate. This example shows that this method may produce point positions that are relatively close to each other.

[0136] b) All four points within the defined finite spatial range are evenly distributed on the projected convex shell. The offset of the points defining the finite spatial range is selected so that the position of the highest point coincides with the position of the highest point in a).

[0137] c) All four points within a finite spatial range are uniformly distributed on the contracted projected convex shell. The offset position of each point is equal to the offset position selected in b). The contraction operation of the projected convex shell is performed towards the centroid of the projected convex shell using a direction-independent stretching factor.

[0138] therefore, Figure 13 An ellipsoidal spatially extended sound source is shown under three different methods for determining the location of points that define a finite spatial range, with four points defining the finite spatial range: a / top) horizontal and vertical extreme points, b / middle) points uniformly distributed on the convex shell, and c / bottom) points uniformly distributed on the contracting convex shell.

[0139] Figure 14 The next example considers a linear spatially extended sound source. While the previous examples considered volumetric spatially extended sound source geometries, this example illustrates that spatially extended sound source geometries can be well chosen as single-dimensional objects within 3D space. Subfigure a) depicts two points defining a finite spatial extent placed at the extreme points of a finite linear spatially extended sound source geometry. b) Two points defining a finite spatial extent are placed at the extreme points of a finite linear spatially extended sound source geometry, and another point is placed in the middle of the line. As described in embodiments of the method or apparatus of the present invention, placing additional points within the spatially extended sound source geometry can help fill large gaps in large spatially extended sound source geometries. c) Considers the same linear spatially extended sound source geometry as in a) and b), but with a different relative angle toward the listener, resulting in a significantly smaller projected length of the linear geometry. As described in the embodiments of the method or apparatus of the present invention above, the reduced size of the projected convex shell can be represented by a reduced number of points defining a finite spatial extent, which in this specific example can be represented by a single point located at the center of the linear geometry.

[0140] therefore, Figure 14 The diagram shows a linear spatial extension sound source, with the locations of points defining a limited spatial range distributed using three different methods: a / top) two extreme points on the projected convex shell; b / middle) two extreme points on the projected convex shell, with an additional point at the center of the line; c / bottom) one or two points defining a limited spatial range at the center of the convex shell, because the convex shell projected by the rotating line is too small to accommodate one or more points.

[0141] Figure 15The next example considers a cuboid spatially extended sound source. The cuboid spatially extended sound source has a fixed size and fixed position, but the listener's relative position changes. Subfigures a) and b) depict different methods of placing four points defining a finite spatial extent on a projected convex shell. The positions of the back-projected points are uniquely determined by the selection on the projected convex shell. c) depicts four points defining a finite spatial extent, whose back-projected positions are not well separated. Instead, the distance between the selected point positions is equal to the distance between the centroids of the spatially extended sound source geometry.

[0142] therefore, Figure 15 The diagram illustrates a cuboid spatially extended sound source, with points defining a finite spatial range distributed using three different methods: a / top) two points defining the finite spatial range on the horizontal axis and two points defining the finite spatial range on the vertical axis; b / middle) two points defining the finite spatial range at the horizontal extreme points of the projected convex shell and two points defining the finite spatial range at the vertical extreme points of the projected convex shell; c / bottom) selecting back-projection points at a distance equal to the centroid of the geometry of the spatially extended sound source.

[0143] Figure 16 The next example considers a spherical spatially extended sound source of fixed size and shape, but at three different distances relative to the listener's position. Points defining a finite spatial range are uniformly distributed on the curve of the convex shell. The number of points defining the finite spatial range is dynamically determined based on the length of the convex shell curve and the minimum distance between possible point locations. a) The spherical spatially extended sound source is at a close distance, thus four points defining the finite spatial range are selected on the projected convex shell. b) The spherical spatially extended sound source is at a medium distance, thus three points defining the finite spatial range are selected on the projected convex shell. a) The spherical spatially extended sound source is at a far distance, thus only two points defining the finite spatial range are selected on the projected convex shell. As described above in embodiments of the method or apparatus of the present invention, the number of points defining the finite spatial range can also be determined based on the range represented in the spherical angular coordinates.

[0144] therefore, Figure 16 The diagram shows spherical spatially extended sound sources of equal size but at different distances: a / top) near distance, where four points defining a finite spatial range are evenly distributed on the projected convex shell; b / middle) medium distance, where three points defining a finite spatial range are evenly distributed on the projected convex shell; c / bottom) far distance, where two points defining a finite spatial range are evenly distributed on the projected convex shell.

[0145] exist Figure 17 and Figure 18The last example considers a spatially expanded sound source in the shape of a piano placed in a virtual world. The user wears a head-mounted display (HMD) and headphones. A virtual reality scene is presented to the user, consisting of a blank canvas and a 3D upright piano model standing on the floor within a freely movable area (see [link to example]). Figure 17 An open-world canvas is a static image of a sphere projected onto a sphere surrounding the user. In this particular case, the open-world canvas depicts a blue sky with white clouds. The user can move around and view and listen to the piano from various angles. In this scene, cues are used to render the piano; these cues represent either a single point source placed at the center of gravity or a spatially extended sound source with three points defining a finite spatial extent on the projected convex shell (see [link to relevant documentation]). Figure 18 ).

[0146] To simplify point calculations, the piano's geometry is abstracted as an ellipsoid with similar dimensions; see [link to relevant documentation]. Figure 17 Place two alternative points at the left and right extremes of the equator, while a third alternative point remains at the North Pole, see [reference needed]. Figure 18 This arrangement ensures an appropriate horizontal source width from all angles, while significantly reducing computational costs.

[0147] therefore, Figure 17 This illustrates a spatially extended sound source in the shape of a piano with an approximate parametric ellipsoidal shape, while Figure 18 A spatially extended sound source in the shape of a piano is shown, having three points defining a finite spatial range. These three points are located at the vertical extreme points and the vertical top of the projected convex shell. Note that, for better visualization, the points defining the finite spatial range are placed on the stretched projected convex shell.

[0148] The described technology can be applied as part of the 6DoF audio VR / AR standard. In this context, a classic encoder / bitstream / decoder (+renderer) scenario is employed:

[0149] • In the encoder, the shape of the spatially extended sound source is encoded together with the "basic" waveform of the spatially extended sound source as auxiliary information. The "basic" waveform of the spatially extended sound source can be a representation of the spatially extended sound source:

[0150] o Mono signal, or

[0151] o stereo signal (preferably fully decorrelated), or

[0152] o or even more recorded signals (preferably also fully decorrelated)

[0153] These waveforms can be encoded at a low bit rate.

[0154] • In the decoder / renderer, as previously described, the spatially extended sound source shape and corresponding waveform are retrieved from the bitstream, and the spatially extended sound source shape and corresponding waveform are used to render the spatially extended sound source.

[0155] Depending on the embodiment used and as alternatives to the described embodiments, it should be noted that the interface may be implemented as an actual tracker or detector for detecting the listener's location. However, the listener's location will typically be received from an external tracker device and fed into the playback device via the interface. However, the interface may simply represent data input for output data from an external tracker, or it may represent the tracker itself.

[0156] As outlined, a bitstream generator can be implemented to generate a bitstream with only one audio signal for spatially expanding the sound source, and the remaining audio signals are generated on the decoder or reproduction side through decorrelation. When only a single signal exists, and the entire space is to be averaged with that single signal, no positional information is needed. However, in this case, at least some additional information about the geometry of the spatially expanding sound source may be useful.

[0157] Depending on the implementation method, preferably in Figure 1a , Figure 1b , Figure 4 , Figure 5 The prompt information provider 200 uses some type of pre-computed data to provide the correct prompt information item for a given environment. This pre-computed data, such as data from... Figure 6 The sector diagram 600, with a set of values ​​for each sector, can be measured and stored to determine, for example, the data within lookup table 210 and the selected HRTF box 220 empirically. In another embodiment, this data can be pre-calculated, or it can be derived through a combination of empirical and pre-calculated methods. A preferred embodiment for calculating this data is then given.

[0158] During the lookup table generation, the IACC, IAPD, and IALD values ​​required for SESS synthesis, as described above, are pre-calculated for multiple source region ranges.

[0159] As mentioned earlier, as the underlying model, SESS is described by an infinite number of decorrelational point sources distributed across the entire source region. This model can be approximated by placing a decorrelational point source at each HRTF dataset location within the desired source region. By convolving these signals with the corresponding HRTF, the resulting left and right ear signals, Y, can be determined. l (ω) and Y r (ω). From this, the values ​​of IACC, IAPD, and IALD can be derived. The corresponding expressions are given below.

[0160] Given N decorrelation signals S n (ω), having the same power spectral density:

[0161]

[0162] in,

[0163]

[0164] Where N equals the number of HRTF dataset points within the desired source region. Therefore, these N input signals are each placed at different HRTF dataset locations, where...

[0165]

[0166]

[0167] It is important to note that: A l,n A r,n Φ l,n And A l,n This usually depends on ω. However, for the sake of notation simplification, this dependency is omitted here. Using equations (16) and (17), Y is respectively... l (ω) and Y r The left and right ear signals of (ω) can be represented as follows:

[0168]

[0169]

[0170] To determine IACC, IALD, and IAPD, the following is first derived: E{|Y l (ω)| 2} and E{|Y r (ω)| 2 The expression for}:

[0171]

[0172]

[0173]

[0174] Using equations (20) to (22), the following expressions for IACC(ω), IALD(ω), and IAPD(ω) can be determined:

[0175]

[0176]

[0177]

[0178] E{|Y} can be normalized by the number of sources and the source power respectively. l (ω)| 2 and E{|Y r (ω)| 2 To determine G respectively l (ω) and G r Left and right ear gains of (ω):

[0179]

[0180]

[0181] As can be seen, all the generated expressions depend only on the selected HRTF dataset and no longer on the input signal.

[0182] To reduce computational complexity during lookup table generation, one possibility is to disregard the location of every available HRTF dataset. In this case, the expected interval is limited. While this process reduces computational complexity during pre-computation, it will also lead to a degradation of the solution to some extent.

[0183] Compared with the prior art, the preferred embodiments of the present invention offer significant advantages.

[0184] Starting from the fact that the proposed method only requires two decorrelation input signals, it offers many advantages compared to current technologies that require a larger number of decorrelation input signals:

[0185] The proposed method exhibits low computational complexity because it requires only one decorrelation. Furthermore, it only requires filtering of two input signals.

[0186] Since pairwise decorrelation is typically higher when generating less decorrelation signal (and allows for the same amount of signal degradation), it is expected to reproduce auditory cues more accurately.

[0187] Similarly, in order to achieve the same amount of pairwise decorrelation and thus the same accuracy of reproduced auditory cues, more signal degradation is expected.

[0188] Subsequently, several interesting features of the embodiments of the present invention are summarized.

[0189] 1. Only two decorrelated input signals are needed (or one input signal plus a decorrelator).

[0190] 2. [Frequency Selectivity] Adjust the binaural cues of these input signals to effectively achieve binaural output signals for spatially extended sound sources (instead of modeling many single-point sources covering the area / volume of the SESS).

[0191] (a) Always adjust the input ICC.

[0192] (b) ICPD / ICTD and ICLD can be adjusted in a dedicated processing step, or ICPD / ICTD and ICLD can be introduced into the signal by using HRIR / HRTF processing that takes advantage of these characteristics.

[0193] 3. [Frequency Selectivity] The target binaural cue is determined from pre-computed storage (lookup table or other way of storing multidimensional data, such as vector codebook or multidimensional function fitting, GMM, SVM) based on the spatial range to be filled (specific example: azimuth range, height range).

[0194] (a) The target IACC is always stored and retrieved / used for synthesis.

[0195] (b) Target IAPD / IATD and IALD can be stored and retrieved / used for synthesis, or can be replaced using HRIR / HRTF processing.

[0196] A preferred embodiment of the present invention can be used as part of MPEG-1 audio 6DoF VR / AR (Virtual Reality / Augmented Reality standard). In this context, there are encoding / bitstream / decoder (plus renderer) applications. In the encoder, the shape of the spatially extended sound source or several spatially extended sound sources is encoded together with one or more “spatial” waveforms of the spatially extended sound sources as auxiliary information. These waveforms, representing the signal input to box 300 (i.e., the audio signal for the spatially extended sound sources), can be encoded at a low bit rate using AAC, EVS, or any other encoder. In the decoder / renderer, where, for example, in… Figure 11 The applications illustrated include a bitstream demultiplexer (parser 180 and audio decoder 190) that retrieves the SESS shape and corresponding waveform from the bitstream and uses them to render the SESS. The process illustrated in this invention provides a high-quality yet low-complexity decoder / renderer.

[0197] Although some aspects have been described in the context of the device, it is clear that these aspects also represent descriptions of the corresponding methods, where boxes or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding boxes, items, or features of the corresponding device.

[0198] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, storing electronically readable control signals that can (or are capable of) cooperating with a programmable computer system to execute the corresponding methods.

[0199] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0200] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.

[0201] Other embodiments include a computer program stored on a machine-readable carrier or non-transitory storage medium for performing one of the methods described herein.

[0202] In other words, therefore, an embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.

[0203] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein.

[0204] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transferred via a data communication connection, such as via a network.

[0205] Another embodiment includes a processing device, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0206] Another embodiment includes a computer having a computer program installed on it for performing one of the methods described herein.

[0207] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Typically, the method is preferably performed by any hardware device.

[0208] The embodiments described above are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent of the invention is limited only by the scope of the pending patent claims, and not by the specific details given in the description and explanation of the embodiments herein.

[0209] References

[0210] [1] J. Blauert, Spatial Hearing: Psychophysics of Human Sound Localization, 3rd ed. Cambridge, Mass: MIT Press, 2001.

[0211] [2]H.Lauridsen, "Experiments Concerning Different Kinds of Room-Acoustics Recording," Ingenioren, 1954.

[0212] [3]G.Kendall, "The Decorrelation of Audio Signals and Its Impact onSpatial Imagery," Computer Music Journal, vol.19, no.4, pp.71–87, 1995.

[0213] [4] C.Faller and F.Baumgarte, "Binaural cue coding-Part II:Schemes and applications," IEEE Transactions on Speech and Audio Processing, vol.11, no.6, pp.520–531, Nov.2003.

[0214] [5] F.Baumgarte and C.Faller, "Binaural cue coding-Part I: Psychoacoustic fundamentals and design principles," IEEE Transactions onSpeech and Audio Processing, vol.11, no.6, pp.509–519, Nov.2003.

[0215] [6]F.Zotter and M.Frank,“Efficient Phantom Source Widening,”Archivesof Acoustics,vol.38,pp.27–37,Mar.2013.

[0216] [7]B.Alary,A.Politis,and V. “Velvet-noise decorrelator,”Proc.DAFx-17,Edinburgh,UK,pp.405–411,2017.

[0217] [8]S.Schlecht,B.Alary,V. and E.Habets,“Optimized velvet-noisedecorrelator,”Sep.2018.

[0218] [9]V.Pulkki,“Uniform spreading of amplitude panned virtual sources,”Proceedings of the 1999 IEEE Workshop on Applications of Signal Processing toAudio and Acoustics.WASPAA’99(Cat.No.99TH8452),pp.187–190,1999.

[0219]

[10] ——,“Virtual Sound Source Positioning Using Vector BaseAmplitude Panning,”Journal of the Audio Engineering Society,vol.45,no.6,pp.456–466,Jun.1997.

[0220]

[11] V.Pulkki,M.-V.Laitinen,and C.Erkut,“Efficient Spatial SoundSynthesis for Virtual Worlds.”Audio Engineering Society,Feb.2009.

[0221]

[12] V.Pulkki,“Spatial Sound Reproduction with Directional AudioCoding,”Journal of the Audio Engineering Society,vol.55,no.6,pp.503–516,Jun.2007.

[0222]

[13] T. O.Santala,and V.Pulkki,“Synthesis of Spatially ExtendedVirtual Source with Time-Frequency Decomposition of Mono Signals,”Journal ofthe Audio Engineering Society,vol.62,no.7 / 8,pp.467–484,Aug.2014.

[0223]

[14] C.Verron,M.Aramaki,R.Kronland-Martinet,and G.Pallone,“A 3-DImmersive Synthesizer for Environmental Sounds,”Audio,Speech,and LanguageProcessing,IEEE Transactions on,vol.18,pp.1550–1561,Sep.2010.

[0224]

[15] G.Potard and I.Burnett,“A study on sound source apparent shapeand wideness,”pp.6–9,Aug.2003.

[0225]

[16] ——,“Decorrelation techniques for the rendering of apparentsound source width in 3D audio displays,”Jan.2004,pp.280–208.

[0226]

[17] J.Schmidt and E.F.Schroeder,“New and Advanced Features for AudioPresentation in the MPEG-4 Standard.”Audio Engineering Society,May 2004.

[0227]

[18] S.Schlecht,A.Adami,E.Habets,and J.Herre,“Apparatus and Method forReproducing a Spatially Extended Sound Source or Apparatus and Method forGenerating a Bitstream from a Spatially Extended Sound Source,”PatentApplication PCT / EP2019 / 085 733.

[0228]

[19] T.Schmele and U.Sayin,“Controlling the Apparent Source Size inAmbisonics Using Decorrelation Filters.”Audio Engineering Society,Jul.2018.

[0229]

[20] F.Zotter,M.Frank,M.Kronlachner,and J.-W.Choi,“Efficient PhantomSource Widening and Diffuseness in Ambisonics,”Jan.2014.

[0230]

[21] C.Borβ,“An Improved Parametric Model for the Design of VirtualAcoustics and its Applications,”Ph.D.dissertation,Ruhr- Bochum,Jan.2011.

Claims

1. An apparatus for synthesizing a spatially extended sound source, comprising: A spatial information interface (100) is used to receive a spatial range indication, which indicates a limited spatial range within a maximum spatial range (600) for spatially extending the sound source; A prompt information provider (200) is configured to provide one or more prompt information items in response to the limited spatial range, wherein the one or more prompt information items include inter-channel correlation values ​​provided in response to the limited spatial range; and An audio processor (300) is configured to process an audio signal representing the spatially extended sound source using the one or more cue information items. The audio signal includes a first audio channel for the spatially extended sound source and a second audio channel for the spatially extended sound source, or the audio signal includes a first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source is derived from the first audio channel by a second channel processor (310). The audio processor (300) is configured to perform correlation processing on the first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source to apply (320) correlation between the first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source using the inter-channel correlation value provided in response to the limited spatial range.

2. The device according to claim 1, The aforementioned prompt information provider (200) is configured to provide at least one of the following as additional prompt information items: inter-channel phase difference item, inter-channel time difference item, inter-channel level difference and gain item, and a first gain information item and a second gain information item; and The audio processor (300) is configured to apply an inter-channel phase difference, inter-channel time difference, or inter-channel level difference or absolute level to the first audio channel and the second audio channel using at least one of the inter-channel phase difference term, the inter-channel time difference term, the inter-channel level difference and the gain term, and the first gain information term and the second gain information term.

3. The device according to claim 2, The audio processor (300) is configured to apply an inter-channel phase difference (330), an inter-channel time difference, or an inter-channel level difference (340) or an absolute level to the first and second audio channels after the correlation processing; or The second channel processor (310) includes a decorrelation filter or a neural network processor, which is used to derive the second audio channel from the first audio channel so that the second audio channel is decorrelated with the first audio channel.

4. The device according to claim 1, The cue information provider (200) includes a filter function provider (220) for providing an audio filter function as one or more cue information items in response to the limited spatial range; and The audio processor (300) includes a filter applicator (350) for applying the audio filtering function to the first audio channel and the second audio channel.

5. The device according to claim 4, For each of the first audio channel and the second audio channel, the audio filtering function includes a head-related transfer function, a head-related impulse response, a binaural room impulse response, or a room impulse response; or The second channel processor (310) includes a decorrelation filter or a neural network processor, which is used to derive the second audio channel from the first audio channel so that the second audio channel is decorrelated with the first audio channel.

6. The device according to claim 4, The filter applicator (350) is configured to apply the audio filtering function to the result of the correlation processing performed by the audio processor (300) using the inter-channel correlation value.

7. The device according to claim 1, The notification information provider (200) includes: Memory (210) is used to store information about different prompt information items related to different finite space ranges; and An output interface is used to retrieve one or more prompt information items associated with the limited space range using the memory (210).

8. The device according to claim 7, The memory (210) mentioned therein includes at least one of a lookup table, a vector codebook, a multidimensional function fitting, a Gaussian mixture model (GMM), and a support vector machine (SVM); and The output interface is configured to retrieve one or more prompt information items by looking up the lookup table, by using the vector codebook, by applying the multidimensional function fit, or by using the GMM or the SVM.

9. The device according to claim 1, The cue information provider (200) is configured to store information about one or more cue information items associated with a set of intervals of candidate spatial ranges, the set of intervals of finite spatial ranges covering the maximum spatial range (600), wherein the cue information provider (200) is configured to: match the finite spatial range with candidate finite spatial ranges (30), wherein the candidate finite spatial ranges define candidate spatial ranges that are closest to a specific finite spatial range defined by the finite spatial ranges, and provide one or more cue information items associated with the matched candidate finite spatial ranges; or The finite spatial range includes information about azimuth, paired elevation angles, information about horizontal distance, information about vertical distance, information about total distance, and at least one of azimuth and paired elevation angles; or The spatial range indication includes codes (S3, S5) that identify the finite spatial range as a specific sector of the maximum spatial range (600), wherein the maximum spatial range (600) includes multiple different sectors.

10. The device of claim 9, wherein the sector in the plurality of different sectors has a first extension in the azimuth or horizontal direction and a second extension in the height or vertical direction, wherein the second extension in the height or vertical direction of the sector is greater than the first extension, or wherein the second extension covers a maximum height or vertical direction range.

11. The device of claim 9, wherein the plurality of different sectors are defined such that the distance between the centers of adjacent sectors in the azimuth or horizontal direction is greater than 5 degrees, or even greater than or equal to 10 degrees.

12. The device according to claim 1, The audio processor (300) is configured to generate a processed first audio channel and a processed second audio channel from the audio signal for binaural rendering, speaker rendering, or active crosstalk reduction speaker rendering.

13. The device according to claim 1, The prompt information provider (200) is configured to provide one or more inter-channel prompt values ​​as the one or more prompt information items in addition to the inter-channel related values; The audio processor (300) is configured to generate the processed first audio channel and the processed second audio channel from the audio signal in such a manner that the processed first audio channel and the processed second audio channel have one or more inter-channel cues controlled by the one or more inter-channel cues values.

14. The device according to claim 1, The cue information provider (200) is configured to provide one or more cue information items for multiple frequency bands in response to the limited spatial range, wherein the one or more cue information items for different frequency bands of the multiple frequency bands are different from each other.

15. The device according to claim 1, The aforementioned prompt information provider (200) is configured to provide one or more prompt information items for multiple different frequency bands; and The audio processor (300) is configured to process the audio signal in the spectral domain, wherein one or more prompt information items for multiple different frequency bands are applied to multiple spectral values ​​of the audio signal in the frequency band.

16. The device according to claim 1, The first audio channel and the second audio channel are decorrelated with each other to a certain extent; The prompt information provider (200) is configured to provide the inter-channel correlation value as one or more prompt information items; and The audio processor (300) is configured to reduce the correlation between the first audio channel and the second audio channel to the amount indicated by the inter-channel correlation value provided by the cue information provider (200).

17. The device according to claim 1, further comprising: An audio signal interface (305) is configured to receive an audio signal representing the spatially extended sound source, wherein the audio signal comprises only the first audio channel, or wherein the audio signal comprises only the first audio channel and the second audio channel, or wherein the audio signal does not include more than the first audio channel and the second audio channel.

18. The device according to claim 1, wherein the spatial information interface (100) is configured to: Receive (100) the listener's location as an indication of the spatial range: Using the listener's location and information about the spatially extended sound source, calculate (120) the projection of the two-dimensional or three-dimensional shell associated with the spatially extended sound source onto the projection plane; or using the listener's location and information about the spatially extended sound source, calculate (120) the projection of the geometry of the spatially extended sound source onto the projection plane as a two-dimensional or three-dimensional shell; and The finite spatial range of the spatially extended sound source is determined (140) from the shell projection data.

19. The device according to claim 18, wherein the spatial information interface (100) is configured to: Using the geometry of the spatially extended sound source as information about the spatially extended sound source, calculate (121) the shell of the spatially extended sound source, and project (122) the shell in the direction toward the listener using the listener position to obtain the projection of the two-dimensional or three-dimensional shell onto the projection plane; or project (123) the geometry of the spatially extended sound source defined by the information about the geometry of the spatially extended sound source in the direction toward the listener position, and calculate (124) the shell of the projected geometry to obtain the projection of the two-dimensional or three-dimensional shell onto the projection plane.

20. The device of claim 18, wherein the spatial information interface (100) is configured to determine the finite spatial range such that the boundary of the sector defined by the finite spatial range is located on the right side of the projection plane relative to the listener and / or on the left side of the projection plane relative to the listener and / or on the upper side of the projection plane relative to the listener and / or on the lower side of the projection plane relative to the listener, or coincides with one of the right, left, upper, and lower boundaries of the projection plane relative to the listener, for example within a tolerance of + / -10%.

21. A method for synthesizing a spatially extended sound source, comprising: Receive a spatial range indication, which indicates a limited spatial range within a maximum spatial range (600) for spatially extending the sound source; In response to the limited spatial range, one or more prompt information items are provided, wherein the one or more prompt information items include inter-channel correlation values ​​provided in response to the limited spatial range; and Using the one or more of the aforementioned prompt information items, the audio signal representing the spatially extended sound source is processed. The audio signal includes a first audio channel for the spatially extended sound source and a second audio channel for the spatially extended sound source, or the audio signal includes a first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source is derived from the first audio channel by processing the second channel. The processing includes performing correlation processing on the first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source to apply a (320) correlation between the first audio channel for the spatially extended sound source and the second audio channel for the spatially extended sound source using the inter-channel correlation value provided in response to the limited spatial range.

22. A computer-readable storage medium having a computer program stored thereon, the computer program being configured to perform the method of claim 21 when executed on a computer or processor.

Citation Information

Patent Citations

  • Sound processing method and sound processing device

    CN108781341A

  • Distance rendering method for audio signal and apparatus for outputting audio signal using same

    US20180077514A1