Audio rendering suitable for a reverberation chamber

The audio processor stabilizes sound quality by adjusting gain compensation based on reverberation effects, addressing the issue of degraded audio reproduction due to listener movement in reverberant environments.

JP2025523679AInactive Publication Date: 2025-07-23FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025501469
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-12
Filing Date
2023-07-07
Publication Date
2025-07-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing audio reproduction systems using loudspeakers are optimized for a narrow range of listener positions and suffer significant quality degradation when the listener moves, with previous methods failing to account for reverberant environments.

Method used

An audio processor that adjusts gain compensation based on a roll-off gain compensation function, considering reverberation effects to stabilize sound quality across varying listener positions by using a listener-loudspeaker distance compensation gain that becomes shallower as the distance increases.

Benefits of technology

Enhances audio rendering stability and accuracy by compensating for reverberation, allowing consistent sound quality regardless of listener position within a realistic playback environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523679000001_ABST
    Figure 2025523679000001_ABST
Patent Text Reader

Abstract

An audio processor for performing audio rendering by generating rendering parameters that determine the derivation of an audio signal of a loudspeaker signal to be reproduced by a set of loudspeakers. The audio processor is configured to perform gain adjustment to obtain reverberation effect information and determine a gain for generating a loudspeaker signal for the loudspeakers from the audio signal based on the listener position. The audio processor is configured to use a roll-off gain compensation function for associating, in the gain adjustment, a listener-loudspeaker distance of at least one loudspeaker with a listener-loudspeaker distance compensation gain for at least one loudspeaker according to the reverberation effect information, and for the roll-off gain compensation function, the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance increases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments according to the present invention relate to audio processors, systems, methods, and computer programs for audio rendering, such as rendering of user-adapted loudspeakers for reverberation chambers, for example.

Background Art

[0002] A common problem in audio reproduction using loudspeakers is that the reproduction is usually optimal only at one or a narrow range of listener positions. Even worse, when the listener changes position or moves, the quality of the audio reproduction varies greatly. The evoked spatial auditory image becomes unstable as the listening position changes away from the sweet spot. The stereo image collapses towards the nearest loudspeaker.

[0003] This problem has been addressed by previous publications, including [1] tracking the listener's position and adjusting the gain and delay to compensate for the derivation from the optimal listening position. [2] shows an extension regarding how to adapt to the spatial radiation characteristics of the loudspeakers being used. Listener tracking is also used in conjunction with crosstalk cancellation (XTC) (see, for example, [3]). XTC requires extremely accurate positioning of the listener, whereby listener tracking becomes almost essential.

[0004] Previous methods for listener-position-adapted gain compensation for loudspeaker signals have assumed that the sound energy (and thus the required compensation gain) tends to roll off at a constant rate with distance. As an example, the theoretical roll-off ("slope") of the acoustic energy with respect to this distance is 6 dB every time the distance to the point source doubles. Other slope values may also apply. However, in practice, these dependencies only hold under very dry conditions (close to an anechoic chamber), which are rarely seen in real-world sound reproduction environments.

[0005] Therefore, for the purpose of optimizing the quality of the output audio signal of the loudspeaker for listeners at different listening positions, it is desirable to obtain a concept with a compensation gain method that can also consider a playback environment including a certain amount of reverberant sound.

[0006] This object is achieved by the subject matter of the independent claims.

[0007] Advantageous embodiments are the subject matter of the dependent claims.

Summary of the Invention

Problems to be Solved by the Invention

[0008] The object of the present invention is to provide a more realistic distance gain compensation taking into account the fact that there is reverberant energy in a realistic playback environment (room / playback space). This difficulty is overcome by considering reverberation effect information in gain adjustment / compensation. In particular, a roll-off gain compensation function for associating the listener-loudspeaker distance with the compensation gain is used, which takes into account, for example, the effect of reverberation. The concept behind the embodiments of the present invention is that due to the presence of reverberation in the sound playback environment, the gain to be compensated does not increase uniformly, i.e., at a constant coefficient, with the increase in the distance of the listener to the loudspeaker. This is based on the recognition that in a realistic room, the acoustic energy rolls off more slowly with the increase in the distance between the position of the loudspeaker and the listener than in an anechoic playback environment. Due to reverberation, the attenuation of the sound energy can, for example, decrease with the increase in the distance of the listener to the loudspeaker. This correlation is reflected by a roll-off gain compensation function that takes into account, for example, that the roll-off compensated by the compensation gain becomes shallower monotonically with the increase in the listener-loudspeaker distance. Using a roll-off gain compensation function in such a way may seem to complicate the calculation compared to gain adjustment that takes into account a certain roll-off of the sound energy, but in fact, this gain adjustment enhances the rendering stability and the accuracy of the sound reproduced by the loudspeaker at the listener position.

Means for Solving the Problem

[0009] Accordingly, an embodiment relates to an audio processor for performing audio rendering by generating rendering parameters, the rendering parameters determining the derivation of an audio signal of a loudspeaker signal to be reproduced by a set of loudspeakers. The audio processor is configured to perform gain adjustment to obtain reverberation effect information and determine a gain for generating a loudspeaker signal for a loudspeaker from the audio signal based on the listener position. The audio processor is configured to use, in the gain adjustment, a roll-off gain compensation function for associating, for at least one loudspeaker, a listener-loudspeaker distance of at least one loudspeaker with a listener-loudspeaker distance compensation gain for at least one loudspeaker according to the reverberation effect information, and for the roll-off gain compensation function, the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance increases. In other words, the roll-off gain compensation function may be configured to compensate for a roll-off of sound energy that becomes monotonically shallower as the listener-loudspeaker distance increases, that is, the roll-off of sound energy decreases as the listener-loudspeaker distance increases. The slope of the roll-off gain compensation function may become monotonically shallower as the listener-loudspeaker distance increases. For example, "shallower" with respect to the compensation gain increases more slowly when the listener-loudspeaker distance is long than when the listener-loudspeaker distance is short, that is, as the listener-loudspeaker distance increases, the rising rate of the compensation gain becomes smaller.

[0010] The reverberation effect information may indicate, for example, the amount of effective reverberation in the playback room of an audio rendering, or whether reverberation is effective in the playback room of an audio rendering. According to one embodiment, the reverberation effect information may comprise a first compensated roll-off slope of a roll-off gain compensation function, a second compensated roll-off slope of the roll-off gain compensation function, a near-field attenuation parameter, a non-near-field attenuation parameter, a critical distance parameter, and / or a near-field - non-near-field transition parameter. The first compensated roll-off slope and the second compensated roll-off slope may indicate a compensated gain per distance or sound energy per distance. The near-field attenuation parameter and the non-near-field attenuation parameter may indicate a roll-off of acoustic energy per distance, and the near-field attenuation parameter may indicate a higher attenuation compared to the non-near-field attenuation parameter. The first compensated roll-off slope may be related to the near-field attenuation parameter, and the second compensated roll-off slope may be related to the non-near-field attenuation parameter. The critical distance parameter may indicate a distance to a certain loudspeaker of a set of loudspeakers, for example a boundary distance, which separates two distance regions associated with different reverberation effects. For example, a first distance region with a shorter distance than the boundary distance, i.e., the near field, may be associated with a higher roll-off of sound energy than a second distance region with a longer distance than the boundary distance, i.e., the non-near field. The critical distance parameter may indicate a distance to a certain loudspeaker of a set of loudspeakers at which the energy of the direct sound becomes equal to the energy of the reverberant sound. The near-field - non-near-field transition parameter may indicate how fast the transition between near-field attenuation and non-near-field attenuation is, for example, how the roll-off gain compensation function transitions from the first distance region to the second distance region.

[0011] The listener position can be defined, for example, by tracking data, by coordinates indicating the position of the listener within the playback space, such as the position of the listener's body, the position of the listener's head, or the position of the listener's ears. The listener position can be described, for example, in Cartesian coordinates, spherical coordinates, or cylindrical coordinates. Instead of the absolute position of the listener, the listener position can indicate the relative position of the listener with respect to, for example, a reference loudspeaker of a set of loudspeakers, or with respect to each loudspeaker of a set of loudspeakers, or with respect to a sweet spot within the playback space, or with respect to any other predetermined position within the playback space.

[0012] A further embodiment relates to a method for audio rendering by generating rendering parameters, the rendering parameters determining the derivation of an audio signal into loudspeaker signals to be reproduced by a set of loudspeakers. The method comprises the step of obtaining reverberation effect information and the step of performing gain adjustment to determine a gain for generating loudspeaker signals for the loudspeakers from the audio signal based on the listener position. Depending on the reverberation effect information, the gain adjustment uses a roll-off gain compensation function for associating, for at least one loudspeaker, the listener-loudspeaker distance of at least one loudspeaker with a listener-loudspeaker distance compensation gain for at least one loudspeaker, and for the roll-off gain compensation function, the compensated roll-off becomes monotonically shallower with an increase in the listener-loudspeaker distance.

[0013] A further embodiment relates to a computer program or a digital storage medium storing it. The computer program has program code for instructing a computer to perform one of the methods described herein when the program is executed on the computer.

[0014] Further embodiments relate to a bitstream or a digital storage medium storing the same, as referred to herein. The bitstream may comprise, for example, reverberation effect information and / or listener position and / or loudspeaker signals and / or audio signals.

[0015] The methods, computer program products, and bitstreams described herein are based on the same considerations as the audio processors described herein. It should be noted that the methods, computer programs, and bitstreams can be completed with all features and / or functions, and these are also described with respect to the audio processors.

[0016] The drawings are not necessarily to scale and it is emphasized that they generally illustrate the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10a

Figure 10b

Figure 10c

Figure 10d

Figure 10e

Figure 10f

Figure 10g

Figure 10h

Figure 10i-1

Figure 10i-2

Figure 11a

Figure 11b

Figure 11c

DETAILED DESCRIPTION OF THE INVENTION

[0018] In the following description, elements that are equal or equivalent, or elements having equal or equivalent functions, are denoted by equal or equivalent reference numerals even when they appear in different figures.

[0019] In the following description, in order to provide a more complete description of embodiments of the present invention, a plurality of details are set forth. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not shown in detail in order to avoid obscuring embodiments of the present invention and are shown in the form of block diagrams. Additionally, features of different embodiments described later in this specification may be combined with each other unless otherwise specifically stated.

[0020] In the following, various examples are described that may help achieve more effective compression when using gain and / or delay adjustments controlled by listener position. The gain adjustment and / or delay adjustment may be added to other parameter adjustments for sound rendering, for example, or provided exclusively.

[0021] To simplify the understanding of the following examples of this application, the description begins by presenting a possible device that conforms to this application, upon which the examples outlined later in this application can be built. The following description begins with an explanation of an embodiment of a device for generating loudspeaker signals for a plurality of loudspeakers. More specific embodiments that can be applied to the device of FIG. 1 individually or collectively are outlined below in this specification together with a detailed description.

[0022] The device of FIG. 1 is generally designated by reference numeral 10 and is for generating loudspeaker signals 12 for a plurality of loudspeakers 14 in such a way as to render at least one audio object at a virtual position where application of the loudspeaker signals 12 to or at the plurality of loudspeakers 14 is intended.

[0023] The device 10 can be configured for an arrangement of loudspeakers 14, i.e., for several positions where a plurality of loudspeakers 14 are located or are located and oriented. However, the device may alternatively be configurable for different loudspeaker arrangements of the loudspeakers 14. Similarly, the number of loudspeakers 14 may be two or more, and the device may be designed for a set number of loudspeakers 14 or may be configurable to accommodate any number of loudspeakers 14.

[0024] The device 10 includes an interface 16 through which the device 10 receives an audio signal 18 representing at least one audio object. The device 10 can be configured to decode the audio signal 18 from, for example, a bitstream. For now, assume that the audio input signal 18 is a mono audio signal representing an audio object such as the sound of a helicopter. Additional examples and further details are given below. Alternatively, the audio input signal 18 can be a stereo audio signal or a multi-channel audio signal. In any case, the audio signal 18 may represent an audio object in the time domain, the frequency domain, or any other domain, and it may represent the audio object in a compressed format or without compression.

[0025] As shown in FIG. 1, the apparatus 10 further comprises an object position input 20 for receiving an intended virtual position 21. That is, at the object position input 20, the apparatus 10 is notified of the intended virtual position 21 at which an audio object is to be virtually rendered by the application of the loudspeaker signal 12 at the loudspeaker 14. That is, the apparatus 10 receives information on the intended virtual position 21 at the input 20, and this information can be given relative to the arrangement / position of the loudspeakers 14, relative to the sweet spot, relative to the position and / or the head orientation of the listener, and / or relative to real-world coordinates. This information can be based, for example, on a Cartesian or polar coordinate system. It can be based on a room-centered or listener-centered coordinate system, for example, as either a Cartesian or polar coordinate system.

[0026] In addition, the apparatus 10 comprises a listener position input 30 for receiving the actual position of the listener. The listener position 31 can be defined, for example, by coordinates indicating the position of the listener within the playback space, such as the position of the listener's body, the position of the listener's head, or the position of the listener's ear, for example, by tracking data, i.e., information on the position of the listener over time. The listener position 31 can be described, for example, in Cartesian, spherical, or cylindrical coordinates. Instead of the absolute position of the listener, the listener position 31 can indicate, for example, the relative position of the listener with respect to a reference loudspeaker of a set of loudspeakers, with respect to the sweet spot within the playback space, or with respect to any other predetermined position within the playback space.

[0027] For example, when the intended virtual position 21 defines the relative position of the audio object with respect to the listener position 31, the apparatus 10 may not necessarily require the listener position input 30 for receiving the listener position 31. This is due to the fact that the intended virtual position 21 already takes into account the listener position 31.

[0028] As shown in FIG. 1, the apparatus 10 may include a gain determiner 40 configured to determine a gain 41 for a plurality of loudspeakers 14 in accordance with an intended virtual position 21 received at input 20 and / or a listener position 31 received at input 30. According to one embodiment, the gain determiner 40 may calculate an amplitude gain for each loudspeaker signal one by one such that the intended virtual position 21 is panned among the plurality of loudspeakers 14 and / or such that the roll-off of sound energy is compensated. As will be described with respect to FIG. 3, the gain 41 provided by the gain determiner 40 may represent a compensation gain. Alternatively, as will be described in more detail with respect to FIG. 2, each panning gain g n is the horizontal component

[0029]

Number

[0030] and the vertical component

[0031]

Number

[0032] , for example

[0033]

Number

[0034] and, optionally, a further component corresponding to the compensation gain (see FIG. 3). The index n represents a positive integer in the range 1 ≦ n ≦ i, where i represents the number of loudspeakers 14. The gain determiner 40 may be configured to determine a respective gain 41 for each loudspeaker.

[0035] Additionally, or alternatively, the apparatus 10 may comprise a delay determiner / controller 50 for determining / controlling a delay 51 for a plurality of loudspeakers 14 in accordance with the intended virtual position 21 received at the input 20 and / or the listener position 31 received at the input 30. The delay determiner 50 may be configured to determine a respective delay 51 for each loudspeaker such that at least one audio object is rendered at the virtual position where the application of the loudspeaker signal 12 to or at the plurality of loudspeakers 14 is intended and / or such that the loudspeaker signals reproduced by the loudspeakers 14 reach the listener simultaneously.

[0036] The apparatus 10 may comprise an audio renderer 11 configured to render the audio signal 18 based on a gain 41 and / or a delay 51 in order to derive the loudspeaker signal 12 from the audio signal 18.

[0037] With respect to FIG. 2, possible 3D panning implemented by the panning gain determiner 40 is described in more detail.

[0038] The loudspeakers 14 may be arranged in one or more horizontal layers 15. As shown in FIG. 2, a first set 141 to 145 of loudspeakers of the plurality of loudspeakers 14 may be arranged in a first horizontal layer 151, and a second set 146 to 148 of loudspeakers of the plurality of loudspeakers 14 may be arranged in a second horizontal layer 152. That is, the first set 141 to 145 of loudspeakers are arranged at an apparently similar height, and the second set 146 to 148 of loudspeakers are arranged at an apparently similar height. The first set 141 to 145 of loudspeakers may be arranged at or near a first height, and the second set 146 to 148 of loudspeakers may be arranged at or near, for example, a second height above the first height. According to the embodiment shown in FIG. 2, the listener position 31 is exemplarily arranged within the first horizontal layer 151.

[0039] In the following, an example case of rendering an object in 3D is described, where an object 1041, for example a sound source, is panned in a direction (as seen from listener 1) between two physically existing loudspeaker layers (which are at different heights). The object 1041 is amplitude panned in the first layer 151 by applying an object signal to the loudspeakers in this layer with different first layer horizontal gains, for example by applying the object signal from loudspeaker 141 to 145 such that the object 1041 is amplitude panned to the lower layer, i.e., the first layer 151 (see the panned first layer position 104'1 in FIG. 2). In this horizontal panning, for each loudspeaker of the first set 141 to 145 of loudspeakers, the horizontal component of the respective panning gain 41

[0040]

Number

[0041] is determined. Similarly, the object 1041 is amplitude panned in the second layer 152 to the panned second layer position 104''1 in FIG. 2. In this horizontal panning, for each loudspeaker of the second set 146 to 148 of loudspeakers, the horizontal component of the respective panning gain 41

[0042]

Number

[0043] is determined. As can be understood, positions 104'1 and 104''1 can be selected such that they overlap perpendicularly to each other and / or such that the vertical projections of the intended position 1041 and positions 104'1 and 104''1 match well. FIG. 2 shows rendering the final object position 1041 by applying amplitude panning between layers 15, i.e., vertical panning. Considering the virtual objects at positions 104'1 and 104''1 as virtual loudspeakers, the amplitude panning by the gain determiner 40 is applied to render the virtual object at the intended position 1041 between the two layers 151 and 152. In this vertical panning, for example, for each loudspeaker of the first set 141 to 145 of loudspeakers and the second set 146 to 148 of loudspeakers, the vertical component of the respective panning gain 41

[0044] [Number]

[0045] is determined. The result of this amplitude panning between layers 151 and 152 is two gain coefficients for each loudspeaker by which the respective loudspeaker signals are weighted for them, i.e., the horizontal component

[0046] [Number]

[0047] and the vertical component

[0048] [Number]

[0049] Therefore, for example, the sound source of the audio signal is panned to the sound source position of the desired audio signal. In addition, this weighting for horizontal panning between the (actual) loudspeaker layers 15 can be frequency-dependent in order to compensate for the effect that different frequency ranges can be perceived at different heights in vertical panning.

[0050] In the following, an example case of rendering an object in 3D will be described for an exemplary case where the object 1042 is panned above or below the outermost layer. The object can have a direction or position 1042 that is not within the range of directions between the two layers 151 and 152, as discussed with respect to the object position 1041. The intended position 1042 of the object is, for example, above or below the (physically existing) layer 15, here below all available layers, specifically below the bottom layer, i.e., below the first layer 151. As an example, the object has a direction / position 1042 below the bottom loudspeaker layer of the loudspeaker setup used as an exemplary setup in FIG. 2, i.e., below the first layer 151. In this case, horizontal amplitude panning is applied by the bottom layer panning gain determiner 40 to render the object 1042 in that layer 151 (see the obtained position 104'2). The obtained position 104'2 can represent a virtual sound source position corresponding to the projection of the sound source position of the desired audio signal (see 1042) onto the nearest loudspeaker layer (see 151). More generally, 2D amplitude panning is applied between the loudspeakers 141 to 145 belonging to the loudspeaker layer closest to the object 1042, i.e., the first layer 151. In this horizontal panning, for example, for each loudspeaker of the first set 141 to 145 of loudspeakers, the horizontal component of the respective panning gain 41

[0051]

Number

[0052] is determined. Then, spectral shaping of the audio signal to provide a rendition of the sound by the loudspeakers 141 to 145 of the nearest loudspeaker layer, i.e., the first layer 151, is performed, and further amplitude panning is applied between the loudspeakers 141 to 145 belonging to the nearest loudspeaker layer, i.e., the first layer 151. This rendition of the sound simulates sound from a further virtual sound source position 104''2 offset from the nearest loudspeaker layer, i.e., the first layer 151, towards the sound source position (see 1042) of the desired audio signal. Since there are no actual loudspeakers in the upper and lower vertical directions, the vertical signal at 104''2 can be equalized to simulate the upper or lower sound coloring, respectively. The vertical signal is then provided to the loudspeakers designated for the up / down direction. To render the final object position 1042, the panning gain determiner 40 applies further amplitude panning between the virtual sound source position 104'2 and the further virtual sound source position 104''2, determines a second panning gain for panning between the virtual sound source position 104'2 and the further virtual sound source position 104''2, and may be configured to provide a rendering of the audio signal by the loudspeakers 141 to 145 of the nearest loudspeaker layer from the sound source position 1042 of the desired audio signal. The spectral shaping of the audio signal may be performed using a first equalization function that simulates the timbre of the lower sound when the sound source position of the desired audio signal is located below one or more loudspeaker layers, i.e., below the first layer 151, and / or the spectral shaping of the audio signal may be performed using a second equalization function that simulates the timbre of the upper sound when the sound source position of the desired audio signal is located above one or more loudspeaker layers, i.e., above the second layer 152.

[0053] Figure 3 shows an embodiment of an audio processor 10 (see audio renderer 11) for performing audio rendering by generating rendering parameter 100, which determines the derivation of an audio signal 18 of a loudspeaker signal 12 that is to be reproduced by a set of loudspeakers 14. The focus of the embodiment shown in Figure 3 is on the gain determiner 40. Optionally, as described with respect to Figure 1, the gain determiner 40 may be combined with a delay determiner 50. The embodiment shown in Figure 3 provides details regarding the determination of a compensation gain 41 using the gain determiner 40. The compensation gain 41 may represent the gain provided by the gain determiner shown in Figure 1. Alternatively, as described with respect to Figure 1, each compensation gain 41 may represent each component of each gain that is to be applied to each loudspeaker.

[0054] The gain determiner 40 is configured to perform gain adjustment to determine a gain 41 for generating the loudspeaker signal 12 for the loudspeaker 14 from the audio signal 18 based on the listener position 31. For example, gain adjustment regarding adjusting the gain associated with the conditions of an anechoic environment so that the effects of reverberation are taken into account. Thus, the gain 41 determined by the gain determiner 40 is more suitable for the real-world sound reproduction environment.

[0055] As shown in Figure 3, the gain determiner 40 of the audio processor 10 acquires reverberation effect information 110. The reverberation effect information 110 may indicate whether reverberation is effective in the reproduction space 112 and / or the reverberation conditions in the reproduction space 112. The audio processor 10 may be configured to derive the reverberation effect information 110 from a bitstream or side information of the bitstream.

[0056] The audio processor 10 is configured to use a roll-off gain compensation function 42 for associating, in gain adjustment, the listener-loudspeaker distance 44 for at least one loudspeaker 14 with a listener-loudspeaker distance compensation gain 46 for at least one loudspeaker 14 in accordance with the reverberation effect information 110. For the roll-off gain compensation function, the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance 44 increases (see also FIGS. 4 and 6). The audio processor 10 is configured to determine, adapt, or select the roll-off gain compensation function 42, for example, in accordance with the reverberation effect information 110. The listener-loudspeaker distance compensation gain 46 determined for at least one loudspeaker 14 may represent the gain 41 provided to the audio renderer 11 for deriving each loudspeaker signal 12 that will be reproduced by each respective loudspeaker 14 from the audio signal 18.

[0057] The listener position 31 may indicate the listener-loudspeaker distance 44 for at least one loudspeaker 14 for which gain adjustment is used. Alternatively, the listener position 31 may comprise the listener-loudspeaker distance 44 for each loudspeaker 14 of a set of loudspeakers 14. Alternatively, it is also possible for the listener position 31 to indicate the absolute position of the listener 1 within the reproduction space 112. In this case, the audio processor 10 may be configured to additionally obtain information about the position of at least one loudspeaker 14, or the positions of all loudspeakers 14, within the reproduction space 112 for which gain adjustment is used. The audio processor 10 may be configured to determine the respective listener-loudspeaker distance for at least one loudspeaker 14 for which gain adjustment is used based on the listener position 31 and the position of each respective loudspeaker 14.

[0058] The roll-off gain compensation function 42 used by the gain determiner 40 is shown in FIG. 4 (42β1 and 42 β2 (see reference) and will be described in more detail with respect to FIG. 6.

[0059] FIG. 4 schematically shows roll-off gain compensation functions for two different near-field - far-field transition parameters beta (42 β1 and 42 β2 (see reference). The larger beta is, the faster the transition between near-field attenuation and far-field attenuation becomes. The near-field - far-field transition parameter beta may be included in the reverberation effect information 110 discussed herein.

[0060] The roll-off gain compensation function 42 β1 and 42 β2 both have the same critical distance 44 12 illustrated, for example, for a distance of 4 meters to a relevant loudspeaker, i.e., a loudspeaker to which the roll-off gain compensation function may be applied. The critical distance 44 12 can shift the roll-off gain compensation function along the axis of the listener-loudspeaker distance (see 44). The greater the amount of effective reverberation in the playback space, the shorter the critical distance 44 12 becomes. FIG. 6 illustratively shows the roll-off gain compensation function for a critical distance 44 of 2 meters. The critical distance 44 12 can be included in the reverberation effect information 110 discussed herein. The critical distance 44 12 can represent the distance at which the energy of the direct sound becomes equal to the energy of the reverberant sound. 12

[0061] Furthermore, a near-field roll-off gain compensation function 43 nf and a far-field roll-off gain compensation function 43 ff are shown. FIG. 4 shows the compensation gain 46 and the listener-loudspeaker distance 44.

[0062] The reverberation effect information 110 may indicate that the sound decays more slowly as the distance to the loudspeaker 14 increases. For example, near each loudspeaker 14, i.e., in the near field (see 441), the sound energy rolls off faster than when it is away from each loudspeaker 14, i.e., than in the non - near field (see 442). The reverberation effect information 110 may comprise a near - field attenuation parameter and a non - near - field attenuation parameter. For example, refer to decay_1_dB and decay_2_dB in FIGS. 10c and 10i. The near - field roll - off gain compensation function 43 nf indicates a compensation gain 46 for compensating the roll - off, i.e., the roll - off of the sound energy, according to the near - field attenuation parameter, and the non - near - field roll - off gain compensation function 43 ff indicates a compensation gain 46 for compensating the roll - off, i.e., the roll - off of the sound energy, according to the non - near - field attenuation parameter. In the determination of the roll - off gain compensation functions (42 β1 and 42 β2 referenced), both the near - field attenuation parameter and the non - near - field attenuation parameter are considered. The roll - off gain compensation functions (42 β1 and 42 β2 referenced) schematically show the total compensation gain 46 for compensating the roll - off of the sound energy over the listener - to - loudspeaker distance 44 in the near field and the non - near field. As can be seen in FIG. 4, the (e.g., total) roll - off gain compensation functions (42 β1 and 42 β2 referenced) transition between the near - field attenuation and the non - near - field attenuation.

[0063] The roll - off gain compensation functions (42 β1 and 42 β2 referenced) indicate a listener - to - loudspeaker distance compensation gain 46, which is to be applied to the loudspeaker signal 12 to compensate for the reverberation - dependent roll - off of the sound energy. As shown in FIG. 4, the roll - off gain compensation functions (42 β1 and 42 β2The reference) is configured such that as the listener - loudspeaker distance 44 increases, the listener - loudspeaker distance compensation gain 46 increases more slowly, that is, as the listener - loudspeaker distance 44 increases, the roll - off gain compensation function 42 becomes monotonically shallower, for example, as the listener - loudspeaker distance 44 increases, the change in compensation gain per unit distance decreases.

[0064] Roll - off gain compensation function (42 β1 and 42 β2 The reference), for example, within the first distance region 441, for example in the near - field, has a first slope (42' β1 and 42' β2 The reference), for example, having a first compensated roll - off slope, and within the second distance region 442, for example in the non - near - field, has a second slope (42'' β1 and 42'' β2 The reference), for example, having a second compensated roll - off slope, where the first slope 421 is greater than the second slope 422, and the first distance region 441 is related to a shorter distance than the second distance region 442. The first slope 421 and / or the second slope 422 can be indicated by the reverberation effect information 110. The reverberation effect information 110 can further indicate a boundary distance, for example a critical distance 44 that separates the first distance region 441 and the second distance region 442. 12 The boundary distance 44 12 can correspond to the distance to the loudspeaker such that the energy of the direct sound becomes equal to the energy of the reverberant sound within the reproduction space 112.

[0065] According to an embodiment, the reverberation effect information 110 can indicate, for the roll - off gain compensation function 42, how the roll - off gain compensation function 42 must transition from the first distance region 441 to the second distance region 442, for example using a near - field - non - near - field transition parameter beta. FIG. 4 shows a roll - off gain compensation function 42 β2 compared to a roll - off gain compensation function 42 with a slower transition β1is exemplarily shown. By being able to consider the specific transition of the reproduction space 112 between the near-field sound energy attenuation and the non-near-field sound energy attenuation, the accuracy of the determination of the compensation gain 41 can be enhanced.

[0066] The audio processor 10 is configured to perform a gain adjustment such that the listener position 31 becomes a sweet spot with respect to the set of loudspeakers 14 in an acoustic or perceptual sense, that is, such that the listener 1 perceives the sound reproduced by the set of loudspeakers 14 as intended by the mixer. Artifacts that may be perceivable by the listener 1 at the position of the listener 1 are reduced by this special gain adjustment.

[0067] Hereinafter, the relationship between the reverberation effect information 110 using the roll-off gain compensation function 42 and the gain adjustment will be described in more detail with respect to FIGS. 3 and 4.

[0068] The reverberation effect information 110 may indicate the amount of effective reverberation in the reproduction chamber, that is, the reproduction space 112, that is, it may indicate how much sound or signal in the reproduction space 112 is reflected, for example, by walls or furniture. The amount of effective reverberation in the reproduction space 112 may indicate how many reflections increase and then decay as the sound is absorbed, for example, by the surface of an object / wall in the reproduction space 112. In this case, the audio processor 10 selects a roll-off gain compensation function (42 in FIG. 3 and 42 β1 and 42 β2 (hereinafter generally referred to using the reference numeral 42 for reference) or may be configured to adapt the roll-off gain compensation function 42 to obtain the roll-off gain compensation function 42, for which the greater the amount of effective reverberation in the reproduction space 112, the stronger the degree to which the compensated roll-off monotonically becomes shallower as the listener-loudspeaker distance 44 increases. FIG. 4 shows the roll-off gain compensation function 42 β2 for the reproduction space 112 with a greater amount of effective reverberation compared to the reproduction space 112 related to the roll-off gain compensation function 42 β1is shown by way of example. The roll-off gain compensation function 42 β1 compensates for the roll-off of sound energy for a playback space 112 where the amount of effective reverberation is less and the sound energy does not roll off or attenuate as quickly as in a playback space 112 with a smaller amount of effective reverberation, and the roll-off gain compensation function 42 for a playback space 112 with a smaller amount of effective reverberation β2 can be used for comparison. Therefore, the roll-off gain compensation function should be adapted or selected by the audio processor 10 such that the greater the amount of effective reverberation in the playback space 112, the smaller the gain (critical distance 44 12 reference), and the listener-speaker distance compensation gain 46 begins to increase more slowly as the listener-speaker distance 44 increases. This is based on the recognition that an increase in the amount of effective reverberation in the playback space 112 reduces the roll-off of sound energy. This adaptation of the roll-off gain compensation function 42 makes it possible to increase the accuracy in determining the compensation gain 41.

[0069] The reverberation effect information 110 can indicate whether the reverberation is effective in the playback space 112. The roll-off gain compensation function 42 described herein, which increases monotonically shallower / slower as the listener-speaker distance 44 increases, can only be used if the reverberation is effective in the playback space 112. If the reverberation effect information 110 indicates that the reverberation is not effective in the playback space 112, the audio processor 10 may be configured to use an additional roll-off gain compensation function where the roll-off to be compensated is constant, for example, in this case, the near-field roll-off gain compensation function 43 nfIt can be used. For example, a further roll-off gain compensation function can be configured to compensate for a predetermined roll-off of acoustic energy, such as 6 dB, every time the listener-loudspeaker distance 44 doubles. Reverberation can result in attenuation of sound energy in the near field of the loudspeaker 14 that is different compared to the non-near field of the loudspeaker. However, if the reverberation is not effective in the playback space 112, there is no need to consider this distinction between the near field and the non-near field. Thus, in such cases, a simpler determination of the compensation gain can be implemented. This enables the determination of the compensation gain for different playback spaces 112 efficiently and with low complexity.

[0070] The concept behind the embodiments of the present invention will be continued to be described. Specifically, the playback chamber 112 has reverberant energy, and thus distance gain compensation (the roll-off gain compensation function 42 in FIG. 3 and 42 in FIG. 4 β1 and 42 β2 reference) is provided that takes into account the fact that the acoustic energy rolls off more slowly as the distance between the loudspeaker position and the listener 1 increases. That is, gain / level adjustment is performed taking into account information about the amount of reverberation present in the playback chamber 112. As an example, the theoretical roll-off ("slope") of acoustic energy over distance is 6 dB every time the distance from a point source doubles. By taking into account the reverberation in the room, the strength (slope) of the gain compensation for the rendering of the loudspeaker adapted to the user becomes shallower (more gradual) as the distance increases (see FIGS. 4 and 6). One parameter for defining this change in roll-off is the so-called "critical distance" or boundary distance 44 12 known in acoustics as the distance at which the energy of the direct sound equals the energy of the reverberant sound. In this rendering method of the loudspeaker adapted to the user, the control parameter related to the critical distance 44 12 is very effective in controlling the appropriate compensation characteristics.

[0071] Thus, according to one embodiment, the above idea results in an audio signal processor 10 as follows. · The value of the slope of the gain compensation for at least one loudspeaker signal depends on the position of the listener / the distance 44 of the listener from this loudspeaker · Optionally, the delay can also be adjusted according to data on reverberation in the playback environment chamber 112 · The slope is smaller (shallower) at longer distances than at shorter distances · Different slope values or ranges of slope values are applied, and there are at least two distance regions 441 and 442 where the slope value of the near (first) region 441 is greater than the slope of the far (second) region 442 · The "critical distance" 44 12 Parameters regarding are used to define the boundary between the near (first) region and the far (second) region · The slope value of the near (first) region 441 is steeper than the slope value of the far (second) region 442 · The slope parameter of the near (first) region 441 is used / accepted / applied in this region · The slope parameter of the far (second) region 442 is used / accepted / applied in this region · Optionally, a transition parameter that determines the transition (e.g., roundness) between these two regions 441 and 442 is defined and applied to the roll-off gain compensation function 42

[0072] One embodiment of the present invention relates to an audio processor 10 configured to generate a set of one or more parameters (which can be, for example, rendering parameters 100 that can affect the delay, level, or frequency response of one or more audio signals) for each of a set of one or more loudspeakers 14, which is based on a listener position 31 (the listener position 31 can be, for example, the position of the entire body of the listener 1 in the same room as the set of one or more loudspeakers 14, i.e., the playback space 112, or, for example, only the position of the head of the listener 1, or even, for example, the position of the ears of the listener 1. The listener position 31 can be, for example, a position relative to the set of one or more loudspeakers 14, such as the distance from the listener's head to the set of one or more loudspeakers 14) and the loudspeaker positions of the set of one or more loudspeakers 14, and determines the derivation of the audio signal 18 of the loudspeaker signal 12 to be reproduced by each loudspeaker 14. The audio processor 10 is configured to make the generation of the set of one or more parameters of the set of one or more loudspeakers 14 based on information about the reverberation characteristics of the playback environment (room), i.e., the reverberation effect information 110. Specifically, the calculation of the level (gain 41) value for the loudspeaker signal 12 is based on information about the level of the reverberant sound present in the playback room 112.

[0073] Taking this information about the level of the reverberant sound into account, the present invention achieves an improved rendering result by utilizing the strength (slope) of the level (gain 41) compensation for user-adaptive loudspeaker rendering that becomes shallower (more gradual) with an increase in the distance, i.e., the listener-loudspeaker distance 44. One important parameter for defining this change in the distance-dependent slope can be related to the so-called "critical distance" (44 12 reference). The "critical distance" 44 12The term is known in acoustics as the distance at which the energy of the direct sound radiated from the sound source becomes equal to the energy of the reverberant sound [4]. In the rendering method of the loudspeaker adapted to the user of the present invention, the critical distance 44 12 The control parameters related to are known to be very effective in controlling appropriate compensation characteristics. Furthermore, for the listener position 31 clearly below the critical distance 44 12 a tilt value, and for the listener position 31 clearly above the critical distance 44 12 a tilt value can be defined and used.

[0074] This can be realized by the audio processor 10. The audio processor 10 obtains information about, for example, the positioning of the listener, i.e., the listener position 31, the positioning of the loudspeaker, i.e., the loudspeaker position, and, for example, the critical distance of the room, the proximity tilt parameter (for example, indicating the first tilt 421), or the far tilt parameter (for example, indicating the second tilt 422), etc., the reverberation characteristics of the playback room, i.e., the reverberation effect information 110. The audio processor 10 can calculate a set of one or more parameters from this information. Using the set of one or more parameters, the input audio, in other words the incoming audio signal 18, can be modified. By this modification of the audio signal 18, the listener 1 receives an optimized audio signal at his position. With this optimized signal, the listener 1 can obtain, for example, an almost or exactly the same auditory sensation at his position as if he were at the ideal listening position of the listener. The ideal listening position is, for example, a position such as a sweet spot where the listener experiences optimal audio perception without modification of the audio signal. This means, for example, that the listener 1 can perceive the audio scene in the manner intended at the production site at this position. The ideal listening position may correspond to a position equidistant from all the loudspeakers 14 (one or more loudspeakers 14) used for playback.

[0075] Accordingly, the audio processor 10 according to the present invention enables listener 1 to change his or her position to different listener positions 31 and obtain the same or at least partially the same listening sensation as when the listener is in an ideal listening position at each of at least some positions.

[0076] In summary, it should be noted that the audio processor 10 is capable of adjusting at least one of the delay, level, or frequency response of one or more audio signals 18 based on the positioning of the listener, the positioning of the loudspeaker, and / or the characteristics of the loudspeaker, for the purpose of achieving optimized audio reproduction for at least one listener 1. The level is adjusted in response to information about the reverberation characteristics 110 of the playback room 112.

[0077] Embodiments of the present invention will now be described with respect to the rendering of adaptive loudspeakers.

[0078] First, some general remarks will be made. Instead of rendering the MPEG-I scene for headphones and binauralizing it, playback via loudspeakers is defined. In this mode of operation, the MPEG-I Spatializer (HRTF-based renderer) is replaced by a dedicated loudspeaker-based renderer described below.

[0079] For a high-quality listening experience, the loudspeaker setup assumes that listener 1 is in a dedicated fixed position, a so-called sweet spot. Usually, in a 6DOF playback situation, listener 1 is moving. Therefore, 3D spatial rendering must adapt instantaneously and continuously to the changing listener position 31. This can be achieved at two hierarchically nested technical levels. 1. The loudspeaker signal 12 is applied with the same gain and delay so as to reach the listener position 31, i.e., the listener position 31 is at the sweet spot. For example, the gain 41 and the delay 51 are applied to the loudspeaker signal 12. Optionally, a high-shelving compensation filter is applied to each loudspeaker signal 12 with respect to the current listener position 31 and the orientation of the loudspeakers with respect to the listener 1. Thus, as the listener 1 moves away from the axis of the loudspeaker 14 or even further away from the loudspeaker 14, the high-frequency loss due to the radiation high-frequency pattern of the loudspeaker is compensated. 2. Due to the 6DoF movement, the angles between the loudspeaker 14, the object, and the listener 1 change as a function of the listener position 31. Thus, for example, a 3D amplitude panning algorithm (see Figure 2) is updated in real time with the changing listener position 31 and the relative positions and angles with respect to a fixed loudspeaker configuration as set in the LSDF. All coordinates (listener position 31, sound source position) can be transformed into the coordinate system of the listening room, i.e., the coordinate system of the playback space 112.

[0080] Physical compensation level (level 1) Figure 5 shows an overview of an embodiment of the level 1 system 10 together with its main components and parameters. The audio processor 10 described with respect to FIGS. 1 to 4 may comprise the features and / or functions described with respect to the embodiment of FIG. 5.

[0081] Level 1: Real-time updated compensation of (frequency-dependent) loudspeaker gain and delay (see Audio Renderer 11) enables "enhanced rendering of content". By utilizing the tracked user location information, e.g., a version with the listener position 31, the listener 1, i.e., the user, can move within a wide "sweet area" (rather than a sweet spot) and experience a stable sound stage within this wide area when listening to, e.g., legacy content (e.g., stereo, 5.1, 7.1+4H). In an immersive format (i.e., not a format for stereo), the sound is not felt to pour into the nearest speaker 14 when walking away from the sweet spot, but rather to move away from the loudspeaker 14, i.e., this has somewhat similar properties to what is known as wave field synthesis, but is for a single-user experience. In stereo playback, this technology provides left-right sound stage stability for a wide range (i.e., the range between the left and right loudspeakers at any distance) of user positions 31.

[0082] The gain compensation at Level 1 is based on, for example, the amplitude decay law. In a free field, the amplitude is proportional to 1 / r, where r is the distance from the listener 1 to the loudspeaker 14 (1 / r corresponds to a 6 dB attenuation every time the distance doubles). In the room 112, due to the presence of acoustic reflections and reverberation, the sound decays more slowly as the distance to the loudspeaker 14 increases. Thus, parameters such as near-field attenuation, non-near-field attenuation, and / or critical distance, which are included in, for example, the reverberation effect information 110, can be used to specify the attenuation rate as a function of the distance to the loudspeaker 14. Additionally, there can be a near-field - non-near-field transition parameter beta, which is included in, for example, the reverberation effect information 110. The larger the beta, the faster the transition between near-field attenuation and non-near-field attenuation. FIG. 6 shows an example of gain compensation as a function of distance, i.e., the roll-off gain compensation function 42 that can be used by the gain determiner 40. In the reverberant field, the gain change is smaller than in the free field.

[0083] The delay compensation at level 1 calculates, for example, the propagation delay from each loudspeaker 14 to the listener position 31 and then applies a delay to each loudspeaker 14 to compensate for the difference in propagation delay between the loudspeakers 14. The delay can be normalized (an offset is added or subtracted) so that the minimum delay applied to the loudspeaker signal 12 becomes zero.

[0084] Object rendering level (level 2) Level 2: Object panning tracked by the user enables the rendering of point sources (objects, channels) within the 6DoF playback space and requires level 1 as essential. Thus, it addresses the use case of "6DoF VR / AR rendering". The following features and / or functions may additionally be included in the level 1 system 10.

[0085] For example, as described with respect to FIG. 2, a 3D amplitude panning algorithm that functions in loudspeaker layers, such as horizontal and height layers, may be used. Each layer may apply a 2D panning algorithm for the projection of the object onto the layer. The final 3D object is rendered by applying amplitude panning between two virtual objects from the 2D panning in the two layers.

[0086] When the object is located above the highest layer, 2D panning is applied in that layer. The final 3D object is rendered by applying amplitude panning between the virtual object from the 2D panning and an object (nonexistent) in the vertical direction above. The signal of the vertical object may be equalized to simulate the timbre of the sound above and equally distributed to the loudspeakers of the highest layer.

[0087] When the object is positioned below the bottom layer, 2D panning is applied in that layer. The final 3D object is rendered by applying amplitude panning between the virtual object from the 2D panning and the (non-existent) object in the vertical direction below. The signal of the vertical object can be equalized to simulate the timbre of the sound below and distributed equally to the loudspeakers of the bottom layer.

[0088] The vertical panning as described is equally applicable to loudspeaker setups with one layer such as 5.1 and loudspeaker setups with multiple layers such as 7.4.6.

[0089] Levels 1 and 2 applied to object rendering render the MPEG-I scene faithfully, as in the case of via headphones. This has a great benefit compared to loudspeakers that render MPEG-I content without applying adaptive tracking (1 and 2).

[0090] Physical compensation level (Level 1) In the following, embodiments of gain and delay adjustment based on the listener position are described using code snippets (see FIGS. 10c to 10i and FIGS. 11b to 11c). The features and / or functions described below regarding gain and / or delay adjustment may be included in the audio processor 10 of FIG. 1 or the Level 1 system 10 of FIG. 5. The audio processor 10 of FIG. 3 may additionally have the features and / or functions described below regarding gain adjustment. Optionally, the audio processor 10 of FIG. 3 may have the features and / or functions described below regarding gain and / or delay adjustment. Optionally, the audio processor 10 of FIG. 1, the audio processor 10 of FIG. 3, and the audio processor 10 of FIG. 3 may have further features and / or functions as described below.

[0091] Data elements and variables Definitions and / or explanations of data elements and variables used below (see FIGS. 7 to 11c) are provided. SFREQ_MIN Minimum sample rate [Hz] = 44100 SFREQ_MAX Maximum sample rate [Hz] = 48000 VSOUND Speed of sound in air [m / s] = 340.0 MAX_DELAY Maximum delay [samples] = 960 OVERHEAD_GAIN Overhead [lin] = 0.25 framesize Number of samples per frame, default: 256 sfreq_Hz Sampling frequency of the input audio, default: 48000 nchan Number of channels (loudspeakers) max_delay Maximum delay [samples], default: MAX_DELAY bypass_on 0: Normal operation, 1: Bypass, default: 0 ref_proc 0: Normal operation, 1: Processing for sweet spot, default: 0 cal_system 0: Normal operation, 1: Calibrated system, default: 0 gain_on 0: Gain off, 1: On, default: 1 delay_on 0: Delay off, 1: On, default: 1 decay_1_dB Proximity field sound attenuation per doubling of distance [dB], default: 8 decay_2_dB Non - proximity field sound attenuation per doubling of distance [dB], default: 0 beta 1: Default proximity - non - proximity transition, > 1: Faster transition crit_dist_m Critical distance [m], default: 4 max_m_s Maximum moving speed [v in m / s], default: 1 max_m_s_s Maximum moving acceleration [a in m / s], default: 1 gain_ms Gain smoothing time constant [ms], default: 40 sweet_spot sweet spot position [m, m, m] spk_pos speaker coordinates [m, m, m] listener_pos listener coordinates [m, m, m]

[0092] All coordinates are relative to the listening room as defined, for example, in an LSDF file.

[0093] These parameters can be stored in the following structure.

[0094] Public data structure typedef struct rendering_gd_cfg { int framesize; float sfreq_Hz; int nchan; float max_delay; } rendering_gd_cfg_t; typedef struct rendering_gd_rt_cfg { int bypass_on; int ref_proc; int cal_system; int gain_on; int delay_on; float decay_1_dB; float decay_2_dB; float crit_dist_m; float beta; float max_m_s; float max_m_s_s; float gain_ms; float sweet_spot[3]; float spk_pos[NCHANMAX][3]; float listener_pos[3]; } rendering_gd_rt_cfg_t;

[0095] Internal parameters calculated from the parameters and states enumerated above are stored, for example, in the following structure.

[0096] Internal data structure typedef struct { / * Static parameters * / float sfreq_Hz; int nchan; int framesize; / * Real-time parameters * / int bypass_on; int gain_on; float delta_gi; float delta_gd; float gain_alpha; float delay_delta; float delay_delta2; / * States * / float delay0[NCHANMAX]; float delay[NCHANMAX]; float gain0[NCHANMAX]; float gain[NCHANMAX]; } rendering_gd_data_t;

[0097] Step description Embodiments of gain and delay adjustment based on listener position are described below using code snippets related to different stages. This embodiment may include an initialization stage (see FIG. 7), a release stage (see FIG. 8), a reset stage (see FIG. 9), a real-time parameter update stage (see FIGS. 10a to 10i), and an audio processing stage (see FIGS. 11a to 11c). The audio processor 10 of FIG. 1, the level 1 system 10 of FIG. 5, and the audio processor 10 of FIG. 3 may include features and / or functions described with respect to one or more of the stages, or individual features and / or functions of one or more of the stages.

[0098] Initialization FIG. 7 illustratively shows a code snippet of the initialization stage.

[0099] The loudspeaker setup can be loaded from the LSDF file.

[0100] The structure of type rendering_gd_cfg_t is initialized with default values, and the nchan field is set to the number of loudspeakers in the loudspeaker setup.

[0101] The structure of type rendering_gd_rt_cfg_t is initialized with default values. The loudspeaker positions from the LSDF file are stored in the field spk_pos. If a ReferencePoint element is given in the LSDF file, its coordinates are stored in the field sweet_spot. If present, the field cal_system is set to the value of the calibrated attribute.

[0102] The above structures are passed to the rendering_gd_init function.

[0103] Release FIG. 8 illustratively shows a code snippet of the release stage.

[0104] Reset FIG. 9 exemplarily shows a code snippet in the reset stage. FIG. 9 shows that all internal buffers are flushed.

[0105] Update real-time parameters In the update thread, the virtual listening position is converted to the coordinate system of the listening room. This is only important for the VR scene, and in the AR scene, these two coordinate systems coincide.

[0106] All further processing occurs in the audio thread.

[0107] The structure of type rendering_gd_rt_cfg_t is updated by setting the listener_pos field to the listener position (in the coordinate system of the listening room) (see FIG. 10a). Then, this structure is passed to the rendering_gd_updatecfg function (see FIG. 10a).

[0108] For each loudspeaker, the compensation gain and delay are calculated. The reference distance r_ref (calculated in FIG. 10a) is the distance at which the gain and delay compensation are 0 (dB, samples). Based on the distance between the loudspeaker and the listener r and the reference distance r_ref, the gain and delay compensation are calculated. The calculation of the listener-loudspeaker distance 44 based on the listener position 31 and each loudspeaker position 32 is shown in FIG. 10b. The listener-loudspeaker distance 44 can represent a version with the listener position 31.

[0109] In a free field, sound attenuates by 6 dB every time the distance doubles. In a room, the attenuation can be approximated by using a smaller attenuation, for example 4 dB every time the distance doubles. Alternatively, the critical distance (hall range) can be considered. When close to the loudspeaker, the attenuation is decay_dB every time the distance doubles. Beyond the critical distance crit_dis_m, the sound only attenuates slowly. To determine the gain compensation that compensates for the gain change due to the described sound attenuation, it is proposed to use the roll-off gain compensation function 42 (see, for example, FIGS. 6, 10c, and 10i).

[0110] The gain compensation can be based on the amplitude attenuation law. In a free field, the amplitude is proportional to 1 / r, where r is the distance from the listener to the loudspeaker (1 / r corresponds to an attenuation of 6 dB every time the distance doubles). In a room, due to the presence of acoustic reflections and reverberation, the sound attenuates more slowly as the distance to the loudspeaker increases. Therefore, the parameters of near-field attenuation, non-near-field attenuation, and critical distance can be used to specify the attenuation rate as a function of the distance to the loudspeaker. Additionally, there is the near-field - non-near-field transition parameter beta 47. The larger beta is, the faster the transition between near-field attenuation and non-near-field attenuation. The roll-off gain compensation function 42 can depend on the near-field - non-near-field transition parameter beta 47. The near-field - non-near-field transition parameter beta 47 can define how fast the roll-off gain compensation function 42 transitions between the near field and the non-near field, that is, how fast the roll-off gain compensation function 42 transitions from a rapid increase in the compensation gain per listener-loudspeaker distance 44 to a shallow / slight increase in the compensation gain per listener-loudspeaker distance 44.

[0111] It should be noted that the situation where the compensated roll-off becomes monotonically shallower with an increase in the listener-loudspeaker distance 44 can be realized by the fact that the slope of the compensated roll-off energy decreases monotonically with an increase in the listener-loudspeaker distance 44 when measured in the logarithmic region.

[0112] The roll-off gain compensation function 42 associates the listener-loudspeaker distance 44 related to the loudspeaker with a listener-loudspeaker distance compensation gain 41 for the loudspeaker related to the listener-loudspeaker distance 44. The roll-off gain compensation function 42 can be configured to compensate for a roll-off that monotonically becomes shallower as the listener-loudspeaker distance 44 increases. As described above, in a reproduction space where reverberation is effective, sound energy can decay differently in the near field than in the far field. Therefore, it is proposed to use a first attenuation parameter 481 (see decay_1_dB) for the near field, i.e., the first distance region, and a second attenuation parameter 482 (see decay_2_dB) for the far field, i.e., the second distance region, where the first distance region is associated with a shorter listener-loudspeaker distance 44 than the second distance region. As can be seen from FIGS. 10c and 10i, the roll-off gain compensation function 42 takes into account different attenuations 481 and 482 for the near field and the far field when determining the compensation gain 47 for a certain listener-loudspeaker distance 44. For example, the roll-off gain compensation function 42 can consider how much sound energy has decayed at the listener-loudspeaker distance 44 according to the first attenuation parameter 481 (see pow_nf) and according to the second attenuation parameter 482 (see pow_ff). The critical distance 44 12 separates the near field and the far field. The sound energy that decays according to the second attenuation parameter 482 (see pow_ff) can be scaled such that the decay of the sound energy according to the first attenuation parameter 481 and the decay of the sound energy according to the second attenuation parameter 482 are equal at the critical distance 44 12 . The first attenuation parameter 481 can exhibit a faster decay of sound energy than the second attenuation parameter 482. Therefore, in the roll-off gain compensation function 42, the compensated roll-off monotonically becomes shallower as the listener-loudspeaker distance 44 increases.

[0113] Furthermore, the roll-off gain compensation function 42 may consider how much sound energy has decayed at the sweet spot (refer to r_ref which is pow_ref at the sweet spot). Thus, gain adjustment is performed so that the listener position becomes the sweet spot for the set of loudspeakers in an acoustic or perceptual sense. The sound energy that decays at the sweet spot may be determined considering both the first attenuation parameter 481 and the second attenuation parameter 482.

[0114] Depending on the distance 44 of the loudspeaker to the listener position, the sound transmission time changes. These variations can be compensated by applying a delay. For example, the offset MAX_DELAY / 2 is added to the compensation delay so that both the offset MAX_DELAY / 2 and the compensation delay are always positive (refer to FIG. 10d). Furthermore, the listener-loudspeaker distance can be considered in the determination / adjustment of the delay together with the distance between the sweet spot and each loudspeaker (refer to r_ref). Thus, delay processing is performed so that the listener position becomes the sweet spot for the set of loudspeakers in an acoustic or perceptual sense.

[0115] FIG. 10d shows that for each loudspeaker, the distance 44 of the listener position to the position of each loudspeaker may be determined, and based on the distance 44, the delay (refer to delay0[i]) for each loudspeaker may be determined.

[0116] As can be seen in FIG. 10d, for each loudspeaker, a separate delay (e.g., an absolute delay) is determined (refer to the index i of the delay variable delay0). Alternatively, the delay processing may determine a reference loudspeaker from the set of loudspeakers and determine the relative delay of the loudspeakers other than the reference loudspeaker with respect to the delay determined for the reference loudspeaker.

[0117] The overhead determined by OVERHEAD_GAIN can be used (see Figure 10e). That is, this system can amplify the signal by a factor of up to 1 / OVERHEAD_GAIN when the listener is away from the loudspeaker. If the gain is replaced by this value, all gains across multiple channels are scaled using the same factor so that the maximum gain becomes 1.0 (0 dB). This corresponds to inter-channel linked limiter action.

[0118] Separate from gain adjustment, and additionally or alternatively, delay adjustment can be performed to reduce artifacts in the audio rendition due to changes in delay.

[0119] According to one embodiment, the control of the delay process may be performed by clipping the speed of the listener or by clipping the delay, and the clipping of the delay and the speed of the listener may be controlled based on the maximum allowable listener speed (see max_m_s). For example, a maximum speed may be defined, at which the change in position by the listener is too fast, so that artifacts due to changes in delay hardly occur in the audio rendition. Figure 10f shows the determination of the maximum delay change (see delay_delta) based on the maximum allowable listener speed. The number of samples of the change in delay per allowed frame is calculated as a function of the maximum allowable movement speed max_m_s. The maximum allowable movement speed max_m_s may be correlated with the maximum rate of change of delay [v in units of m / s].

[0120] According to an alternative embodiment, the control of the delay processing may be implemented by clipping the listener's acceleration or by clipping the time rate of change of the delay, and the clipping of the time rate of change of the delay and the listener's acceleration may be controlled based on the maximum allowable listener acceleration (see max_m_s_s). For example, a maximum acceleration may be defined such that at the maximum acceleration, the change in position by the listener is too fast and artifacts due to the change in delay hardly occur in the audio rendition. FIG. 10g shows the determination of the maximum time rate of change of the delay (see delay_delta2) based on the maximum allowable listener acceleration. The number of samples of the change in the delay change per allowed frame is calculated as a function of the maximum allowable moving acceleration max_m_s_s. The maximum allowable moving acceleration max_m_s_s may be correlated with the second maximum rate of change of the delay (a in m / s units).

[0121] The two examples shown in FIGS. 10f and 10g perform delay processing such that the delay compensates for the variation in the listener-loudspeaker distance between the loudspeakers.

[0122] Auditory roughness can be reduced by the following measures. · Update the VDL with an interpolated target delay value of sample accuracy (linear interpolation from the current value towards the target delay value at the end of each processing block) · The returned delay value for each output channel is used as the target value for the associated variable delay line, and this delay line applies the appropriate delay to the corresponding output signal. These output delay lines use the same implementation as the VDL used in distance rendering within MPEG-I.

[0123] Optionally, the gain is smoothed using single-pole averaging (see FIG. 10h). The averaging constant is calculated as a function of the smoothing time constant gain_ms.

[0124] If a system or an audio processor is already configured to optimize delay and / or gain without considering the near field and the far field in a reverberant playback space, it is proposed that the system or the audio processor can be configured to calibrate the adjustment of the gain and / or the delay. When dealing with a system that already applies the specific optimal gain and delay (etc.) for the sweet spot, the calibrated system option cal_system can be used. In this case (see Fig. 10i), the compensation for the gain and delay of the sweet spot is additionally calculated (in the above (see Fig. 10c), these were calculated for the listener position). In this case, the difference between the two calculation results is applied. In addition to this difference, the determination of the compensated gain shown in Fig. 10i is based on the same considerations as described for Fig. 10c (the same features are indicated by the same reference numbers).

[0125] Audio processing For example, after rendering_gd_updatecfg is called, the function rendering_gd_process is called to specify the input buffer and the output buffer (see Fig. 11a).

[0126] Optionally, gain is applied with the monopole average (see Fig. 11b). For example, the audio processor 10 described herein may be configured to perform gain adjustment to determine a gain 41 based on the listener position. This gain adjustment may be performed by considering a target value (see gain0[ch]). The target value may represent the maximum allowable compensation gain, which can be determined, for example, using the roll-off gain compensation function described herein (see Figs. 4, 6, 10c, and 10i). The current gain 41a, for example, the gain determined for each loudspeaker without considering the attenuation of the sound energy being different in the near field and the non-near field of each loudspeaker, is adjusted towards the target value, i.e., gain0[ch], with a limited change per unit of time, i.e., per sample. The attenuation of the different sound energies in the near field and the non-near field of each loudspeaker is considered when determining the target value. Since the gain changes only slightly per sample, this prevents artifacts. The target value limits the gain change and prevents too fast or incorrect gain changes due to irregular or too fast changes in the listener position.

[0127] According to an embodiment, a delay for an external delay line can be calculated (see FIG. 11c). To reduce artifacts and pitch shifts, the delay change per frame and / or the second order delay change per frame are limited. For example, the audio processor 10 described herein can be configured to perform delay processing to determine a delay 51 based on the listener position. This delay processing can be performed by considering a target value (see delay0[ch]). The target value can represent the delay for each loudspeaker without boundary conditions, for example, the delay for the actual current listener position, without considering the possibility of irregular or too fast changes in the listener position. The target value can be determined as described with respect to FIG. 10d. The delay determined in the delay processing for each loudspeaker can be smoothed. For example, the audio processor can be configured to perform smoothing in the delay processing by determining a smooth transition from the delay determined for each loudspeaker for the previous frame, i.e., the frame preceding the current frame (see reference number 51a), to the delay for the current frame, for example, the target value. Assuming that the speed and acceleration of the listener should not exceed a certain value, a smoothed delay (see reference number 51) is calculated (see the consideration of delay_delta in the limitation of the delay change and / or the consideration of delay_delta2 in the limitation of the second order delay change). It may not be necessary to consider both limitations, but considering both limitations can more efficiently reduce artifacts. The variable delay_delta represents the maximum number of samples of the change in delay per allowed frame and can be determined as described with respect to FIG. 10f. The variable delay_delta2 represents the maximum number of samples of the change in the delay change per allowed frame and can be determined as described with respect to FIG. 10g. Thereby, the maximum rate of change of the delay and / or the maximum second order rate of change of the delay are limited for the purpose of minimizing artifacts.

[0128] The returned delay value for each output channel is used as the target value for the associated variable delay line, which applies the appropriate delay to the corresponding output signal. These output delay lines use the same implementation form as the VDL.

[0129] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step represent descriptions of corresponding blocks or corresponding items or features of an apparatus.

[0130] The encoded audio signal of the present invention may be stored on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.

[0131] Depending on the requirements of a particular implementation, embodiments of the present invention may be implemented in hardware or software. This implementation may be carried out using a digitally readable control signal stored on a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which cooperates with (or is capable of cooperating with) a programmable computer system such that each method is carried out.

[0132] Some embodiments according to the present invention comprise a data carrier having an electrically readable control signal capable of cooperating with a programmable computer system such that one of the methods described herein is carried out.

[0133] In general, embodiments of the present invention may be implemented as a computer program product having program code, which is operative to carry out one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.

[0134] Other embodiments comprise a computer program for implementing one of the methods described herein, stored on a machine-readable carrier.

[0135] In other words, certain embodiments of the method of the present invention are, accordingly, computer programs having program code for implementing one of the methods described herein when the computer program is executed on a computer.

[0136] Further embodiments of the method of the present invention are, accordingly, a data carrier (or digital storage medium, or computer-readable medium) on which is recorded a computer program for implementing one of the methods described herein.

[0137] Further embodiments of the method of the present invention are, accordingly, a data stream or sequence of signals representing a computer program for implementing one of the methods described herein. The data stream or sequence of signals can be configured to be transmitted via, for example, a data communication connection, for example via the Internet.

[0138] Further embodiments comprise a processing means, for example, a computer or a programmable logic device configured or adapted to implement one of the methods described herein.

[0139] Further embodiments comprise a computer on which is installed a computer program for implementing one of the methods described herein.

[0140] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to implement some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to implement one of the methods described herein. Generally, the methods are preferably implemented by any hardware device.

[0141] Although some aspects have been described in the context of apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, and that a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent descriptions of corresponding blocks or corresponding items or features of a device.

[0142] The encoded audio signals of the present invention may be stored on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.

[0143] Depending on the requirements of a particular implementation, embodiments of the present invention may be implemented in hardware or software. This implementation may be carried out using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which stores electronically readable control signals that cooperate with (or are capable of cooperating with) a programmable computer system so that respective methods are implemented.

[0144] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals capable of cooperating with a programmable computer system so that one of the methods described herein is implemented.

[0145] In general, an embodiment of the present invention may be implemented as a computer program product having program code, which is operable to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, in a machine-readable carrier.

[0146] Other embodiments comprise a computer program stored in a machine-readable carrier for performing one of the methods described herein.

[0147] In other words, certain embodiments of the method of the present invention are thus computer programs having program code for performing one of the methods described herein when the computer program is executed on a computer.

[0148] A further embodiment of the method of the present invention is thus a data carrier (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods described herein is recorded.

[0149] A further embodiment of the method of the present invention is thus a sequence of data streams or signals representing a computer program for performing one of the methods described herein. The sequence of data streams or signals may be configured to be transmitted via a data communication connection, for example via the Internet.

[0150] A further embodiment comprises a processing means, for example a computer or a programmable logic device configured or adapted to perform one of the methods described herein.

[0151] A further embodiment comprises a computer on which a computer program for performing one of the methods described herein is installed.

[0152] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to implement some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to implement one of the methods described herein. In general, the methods are preferably implemented by any hardware device.

[0153] The embodiments described above are merely for illustrating the principles of the present invention. It is understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the following claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.

[0154] (References) [1] “Adaptively Adjusting the Stereophonic Sweet Spot to the Listener’s Position”, Sebastian Merchel and Stephan Groth, J. Audio Eng. Soc., Vol. 58, No. 10, October 2010 [2] “AUDIO PROCESSOR, SYSTEM, METHOD AND COMPUTER PROGRAM FOR AUDIO RENDERING”, WO 2018 / 202324 A1 [3] https: / / www.princeton.edu / 3D3A / PureStereo / Pure_Stereo.html [4] https: / / en.wikipedia.org / wiki / Critical_distance

Explanation of Signs

[0155] 1 Listener 10 Device, Audio Processor, Level 1 System 11 Audio renderer 12 Loudspeaker signal 14 Loudspeaker 15 Horizontal layer, loudspeaker layer 16 Interface 18 Audio signal 20 Object position input 21 Virtual position 30 Listener position input 31 Listener position 32 Loudspeaker position 40 Gain determiner 41 Gain 42 Roll-off gain compensation function 421 First compensated roll-off slope 422 Second compensated roll-off slope 44 Listener-loudspeaker distance 44 12 Critical distance 441 First distance region 442 Second distance region 46 Listener-loudspeaker distance compensation gain 47 Near-field - far-field transition parameter beta 481 First attenuation parameter 482 Second attenuation parameter 50 Delay determiner / controller 51 Delay 100 Rendering parameter 104 Object 110 Reverberation effect information 112 Reproduction room, reproduction space

Claims

1. An audio processor (10) for performing audio rendering by generating rendering parameters (100), wherein the rendering parameters (100) determine the derivation of an audio signal (18) of a loudspeaker signal (12) to be reproduced by a set of loudspeakers (14), and the audio processor (10) performs gain adjustment to determine a gain (41) for generating the loudspeaker signal (12) for the loudspeaker (14) from the audio signal (18) based on a listener position (31), and obtains reverberation effect information (110). configured to do so, The audio processor (10) is configured to use a roll-off gain compensation function (42) to associate the listener-loudspeaker distance (44) of the at least one loudspeaker (14) with a listener-loudspeaker distance compensation gain (46) for the at least one loudspeaker (14) in the gain adjustment according to the reverberation effect information (110). For the roll-off gain compensation function (42), the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance (44) increases. Audio processor (10).

2. The reverberation effect information (110) indicates the amount of effective reverberation in the playback room (112) of the audio rendering, The audio processor (10) according to claim 1, wherein the roll-off gain compensation function (42) is adapted such that the greater the amount of effective reverberation in the playback space (112), the stronger the degree to which the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance (44) increases.

3. The reverberation effect information (110) indicates whether reverberation is effective in the playback room (112) of the audio rendering. When the reverberation effect information (110) indicates that reverberation is effective in the playback room (112) of the audio rendering, the compensated roll-off uses the roll-off gain compensation function (42) that monotonically becomes shallower as the listener-loudspeaker distance (44) increases. When the reverberation effect information (110) indicates that reverberation is not effective in the playback room (112) of the audio rendering, the audio processor (10) according to claim 1 or 2 is configured to use a further roll-off gain compensation function in which the compensated roll-off is constant.

4. The roll-off gain compensation function (42) has a first compensated roll-off slope (42 1 ) within a first distance region (44 1 ), and a second compensated roll-off slope (42 2 ) within a second distance region (44 2 ), the first compensated roll-off slope (42 1 ) being greater than the second compensated roll-off slope (42 2 ), and the first distance region (44 1 ) being related to a shorter distance than the second distance region (44 2 ), the audio processor (10) according to any one of claims 1 to 3.

5. A boundary distance (44 12 ) for separating the first distance region and the second distance region from the reverberation effect information (110), the audio processor (10) according to claim 4, configured to derive

6. The first compensated roll-off slope (42 1 ) and / or the second compensated roll-off slope (42 2 ) is derived from the reverberation effect information (110), and the audio processor (10) according to claim 4 or 5 is configured to do so.

7. The audio processor (10) according to any one of claims 4 to 6 is configured to derive information from the reverberation effect information (110) about how the roll-off gain compensation function (42) transitions from the first distance region to the second distance region.

8. The audio processor (10) according to any one of claims 1 to 7 is configured to perform the gain adjustment so that the listener position (31) becomes a sweet spot for the set of loudspeakers (14) in an acoustic or perceptual sense.

9. The audio processor (10) according to any one of claims 1 to 8 is configured to perform delay processing to determine a delay (51) for generating the loudspeaker signal (12) for the loudspeaker (14) from the audio signal (18) based on the listener position (31).

10. The audio processor (10) according to claim 9 is configured to perform the delay processing so that the delay (51) compensates for variations in the listener-loudspeaker distance (44) between the loudspeakers (14).

11. The audio processor (10) according to claim 9 or 10 is configured to perform the delay processing so that the listener position (31) becomes a sweet spot for the set of loudspeakers (14) in an acoustic or perceptual sense.

12. the audio processor (10) performing the delay processing by determining the delay (51) for each loudspeaker (14) independently of the delays (51) determined for any other loudspeaker (14) in the set of loudspeakers (14), or performing the delay processing by determining a reference loudspeaker from among the set of loudspeakers (14) and determining the delay (51) of the loudspeakers (14) other than the reference loudspeaker relative to the delay (51) determined for the reference loudspeaker The audio processor (10) according to any one of claims 9 to 11, configured as described above.

13. the set of loudspeakers (14) belongs to one or more loudspeaker layers (15), and the audio processor (10) The sound source position (104 1 ) of the desired audio signal is between two loudspeaker layers (15), For each of the two loudspeaker layers (15), a virtual sound source position (104'' 1 ), 104' 1 ) corresponding to the projection of the sound source position (104 1 ) of the desired audio signal onto the respective loudspeaker layer (15), determine a first panning gain (41) for the rendering of the audio signal (18) by the loudspeakers (14) belonging to the respective loudspeaker layer (15) from the virtual sound source position, and apply 2D amplitude panning between the loudspeakers (14) of the respective loudspeaker layer (15) so as to determine the first panning gain (41) for the loudspeakers (14) belonging to the respective loudspeaker layer (15). When applied in addition to the first panning gain (41), a second panning gain (41) for rendering the audio signal (18) by the loudspeakers (14) of the two loudspeaker layers from the sound source position (104 1 ) of the desired audio signal is determined for the loudspeaker layer (15) by applying amplitude panning between the virtual sound source positions (104' 1 , 104" 1 ) of the two loudspeaker layers (15) and The audio processor (10) according to any one of claims 1 to 12, configured as described above.

14. the set of loudspeakers (14) belongs to one or more loudspeaker layers (15), and the audio processor (10) The sound source position (104 2 ) of the desired audio signal is located outside the one or more loudspeaker layers (15), Among the one or more loudspeaker layers (15), for the loudspeaker (14) of the closest loudspeaker layer (15) to the sound source position (104 2 ) of the desired audio signal, the virtual sound source position (104' 2 ) corresponding to the projection of the sound source position (104 2 ) of the desired audio signal onto the closest loudspeaker layer (15), applying 2D amplitude panning among the loudspeakers (14) belonging to the closest loudspeaker layer (15) so as to determine the first panning gain (41) for the rendering of the audio signal (18) by the loudspeaker (14) of the closest loudspeaker layer (15), the sound from a further virtual sound source position (104'' 2 ) offset from the nearest loudspeaker layer (15) towards the sound source position (104 2 ) of the desired audio signal, such that the loudspeaker (14) of the nearest loudspeaker layer (15) provides a rendering of the sound, applying further amplitude panning between the loudspeakers (14) belonging to the nearest loudspeaker layer (15) together with spectral shaping of the audio signal (18); the sound source position (104 of the desired audio signal 2 ) to cause rendering of the audio signal (18) by the loudspeaker (14) of the nearest loudspeaker layer from the 2 ) and a second panning gain (41) for panning between the virtual sound source position (104' 2 ) and the further virtual sound source position (104'' 2 ) to determine, and further applying additional amplitude panning between the virtual sound source position (104' 2 ) and the further virtual sound source position (104'' The audio processor (10) according to any one of claims 1 to 13, configured as described above.

15. when the sound source position (104 2 ) of the desired audio signal is located below the one or more loudspeaker layers (15), performing the spectral shaping of the audio signal (18) using a first equalization function that simulates the timbre of the lower sound, and / or, when the sound source position (104 2 ) of the desired audio signal is located above the one or more loudspeaker layers (15), performing the spectral shaping of the audio signal (18) using a second equalization function that simulates the timbre of the upper sound, the audio processor (10) according to claim 14, configured as such.

16. The audio processor (10) according to any one of claims 1 to 15, configured to derive the reverberation effect information (110) from a bitstream.

17. The audio processor (10) according to any one of claims 1 to 16, configured to derive the reverberation effect information (110) from side information of a bitstream and decode the audio signal (18) from the bitstream.

18. A method for audio rendering by generating rendering parameters (100), the rendering parameters (100) determining the derivation of an audio signal (18) of a loudspeaker signal (12) to be reproduced by a set of loudspeakers (14), the method Performing gain adjustment to determine a gain (41) for generating the loudspeaker signal (12) for the loudspeaker (14) from the audio signal (18) based on the listener position (31); Obtaining reverberation effect information (110); comprising; wherein, in response to the reverberation effect information (110), the gain adjustment uses a roll-off gain compensation function (42) for associating the listener-loudspeaker distance (44) of the at least one loudspeaker with a listener-loudspeaker distance compensation gain (46) for the at least one loudspeaker, and for the roll-off gain compensation function (42), the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance (44) increases.

19. A computer program having program code for instructing the computer to perform the method according to claim 18 when executed on a computer.

20. A bitstream (or a digital storage medium storing the bitstream) referred to in any one of claims 1 to 19.