Delay Processing in Audio Rendering
The audio processor stabilizes audio playback by modifying listener position data to control delay processing, addressing issues of dynamic movements and tracking errors, ensuring high-quality sound across a wide listening area.
Patent Information
- Application Number
- JP2025501467
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-12
- Filing Date
- 2023-07-07
- Publication Date
- 2025-07-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio playback systems using loudspeakers are optimized for a narrow range of listener positions and suffer significant quality degradation when the listener moves, with tracking errors and dynamic movements causing artifacts in audio reproduction.
An audio processor that modifies listener position data through smoothing, clipping, and scaling to control delay processing, ensuring stable audio reproduction even with dynamic and inaccurate tracking data, using a modified form of listener position or intermediate values to reduce artifacts.
Enables stable audio reproduction across a wide listening area by minimizing artifacts due to listener movement, maintaining high-quality sound even with sudden direction changes and tracking errors.
Smart Images

Figure 2025524641000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments according to the present invention relate to, for example, an audio processor, a system, a method, and a computer program for audio rendering such as rendering of a loudspeaker adapted to a user using a tracking device.
Background Art
[0002] A common problem in audio playback using loudspeakers is that playback is usually optimal only at one or a narrow range of listener positions. Even worse, when the listener changes position or moves, the quality of the audio playback varies greatly. The evoked spatial auditory image becomes unstable as the listening position changes away from the sweet spot. Stereo sound collapses towards the nearest loudspeaker.
[0003] This problem has been addressed by previous publications, including [1] tracking the listener's position and adjusting the gain and delay to compensate for the derivation from the optimal listening position. [2] shows an extension of how to also adapt to the spatial radiation characteristics of the loudspeakers being used. Tracking of the listener is also used together with crosstalk cancellation (XTC) (see, for example, [3]). XTC requires extremely accurate positioning of the listener, whereby tracking of the listener becomes almost essential.
[0004] Previous methods for listener position-controlled delay adjustment / compensation for loudspeaker signals assume that the tracking data is smooth and accurate in order to control a variable delay line (VDL) to adjust the delay of the loudspeaker signal. However, in reality, the listener's movement can be highly dynamic and may include sudden direction changes. In addition, the acquisition of position data can be impaired by tracking errors, time jitter, and too low or irregular position update frequencies.
[0005] Therefore, it is desirable to devise a concept with a robust delay adjustment method that can minimize perceptual artifacts through dynamic delay adjustment in consideration of highly dynamic or inaccurate tracking data input.
[0006] This object is achieved by the subject matter of the independent claims.
[0007] Advantageous embodiments are the subject matter of the dependent claims.
Summary of the Invention
Means for Solving the Problems
[0008] It has been found that smooth and accurate tracking data of a listener for controlling a variable delay line (VDL) is not always available. However, delays determined based on suboptimal listener position information can introduce artifacts in audio playback. Accordingly, an object of the present invention is to provide perceptually high-quality delay adjustment / compensation taking into account the fact that the movements of a real-world listener can be highly dynamic and may include sudden direction changes, and the fact that the acquisition of position data can be impaired by tracking errors, time jitter, and low position update frequencies. This difficulty is overcome in delay processing by using a modified form of listener position or a modified intermediate value determined by the delay processing. The concept underlying the present invention is that modifications such as controlling, limiting, smoothing, or scaling the input listener tracking data or the derived values can be performed to avoid artifacts in adaptive rendering. This is based on the recognition that such modifications can reduce the variability / noise of the listener position information or the delay determined based on the listener information. Controlling the delay processing using this modification avoids or at least reduces overly rapid and error-prone changes in the delay, thus reducing artifacts even for very deterministic sound materials such as tonal sounds (high-frequency sine waves, pitch pipes, glockenspiels). This enables delay processing adapted to the listener even for the highly dynamic and sudden direction-changing movements of a real-world listener. Accordingly, when listening to sound reproduced by a set of loudspeakers based on signals or parameters obtained by controlled delay processing, a listener can move within a wide "sweet area" (not a sweet spot) and experience a stable sound stage in this wide area.
[0009] Accordingly, an embodiment relates to an audio processor for performing audio rendering by generating rendering parameters, the rendering parameters determining the derivation of an audio signal into a loudspeaker signal to be reproduced by a set of loudspeakers. The audio processor is configured to perform delay processing to determine a delay for generating a loudspeaker signal for a loudspeaker from an audio signal based on a listener position. For example, each delay determined by the delay processing may be associated with one of the loudspeakers, for example, the audio processor may be configured to determine for each loudspeaker a delay according to which the respective loudspeaker signal may be derived accordingly. Further, the audio processor is configured to control the delay processing by modifying a listener position of a certain form, or any intermediate value determined by the delay processing based on the listener position, so as to reduce artifacts in audio reproduction due to changes in the delay, and the delay processing is initiated based on a listener position of a certain form.
[0010] A listener position of a certain form, a listener position of a certain form, for example, the coordinates of the listener, the distance of the listener to one or more of the loudspeakers of the set of loudspeakers, the speed of the listener, the acceleration of the listener, or the position change between a previous listener position and the current listener position, may be modified by adapting, smoothing, clipping, or scaling.
[0011] Any intermediate value determined by the delay processing based on the listener position may be modified by adapting, smoothing, clipping, or scaling the intermediate value. The intermediate value may be an intermediate delay value. For example, for a certain loudspeaker of the set of loudspeakers, the distance between the listener and the loudspeaker may be calculated, and then the distance may be converted into an intermediate delay value. Alternatively, the intermediate value may be the time rate of change of the delay, for example, the intermediate delay value, or the rate of change of the time rate of change of the delay, for example, the intermediate delay value.
[0012] The listener position can be defined, for example, by tracking data, by coordinates indicating the position of the listener within the playback space, such as the position of the listener's body, the position of the listener's head, or the position of the listener's ear. The listener position can be described, for example, in Cartesian coordinates, spherical coordinates, or cylindrical coordinates. Instead of the absolute position of the listener, the listener position can indicate the relative position of the listener with respect to, for example, a reference loudspeaker among a set of loudspeakers, or with respect to each loudspeaker of a set of loudspeakers, or with respect to a sweet spot within the playback space, or with respect to any other predetermined position within the playback space.
[0013] The velocity of the listener, the acceleration of the listener, the rate of change of time of the distance of the listener position to one or more of a set of loudspeakers, and the rate of change of the rate of change of time of the distance of the listener position to one or more of a set of loudspeakers can represent a form of listener position.
[0014] The sweet spot may describe the focus between loudspeakers, where the listener can perceive the sound reproduced by the loudspeakers in a manner, for example, as it was intended to be heard by a mixer. The sweet spot can define a position within the playback space where all wavefronts emitted by a set of loudspeakers arrive simultaneously. The sweet spot can alternatively be referred to as a reference listening point.
[0015] According to an embodiment, the audio processor the position of the listener, the velocity of the listener, the velocity of the listener towards one or more of a set of loudspeakers, the acceleration of the listener, the acceleration of the listener towards one or more of a set of loudspeakers, the distance of the listener position to one or more of a set of loudspeakers, The rate of change over time of the distance to one or more listener positions among a set of loudspeakers The rate of change of the rate of change over time of the distance to one or more listener positions among a set of loudspeakers The delay with respect to one or more of a set of loudspeakers The rate of change over time of the delay with respect to one or more of a set of loudspeakers, and The rate of change of the rate of change over time of the delay with respect to one or more of a set of loudspeakers To one or more of Smoothing Clipping, and Scaling using a monotonically increasing function having a monotonically decreasing slope By performing one or more of, it is configured to implement control of delay processing.
[0016] For example, a certain form of listener position can be corrected by smoothing, clipping, and / or scaling the listener position, the listener's velocity, the velocity of the listener towards one or more of a set of loudspeakers, the listener's acceleration, the acceleration of the listener towards one or more of a set of loudspeakers, the distance to one or more listener positions among a set of loudspeakers, the rate of change over time of the distance to one or more listener positions among a set of loudspeakers, and / or the rate of change of the rate of change over time of the distance to one or more listener positions among a set of loudspeakers. The intermediate value determined by delay processing based on the listener position can be corrected by smoothing, clipping, and / or scaling the delay with respect to one or more of a set of loudspeakers, such as an intermediate delay value, the rate of change over time of the delay with respect to one or more of a set of loudspeakers, and / or the rate of change of the rate of change over time of the delay with respect to one or more of a set of loudspeakers.
[0017] Such corrections make it possible to limit or control sudden and / or error-prone changes in the listener position or the delay. For example, smoothing may be applied to the listener position, which may be determined by tracking data. Thus, smoothing makes it possible to reduce, for example, tracking error and time jitter and to obtain smooth position data even with a low position update frequency. Clipping constrains or limits the value, for example, so that sudden changes are limited. For example, the listener's velocity or the listener's acceleration may be clipped so that they do not exceed a threshold. This also reduces the dynamic delay adjustment determined based on the listener position, thus reducing possible artifacts due to too-fast delay changes. The same applies to clipping the rate of change of the rate of change of the time of the distance to one or more of the set of loudspeakers or the rate of change of the distance to one or more of the set of loudspeakers with respect to the listener position. The threshold may correspond to a value such that a pitch shift caused by too-fast listener movement or an instantaneous change in the listener's movement is perceptible by the listener. Further, it is possible to directly clip the delay for one or more of the set of loudspeakers, the rate of change of the delay for one or more of the set of loudspeakers, or the rate of change of the rate of change of the delay for one or more of the set of loudspeakers so that the phase modulation caused by the delay is constrained / reduced. Additionally or alternatively, for example, scaling may be applied to reduce or attenuate values, particularly those related to fast changes in the listener position or fast changes in the delay. Using a monotonically increasing function with a monotonically decreasing slope in scaling makes it possible to scale larger values more than smaller values. Thus, large or fast changes are scaled more than small or slow changes, and the monotonically decreasing slope makes it possible to scale high velocity, acceleration, delay, rate of change of delay, or rate of change of rate of change of delay substantially more than low velocity, acceleration, delay, rate of change of delay, or rate of change of rate of change of delay. This enables an advantageous reduction of artifacts.Optionally, clipping and scaling may be combined. For example, the value may first be scaled, and if the scaled value still exceeds a threshold, the scaled value may be clipped. Optionally, smoothing may be combined with clipping and / or scaling, for example, by smoothing the listener position and then scaling and / or clipping the smoothed listener position or a value derived therefrom. Alternatively, it is also possible to first scale and / or clip the value and then smooth the scaled and / or clipped value, for example, by comparing it to a previous value, so that, for example, a smooth transition from the previous value to the current value is obtained.
[0018] According to one embodiment, the audio processor is configured to control delay processing according to control information and perform a correction according to the control information.
[0019] According to one embodiment, the audio processor is configured to derive one or more of information about the strength of smoothing, information about a clipping threshold for clipping, and information about the parameterization of a monotonically increasing function having a monotonically decreasing slope from control information. This optimizes artifact reduction because the delay processing can be controlled individually for different environments and loudspeaker setups, and control information can be provided individually for each environment or loudspeaker setup.
[0020] A further embodiment relates to a method for audio rendering by generating rendering parameters, the rendering parameters determining the derivation of an audio signal into a loudspeaker signal to be reproduced by a set of loudspeakers. The method comprises performing a delay process to determine a delay for generating a loudspeaker signal for the loudspeakers from the audio signal based on the listener position. Further, the method comprises controlling the delay process by modifying a listener position of a certain type, or any intermediate value determined by the delay process based on the listener position, so as to reduce artifacts in the audio reproduction due to changes in the delay, the delay process being initiated based on a listener position of a certain type.
[0021] A further embodiment relates to a computer program or a digital storage medium storing the same. The computer program has program code for instructing a computer to perform one of the methods described herein when the program is executed on the computer.
[0022] A further embodiment relates to a bitstream or a digital storage medium storing the same. The bitstream may comprise, for example, control information and / or an audio signal and / or rendering parameters and / or a delay and / or a listener position and / or a loudspeaker signal and / or an audio signal.
[0023] The methods, computer program products, and bitstreams described herein are based on the same considerations as the audio processor described herein. It should be noted that the methods, computer programs, and bitstreams can be completed with all features and / or functions, which are also described with respect to the audio processor.
[0024] The drawings are not necessarily to scale and it is emphasized that they generally illustrate the principles of the present invention. In the following description, various embodiments of the present invention are described with reference to the following drawings.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9a
Figure 9b
Figure 9c
Figure 9d
Figure 9e
Figure 9f
Figure 9g
Figure 9h
Figure 9i(1)
Figure 9i(2)
Figure 10a
Figure 10b
Figure 10c
[0026] In the following description, elements that are equal or equivalent, or elements having equal or equivalent functions, are denoted by equal or equivalent reference numerals even when they appear in different figures.
[0027] In the following description, in order to provide a more complete description of the embodiments of the present invention, a plurality of details are revealed. However, it will be apparent to those skilled in the art that the embodiments of the present invention can be practiced without these specific details. In other instances, well-known structures and devices are not shown in detail in order to avoid obscuring the embodiments of the present invention, and are shown in the form of block diagrams. In addition, the features of different embodiments described later in this specification may be combined with each other unless otherwise specifically stated.
[0028] In the following, various examples are described that can help achieve more effective compression when using gain and / or delay adjustments controlled by the listener's position. The gain adjustment and / or delay adjustment may be added to other parameter adjustments for sound reproduction, for example, or may be provided exclusively.
[0029] To simplify the understanding of the following examples of the present application, the description begins by presenting a possible apparatus according to the present application, to which the examples outlined later in the present application can be built. The following description begins with an explanation of an embodiment of an apparatus for generating loudspeaker signals for a plurality of loudspeakers. More specific embodiments that can be applied to the apparatus of FIG. 1, either individually or in groups, are outlined below in this specification together with the detailed description.
[0030] The apparatus of FIG. 1 is generally designated by reference numeral 10 and is for generating loudspeaker signals 12 for a plurality of loudspeakers 14 in such a way as to render at least one audio object at a virtual position where application of the loudspeaker signals 12 to or in the plurality of loudspeakers 14 is intended.
[0031] The apparatus 10 can be configured for an arrangement of the loudspeakers 14, i.e., for several positions where the plurality of loudspeakers 14 are located or are located and oriented. However, the apparatus may alternatively be configurable for different loudspeaker arrangements of the loudspeakers 14. Similarly, the number of loudspeakers 14 may be two or more, and the apparatus may be designed for a set number of loudspeakers 14 or may be configurable to accommodate any number of loudspeakers 14.
[0032] Device 10 comprises an interface 16 through which device 10 receives an audio signal 18 representing at least one audio object. For the purposes of this example, assume that the audio input signal 18 is a mono audio signal representing an audio object such as the sound of a helicopter. Additional examples and further details are provided below. Alternatively, the audio input signal 18 can be a stereo audio signal or a multi-channel audio signal. In either case, the audio signal 18 may represent the audio object in the time domain, the frequency domain, or any other domain, and it may represent the audio object in a compressed format or without compression.
[0033] As shown in FIG. 1, device 10 further comprises an object position input 20 for receiving an intended virtual position 21. That is, at the object position input 20, device 10 is notified of the intended virtual position 21 at which the audio object is to be virtually rendered by application of the loudspeaker signal 12 at the loudspeaker 14. That is, device 10 receives information on the intended virtual position 21 at the input 20, and this information can be given relative to the arrangement / position of the loudspeaker 14, relative to the sweet spot, relative to the position and / or orientation of the listener's head, and / or relative to real-world coordinates. This information can be based, for example, on a Cartesian coordinate system or a polar coordinate system. It can be based on a coordinate system centered on the room or on the listener, either as a Cartesian coordinate system or a polar coordinate system.
[0034] In addition, the apparatus 10 comprises a listener position input 30 for receiving the actual position of the listener. The listener position 31 can be defined, for example, by coordinates indicating the position of the listener within the playback space, such as the position of the listener's body, the position of the listener's head, or the position of the listener's ear, by tracking, for example, data, i.e., information on the position of the listener, over time. The listener position 31 can be described, for example, in Cartesian coordinates, spherical coordinates, or cylindrical coordinates. Instead of the absolute position of the listener, the listener position 31 can indicate, for example, the relative position of the listener with respect to a reference loudspeaker of a set of loudspeakers, or with respect to a sweet spot within the playback space, or with respect to any other predetermined position within the playback space.
[0035] For example, when the intended virtual position 21 defines the position of the audio object relative to the listener position 31, the apparatus 10 may not necessarily require a listener position input 30 for receiving the listener position 31. This is due to the fact that the intended virtual position 21 already takes into account the listener position 31.
[0036] As shown in FIG. 1, the apparatus 10 may comprise a gain determiner 40 configured to determine a gain 41 for a plurality of loudspeakers 14 in accordance with the intended virtual position 21 received at the input 20 and / or the listener position 31 received at the input 30. According to one embodiment, the gain determiner 40 may calculate an amplitude gain, one by one, for each loudspeaker signal such that the intended virtual position 21 is panned among the plurality of loudspeakers 14 and / or such that the roll-off of the sound energy is compensated. As will be outlined in more detail below (see FIG. 2), the gain 41 may comprise a pan gain. The pan gain g to be applied to each loudspeaker signal n has a horizontal component
Number
Number
[0037] In addition or alternatively, the apparatus 10 may comprise a delay determiner / controller 50 for determining / controlling a delay 51 for the plurality of loudspeakers 14 in accordance with the intended virtual position 21 received at the input 20 and / or the listener position 31 received at the input 30. The delay determiner 50 may be configured to determine a respective delay 51 for each loudspeaker such that the application of the loudspeaker signal 12 to or at the plurality of loudspeakers 14 renders at least one audio object at the intended virtual position and / or such that the loudspeaker signals reproduced by the loudspeakers 14 reach the listener simultaneously.
[0038] The apparatus 10 may comprise an audio renderer 11 configured to render the audio signal 18 based on the gain 41 and / or the delay 51 in order to derive the loudspeaker signal 12 from the audio signal 18.
[0039] With reference to FIG. 2, possible 3D panning performed by the pan gain determiner 40 will be described in more detail.
[0040] The loudspeaker 14 can be arranged in one or more horizontal layers 15. As shown in FIG. 2, the first set 141 to 145 of loudspeakers among the plurality of loudspeakers 14 may be arranged in the first horizontal layer 151, and the second set 146 to 148 of loudspeakers among the plurality of loudspeakers 14 may be arranged in the second horizontal layer 152. That is, the first set 141 to 145 of loudspeakers are arranged at an externally similar height, and the second set 146 to 148 of loudspeakers are arranged at an externally similar height. The first set 141 to 145 of loudspeakers may be arranged at or near the first height, and the second set 146 to 148 of loudspeakers may be arranged at or near, for example, a second height higher than the first height. According to the embodiment shown in FIG. 2, the listener position 31 is exemplarily arranged within the first horizontal layer 151.
[0041] In the following, an example case of rendering an object in 3D is described for an exemplary case where an object 1041, for example a sound source, is panned in a direction (as seen from listener 1) between two physically existing loudspeaker layers (which are at different heights). The object 1041 is amplitude panned in the first layer 151 (see the panned first layer position 104'1 in FIG. 2) by applying an object signal to the loudspeakers in this layer with different horizontal gains of the first layer, for example, by applying the object signal to the loudspeakers 141 to 145 such that the object 1041 is amplitude panned to the lower layer, i.e., the first layer 151. In this horizontal pan, for example, for each loudspeaker of the first set 141 to 145 of loudspeakers, the horizontal component of the respective pan gain 41
Number
Number
Number
Number
Number
[0042] Hereinafter, an example case of rendering an object in 3D will be described for an exemplary case where the object 1042 is panned above or below the outermost layer. The object can have a direction or position 1042 that is not within the range of directions between the two layers 151 and 152, as discussed with respect to the object position 1041. The intended position 1042 of the object is, for example, above or below the (physically existing) layer 15, here below all available layers, specifically below the bottom layer, i.e., below the first layer 151. As an example, the object has a direction / position 1042 below the bottom loudspeaker layer of the loudspeaker setup used as an exemplary setup in FIG. 2, i.e., below the first layer 151. In this case, a horizontal amplitude pan is applied by the bottom layer pan gain determiner 40 to render the object 1042 in that layer 151 (see the obtained position 104'2). The obtained position 104'2 can represent a virtual sound source position corresponding to the projection of the sound source position of the desired audio signal (see 1042) onto the nearest loudspeaker layer (see 151). More generally, a 2D amplitude pan is applied between the loudspeakers 141 to 145 belonging to the first layer 151, which is the loudspeaker layer closest to the object 1042. In this horizontal pan, for example, for each loudspeaker of the first set 141 to 145 of loudspeakers, the horizontal component of the respective pan gain 41 [Number] is determined. Then, in addition to spectral shaping of the audio signal to effect reproduction of sound by the loudspeakers 141 to 145 of the nearest loudspeaker layer, i.e., the first layer 151, further amplitude panning is applied between the loudspeakers 141 to 145 belonging to the nearest loudspeaker layer, i.e., the first layer 151, and this reproduction of sound simulates sound from a further virtual sound source position 104''2 offset from the nearest loudspeaker layer, i.e., the first layer 151, towards the sound source position (see 1042) of the desired audio signal. Since there are no actual loudspeakers in the upper and lower vertical directions, the vertical signal at 104''2 can be equalized to simulate the characteristics of upper or lower sound, respectively. The vertical signal is then applied to the loudspeakers designated for the up / down direction. To render the final object position 1042, the pan gain determiner 40 applies a further amplitude pan between the virtual sound source position 104'2 and the further virtual sound source position 104''2, determines a second pan gain for the pan between the virtual sound source position 104'2 and the further virtual sound source position 104''2, and may be configured to effect rendering of the audio signal by the loudspeakers 141 to 145 of the nearest loudspeaker layer from the sound source position 1042 of the desired audio signal. The spectral shaping of the audio signal may be implemented using a first equalization function that simulates the timbre of lower sound when the sound source position of the desired audio signal is located below one or more loudspeaker layers, i.e., below the first layer 151, and / or the spectral shaping of the audio signal may be implemented using a second equalization function that simulates the timbre of upper sound when the sound source position of the desired audio signal is located above one or more loudspeaker layers, i.e., above the second layer 152.
[0043] Figure 3 shows an embodiment of an audio processor 10 (see audio renderer 11) for performing audio rendering by generating rendering parameter 100, which determines the derivation of an audio signal 18 of a loudspeaker signal 12 to be reproduced by a set of loudspeakers 14. The focus of the embodiment shown in Figure 3 is on delay determiner 50. Optionally, as described with respect to Figure 1, delay determiner 50 may be combined with gain determiner 40.
[0044] Audio processor 10 is configured to perform delay processing to determine a delay 51 from audio signal 18 for a loudspeaker signal 12 for loudspeakers 14 based on listener position 31 (see delay determiner 50). Audio processor 10 is configured to control the delay processing by modifying (52') a form of listener position 31 or by modifying (52") any intermediate value 54 determined by the delay processing based on listener position 31, in order to reduce artifacts in audio reproduction due to changes in delay 51, and the delay processing is initiated based on a form of listener position 31 (see controller 52). Modifications 52' or 52" may be performed by smoothing, clipping, and / or scaling the respective inputs. Scaling may be performed using a monotonically increasing function having a monotonically decreasing slope.
[0045] A form of listener position 31 may correspond to an absolute listener position within the playback space 112, the listener's velocity, the listener's velocity towards one or more of the set of loudspeakers 14, the listener's acceleration, the listener's acceleration towards one or more of the set of loudspeakers 14, the distance of listener 1 to one or more of the set of loudspeakers 14, the time rate of change of the distance of listener 1 to one or more of the set of loudspeakers 14, and / or the rate of change of the time rate of change of the distance of listener 1 to one or more of the set of loudspeakers 14.
[0046] The intermediate value 54 may correspond to a delay for one or more of the sets of loudspeakers 14, a rate of change of the delay for one or more of the sets of loudspeakers 14, and / or a rate of change of the rate of change of the delay for one or more of the sets of loudspeakers 14.
[0047] The concept behind the embodiments of the present invention is that restricting the variability (noise) of the listener position 31, for example the input listener tracking data, or of a derived value, for example the intermediate value 54, can be used to avoid artifacts in adaptive rendering, specifically due to the effect of the variable delay line (VDL). It is also possible to react fast enough to the movement of the listener while avoiding artifacts related to delay adjustment. For example, the speed and / or acceleration of the movement can be used to control changes in the VDL operation, for example by the controller 52.
[0048] Thus, according to one embodiment, the above concept results in an audio processor that, for example, · means (see controller 52) for controlling the delay processing (aspect: "what to control / smooth?") of the loudspeaker audio signals 12 · a control criterion for instructing the above delay control / smoothing (aspect: "how to control / depend on which criterion?") is provided. For example, the controller 52 may acquire control parameters or control information and control the corrections 52' and / or 52".
[0049] According to one embodiment, the audio processor 10 is configured to derive one or more of information about the strength of smoothing, information about a clipping threshold for clipping, and information about the parameterization of a monotonically increasing function having a monotonically decreasing slope from the control information.
[0050] Means for controlling delay processing: Inside a renderer adapted to the user (see audio processor 10), according to one embodiment, the following calculations are performed for each loudspeaker 14. From the tracked user position 31 (which may be jittery), the distance between the user, i.e., the listener 1 and each loudspeaker 14, is calculated, for example, by the delay processing unit 55 of the delay determiner 50. This distance can then be converted to the respective delay that would normally need to be applied to the loudspeaker feed signal, for example, an intermediate value 54 (specified either in sample units at the system's sampling rate or in seconds / milliseconds). This target delay can then normally be used to control the variable delay line (VDL) of the system. However, too fast or error-prone changes in the delay can result in artifacts in the audio playback. Therefore, for example, it is proposed to input such a delay as the intermediate value 54 to the controller 52, or to directly input a certain form of listener position 31 to the controller 52, because too fast a change in the listener position 31 or an error-prone listener position 31 will result in too fast or error-prone changes in the delay. As a result, there are several possibilities for limiting the possible impact of too fast and error-prone changes in the delay. · The delay calculated in each processing frame can be smoothed / controlled, for example, by modification 52" (the most preferred variant). · In this preferred variant, the change (difference) in the sample delay per frame is limited / controlled. · Further, for example, additionally or alternatively, the user-loudspeaker distance calculated in each processing frame can be smoothed / controlled, for example, by modification 52'. · In this preferred variant, the change (difference) in the user-loudspeaker distance per frame is limited / controlled. · Further, for example, additionally or alternatively, the tracked user position used in each processing frame can be smoothed / controlled, for example, by modification 52' (which is at least a preferred variant). · Specifically, the change (difference) in user position for each frame can be restricted / controlled.
[0051] Control criteria: To control the delay processing (as described above), some intelligent criteria can be utilized to ensure that the normal operation of the processor is not hindered (i.e., the delay adjustment ideally provides the same audio quality to listener 1 as in the sweet spot and functions in an optimal way such that it responds fast enough when listener 1 moves within the room, i.e., playback space 112, and relative to the loudspeaker setup. Also, at the same time, there should be no artifacts generated by the time-varying delay for very deterministic sound materials such as tonal sounds (high-frequency sine wave sounds, pitch pipes, glockenspiels)).
[0052] One or more parameters including the following can be used as intelligent control criteria for delay processing. · The estimated listener speed (expressed in m / s or other equivalent units). This can be · The speed in 3D space, or · The speed in the direction towards the loudspeaker (possible but less preferred) measured as any of these · The estimated listener acceleration (expressed in m / s 2 or other equivalent units) · The acceleration in 3D space, or · The acceleration in the direction towards the loudspeaker (possible but less preferred) · Alternatively, as a simple but less effective approach, the control of the VDL activity can be implemented by directly applying temporal smoothing to the variables themselves in the computational chain (see above), i.e., · VDL delay · User-loudspeaker distance · Tracked user position by directly applying temporal smoothing to these.
[0053] Therefore, the audio signal processor 10 · Means for controlling the delay processing of the loudspeaker audio signal 18 (Aspect: "What to control?"), for example, the controller 52 · The above references (VDL delay, user loudspeaker distance, tracked user position) · Control criteria for instructing the above delay control (Aspect: "How to control / What criteria to depend on?") · The above references (speed, acceleration) · Alternative (the above reference): · VDL delay · User loudspeaker distance · Tracked user position direct smoothing of may be provided.
[0054] One embodiment of the present invention relates to an audio processor 10 configured to generate a set of one or more parameters, namely rendering parameters 100, for each of one or more sets of loudspeakers 14 (which may be parameters that can affect, for example, the delay, level, or frequency response of one or more audio signals 18). This is based on the listener position 31 (the listener position 31 can be, for example, the position of the entire body of the listener 1 in the same room as one or more sets of loudspeakers 14, that is, the reproduction space 112, or, for example, only the position of the head of the listener 1, or even, for example, the position of the ear of the listener 1. The listener position 31 can be, for example, a position based on one or more sets of loudspeakers 14, for example, the distance from the listener's head to one or more sets of loudspeakers 14) and the loudspeaker positions of one or more sets of loudspeakers 14, and determines the derivation of the audio signal 18 of the loudspeaker signal 12 to be reproduced by each loudspeaker 14.
[0055] Loudspeaker signal delay adjustment can be implemented by a variable (partial) delay line (VDL). Although adjustment of the steady state of the VDL is not important, its dynamic behavior while interactively adjusting the VDL delay in response to the user's movement via a delay control signal should be carefully constrained to avoid perceptual defects. A possible perceptual defect is due to the fact that the dynamically adjusted delay line performs phase modulation on the audio signal 18 processed in that delay line which is operated by the control signal.
[0056] Unconstrained phase modulation can cause audible roughness of the tonal signal and / or a perceivable pitch shift. Audible roughness is caused by, for example, high-speed modulation in the control signal due to position tracking time jitter, or sample-and-hold behavior of overly slow or unstable position data acquisition. A perceivable pitch shift or an instantaneous jump in pitch can be caused by overly fast user movement or an instantaneous change in the user's movement.
[0057] Therefore, one or more of the following countermeasures are perceptually beneficial. · Constraining acceptable delay changes, e.g., by modification 52", limits audible pitch offsets through limiting instantaneous frequency corrections · Constraining acceptable changes in delay changes, e.g., by modification 52", avoids audible pitch jumps through limiting instantaneous frequency jumps · Adjusting the absolute delay rather than the relative delay (with respect to the dynamically selected reference channel with the minimum delay), e.g., by modification 52", avoids unnecessary instantaneous frequency jumps due to listener movement, especially near the sweet spot, where otherwise the reference channel would suddenly change · For example, by modification 52", instead of adjusting the relative delay (with respect to the dynamically selected reference channel with the minimum delay), adjusting the absolute delay minimizes the bias in the perceived pitch offset for the sum of all channels, which can stay close to the true pitch because the pitch offset can be above in one channel with a pitch offset and below in another channel, and otherwise the reference channel remains unmodified and only the other channels are modified in one direction.
[0058] Note that in addition to smoothing, clipping, and / or scaling using a monotonically increasing function with a monotonically decreasing slope, interpolation between frames can be applied by the controller 52, for example, in modification 52' or 52". In other words, smoothing, clipping, and / or scaling using a monotonically increasing function with a monotonically decreasing slope may be performed on a frame-by-frame basis, similar to other tasks such as gain adaptation and panning, and interpolation between successive smoothed / clipped / scaled values, i.e., the values of successive frames, may be used to vary the delay in units finer than a frame, thereby connecting the value of the previous frame to the value of the current frame.
[0059] Now, embodiments of the present invention will be described here with respect to the rendering of adaptive loudspeakers.
[0060] First, general notes shall be made. Instead of rendering the MPEG-I scene for binauralization for headphones, playback via loudspeakers is defined. In this mode of operation, the MPEG-I Spatializer (HRTF-based renderer) is replaced by a dedicated loudspeaker-based renderer described below.
[0061] For a high-quality listening experience, assume that the loudspeaker setup has the listener 1 in a dedicated fixed position, the so-called sweet spot. Usually, in a 6DOF playback situation, the listener 1 is moving. Therefore, the 3D spatial rendering must instantaneously and continuously adapt to the changing listener position 31. This can be achieved at two hierarchically nested technical levels. 1. The loudspeaker signals 12 are made to reach the listener position 31 with the same gain and delay, i.e., as if the listener position 31 were at the sweet spot. For example, the gain 41 and the delay 51 are applied to the loudspeaker signals 12. Optionally, a high-shelving compensation filter is applied to each loudspeaker signal 12 with respect to the current listener position 31 and the orientation of the loudspeaker with respect to the listener 1. Thus, as the listener 1 moves off the axis of the loudspeaker 14 or even further away from the loudspeaker 14, the high-frequency loss due to the radiation high-frequency pattern of the loudspeaker is compensated. 2. Due to the 6DoF movement, the angles between the loudspeaker 14, the object, and the listener 1 change as a function of the listener position 31. Therefore, for example, a 3D amplitude panning algorithm (see Figure 2) is updated in real time with the changing listener position 31 and the relative positions and angles with respect to a fixed loudspeaker configuration as set in the LSDF. All coordinates (listener position 31, sound source position) can be transformed into the coordinate system of the listening room, i.e., the coordinate system of the playback space 112.
[0062] Physical compensation level (level 1) Figure 4 shows an overview of an embodiment of the level 1 system 10 together with its main components and parameters. The audio processor 10 described with respect to Figures 1 to 3 may comprise the features and / or functions described with respect to the embodiment of Figure 4.
[0063] Level 1: Real-time updated compensation of (frequency-dependent) loudspeaker gain and delay (see Audio Renderer 11) enables "enhanced rendering of content". By utilizing tracked user location information, such as a form of listener position 31, the listener 1, i.e., the user, can move within a wide "sweet area" (rather than a sweet spot) and experience a stable sound stage within this wide area when listening to, for example, legacy content (e.g., stereo, 5.1, 7.1+4H). In an immersive format (i.e., not a format for stereo), the sound is not felt to flow towards the nearest speaker 14 when walking away from the sweet spot, but rather is felt to move away from the loudspeaker 14, i.e., this has somewhat similar properties to what is known as wave field synthesis, but is for a single-user experience. In stereo playback, this technology provides left-right sound stage stability for a wide range (i.e., the range between the left and right loudspeakers at any distance) of user positions 31.
[0064] The gain compensation at Level 1 is based on, for example, the amplitude decay law. In a free field, the amplitude is proportional to 1 / r, where r is the distance from the listener 1 to the loudspeaker 14 (1 / r corresponds to a 6 dB attenuation every time the distance doubles). In the room 112, due to the presence of acoustic reflections and reverberation, the sound attenuates more slowly as the distance to the loudspeaker 14 increases. Thus, for example, parameters of near-field attenuation, non-near-field attenuation, and / or critical distance, included in the reverberation effect information 110, can be used to specify the attenuation rate as a function of the distance to the loudspeaker 14. In addition, there can be a near-field - non-near-field transition parameter beta, for example, included in the reverberation effect information 110. The larger beta is, the faster the transition between near-field attenuation and non-near-field attenuation. FIG. 5 shows an example of gain compensation as a function of distance, i.e., the roll-off gain compensation function 42, that can be used to determine the applicable compensation gain by the Audio Renderer 11. In the reverberant field, the gain change is smaller than in the free field.
[0065] The delay compensation at level 1, for example, calculates the propagation delay from each loudspeaker 14 to the listener position 31 and then applies the delay to each loudspeaker 14 to compensate for the difference in propagation delay between the loudspeakers 14. The delay can be normalized (an offset is added or subtracted) so that the minimum delay applied to the loudspeaker signal 12 becomes zero.
[0066] Object rendering level (level 2) Level 2: Object panning tracked by the user enables the rendering of point sources (objects, channels) in the 6DoF playback space and requires level 1 as essential. Thus, it addresses the use case of "6DoF VR / AR rendering". The following features and / or functions may additionally be included in the level 1 system 10.
[0067] For example, as described with respect to FIG. 2, a 3D amplitude panning algorithm that functions in loudspeaker layers, such as horizontal and height layers, may be used. Each layer may apply a 2D panning algorithm for the projection of the object onto the layer. The final 3D object is rendered by applying amplitude panning between two virtual objects from the 2D pans in the two layers.
[0068] When the object is located above the highest layer, the 2D pan is applied in that layer. The final 3D object is rendered by applying amplitude panning between the virtual object from the 2D pan and an object (not present) in the vertical direction above. The signal of the vertical object may be equalized to simulate the timbre of the sound above and evenly distributed to the loudspeakers of the highest layer.
[0069] When the object is located below the bottom layer, 2D panning is applied in that layer. The final 3D object is rendered by applying amplitude panning between the virtual object from the 2D panning and the (non-existent) object in the vertical direction below. The signal of the vertical object can be equalized to simulate the timbre of the lower sound and distributed equally to the loudspeakers of the bottom layer.
[0070] The vertical panning as described is equally applicable to loudspeaker setups with one layer such as 5.1 and loudspeaker setups with multiple layers such as 7.4.6.
[0071] Levels 1 and 2 applied to object rendering faithfully render the MPEG-I scene, as in the case of via headphones. This has a great advantage compared to loudspeakers that render MPEG-I content without applying adaptive tracking (1 and 2).
[0072] Physical compensation level (Level 1) In the following, embodiments of gain and delay adjustment based on the listener position are described using code excerpts (see FIGS. 9c to 9i and FIGS. 10b and 10c). The features and / or functions described below regarding gain and / or delay adjustment may be included in the audio processor 10 of FIG. 1 or the Level 1 system 10 of FIG. 4. The audio processor 10 of FIG. 3 may additionally include the features and / or functions described below regarding delay adjustment. Optionally, the audio processor 10 of FIG. 3 may include the features and / or functions described below regarding gain and / or delay adjustment. Optionally, the audio processor 10 of FIG. 1, the audio processor 10 of FIG. 3, and the audio processor 10 of FIG. 3 may include further features and / or functions as described below.
[0073] Data elements and variables Definitions and / or explanations of the data elements and variables used below (see FIGS. 6 to 10c) are provided. SFREQ_MIN Minimum sample rate [Hz] = 44100 SFREQ_MAX Maximum sample rate [Hz] = 48000 VSOUND Speed of sound in air [m / s] = 340.0 MAX_DELAY Maximum delay [samples] = 960 OVERHEAD_GAIN Overhead [lin] = 0.25 framesize Number of samples per frame, default: 256 sfreq_Hz Sampling frequency of the input audio, default: 48000 nchan Number of channels (loudspeakers) max_delay Maximum delay [samples], default: MAX_DELAY bypass_on 0: Normal operation, 1: Bypass, default: 0 ref_proc 0: Normal operation, 1: Processing for sweet spot, default: 0 cal_system 0: Normal operation, 1: Calibrated system, default: 0 gain_on 0: Gain off, 1: On, default: 1 delay_on 0: Delay off, 1: On, default: 1 decay_1_dB Proximity field sound attenuation [dB] every time the distance doubles, default: 8 decay_2_dB Non-proximity field sound attenuation [dB] every time the distance doubles, default: 0 beta 1: Default proximity - non - proximity transition, > 1: Faster transition crit_dist_m Critical distance [m], default: 4 max_m_s Maximum moving speed [v in m / s units], default: 1 max_m_s_s Maximum moving acceleration [a in m / s units], default: 1 gain_ms Gain smoothing time constant [ms], default: 40 sweet_spot sweet spot position [m, m, m] spk_pos speaker coordinates [m, m, m] listener_pos listener coordinates [m, m, m]
[0074] All coordinates are relative to the listening room as defined, for example, in an LSDF file.
[0075] These parameters can be stored in the following structures.
[0076] Public data structure typedef struct rendering_gd_cfg { int framesize; float sfreq_Hz; int nchan; float max_delay; } rendering_gd_cfg_t; typedef struct rendering_gd_rt_cfg { int bypass_on; int ref_proc; int cal_system; int gain_on; int delay_on; float decay_1_dB; float decay_2_dB; float crit_dist_m; float beta; float max_m_s; float max_m_s_s; float gain_ms; float sweet_spot[3]; float spk_pos[NCHANMAX][3]; float listener_pos[3]; } rendering_gd_rt_cfg_t;
[0077] Internal parameters calculated from the parameters and states enumerated above are stored, for example, in the following structure.
[0078] Internal data structure typedef struct { / * Static parameters * / float sfreq_Hz; int nchan; int framesize; / * Real-time parameters * / int bypass_on; int gain_on; float delta_gi; float delta_gd; float gain_alpha; float delay_delta; float delay_delta2; / * States * / float delay0[NCHANMAX]; float delay[NCHANMAX]; float gain0[NCHANMAX]; float gain[NCHANMAX]; } rendering_gd_data_t;
[0079] Step description Embodiments of gain and delay adjustment based on listener position are described below using code excerpts related to different stages. This embodiment may include an initialization stage (see FIG. 6), a release stage (see FIG. 7), a reset stage (see FIG. 8), a real-time parameter update stage (see FIGS. 9a to 9i), and an audio processing stage (see FIGS. 10a to 10c). The audio processor 10 of FIG. 1, the level 1 system 10 of FIG. 4, and the audio processor 10 of FIG. 3 may include features and / or functions described with respect to one or more of the stages, or individual features and / or functions of one or more of the stages.
[0080] Initialization FIG. 6 illustratively shows a code excerpt of the initialization stage.
[0081] The loudspeaker setup can be loaded from the LSDF file.
[0082] A structure of type rendering_gd_cfg_t is initialized with default values, and the nchan field is set to the number of loudspeakers in the loudspeaker setup.
[0083] A structure of type rendering_gd_rt_cfg_t is initialized with default values. The loudspeaker positions from the LSDF file are stored in the field spk_pos. If a ReferencePoint element is given in the LSDF file, its coordinates are stored in the field sweet_spot. The field cal_system is set to the value of the calibrated attribute if it exists.
[0084] The above-mentioned structures are passed to the rendering_gd_init function.
[0085] Release FIG. 7 illustratively shows a code excerpt of the release stage.
[0086] Reset FIG. 8 exemplarily shows a code excerpt of the reset phase. FIG. 8 shows that all internal buffers are flushed.
[0087] Update real-time parameters In the update thread, the virtual listening position is converted to the coordinate system of the listening room. This is only important for the VR scene, and in the AR scene, these two coordinate systems coincide.
[0088] All further processing occurs in the audio thread.
[0089] The structure of type rendering_gd_rt_cfg_t is updated by setting the listener_pos field to the listener position (in the coordinate system of the listening room) (see FIG. 9a). Then, this structure is passed to the rendering_gd_updatecfg function (see FIG. 9a).
[0090] For each loudspeaker, the compensation gain and delay are calculated. The reference distance r_ref (calculated in FIG. 9a) is the distance at which the gain and delay compensation are 0 (dB, samples). Based on the distance from the loudspeaker to the listener r and the reference distance r_ref, the gain and delay compensation are calculated. The calculation of the listener-loudspeaker distance 44 based on the listener position 31 and each loudspeaker position 32 is shown in FIG. 9b. The listener-loudspeaker distance 44 can represent a certain form of listener position 31.
[0091] In a free field, sound attenuates by 6 dB every time the distance doubles. In a room, the attenuation can be approximated by using a smaller attenuation, for example 4 dB every time the distance doubles. Alternatively, the critical distance (hall range) can be considered. When close to the loudspeaker, the attenuation is decay_dB every time the distance doubles. Beyond the critical distance crit_dis_m, the sound only attenuates slowly. To determine the gain compensation that compensates for the gain change due to the described sound attenuation, it is proposed to use the roll-off gain compensation function 42 (see FIGS. 5, 9c, and 9i).
[0092] The gain compensation can be based on the amplitude attenuation law. In a free field, the amplitude is proportional to 1 / r, where r is the distance from the listener to the loudspeaker (1 / r corresponds to an attenuation of 6 dB every time the distance doubles). In a room, due to the presence of acoustic reflections and reverberation, the sound attenuates more slowly as the distance to the loudspeaker increases. Therefore, the parameters of near-field attenuation, non-near-field attenuation, and critical distance can be used to specify the attenuation rate as a function of the distance to the loudspeaker. In addition, there is a near-field - non-near-field transition parameter beta 47. The larger beta is, the faster the transition between near-field attenuation and non-near-field attenuation. The roll-off gain compensation function 42 can depend on the near-field - non-near-field transition parameter beta 47. The near-field - non-near-field transition parameter beta 47 can define how fast the roll-off gain compensation function 42 transitions between the near field and the non-near field, that is, how fast the roll-off gain compensation function 42 transitions from a sharp increase in the compensation gain per listener-loudspeaker distance 44 to a shallow / slight increase in the compensation gain per listener-loudspeaker distance 44.
[0093] It should be noted that the situation where the compensated roll-off becomes monotonically shallower as the listener-loudspeaker distance 44 increases can be realized by the fact that the slope of the compensated roll-off energy decreases monotonically as the listener-loudspeaker distance 44 increases when measured in the logarithmic region.
[0094] The roll-off gain compensation function 42 maps the listener loudspeaker distance 44 related to the loudspeaker to the listener loudspeaker distance compensation gain 41 for the loudspeaker related to the listener loudspeaker distance 44. The roll-off gain compensation function 42 can be configured to compensate for a roll-off that monotonically becomes shallower as the listener loudspeaker distance 44 increases. As described above, in a playback space where reverberation is effective, sound energy can decay differently in the near field than in the far field. Therefore, it is proposed to use a first attenuation parameter 481 (see decay_1_dB) for the near field, i.e., the first distance region, and a second attenuation parameter 482 (see decay_2_dB) for the far field, i.e., the second distance region, where the first distance region is associated with a shorter listener loudspeaker distance 44 than the second distance region. As can be seen from FIGS. 9c and 9i, the roll-off gain compensation function 42 takes into account different attenuations 481 and 482 for the near field and the far field when determining the compensation gain 47 for a certain listener loudspeaker distance 44. For example, the roll-off gain compensation function 42 can consider how much sound energy has decayed at the listener loudspeaker distance 44 according to the first attenuation parameter 481 (see pow_nf) and according to the second attenuation parameter 482 (see pow_ff). Critical distance 44 12 separates the near field from the far field. The sound energy that decays according to the second attenuation parameter 482 (see pow_ff) is such that the decay of the sound energy according to the first attenuation parameter 481 and the decay of the sound energy according to the second attenuation parameter 482 are equal at the critical distance 44 12 can be scaled. The first attenuation parameter 481 can exhibit a faster decay of sound energy than the second attenuation parameter 482. Therefore, in the roll-off gain compensation function 42, the compensated roll-off becomes monotonically shallower as the listener loudspeaker distance 44 increases.
[0095] Furthermore, the roll-off gain compensation function 42 may consider how much sound energy has decayed at the sweet spot (see r_ref which is pow_ref at the sweet spot). Therefore, gain adjustment is performed so that the listener position becomes the sweet spot for the set of loudspeakers in an acoustic or perceptual sense. The sound energy that decays at the sweet spot can be determined considering both the first attenuation parameter 481 and the second attenuation parameter 482.
[0096] The transmission time of sound changes according to the distance 44 of the loudspeaker to the listener position. These variations can be compensated by applying a delay. For example, the offset MAX_DELAY / 2 is added to the compensation delay so that both the offset MAX_DELAY / 2 and the compensation delay are always positive (see FIG. 9d). Furthermore, the listener loudspeaker distance can be considered in the determination / adjustment of the delay together with the distance between the sweet spot and each loudspeaker (see r_ref). Therefore, delay processing is performed so that the listener position becomes the sweet spot for the set of loudspeakers in an acoustic or perceptual sense.
[0097] FIG. 9d shows that for each loudspeaker, the distance 44 of the listener position to the position of each loudspeaker may be determined, and based on the distance 44, the delay (see delay0[i]) for each loudspeaker may be determined.
[0098] As can be seen in FIG. 9d, for each loudspeaker, a separate delay (e.g., an absolute delay) is determined (see index i of the delay variable delay0). Alternatively, the delay processing may determine a reference loudspeaker from the set of loudspeakers and determine the relative delay of the loudspeakers other than the reference loudspeaker with respect to the delay determined for the reference loudspeaker.
[0099] The overhead determined by OVERHEAD_GAIN can be used (see Figure 9e). That is, this system can amplify the signal by a factor of up to 1 / OVERHEAD_GAIN when the listener is away from the loudspeaker. If the gain is replaced by this value, all gains across multiple channels are scaled using the same factor such that the maximum gain becomes 1.0 (0 dB). This corresponds to inter-channel linked limiter action.
[0100] Separate from gain adjustment, and additionally or alternatively, delay adjustment may be performed to reduce artifacts in audio playback due to changes in delay.
[0101] According to one embodiment, the control of the delay process may be performed by clipping the speed of the listener or by clipping the delay, and the clipping of the delay and the speed of the listener may be controlled based on the maximum allowable listener speed (see max_m_s). For example, a maximum speed may be defined such that at the maximum speed, the change in position by the listener is too fast for artifacts due to changes in delay to occur hardly in audio playback. Figure 9f shows the determination of the maximum delay change (see delay_delta) based on the maximum allowable listener speed. The number of samples of the change in delay per allowed frame is calculated as a function of the maximum allowable movement speed max_m_s. The maximum allowable movement speed max_m_s may be correlated with the maximum rate of change of delay [v in units of m / s].
[0102] According to an alternative embodiment, the control of the delay processing may be implemented by clipping the listener's acceleration or by clipping the time rate of change of the delay, and the clipping of the time rate of change of the delay and the listener's acceleration may be controlled based on the maximum allowable listener acceleration (see max_m_s_s). For example, a maximum acceleration may be defined, at which the change in position by the listener is too fast for artifacts due to the change in delay to occur hardly in audio playback. FIG. 9g shows the determination of the maximum time rate of change of the delay (see delay_delta2) based on the maximum allowable listener acceleration. The number of samples for which the delay change is allowed to vary per frame is calculated as a function of the maximum allowable moving acceleration max_m_s_s. The maximum allowable moving acceleration max_m_s_s may be correlated with the second maximum rate of change of the delay (a in m / s units).
[0103] The two examples shown in FIGS. 9f and 9g perform delay processing such that the delay compensates for the variation in the listener loudspeaker distance between the loudspeakers.
[0104] Auditory roughness can be reduced by the following measures. · Update the VDL with the interpolated target delay value of the sample accuracy (linear interpolation from the current value towards the target delay value at the end of each processing block) · The returned delay value for each output channel is used as the target value for the associated variable delay line, which applies the appropriate delay to the corresponding output signal. These output delay lines use the same implementation form as the VDL used in the distance rendering within MPEG-I.
[0105] Optionally, the gain is smoothed using single-pole averaging (see FIG. 9h). The averaging constant is calculated as a function of the smoothing time constant gain_ms.
[0106] If a system or an audio processor is already configured to optimize delay and / or gain without considering the near field and the non - near field in a reverberant playback space, it is proposed that the system or the audio processor can be configured to calibrate the adjustment of the gain and / or the delay. When dealing with a system that already applies the specific optimal gain and delay (etc.) for the sweet spot, the calibrated system option cal_system can be used. In this case (see Fig. 9i), the compensation for the gain and delay of the sweet spot is additionally calculated (in the above (see Fig. 9c), these were calculated for the listener position). In this case, the difference between the two calculation results is applied. In addition to this difference, the determination of the compensated gain shown in Fig. 9i is based on the same considerations as those described with respect to Fig. 9c (the same features are indicated by the same reference numbers).
[0107] Audio processing For example, after rendering_gd_updatecfg is called, the function rendering_gd_process is called, specifying the input buffer and the output buffer (see Fig. 10a).
[0108] Optionally, gain is applied with the monopole average (see FIG. 10b). For example, the audio processor 10 described herein may be configured to perform gain adjustment to determine a gain 41 based on the listener position. This gain adjustment may be performed by considering a target value (see gain0[ch]). The target value may represent the maximum allowable compensation gain, which can be determined, for example, using the roll-off gain compensation function described herein (see FIGS. 5, 9c, and 9i). The current gain 41a, for example, the gain determined for each loudspeaker without considering the attenuation of the sound energy being different in the near field and the non-near field of each loudspeaker, is adjusted towards the target value, i.e., gain0[ch], with a limited change per unit of time, i.e., per sample. When determining the target value, the attenuation of the different sound energies in the near field and the non-near field of each loudspeaker is considered. Since the gain changes only slightly per sample, this prevents artifacts. The target value limits the gain change and prevents too fast or incorrect gain changes due to irregular or too fast changes in the listener position.
[0109] According to one embodiment, a delay for an external delay line can be calculated (see FIG. 10c). To reduce artifacts and pitch shift, the delay change per frame and / or the second-order delay change per frame are limited. For example, the audio processor 10 described herein may be configured to perform delay processing to determine a delay 51 based on the listener position. This delay processing may be performed by considering a target value (see delay0[ch]). The target value may represent a delay for each loudspeaker without boundary conditions, for example, a delay for the actual current listener position, not considering the possibility of an irregular or too fast change in the listener position. The target value may be determined as described with respect to FIG. 9d. The delay determined in the delay processing for each loudspeaker may be smoothed. For example, the audio processor may perform smoothing in the delay processing by determining a smooth transition from the delay determined for each loudspeaker for the previous frame, i.e., the frame preceding the current frame (see reference number 51a), to the delay for the current frame, e.g., the target value. Assuming that the speed and acceleration of the listener should not exceed a certain value, a smoothed delay (see reference number 51) is calculated (see the consideration of delay_delta in the limitation of the delay change and / or the consideration of delay_delta2 in the limitation of the second-order delay change). It may not be necessary to consider both limitations, but considering both limitations can more efficiently reduce artifacts. The variable delay_delta represents the maximum number of samples of the delay change per allowed frame and may be determined as described with respect to FIG. 9f. The variable delay_delta2 represents the maximum number of samples by which the delay change is allowed to change per frame and may be determined as described with respect to FIG. 9g. Thereby, the maximum rate of change of the delay and / or the maximum second-order rate of change of the delay are limited for the purpose of minimizing artifacts.
[0110] The returned delay values for each output channel are used as target values for the associated variable delay lines, which apply the appropriate delay to the corresponding output signals. These output delay lines use the same implementation form as the VDL.
[0111] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, and that a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step represent descriptions of corresponding blocks or corresponding items or features of a device.
[0112] The encoded audio signals of the present invention may be stored on a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.
[0113] Depending on the requirements of a particular implementation, embodiments of the present invention may be implemented in hardware or software. This implementation may be carried out using a digitally-stored control signal which is stored on a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, which cooperates with (or is capable of cooperating with) a programmable computer system so that respective methods are executed.
[0114] Some embodiments according to the present invention comprise a data carrier having an electronically-readable control signal which is capable of cooperating with a programmable computer system so that one of the methods described herein is executed.
[0115] In general, embodiments of the present invention may be implemented as a computer program product having program code, the program code being operative to perform one of the methods when the computer program product is executed on a computer. The program code may be stored, for example, on a machine-readable carrier.
[0116] Other embodiments comprise a computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0117] In other words, certain embodiments of the method of the present invention are, accordingly, computer programs having program code for performing one of the methods described herein when the computer program is executed on a computer.
[0118] Further embodiments of the method of the present invention are, accordingly, data carriers (or digital storage media, or computer-readable media) on which a computer program for performing one of the methods described herein is recorded.
[0119] Further embodiments of the method of the present invention are, accordingly, a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals can be configured to be transmitted, for example, via a data communication connection, such as via the Internet.
[0120] Further embodiments comprise processing means, such as a computer or a programmable logic device configured or adapted to perform one of the methods described herein.
[0121] Further embodiments comprise a computer on which a computer program for performing one of the methods described herein is installed.
[0122] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to implement some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods are preferably executed by any hardware device.
[0123] The embodiments described above are only for illustrating the principles of the present invention. It is understood that modifications and variations of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the following claims, rather than by the specific details presented as descriptions and explanations of the embodiments herein.
[0124] References [1] “Adaptively Adjusting the Stereophonic Sweet Spot to the Listener’s Position”, Sebastian Merchel and Stephan Groth, J. Audio Eng. Soc., Vol. 58, No. 10, October 2010 [2] “AUDIO PROCESSOR, SYSTEM, METHOD AND COMPUTER PROGRAM FOR AUDIO RENDERING”, WO 2018 / 202324 A1 [3] https: / / www.princeton.edu / 3D3A / PureStereo / Pure_Stereo.html
Explanation of Signs
[0125] 1 Listener 10 Device, Audio Processor, Level 1 System 11 Audio Renderer 12 Loudspeaker Signal 14 Loudspeaker 15 Horizontal layer, loudspeaker layer 16 Interface 18 Audio signal 20 Object position input 21 Virtual position 30 Listener position input 31 Listener position 32 Loudspeaker position 40 Gain determiner 41 Gain 42 Roll-off gain compensation function 421 First compensated roll-off slope 422 Second compensated roll-off slope 44 Distance between listener and loudspeaker 44 12 Critical distance 441 First distance region 442 Second distance region 46 Distance compensation gain between listener and loudspeaker 47 Near-field - far-field transition parameter beta 481 First attenuation parameter 482 Second attenuation parameter 50 Delay determiner / controller 51 Delay 52 Controller 54 Median value 55 Delay processing unit 100 Rendering parameter 104 Object 110 Reverberation effect information 112 Reproduction space
Claims
1. An audio processor (10) for performing audio rendering by generating rendering parameters (100), wherein the rendering parameters (100) determine the derivation of a loudspeaker signal (12) from an audio signal (18) to be reproduced by a set of loudspeakers (14), the audio processor (10) is configured to perform a delay process to determine a delay (51) for generating the loudspeaker signal (12) for the loudspeakers (14) from the audio signal (18) based on a listener position (31), the audio processor (10) is configured to control the delay process by modifying (52', 52") a certain form of the listener position (31) or any intermediate value (54) determined by the delay process based on the listener position (31) so as to reduce artifacts in audio reproduction due to changes in the delay (51), and the delay process is started based on the certain form of the listener position (31), the audio processor (10).
2. The audio processor (10) is the listener position (31), the speed of the listener, the speed of the listener towards one or more of the set of loudspeakers (14), the acceleration of the listener, the acceleration of the listener towards one or more of the set of loudspeakers (14), the distance of the listener position (31) to one or more of the set of loudspeakers (14), the time rate of change of the distance of the listener position (31) to one or more of the set of loudspeakers (14), the rate of change of the time rate of change of the distance of the listener position (31) to one or more of the set of loudspeakers (14), the delay to one or more of the set of loudspeakers (14), the time rate of change of the delay to one or more of the set of loudspeakers (14), and the rate of change of the time rate of change of the delay to one or more of the set of loudspeakers (14) to one or more of smoothing, clipping, and scaling using a monotonically increasing function having a monotonically decreasing slope The audio processor (10) according to claim 1, configured to perform one or more of the above to implement the control of the delay process.
3. The audio processor (10) according to claim 1 or 2, wherein the audio processor (10) is configured to perform the delay process such that the delay (51) compensates for variations in the listener loudspeaker distance (44) between the loudspeakers (14).
4. The audio processor (10) according to any one of claims 1 to 3, wherein the audio processor (10) is configured to perform the delay process such that the listener position (31) becomes a sweet spot with respect to the set of loudspeakers (14) in an acoustic or perceptual sense.
5. The audio processor (10) according to any one of claims 1 to 4, wherein the audio processor (10) is configured to perform gain adjustment to determine a gain (41) for generating the loudspeaker signal (12) for the loudspeaker (14) from the audio signal (18) based on the listener position (31).
6. The audio processor (10) according to any one of claims 1 to 5, wherein the audio processor (10) is configured to perform gain adjustment by using a roll-off gain compensation function (42) for mapping the listener loudspeaker distance (44) of each loudspeaker to a listener loudspeaker distance compensation gain (46) for each respective loudspeaker.
7. The audio processor (10) according to claim 6, wherein the audio processor (10) is configured to perform the gain adjustment such that the listener position (31) becomes a sweet spot with respect to the set of loudspeakers (14) in an acoustic or perceptual sense.
8. The set of loudspeakers (14) belongs to one or more loudspeaker layers (15), and the audio processor (10) Desired sound source position of the audio signal (104 1 ) is between two loudspeaker layers (15), For each of the two loudspeaker layers (15), for the loudspeakers (14) belonging to the respective loudspeaker layer (15), a sound source position (104 1 ) of a desired audio signal to the respective loudspeaker layer (15), a virtual sound source position (104' 1 , 104'' 1 ) corresponding to the projection of) from the respective loudspeaker layer (15) belonging to the respective loudspeaker layer (15), apply 2D amplitude panning among the loudspeakers (14) of the respective loudspeaker layer (15) to determine a first pan gain (41) for the rendering of the audio signal (18) by the loudspeakers (14) belonging to the respective loudspeaker layer (15), When applied in addition to the first panning gain (41) to the loudspeaker layer (15), the sound source position (104 1 ) of the desired audio signal from the loudspeakers (14) of the two loudspeaker layers for rendering of the audio signal (18). Determine a second panning gain (41), the virtual sound source positions (104', 1 104'' 1 ) of the two loudspeaker layers (15) are configured to apply amplitude panning between them. The audio processor (10) according to any one of claims 1 to 7.
9. The set of loudspeakers (14) belongs to one or more loudspeaker layers (15), and the audio processor (10) The sound source position (104 2 ) of the desired audio signal is located outside the one or more loudspeaker layers (15), Among the one or more loudspeaker layers (15), the sound source position (104 of the desired audio signal 2 ) closest to the loudspeaker (14) of the closest loudspeaker layer (15) to the sound source position (104 of the desired audio signal to the closest loudspeaker layer (15) 2 ) corresponding to the projection of the virtual sound source position (104' 2 ) Apply 2D amplitude panning among the loudspeakers (14) belonging to the closest loudspeaker layer (15) to determine a first pan gain (41) for rendering the audio signal (18) by the loudspeaker (14) of the closest loudspeaker layer (15) from from the nearest loudspeaker layer (15) to a further virtual sound source position (104'' 2 ) offset towards the sound source position (104 2 ) of the desired audio signal, causing the loudspeaker (14) of the nearest loudspeaker layer (15) to reproduce sound from the further virtual sound source position (104''), applying further amplitude panning among the loudspeakers (14) belonging to the nearest loudspeaker layer (15) together with spectral shaping of the audio signal (18), the sound source position (104 of the desired audio signal 2 ) to cause rendering of the audio signal (18) by the loudspeaker (14) of the nearest loudspeaker layer from said 2 ), a second pan gain (41) for panning between the virtual sound source position (104' 2 ) and the further virtual sound source position (104'' 2 ) is determined, and further amplitude panning is applied between the virtual sound source position (104' 2 ), an audio processor (10) according to any one of claims 1 to 8, configured to apply further amplitude panning between the virtual sound source position (104' 2 ) and the further virtual sound source position (104'' 2 ).
10. The audio processor (10) performs the spectral shaping of the audio signal (18) using a first equalization function that simulates the timbre of the lower sound when the sound source position (104 2 ) of the desired audio signal is located below the one or more loudspeaker layers (15), and / or the sound source position (104 2 ) of the desired audio signal is located above the one or more loudspeaker layers (15), and performs the spectral shaping of the audio signal (18) using a second equalization function that simulates the timbre of the upper sound. The audio processor (10) according to claim 9, which is configured to perform the above.
11. The audio processor (10) performs the delay processing by determining the delay for each loudspeaker independently of the delay determined for any other loudspeaker in the set of loudspeakers (14), or determines a reference loudspeaker from the set of loudspeakers (14) and determines the delay (51) of the loudspeakers (14) other than the reference loudspeaker relative to the delay determined for the reference loudspeaker, and is configured to perform the delay processing, the audio processor (10) according to any one of claims 1 to 10. **Claim 12** The audio processor (10) is configured to perform the delay processing by determining the delay for each loudspeaker independently of the delay determined for any other loudspeaker in the set of loudspeakers (14) so as to obtain an absolute delay for each loudspeaker, The audio processor (10) the absolute delay for one or more of the loudspeakers (14) in the set of loudspeakers, the rate of change over time of the absolute delay for one or more of the loudspeakers (14) in the set of loudspeakers, and the rate of change of the rate of change over time of the absolute delay for one or more of the loudspeakers (14) in the set of loudspeakers for one or more of smoothing, clipping, and scaling using a monotonically increasing function having a monotonically decreasing slope performs one or more of to perform the control of the delay processing, the audio processor (10) according to any one of claims 1 to 11. **Claim 13** The audio processor (10) is configured to perform the delay processing by determining, for each loudspeaker, the distance of the listener position (31) to the position of each loudspeaker and determining the delay for each loudspeaker based on the distance, the audio processor (10) according to any one of claims 1 to 12. **Claim 14** The audio processor (10) the distance of the listener position (31) to one or more of the loudspeakers (14) in the set of loudspeakers, the rate of change over time of the distance to the listener position (31) for one or more of the set of loudspeakers (14), the rate of change of the rate of change over time of the distance to the listener position (31) for one or more of the set of loudspeakers (14), to one or more of: smoothing, clipping, and scaling using a monotonically increasing function having a monotonically decreasing slope The audio processor (10) according to claim 13, configured to perform one or more of the above to implement the control of the delay processing. **Claim 15** The audio processor (10) according to any one of claims 1 to 14, configured to control the delay processing according to control information and perform the correction according to the control information. **Claim 16** The audio processor (10) is configured to: derive from the control information one or more of: information about the strength of the smoothing, information about a clipping threshold for the clipping, information about the parameterization of the monotonically increasing function having a monotonically decreasing slope The audio processor (10) according to any one of claims 2, 12, and 14. **Claim 17** The audio processor (10) according to claim 15 or 16, configured to derive the control information from a bitstream. **Claim 18** The audio processor (10) according to claim 15 or 16, configured to derive the control information from side information of a bitstream and decode the audio signal (18) from the bitstream. **Claim 19** A method for audio rendering by generating rendering parameters (the rendering parameters (100) determine the derivation of loudspeaker signals (12) to be reproduced by a set of loudspeakers (14) from an audio signal (18), and the method includes: performing a delay process to determine a delay (51) for generating the loudspeaker signal (12) for the loudspeaker (14) from the audio signal (18) based on a listener position (31); A step of controlling the delay process by modifying a certain form of listener position (31) or any intermediate value (54) determined by the delay process based on the listener position (31) so as to reduce artifacts in audio playback due to changes in the delay (51), the delay process being started based on the listener position (31) of the certain form, and A method comprising. **Claim 20** A computer program (or a digital storage medium storing the computer program), the computer program having program code for instructing the computer to execute the method according to claim 19 when executed on the computer. **Claim 21** A bitstream according to any one of claims 1 to 20 (or a digital storage medium storing the bitstream).