Pre-compensation audio processing apparatus and method
The audio processing apparatus addresses sound field inaccuracies in vehicles by upmixing and applying frequency-specific pre-compensation filters, optimizing for individual speaker positions and signal types, enhancing clarity and authenticity.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-01-16
- Publication Date
- 2026-07-23
AI Technical Summary
Current audio reproduction systems, particularly in vehicles, fail to accurately recreate the original sound field due to issues like crosstalk, frequency response distortions, and nonlinearities, especially when listeners move from their designated seating positions, and conventional pre-compensation methods impose unnecessary restrictions leading to suboptimal results.
An audio processing apparatus that upmixes input signals into multiple channels, applies frequency-specific pre-compensation filters, and optimizes for individual speaker positions and signal types, allowing for more flexible and effective sound reproduction by distinguishing between ambient and non-ambient signals, and accounting for head shadowing effects.
Enhances audio clarity and authenticity by optimizing sound reproduction for different listener positions and signal types, improving stage resolution and reducing frequency and phase artifacts, resulting in a more immersive listening experience.
Smart Images

Figure EP2025051011_23072026_PF_FP_ABST
Abstract
Description
[0001] PRE-COMPENSATION AUDIO PROCESSING APPARATUS AND METHOD
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to audio processing in general. More specifically, the disclosure relates to an audio processing apparatus and method for pre-compensation control of an audio rendering system, such as a spatial-surround sound system, in particular inside of a vehicle.
[0004] BACKGROUND
[0005] Loudspeaker based audio reproduction suffers from a wide variety of problems. The most important ones are that their frequency response is not flat, they do not behave linear but cause distortions and they cannot recreate the original sound field accurately. Mechanical improvements of loudspeakers over the last decades resulted in that the problems of the frequency response and nonlinearities were reduced drastically, but there is still no loudspeaker based sound system, regardless if it is a stereo or 9.1 setup, which can recreate the original sound field accurately. Assuming that currently most audio reproductions are based on stereo signals, the focus herein will be mainly on stereo signals, but the same solution can be applied for multichannel- or object-based audio signal encodings similarly.
[0006] Regardless of the input format, there are two main problems causing the sound field reproduction to fail. One is, that the sound emitted from a loudspeaker is not only audible on one ear but on both. So, if an audio engineer wants to create a sound that is audible only on the left side, the best he can do is mixing the sound to the left channel. If this is then reproduced with a standard stereo loudspeaker setup, the sound will not be perceived from the left, but from the left speaker which usually is only at a 30-degree angle. Multichannel systems are able to reduce the impact from this effect but also cannot completely solve it.
[0007] Things get even worse if a signal in front of the listener shall be created. Audio engineers mix such sounds to both speakers. If the listener is placed in a centered position between both, he will perceive the same signal on the left and right ear, so it seems everything should be fine. But, as mentioned before, the signal sent to the left speaker does not only arrive at the left ear, but with some delay on the right ear. The right speaker signal will do the same, so first arrive at the right ear and with a delay at the left ear. What the listener perceives is thus not the original signal any more, but this signal and a delayed version. Even if the delay is very short, the frequency response is changed which is audible. The problem, that the sound of one speaker can be perceived at both ears, can be solved by emitting inverse signals with some time delay to cancel out the undesired signals. This is a well-known principle which is called crosstalk cancellation. The main reason, why it is practically never used in any audio system, is, that it requires that the position of the speakers and the listener are known and static. If the listener just moves to another seat, the delays do not match any more and the compensation will fail.
[0008] Crosstalk relates to the sound that arrives at the ears directly but not through any wall or other surface reflection. Such reflections have an even stronger impact on the sound field reproduction. The human auditory system will perceive such reflections and by evaluating them create an impression of the room. So, listening especially in small environments like a car cabin, one can clearly perceive that the sound is played back in that room but not get the impression to be in the original recording environment which might be an opera, jazz club or even just a recording studio.
[0009] To improve the sound field reproduction, several technologies have been developed. They usually upmix the audio signal into a multichannel format matching the speaker layout and then measure the sound field by placing a plurality of microphones at different measurement positions and then run some optimization towards a set of desired target functions at the different measurement positions. This results in having a set of so-called pre-compensation filters. Filtering the signals that shall be sentto the speakers with those pre-compensation filters will then improve the sound. It is also possible to add to the signals, that are sent to the speakers, filtered signals from other channels, i.e. to sum up the pre-compensation filter outputs for individual speaker channels. The approach of measuring the sound field with arrays of microphones usually does not distinguish the direction of arrival of different sounds which is important. As an example, a sound at one measurement position might be optimized to be emitted from a speaker on the right side, but if the measurement position corresponds to the left ear of a listener, then in reality it will be damped in level due to the head of the passenger being in the transmission path which then could lead to the signal being perceived louder on the right ear. This problem can be avoided by also optimizing for the direction of the arrival of the sound. But then things get very complicated in that the head- and body shadowing effects still cause issues and only a very suboptimal result may be achieved.
[0010] There is one solution to this problem which in theory is perfect. This is to place a dummy head and corso at the listening place and to capture the real sound perceived at both ears. By measuring all transfer functions from all speakers to both ears, an optimization is possible with nearly arbitrary target functions defining how the left input signal should arrive at the left and right ear and vice versa for the right signal. If in such target functions the crosstalk components and surface reflections are removed, then these will be eliminated through the optimization. But, in reality, this causes a lot of new problems. The main problem is that the optimization will normally only work in case the real listener is seated always at exactly the same place. A sine signal of 2kHz has a wavelength of around 17cm in air, so moving the listener position by just 8.5 cm would invert the signal completely and this way destroy the optimization.
[0011] Another problem current audio pre-compensation implementations have is that they just use one upmixer which usually matches the speaker layout. The outputs from such upmixer are then used as the basic signals to be sent to the speakers but also for the pre-compensation filtering. This causes a lot of problems. For instance, a signal whose position is not matching one of the real speakers needs to be panned somehow to the existing speakers. Assuming that a signal whose angle in reality is between a center speaker and a left speaker is panned half to the center and half to the left speaker, this can create the sensation that there are two sound sources. Even if level and time correction are implemented, then the speaker locations will still cause improper summation at the ears, the real transfer function such a sound source would create cannot be met. The reason for this is that the transfer functions from both speakers to the ears are not identical, so some frequency and phase artifacts occur. Another problem relates to the ambience present in most recordings. This should not be panned at all, because it is a diffuse sound and thus should stay diffuse. Assigning this to speakers and applying pre-compensation will cause coloration of the ambient signal. Another problem is related to the fact that all optimizations, also the one to calculate audio pre-compensation filters, work better the less unnecessary restrictions are applied during the optimization process. Just panning signals to the speakers and then trying to optimize this causes such unnecessary restrictions to be applied. A good example for this is a signal which can only be audible on one side, for such signals the phase relation between the left side and right side can be ignored. Conventional algorithms do not exploit this effect and thus can only deliver suboptimal results.
[0012] SUMMARY
[0013] It is an objective to provide an improved audio processing apparatus and method for pre-compensation control of an audio rendering system, such as a spatial-surround sound system, in particular inside of a vehicle.
[0014] The foregoing and other objectives are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0015] As used herein a speaker or loudspeaker may comprise a set of woofer, midrange and tweeter drivers which can be arranged in one or more enclosures at different locations. In the automotive industry it is often the case that what is called a front leftspeaker is a combination of a front left woofer driver which is installed in the front left door panel in the bottom, a front left midrange driver installed in the upper front part of the left door and a front left tweeter driver installed in the A-pillar of the car.
[0016] As used herein a loudspeaker driver may be a single loudspeaker chassis, for instance, a subwoofer driver, a woofer driver, a midrange driver or a tweeter driver. To avoid the need to specify drivers by writing all their properties like “Front Left Woofer”, instead abbreviations are used consisting of the first letter of the words. Thus, a front left woofer is abbreviated as FLW. Thus, the following abbreviations will be used herein for loudspeaker drivers:
[0017] F Front
[0018] B Back (this usually refers to speakers installed in the back doors of a vehicle)
[0019] S Surround (this is usually the location behind the rear passengers like on the parcel shelf)
[0020] L Left
[0021] R Right
[0022] C Center
[0023] W Woofer
[0024] M Midrange
[0025] T Tweeter
[0026] A subwoofer driver is herein referred to as “SUB”. Filters that are applied will use the common technical letter indicating a transfer function which is “H”:
[0027] H Transfer function
[0028] Thus, a filter for the front left woofer will be abbreviated as “HFLW”. A filter like “HLR” indicates a filter that is identical for the left and right channels.
[0029] According to a first aspect an audio processing apparatus is provided for generating a plurality of output signals for driving a plurality of loudspeakers based on one or more input signals. The audio processing apparatus according to the first aspect is configured to upmix the one or more input signals into a first plurality of upmixed input signals and to upmix the one or more input signals into a second plurality of upmixed input signals. Moreover, the audio processing apparatus according to the first aspect is configured to filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with one or more pre-compensation filters. The audio processing apparatus according to the first aspect is further configured, for each of the plurality of loudspeakers, to add the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective output signal. Thus, the audio processing apparatus according to the first aspect allows to calculate and use more suitable pre-compensation filters. As described in the background section, the optimization process of conventional systems is not really matching the required properties, they usually only optimize for a given number of loudspeaker positions with identical parameters for all of them. With the approach implemented by the audio processing apparatus according to the first aspect, an optimization for any number of audio sources is possible, not only for those where a physical speaker can be found. For different positions different optimization criteria may be used. The optimization may also take signal properties into account, like distinguishing ambient and non-ambient signals. This better and less restrictive match of the resulting compensation filters causes a better compensation. This means the clarity of the sound reproduction is increased, the stage resolution is improved, so that a much more authentic audio reproduction becomes possible by the audio processing apparatus according to the first aspect.In a further possible implementation form, the audio processing apparatus according to the first aspect is configured to upmix the one or more input signals into the first plurality of upmixed input signals based on a configuration of the plurality of loudspeakers and / or the audio processing apparatus is configured to upmix the one or more input signals into the second plurality of upmixed input signals based on one or more characteristics of the one or more input signals.
[0030] In a further possible implementation form, the audio processing apparatus according to the first aspect is configured to filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more precompensation filters by applying one or more frequency range optimized pre-compensation filters in a plurality of different frequency ranges to the first plurality of upmixed input signals and / or the second plurality of upmixed input signals.
[0031] In a further possible implementation form, the different frequency ranges are selected to match the frequency ranges of the plurality of loudspeakers.
[0032] In a further possible implementation form, the audio processing apparatus according to the first aspect is further configured to bandpass filter different copies of the first plurality of upmixed input signals and / or the second plurality of upmixed input signals.
[0033] In a further possible implementation form, the audio processing apparatus according to the first aspect is configured to bandpass filter the different copies of the first plurality of upmixed input signals and / or the second plurality of upmixed input signals before filtering the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more pre-compensation filters.
[0034] In a further possible implementation form, the one or more pre-compensation filters are based on one or more impulse responses.
[0035] In a further possible implementation form, the one or more impulse responses are measured taking the head shadowing effect into account.
[0036] In a further possible implementation form, the plurality of loudspeakers comprises one or more front loudspeakers, wherein for generating the one or more output signals for the one or more front loudspeakers the audio processing apparatus is configured to filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more pre-compensation filters such that a symmetric front stage is achieved.
[0037] In a further possible implementation form, the one or more pre-compensation filters are based on an optimization of a selected target of the individual compensation filter which is defined by the frequency band dependent listening properties, signal type classifications, loudspeaker types and / or loudspeaker position information.
[0038] In a further possible implementation form, the one or more pre-compensation filters are configured such that phase differences are distinguishable at both ears of the listener up to about 800 Hz and / or such that time structures at higher frequencies than about 2 kHz are indistinguishable.
[0039] In a further possible implementation form, the audio processing apparatus according to the first aspect is configured to upmix the one or more input signals into the first plurality of upmixed input signals, including a center, a left, a right, a far left, a far right, an ambient left, an ambient right, a surround left and / or a surround right signal.In a further possible implementation form, the one or more pre-compensation filters are configured to perform an optimization for a respective frequency range of the respective loudspeaker.
[0040] In a further possible implementation form, the one or more pre-compensation filters are optimized with respect to a respective position of each loudspeaker.
[0041] In a further possible implementation form, the configuration of the plurality of loudspeakers includes a subwoofer, a woofer, a midrange and / or a tweeter loudspeaker configuration.
[0042] In a further possible implementation form, the configuration of the plurality of loudspeakers comprises a front center, front left, front right, back left, back right, surround left and / or surround right loudspeaker configuration.
[0043] In a further possible implementation form, for upmixing the one or more input signals into the first plurality of upmixed input signals the audio processing apparatus according to the first aspect is configured to generate a center signal by summing up a left signal and a right signal.
[0044] In a further possible implementation form, for upmixing the one or more input signals into the second plurality of upmixed input signals the audio processing apparatus according to the first aspect is configured to upmix the one or more input signals into a far left, a left, a center, a right, a far right, a surround left, a surround right, an ambience left, an ambience right and / or a subwoofer signal.
[0045] In a further possible implementation form, the one or more bandpass filters comprise a bandpass filter in a subwoofer frequency range, a bandpass filter in a woofer, i.e. midrange frequency range, and / or a bandpass filter tweeter frequency range.
[0046] In a further possible implementation form, the audio processing apparatus according to the first aspect comprises a crossover filter configured to split the output of the bandpass filter in the woofer, i.e. midrange frequency range into one or more further frequency ranges.
[0047] In a further possible implementation form, for upmixing the one or more input signals into the first plurality of upmixed input signals the audio processing apparatus according to the first aspect comprises a delay unit configured to delay one or more of the upmixed input signals for alignment with the output of pre-compensation filters of the second upmixer output.
[0048] According to a second aspect an audio rendering system, in particular an automotive audio, i.e. sound rendering system, is provided, wherein the audio rendering system according to the first aspect comprises a plurality of loudspeakers and an audio processing apparatus according to the first aspect for generating a plurality of output signals for driving the plurality of loudspeakers based on one or more input signals.
[0049] According to a third aspect an audio processing method is provided for generating a plurality of output signals for driving a plurality of loudspeakers based on one or more input signals. The audio processing method according to the third aspect comprises:
[0050] upmixing the one or more input signals into a first plurality of upmixed input signals;
[0051] upmixing the one or more input signals into a second plurality of upmixed input signals;
[0052] filtering the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with a precompensation filter; andfor each of the plurality of loudspeakers adding the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective output signal.
[0053] According to a fourth aspect a computer program product is provided, comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method according to the third aspect, when the program code is executed by the computer or the processor.
[0054] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0055] BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In the following, embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which:
[0057] Fig. 1 is a schematic diagram illustrating an audio processing apparatus according to an embodiment;
[0058] Fig. 2 is a schematic diagram illustrating the principle of a mid-side decomposition upmixer implemented by an audio processing apparatus according to an embodiment;
[0059] Fig. 3 is a schematic diagram illustrating in more detail an audio processing apparatus according to an embodiment with a post processing stage after an asymmetric audio pre-compensation stage;
[0060] Fig. 4 is a schematic diagram illustrating in more detail an audio processing apparatus according to an embodiment with a post processing stage after a symmetric audio pre-compensation stage;
[0061] Fig. 5 is a schematic diagram illustrating in more detail an audio processing apparatus according to an embodiment with a post processing stage prior to a symmetric audio pre-compensation stage; and
[0062] Fig. 6 is a flow diagram illustrating steps of an audio processing method according to an embodiment.
[0063] In the following, identical reference signs refer to identical or at least functionally equivalent features.
[0064] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0066] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform thedescribed one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.
[0067] Figure 1 is a schematic diagram illustrating an audio processing apparatus 100 according to an embodiment configured to generate a plurality of output signals (referred to as the loudspeaker output signals 1 to N in figure 1 ) for driving a plurality of loudspeakers based on one or more input signals (referred to as the audio input in figure 1 ). The audio processing apparatus may be a component of an audio rendering system further comprising the plurality of loudspeakers. In an embodiment, the audio rendering system may be a spatial surround sound system, in particular an automotive spatial surround sound system, i.e. configured to be implemented and operated for generating surround sound, i.e. a soundfield in the interior of a vehicle, e.g. car. The audio processing apparatus 100 may comprise processing circuitry for generating the plurality of loudspeaker output signals 1 to N for driving the plurality of loudspeakers based on the one or more input signals, which may be implemented in hardware and / or software and may comprise digital circuitry, or both analog and digital circuitry. Digital circuitry may comprise components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or general-purpose processors. The audio processing apparatus 100 may further comprise a memory configured to store executable program code which, when executed by the processing circuitry, causes the audio processing apparatus 100 to perform the functions and methods described herein.
[0068] As will be described in more detail below, in particular in the context of the embodiments shown in figures 3, 4 and 5, the audio processing apparatus 100 is configured to implement the following functionality: upmix the one or more input signals into a first plurality of upmixed input signals and a second plurality of upmixed input signals; filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with one or more pre-compensation filters; and, for each of the plurality of loudspeakers, to add the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective loudspeaker output signal.
[0069] In the embodiment shown in figure 1 this functionality is implemented by the processing stages 110, 120 and 130. More specifically, in the embodiment of the data processing apparatus 100 shown in figure 1 the one or more audio input signals are provided to a delay stage 105 which is configured to compensate for any delay caused by the downstream pre-compensation filters. The output of the delay stage 150 is then upmixed by the first upmixer stage 110 into the first plurality of upmixed input signals (referred to as Signals 1 to J in figure 1 ). As will be described in more detail below, the first upmixer stage 110 may be configured to perform the upmixing into the J signals based on a configuration of the plurality of loudspeakers, e.g. for matching one or more properties of the plurality of loudspeakers, such as a frequency ranges and / or mounting positions of the plurality of loudspeakers. Moreover, the one or more audio input signals are provided to the second upmixer stage 120 for upmixing the one or more audio input signals into the second plurality of upmixed signals (referred to as Signals 1 to I in figure 1 ). As will be described in more detail below, the second upmixer stage 120 may be configured to perform the upmixing into the I signals based on one or more characteristics of the audio input signals, such as being, for instance, a centered, far left, left, ambient, or surround audio input signal. Then, based on the outputs of the second upmixer stage 120, pre-compensation filters are applied by the pre-compensation stage 130 to each combination of a signal from the first upmixer stage 110 and the second upmixer stage 120 and the results for each signal of the first upmixer stage 110 summed up. These signals may then be provided to apostprocessing stage 140 for some optional postprocessing, such as remixing, filtering, compressing, limiting, linearizing and the like.
[0070] As will be described in more detail below, the audio processing apparatus 100 according to embodiments disclosed herein allows providing a more authentic audio reproduction by using a special pre-processing of the input audio signal(s) to separate signal components, a special optimization approach, bandlimited processing and taking psychoacoustic effects into account. A first basic principle implemented by the data processing apparatus 100 according to an embodiment is to provide a first and a second upmixer stage 110, 120, wherein the first upmixer stage 110 is configured to determine a first set of output signals, which in an embodiment may correspond to the given loudspeaker layout, i.e. configuration, while the second upmixer stage 120 is configured to generate another set of upmixed signals, which in an embodiment may be optimized to have similar properties. As will be appreciated, the second plurality of upmixed input signals generated by the second upmixer stage 120 does not need to match the loudspeaker layout used by the first upmixer stage 110. As already described above, each of the output signals of the second upmixer stage 120 may then be filtered with frequency band-specific filters for each of the output signals of the first upmixer stage 110 and the results summed up for each loudspeaker.
[0071] In the following more detailed embodiments of the first upmixer stage 110 and the second upmixer stage 120 as well as of the pre-compensation filtering stage 130 of the audio processing apparatus 100 are described. Moreover, an approach for determining suitable pre-compensation filters for the pre-compensation filtering stage 130 will be described.
[0072] In an embodiment, the first upmixer stage 210 is configured to create a symmetric sound stage, as illustrated in figure 2. As can be taken from figure 2, in this embodiment, the perceived front sound zone for listener 1 ranges approximately from the left speaker 220a to the right speaker 220c. In a car the driver will be sitting at the position of listener 1 , i.e. left of the middle of the car. Listener 1 will perceive that the left angle is steeper than the right angle. But Listener 1 will perceive a centered signal directly in front of him. The co-driver on the right side will also perceive the stage from the left speaker 220a to the right speaker 220c. Listener 2 will also perceive the center directly in front of him, but in his case the right angle is steeper than the left angle. Both will perceive left sounds from the left and right sounds from the right. The upmixer 210 in this case is a very simple one where the center signal is created by just summing the left and right signals. This is a very widely used approach, which achieves a good front stage impression, just has the disadvantage that the width of the stage is limited to the locations of the left and right speakers.
[0073] The upmixer 21 Oof figure 2 is a first possible embodiment of the first upmixer stage 110 of the audio processing apparatus 100. According to an embodiment, the first upmixer 210 preferably does not apply nonlinearities (a wide class of algorithms use so called “steering” which tries to evaluate signal directions and then boosts those in a nonlinear way). According to a further embodiment, the output signals may match the loudspeaker configuration. As will be appreciated, the upmixer 21 Oillustrated in figure 2 does not try to split up the input signal into signals like left / right / center / ambience with perfect reconstruction properties, it also does not try to separate different stems (instruments, voices etc.), it only focuses on creating an upmix of the input signals that is suitable to achieving a symmetric soundstage. One of the first implementations for such an upmixer 21 Owas based on the so-called M / S = mid-side decomposition. For a M / S upmixer 210, the “mid” signal is the sum of the left and right, the “side” signal is the difference (left-right). These mid and side signals are then further processed and mixed towards the different speaker channels. One straightforward solution is to use (M+S) for the left speaker, M for the center, (M-S) for the right speaker, S for the left surround and -S for the right surround. This already achieves a nice surround sound impression with a symmetric soundstage and can be realized with just 5 speakers. More modem upmixers use further properties like transient detection and coherence evaluation for such signal decompositions. These are the preferred implementations for the upmixer 210, as they achieve a better decoupling of the signal components. The filters HS, HM of the filter stage 215 illustrated in figure 2 are used to generate the desired frequency response but also to achieve a symmetric front zone.As already described above, according to an embodiment, the second upmixer stage 120 is configured to separate the input audio signal(s) into signal components with similar signal properties (i.e. having a different purpose than the first upmixer stage 110, which is matching the loudspeaker setup). One possible embodiment of the second upmixer stage 120 is an upmixer configured to decompose the audio input signal(s) into front far left / left / center / right / far right, back left / right, surround left / right and ambience left / right. Further possible embodiments of the second upmixer stage 120 are the publicly available Al based upmixers like the Spleeter from Deezer or OpenUnmix, which are able to split-up the input signal(s) into different stems representing the different instruments, effect sounds and voices. For the second upmixer stage 120 the perfect reconstruction property is desired (but not mandatory). This decomposition is used to group signals with similar properties together and then use these signals as the source for audio pre-compensation filtering implemented by the processing stage 130. Audio precompensation in this case means that the upmixed signals are filtered in some way and then added to the signals before sending them to the loudspeakers. As these upmixer signals share the same properties, the pre-compensation filters can be individually optimized for these properties. As one example, extracting a centered signal, the main property is that it is centered. The optimization criteria for centered sounds could be that it arrives at both ears of a listener with the same phase, group delay and frequency response (in reality, the criteria are more complex as, for instance, also time resolution and frequency dependent hearing effects may influence the result). Another example is a signal that shall be audible from the left side. For such a signal the relation of phase, group delay and frequency response on both ears are not that relevant any more as the sound is mostly only audible on the left ear. Thus, in this case, the main optimization criteria used is to damp the sound on the right side. This decomposition thus allows to use individual optimization criteria for each signal which are limited to only do the mandatory optimizations. This relaxes the optimization constraints and, in this way, leads to significantly better optimization of the individually desired targets.
[0074] In an embodiment, the audio processing apparatus 100 is configured to make advantageous use of the behavior of the human auditory system, which evaluates sounds differently in different frequency bands. For very low frequencies (e.g. below around 150 Hz) the human auditory system is not able to perceive a direction information. Thus, for such frequencies, compensation filtering is not really required. In the range from 150 Hz to around 900 Hz, the human auditory system is sensitive to amplitude and phase, so both should be considered for the compensation filtering implemented by the pre-compensation stage 130 of the audio processing apparatus 100. Above 900 Hz, the human auditory system mostly relies on time-of-arrival and amplitude evaluation so that the phase can be mostly ignored. The values mentioned here are only examples, as the transitions are not happening abruptly and even differ from person to person and for different settings. Thus, in an embodiment, the precompensation stage 130 of the audio processing apparatus 100 is configured to use different optimization criteria for different frequency bands. To achieve such a frequency band dependent optimization, one way is to calculate band-limited compensation filters and to apply those on the full-range compensation signals from the second upmixer stage 120. Alternatively, the signals from the second upmixer stage 120 may be band-limited and then send to compensation filters matching the specific band. According to further embodiments, the individual filters may be summed up, first full range filters may be used and then the results band limited or all these variants even implemented in the frequency domain by fast convolution. In practice, also pre-and postprocessing filters may influence the operation. The output of the first upmixer stage 110 may be post-processed by audio processing modules like equalizers, delays and gains. The compensation filters may be applied before such postprocessing or after the post processing.
[0075] Figure 3 illustrates in more detail the audio processing apparatus 100 according to an embodiment with a post processing stage after an asymmetric audio pre-compensation stage. In other words, in the embodiment of figure 3 the bandpass filtering is done in the postprocessing after adding the pre-compensation filter results, while no symmetry constraints are applied to the audio pre-compensation filtering.As can be taken from figure 3, in this embodiment the first upmixer stage 310 is a simple upmixer where the center signal is created by summing the left and right input signals. The input to this upmixer 310 is delayed by the delay element 305 causing a N-sample delay. This may be required to compensate for the delay caused by the compensation filtering. To simplify the drawing, only the outputs for the front speakers are shown and the upmixer is shown individually for both frequency bands, even if it only needs to be computed once. To the outputs of the first upmixer 310, which is basically the passthrough of the left and right delayed signals and the summation of both to form the center signal, the outputs of the pre-compensation filters 330a and 330b of the second upmixer 320 are added. The precompensation filters 330a are used for filtering the first frequency band, the precompensation filters 330b are used for filtering the second frequency band. Finally, the postprocessing stage 340a implements, for instance, bandpass filtering by means of the filters HLO, HM0, HR0, HL 1 , HM1 and HR1. In case full-band operation is desired, then the generation of the signals L 1 , C 1 and R1 may be omitted.
[0076] In the embodiment shown in figure 3, the outputs L0, CO and R0 might be signals covering any frequency range. If the frequency ranges do not match those of the loudspeaker drivers, then the relevant bands may need to be summed up and, if this causes the frequency range to be too wide, even bandpass filtered. In an embodiment, the frequency bands may be set to match the speaker drivers or a group of speaker drivers like woofer and midrange. The frequency range of the filters HLO and HM0 may also differ if the center speaker cannot cover the same frequency range as the speakers on the sides (which is often the case).
[0077] The same principle may be used to also do the filtering in the frequency domain. In such a case, blocks of the input signals L and R, which may overlap and use windowing, may be first transformed to the frequency domain using e.g. a windowed DFT. Then from a desired number of subsequent transformed block results for each frequency bin the processing according to one of the bands may performed. Assuming the filtering is done on 8 subsequent blocks, 8 frequency bin values for the left and 8 frequency bin values for the right channel may be processed for band 0, so the band that will compute L0, CO and R0. The band processing is then similar to the band processing of figure 3, this means that the center signal can be created by summation and the filtering then also performed using filters of length 8. The main difference is that, in case of the DFT, the values are complex, so complex multiplications and additions have to be used. At the output side, the resulting filtered output values are transformed back into time domain and, if needed, additional windowing and overlapping performed. In this case, the post-processing filters may be efficiently executed in the frequency domain, i.e. before the inverse transform. In this way the convolution turns into a simple multiplication.
[0078] Thus, the embodiment of the audio processing apparatus 100 of figure 3 implements a M / S upmixing on the top 310, a signal property based upmixing 320 on the left, a band separation according to the loudspeaker capabilities. The post processing 340a and 340b is done after the pre-compensation filtering stage 330a and 330b and a non-symmetric pre-compensation filtering 330a and 330b is used. As will be appreciated, the embodiment of the audio processing apparatus 100 of figure 3 is very versatile, may be optimized for a single seat or multiple seats, and is very efficient due to the band limiting within the post processing filters 340a, b. In comparison with conventional pre-compensation filtering approaches the embodiment of the audio processing apparatus 100 of figure 3 may achieve a substantial improvement of the audio quality in terms of clarity, stage width perception, stage depth perception, tightness of the bass and the like. This is mainly achieved by using specific optimization targets for the different sound components which allows for more efficient and psycho-acoustically matched optimization.
[0079] As will be appreciated, in the embodiment of the audio processing apparatus 100 illustrated in figure 3, no symmetry constraints are applied. This is ideal if an optimization for just one listener has to be done, because this may achieve the best possible optimization. In a car environment, the loudspeaker positions and seat positions are symmetric. Minor deviations may occur due to having the steering wheel only on one side and passengers adjusting the seat positions differently. These changes are sufficiently small to still allow to exploit this symmetry. This can be done by forcing the optimization to use a symmetric filtersetup as shown in the embodiment of the audio processing apparatus 100 illustrated in figure 4, which implements post processing after symmetric audio pre-compensation 430a and 430b.
[0080] As indicated in figure 4, in this embodiment the filters may be mirrored around the center. This will then lead to a symmetric sound stage for both the driver and the co-driver, i.e. the person sitting next to the driver. But again, the extra symmetry constraint may reduce the compensation quality. In practice, the symmetric optimization may thus only be used if two passengers are sitting in the front. If the driver is alone, then the asymmetric optimization may be chosen. This may be realized by either allowing a manual selection of the preferred compensation or by an automatic detection of the occupancy. Modem cars have sensors in the seats which detect which seats are occupied. This information may directly be used by the audio processing apparatus 100 to automatically select if the symmetric or asymmetric compensation shall be activated.
[0081] As already mentioned above, in the embodiment of the data processing apparatus 100 illustrated in figure 4 all filters are symmetric to the center. This restriction forces a symmetric sound stage. It also allows further mathematical optimizations, because, for instance, the L and R signals are both added to the same output using the filter W5 twice. As will be appreciated, this operation may be simplified by first summing up the channels and then only filtering once. Otherwise, the operation is pretty identical to the one of figure 3. This means there is a pre-delay 405 to time align the signals of the first upmixer 410 with the results of the pre-compensation filter stages 430a and 430b which are working on the outputs of the second upmixer 420. The filter stage 430a again covers one frequency band, filter stage 430b another frequency band.
[0082] The embodiments of the audio processing apparatus 100 of figures 3 and 4 apply the post-processing filters like HS0 and HM0 after the audio pre-compensation stage. But it is also possible to apply them directly after the first upmixer stage. This approach is used in the embodiment of the audio processing apparatus 100 illustrated in figure 5, which, thus, implements post processing before a symmetric audio pre-compensation. Otherwise, the operation is again pretty identical to the one of figure 3. This means there is a pre-delay 505 to time align the signals of the first upmixer 510 with the results of the pre-compensation filter stages 530a and 530b which are working on the outputs of the second upmixer 520. The filter stage 530a again covers the first frequency band, filter stage 530b the second frequency band. The band filtering is done in the filters 545a for the first frequency band and in filters 545b for the second frequency band.
[0083] The embodiment of the audio processing apparatus of figure 5 has some advantages and some disadvantages compared to the embodiment of figure 4. One major disadvantage is that the bandpass filtering, which is done in the filters HS0, HM0, HS1 and HM1 after summing up the filtered upmixer outputs from stage 540a and 540b with the outputs of the pre-compensation filters 530a and 530b, is now not any more in the signal chain of the pre-compensation signals as it is moved above. This is addressed by the embodiment of figure 5 by adding bandpass filters to the outputs of the second upmixer stage before applying the precompensation filtering and summation. As mentioned before, a variety of similar solutions are possible like integrating the bandpass filtering into the compensation filters. Another major disadvantage also results from not having the filters HS0 and the like at the signal output. These filters are used to flatten the frequency response and also already perform some timealignment on the signals. The pre-compensation filter now has to do all those tasks on top, so that the filter will become more complicated, e.g. require more filter taps. Depending on the optimization approach, the additional complexity may also cause the optimizer to achieve worse overall results.
[0084] An advantage of the embodiment of figure 5 is that the pre-compensation filter outputs are not delayed by the filters HS0 and the like, which means that at similar overall latencies the compensation quality, from a theoretical point, might be superior. This also becomes obvious, as the pre-compensation filters might also need to compensate for latencies or frequency response issues caused by the filters HS0 and the like. In practice, the approach from figure 4 is beneficial if the processing power available is critical. Otherwise, the achievable results from both approaches may be measured and the better one may be used.In an embodiment, the audio processing apparatus 100 may use for each frequency band filters with frequency-band specific optimization criteria. Moreover, for each speaker type and location, individually optimized filters may be applied. As will be appreciated, separating uncorrelated signals like the ambience simplifies the filter design. Moreover, separating the inputs signals into signal components with similar properties allows to develop filters specialized for such properties. Thus, the filters of the audio processing apparatus 100 may be designed with a limited set of requirements compared to a conventional approach. By having less requirements, the filter optimization becomes easier and can achieve better results for the reduced requirements. Conventional approaches usually define a single target and then run one optimization for all filters at once to reach the target.
[0085] How the targets for the filter optimization are defined may depend on the available loudspeaker setup, the used upmixers and also on further factors like the number of sitting locations to which it shall be optimized. In the following, an example is described for an optimization for a loudspeaker setup in a car with a 5.1 style setup, i.e. in each front- and back door a combination of a woofer, a midrange and a tweeter driver, in the center of the dashboard a combination of a midrange and tweeter driver and a subwoofer usually in the trunk. For this example, the optimization is performed for a symmetric setup, post filtering after pre-compensation filtering according to the embodiment of figure 2. The output signals L0, CO and RO are used for the speaker setup consisting of woofer and midrange, the output signals LI, Cl and R1 are sent to the tweeters. The upmixers are assumed to be identical to the ones of the embodiment of figure 5. In this case, the optimization may be performed in the following way.
[0086] An ideal optimization for the filters W6, W7 and W8 would aim at achieving that a far left signal shall be audible on the left ear and right ear according to a free-field measurement for a sound coming from the left side, so that the right ear’s sound will be damped significantly and also changed in frequency response and time structure due to the head shadowing effect. In practice it might be beneficial to just optimize for the target that the right side is damped by around 6 dB which roughly corresponds to the damping of the head-shadowing, but not to place constraints on the detailed frequency response and time structure of the signal arriving at the right ear. The frequency response and time structure are only important for the strong signal arriving at the left ear. Thus, this may be a second, but lower priority optimization target. The time structure should be in a way that the arrival time matches those from the other speakers and that the energy of the impulse response decays as fast as possible over all frequencies.
[0087] For the filters W3, W4 and W5 the priority shifts already a little as these signals are not any more audible just on one side. Thus, the separation might only need to achieve a 4dB separation, the other requirements are identical to those of filters W6, W7 and W8. If weights for the targets are used, one might increase the weight for the frequency response and time structure for the right ear as this signal becomes more audible.
[0088] The filters W1 and W2 then have different optimization criteria. Their purpose is to handle a centered signal. For such a signal, the main target is that the frequency response and time structure on both ears is identical, otherwise the sound would not be perceived in the center. As centered signals are often voice signals, it is pretty important that the frequency response matches the target response. Last but not least, the time structure targets from the previous points should also be met.
[0089] Filters W14, W15 and W16 are again working on far left / far right signals, but here for the tweeter drivers. The human auditory system is not able to evaluate the fine time structure or phase relations of tweeter signals, it may only detect the time of arrival and a somehow smoothed frequency response. Thus, the main target is damping a far left signal on the right ear by around 6 dB, achieving the desired time of arrival, having a fast decay and having the smoothed frequency response matching the target response.F or filters W 11 , W12 and W 13 nearly the same targets as for filters W3 , W4 and W 5 are valid. Again, due to the ear not being able to evaluate the fine time structure and phase relations for such frequencies, only the smoothed frequency response needs to be considered.
[0090] For filters W9 and W10 the fine time structure can again be ignored and also only the smoothed frequency response evaluated. Thus, the targets are that the smoothed frequency response on both ear matches, the smoothed frequency response matches the target frequency response, the arrival time matches those of the other speakers and the decay to an impulse is as short as possible.
[0091] As will be appreciated, the example above just focuses on the speakers in the front, i.e. ignored the back speakers and the subwoofer from the first upmixer, and does not describe explicitly how to use the remaining signals from the second upmixer, such as the surround and ambience signals, which, however, may be implemented in a straightforward manner by the person skilled in the art. For instance, if a surround signal component detected by the second upmixer 120 shall not be audible on the front speakers, another row for the surround processing may be added and then the target set to simply remove this signal from the front outputs. In an embodiment, weights may be defined for all combinations of the signal bands, the first upmixer signal output groups, the second upmixer outputs and the different possible optimization criteria.
[0092] As will be appreciated, embodiments of the audio processing apparatus 100 disclosed herein address one or more of the following shortcomings of conventional approaches.
[0093] In conventional devices the measurement is done using arrays of microphones placed at each seat. This causes a lot of problems. First, the microphones do not capture what a passenger would really hear as the measurement is done without a passenger being in the car. As every object present in a room has an influence on the sound transmission, the effect of the passengers cannot be modeled. This especially relates to the head shadowing effect, where a sound is blocked by the head itself and thus is weaker on one side. This can only be approximated by other solutions in a way, that for sound transmission and compensation only those speakers are used, which are in the direction where the sound comes from. In this way again, some shadowing can be achieved. But in a car, the loudspeaker positions are often placed at bad locations which limits this effect significantly. This may be addressed by the audio processing apparatus 100 according to embodiments disclosed herein, because the real impulse arriving at the ears may be used for the optimization.
[0094] For conventional devices the optimization target is usually to achieve for a somehow weighted average of all measurement positions a given target response. By forcing this target, the optimization can only achieve suboptimal results. The audio processing apparatus 100 according to embodiments disclosed herein overcomes this limitation by splitting up the optimization task into a group of tasks with specific properties and defining for each group a specific optimization target that does not contain requirements which would not lead to a sound improvement. Each group may be defined by one frequency band, a target speaker driver set and a second upmixer signal output.
[0095] Conventional devices only use one upmixer and then use the outputs of this upmixer for the speaker outputs as well as to calculate the compensation signals. In this way, during the optimization, it is not possible to define optimization criteria based on input signal properties independently from the speaker signals. This again causes the requirements during the optimization to be more restricted and this leads to suboptimal results. The audio processing apparatus 100 according to embodiments disclosed herein overcomes this limitation by using a first upmixer stage 110 and a second upmixer stage 120, one optimized for the speaker setup and one optimized for signal properties.Conventional devices do not allow optimizing the sound field and at the same time achieving crosstalk cancellation. This thus limits the width of the front stage perception, i.e. sounds cannot be rendered clearly to the sides. The root cause for this is again that the attempt is made to optimize towards arrays of microphones, which cannot focus on the real ear positions. The audio processing apparatus 100 according to embodiments disclosed herein overcomes this limitation by optimizing for the real ear positions and, thus, allowing to apply an optimization target to achieve crosstalk cancellation.
[0096] Figure 6 is a flow diagram illustrating an audio processing method 600 for generating a plurality of output signals for driving a plurality of loudspeakers based on one or more input signals. The audio processing method 600 comprises a step 601 of upmixing the one or more input signals into a first plurality of upmixed input signals and a step 603 of upmixing the one or more input signals into a second plurality of upmixed input signals. Moreover, the audio processing method 600 comprises a step 605 of filtering the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with a precompensation filter. For each of the plurality of loudspeakers the audio processing method 600 comprises a further step 607 of adding the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective output signal. As will be appreciated, although in figure 6 the step 603 is performed after the step 601, in further embodiments of the audio processing method 600 the step 603 may be performed before or at least partially in parallel with the step 601.
[0097] The method 600 can be performed by the audio processing apparatus 100 according to an embodiment. Thus, further features of the method 600 result directly from the functionality of the audio processing apparatus 100 as well as its different embodiments described above and below.
[0098] The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the present disclosure (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step).
[0099] In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described embodiment of an apparatus is merely exemplary. For example, the unit division is merely logical function division and may be another division in an actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0100] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
[0101] In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.
Claims
CLAIMS1. An audio processing apparatus (100) for generating a plurality of output signals for driving a plurality of loudspeakers (220a-c) based on one or more input signals, wherein the audio processing apparatus (100) is configured to:upmix the one or more input signals into a first plurality of upmixed input signals;upmix the one or more input signals into a second plurality of upmixed input signals;filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with one or more pre-compensation filters; andfor each of the plurality of loudspeakers (220a-c), add the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective output signal.
2. The audio processing apparatus (100) of claim 1, wherein the audio processing apparatus (100) is configured to upmix the one or more input signals into the first plurality of upmixed input signals based on a configuration of the plurality of loudspeakers (220a-c) and / or wherein the audio processing apparatus (100)is configured to upmix the one or more input signals into the second plurality of upmixed input signals based on one or more characteristics of the one or more input signals.
3. The audio processing apparatus (100) of claim 1 or 2, wherein the audio processing apparatus (100) is configured to filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more precompensation filters by applying the one or more frequency range optimized pre-compensation filters to the first plurality of upmixed input signals and / or the second plurality of upmixed input signals in a plurality of different frequency ranges.
4. The audio processing apparatus (100) of claim 3, wherein the plurality of different frequency ranges are selected to match the frequency ranges of the plurality of loudspeakers (220a-c).
5. The audio processing apparatus (100) according to claim 3 or 4, wherein the audio processing apparatus (100) is further configured to bandpass filter further copies of the first plurality of upmixed input signals and / or the second plurality of upmixed input signals.
6. The audio processing apparatus (100) of claim 5, wherein the audio processing apparatus (100) is configured to bandpass filter further copies of the first plurality of upmixed input signals and / or the second plurality of upmixed input signals before filtering the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more pre-compensation filters.
7. The audio processing apparatus (100) of any one of the preceding claims, wherein the one or more pre-compensation filters are based on one or more impulse responses.
8. The audio processing apparatus (100) of claim 7, wherein the one or more impulse responses are measured taking the head shadowing effect into account.
9. The audio processing apparatus (100) of any one of the preceding claims, wherein the plurality of loudspeakers (220a-c) comprises one or more front loudspeakers and wherein for generating the one or more output signals for the one or more front loudspeakers the audio processing apparatus (100) is configured to filter the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with the one or more pre-compensation filters such that a symmetric front stage is achieved.
10. The audio processing apparatus (100) of any one of the preceding claims, wherein the one or more pre-compensation filters are based on an optimization of a selected target of the individual compensation filter.
11. The audio processing apparatus (100) of any one of the preceding claims, wherein the one or more pre-compensation filters are configured such that phase differences are distinguishable up to about 800 Hz and / or such that time structures at higher frequencies than about 2 kHz are not distinguishable.
12. The audio processing apparatus (100) of any one of the preceding claims, wherein the audio processing apparatus (100) is configured to upmix the one or more input signals into the first plurality of upmixed input signals, including a center, a left, a right, a far left, a far right, an ambient left, an ambient right, a surround left and / or a surround right signal.
13. The audio processing apparatus (100) according to any one of the preceding claims, wherein the one or more precompensation filters are configured to perform an optimization for a respective frequency range of the respective loudspeaker (210a-c).
14. The audio processing apparatus (100) according to any one of the preceding claims, wherein the one or more precompensation filters are optimized with respect to a respective position of each loudspeaker (220a-c).
15. The audio processing apparatus (100) according to any one of the preceding claims, wherein the configuration of the plurality of loudspeakers (220a-c) includes a subwoofer, a woofer, a midrange and / or a tweeter loudspeaker configuration.
16. The audio processing apparatus (100) according to any one of the preceding claims, wherein the configuration of the plurality of loudspeakers (220a-c) comprises a front center, front left, front right, back left, back right, surround left and / or surround right loudspeaker configuration.
17. The audio processing apparatus (100) according to any one of the preceding claims, wherein for upmixing the one or more input signals into the first plurality of upmixed input signals the audio processing apparatus (100) is configured to generate a center signal by summing up a left signal and a right signal.
18. The audio processing apparatus (100) according to any one of the preceding claims, wherein for upmixing the one or more input signals into the second plurality of upmixed input signals the audio processing apparatus (100) is configured to upmix the one or more input signals into a far left, a left, a center, a right, a far right, a surround left, a surround right, an ambience left, an ambience right and / or a subwoofer signal.
19. The audio processing apparatus (100) of any one of the preceding claims, wherein the one or more bandpass filters comprise a bandpass filter in a subwoofer frequency range, a bandpass filter in a woofer frequency range, and / or a bandpass filter in a tweeter frequency range.
20. The audio processing apparatus (100) of claim 19, wherein the audio processing apparatus (100) comprises a crossover filter configured to split the output of the bandpass filter in the woofer frequency range into one or more further frequency ranges.
21. The audio processing apparatus (100) of any one of the preceding claims, wherein for upmixing the one or more input signals into the first plurality of upmixed input signals the audio processing apparatus (100) comprises a delay unit (105; 305, 405; 505) configured to delay one or more of the upmixed input signals for alignment with the output of pre-compensation filters of the second upmixer output.
22. An audio rendering system, comprising:a plurality of loudspeakers (220a-c); andan audio processing apparatus (100) according to any one of the preceding claims for generating the plurality of output signals for driving the plurality of loudspeakers (210a-c) based on the one or more input signals.
23. An audio processing method (600) for generating a plurality of output signals for driving a plurality of loudspeakers (220a-c) based on one or more input signals, wherein the audio processing method (600) comprises:upmixing (601) the one or more input signals into a first plurality of upmixed input signals;upmixing (603) the one or more input signals into a second plurality of upmixed input signals;filtering (605) the first plurality of upmixed input signals and / or the second plurality of upmixed input signals with a pre-compensation filter; andfor each of the plurality of loudspeakers adding (607) the respective filtered or unfiltered upmixed input signal of the second plurality of upmixed input signals to the respective filtered or unfiltered upmixed input signal of the first plurality of upmixed input signals for generating the respective output signal.
24. A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (600) of claim 23 when the program code is executed by the computer or the processor.17