Audio processing apparatus and method for enhancing an audio signal in a vehicle

EP4744153A1Pending Publication Date: 2026-05-20YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
YINWANG INTELLIGENT TECHNOLOGIES CO LTD
Filing Date
2024-02-26
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing audio processing systems in vehicles suffer from perceptible amplitude modulation, insufficient gain compensation, audible frequency response alteration, undesired gain compensation at low playback volumes, and dynamical system instabilities, making them imperfect in enhancing audio signals in noisy environments.

Method used

An audio processing apparatus and method that includes a loudness adjustment stage, an equalizer stage, and a control unit to estimate ambient noise and desired acoustic signal levels, using dynamic processing units and machine learning for content type detection, to control loudness and frequency-dependent gain adjustments, avoiding feedback loops and using a cascade of compressors with different time constants.

Benefits of technology

The solution provides improved noise compensation, prevents audio modulation, ensures stability, and adapts to signal content, allowing high compensation gains while maintaining audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024054758_04092025_PF_FP_ABST
    Figure EP2024054758_04092025_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing apparatus (100) is disclosed for enhancing an audio signal inside of a vehicle having ambient noise. The audio processing apparatus (100) comprises a loudness adjustment stage (110, 120) configured to adjust a loudness of the audio signal and an equalizer stage (130) configured to equalize the audio signal processed by the loudness adjustment stage (110, 120) or a further processed audio signal based on the audio signal processed by the loudness adjustment stage (110, 120). Moreover, the audio processing apparatus (100) comprises one or more transducers configured to generate one or more acoustic signals based on the audio signal and a volume setting. The audio processing apparatus further comprises a control unit (140, 150, 160, 170) configured to (a) estimate an ambient noise level based on a speed of the vehicle and (b) estimate an level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal, wherein the control unit (140, 150, 160, 170) is further configured to control the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUDIO PROCESSING APPARATUS AND METHOD FOR ENHANCING AN AUDIO SIGNAL IN A VEHICLE

[0002] TECHNICAL FIELD

[0003] The present disclosure relates to audio processing in general. The disclosure relates to an audio processing apparatus and method for enhancing an audio signal in a listening environment inside of a vehicle having ambient noise (also known as Vehicle Noise Compensation, VNC).

[0004] BACKGROUND

[0005] Several approaches are known for adjusting sound to be played back based on environmental data. One approach is to adjust the volume of the sound to be played back in a car by changing the output gain (“volume”) depending on the driving noise. The driving noise may be detected using microphones or can be estimated based on parameters like the velocity of the car, the fan setting or the window opening state, or combinations thereof. This is generally referred to as speed-dependent volume control.

[0006] An extension of this approach does not simply adjust the output gain but adjusts it in different frequency bands, using an equalizer. In this way, the sound playback may be more accurately adjusted to the environmental noise situation, e.g. in a frequency-dependent fashion.

[0007] Another extension also takes some psychoacoustic effects, e.g. the way sound is actually perceived into account. In that case the gain is usually determined by comparing the loudness estimate of the audio signal and the loudness estimate of the environmental noise. This is also possible in different frequency bands.

[0008] Even if the above algorithms, some of which are based on psychoacoustic models, already provide pretty good results, they are still not perfect. Typical issues of these approaches are perceptible amplitude modulation (e.g. an unwanted gain variance over time) of the audio signal, insufficient gain compensation, audible frequency response alteration, undesired gain compensation at low playback volumes, and even (dynamical system) instabilities. Furthermore, it can become complicated and thus timeconsuming to ‘Tune”, e.g. adjust control parameters by applying expert knowledge to these parameters.

[0009] SUMMARY

[0010] It is an objective to provide an improved audio processing apparatus and method for enhancing an audio signal in a listening environment inside of a vehicle having ambient noise.

[0011] The foregoing and other objectives are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.

[0012] According to a first aspect an audio processing apparatus is provided for enhancing an audio signal in a listening environment inside of a vehicle having ambient noise. The audio processing apparatus according to the first aspect comprises a loudness adjustment stage configured to adjust a loudness of the audio signal e.g. for obtaining a processed loudness adjusted audio signal. Moreover, the audio processing apparatus according to the first aspect comprises an equalizer stage configured to equalize (e.g. to perform a frequency-dependent gain adjustment using a plurality of frequency-dependent gains of) the audio signal processed by the loudness adjustment stage or a further processed audio signal based on the audio signal processed by the loudness adjustment stage (e.g. for obtaining a further processed equalized loudness adjusted audio signal). The audio processing apparatus according to the further aspect further comprises one or more transducers (e.g. loudspeakers) configured to generate one or more acoustic signals based on (e.g. the further processed equalized loudness adjusted) audio signal and a volume setting (e.g. a volume knob setting). Moreover, the audio processing apparatus according to the first aspect further comprises a control unit configured to (a) estimate an ambient noise level based on at least a current speed of the vehicle and (b) estimate an level of the desired one or more acoustic signals inside of the vehicle based on the volume setting and not based on an analysis of the original audio signal, wherein the control unit is further configured to control the loudness adjustment stage and the equalizer stage based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.

[0013] The audio processing apparatus according to the first aspect allows for an improved compensation of the ambient noise. Moreover, it prevents audio modulation, is implicitly stable, adapts to the signal content and allows high compensation gains.

[0014] In a further possible implementation form, the loudness adjustment stage comprises a plurality of dynamic processing units for adjusting the loudness of the audio signal, wherein the control unit is configured to control a first dynamic processing unit based on a measured loudness range, LRA, of the audio signal and to control a second dynamic processing unit based on a measured loudness, LUFS, of the audio signal. Using a plurality of dynamic processing units allows to first compress the dynamic range and then the loudness. In this way modulation effects may be kept to a minimum.

[0015] In a further possible implementation form, the first dynamic processing unit comprises a loudness range, LRA, reduction unit configured to adjust the loudness range of the audio signal. In an implementation form, the LRA may be reduced with a slow time constant by the LRA reduction unit. Reducing the loudness range especially with a slow time constant allows compensating slow changes of the loudness of an audio signal. Furthermore, this compensates level differences of different input sources (FM radio usually has a low level due to some headroom requirement, digital transmission instead uses the full headroom) and even audio tracks.

[0016] In a further possible implementation form, the second dynamic processing unit comprises a loudness units full scale, LUFS, compression unit configured to adjust the loudness of the audio signal. In an implementation form, the LUFS compression unit may be configured to operate on a faster time constant than the LRA reduction unit. This allows boosting the level of the audio input signal whose loudness range has already been compressed.

[0017] In a further possible implementation form, the control unit is configured to determine an audio content type based on the input audio signal and to control the loudness adjustment stage and the equalizer stage based on the audio content type. Different audio contents cause different loudness perception and intelligibility. For instance, a speech signal usually requires around 3dB more level boost to keep the same loudness impression and intelligibility compared to a music signal. By knowing the content type this psychoacoustical effect can be taken into account.

[0018] In a further possible implementation form, the control unit is configured to implement a machine learning, ML, model, wherein the ML model is configured to determine the content type of the audio signal. Machine learning models for content type estimation usually achieve better detection quality than the classical signal processing-based content type detection methods. Thus, by using such an Al model, the content type detection is more accurate which will lead to better compensation results.

[0019] In a further possible implementation form, the control unit is configured to control the loudness adjustment stage and the equalizer stage based on the estimate of the ambient noise level and the estimate of the level of the desired one or more acoustic signals by determining one or more frequency-dependent gains, e.g. gain factors. Using the estimates instead of real measured values has the advantage to avoid modulation of the audio signal. In a further possible implementation form, the control unit is configured to control the loudness adjustment stage and the equalizer stage based on the estimate of the ambient noise level and the estimate of the level of the desired one or more acoustic signals by using one or more matrices for mapping the estimate of the ambient noise level and / or the estimate of the level of the desired one or more acoustic signals to one or more frequency-dependent gains for the equalizer stage and / or one or more loudness adjustment parameters, in particular one or more threshold values and / or ratios, wherein the control unit is configured to control the loudness adjustment stage based on the one or more frequency-dependent gains and / or the one or more loudness adjustment parameters, in particular the one or more threshold values and / or ratios. Using matrices allows tuning the behavior very accurately and easily. This will thus lead to very accurate tuning results. Adjusting the compressor parameters like threshold and ratio directly instead of mixing the result of the compressed with the uncompressed signal not only reduces the processing load but also reduces modulation artifacts.

[0020] In a further possible implementation form, each of the one or more matrices comprises a plurality of matrix values, wherein the control unit is configured to interpolate the plurality of matrix values for mapping the estimate of the ambient noise level and / or the estimate of the level of the desired one or more acoustic signals to the one or more frequency-dependent gains for the equalizer stage and / or the one or more loudness adjustment parameters. The interpolation allows for smooth transition between the matrix elements, which otherwise could cause audible clicks.

[0021] In a further possible implementation form, the audio processing apparatus according to the first aspect further comprises a microphone configured to detect, e.g. sense ambient noise, wherein the control unit is configured to estimate the ambient noise level based on the current speed of the vehicle and the ambient noise detected by the microphone. Using the microphone also ambient noise sources like the noise close to a construction site or highway can be detected and compensated for.

[0022] In a further possible implementation form, the control unit is configured to obtain a current fan speed of a fan of the vehicle and to estimate the ambient noise level based on the current speed of the vehicle and the current fan speed. As the fan of the car creates some noise, especially at higher fan settings, it is desired to compensate for this noise. The noise level is directly related to the speed of the fan, this way it can be determined using this parameter.

[0023] According to a second aspect an audio processing method is provided for enhancing an input audio signal in a listening environment inside of a vehicle having ambient noise. The audio processing method comprises the following steps: adjusting a loudness of the audio signal by a loudness adjustment stage; equalizing by an equalizer stage the audio signal processed by the loudness adjustment stage or a further processed audio signal based on the audio signal processed by the loudness adjustment stage; and generating one or more acoustic signals based on the audio signal and a volume setting.

[0024] The audio processing method further comprises: estimating an ambient noise level based on a speed of the vehicle; estimating a level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal; and controlling the loudness adjustment stage and the equalizer stage based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.

[0025] The audio processing method according to the second aspect can be performed by the audio processing apparatus according to the first aspect. Thus, further features of the audio processing method according to the second aspect result directly from the functionality of the audio processing apparatus according to the first aspect as well as its different implementation forms and embodiments described above and below. According to a third aspect a computer program product is provided, comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method according to the second aspect, when the program code is executed by the computer or the processor.

[0026] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.

[0027] BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In the following, embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which:

[0029] Fig. 1 is a schematic diagram illustrating an audio processing apparatus according to an embodiment for enhancing an input audio signal in a listening environment inside of a vehicle having ambient noise;

[0030] Fig. 2 is a schematic diagram illustrating a more detailed embodiment of the audio processing apparatus according to an embodiment for enhancing an input audio signal in a listening environment inside of a vehicle having ambient noise;

[0031] Fig. 3 is a schematic diagram illustrating a variant of the embodiment of the audio processing apparatus of figure 2 making use of a matrix instead of various tables;

[0032] Fig. 4 is a schematic diagram illustrating a variant of the embodiment of the audio processing apparatus of figure 2 making use of two matrices instead of various tables for independently controlling loudness compression and equalization effect intensities; and

[0033] Fig. 5 is a flow diagram illustrating steps of an audio processing method according to an embodiment for enhancing an input audio signal in a listening environment inside of a vehicle having ambient noise.

[0034] In the following, identical reference signs refer to identical or at least functionally equivalent features.

[0035] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims.

[0037] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.

[0038] Figure 1 is a schematic diagram illustrating an audio processing apparatus 100 according to an embodiment for enhancing an audio input signal in a listening environment inside of a vehicle having, e.g. experiencing ambient noise.

[0039] The audio processing apparatus 100 comprises a loudness adjustment stage 110, 120 configured to adjust a loudness of the audio input signal, which in the embodiment shown in figure 1 comprises a plurality of dynamic processing units 110, 120 for adjusting the loudness of the audio input signal, wherein a first dynamic processing unit 120 comprises a loudness range, LRA, reduction unit 120 configured to adjust the loudness range of the audio input signal and wherein a second dynamic processing unit 110 comprises a loudness units full scale, LUFS, compression unit 110 configured to adjust the loudness of the audio input signal. In an embodiment, the LRA reduction unit 120 may be configured to operate based on a slow time constant, e.g. with long time samples of the audio signal, while the LUFS compression unit 111 may be configured to operate on a faster time constant, e.g. with shorter time samples.

[0040] As illustrated in figure 1, the audio processing apparatus 100 further comprises an equalizer stage 130 configured to equalize, e.g. to perform a frequency-dependent gain adjustment using a plurality of frequency-dependent gains of the audio input signal processed by the loudness adjustment stage 110, 120 for obtaining a further processed, e.g. equalized loudness adjusted audio signal. As will be appreciated, the audio processing apparatus 100 may comprise further audio processing blocks not illustrated in figure 1. In particular, one or more further processing blocks may be implemented between the output of the LRA reduction unit 120 and the input of the equalizer stage 130. As will be further appreciated, for such embodiments the equalizer stage 130 takes as input a further processed audio signal based on the audio input signal processed by the loudness adjustment stage 110, 120.

[0041] The audio processing apparatus 100 further comprises one or more transduces, e.g. loudspeakers (not shown in figure 1) configured to generate one or more acoustic signals, e.g. the audio output signal(s) based on the audio signal processed by the loudness adjustment stage 110, 120 and the equalizer stage 130 and based on a volume setting, for instance the setting of a volume knob of the audio processing apparatus 100. The one or more transducers, e.g. loudspeakers may be arranged at suitable positions inside of the vehicle.

[0042] As illustrated in figure 1, the audio processing apparatus 100 may comprise a plurality of control components 140, 150, 160, 170, which will be described in more detail in the following and constitute a control unit 140, 150, 160, 170, which is generally configured to (a) estimate an ambient noise level based on a speed of the vehicle and (b) estimate a level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal, as will be described in more detail in the following. The control unit 140, 150, 160, 170, in particular the control component 140 thereof is further configured to control the loudness adjustment stage 110, 120 and the equalizer stage 130 based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals. In an embodiment, the control component 140 is configured to control the loudness adjustment stage 110, 120 and the equalizer stage 130 based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals by determining one or more frequency-dependent gains.

[0043] In a further embodiment, the control component 140 of the control unit 140, 150, 160, 170 is configured to control the first dynamic processing unit 120, e.g. the LRA reduction unit 120 based on a measured LRA of the audio signal and to control the second dynamic processing unit 110, e.g. the LUFS compression unit 110 based on a measured LUFS of the audio signal. The audio processing apparatus 100 may further comprise a microphone configured to detect ambient noise, wherein the control unit 140, 150, 160, 170, in particular the control component 160 thereof is configured to estimate the ambient noise level based on the speed of the vehicle and the ambient noise detected by the microphone. In an embodiment, the control unit 140, 150, 160, 170, in particular the control component 160 thereof is further configured to obtain a fan speed of a fan of the vehicle and to estimate the ambient noise level based on the speed of the vehicle, the fan speed and / or the ambient noise detected by the microphone. In an embodiment, as will be described in more detail in the following, the 1.41 I S compression 110 and the LRA reduction unit 120 may be adjusted by the control unit based on the scaled difference of the ambient noise level and the expected audio level, wherein this value may be furthermore scaled based on the volume setting and corrected based on the content type.

[0044] In an embodiment, the control unit 140, 150, 160, 170, in particular the control component 150 thereof is configured to determine an audio content type, such as speech, pop music, classical music, and the like, based on the audio input signal and to control the loudness adjustment stage 110, 120 and the equalizer stage 130 based on the audio content type. To this end, in an embodiment the control component 150 may implement a machine learning, ML, model, wherein the ML model is configured to determine the audio content type of the audio input signal.

[0045] As will be appreciated, according to an embodiment the audio processing apparatus 100 is configured to process the audio input signal with a cascade of three different audio adjusting stages, e.g. modules 110, 120, 130. As already described above, the first stage 110 compresses the loudness, the second stage 120 reduces the loudness range and the third stage 130 performs frequency-dependent gain adjustments. All these audio processing functions are adjusted by the control unit, in particular the control component 140. The control component 140 determines the control data based on information about the content type, the ambient noise, and the estimated music level. This data is obtained via the other control components 150, 160, 170 for evaluating the audio signal, as well as car parameters, such as which audio source is selected for playback (e.g. music, telephone, navigation etc.), the car speed, the fan level, microphone data, the volume setting and the like.

[0046] The audio processing apparatus 100 according to an embodiment is configured to generate different compression and level adjustments for different audio content types. As an example, speech signals (such as a news broadcast, or a telephone call) may require about 3dB more boost, in order to keep the same impression (compared with a music signal). This is achieved by the control component 150 for detecting the content type using e.g. machine-leaming-based content type detection and evaluating the audio source type (telephone, music, navigation).

[0047] In conventional VNC systems usually only the loudness compression is used but the loudness range is not additionally modified. By using a cascade consisting of a loudness compressor and a loudness range reduction module, the audio processing apparatus 100 according to an embodiment allows implementing more extensive loudness control. This approach reduces dynamic-range-compression-related artifacts for music with a high LRA in scenarios in which a simple LUFS-based compressor would cause compression artifacts due to the relatively smaller time constants.

[0048] Many conventional VNC algorithms use a feedback loop where the gain to be applied is added to the gain that would occur without VNC, and then microphones are used for determining the real level of the audio signal that is being played. This feedback loop causes a variety of issues ranging from level overshoots to instability. To avoid such unwanted side effects, those algorithms limit the reaction time and possible amplification gains. Embodiments of the audio processing apparatus 100 disclosed herein do not require this kind of feedback, and thus overcome these issues.

[0049] Psychoacoustic-masking-based algorithms may achieve better intelligibility, but due to continuous adaptation of the audio signal, may also cause perceptible modulation of the signal. This is acceptable for speech signals, but unwanted for high-end music playback. Embodiments of the audio processing apparatus 100 disclosed herein use only an estimate of the music level based on the volume playback setting, as an input into the control operation, which also models psychoacoustic effects and thereby overcome this problem, as no modulation is caused any more.

[0050] Figure 2 shows a more detailed embodiment of the more general embodiment of the audio processing apparatus 100 of figure 1. In the embodiment of figure 2, the audio processing apparatus 100 comprises a respective processing block 231, 233, 235 configured to convert the speed of the car, the fan speed of the car and the ambient noise detected in the microphone signal to specific-source-related noise level estimates. These levels are then combined into an overall noise level by the noise summation block 237. The volume setting of the sound apparatus is converted into an expected audio playback level by a processing block 241 configured to map volume settings into dBs, as indicated by the table shown near the processing block 241. The estimated noise level is then subtracted from the estimated audio level with the result being considered as an audio-to-noise-ratio (ANR). Based on psychoacoustic effects, different ANR values are required for different types of audio content, to deliver results comparable in terms of sound impression. As one example, for a speech signal, a 3dB higher ANR is desired than for a pop music signal. To take this effect into account, the content type of the audio signal is determined by the processing block 221, and this is mapped to a compensation factor by the processing block 223, which is then added to the ANR with the result referred to herein as “ANR modified” (ANRM) value.

[0051] The ANRM value is then mapped to a desired boost by the processing block 243. For each frequency band, the desired boost is then scaled again by the processing block 247. To avoid that at very low volume values or very high- volume values undesired boosting occurs, the desired boost is scaled again depending on the volume by the processing block 251 , to reflect this property. This scaled boost is then applied to the audio output signal.

[0052] On top of this processing also the EBU R128 LRA (or similar) and LUFS (or similar) values are determined by the processing blocks 201 and 211. An audio signal where the signal level varies significantly over time will have a large LRA. LUFS instead just indicates the loudness of the signal, so this will react much faster.

[0053] If the environmental noise can already affect the audio perception, it is desired to reduce the loudness range (LRA) but to increase the loudness (LUFS). For this purpose, the LRA values are mapped to compressor threshold and ratio values by the processing block 213 and a slow audio compressor 215 is used to compress the LRA. The LUFS value is then also mapped to compressor threshold and ratio parameters by the processing block 203, which are then applied using a fast audio compressor to increase LUFS 205. As already mentioned, such compression may only be applied, if the noise level is relatively high, to the point where it already degrades the audio perception quality. An indicator for this is the ANRM value. For high ANRM values no compression is required, as the audio signal can be estimated to be clearly audible. Thus, to only activate this compression if the ANRM value is low, the ANRM value is mapped to an intensity factor by the processing block 245 defining how much compression shall be used. In this embodiment, the intensity factor is simply used to crossfade among the uncompressed and compressed signals (a more sophisticated embodiment could directly adjust the compression parameters based on the ANRM value). Using the default boost compared to the boost where the volume mapping is already applied has the advantage that an individual volume mapping curve can be applied, where, for example, compression is also allowed at high volume values (this will increase the maximum sound pressure level for highly dynamic signals). The volume values are furthermore mapped to a respective intensity in processing block 255. The output value of the processing block 255 is multiplied with the output of processing block 245 and then used in the gain modules to adjust the intensity of the loudness. In processing block 253 the desired gain in dB is converted to a linear value which is then multiplied with the audio signal.

[0054] Using a compressor on the input audio signal, which is a common approach, has the disadvantage that it causes permanent audio modulation which is audible, and which deteriorates the audio quality. The embodiment of the audio processing apparatus 100 shown in figure 2 overcomes this problem in two ways. Firstly, not a single compressor is used but a cascade of 2 compressors which differ in reaction time and the input signal domain. The slower one is used to reduce the loudness range, the faster one - to directly increase the loudness by compression. In addition to independently controlling the loudness range and the loudness itself, the dual cascade compression furthermore may be adjusted in intensity from the control operation block. The output of the cascade is mixed with the original, uncompressed signal. The mixing gain is determined from the control block, so at sufficiently high sound pressure the control block will adjust the mixing gains so that the original, unmodified signal is used. This will prevent perceivable audio modulation. In the embodiment shown in figure 2, the content type may be determined using a machine-leaming-based algorithm. Based on the detected content type, then, an offset to the audio-to-noise ratio is added, the offset is retrieved from a table which maps the different content types to the relevant gains. This has the advantage that for each content type an optimal compensation can be achieved. Using the content detection on the audio signal allows to even detect speech signals that are transmitted in a stream where usually music is expected, e.g. in a radio transmission when the news is presented. Furthermore, different music styles can also be handled differently, when mapped to content types. A classical recording e.g. can be compressed more than pop-music which is usually already heavily compressed.

[0055] As will be appreciated from the embodiment shown in figure 2, the main control operation is feedback-free (other than the - optional - mapping of the microphone signal, which may use preprocessing to remove the music signal). Other implementations use a feedback of the calculated gain values. This is usually done as adding gain improves the audio-to-noise ratio, so this value has to be corrected to match the psychoacoustic model used by other implementations. This is overcome here by having access to the control tables, this way e.g. the steepness can be reduced, which has a similar effect as feeding back the gain signal, but which avoids the drawbacks of feedback (potential overshoots, level oscillation), which require reducing the reaction time and gain in order to overcome these problems. So, the disclosed embodiment does not suffer from any control system instability, level overshoot or even oscillation issues.

[0056] Figure 3 shows a further embodiment of the audio processing apparatus 100. As will be appreciated, the embodiment shown in figure 3 comprises a lot of the processing blocks already described in the context of the embodiment shown in figure 2. The main difference of the embodiment shown in figure 3 with the embodiment shown in figure 2 is that instead of using a variety of functions to map the noise and audio level to gains, in the embodiment shown in figure 3 a matrix 301 is used. To avoid discontinuities, values in the matrix may be interpolated in the x (row) and y (column) directions, in order to achieve a smooth gain change. Furthermore, the output gain from the matrix may be scaled by a user-adjustable intensity level. This allows the user to deactivate the operation, reduce the effect intensity or even increase the intensity above a normal level. This intensity adjustment approach can also be used in the embodiment shown in figure 2.

[0057] As illustrated in figure 3, the gain adjustment may be done in frequency sub-bands. For this purpose, the input audio signal may be split up either between the first and second compressor, or after the second compressor into sub-band signals. In a straightforward implementation, then, for each sub-band one matrix 301 may be used to map the noise and audio levels to the gain for the relevant sub-band. To reduce the number of matrices, it is possible to use additional interpolation in the frequency domain (between matrices for sub-bands).

[0058] Using a matrix 301 to map the noise and audio levels to desired gains significantly simplifies the tuning process of the embodiment illustrated in figure 3. The clear assignment of a parameter pair consisting of a noise level and an audio level to a target gain allows for very accurate adjustment. This can be done by e.g. driving a car on a test track at such a speed that the noise level for which the matrix parameter shall be adjusted is reached, then setting the volume to the audio level that shall be adjusted, and then simply updating the corresponding matrix value, so that the desired compensation for this parameter combination is achieved. This will ensure very consistent results. As the full permutation of audio and noise parameters is handled, it is also ensured that the desired behavior is achieved for all predicted situations. In the approach according to the embodiment illustrated in figure 2 there are more interdependencies, so changing one table while driving at a given speed and listening at a given audio level might impact also the behavior of the system in another driving / listening situation. The extension towards a user-selectable effect intensity has the advantage that the end users may want to adjust the intensity of the effect towards their personal preference. Some end users prefer a stronger compensation, while some - a weaker compensation. As already mentioned, this intensity adjustment feature may also be used in the embodiment shown in figure 2.

[0059] Figure 4 shows a further embodiment of the audio processing apparatus 100. As will be appreciated, the embodiment shown in figure 4 comprises a lot of the processing blocks already described in the context of the embodiments shown in figure 2 and figure 3. The main difference of the embodiment shown in figure 4 with the embodiment shown in figure 3 is that instead of deriving the intensity for the compression path by mapping the boost from the first table to a linear intensity, in the embodiment of figure 4 an individual matrix, e.g. table 401 is used to determine this intensity. The table inputs are again the estimated noise level and the estimated audio level, the output is the compression intensity. As with the other matrix, e.g. table 301 of the embodiment of figure 3, the matrix, e.g. table 401 may also be specified for every subband (this then also requires to do the compression in subbands). The output gain from both matrices 301, 401 are scaled by a user-adjustable intensity level. This allows the user to deactivate the operation, reduce the effect intensity or even increase the intensity above a normal level.

[0060] As indicated in figure 4, both the dynamic range modifications and the gain adjustment can be done in frequency sub-bands. For this purpose, the input audio signal must be split up either between the first and second compressor, or after the second compressor into sub-band signals. In a straightforward implementation, then, for each sub-band one pair of matrices 301, 401 can be used to map the noise and audio levels to the loudness modification intensity and the gain for the relevant sub-band. To reduce the number of matrices, it is possible to use additional interpolation in the frequency domain (between matrices for subbands).

[0061] The main advantage of the embodiment shown in figure 4 results from the possibility to independently adjust the dynamic range and gain for the different frequency bands. This is also visible in the matrices 301, 401 indicated in figure d. For example, at an audio level of 20dB(A) no gain is applied but still the dynamic range modification is performed. The motivation behind this is that at low volume values no extra boosting of normal signals is desired. But if the audio signal in this case has a high dynamic range or high loudness range, then compression is still desired.

[0062] Instead of using a real LUFS measurement according to the EBU R128 standard, also other methods to obtain an estimate of the perceived loudness may be used for the embodiments illustrated in figures 1, 2, 3 and 4. Typical implementations may be based on dB(A), dB(C) measurement standards as well as loudness models like the so called “Zwicker loudness”. For the loudness range also simpler approximations are possible, the LRA measurement according to the EBU R128 standard described for the embodiments above is just one possible implementation. An alternative embodiment may simply observe the variation of the dB(A) or dB(C) values over time. For content type detection instead of machine-leaming-based algorithms also classical signal processing approaches may be used, or even just the information about the connected source or any other information about the track played may be used. Instead of mixing the output of the compressor cascade with the original signal, it is also possible to adjust the compression parameters so that they become transparent (the static characteristic of the compressor is linear) when the audio signal level is sufficiently high.

[0063] Figure 5 is a flow diagram illustrating an audio processing method 500 for enhancing an audio signal inside of a vehicle having ambient noise. The audio processing method 500 comprises a step 501 of adjusting a loudness of the audio signal by a loudness adjustment stage 110, 120. Moreover, the audio processing method 500 comprises a step 503 of equalizing by an equalizer stage 130 the audio signal processed by the loudness adjustment stage 110, 120 or a further processed audio signal based on the audio signal processed by the loudness adjustment stage 110, 120. The method 500 further comprises a step 505 of generating one or more acoustic signals based on the audio signal and a volume setting. The audio processing method 500 further comprises a step 507 of estimating an ambient noise level based on a speed of the vehicle and a step 509 of estimating a level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal. The audio processing method 500 further comprises a step 509 of and controlling the loudness adjustment stage and the equalizer stage based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.

[0064] The audio processing method 500 can be performed by the audio processing apparatus 100 according to an embodiment. Thus, further features of the audio processing method 500 result directly from the functionality of the audio processing apparatus 100 as well as its different embodiments described above and below. As will be further appreciated, the steps of the audio processing method 500 may be implemented in a different order than the one illustrated in figure 5.

[0065] The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the present disclosure (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step).

[0066] In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described embodiment of an apparatus is merely exemplary. For example, the unit division is merely logical function division and may be another division in an actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0067] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.

[0068] In addition, functional units in the embodiments of the invention may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

Claims

CLAIMS1. An audio processing apparatus (100) for enhancing an audio signal inside of a vehicle having ambient noise, wherein the audio processing apparatus (100) comprises: a loudness adjustment stage (110, 120) configured to adjust a loudness of the audio signal; an equalizer stage (130) configured to equalize the audio signal processed by the loudness adjustment stage (110, 120) or a further processed audio signal based on the audio signal processed by the loudness adjustment stage (110, 120); one or more transducers configured to generate one or more acoustic signals based on the audio signal and a volume setting; and a control unit (140, 150, 160, 170) configured to (a) estimate an ambient noise level based on a speed of the vehicle and (b) estimate an level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal, wherein the control unit (140, 150, 160, 170) is further configured to control the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.

2. The audio processing apparatus (100) of claim 1, wherein the loudness adjustment stage (110, 120) comprises a plurality of dynamic processing units (110, 120) for adjusting the loudness of the audio signal, wherein the control unit (140) is configured to control a first dynamic processing unit (120) based on a measured loudness range, LRA, of the audio signal and to control a second dynamic processing unit (110) based on a measured loudness, LUFS, of the audio signal.

3. The audio processing apparatus (100) of claim 2, wherein the first dynamic processing unit (120) comprises a loudness range, LRA, reduction unit (120) configured to adjust the loudness range of the audio signal.

4. The audio processing apparatus (100) of claim 2 or 3, wherein the second dynamic processing unit (110) comprises a loudness units full scale, LUFS, compression unit (110) configured to adjust the loudness of the audio signal.

5. The audio processing apparatus (100) of any one of the preceding claims, wherein the control unit (150) is configured to determine an audio content type based on the audio signal and to control the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the audio content type.

6. The audio processing apparatus (100) according to claim 5, wherein the control unit (150) is configured to implement a machine learning, ML, model, wherein the ML model is configured to determine the audio content type of the audio signal.

7. The audio processing apparatus (100) of any one of the preceding claims, wherein the control unit (150) is configured to control the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals by determining one or more frequency-dependent gains.

8. The audio processing apparatus (100) of any one of claims 1 to 7, wherein the control unit (150) is configured to control the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals by using one or more matrices for mapping the estimate of the ambient noise level and / or the estimate of the level of the one or more acoustic signals to one or more frequency-dependent gains and / or one or more loudness adjustment parameters, wherein the control unit (150) is configured to control the equalizer stage (130) and / or the loudness adjustment stage (110, 120) based on the one or more frequency-dependent gains and / or the one or more loudness adjustment parameters.

9. The audio processing apparatus (100) of claim 8, wherein each of the one or more matrices comprises a plurality of matrix values and wherein the control unit (150) is configured to interpolate the plurality of matrix values for mapping the estimate of the ambient noise level and / or the estimate of the level of the one or more acoustic signals to the one or more frequency-dependent gains and / or the one or more loudness adjustment parameters.

10. The audio processing apparatus (100) of any one of the preceding claims, wherein the audio processing apparatus (100) further comprises a microphone configured to detect ambient noise and wherein the control unit (160) is configured to estimate the ambient noise level based on the speed of the vehicle and the ambient noise detected by the microphone.

11. The audio processing apparatus (100) of any one of the preceding claims, wherein the control unit (160) is configured to obtain a fan speed of a fan of the vehicle and to estimate the ambient noise level based on the speed of the vehicle and the fan speed.

12. An audio processing method (500) for enhancing an audio signal inside of a vehicle having ambient noise, wherein the audio processing method (500) comprises: adjusting (501) a loudness of the audio signal by a loudness adjustment stage (110, 120); equalizing (503) by an equalizer stage (130) the audio signal processed by the loudness adjustment stage (110, 120) or a further processed audio signal based on the audio signal processed by the loudness adjustment stage (110, 120); generating (505) one or more acoustic signals based on the audio signal and a volume setting; wherein the audio processing method (500) further comprises: (a) estimating (507) an ambient noise level based on a speed of the vehicle and (b) estimating (509) an level of the one or more acoustic signals based on the volume setting and not based on an analysis of the audio signal, and controlling (511) the loudness adjustment stage (110, 120) and the equalizer stage (130) based on the estimate of the ambient noise level and the estimate of the level of the one or more acoustic signals.

13. A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (500) of claim 12 when the program code is executed by the computer or the processor.