Audio processing apparatus and method for enhancing audio signals within a vehicle

CN122804371APending Publication Date: 2026-09-22YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480088690.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

这些方法的典型问题是音频信号的可感知的幅度调制(例如,随时间出现的不希望的增益变化)、不充分的增益补偿、可听频率响应改变、在低播放音量下不希望的增益补偿,以及甚至(动态系统)不稳定

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804371A_ABST
    Figure CN122804371A_ABST
Patent Text Reader

Abstract

An audio processing apparatus (100) is disclosed for enhancing an audio signal inside a vehicle with ambient noise. The audio processing apparatus (100) includes: a loudness adjustment stage (110, 120) for adjusting the loudness of the audio signal; and an equalizer stage (130) for equalizing the audio signal processed by the loudness adjustment stage (110, 120), or an audio signal further processed based on the audio signal processed by the loudness adjustment stage (110, 120). Furthermore, the audio processing apparatus (100) includes one or more transducers for generating one or more acoustic signals based on the audio signal and a volume setting. The audio processing device further includes control units (140, 150, 160, 170) for (a) estimating the ambient noise level based on the speed of the vehicle, and (b) estimating the level of the one or more acoustic signals based on the volume setting rather than on the analysis of the audio signals, wherein the control units (140, 150, 160, 170) are also used to control the loudness adjustment level (110, 120) and the equalizer level (130) based on the estimated ambient noise level and the estimated level of the one or more acoustic signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to audio processing. Specifically, it relates to an audio processing apparatus and a method for enhancing audio signals in a listening environment inside a vehicle with ambient noise (also known as vehicle noise compensation (VNC)). Background Technology

[0002] Several methods are known for adjusting the volume of a sound played in a car based on environmental data. One method is to adjust the volume of the sound played in the car by changing the output gain (“volume”) according to driving noise. Driving noise can be detected using a microphone or estimated based on parameters such as car speed, fan settings, or whether windows are open, or a combination thereof. This is often referred to as speed-dependent volume control.

[0003] This method is extended beyond simply adjusting the output gain; instead, it uses an equalizer to adjust the output gain across different frequency bands. In this way, for example in a frequency-dependent manner, sound playback can be adjusted more accurately to suit ambient noise conditions.

[0004] Another extension also considers some psychoacoustic effects, such as how sound is actually perceived. In this case, gain is typically determined by comparing the loudness estimate of the audio signal with the loudness estimate of the ambient noise. This may also be achieved in different frequency bands.

[0005] Even though the algorithms mentioned above (some of which are based on psychoacoustic models) have provided fairly good results, they are still not perfect. Typical problems with these methods include perceptible amplitude modulation of the audio signal (e.g., unwanted gain changes over time), insufficient gain compensation, altered audible frequency response, unwanted gain compensation at low playback volumes, and even instability (in dynamic systems). Furthermore, “tuning” can become complex and therefore time-consuming, for example, by applying expert knowledge to adjust control parameters. Summary of the Invention

[0006] The aim is to provide an improved audio processing apparatus and method for enhancing audio signals in a listening environment inside a vehicle with ambient noise.

[0007] The foregoing and other objectives are achieved through the subject matter claimed in the independent claims. Other implementations are apparent from the dependent claims, the specification, and the drawings.

[0008] According to a first aspect, an audio processing apparatus is provided for enhancing an audio signal in a listening environment inside a vehicle with ambient noise. The audio processing apparatus according to the first aspect includes a loudness adjustment stage for adjusting the loudness of the audio signal, for example, for acquiring a processed loudness-adjusted audio signal. Furthermore, the audio processing apparatus according to the first aspect includes an equalizer stage for equalizing the audio signal processed by the loudness adjustment stage, or an audio signal further processed by the audio signal processed by the loudness adjustment stage (e.g., for acquiring a further processed equalized loudness-adjusted audio signal) (e.g., performing frequency-related gain adjustment using multiple frequency-related gains of the audio signal or the further processed audio signal).

[0009] The audio processing apparatus according to another aspect further includes one or more transducers (e.g., speakers) for generating one or more acoustic signals based on (e.g., further processed equalization loudness adjustment) audio signals and volume settings (e.g., volume knob settings). Furthermore, the audio processing apparatus according to the first aspect includes a control unit for (a) estimating an ambient noise level based at least on the current speed of the vehicle, and (b) estimating the levels of one or more desired acoustic signals inside the vehicle based on the volume settings rather than on the analysis of the original audio signal, wherein the control unit is also configured to control the loudness adjustment level and the equalizer level based on the estimated ambient noise level and the estimated levels of the one or more acoustic signals.

[0010] The audio processing apparatus according to the first aspect can improve compensation for ambient noise. Furthermore, it can prevent audio modulation, is implicitly stable, adapts to signal content, and supports high compensation gain.

[0011] In another possible implementation, the loudness adjustment level includes multiple dynamic processing units for adjusting the loudness of the audio signal, wherein the control unit controls a first dynamic processing unit based on a measured loudness range (LRA) of the audio signal and a second dynamic processing unit based on a measured loudness LUFS of the audio signal. Using multiple dynamic processing units allows for compression of the dynamic range first, followed by compression of the loudness. In this way, modulation effects can be minimized.

[0012] In another possible implementation, the first dynamic processing unit includes a loudness range (LRA) reduction unit for adjusting the loudness range of the audio signal. In one implementation, the LRA can be reduced by the LRA reduction unit with a slow time constant. Reducing the loudness range, especially using a slow time constant, can compensate for slow changes in the loudness of the audio signal. Furthermore, this compensates for level differences between different input sources and even audio tracks (FM radios typically have low levels due to margin requirements, while digital transmissions use full margin).

[0013] In another possible implementation, the second dynamic processing unit includes a loudness unit full scale (LUFS) compression unit for adjusting the loudness of the audio signal. In one implementation, the LUFS compression unit can be used to operate on a faster time constant than the LRA reduction unit. This can increase the level of the audio input signal whose loudness range has already been compressed.

[0014] In another possible implementation, the control unit is used to determine the audio content type based on the input audio signal and to control the loudness adjustment level and the equalizer level based on the audio content type. Different audio content results in different perceived loudness and clarity. For example, speech signals typically require approximately 3 dB more level enhancement compared to music signals to maintain the same loudness impression and clarity. This psychoacoustic effect can be considered by understanding the content type.

[0015] In another possible implementation, the control unit is used to implement a machine learning (ML) model, which is used to determine the content type of the audio signal. Machine learning models for content type estimation typically achieve better detection quality than content type detection methods based on classical signal processing. Therefore, by using such an AI model, content type detection is more accurate, resulting in better compensation results.

[0016] In another possible implementation, the control unit is used to control the loudness adjustment level and the equalizer level by determining one or more frequency-dependent gains (e.g., gain factors) based on the estimated values ​​of the ambient noise level and the estimated values ​​of the levels of the desired one or more acoustic signals. Using estimated values ​​instead of actual measurements has the advantage of avoiding audio signal modulation.

[0017] In another possible implementation, the control unit is used to control the loudness adjustment level and the equalizer level based on the estimated ambient noise level and the estimated levels of the desired one or more acoustic signals by using one or more matrices to map the estimated ambient noise level and / or the estimated levels of the desired one or more acoustic signals to one or more frequency-dependent gains and / or one or more loudness adjustment parameters, particularly one or more thresholds and / or ratios. The control unit is used to control the equalizer level and / or the loudness adjustment level based on the one or more frequency-dependent gains and / or the one or more loudness adjustment parameters, particularly the one or more thresholds and / or ratios. Using matrices allows for very accurate and easy tuning of the behavior. Therefore, this produces very accurate tuning results. Directly adjusting compressor parameters (such as thresholds and ratios), rather than mixing the compressed result with the uncompressed signal, not only reduces the processing load but also reduces modulation artifacts.

[0018] In another possible implementation, each of the one or more matrices includes a plurality of matrix values, wherein the control unit is used to interpolate the plurality of matrix values ​​to map the estimated values ​​of the ambient noise level and / or the estimated values ​​of the levels of the desired one or more acoustic signals to the one or more frequency-dependent gain and / or the one or more loudness adjustment parameters of the equalizer level. Interpolation supports smooth transitions between matrix elements that would otherwise result in audible clicking sounds.

[0019] In another possible implementation, the audio processing apparatus according to the first aspect further includes a microphone for detecting (e.g., sensing) ambient noise, wherein the control unit is used to estimate the ambient noise level based on the vehicle's current speed and the ambient noise detected by the microphone. The microphone can also be used to detect and compensate for ambient noise sources, such as noise from near construction sites or highways.

[0020] In another possible implementation, the control unit is used to acquire the current fan speed of the vehicle's fan and estimate the ambient noise level based on the vehicle's current speed and the current fan speed. Since the vehicle's fan generates some noise, especially at higher fan settings, this noise needs to be compensated for. The noise level is directly related to the fan speed, so this parameter can be used to determine the noise level.

[0021] According to a second aspect, an audio processing method is provided for enhancing an input audio signal in a listening environment inside a vehicle with ambient noise. The audio processing method includes the following steps: The loudness of the audio signal can be adjusted by using a loudness adjustment level; The audio signal processed by the loudness adjustment level, or the audio signal further processed based on the audio signal processed by the loudness adjustment level, is equalized through the equalizer level. One or more acoustic signals are generated based on the audio signal and volume settings.

[0022] The audio processing method further includes: The ambient noise level is estimated based on the vehicle's speed. The level of the one or more acoustic signals is estimated based on the volume setting rather than on the analysis of the audio signal; The loudness adjustment level and the equalizer level are controlled based on the estimated levels of the ambient noise level and the estimated levels of the one or more acoustic signals.

[0023] The audio processing method according to the second aspect can be executed by the audio processing apparatus according to the first aspect. Therefore, other features of the audio processing method according to the second aspect arise directly from the functionality of the audio processing apparatus described in the first aspect and in its various implementations and embodiments as described above and below.

[0024] According to a third aspect, a computer program product is provided, the computer program product including a computer-readable storage medium for storing program code, which, when executed by a computer or processor, causes the computer or processor to perform the method according to the second aspect.

[0025] One or more embodiments will be described in detail in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description

[0026] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. In the drawings: Figure 1 This is a schematic diagram illustrating an audio processing apparatus provided in an embodiment for enhancing an input audio signal in a listening environment inside a vehicle with ambient noise. Figure 2 This is a schematic diagram illustrating a more detailed embodiment of an audio processing apparatus provided in the embodiments for enhancing input audio signals in a listening environment inside a vehicle with ambient noise. Figure 3 It is shown Figure 2 A schematic diagram of a variant of an embodiment of an audio processing device, which utilizes matrices instead of various tables; Figure 4 It is shown Figure 2A schematic diagram of a variant of an embodiment of an audio processing device, which uses two matrices instead of various tables to independently control the intensity of loudness compression and equalization effects; Figure 5 This is a flowchart illustrating the steps of an audio processing method provided in an embodiment for enhancing an input audio signal in a listening environment inside a vehicle with ambient noise.

[0027] In the following text, the same reference numerals refer to the same or at least functionally equivalent features. Detailed Implementation

[0028] In the following description, reference is made to the accompanying drawings, which form part of this invention and illustrate specific aspects of embodiments of the invention or to the drawings in which specific aspects of embodiments of the invention may be used. It should be understood that embodiments of the invention may be used in other aspects and may include structural or logical variations not depicted in the drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.

[0029] For example, it should be understood that the disclosure relating to the described method may also apply to the corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., functional units) to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units perform one or more of multiple steps respectively), even if the one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific apparatus is described based on one or more units (e.g., functional units), the corresponding method may include a step to perform the function of one or more units (e.g., one step performs the function of one or more units, or multiple steps perform the function of one or more of multiple units respectively), even if the one or more steps are not explicitly described or shown in the drawings. Furthermore, it should be understood that, unless otherwise stated, features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0030] Figure 1 This is a schematic diagram illustrating an audio processing device 100 provided in an embodiment for enhancing an audio input signal in a listening environment inside a vehicle with (e.g., experiencing) ambient noise.

[0031] The audio processing device 100 includes loudness adjustment levels 110 and 120 for adjusting the loudness of the audio input signal. Figure 1In the illustrated embodiment, the audio processing device 100 includes multiple dynamic processing units 110 and 120 for adjusting the loudness of the audio input signal. The first dynamic processing unit 120 includes a loudness range (LRA) reduction unit 120 for adjusting the loudness range of the audio input signal, and the second dynamic processing unit 110 includes a loudness unit full-scale LUFS compression unit 110 for adjusting the loudness of the audio input signal. In one embodiment, the LRA reduction unit 120 may operate based on a slow time constant (e.g., using long-time samples of the audio signal), while the LUFS compression unit 111 may operate based on a faster time constant (e.g., using shorter-time samples).

[0032] like Figure 1 As shown, the audio processing device 100 also includes an equalizer stage 130 for performing equalization, such as frequency-dependent gain adjustment using multiple frequency-dependent gains of the audio input signal processed by the loudness adjustment stages 110, 120, to obtain an audio signal for further processing (e.g., equalization loudness adjustment). It should be understood that the audio processing device 100 may include... Figure 1 Other audio processing blocks not shown. In particular, one or more other processing blocks may be implemented between the output of the LRA reduction unit 120 and the input of the equalizer stage 130. It should be further understood that, in this embodiment, the equalizer stage 130 takes as input an audio signal that has been further processed based on the audio input signal processed by the loudness adjustment stages 110, 120.

[0033] The audio processing device 100 also includes one or more transducers, such as speakers. Figure 1 (Not shown in the image) is used to generate one or more acoustic signals, such as one or more audio output signals, based on audio signals processed by loudness adjustment levels 110, 120 and equalizer level 130 and based on volume settings (e.g., settings of the volume knob of the audio processing device 100). One or more transducers (e.g., speakers) may be arranged in suitable locations within the vehicle interior.

[0034] like Figure 1As shown, the audio processing device 100 may include a plurality of control components 140, 150, 160, and 170, which will be described in more detail below and constitute control units 140, 150, 160, and 170. These control units are typically used to (a) estimate the ambient noise level based on the vehicle's speed, and (b) estimate the level of one or more acoustic signals based on volume settings rather than on the analysis of the audio signals, which will be described in more detail below. Control units 140, 150, 160, and 170, particularly their control component 140, are also used to control loudness adjustment levels 110 and 120 and equalizer level 130 based on estimates of the ambient noise level and estimates of the levels of one or more acoustic signals. In one embodiment, control component 140 is used to control loudness adjustment levels 110 and 120 and equalizer level 130 based on estimates of the ambient noise level and estimates of the levels of one or more acoustic signals by determining one or more frequency-dependent gains.

[0035] In another embodiment, control component 140 of control units 140, 150, 160, and 170 controls a first dynamic processing unit 120 (e.g., LRA reduction unit 120) based on a measured LRA of the audio signal, and controls a second dynamic processing unit 110 (e.g., LUFS compression unit 110) based on a measured LUFS of the audio signal. The audio processing device 100 may also include a microphone for detecting ambient noise, wherein control units 140, 150, 160, and 170, particularly their control component 160, are used to estimate the ambient noise level based on the vehicle's speed and the ambient noise detected by the microphone. In one embodiment, control units 140, 150, 160, and 170, particularly their control component 160, are also used to acquire the fan speed of the vehicle's fan and estimate the ambient noise level based on the vehicle's speed, fan speed, and / or the ambient noise detected by the microphone. In one embodiment, as will be described in more detail below, the LUFS compression unit 110 and the LRA reduction unit 120 can be adjusted by a control unit based on a scaling difference between the ambient noise level and the desired audio level, wherein the value can also be scaled based on the volume setting and corrected based on the content type.

[0036] In one embodiment, control units 140, 150, 160, and 170, particularly their control component 150, are used to determine the audio content type (e.g., speech, pop music, classical music, etc.) based on the audio input signal, and to control loudness adjustment levels 110 and 120 and equalizer level 130 based on the audio content type. To this end, in one embodiment, control component 150 may implement a machine learning (ML) model, wherein the ML model is used to determine the audio content type of the audio input signal.

[0037] It should be understood that, according to one embodiment, the audio processing device 100 is used to process the audio input signal using a cascade of three different audio adjustment levels (e.g., modules 110, 120, 130). As described above, the first level 110 compresses loudness, the second level 120 reduces the loudness range, and the third level 130 performs frequency-dependent gain adjustment. All these audio processing functions are adjusted by a control unit, specifically control component 140. Control component 140 determines control data based on information about the content type, ambient noise, and estimated music level. This data is acquired by other control components 150, 160, 170 and used to estimate the audio signal and vehicle parameters, such as which audio source to play (e.g., music, telephone, navigation, etc.), vehicle speed, fan level, microphone data, volume settings, etc.

[0038] One embodiment provides an audio processing apparatus 100 for generating different compression and level adjustments for different audio content types. For example, a voice signal (e.g., a news broadcast or telephone call) may require an additional boost of approximately 3 dB to maintain the same impression (compared to a music signal). This is achieved through a control component 150, which detects the content type using, for example, machine learning-based content type detection and evaluation of the audio source type (telephone, music, navigation).

[0039] In traditional VNC systems, loudness compression is typically used without additional modification to the loudness range. By using a cascade of loudness compressors and loudness range reduction modules, an audio processing apparatus 100 provided in one embodiment can achieve a wider range of loudness control. Simple LUFS-based compressors suffer from compression artifacts due to their relatively small time constants; in this scenario, this approach reduces dynamic range compression-related artifacts in music with high LRA.

[0040] Many conventional VNC algorithms use a feedback loop where the gain to be applied is added to the gain that would occur without VNC, and then a microphone is used to determine the true level of the audio signal being played. This feedback loop can cause various problems, ranging from level overshoot to instability. To avoid this undesirable side effect, these algorithms limit response time and possible amplification gain. The embodiments of the audio processing apparatus 100 disclosed herein do not require this feedback, thus overcoming these problems.

[0041] Algorithms based on psychoacoustic masks can achieve better clarity, but they can also lead to perceptible modulation of the signal due to continuous adjustment of the audio signal. This is acceptable for speech signals, but undesirable for high-end music playback. The embodiment of the audio processing apparatus 100 disclosed herein overcomes this problem by using only an estimate of the music level based on the volume playback setting as input to the control operation, which also models the psychoacoustic effect, thus eliminating modulation.

[0042] Figure 2 It shows Figure 1 A more detailed embodiment of a more general embodiment of the audio processing device 100. In Figure 2 In this embodiment, the audio processing device 100 includes corresponding processing blocks 231, 233, and 235 for converting the vehicle's speed, the vehicle's fan speed, and ambient noise detected in the microphone signal into source-specific noise level estimates. These levels are then combined into a total noise level by a noise summing block 237. The volume setting of the sound device is converted into a desired audio playback level by processing block 241, which maps the volume setting to dB (as indicated by the table shown near processing block 241). The estimated noise level is then subtracted from the estimated audio level, and the result is considered the audio-to-noise ratio (ANR). Based on psychoacoustic effects, different types of audio content require different ANR values ​​to provide comparable results in terms of sound impression. As an example, for a speech signal, an ANR 3 dB higher than that for a pop music signal is expected. To account for this effect, the content type of the audio signal is determined by processing block 221 and mapped to a compensation factor by processing block 223, which is then added to the ANR, the result of which is referred to herein as an "ANR modified" (ANRM) value.

[0043] Then, processing block 243 maps the ANRM values ​​to the desired enhancement. Then, for each frequency band, processing block 247 scales the desired enhancement again. To avoid unwanted enhancement at very low or very high volume values, processing block 251 scales the desired enhancement again according to the volume to reflect this characteristic. This scaled enhancement is then applied to the audio output signal.

[0044] Above this processing, processing blocks 201 and 211 also determine the EBU R128 LRA (or similar parameter) and LUFS (or similar parameter) values. Audio signals whose signal level changes significantly over time will have a larger LRA. LUFS only indicates the loudness of the signal, so the response time is much faster.

[0045] If ambient noise already affects audio perception, it is necessary to reduce the loudness range (LRA) but increase the loudness (LUFS). To do this, processing block 213 maps the LRA value to compressor threshold and ratio values ​​and uses a slow audio compressor 215 to compress the LRA. Then, processing block 203 also maps the LUFS value to compressor threshold and ratio parameters and applies these parameters using a fast audio compressor to increase the LUFS 205. As mentioned above, this compression can only be applied when the noise level is relatively high, to the point that it has already reduced the perceived quality of the audio. The metric for this is the ANRM value. For high ANRM values, compression is unnecessary because the audio signal can be estimated to be clearly audible. Therefore, to activate the compression only when the ANRM value is low, processing block 245 defines how much compression should be used by mapping the ANRM value to an intensity factor. In this embodiment, the intensity factor is only used for cross-gradient between the uncompressed and compressed signals (more complex embodiments may directly adjust the compression parameters based on the ANRM value). Compared to enhancements that already have volume mapping applied, the advantage of using default enhancement is that a single volume mapping curve can be applied, allowing compression even at high volume values ​​(which will increase the maximum sound pressure level of high dynamic range signals). Furthermore, volume values ​​are mapped to corresponding intensities in processing block 255. The output value of processing block 255 is multiplied by the output of processing block 245 and then used in the gain module to adjust the loudness intensity. In processing block 253, the desired gain in dB is converted to a linear value, which is then multiplied by the audio signal.

[0046] The disadvantage of using a compressor on the input audio signal (which is a common method) is that it results in permanent audio modulation, which is audible and degrades audio quality. Figure 2 The embodiment of the audio processing apparatus 100 shown overcomes this problem in two ways. First, instead of using a single compressor, it uses a cascade of two compressors that differ in response time and input signal domain. The slower compressor is used to reduce the loudness range, while the faster compressor is used to directly increase loudness through compression. In addition to independently controlling the loudness range and loudness itself, the intensity of the dual-cascade compression can be adjusted from a control block. The cascaded output is mixed with the original uncompressed signal. The mixing gain is determined by the control block, so that at sufficiently high sound pressure levels, the control block will adjust the mixing gain to utilize the original, unmodified signal. This prevents perceptible audio modulation. Figure 2In the illustrated embodiment, a machine learning-based algorithm can be used to determine the content type. Then, based on the detected content type, an offset relative to the audio-to-noise ratio is added, and this offset is retrieved from a table that maps different content types to relevant gains. The advantage of this approach is that optimal compensation can be achieved for each content type. Using content detection on audio signals can even support the detection of speech signals transmitted in streams where music is typically expected, such as in radio transmissions presenting news. Furthermore, different musical styles can be processed differently when mapped to content type. For example, classic recordings can be compressed more than popular music, which is typically already heavily compressed.

[0047] from Figure 2 As can be understood from the illustrated embodiments, the main control operation is feedback-free (except for the (optional) mapping of the microphone signal, which can be preprocessed to remove the music signal). Other implementations use feedback with a calculated gain value. This is typically done because increasing the gain increases the audio-to-noise ratio, so the value must be corrected to match the psychoacoustic model used in other implementations. Here, this is overcome by accessing a control table, which, for example, reduces the steepness, having a similar effect to feedback gain signals, but avoids the drawbacks of feedback (potential overshoot, level oscillations), which require reduced response time and gain to overcome. Therefore, the disclosed embodiments do not suffer from any control system instability, level overshoot, or even oscillation problems.

[0048] Figure 3 Another embodiment of the audio processing device 100 is shown. It should be understood that... Figure 3 The illustrated embodiments include Figure 2 The multiple processing blocks have been described in the context of the illustrated embodiment. Figure 3 The illustrated embodiments and Figure 2 The main difference in the illustrated embodiments is that, Figure 3 In the illustrated embodiment, matrix 301 is used instead of various functions to map noise and audio levels to gain. To avoid discontinuities, the values ​​in the matrix can be interpolated in the x (row) and y (column) directions to achieve smooth gain changes. Furthermore, the output gain of the matrix can be scaled by a user-adjustable intensity level. This allows users to deactivate the effect, reduce its intensity, or even increase it above normal levels. This intensity adjustment method can also be used for... Figure 2 The illustrated embodiment.

[0049] like Figure 3As shown, gain adjustment can be performed in frequency subbands. For this purpose, the input audio signal can be split into subband signals between the first and second compressors or after the second compressor. In a simple implementation, for each subband, a matrix 301 can be used to map the noise level and audio level to the gain of the relevant subband. To reduce the number of matrices, additional interpolation can be used in the frequency domain (between the subband matrices).

[0050] Using matrix 301 to map noise levels and audio levels to the desired gain greatly simplifies the process. Figure 3 The tuning process of the illustrated embodiment allows for very precise adjustment by explicitly assigning a parameter pair consisting of noise level and audio level to the target gain. This can be achieved, for example, by driving a car on a test track at a speed that reaches the noise level at which the matrix parameters should be adjusted, then setting the volume to the audio level to be adjusted, and then directly updating the corresponding matrix values ​​to achieve the desired compensation for that parameter combination. This ensures very consistent results.

[0051] When handling complete permutations of audio and noise parameters, it also ensures that the desired behavior is achieved in all prediction scenarios. Figure 2 The methods provided in the illustrated embodiments have more interdependencies, so changing one table when driving at a given speed and listening at a given audio level may also affect the system's behavior in another driving / listening situation. The advantage of extending to user-selectable effect intensity is that end users may want to adjust the intensity of the effect according to their personal preferences. Some end users prefer stronger compensation, while others prefer weaker compensation. As previously mentioned, this intensity adjustment feature can also be used... Figure 2 The illustrated embodiment.

[0052] Figure 4 Another embodiment of the audio processing device 100 is shown. It should be understood that... Figure 4 The illustrated embodiments include Figure 2 and Figure 3 The multiple processing blocks have been described in the context of the illustrated embodiment. Figure 4 The embodiments shown are similar to Figure 3 The main difference between the embodiments shown is that, Figure 4 In this embodiment, a single matrix (e.g., Table 401) is used to determine the intensity, rather than deriving the intensity of the compression path by mapping the enhancement from a first table to a linear intensity. The table inputs are also the estimated noise level and the estimated audio level, and the output is the compression intensity. Compared with other matrices (e.g.... Figure 3Similar to Table 301 in the embodiment, matrices (e.g., Table 401) can also be specified for each subband (this also requires compression within the subband). The output gain from the two matrices 301 and 401 is scaled by a user-adjustable intensity level. This allows users to deactivate the effect, reduce its intensity, or even increase the intensity above normal levels.

[0053] like Figure 4 As shown, dynamic range modification and gain adjustment can both be performed within the frequency subbands. For this, the input audio signal must be split into subband signals between the first and second compressors or after the second compressor. In a simple implementation, for each subband, a pair of matrices 301 and 401 can be used to map the noise level and audio level to the loudness modification intensity and gain of the relevant subband. To reduce the number of matrices, additional interpolation can be used in the frequency domain (between the subband matrices).

[0054] Figure 4 The main advantage of the illustrated embodiment comes from the ability to independently adjust the dynamic range and gain of different frequency bands. This is in Figure 4 This is also visible in matrices 301 and 401. For example, at an audio level of 20 dB(A), no gain is applied, but dynamic range modification is still performed. The motivation behind this is that at low volume values, no additional enhancement of the normal signal is needed. However, if the audio signal has a high dynamic range or a high loudness range in this case, compression is still required.

[0055] for Figure 1 , Figure 2 , Figure 3 and Figure 4 The illustrated embodiment can also utilize other methods for obtaining estimates of perceived loudness, instead of actual LUFS measurements according to the EBU R128 standard. Typical implementations can be based on dB(A), dB(C) measurement standards and loudness models, such as the so-called "Zwicker loudness." Simpler approximations of the loudness range can also be made; LRA measurements according to the EBU R128 standard described for the above embodiments are just one possible implementation. Alternative embodiments can directly observe changes in dB(A) or dB(C) values ​​over time. For content type detection, classical signal processing methods can be used instead of machine learning-based algorithms, or even information about the connection source or any other information about the track being played. Compression parameters can also be adjusted to become transparent when the audio signal level is sufficiently high (the compressor's static characteristics are linear), rather than mixing the output of the cascaded compressor with the original signal.

[0056] Figure 5This is a flowchart illustrating an audio processing method 500 for enhancing an audio signal inside a vehicle with ambient noise. The audio processing method 500 includes step 501: adjusting the loudness of the audio signal via loudness adjustment levels 110, 120. Furthermore, the audio processing method 500 includes step 503: equalizer level 130 equalizes the audio signal processed by loudness adjustment levels 110, 120, or an audio signal further processed based on the audio signal processed by loudness adjustment levels 110, 120. The method 500 also includes step 505: generating one or more acoustic signals based on the audio signal and volume settings. The audio processing method 500 further includes step 507: estimating the ambient noise level based on the vehicle's speed; and step 509: estimating the levels of one or more acoustic signals based on volume settings rather than on analysis of the audio signal. The audio processing method 500 also includes step 509: controlling the loudness adjustment level and equalizer level based on the estimated ambient noise level and the estimated levels of one or more acoustic signals.

[0057] The audio processing method 500 can be executed by the audio processing apparatus 100 provided in the embodiments. Therefore, further features of the audio processing method 500 directly derive from the functionality of the audio processing apparatus 100 and its various embodiments described above and below. It should be further understood that the steps of the audio processing method 500 can be different from those described below. Figure 5 Implement it in the order shown.

[0058] Those skilled in the art will understand that “blocks” (“units”) in the various figures (methods and apparatuses) represent or describe the functionality of embodiments of the invention (and are not necessarily independent “units” in hardware or software), thereby equally describing the functionality or features (unit = step) of apparatus embodiments and method embodiments.

[0059] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the described apparatus embodiments are merely exemplary. For example, the unit division is merely a logical functional division, and in actual implementation, it can be another division. For example, multiple units or components can be merged or integrated into another system, or some features can be ignored or not performed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be implemented through some interface. Indirect coupling or communication connection between devices or units can be implemented electronically, mechanically, or otherwise.

[0060] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiment solution according to actual needs.

[0061] In addition, the functional units in the embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

Claims

1. An audio processing apparatus (100) for enhancing audio signals inside a vehicle with ambient noise, characterized in that, The audio processing device (100) includes: Loudness adjustment levels (110, 120) are used to adjust the loudness of the audio signal; An equalizer level (130) is used to equalize the audio signal processed by the loudness adjustment levels (110, 120), or an audio signal further processed based on the audio signal processed by the loudness adjustment levels (110, 120); One or more transducers for generating one or more acoustic signals based on the audio signal and volume setting; Control units (140, 150, 160, 170) are configured to (a) estimate the ambient noise level based on the speed of the vehicle, and (b) estimate the level of the one or more acoustic signals based on the volume setting rather than on the analysis of the audio signals, wherein the control units (140, 150, 160, 170) are further configured to control the loudness adjustment level (110, 120) and the equalizer level (130) based on the estimated ambient noise level and the estimated level of the one or more acoustic signals.

2. The audio processing apparatus (100) according to claim 1, characterized in that, The loudness adjustment levels (110, 120) include a plurality of dynamic processing units (110, 120) for adjusting the loudness of the audio signal, wherein the control unit (140) is used to control a first dynamic processing unit (120) based on a measured loudness range of the audio signal and to control a second dynamic processing unit (110) based on a measured loudness of the audio signal.

3. The audio processing apparatus (100) according to claim 2, characterized in that, The first dynamic processing unit (120) includes a loudness range reduction unit (120) for adjusting the loudness range of the audio signal.

4. The audio processing apparatus (100) according to claim 2 or 3, characterized in that, The second dynamic processing unit (110) includes a loudness unit full-scale compression unit (110) for adjusting the loudness of the audio signal.

5. The audio processing apparatus (100) according to any one of the preceding claims, characterized in that, The control unit (150) is used to determine the audio content type based on the audio signal, and to control the loudness adjustment level (110, 120) and the equalizer level (130) based on the audio content type.

6. The audio processing apparatus (100) according to claim 5, characterized in that, The control unit (150) is used to implement a machine learning model, wherein the ML model is used to determine the audio content type of the audio signal.

7. The audio processing apparatus (100) according to any one of the preceding claims, characterized in that, The control unit (150) is used to control the loudness adjustment level (110, 120) and the equalizer level (130) by determining one or more frequency-related gains based on the estimated value of the ambient noise level and the estimated value of the level of the one or more acoustic signals.

8. The audio processing apparatus (100) according to any one of claims 1 to 7, characterized in that, The control unit (150) is configured to control the loudness adjustment level (110, 120) and the equalizer level (130) based on the estimated value of the ambient noise level and the estimated value of the level of the one or more acoustic signals by using one or more matrices to map the estimated value of the ambient noise level and / or the estimated value of the level of the one or more acoustic signals to one or more frequency-dependent gains and / or one or more loudness adjustment parameters, wherein the control unit (150) is configured to control the equalizer level (130) and / or the loudness adjustment level (110, 120) based on the one or more frequency-dependent gains and / or the one or more loudness adjustment parameters.

9. The audio processing apparatus (100) according to claim 8, characterized in that, Each of the one or more matrices includes a plurality of matrix values, and the control unit (150) is used to interpolate the plurality of matrix values ​​to map the estimated values ​​of the ambient noise level and / or the estimated values ​​of the levels of the one or more acoustic signals to the one or more frequency-related gain and / or the one or more loudness adjustment parameters.

10. The audio processing apparatus (100) according to any one of the preceding claims, characterized in that, The audio processing device (100) further includes a microphone for detecting ambient noise, and the control unit (160) is used to estimate the ambient noise level based on the vehicle's speed and the ambient noise detected by the microphone.

11. The audio processing apparatus (100) according to any one of the preceding claims, characterized in that, The control unit (160) is used to acquire the fan speed of the vehicle's fan and estimate the ambient noise level based on the vehicle's speed and the fan speed.

12. An audio processing method (500) for enhancing audio signals inside a vehicle with ambient noise, characterized in that, The audio processing method (500) includes: The loudness of the audio signal (501) is adjusted by loudness adjustment levels (110, 120); The audio signal processed by the loudness adjustment levels (110, 120) or the audio signal further processed based on the audio signal processed by the loudness adjustment levels (110, 120) is equalized (503) by the equalizer level (130). Generate (505) one or more acoustic signals based on the audio signal and volume setting; The audio processing method (500) further includes: (a) estimating (507) the ambient noise level based on the speed of the vehicle, and (b) estimating (509) the level of the one or more acoustic signals based on the volume setting rather than on the analysis of the audio signals, and controlling (511) the loudness adjustment level (110, 120) and the equalizer level (130) based on the estimated value of the ambient noise level and the estimated value of the level of the one or more acoustic signals.

13. A computer program product, characterized in that, Includes a computer-readable storage medium for storing program code, which, when executed by a computer or processor, causes the computer or processor to perform the method (500) according to claim 12.