An audio loudness adjustment method and device, a terminal device, and a storage medium
Patent Information
- Application Number
- CN202211652388.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-21
AI Technical Summary
受限于尺寸和输出功率,微扬声器往往不具备大尺寸高保真(High-Fidelity,Hi-Fi)音箱所拥有的全频段相对平坦的频响曲线,且其只能输出较低的总响度水平,不利于音乐的真实准确回放
[0019]本申请实施例提供一种音频响度调节方法、终端设备及存储介质,该方法包括:获取多条目标候选群时延响应曲线,并确定多条目标候选群时延响应曲线对应的多组滤波器参数;将多组滤波器参数分别赋值给多个全通滤波器组,其中,一组滤波器参数与一个全通滤波器组对应;将预设音源信号输入多个全通滤波器组中进行滤波处理,得到每个全通滤波器组对应的滤波音源信号;对滤波音源信号进行响度预测,确定滤波音源信号对应的全通滤波器组的响度预测值;采用响度预测值最大的一组全通滤波器对待播放音频信号进行响度调节。采用上述实现方案,在对音频响度进行调节的过程中,引入全通滤波器组,并根据多个全通滤波器组对应的多组滤波器参数对输入的音源信号进行滤波处理,基于全通滤波器能够改变时延的特性,针对于不同频率,调整音源信号在该频率下对应的群时延,错开不同频率正弦波的峰值,使得同一时刻不同频率的峰值不会被叠加,从而维持每个单频正弦波的形状不被改变,避免了截顶失真,从而减小因谐波失真导致的音源信号受损,且为了保证响度增强效果更优,分别对每个全通滤波器组滤波后的音源信号进行响度值预测,进一步确定响度增强值最大的滤波器组,利用该全通滤波器组不仅能够减少谐波失真对音源音色的影响,还能够显著提升微扬声器重放音频的响度。
Smart Images

Figure CN115811690B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio loudness adjustment technology, and in particular to an audio loudness adjustment method, apparatus, terminal device and storage medium. Background Technology
[0002] To meet consumer demand for lightweight, thin, and small portable devices, mobile phones, tablets, and other devices are typically equipped with 2 to 4 miniature speakers with small voice coil amplitudes and a maximum power of only about 1 watt for external audio output. Limited by size and output power, these miniature speakers often lack the relatively flat frequency response curve across the entire frequency range found in larger high-fidelity (Hi-Fi) speakers, and they can only output a lower overall loudness level, which is detrimental to the accurate and realistic reproduction of music.
[0003] In existing technologies, dynamic range control (DRC) is used to process audio source signals in order to improve the reproduction loudness of micro loudspeakers. However, the loudness improvement effect is limited for audio source signals that contain few low-frequency and high-frequency components. Furthermore, the non-linear processing of audio source signals using limiting and dynamic compression introduces significant harmonic distortion when the limiting and compression are performed on a large scale, resulting in changes to the timbre of the audio source. Summary of the Invention
[0004] In view of this, the embodiments of this application aim to provide an audio loudness adjustment method, apparatus, terminal device and storage medium that can significantly improve the loudness of the audio reproduced by the micro-speaker, while reducing the impact of harmonic distortion on the timbre of the sound source.
[0005] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide an audio loudness adjustment method, applied to an audio loudness adjustment device, the audio loudness adjustment device including a plurality of all-pass filter groups, each all-pass filter group having a set of all-pass filters connected in series, the method including:
[0007] Obtain multiple target candidate group delay response curves and determine multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves;
[0008] Multiple sets of filter parameters are assigned to multiple all-pass filter groups, where each set of filter parameters corresponds to one all-pass filter group.
[0009] The preset audio source signal is input into multiple all-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each all-pass filter group;
[0010] The loudness of the filtered audio source signal is predicted to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal; the loudness of the audio signal to be played is adjusted using the all-pass filter group with the largest loudness prediction value.
[0011] Secondly, embodiments of this application provide an audio loudness adjustment device, the audio loudness adjustment device comprising:
[0012] The acquisition module is used to acquire multiple target candidate group delay response curves and determine multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves;
[0013] The assignment module is used to assign multiple sets of filter parameters to multiple all-pass filter groups, where each set of filter parameters corresponds to one all-pass filter group.
[0014] The filtering module is used to input the preset audio source signal into multiple all-pass filter groups for filtering processing, so as to obtain the filtered audio source signal corresponding to each all-pass filter group.
[0015] The determination module is used to predict the loudness of the filtered sound source signal and determine the predicted loudness value of the all-pass filter bank corresponding to the filtered sound source signal.
[0016] The adjustment module is used to adjust the loudness of the audio signal to be played by using a set of all-pass filters with the largest loudness prediction value.
[0017] Thirdly, embodiments of this application provide a terminal device, which includes a processor, a memory, and a communication bus; the processor executes the running program stored in the memory to implement the above-mentioned audio loudness adjustment method.
[0018] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described audio loudness adjustment method.
[0019] This application provides an audio loudness adjustment method, terminal device, and storage medium. The method includes: acquiring multiple target candidate group delay response curves and determining multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves; assigning the multiple sets of filter parameters to multiple all-pass filter groups, wherein one set of filter parameters corresponds to one all-pass filter group; inputting a preset audio source signal into the multiple all-pass filter groups for filtering processing to obtain a filtered audio source signal corresponding to each all-pass filter group; performing loudness prediction on the filtered audio source signal to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal; and using the all-pass filter group with the largest loudness prediction value to adjust the loudness of the audio signal to be played. The above implementation scheme introduces an all-pass filter bank during the adjustment of audio loudness. Multiple filter parameters corresponding to these all-pass filter banks are used to filter the input audio signal. Based on the characteristic that all-pass filters can change time delay, the group delay of the audio signal at different frequencies is adjusted to stagger the peak values of sine waves at different frequencies. This prevents the peak values of different frequencies from being superimposed at the same time, thus maintaining the shape of each single-frequency sine wave and avoiding truncated distortion. This reduces damage to the audio signal caused by harmonic distortion. Furthermore, to ensure better loudness enhancement, loudness value prediction is performed on the audio signal filtered by each all-pass filter bank to further determine the filter bank with the largest loudness enhancement value. Using this all-pass filter bank not only reduces the impact of harmonic distortion on the timbre of the audio source but also significantly improves the loudness of the audio reproduced by the micro-speaker. Attached Figure Description
[0020] Figure 1 This application provides a flowchart of an audio loudness adjustment method. Figure 1 ;
[0021] Figure 2 A flowchart illustrating the process of traversing and searching for the filter parameters of an all-pass filter, as provided in this embodiment of the application;
[0022] Figure 3 A schematic diagram of an all-pass filter bank connection is provided for an embodiment of this application;
[0023] Figure 4 A schematic diagram of a family of candidate group delay curves provided in an embodiment of this application;
[0024] Figure 5 A schematic diagram of the zeros and poles of a second-order all-pass filter in the Z-plane provided for an embodiment of this application;
[0025] Figure 6 This application provides an exemplary schematic diagram illustrating the correspondence between the sampling frequency and delay of an all-pass filter.
[0026] Figure 7 A schematic diagram of the original amplitude frequency response curve of an exemplary micro loudspeaker provided in this application embodiment;
[0027] Figure 8 A schematic diagram of an exemplary loudness prediction module provided in an embodiment of this application;
[0028] Figure 9 An exemplary audio loudness adjustment method flow provided in this application embodiment Figure 2 ;
[0029] Figure 10 An exemplary audio loudness adjustment method flow provided in this application embodiment Figure 3 ;
[0030] Figure 11 An exemplary audio loudness adjustment method flow provided in this application embodiment Figure 4 ;
[0031] Figure 12 This is a schematic diagram of the structure of an audio loudness adjustment device 1 provided in an embodiment of this application;
[0032] Figure 13 This is a schematic diagram of the structure of a terminal device 2 provided in an embodiment of this application. Detailed Implementation
[0033] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the technical solution of this application will be further described in detail below with reference to the accompanying drawings and specific embodiments. The accompanying drawings are for reference only and are not intended to limit the embodiments of this application.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be limiting of this application.
[0035] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. It should also be noted that the terms "first / second / third" used in the embodiments of this application are merely for distinguishing similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.
[0036] In existing technologies, the following technical solutions are typically used to enhance the loudness of a frequency band:
[0037] (1) Loudness enhancement based on frequency band compression
[0038] Loudness enhancement methods based on frequency band compression first compress and attenuate the low-frequency and high-frequency bands of the sound source signal before sending it to the loudspeaker for playback. According to psychoacoustic equal-loudness curves, the human ear is less sensitive to low-frequency loudness than mid- and high-frequency loudness at the same sound pressure level. Furthermore, miniature loudspeakers have weak reproduction capabilities in the low-frequency range below 200Hz and the extremely high-frequency range above 10kHz. Therefore, given limited power, the loudspeaker's output energy can be allocated as much as possible to the mid- and high-frequency ranges where the loudspeaker has strong reproduction capabilities and where the human ear is more sensitive, while attenuating the amplitude of low-frequency and extremely high-frequency signals that contribute less to the overall loudness. This method is a common technique in miniature loudspeaker tuning, but its loudness enhancement effect is limited for some sound sources that inherently have few low- and high-frequency components.
[0039] (2) Loudness enhancement based on the combination of dynamic range compression (DRC) and automatic gain control (AGC)
[0040] A common commercial loudness enhancement plugin, OneKnob Louder, combines a limiter, dynamic range control (DRC), and automatic gain compensation (AGC). It first clips peak signals exceeding a threshold in the input signal, limiting the overall amplitude of the input signal to below a fixed threshold. Then, it applies a variable gain to compress large-amplitude signals and boost small-amplitude signals. Since both limiting and dynamic compression are non-linear processes, significant limiting and compression can introduce noticeable harmonic distortion, altering the timbre.
[0041] To address the technical problems in the prior art, this application provides an audio loudness adjustment method, such as... Figure 1 As shown, this method is applied to an audio loudness adjustment device, which includes multiple all-pass filter banks, each containing a set of all-pass filters connected in series. The method may include:
[0042] S101. Obtain multiple candidate group delay response curves and determine a set of filter parameters corresponding to each candidate group delay response curve.
[0043] In this embodiment of the application, when adjusting the loudness of the audio, multiple full-pass filters are connected in series to form a full-pass filter group, and the loudness of the audio is adjusted linearly and dynamically using the full-pass filter group.
[0044] In the embodiments of this application, the all-pass filter does not change the frequency characteristics of the input signal, but it does change the phase of the input signal. Utilizing this characteristic, the all-pass filter can be used as a delay device, delay equalizer, etc.
[0045] In this embodiment of the application, before adjusting the loudness of the audio source signal, it is necessary to obtain the filter parameters corresponding to each full-pass filter in each full-pass filter group, and use the obtained filter parameters corresponding to each full-pass filter in each full-pass filter group to filter the audio source signal.
[0046] In this embodiment, before obtaining multiple target candidate group delay curves, a search can be performed on a large number of loudspeakers using a traversal search method to search for the filter parameters corresponding to each filter in each filter bank. Specifically, before obtaining multiple target candidate group delay response curves, within a first preset search range, according to a first preset search step size, the pole amplitude values corresponding to each full-pass filter in each full-pass filter bank are adjusted to determine multiple target pole amplitude values corresponding to each full-pass filter; within a second preset search range, according to a second preset search step size, the pole angle values corresponding to each full-pass filter are adjusted to determine multiple target pole angle values corresponding to each full-pass filter; based on the multiple target pole amplitude values and multiple target pole angle values, multiple sets of initial filter parameters corresponding to multiple full-pass filter banks are determined; based on the multiple sets of initial filter parameters, multiple candidate group delay response curves are determined.
[0047] In this embodiment of the application, the filter parameters corresponding to each full-pass filter in multiple full-pass filter banks can be obtained by pre-determining the filter parameters corresponding to each full-pass filter in each full-pass filter bank using an offline traversal search method.
[0048] It should be noted that, considering the possibility of finding several filter banks with similar or identical loudness enhancement effects, the top M filter banks with the best performance can be selected offline as candidate combinations for subsequent real-time filtering. Specifically, the process of obtaining the filter parameters corresponding to each full-pass filter in multiple full-pass filter banks through a traversal search is as follows: Figure 2 As shown:
[0049] 1. Specify the search range of parameter r for each all-pass filter (APF) as [r1, r2] and the search step size as dr. Let there be k1 possible choices for r.
[0050] 2. Specify the search range of parameter w for each all-pass filter (APF) as [w1, w2] and the search step size as dw. Let there be k2 possible choices for w.
[0051] 3. Within the search range specified by r and w, continuously change the parameters r and w of each all-pass filter APF according to the search step size to generate N new all-pass filters;
[0052] 4. Connect N new all-pass filters in series to form a filter group, and repeat the above steps 1-3 until K filter groups are determined.
[0053] In this application embodiment, exemplarily, a filter bank is cascaded as follows: Figure 3 As shown, a filter bank can contain several all-pass filters.
[0054] It should be noted that, since changing the order of each sub-filter within the N filters does not change the final overall filtering effect when N filters are used in series, according to the permutation and combination formula, there are a total of K = (k1*k2+N-1)! / N! / k1*k2)! different ways to combine N filters in series.
[0055] It should be noted that the value of N can be selected according to the actual situation, and no specific limitation is made in this application. K is the number of filter banks obtained by adjusting the r and w parameters. These K combinations are the multiple all-pass filter banks determined.
[0056] In this embodiment, based on the above-mentioned traversal search method, an offline traversal search is performed in advance for a large number of different types of loudspeakers to search for the filter parameters corresponding to each full-pass filter in the full-pass filter bank for loudness enhancement, thereby obtaining the filter parameters corresponding to each full-pass filter in multiple full-pass filter banks.
[0057] In this embodiment of the application, after obtaining the filter parameters corresponding to each full-pass filter in multiple full-pass filter banks, multiple sets of initial filter parameters corresponding to multiple full-pass filter banks are obtained, and then multiple initial candidate group delay response curves are determined based on the multiple sets of initial filter parameters.
[0058] In this embodiment of the application, determining multiple initial candidate group delay response curves based on multiple sets of initial filter parameters can be achieved by using multiple sets of initial filter parameters to determine multiple transfer functions of multiple all-pass filter banks; using multiple transfer functions to determine multiple phase frequency response parameters of multiple all-pass filter banks; and using multiple phase frequency response parameters to determine multiple initial candidate group delay response curves.
[0059] In this embodiment of the application, to adjust the loudness of the audio signal, it is necessary to obtain the filter parameters corresponding to each filter in the all-pass filter bank. The filter parameters can be determined based on the group delay response curves of multiple target candidate groups and the group delay response curve corresponding to the preset equalization filter.
[0060] In this embodiment of the application, before determining the filter parameters, it is first necessary to obtain the group delay response curves of multiple target candidate groups and the group delay response curves corresponding to the preset equalization filter.
[0061] In this embodiment of the application, to obtain the target candidate group delay response curve, multiple sets of initial filter parameters need to be determined first by offline traversal search. Then, based on the multiple sets of initial filter parameters, multiple initial candidate group delay response curves are determined. From the multiple initial candidate group delay response curves, one corresponding APF filter combination with the best loudness enhancement effect is selected.
[0062] In this embodiment of the application, considering that the actual process may find several filter groups with similar or identical loudness enhancement effects, the above-mentioned offline traversal method can be used to select the top M filter groups with the best effect as candidate combinations for subsequent real-time filtering.
[0063] It should be noted that the number of M can be selected according to the actual situation, and no specific limit is made in this application.
[0064] It should be noted that in the embodiments of this application, multiple second-order IIR all-pass filters (APFs) are connected in series to fine-tune the group delay of the input signal. The APFs do not change the characteristics of the signal amplitude and frequency response, and can keep the amplitude of each single-frequency sinusoidal signal in the original signal unchanged. Moreover, the implementation is simple and the total computational load is very small even when multiple are connected in series.
[0065] In this embodiment of the application, after determining the values of multiple initial filter parameters r and w, the transfer function of the all-pass filter can be determined using formula (1), as shown below:
[0066]
[0067] It should be noted that the transfer function of an all-pass filter is determined by two coefficients, r and w.
[0068] Where Z represents the Z-plane, and any point in the Z-plane can be... The diagram showing the poles and zeros of a second-order all-pass filter in the Z-plane is shown below. Figure 5 As shown, the two x's represent two conjugate symmetric poles, and the two o's represent two conjugate symmetric zeros. r represents the amplitude of the pole corresponding to this all-pass filter in the Z-plane, with a value range of (0, 1); w represents the angle between the pole and the positive z-axis, in radians (rad).
[0069] In this embodiment of the application, after calculating A(z), the phase frequency response φ(e) of the system is obtained using the expression for A(z). j ωThe calculation expression is shown in formula (2):
[0070]
[0071] Where atan represents the arctangent function, imag represents the imaginary part of the expression within the parentheses, and real represents the real part of the expression within the parentheses.
[0072] In this embodiment of the application, given the calculated phase frequency response, the known phase frequency response φ(e) is used. jω Find the group delay frequency response parameter GD(e) jω Its calculation expression is shown in formula (3):
[0073]
[0074] It should be noted that the group delay frequency response curve is equal to the negative value of the derivative of the phase frequency response with respect to the angular frequency ω.
[0075] In this embodiment of the application, the candidate group delay response curve corresponding to an all-pass filter bank can be calculated according to the above formula. The calculation method for the candidate group delay response curves corresponding to the remaining all-pass filter banks is the same as above, and will not be repeated here.
[0076] In this embodiment of the application, after determining multiple initial candidate group delay response curves, statistical analysis is performed on the multiple initial candidate group delay response curves based on a preset algorithm to determine the group delay response curve cluster; based on the group delay response curve cluster and through manual listening identification, multiple target candidate group delay response curves corresponding to multiple all-pass filter banks are determined from the group delay response curve cluster.
[0077] In this embodiment, after obtaining the filter parameters corresponding to the all-pass filters in the all-pass filter bank for a large number of different types of loudspeakers, the features of the curves can be statistically summarized by trained neural networks, machine learning and other methods. Several sets of better initial candidate group delay curves are selected from the initial candidate group delay curves corresponding to a large number of filter parameters as the group delay response curve cluster.
[0078] It should be noted that the number of optimal initial candidate group delay curves can be selected according to the actual situation, and no specific limit is made in this application.
[0079] In this embodiment of the application, after determining the group delay response curve cluster, and through human hearing test, multiple target candidate group delay response curves corresponding to multiple all-pass filter banks with the best loudness enhancement effect are selected from the group delay response curve cluster.
[0080] In this embodiment of the application, after obtaining multiple target candidate group delay response curves with the best enhancement effect, the obtained multiple target candidate group delay response curves and the group delay response curves corresponding to the preset equalization filter can be used to further determine a set of filter parameters corresponding to multiple all-pass filter groups.
[0081] In this embodiment, determining multiple sets of filter parameters corresponding to multiple candidate target group delay curves is calculated based on the delay response curves of multiple candidate target groups and the group delay response curves corresponding to a preset equalization filter. Specifically, the delay response curves corresponding to the preset equalization filter can be used to process the delay response curves of multiple candidate target groups corresponding to multiple all-pass filter banks by at least subtraction, resulting in multiple target group delay response curves corresponding to multiple all-pass filter banks. Based on the multiple target group delay response curves, multiple target phase frequency response parameters are determined. Using the multiple target phase frequency response parameters, multiple target transfer functions are determined. Using the multiple target transfer functions, a set of filter parameters corresponding to the delay response curves of multiple candidate target groups is determined.
[0082] In this embodiment of the application, to adjust the loudness of an audio signal using an all-pass filter bank, it is necessary to obtain the filter parameters corresponding to each filter in the all-pass filter bank. The filter parameters can be determined based on the target group delay response curve. One target group delay response curve corresponds to a set of filter parameters. That is, when multiple sets of target group delay response curves are obtained, the filter parameters corresponding to each filter in multiple all-pass filter banks can be determined based on the multiple sets of target group delay response curves.
[0083] In this embodiment of the application, after obtaining multiple target candidate group delay response curves corresponding to multiple all-pass filter banks, the multiple target candidate group delay response curves are subtracted from the group delay response curves corresponding to the obtained preset equalization filter to obtain multiple sets of target group delay response curves.
[0084] It should be noted that, in the embodiments of this application, obtaining multiple sets of target group delay response curves corresponding to multiple all-pass filter banks can be achieved by subtracting the multiple target candidate group delay response curves corresponding to multiple all-pass filter banks from the group delay response curve corresponding to the preset equalization filter to obtain multiple sets of target group delay response curves. Alternatively, any method other than subtraction can be used to obtain the target group delay response curves, all of which fall within the scope of protection of this application. The specific calculation method can be selected according to the actual situation, and no specific limitation is made in this application.
[0085] It should be noted that the calculation method for the group delay response curve corresponding to the preset equalization filter can refer to the calculation method in the existing technology, and will not be repeated here.
[0086] In the embodiments of this application, after the target group delay response curve is determined, one target group delay response curve corresponds to a set of filter parameters. That is, when multiple sets of target group delay response curves are obtained, the filter parameters corresponding to each filter in multiple all-pass filter banks can be determined based on the multiple sets of target group delay response curves.
[0087] It should be noted that when obtaining the target group delay response curve, the loudness enhancement filter parameters can be searched offline for a large number of different loudspeakers to obtain multiple optimal filter groups. Then, through human hearing tests, the candidate group delay characteristic curve family with the best loudness enhancement effect can be statistically summarized. Subsequently, for any micro loudspeaker, the candidate group delay curve with the best effect and the corresponding filter group parameters can be selected from the candidate group delay characteristic curve family. There is no need to search for the all-pass filter group with the best loudness enhancement effect for each loudspeaker through traversal search.
[0088] It should be noted that the method of performing offline traversal search of filter parameters in advance can refer to the traversal search method in the above embodiment, and will not be repeated here.
[0089] In the embodiments of this application, the candidate group delay curve family is as follows: Figure 4 As shown, the group delay of each curve in the candidate group delay curve family roughly decreases monotonically with increasing frequency. This aligns with the human ear's auditory characteristics: it is insensitive to changes in the phase and delay differences of low-frequency signals but sensitive to changes in the phase and delay differences of high-frequency signals. The approach of modifying only the delay in the mid-to-low frequency range while maintaining the original delay in the mid-to-high frequency range avoids the human ear perceiving the influence of the all-pass filter (APF) on the signal, thus preserving the accurate reproduction of the original sound source signal's timbre as much as possible.
[0090] It should be noted that, preferably, the shape of the curves formed by the above candidate group delay characteristic curve family can be adjusted locally, such as by increasing the delay at a few frequency points and changing the shape of the curves formed by the basic group delay characteristic parameters, thereby obtaining a new family of group delay characteristic curves.
[0091] In this embodiment, by using the group delay response curves corresponding to the preset equalization filter and processing the multiple candidate group delay response curves corresponding to multiple all-pass filter banks through the subtraction method, multiple sets of target group delay response curves corresponding to multiple all-pass filter banks can be obtained. Then, the target phase frequency response parameter can be determined based on the obtained target group delay response curves. Subsequently, the target transfer function can be determined based on the target phase frequency response parameter. Finally, the filter parameters corresponding to the all-pass filter banks can be determined based on the target transfer function.
[0092] Specifically, the calculation process of formulas (1)-(3) above can be used to determine the filter parameters corresponding to each all-pass filter group. The filter parameters can be determined by reverse calculation using formulas (1)-(3). The specific calculation process will not be described in detail in this application.
[0093] It should be noted that after summarizing and generalizing multiple candidate group delay response curve families, the computational cost and time consumption of subsequent optimal APF search based on the candidate group delay response curve families can be significantly reduced.
[0094] The specific execution process for determining the target group delay response curve can be as follows: first, import multiple candidate group delay response curves (GD). i (i = 1, 2, 3, ..., Q), and the group delay response curve GD0 of the equalization filter H1, for each candidate group delay response curve GD i With GD i -GD0 is the target group delay response curve of the i-th APF filter bank. The r, w, and other parameters of the N APFs that satisfy the curve formed by this target group delay parameters are solved using a common approximation algorithm.
[0095] It should be noted that the r and w parameters of the N APFs corresponding to the other target group delay response curves can be obtained by referring to the above solution process, and will not be repeated here.
[0096] It should be noted that, using the above embodiments, a set of filter parameters corresponding to multiple all-pass filter banks are determined based on the time delay response curves of multiple target groups, that is, the filter parameters corresponding to each filter in multiple all-pass filter banks are determined, thereby obtaining multiple filter banks that can be applied to audio loudness adjustment.
[0097] It should be noted that the method described above for determining a set of filter parameters corresponding to multiple all-pass filter banks is to first obtain the filter parameters by priority traversal search, and then determine the corresponding candidate group delay curve cluster based on the obtained filter parameters. When necessary, the candidate group delay curve cluster can be used directly to obtain the filter parameters.
[0098] It should be noted that in order to change the relative delay of sinusoidal signals of different frequencies in this embodiment, and since group delay can characterize the time dispersion of the filter for signals of different frequencies, once the group delay curve is determined, the APF filter parameters corresponding to this candidate group delay curve can be quickly solved by the optimal approximation method.
[0099] It should be noted that because the statistical characteristics of actual music signals are randomly changing, the number of frequencies, the initial relative delay of each frequency, and the amplitude of each frequency are all uncertain. Therefore, the optimal group delay change required for different music segments is also different. Even for the same music segment, different EQ equalizers selected when tuning different microspeakers will lead to changes in the initial relative delay of the music signal. Therefore, before actual use, one or more sets of APF parameters should be selected through experiments for the specified microspeaker that can significantly improve the loudness of most types of music. Alternatively, the inherent law or formula between the characteristics of the APF group delay curve and the loudness improvement of random music should be statistically derived or theoretically derived. This will facilitate the rapid determination of the optimal set of filter parameters corresponding to the APF filter bank for any subsequent microspeaker.
[0100] Optionally, in the embodiments of this application, when determining the filter parameters, the sound source signal can be filtered by traversing and searching for parameters while using the filtered parameters obtained from the traversal search. The efficiency of this method is lower than that of first traversing and summarizing the cluster of time delay curves.
[0101] In this embodiment of the application, before assigning multiple sets of filter parameters to multiple all-pass filter groups, the method for obtaining a set of filter parameters corresponding to each of the multiple all-pass filter groups can be as follows: within a first preset search range, according to a first preset search step size, adjust the pole amplitude value corresponding to each all-pass filter in each all-pass filter group to determine multiple target pole amplitude values corresponding to each all-pass filter; within a second preset search range, according to a second preset search step size, adjust the pole angle value corresponding to each all-pass filter to determine multiple target pole angle values corresponding to each all-pass filter; and based on the multiple target pole amplitude values and the multiple target pole angle values, determine a set of filter parameters corresponding to each of the multiple all-pass filter groups.
[0102] In this embodiment, determining the set of filter parameters corresponding to multiple all-pass filter banks can be achieved by directly selecting the APF filter bank with the best loudness enhancement effect through offline traversal search. Considering that several filter banks with similar or identical loudness enhancement effects may be found in practice, the top M filter banks with the best performance can be selected offline as candidate combinations for subsequent real-time filtering. Specifically, the traversal search process for obtaining multiple sets of filter parameters corresponding to multiple all-pass filter banks is as described above. Figure 2 The implementation process is as follows:
[0103] 1. Specify the search range of parameter r for each all-pass filter (APF) as [r1, r2] and the search step size as dr. Let there be k1 possible choices for r.
[0104] 2. Specify the search range of parameter w for each all-pass filter (APF) as [w1, w2] and the search step size as dw. Let there be k2 possible choices for w.
[0105] 3. Within the search range specified by r and w, continuously change the parameters r and w of each all-pass filter APF according to the search step size to generate N new all-pass filters;
[0106] 4. Connect N new all-pass filters in series to form a filter group, and repeat the above steps 1-3 until K filter groups are determined.
[0107] In this application embodiment, exemplarily, a filter bank is cascaded as described above. Figure 3 As shown, a filter bank can contain several all-pass filters.
[0108] It should be noted that, since changing the order of each sub-filter within the N filters does not change the final overall filtering effect when N filters are used in series, according to the permutation and combination formula, there are a total of K = (k1*k2+N-1)! / N! / k1*k2)! different ways to combine N filters in series.
[0109] It should be noted that the value of N can be selected according to the actual situation, and no specific limitation is made in this application. K is the number of filter banks obtained by adjusting the r and w parameters. These K combinations are the multiple all-pass filter banks determined.
[0110] It should be noted that in the embodiments of this application, multiple second-order IIR all-pass filters (APFs) are connected in series to fine-tune the group delay of the input signal. The APFs do not change the characteristics of the signal amplitude and frequency response, and can keep the amplitude of each single-frequency sinusoidal signal in the original signal unchanged. Moreover, the implementation is simple and the total computational load is very small even when multiple are connected in series.
[0111] It should be noted that when selecting multiple all-pass filter banks, the total number N of all-pass filters constituting the filter bank and the number M of candidate optimal filter combinations can be specified in advance.
[0112] It should be noted that the total number of all-pass filters N can be chosen by the sound engineer based on a trade-off between effect and computational load.
[0113] It should be noted that, as Figure 6 The diagram showing the correspondence between the sampling frequency and delay of the all-pass filter APF is as follows: Figure 6As shown, a single APF with r = 0.99 and w = 0.0131 will introduce a maximum delay of approximately 5.2 ms in the 100-200 Hz frequency band (250 samples at a 48 kHz sampling rate), while having virtually no effect on frequencies above 500 Hz. It is evident that the primary effective frequency band of a single APF is relatively narrow, and the resulting change in group delay is also small. Therefore, multiple APFs need to be cascaded as a filter bank to achieve multi-band delay, thereby significantly altering the mutual delay between individual frequencies in a multi-frequency music signal. Therefore, in this embodiment, the total number of APFs N in the filter bank needs to be specified in advance.
[0114] It should be noted that while the traversal search method can ensure the acquisition of M filter banks with optimal loudness enhancement for any loudspeaker, the computational load and time consumption of the traversal are significant, making it unsuitable for quickly and conveniently obtaining the optimal filter parameters in practical applications. Therefore, a better approach is to pre-search for loudness enhancement parameters using the aforementioned traversal method on a large number of different types of loudspeakers to obtain multiple optimal filter banks. Then, through human listening experiments, the family of group delay characteristic curves with optimal loudness enhancement can be statistically summarized. Subsequently, for any micro-loudspeaker, the best-performing group delay curve and the corresponding filter parameters in the corresponding filter bank can be selected from this family of curves.
[0115] S102. Assign multiple sets of filter parameters to multiple all-pass filter groups respectively, wherein one set of filter parameters corresponds to one all-pass filter group.
[0116] In this embodiment of the application, after determining multiple sets of filter parameters through the above implementation process, each set of filter parameters is assigned to each filter in a corresponding filter group, and the filter group formed by connecting each filter in series is used for filtering.
[0117] S103. Input the preset audio source signal into multiple all-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each all-pass filter group.
[0118] In this embodiment, either by directly performing a traversal search to obtain a set of filter parameters corresponding to each of the multiple all-pass filter groups, or by obtaining multiple target group delay response curves corresponding to each of the multiple all-pass filter groups through pre-statistically summarized candidate group delay curves, and after determining a set of filter parameters corresponding to each of the multiple all-pass filter groups based on the multiple target group delay response curves, the preset audio source signal is input into the multiple all-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each all-pass filter group.
[0119] It should be noted that each set of all-pass filters connected in series in each all-pass filter bank corresponds to a set of filter parameters.
[0120] In this embodiment, a preset audio source signal is input into each full-pass filter group, and the preset audio source signal is filtered using a set of filter parameters corresponding to each full-pass filter group to obtain the filtered audio source signal corresponding to each full-pass filter group.
[0121] In this embodiment of the application, in order to fully evaluate the loudness enhancement effect of the all-pass filter bank and select the optimal filter bank, it is necessary to select several different types of music as preset sound source signals. Each sound source signal is preprocessed by equalizer H1 to obtain a loudness evaluation sound source database S, and the number of sound sources in the database is denoted as P.
[0122] It should be noted that the loudness evaluation sound source database in this application embodiment contains more than 100 sound source signals, including transient short segments of various musical instruments such as bass, snare drum, drum kit, guitar, piano, etc., as well as common songs of various styles (pop, classical, rock, jazz, etc.), and several movie clips and game music clips.
[0123] In this embodiment, the equalization filter is selected based on different micro-speakers, such as... Figure 7 The figure shows the original amplitude frequency response curve of a micro loudspeaker. The horizontal axis represents frequency in Hz, and the vertical axis represents amplitude in dB. The selection of the equalization filter can be obtained based on the original amplitude frequency response curve of the micro loudspeaker.
[0124] It should be noted that the amplitude frequency response curve can be obtained by testing and solving using well-known methods such as sweep frequency signals or pseudo-random sequence MLS signals.
[0125] In the embodiments of this application, from Figure 7 It can be seen that its amplitude below 400Hz and above 10kHz is significantly lower than that of the mid-frequency band, and there is a peak in the frequency response near 5.5kHz, which is not completely flat. To address this, several high-pass filters with cutoff frequencies of 100-200Hz and low-pass filters with cutoff frequencies of 10-15kHz are used to attenuate low-frequency and extremely high-frequency signals that the speaker cannot reproduce; and peak filters and notch filters are used to compensate for local peaks and valleys on the frequency response curve, so that the frequency response curve in the mid-high frequency range is as flat as possible. The equalization filter group, denoted as H1, is composed of several filters consisting of high-pass filters, low-pass filters, peak filters, and peak filters.
[0126] It should be noted that the types of equalization filters are not limited to the four types mentioned above. Any form of digital filter can be used as an equalization filter. Specifically, the selection of equalization filter banks can be based on the original amplitude frequency response curve of the actual micro-speaker. This application does not make any specific limitations.
[0127] In this embodiment, the preset sound source signal is subjected to equalization filtering processing using the equalization filter group corresponding to the micro loudspeaker. The processed sound source signal is then input into each full-pass filter group. The filter parameters corresponding to each filter in each full-pass filter group are used to filter the sound source signal obtained after processing by the equalization filter, thereby obtaining the filtered sound source signal corresponding to each full-pass filter group.
[0128] In this embodiment, the K predetermined filter banks are used to filter all preset audio source signals S in the database to obtain the filtered signal S. ij , i represents the i-th filter combination, and j represents the j-th preset audio source signal in P.
[0129] For example, suppose there are 5 filter combinations and 3 songs in the database. Suppose the first song is input into the first filter combination, and the filtered signal can be represented by S. 11 This means that the first song is input into the second filter combination, and the filtered signal can be represented by S. 21 express.
[0130] It should be noted that in existing technologies, to improve the perceived loudness of audio playback, amplitude limiting and compression are first used to reduce the peak value of the digital signal's time-domain waveform, i.e., to reduce the peak-to-average power ratio (PAPR). Then, automatic gain compensation (AGC) is used to amplify the overall signal gain, thereby reducing the dynamic range and enhancing loudness. This method lossily discards some information from the peak segments of the original signal, thus damaging the timbre of the music signal. As is well known, instrument and vocal signals are composed of their fundamental frequency and multiple harmonics. The music played by consumers is composed of various instruments and vocals, indicating that the music signal is a random signal rich in multiple frequencies. Within a short time window (below 10ms to 20ms), the music signal can be considered as a steady-state signal composed of multiple sine waves of different frequencies. The peaks of the sine waves of each frequency superimpose to form the peak of the music signal, i.e., the amplitude maxima segment. To reduce the peak amplitude of a music signal, the aforementioned limiting and compression techniques can be used to directly reduce the amplitude of all frequency sine waves within the peak segment of the music signal. For each single-frequency sine wave, the compression of its peak waveform will make the sine wave no longer standard, which is equivalent to introducing harmonic distortion.
[0131] In this embodiment, the phase difference or group delay difference of different frequencies is adjusted to stagger the peak values of sine waves of different frequencies. The shape of each single-frequency sine wave will remain unchanged. Slight phase changes and delays have little effect on the human ear's perception of timbre, thereby achieving a near "lossless" reduction of the total peak value of the music signal. Combined with AGC to increase the total signal gain, loudness enhancement is achieved.
[0132] S104. Perform loudness prediction on the filtered audio source signal to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal; use the all-pass filter group with the largest loudness prediction value to adjust the loudness of the audio signal to be played.
[0133] In this embodiment of the application, the filtered sound source signal is input into the loudness prediction module to predict the loudness value of the filtered sound source signal.
[0134] In this embodiment of the application, when predicting the loudness value, the P filtered signals Si(S) corresponding to the i-th filter group are sequentially processed. i1 ~S ip The input is fed into the loudness prediction module and the sum of the loudness values of the P output signals, Li, i = 1 to K, is recorded.
[0135] For example, assuming there are 4 filter banks and 3 preset audio source signals, when predicting loudness values, the 3 preset audio source signals are sequentially input into the first filter bank to obtain 3 filtered signals. These 3 filtered signals are then input into the loudness prediction module to predict loudness values, and the sum of these 3 loudness values is calculated to obtain the loudness value corresponding to the first filter bank. This process is repeated to calculate the loudness values corresponding to the 2nd to 4th filter banks.
[0136] In this embodiment of the application, after calculating multiple loudness prediction values corresponding to multiple all-pass filters, the largest loudness prediction value can be selected from the multiple loudness prediction values corresponding to multiple all-pass filter groups, and the all-pass filter group corresponding to the largest loudness prediction value can be determined as the all-pass filter group for loudness adjustment of the audio signal to be played.
[0137] It should be noted that the number of all-pass filter banks can be predetermined. The number of all-pass filter banks can be selected according to the actual situation, and can be M (M>1). Specifically, this application does not limit the number.
[0138] It should be noted that when selecting the maximum loudness prediction value, the K loudness prediction values Li can be sorted from largest to smallest, and the filter bank corresponding to the one with the largest loudness or the top M combinations can be selected for subsequent real-time audio loudness enhancement processing.
[0139] In the embodiments of this application, the loudness prediction value of a corresponding filter group is determined by using the filtered sound source signal. The filtered sound source signal can be processed by dynamic range control (DRC), automatic gain compensation (AGC), and loudspeaker amplitude protection at least sequentially to obtain the signal to be evaluated corresponding to the filtered sound source signal; and the loudness prediction value of a corresponding filter group is determined based on the signal to be evaluated.
[0140] It should be noted that the processing of the filtered audio source signal is not limited to the three processing methods mentioned above: DRC processing, AGC processing, and speaker amplitude protection processing. Processing methods can be added or removed according to the actual situation. Specifically, the selection can be made according to the actual situation, and no specific limitation is made in this application.
[0141] In this embodiment, the loudness prediction module can consist of four sub-processes: Dynamic Range Control (DRC), Automatic Gain Compensation (AGC), Excursion Limit algorithm for loudspeaker amplitude protection, and loudness calculation. For example... Figure 8 As shown, in the loudness prediction module, the processing flow can be to process P input signals S i1 ~S ip After sequential processing by DRC, AGC, and loudspeaker protection algorithm, P signals S2 to be evaluated are obtained. The loudness value Li of S2 is calculated according to a certain loudness calculation criterion.
[0142] It should be noted that the loudness prediction module is not limited to the processing sub-processes mentioned above. Specifically, the sub-processes can be selected according to the actual situation, and no specific limitation is made in this application.
[0143] In this embodiment of the application, determining the loudness prediction value of a corresponding filter group based on the signal to be evaluated can be achieved by calculating the energy value corresponding to the signal to be evaluated; based on the energy value, the loudness prediction value of the corresponding filter group can be obtained at least by summation.
[0144] In this embodiment, the energy Ei of each signal S2 to be evaluated is calculated, and the sum of the energies of the P signals is used as the loudness value Li of this filter group.
[0145] It should be noted that the method of determining the loudness prediction value of the filter bank based on the energy value is not limited to the summation method. Other methods of determining the loudness prediction value of the filter bank based on the energy value are all within the scope of protection of this application. Specifically, the method can be selected according to the actual situation, and no specific limitation is made in this application.
[0146] In another alternative embodiment, determining the loudness prediction value of a corresponding filter group based on the signal to be evaluated can be achieved by inputting the signal to be evaluated into a preset psychological auditory loudness model to obtain a first loudness value; and based on the first loudness value, obtaining the loudness prediction value of the corresponding filter group at least by summation.
[0147] In the embodiments of this application, each signal S2 to be evaluated is substituted into the psychoacoustic loudness model of the ISO 532A or ISO 532B standard to obtain the average loudness, and then the average loudness of all the signals to be tested is summed as the loudness value Li of this filter group.
[0148] It should be noted that the method of determining the loudness prediction value of the filter bank based on the first loudness value is not limited to the summation method. Other methods of determining the loudness prediction value of the filter bank based on the first loudness value are all within the scope of protection of this application. Specifically, the method can be selected according to the actual situation, and no specific limitation is made in this application.
[0149] It should be noted that the method of calculating loudness value according to the loudness calculation criteria is not limited to the two calculation methods mentioned in this application. Specifically, the method can be selected according to the actual situation, and no specific limitation is made in this application.
[0150] In this embodiment, M optimal APF filter banks are obtained through offline search, and these M optimal filters are used for real-time audio playback in music, movie, and game scenarios on electronic devices such as mobile phones and tablets. In the real-time audio processing chain, the loudness enhancement filter bank (LEF) is placed after the speaker amplitude frequency response adjustment (EQ) equalizer and before the dynamic range control (DRC) pre-amplitude filter bank. Figure 9 As shown, the basic processing chain including the loudness enhancement module is as follows: input → equalization filter bank equalization processing → loudness enhancement filter bank processing → dynamic range control processing → automatic gain compensation processing → loudspeaker amplitude protection algorithm processing → output.
[0151] It should be noted that this application introduces a linear dynamic range adjustment method for an all-pass filter bank, which, combined with existing frequency band compression processing, DRC processing, and AGC processing, significantly improves the loudness of the audio reproduced by the micro-speaker while reducing the damage of harmonic distortion to the timbre.
[0152] In this embodiment of the application, when selecting a loudness enhancement filter bank, in order to reduce the computational load caused by LEF, the optimal first filter bank can be selected from the M optimal combinations obtained offline and substituted into the audio processing link for fixed loudness enhancement filtering processing.
[0153] Optionally, in this embodiment of the application, when selecting the loudness enhancement filter group, the preset sound source signal is input into multiple full-pass filter groups for filtering processing to obtain the filtered sound source signal corresponding to each full-pass filter group. Alternatively, multiple initial sound source signals can be preprocessed using a preset equalization filter to obtain multiple preset sound source signals; the preset sound source signals are divided into multiple sound source signals; each sound source signal is copied according to the number of groups in the multiple full-pass filter groups to obtain multiple sound source signals corresponding to each sound source signal; and the multiple sound source signals corresponding to each sound source signal are sequentially input into multiple full-pass filter groups for filtering processing to obtain the filtered sound source signal corresponding to each filter group.
[0154] In the embodiments of this application, such as Figure 10As shown, the input signal in the link can first be processed by framing, for example, with a frame duration of 10ms. The current frame signal is then output to the EQ module. The EQ-processed signal is then copied to M channels and output to M optimal filter combinations for filtering and enhancement. The M processed signals are then output to the loudness prediction module for loudness comparison. The signal with the best loudness is used as the output of the loudness enhancement module and transmitted to the DRC module, AGC module, and speaker amplitude protection algorithm module for subsequent processing.
[0155] It should be noted that since the optimal filtering parameters selected by the loudness enhancement module change dynamically in real time, the connection between the output signals of the current frame and the previous frame needs to be smoothed out to avoid loss of musical sound quality due to discontinuity between the signals of the previous and next frames.
[0156] It should be noted that linear dynamic range adjustment of audio signals based on all-pass filters can significantly reduce the problem of harmonic distortion introduced by existing DRC and other technologies while compressing peak values.
[0157] It should be noted that the selection of loudness enhancement filter banks is not limited to the two selection methods in this application. Specifically, the selection can be made according to the actual situation, and no specific limitation is made in this application.
[0158] It is understood that the audio loudness adjustment method provided in this application introduces an all-pass filter bank during the process of adjusting audio loudness. The input audio source signal is filtered according to multiple sets of filter parameters corresponding to multiple all-pass filter banks. Based on the characteristic that the all-pass filter can change the time delay, the group time delay corresponding to the audio source signal at different frequencies is adjusted to stagger the peak values of sine waves at different frequencies, so that the peak values of different frequencies at the same time are not superimposed, thereby maintaining the shape of each single-frequency sine wave without changing it, avoiding truncated distortion, and reducing the damage to the audio source signal caused by harmonic distortion. In order to ensure better loudness enhancement effect, the loudness value of the audio source signal after filtering by each all-pass filter bank is predicted, and the filter bank with the largest loudness enhancement value is further determined. Using this all-pass filter bank can not only reduce the impact of harmonic distortion on the timbre of the audio source, but also significantly improve the loudness of the audio reproduced by the micro-speaker.
[0159] Based on the above embodiments, this application provides an audio loudness adjustment method, such as... Figure 11 As shown, the specific steps include:
[0160] Step 1: Determine the time delay response curves of multiple target candidate groups from multiple all-pass filter banks;
[0161] Step 2: Obtain the group delay response parameter curve corresponding to the preset equalization filter;
[0162] Step 3: Using the group delay response curves corresponding to the preset equalization filter, process the multiple target candidate group delay response curves corresponding to multiple all-pass filter groups at least by subtraction to obtain multiple target group delay response curves corresponding to multiple all-pass filter groups.
[0163] Step 4: Based on the time delay response curves of multiple target groups, determine a set of filter parameters corresponding to each of the multiple all-pass filter banks;
[0164] Step 5: Input the preset audio source signal into each full-pass filter group, and use a set of filter parameters corresponding to the full-pass filter group to filter the preset audio source signal to obtain the filtered audio source signal corresponding to each full-pass filter group;
[0165] Step 6: Pass the filtered audio source signal through Dynamic Range Control (DRC), Automatic Gain Compensation (AGC), and Speaker Amplitude Protection processes at least sequentially to obtain the signal to be evaluated corresponding to the filtered audio source signal;
[0166] Step 7: Calculate the energy value corresponding to the signal to be evaluated; use the energy value to obtain the loudness prediction value of the corresponding filter bank;
[0167] Step 8: Input the signal to be evaluated into the preset psychological auditory loudness model to obtain the first loudness value; use the first loudness value to obtain the loudness prediction value of the corresponding filter group;
[0168] Step 9: Select the group of all-pass filters with the largest loudness prediction value from multiple all-pass filter groups; use this group of all-pass filters to adjust the loudness of the audio signal to be played.
[0169] It should be noted that in this application, the input audio signal is first subjected to attenuation filtering in the low-frequency and ultra-high-frequency bands, then the signal is linearly dynamically adjusted based on the all-pass filter bank, then a small amount of dynamic range control (DRC) processing is applied, and finally automatic gain compensation (AGC) processing is applied. The loudness is improved more significantly, and the harmonic distortion caused by DRC is reduced.
[0170] Based on the above embodiments, another embodiment of this application provides an audio loudness adjustment device 1, such as... Figure 12 As shown, the audio loudness adjustment device includes multiple full-pass filter groups, and each full-pass filter group has a set of full-pass filters connected in series; the audio loudness adjustment device includes:
[0171] The acquisition module 10 is used to acquire multiple target candidate group delay response curves and determine multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves.
[0172] The assignment module 11 is used to assign multiple sets of filter parameters to multiple all-pass filter groups respectively, wherein one set of filter parameters corresponds to one all-pass filter group.
[0173] The filtering module 12 is used to input the preset audio source signal into multiple all-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each all-pass filter group.
[0174] The determination module 13 is used to predict the loudness of the filtered sound source signal and determine the predicted loudness value of the all-pass filter bank corresponding to the filtered sound source signal.
[0175] The adjustment module 14 is used to adjust the loudness of the audio signal to be played by using a set of all-pass filters with the largest loudness prediction value.
[0176] Optionally, the audio loudness adjustment device may further include: a processing module;
[0177] The processing module is used to process the filtered audio source signal through DRC processing, AGC processing and speaker amplitude protection processing at least sequentially to obtain the signal to be evaluated corresponding to the filtered audio source signal.
[0178] Optionally, the determining module 13 is further configured to determine the loudness prediction value of a corresponding filter bank based on the signal to be evaluated.
[0179] Optionally, the audio loudness adjustment device 1 may further include: a calculation module;
[0180] The calculation module is used to calculate the energy value corresponding to the signal to be evaluated; based on the energy value, the loudness prediction value of the corresponding filter group is obtained by at least summing.
[0181] Optionally, the audio loudness adjustment device 1 may further include: an input module;
[0182] The input module is used to input the signal to be evaluated into a preset psychological auditory loudness model to obtain a first loudness value; based on the first loudness value, at least the loudness prediction value of a corresponding filter group is obtained by summation.
[0183] Optionally, the audio loudness adjustment device 1 may further include: a preprocessing module;
[0184] The preprocessing module is used to preprocess multiple initial audio source signals using a preset equalization filter to obtain multiple preset audio source signals.
[0185] Optionally, the audio loudness adjustment device 1 may further include: a division module;
[0186] The segmentation module is used to divide the preset audio source signal into multiple audio source signals.
[0187] Optionally, the audio loudness adjustment device 1 may further include: a copying module;
[0188] The copying module is used to copy each audio source signal according to the number of groups of multiple all-pass filter banks, so as to obtain multiple audio source signals corresponding to each audio source signal.
[0189] Optionally, the filtering module 12 is also used to sequentially input multiple audio source signals corresponding to each audio source signal into multiple all-pass filter groups for filtering processing, so as to obtain the filtered audio source signal corresponding to each filter group.
[0190] Optionally, the determining module 13 is further configured to, within a first preset search range, adjust the pole amplitude values corresponding to each full-pass filter in each full-pass filter bank according to a first preset search step size, and determine multiple target pole amplitude values corresponding to each full-pass filter.
[0191] Optionally, the determining module 13 is further configured to, within a second preset search range, adjust the pole angle value corresponding to each all-pass filter according to the second preset search step size, and determine multiple target pole angle values corresponding to each all-pass filter.
[0192] Optionally, the determining module 13 is further configured to determine multiple sets of initial filter parameters corresponding to multiple all-pass filter banks based on multiple target pole amplitude values and multiple target pole angle values; and to determine multiple initial candidate group delay response curves based on the multiple sets of initial filter parameters.
[0193] Optionally, the determining module 13 is further configured to perform statistical analysis on multiple initial candidate group delay response curves based on a preset algorithm to determine the group delay response curve cluster; based on the group delay response curve cluster and through manual listening recognition, determine multiple target candidate group delay response curves corresponding to multiple all-pass filter banks from the group delay response curve cluster.
[0194] Optionally, the processing module is further configured to use the group delay response curves corresponding to the preset equalization filter to process the multiple target candidate group delay response curves corresponding to the multiple all-pass filter groups at least by subtraction, so as to obtain the multiple target group delay response curves corresponding to the multiple all-pass filter groups.
[0195] Optionally, the determining module 13 is also used to determine multiple target phase frequency response parameters based on multiple target group time delay response curves.
[0196] Optionally, the determining module 13 is also used to determine multiple target transfer functions using multiple target phase frequency response parameters.
[0197] Optionally, the determining module 13 is also used to determine a set of filter parameters corresponding to the delay response curves of multiple target candidate groups using multiple target transfer functions.
[0198] Optionally, the determining module 13 is also used to determine multiple transfer functions of multiple all-pass filter banks using multiple sets of initial filter parameters.
[0199] Optionally, the determining module 13 is also used to determine multiple phase frequency response parameters of multiple all-pass filter banks using multiple transfer functions.
[0200] Optionally, the determining module 13 is also used to determine multiple initial candidate group delay response curves using multiple phase frequency response parameters.
[0201] This application provides an audio loudness adjustment device that acquires multiple target candidate group delay response curves and determines multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves; assigns the multiple sets of filter parameters to multiple all-pass filter groups, wherein one set of filter parameters corresponds to one all-pass filter group; inputs a preset audio source signal into the multiple all-pass filter groups for filtering processing to obtain a filtered audio source signal corresponding to each all-pass filter group; performs loudness prediction on the filtered audio source signal to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal; and uses the all-pass filter group with the largest loudness prediction value to adjust the loudness of the audio signal to be played. Therefore, the audio loudness adjustment device proposed in this application introduces an all-pass filter bank during the process of adjusting audio loudness. Based on the multiple sets of filter parameters corresponding to the multiple all-pass filter banks, the input audio source signal is filtered. Leveraging the characteristic of all-pass filters to change time delay, the group delay of the audio source signal at different frequencies is adjusted to avoid the peak values of different frequency sine waves. This prevents the peak values of different frequencies from being superimposed at the same time, thus maintaining the shape of each single-frequency sine wave and avoiding truncated distortion. This reduces the damage to the audio source signal caused by harmonic distortion. Furthermore, to ensure a better loudness enhancement effect, the loudness value of the audio source signal filtered by each all-pass filter bank is predicted, further determining the filter bank with the largest loudness enhancement value. Using this all-pass filter bank not only reduces the impact of harmonic distortion on the timbre of the audio source but also significantly improves the loudness of the audio reproduced by the micro-speaker.
[0202] Figure 13 This is a schematic diagram of the composition structure of a terminal device 2 provided in an embodiment of this application. In practical applications, based on the same disclosed concept of the above embodiments, such as... Figure 13 As shown, the terminal device 2 in this embodiment includes a processor 20, a memory 21, and a communication bus 22.
[0203] In specific embodiments, the aforementioned acquisition module 10, assignment module 11, filtering module 12, determination module 13, adjustment module 14, processing module, calculation module, input module, preprocessing module, partitioning module, and copying module can be implemented by the processor 20 located on the terminal device 2. The processor 20 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), CPU, controller, microcontroller, and microprocessor. It is understood that for different devices, the electronic devices used to implement the above processor functions can also be other types; this embodiment does not impose specific limitations.
[0204] In this embodiment, the communication bus 22 is used to realize the connection communication between the processor 20 and the memory 21; when the processor 20 executes the running program stored in the memory 21, it implements the following audio loudness adjustment method:
[0205] Obtain multiple delay response curves of target candidate groups and determine multiple sets of filter parameters corresponding to these curves. Assign these filter parameters to multiple all-pass filter groups, with each set of filter parameters corresponding to one all-pass filter group. Input the preset audio source signal into the multiple all-pass filter groups for filtering to obtain the filtered audio source signal corresponding to each all-pass filter group. Perform loudness prediction on the filtered audio source signal to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal. Use the all-pass filter group with the largest loudness prediction value to adjust the loudness of the audio signal to be played.
[0206] Furthermore, the processor 20 is also used to process the filtered audio source signal through dynamic range control (DRC), automatic gain compensation (AGC), and loudspeaker amplitude protection processes at least sequentially to obtain the signal to be evaluated corresponding to the filtered audio source signal; and to determine the loudness prediction value of a corresponding filter group based on the signal to be evaluated.
[0207] Furthermore, the processor 20 is also used to calculate the energy value corresponding to the signal to be evaluated; based on the energy value, at least the loudness prediction value of a corresponding filter bank is obtained by summation.
[0208] Furthermore, the processor 20 is also used to input the signal to be evaluated into a preset psychological auditory loudness model to obtain a first loudness value; based on the first loudness value, at least by summing, the loudness prediction value of a corresponding filter group is obtained.
[0209] Furthermore, the processor 20 is also used to preprocess multiple initial sound source signals using a preset equalization filter to obtain multiple preset sound source signals; divide the preset sound source signals into multiple sound source signals; copy each sound source signal according to the number of groups in the multiple all-pass filter groups to obtain multiple sound source signals corresponding to each sound source signal; sequentially input the multiple sound source signals corresponding to each sound source signal into multiple all-pass filter groups, and perform filtering processing using a set of filter parameters corresponding to each of the multiple all-pass filter groups to obtain the filtered sound source signal corresponding to each filter group.
[0210] Furthermore, the processor 20 is also configured to, within a first preset search range, adjust the pole amplitude value corresponding to each full-pass filter in each full-pass filter bank according to a first preset search step size, to determine multiple target pole amplitude values corresponding to each full-pass filter; within a second preset search range, adjust the pole angle value corresponding to each full-pass filter according to a second preset search step size, to determine multiple target pole angle values corresponding to each full-pass filter; determine multiple sets of initial filter parameters corresponding to multiple full-pass filter banks based on the multiple target pole amplitude values and the multiple target pole angle values; and determine multiple initial candidate group delay response curves based on the multiple sets of initial filter parameters.
[0211] Furthermore, the processor 20 is also used to perform statistical analysis on multiple initial candidate group delay response curves based on a preset algorithm to determine the group delay response curve cluster; based on the group delay response curve cluster and through manual listening recognition, to determine multiple target candidate group delay response curves corresponding to multiple all-pass filter banks from the group delay response curve cluster.
[0212] Furthermore, the processor 20 is also used to process multiple target candidate group delay response curves corresponding to multiple all-pass filter banks by at least subtraction using the group delay response curves corresponding to the preset equalization filter banks, to obtain multiple target group delay response curves corresponding to multiple all-pass filter banks; to determine multiple target phase frequency response parameters based on the multiple target group delay response curves; to determine multiple target transfer functions using the multiple target phase frequency response parameters; and to determine a set of filter parameters corresponding to the multiple target candidate group delay response curves using the multiple target transfer functions.
[0213] Furthermore, the processor 20 is also used to determine multiple transfer functions of multiple all-pass filter banks using multiple sets of initial filter parameters; to determine multiple phase frequency response parameters of multiple all-pass filter banks using multiple transfer functions; and to determine multiple initial candidate group delay response curves using multiple phase frequency response parameters.
[0214] Based on the above embodiments, this application provides a storage medium storing a computer program thereon. The computer-readable storage medium stores one or more programs, which can be executed by one or more processors and applied in a terminal device. The computer program implements the data processing method described above.
[0215] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0216] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause an image display device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0217] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for adjusting audio loudness, characterized in that, An audio loudness adjustment device, comprising multiple full-pass filter banks, each full-pass filter bank containing a set of full-pass filters connected in series, the method comprising: Multiple target candidate group delay response curves are obtained, and multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves are determined based on the multiple target candidate group delay response curves; wherein, the multiple target candidate group delay response curves are selected from multiple initial candidate group delay response curves determined using the obtained multiple sets of initial filter parameters, and the target candidate group delay response curves are the multiple initial candidate group delay response curves with the best loudness enhancement effect; The multiple sets of filter parameters are respectively assigned to the multiple all-pass filter groups, wherein one set of filter parameters corresponds to one all-pass filter group; The preset audio source signal is input into the multiple all-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each all-pass filter group; The loudness prediction of the filtered audio source signal is performed to determine the loudness prediction value of the all-pass filter group corresponding to the filtered audio source signal; the loudness of the audio signal to be played is adjusted using the all-pass filter group with the largest loudness prediction value. The step of determining multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves based on the multiple target candidate group delay response curves includes: Using the group delay response curve corresponding to the preset equalization filter, the multiple target candidate group delay response curves corresponding to multiple all-pass filter groups are processed by at least the subtraction method to obtain the multiple target group delay response curves corresponding to the multiple all-pass filter groups. Based on the time delay response curves of the multiple target groups, multiple target phase frequency response parameters are determined; Using the multiple target phase frequency response parameters, the transfer functions of multiple targets are determined; Using the multiple target transfer functions, a set of filter parameters corresponding to the delay response curves of the multiple target candidate groups are determined.
2. The method according to claim 1, characterized in that, The step of performing loudness prediction on the filtered sound source signal to determine the loudness prediction value of the all-pass filter bank corresponding to the filtered sound source signal includes: The filtered audio source signal is processed at least sequentially through Dynamic Range Control (DRC), Automatic Gain Compensation (AGC), and Speaker Amplitude Protection to obtain the signal to be evaluated corresponding to the filtered audio source signal. Based on the signal to be evaluated, determine the loudness prediction value of a corresponding all-pass filter bank.
3. The method according to claim 2, characterized in that, The step of determining the loudness prediction value of a corresponding all-pass filter bank based on the signal to be evaluated includes: Calculate the energy value corresponding to the signal to be evaluated; Based on the energy value, the loudness prediction value of the corresponding all-pass filter bank can be obtained by at least summing.
4. The method according to claim 2, characterized in that, The step of determining the loudness prediction value of a corresponding all-pass filter bank based on the signal to be evaluated includes: The signal to be evaluated is input into a preset psychological auditory loudness model to obtain a first loudness value; Based on the first loudness value, at least one loudness prediction value for an all-pass filter bank can be obtained by summation.
5. The method according to claim 1, characterized in that, The step of inputting the preset audio source signal into the plurality of full-pass filter groups for filtering processing to obtain the filtered audio source signal corresponding to each full-pass filter group includes: Multiple initial sound source signals are preprocessed using a preset equalization filter to obtain multiple preset sound source signals; The preset sound source signal is divided into multiple sound source signals; According to the number of groups of the multiple all-pass filter groups, each audio source signal is copied to obtain multiple audio source signals corresponding to each audio source signal; Each segment of the audio source signal is sequentially input into the multiple all-pass filter groups for filtering processing, thereby obtaining the filtered audio source signal corresponding to each filter group.
6. The method according to claim 1, characterized in that, Before obtaining the delay response curves of the multiple target candidate groups, the method further includes: Within a first preset search range, based on a first preset search step size, the pole amplitude values corresponding to each full-pass filter in each full-pass filter bank are adjusted to determine multiple target pole amplitude values corresponding to each full-pass filter. Within the second preset search range, the pole angle values corresponding to each all-pass filter are adjusted according to the second preset search step size to determine multiple target pole angle values corresponding to each all-pass filter. Based on the magnitude values and angle values of the multiple target poles, determine multiple sets of initial filter parameters corresponding to multiple all-pass filter banks; Based on the multiple sets of initial filter parameters, multiple initial candidate group delay response curves are determined.
7. The method according to claim 6, characterized in that, After determining multiple initial candidate group delay response curves, the method further includes: Based on a preset algorithm, statistical analysis is performed on the multiple initial candidate group delay response curves to determine the outgoing delay response curve cluster. Based on the cluster of group delay response curves and through manual auditory recognition, multiple target candidate group delay response curves corresponding to multiple all-pass filter banks are determined from the cluster of group delay response curves.
8. The method according to claim 6, characterized in that, The step of determining multiple initial candidate group delay response curves based on the multiple sets of initial filter parameters includes: Using the multiple sets of initial filter parameters, multiple transfer functions of the multiple all-pass filter banks are determined; Using the multiple transfer functions, multiple phase frequency response parameters of the multiple all-pass filter banks are determined; Using the aforementioned multiple phase frequency response parameters, multiple initial candidate group delay response curves are determined.
9. An audio loudness adjustment device, characterized in that, The audio loudness adjustment device includes multiple full-pass filter groups, each full-pass filter group having a set of full-pass filters connected in series. The audio loudness adjustment device includes: The acquisition module is used to acquire multiple target candidate group delay response curves and determine multiple sets of filter parameters corresponding to the multiple target candidate group delay response curves based on the multiple target candidate group delay response curves; wherein, the multiple target candidate group delay response curves are selected from multiple initial candidate group delay response curves determined using the acquired multiple sets of initial filter parameters, and the target candidate group delay response curves are the multiple initial candidate group delay response curves with the best loudness enhancement effect; The assignment module is used to assign the multiple sets of filter parameters to the multiple all-pass filter groups respectively, wherein one set of filter parameters corresponds to one all-pass filter group; The filtering module is used to input the preset audio source signal into the plurality of all-pass filter groups for filtering processing, so as to obtain the filtered audio source signal corresponding to each all-pass filter group. The determination module is used to predict the loudness of the filtered sound source signal and determine the predicted loudness value of the all-pass filter group corresponding to the filtered sound source signal. The adjustment module is used to adjust the loudness of the audio signal to be played by using a set of all-pass filters with the largest loudness prediction values. The processing module is used to process multiple target candidate group delay response curves corresponding to multiple all-pass filter groups by at least subtraction using the group delay response curves corresponding to the preset equalization filter groups, so as to obtain multiple target group delay response curves corresponding to the multiple all-pass filter groups. The determination module is also used to determine multiple target phase frequency response parameters based on the multiple target group time delay response curves; The determining module is also used to determine the transfer functions of multiple targets using the multiple target phase frequency response parameters; The determination module is also used to determine a set of filter parameters corresponding to the delay response curves of the multiple target candidate groups using the multiple target transfer functions.
10. A terminal device, characterized in that, The terminal device includes: a processor, a memory, and a communication bus; when the processor executes the running program stored in the memory, it implements the method as described in any one of claims 1 to 8.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.