Audio equalization method and device, storage medium and program product

By analyzing the spectral differences between the target audio and the reference audio, compression parameters were determined and EQ curves were optimized, thus solving the problem of audio signal distortion and achieving more efficient and smoother audio equalization processing.

CN120853591APending Publication Date: 2025-10-28TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511170459.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies directly determine the EQ curve gain by the ratio of the reference spectrum to the target audio spectrum when equalizing audio, which may lead to distortion of the equalized audio signal.

Method used

By analyzing the spectral differences between the target audio and the reference audio, the statistical characteristics of the spectral differences are determined. Based on these characteristics, compression parameters are determined, and the soft-knee compression algorithm is used to optimize the EQ curve, generating a smooth and optimized EQ curve.

Benefits of technology

It reduces the possibility of spectral distortion in the equalized audio signal, improves the smoothness and spectral fidelity of the EQ curve, and enhances processing efficiency and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853591A_ABST
    Figure CN120853591A_ABST
Patent Text Reader

Abstract

The invention discloses an audio equalization method and device, a storage medium and a program product, and belongs to the technical field of audio. The method comprises the steps of obtaining a frequency spectrum of a target audio and a frequency spectrum of a reference audio, generating an EQ curve representing a frequency spectrum difference between the target audio and the reference audio, determining a statistical characteristic of the frequency spectrum difference, determining a compression parameter for performing range compression on the frequency spectrum difference based on the statistical characteristic, and performing range compression on the frequency spectrum difference based on the compression parameter and a soft knee compression algorithm. And performing optimization processing on the EQ curve to obtain an optimized EQ curve, and performing equalization processing on the audio signal of the target audio based on the optimized EQ curve to obtain an equalized audio signal of the target audio. Through soft knee compression processing, the spectrum distortion possibility of the balanced audio signal is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio, and in particular to a method, apparatus, storage medium, and program product for equalizing audio. Background Technology

[0002] In the field of audio technology, equalization processing is often required to optimize the overall balance and expressiveness of the sound by adjusting the gain of different frequencies.

[0003] Currently, when performing equalization processing on the target audio, for each frequency point, the ratio of the amplitude of that frequency point in the reference spectrum to the amplitude of that frequency point in the target audio spectrum is determined as the gain of that frequency point on the EQ (Equalization) curve. Then, based on this EQ curve, the coefficients of the FIR (Finite Impulse Response) filter are obtained. Finally, the FIR coefficients are used to process the target audio signal to obtain the equalized audio signal.

[0004] When generating an EQ curve, directly determining the gain of the frequency point in the EQ curve as the ratio of the amplitude of the reference frequency point to the amplitude of the target audio frequency point in the spectrum may cause distortion of the equalized audio signal. Summary of the Invention

[0005] This application provides a method, apparatus, storage medium, and program product for equalizing audio, which can reduce the possibility of distortion in the equalized audio signal. The technical solution is as follows:

[0006] Firstly, a method for equalizing audio is provided, the method comprising:

[0007] Obtain the EQ curve representing the spectral difference between the target audio and the reference audio;

[0008] Based on the spectral differences, determine the statistical characteristics of the spectral differences;

[0009] Based on the statistical characteristics, compression parameters for range compression of the spectral differences are determined;

[0010] Based on the compression parameters and soft knee compression algorithm, the EQ curve is optimized to obtain an optimized EQ curve;

[0011] Based on the optimized EQ curve, the audio signal of the target audio is subjected to equalization processing to obtain the equalized audio signal of the target audio.

[0012] In one alternative approach, the statistical features include mean, standard deviation, and kurtosis, and determining compression parameters for range compression of the spectral differences based on the statistical features includes:

[0013] Based on the mean and the standard deviation, determine the target range value of the compression parameters;

[0014] Based on the standard deviation, the compression threshold and soft knee width in the compression parameters are determined;

[0015] Based on the kurtosis, the compression ratio in the compression parameters is determined.

[0016] In this way, the statistical characteristics of the audio are used to determine the compression parameters, so that the parameters used for compression match the target audio.

[0017] In an alternative approach, the statistical feature further includes spectral flatness; the method further includes:

[0018] The target range value is adjusted based on the spectral flatness.

[0019] Thus, spectral flatness reflects the characteristics of spectral energy distribution. Therefore, spectral flatness is used to fine-tune the target range value to constrain the boundary of gain in the optimized EQ curve, so as to make the compression as smooth as possible and the range controllable.

[0020] In an alternative approach, the method further includes:

[0021] For the k-th frequency point on the EQ curve, based on the gain from the kP-th to the (k+P)-th frequency points on the EQ curve, the local energy value and global energy estimate corresponding to the k-th frequency point are determined. Based on the local energy value and the global energy estimate, the compression threshold is adjusted to obtain the compression threshold of the k-th frequency point, where k is an integer greater than or equal to 1 and P is an integer greater than 1.

[0022] In this way, adjusting the compression threshold based on local energy for each frequency point can enhance the adaptability to local spectral features.

[0023] In one alternative approach, the compression parameters include a compression threshold, soft knee width, compression ratio, and target range value;

[0024] The optimization process based on the compression parameters and the soft-knee compression algorithm to obtain the optimized EQ curve includes:

[0025] For each frequency point on the EQ curve, if the absolute value of the gain of the frequency point is less than or equal to a first value, and the absolute value is less than or equal to the target range value, then the gain of the frequency point is determined as the gain of the frequency point on the optimized EQ curve, wherein the first value is equal to the difference between the compression threshold and half of the soft knee width.

[0026] If the absolute value of the gain at the frequency point is greater than the first value and less than or equal to the second value, then the gain at the frequency point on the EQ curve is added to the third value to obtain a fourth value. If the absolute value of the fourth value is less than or equal to the target range value, the fourth value is determined as the gain at the frequency point on the optimized EQ curve. The second value is equal to the sum of the compression threshold and half of the soft knee width, and the third value is equal to the square of the difference between the gain at the frequency point on the EQ curve and the first value divided by twice the soft knee width.

[0027] If the absolute value of the gain at the frequency point is greater than the second value, then the second value and the fifth value are added to obtain a sixth value. If the absolute value of the sixth value is less than or equal to the target range value, the sixth value is determined as the gain of the frequency point on the optimized EQ curve. The fifth value is equal to the product of the reciprocal of the compression ratio and the difference between the gain of the frequency point on the EQ curve and the second value.

[0028] In this way, processing the EQ curve in segments can smoothly compress the EQ curve.

[0029] In an alternative approach, before performing equalization processing on the audio signal of the target audio based on the optimized EQ curve, the method further includes:

[0030] Set the gain of frequencies below the cutoff frequency on the optimized EQ curve to 0.

[0031] This helps to suppress excessive adjustment in the low-frequency range and avoid muddy sound quality.

[0032] In one alternative approach, obtaining the EQ curve characterizing the spectral difference between the target audio and the reference audio includes:

[0033] Determine the first average spectrum of the M frames with the highest energy in the target audio, and determine the second average spectrum of the M frames with the highest energy in the reference audio, where M is an integer greater than 1;

[0034] Generate an EQ curve that characterizes the difference between the first average spectrum and the second average spectrum.

[0035] Thus, high energy generally indicates human voice, and retaining only the human voice portion can reduce the impact of random noise.

[0036] In an alternative approach, before performing equalization processing on the audio signal of the target audio based on the optimized EQ curve, the method further includes:

[0037] Display the optimized EQ curve;

[0038] Receive user instructions to adjust the gain of the target frequency band in the optimized EQ curve;

[0039] Based on the adjustment command, the gain of the target frequency band in the optimized EQ curve is adjusted.

[0040] This provides users with a way to adjust the gain, resulting in better audio quality after equalization.

[0041] In a second aspect, an apparatus for equalizing audio is provided, the apparatus comprising one or more modules for implementing the method described in the first aspect or any alternative manner of the first aspect.

[0042] Thirdly, this application provides a computer device including a processor and a memory, the memory storing at least one instruction that is loaded and executed by the processor to perform the operation of the equalization audio method as described in the first aspect or any alternative method of the first aspect.

[0043] Fourthly, this application provides a computer-readable storage medium storing at least one instruction that is loaded and executed by a processor to perform the operations of the equalization audio method as described in the first aspect or any alternative method of the first aspect.

[0044] Fifthly, this application provides a computer program product storing at least one instruction that is loaded and executed by the processor to perform the operation of the equalizing audio method as described in the first aspect or any alternative method of the first aspect.

[0045] The beneficial effects of the technical solutions provided in this application are:

[0046] By analyzing the spectral differences between the target audio and the reference audio, the statistical characteristics of these differences are determined. Based on these statistical characteristics, compression parameters are determined to match the target audio. Then, the compression parameters and a soft-knee compression algorithm are used to optimize the EQ curve, thereby improving the smoothness of the optimized EQ curve. Therefore, when using the optimized EQ curve to obtain the equalized audio signal of the target audio, the possibility of spectral distortion of the equalized audio signal can be reduced. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart of a method for equalizing audio provided in an embodiment of this application;

[0049] Figure 2 This is a flowchart of another method for equalizing audio provided in an embodiment of this application;

[0050] Figure 3 This is a schematic diagram of the equalized audio structure provided in an embodiment of this application;

[0051] Figure 4 This is a schematic diagram of the terminal structure provided in the embodiments of this application;

[0052] Figure 5 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0054] In the field of audio technology, equalization processing is often required to optimize the overall balance and expressiveness of the sound by adjusting the gain or attenuation of different frequencies.

[0055] Currently, in one equalization method, the gain of each frequency band is adjusted manually based on experience, which is inefficient. In another equalization method, for each frequency point, the ratio of the amplitude of that frequency point in the reference spectrum to the amplitude of that frequency point in the target audio spectrum is determined as the gain of that frequency point in the EQ curve. Then, the FIR coefficients are obtained based on this EQ curve. Finally, the FIR coefficients are used to process the target audio signal to obtain the equalized audio signal. However, directly determining the gain of the frequency points in the EQ curve using the aforementioned ratio during EQ curve generation may distort the equalized audio signal.

[0056] Based on this, embodiments of this application provide a method for equalizing audio. In this method, the spectral difference between the target audio and the reference audio is automatically analyzed. Based on the statistical characteristics of the spectral difference, compression parameters for range compression of the spectral difference are determined. Then, a soft-knee compression algorithm is used to smoothly adjust the spectral difference to optimize the EQ curve, thereby reducing the possibility of frequency band gain abrupt changes and thus reducing the possibility of spectral distortion of the equalized audio signal.

[0057] In this embodiment, the entity executing the audio equalization method is an audio equalization device. Optionally, the device is a hardware device, such as a computer device, which is a terminal or server, including but not limited to mobile phones, laptops, or desktop computers. Optionally, the device is a software device, such as a software program installed on a computer device.

[0058] In the embodiments of this application, the audio equalization method can be applied to the timbre matching process in music production, or to the speech processing process, such as environment matching in speech processing, or to the audio editing process, such as batch equalization processing of audio. These are merely a few possible examples, and the embodiments of this application do not limit the scenarios for audio equalization.

[0059] The following explanation uses the terminal as the execution subject as an example to illustrate the method of equalizing audio. (See...) Figure 1 Steps S101 to S105.

[0060] Step S101: Obtain the EQ curve representing the spectral difference between the target audio and the reference audio.

[0061] In this context, the target audio is any audio file to be equalized. The reference audio is the audio file used as a reference during the equalization process of the target audio. For example, in timbre matching, the reference audio is an audio file with good timbre quality.

[0062] In this embodiment, the terminal stores the spectrum of the target audio and the spectrum of the reference audio, and can directly obtain the stored spectrum of the target audio and the spectrum of the reference audio. Alternatively, the terminal performs frame-by-frame windowing on the audio signal of the target audio. The windowed audio signal is divided into frames. The windows used in the windowing process include, but are not limited to, Hanning windows and Hamming windows. The Hanning window will be used as an example in the following description. The terminal performs FFT (Fast Fourier Transform) processing on the frame-by-frame windowed audio signal to obtain the frequency domain amplitude, as shown in formula (1).

[0063] X(k)=FFT[x(n)·w(n)] (1)

[0064] Where X(k) is the frequency domain amplitude of the k-th frequency point, x(n) represents the discrete-time signal sequence of the target audio, n is the index of the sampling point, usually ranging from n = 0, 1, ..., N-1, N is the number of sampling points per frame after the target audio is divided into frames, and w(n) is the window function, taking the Hanning window as an example here. x(n)·w(n) represents multiplying each sample point in the current frame by the window function value.

[0065] Then, the average amplitude of the frequency domain at each frequency point is calculated to obtain the average amplitude at each frequency point, thus obtaining the spectrum of the target audio. Alternatively, the spectrum of the reference audio can be determined by frame-by-frame windowing followed by FFT processing.

[0066] For each frequency point, the terminal calculates the ratio of the amplitude in the spectrum of the reference audio to the amplitude in the spectrum of the target audio, thus obtaining the spectral difference for each frequency point.

[0067] Alternatively, the terminal can obtain an EQ curve representing the spectral difference between the target audio and the reference audio from other devices.

[0068] It should be noted that in the embodiments of this application, the frame length and frame shift are not limited when framing. The frame length is the length of a single frame, and the frame shift refers to the time interval or sampling point interval between the start points of two adjacent frames.

[0069] In one alternative approach, high energy generally represents human voice, and retaining only the human voice portion can reduce the impact of random noise. Therefore, to reduce the impact of random noise, the average spectrum of high-energy frames is calculated, and the difference is calculated using the average spectrum. The terminal sorts the energy of each frame in the target audio signal from highest to lowest, obtains the top M frames, and determines the frequency domain amplitude of the M frames after FFT processing. The value of M is set based on the length of the target audio and empirical values. For each frequency point, the frequency domain amplitudes corresponding to the M frames are averaged to obtain the average amplitude of each frequency point, thus obtaining the average spectrum of the M frames, called the first average spectrum, denoted as . This is the average frequency domain amplitude of the k-th frequency point in the first average spectrum. Following this method, the average spectrum of the M-th frame with the highest energy in the reference audio is determined, called the second average spectrum, denoted as... It is the average frequency domain amplitude of the k-th frequency point in the second average spectrum.

[0070] For each frequency point, the ratio of the amplitude in the second average spectrum to the amplitude in the first average spectrum is determined, and the difference between the spectrum of the target audio and the spectrum of the reference audio is obtained, expressed as formula (2).

[0071]

[0072] In formula (2), M(k) is the difference at the k-th frequency point on the EQ curve, also known as the gain at the k-th frequency point.

[0073] Step S102: Based on the spectral difference, determine the statistical characteristics of the spectral difference.

[0074] In this embodiment, the terminal determines the statistical characteristics of the spectral difference. These statistical characteristics include, but are not limited to, mean, standard deviation, kurtosis, and SF (spectral flatness). Kurtosis is a statistical measure used to describe the shape of the energy distribution in the signal spectrum, and spectral flatness is used to quantify the flatness of the energy distribution in the frequency domain. The mean, standard deviation, kurtosis, and spectral flatness correspond to formulas (3) to (6), respectively.

[0075] Mean:

[0076] Standard deviation:

[0077] Kuroshi:

[0078] Spectral flatness:

[0079] In formulas (3) to (6), K is the number of frequency points on the EQ curve.

[0080] Step S103: Based on the statistical characteristics, determine the compression parameters used to compress the range of spectral differences.

[0081] In this embodiment, the compression parameter is a range compression parameter used to compress spectral differences, and it is the compression parameter required for soft-knee compression processing. The terminal uses statistical features to calculate the compression parameter required for soft-knee compression.

[0082] In one alternative approach, the compression parameters required for soft knee compression include, but are not limited to, target range values, compression threshold, soft knee width, and compression ratio.

[0083] The target range value is used to provide a safety boundary for the compression result. It can also be understood as indicating the degree of compression and defining a reasonable upper limit for gain. The target range value is denoted as T, and the formula is: T = 0.8σ + 0.2|μ|.

[0084] The compression threshold is the critical point that triggers compression. Compression is performed after the compression threshold is exceeded. It is denoted as K1 and the formula is: K1 = 1.2σ.

[0085] The soft knee width is a core parameter of dynamic range compression, referring to the width of the gradual transition region where the compression ratio smoothly changes near the compression threshold. The soft knee width is denoted as W, and the formula is: W = 0.3σ.

[0086] Compression ratio is used to control the attenuation of gain. The higher the value, the stronger the compression. It is denoted as R and the formula is: R = 0.5Kurt + 1.5.

[0087] Optionally, SF reflects the spectral energy distribution characteristics. The larger the SF, the flatter the spectrum and the more uniform the energy distribution; conversely, the spectrum is steeper and the energy is concentrated in a few frequency bands. To provide a safe boundary for dynamic adaptation of the compression results, the target range value can be fine-tuned based on SF to constrain the gain range in the optimized EQ curve, thereby ensuring smooth compression and controllable range as much as possible. The adjusted target range value is denoted as T′, and the formula is: T′=T(1+λ·(SF-0.7)), where λ is a coefficient that can be set according to empirical values, such as λ=0.3.

[0088] Alternatively, local energy can be used to adjust the compression threshold to enhance adaptability to local spectral features, as follows:

[0089] For the k-th frequency point on the EQ curve, the terminal obtains the gain from the kP-th to the (k+P-th)-th frequency points on the EQ curve. Based on the gain from the kP-th to the (k+P-th)-th frequency points, it calculates the local energy value and the global energy estimate corresponding to the k-th frequency point. Based on the local energy value and the global energy estimate, the compression threshold is adjusted to obtain the compression threshold for the k-th frequency point. The value of P is set based on empirical values.

[0090] For example, for the k-th frequency point, the local energy value is expressed as formula (7), and the global energy estimate is expressed as formula (8) and formula (9).

[0091]

[0092] In formulas (7) to (9), E local (k) is the local energy value at the k-th frequency point, M(i) is the gain at the i-th frequency point, P is the window radius when determining the local energy value, P is an empirical value, and K is the number of frequency points on the EQ curve.

[0093] Then based on E local (k), μ E and σ E The compression threshold is updated, as expressed in formula (10).

[0094]

[0095] Where K1 is the compression threshold before the update, K′(k) represents the compression threshold of the k-th frequency point after adjustment, and α is the scaling factor, which is obtained based on empirical values ​​or simulations, such as α = 0.3.

[0096] In this way, since the compression threshold is adjusted based on local and global energy estimates for each frequency point, the adaptability to local spectral features can be enhanced.

[0097] Step S104: Based on the compression parameters and the soft knee compression algorithm, the EQ curve is optimized to obtain the optimized EQ curve.

[0098] In this embodiment, after determining the compression parameters, the piecewise function in the soft knee compression process is used to smooth the compression spectrum curve to obtain the optimized difference, see formula (11).

[0099]

[0100] In formula (11), C(k) is the gain at the k-th frequency point on the optimized EQ curve, M(k) is the gain at the k-th frequency point on the EQ curve, and R is the compression ratio.

[0101] Furthermore, after obtaining C(k), if the absolute value of C(k) |C(k)| is less than or equal to T′, then C(k) remains unchanged. If |C(k)| is greater than T′, then a secondary adjustment is triggered on C(k), such as linear attenuation of C(k), limiting the gain to a range less than or equal to T′. This application does not limit the method of triggering the secondary adjustment of C(k) in this embodiment; for example, linear attenuation of C(k) can be achieved by multiplying the value of C(k) by a corresponding coefficient. In this way, by constraining the output range of C(k) by T′, the compressed spectral gain can be kept within a reasonable range defined by T′.

[0102] It should be noted that if T′ is not obtained, T can be used to constrain C(k).

[0103] It should also be noted that if the local energy value and the global estimated energy value of the kth frequency point are used to update K1 to obtain K′(k), then in formula (11), K1 is K′(k).

[0104] By using a piecewise function to smooth the spectrum curve, the occurrence of sudden gain changes can be reduced.

[0105] In one alternative approach, after obtaining the optimized difference, in order to reduce the muddiness of the sound quality, the over-adjustment of the low-frequency band can be suppressed, and the optimized difference of the low-frequency band can be adjusted using formula (12).

[0106]

[0107] In formula (12), C final (k) represents the gain at the k-th frequency point after suppression adjustment of the low-frequency band, and fc is the cutoff frequency. Formula (12) means that when the frequency of the k-th frequency point is less than fc, the gain of the k-th frequency point is set to 0; when the frequency of the k-th frequency point is greater than or equal to fc, the gain of the k-th frequency point is the optimized gain. fc can be an empirical value, such as f c =150Hz. Alternatively, fc can be calculated using a dynamic formula, see formula (13).

[0108]

[0109] In formula (13), f c0 For example, f c0 =150Hz, γ is a coefficient, an empirical value, such as γ=0.5. E low Low-frequency energy refers to the total energy at frequencies below a certain threshold. This threshold is related to the frequency range of the audio signal and is an empirical value. E total This represents the total energy across the entire frequency band.

[0110] In one alternative approach, to allow users to adjust the gain, the terminal can also display the optimized spectral differences according to the EQ curve. If the user feels the gain of the target frequency band is unsuitable, they can adjust the gain of the target frequency band. The terminal will then detect the adjustment command and adjust the gain of the target frequency band to the gain indicated by the adjustment command. In this way, if the user feels the gain is unsuitable, they can adjust the gain to obtain a suitable gain.

[0111] Optionally, when outputting the EQ curve, a first average spectrum and a second average spectrum can also be output as a reference.

[0112] Step S105: Based on the optimized EQ curve, perform equalization processing on the audio signal of the target audio to obtain the equalized audio signal of the target audio.

[0113] In this embodiment, the terminal performs IFFT (Inverse Fast Fourier Transform) processing on the gain of each frequency point on the optimized EQ curve using formula (14) to obtain the time-domain filter coefficients, which are FIR filter coefficients.

[0114] h(n) = IFTC final [(k)]·w(n) (14)

[0115] In formula (14), h(n) are discrete time-domain filter coefficients, which are a sequence, C final (k) represents the optimized gain at the k-th frequency point, and w(n) represents the window function.

[0116] Then the terminal performs a convolution operation on the time-domain filter coefficients and the target audio using formula (15) to obtain the equalized audio of the target audio.

[0117] y(n)=x(n)*h(n) (15)

[0118] In formula (15), y(n) represents the time signal sequence of the target audio after equalization, that is, the equalized audio signal of the target audio.

[0119] Alternatively, FFT can be used to accelerate the convolution operation process in order to quickly obtain the equalized audio.

[0120] Optionally, after obtaining the equalized audio, the equalized audio signal and time-domain filter coefficients can also be output.

[0121] In addition, to better understand the solutions of the embodiments of this application, the following are also provided. Figure 2 The process shown is as follows: Figure 2For detailed explanations of the process shown, please refer to the previous description, which will not be repeated here.

[0122] In this embodiment, compression parameters are determined through statistical characteristics of differences, eliminating the need for manual intervention and achieving automation, thus improving processing efficiency by over 70%. Furthermore, a soft-knee compression algorithm is used to smooth the EQ curve, improving its smoothness by over 30% and achieving a spectral fidelity greater than 98%. The audio equalization method in this embodiment is robust; specifically, by determining compression parameters based on statistical characteristics, it can adapt to different audio types and achieves a high success rate of 95%. In addition, the computational complexity in this embodiment is relatively low, with a real-time processing latency of less than 20ms.

[0123] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0124] Based on the same technical concept, embodiments of this application also provide an audio equalization device, see [link to relevant documentation]. Figure 3 , Figure 3 The illustrated device can be implemented as part or all of the apparatus through software, hardware, or a combination of both. This device is used to implement the method flow executed by the terminal in the embodiments of this application. The device includes:

[0125] The acquisition module 310 is used to acquire an EQ curve that represents the spectral difference between the target audio and the reference audio.

[0126] The determination module 320 is used for:

[0127] Based on the spectral differences, determine the statistical characteristics of the spectral differences;

[0128] Based on the statistical characteristics, compression parameters for range compression of the spectral differences are determined;

[0129] Based on the compression parameters and soft knee compression algorithm, the EQ curve is optimized to obtain an optimized EQ curve;

[0130] The transformation module 330 is used to perform equalization processing on the audio signal of the target audio based on the optimized EQ curve to obtain the equalized audio signal of the target audio.

[0131] In an alternative approach, the statistical characteristics include mean, standard deviation, and kurtosis, and the determining module 320 is used to:

[0132] Based on the mean and the standard deviation, determine the target range value of the compression parameters;

[0133] Based on the standard deviation, the compression threshold and soft knee width in the compression parameters are determined;

[0134] Based on the kurtosis, the compression ratio in the compression parameters is determined.

[0135] In an alternative approach, the statistical feature further includes spectral flatness; the determining module 320 is also configured to: adjust the target range value based on the spectral flatness.

[0136] In an alternative embodiment, the determining module 320 is further configured to: for the k-th frequency point on the EQ curve, based on the gain of the kP-th to k+P-th frequency points on the EQ curve, determine the local energy value and global energy estimate corresponding to the k-th frequency point, and adjust the compression threshold based on the local energy value and the global energy estimate to obtain the compression threshold of the k-th frequency point, where k is an integer greater than or equal to 1 and P is an integer greater than 1.

[0137] In one alternative approach, the compression parameters include a compression threshold, soft knee width, compression ratio, and target range value;

[0138] The determining module 320 is used for:

[0139] For each frequency point on the EQ curve, if the absolute value of the gain of the frequency point is less than or equal to a first value, and the absolute value is less than or equal to the target range value, then the gain of the frequency point is determined as the gain of the frequency point on the optimized EQ curve, wherein the first value is equal to the difference between the compression threshold and half of the soft knee width.

[0140] If the absolute value of the gain at the frequency point is greater than the first value and less than or equal to the second value, then the gain at the frequency point on the EQ curve is added to the third value to obtain a fourth value. If the absolute value of the fourth value is less than or equal to the target range value, the fourth value is determined as the gain at the frequency point on the optimized EQ curve. The second value is equal to the sum of the compression threshold and half of the soft knee width, and the third value is equal to the square of the difference between the gain at the frequency point on the EQ curve and the first value divided by twice the soft knee width.

[0141] If the absolute value of the gain at the frequency point is greater than the second value, then the second value and the fifth value are added to obtain a sixth value. If the absolute value of the sixth value is less than or equal to the target range value, the sixth value is determined as the gain of the frequency point on the optimized EQ curve. The fifth value is equal to the product of the reciprocal of the compression ratio and the difference between the gain of the frequency point on the EQ curve and the second value.

[0142] In an alternative embodiment, the determining module 320 is further configured to set the gain of frequencies less than the cutoff frequency on the optimized EQ curve to 0 before performing equalization processing on the audio signal of the target audio based on the optimized EQ curve.

[0143] In an alternative embodiment, the determining module 320 is configured to: determine a first average spectrum of the M frames with the highest energy in the target audio, and determine a second average spectrum of the M frames with the highest energy in the reference audio, wherein M is an integer greater than 1;

[0144] Generate an EQ curve that characterizes the difference between the first average spectrum and the second average spectrum.

[0145] In an alternative embodiment, the apparatus further includes an interaction module for displaying the optimized EQ curve before equalizing the audio signal of the target audio based on the optimized EQ curve.

[0146] Receive user instructions to adjust the gain of the target frequency band in the optimized EQ curve;

[0147] The determining module 320 is further configured to: adjust the gain of the target frequency band in the optimized EQ curve based on the adjustment instruction.

[0148] It should be noted that the audio equalization device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio equalization device and the audio equalization method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0149] Figure 4 This illustration shows a structural block diagram of a terminal 400 provided in an exemplary embodiment of this application. The terminal 400 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 400 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0150] Typically, terminal 400 includes a processor 401 and a memory 402.

[0151] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0152] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 402 are used to store at least one instruction, which is executed by processor 401 to implement the equalization audio method provided in the method embodiments of this application.

[0153] In some embodiments, the terminal 400 may also optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, memory 402, and peripheral device interface 403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, a positioning assembly 408, and a power supply 409.

[0154] Peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 401 and memory 402. In some embodiments, processor 401, memory 402 and peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 401, memory 402 and peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0155] The radio frequency (RF) circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 404 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0156] Display screen 405 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 401 for processing. In this case, display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 405, disposed on the front panel of terminal 400; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 400 or in a folded design; in other embodiments, display screen 405 may be a flexible display screen, disposed on a curved or folded surface of terminal 400. Furthermore, display screen 405 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 405 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0157] The camera assembly 406 is used to acquire images or videos. Optionally, the camera assembly 406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0158] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 401 for processing, or to the radio frequency circuit 404 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 400. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.

[0159] The positioning component 408 is used to determine the current geographical location of the terminal 400 in order to enable navigation or LBS (Location Based Service). The positioning component 408 can be a positioning component based on GPS (Global Positioning System), BeiDou system, or Galileo system.

[0160] Power supply 409 is used to power the various components in terminal 400. Power supply 409 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 409 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0161] In some embodiments, the terminal 400 further includes one or more sensors 410. The one or more sensors 410 include, but are not limited to: an accelerometer 411, a gyroscope 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.

[0162] Accelerometer 411 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 400. For example, accelerometer 411 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 401 can control display screen 405 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 411. Accelerometer 411 can also be used for games or for acquiring user motion data.

[0163] The gyroscope sensor 412 can detect the orientation and rotation angle of the terminal 400. The gyroscope sensor 412, in conjunction with the accelerometer sensor 411, can collect 3D motion data from the user on the terminal 400. Based on the data collected by the gyroscope sensor 412, the processor 401 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0164] The pressure sensor 413 can be disposed on the side bezel of the terminal 400 and / or on the lower layer of the display screen 405. When the pressure sensor 413 is disposed on the side bezel of the terminal 400, it can detect the user's grip signal on the terminal 400, and the processor 401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 413. When the pressure sensor 413 is disposed on the lower layer of the display screen 405, the processor 401 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0165] The fingerprint sensor 414 is used to collect the user's fingerprint. The processor 401 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 414, or the fingerprint sensor 414 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 401 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 414 can be located on the front, back, or side of the terminal 400. When the terminal 400 has physical buttons or a manufacturer's logo, the fingerprint sensor 414 can be integrated with the physical buttons or manufacturer's logo.

[0166] An optical sensor 415 is used to collect ambient light intensity. In one embodiment, the processor 401 can control the display brightness of the display screen 405 based on the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the display screen 405 is increased; when the ambient light intensity is low, the display brightness of the display screen 405 is decreased. In another embodiment, the processor 401 can also dynamically adjust the shooting parameters of the camera assembly 406 based on the ambient light intensity collected by the optical sensor 415.

[0167] The proximity sensor 416, also known as a distance sensor, is typically located on the front panel of the terminal 400. The proximity sensor 416 is used to detect the distance between the user and the front of the terminal 400. In one embodiment, when the proximity sensor 416 detects that the distance between the user and the front of the terminal 400 is gradually decreasing, the processor 401 controls the display screen 405 to switch from a screen-on state to a screen-off state; when the proximity sensor 416 detects that the distance between the user and the front of the terminal 400 is gradually increasing, the processor 401 controls the display screen 405 to switch from a screen-off state to a screen-on state.

[0168] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on terminal 400 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0169] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 500 can vary significantly due to differences in configuration or performance, and may include one or more CPUs 501 and one or more memories 502. The memory 502 stores at least one instruction, which is loaded and executed by the processor 501 to implement the methods provided in the above-described method embodiments. Of course, the computer device may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device may also include other components for implementing device functions, which will not be elaborated upon here.

[0170] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to perform the audio equalization method described above. This computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.

[0171] In an exemplary embodiment, a computer program product is also provided, which stores at least one instruction that is loaded and executed by the processor to perform the operations performed by the equalization audio method described above.

[0172] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices) involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the reference audio involved in this application was obtained with full authorization.

[0173] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0174] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for equalizing audio, characterized in that, The method includes: Obtain the equalization EQ curve that represents the spectral difference between the target audio and the reference audio; Based on the spectral differences, determine the statistical characteristics of the spectral differences; Based on the statistical characteristics, compression parameters for range compression of the spectral differences are determined; Based on the compression parameters and soft knee compression algorithm, the EQ curve is optimized to obtain an optimized EQ curve; Based on the optimized EQ curve, the audio signal of the target audio is subjected to equalization processing to obtain the equalized audio signal of the target audio.

2. The method according to claim 1, characterized in that, The statistical features include mean, standard deviation, and kurtosis. The step of determining compression parameters for range compression of the spectral differences based on the statistical features includes: Based on the mean and the standard deviation, determine the target range value of the compression parameters; Based on the standard deviation, the compression threshold and soft knee width in the compression parameters are determined; Based on the kurtosis, the compression ratio in the compression parameters is determined.

3. The method according to claim 2, characterized in that, The statistical feature also includes spectral flatness; the method further includes: The target range value is adjusted based on the spectral flatness.

4. The method according to claim 2, characterized in that, The method further includes: For the k-th frequency point on the EQ curve, based on the gain from the kP-th to the (k+P)-th frequency points on the EQ curve, the local energy value and global energy estimate corresponding to the k-th frequency point are determined. Based on the local energy value and the global energy estimate, the compression threshold is adjusted to obtain the compression threshold of the k-th frequency point, where k is an integer greater than or equal to 1 and P is an integer greater than 1.

5. The method according to any one of claims 1 to 4, characterized in that, The compression parameters include compression threshold, soft knee width, compression ratio, and target range value; The optimization process based on the compression parameters and the soft-knee compression algorithm to obtain the optimized EQ curve includes: For each frequency point on the EQ curve, if the absolute value of the gain of the frequency point is less than or equal to a first value, and the absolute value is less than or equal to the target range value, then the gain of the frequency point is determined as the gain of the frequency point on the optimized EQ curve, wherein the first value is equal to the difference between the compression threshold and half of the soft knee width. If the absolute value of the gain at the frequency point is greater than the first value and less than or equal to the second value, then the gain at the frequency point on the EQ curve is added to the third value to obtain a fourth value. If the absolute value of the fourth value is less than or equal to the target range value, the fourth value is determined as the gain at the frequency point on the optimized EQ curve. The second value is equal to the sum of the compression threshold and half of the soft knee width, and the third value is equal to the square of the difference between the gain at the frequency point on the EQ curve and the first value divided by twice the soft knee width. If the absolute value of the gain at the frequency point is greater than the second value, then the second value and the fifth value are added to obtain a sixth value. If the absolute value of the sixth value is less than or equal to the target range value, the sixth value is determined as the gain of the frequency point on the optimized EQ curve. The fifth value is equal to the product of the reciprocal of the compression ratio and the difference between the gain of the frequency point on the EQ curve and the second value.

6. The method according to any one of claims 1 to 4, characterized in that, Before performing equalization processing on the audio signal of the target audio based on the optimized EQ curve, the method further includes: Set the gain of frequencies below the cutoff frequency on the optimized EQ curve to 0.

7. The method according to any one of claims 1 to 4, characterized in that, The acquisition of the EQ curve representing the spectral difference between the target audio and the reference audio includes: Determine the first average spectrum of the M frames with the highest energy in the target audio, and determine the second average spectrum of the M frames with the highest energy in the reference audio, where M is an integer greater than 1; Generate an EQ curve that characterizes the difference between the first average spectrum and the second average spectrum.

8. The method according to any one of claims 1 to 4, characterized in that, Before performing equalization processing on the audio signal of the target audio based on the optimized EQ curve, the method further includes: Display the optimized EQ curve; Receive user instructions to adjust the gain of the target frequency band in the optimized EQ curve; Based on the adjustment command, the gain of the target frequency band in the optimized EQ curve is adjusted.

9. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction that is loaded and executed by the processor to perform the operations of the equalizing audio method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to perform the operation of the equalizing audio method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, The computer program product stores at least one instruction, which is loaded and executed by a processor to perform the operation of the equalization audio method as described in any one of claims 1 to 8.