Audio processing method, device and non-volatile computer readable storage medium
Patent Information
- Application Number
- CN202210885907.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-07-26
AI Technical Summary
[0004]本公开的发明人发现上述相关技术中存在如下问题:容易出现严重的截幅失真,导致响度均衡效果差
[0036] In the above embodiments, the gain of each sampling point is determined based on the pre-estimated loudness peak value. In this way, loudness equalization processing based on the loudness peak value can solve the technical problem of audio clipping distortion and eliminate phenomena such as fluctuating loudness and over-amplification, thereby improving the effect of loudness equalization.
Smart Images

Figure CN117499838B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of signal processing technology, and in particular to an audio processing method, an audio processing apparatus, and a non-volatile computer-readable storage medium. Background Technology
[0002] The loudness of different audio and video recordings often varies, requiring users to frequently adjust the volume. Furthermore, this loudness "battle" can cause hearing damage. Loudness equalization technology ensures that all audio loudness is within a preset range during video playback, allowing users to maintain a more ideal listening experience without constantly adjusting the volume, thus providing some protection for the listener's hearing.
[0003] In related technologies, global loudness is calculated, and the loudness value or the maximum loudness value is used to perform uniform gain processing on the audio. Summary of the Invention
[0004] The inventors of this disclosure have discovered the following problem in the above-mentioned related technologies: severe clipping distortion is prone to occur, resulting in poor loudness equalization effect.
[0005] In view of this, this disclosure proposes an audio processing technology that can improve the loudness equalization effect.
[0006] According to some embodiments of this disclosure, an audio processing method is provided, including: estimating the loudness peak within a preset time length based on the gain corresponding to the current sampling point of the current frame in the audio; determining the gain corresponding to the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point; and performing loudness equalization processing on the current frame using each sampling point of the current frame and its corresponding gain.
[0007] In some embodiments, determining the gain corresponding to the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point includes: determining the gain corresponding to the next sampling point based on whether the loudness peak is less than a peak threshold.
[0008] In some embodiments, determining the gain corresponding to the next sampling point based on whether the loudness peak value is less than the peak threshold includes: if the loudness peak value is less than the peak threshold, determining the gain corresponding to the next sampling point based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point; if the loudness peak value is greater than or equal to the peak threshold, determining the gain corresponding to the next sampling point based on the difference between the loudness peak value and the target loudness.
[0009] In some embodiments, determining the gain corresponding to the next sampling point includes: if the loudness peak value is less than a peak threshold, determining the target gain of the next sampling point based on the difference between the loudness of the current sampling point and the target loudness, and the gain corresponding to the current sampling point; if the loudness peak value is greater than or equal to the peak threshold, determining the target gain of the next sampling point based on the difference between the loudness peak value and the target loudness, and the gain corresponding to the current sampling point; and determining the gain corresponding to the next sampling point based on the gain corresponding to the current sampling point and the target gain, wherein the gain corresponding to the next sampling point is less than the target gain.
[0010] In some embodiments, determining the gain corresponding to the next sampling point based on the gain corresponding to the current sampling point and the target gain includes: if the loudness peak value is less than the peak threshold, determining whether to use a first convergence rate factor or a second convergence rate factor based on whether the difference between the loudness of the current sampling point and the target loudness is positive or negative, wherein the first convergence rate factor and the second convergence rate factor are different; and determining the gain corresponding to the next sampling point based on the first convergence rate factor or the second convergence rate factor and the gains corresponding to historical sampling points.
[0011] In some embodiments, determining the gain corresponding to the next sampling point based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point includes: adjusting the loudness of the current sampling point using the gain corresponding to the current sampling point; and determining the gain corresponding to the next sampling point based on the difference between the adjusted loudness and the target loudness.
[0012] In some embodiments, determining the gain corresponding to the next sampling point of the current frame includes: determining a candidate gain for the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point; determining the candidate gain as the gain corresponding to the next sampling point if the candidate gain does not exceed a gain threshold; and determining the gain corresponding to the next sampling point based on the gain threshold if the candidate gain exceeds a gain threshold.
[0013] In some embodiments, loudness equalization processing of the current frame includes: adjusting the loudness of the current sampling point using the gain corresponding to the current sampling point; if the loudness of the audio does not exceed a first loudness threshold, using the adjusted loudness as the output loudness of the current sampling point; if the loudness of the audio exceeds the first loudness threshold, determining the output loudness of the current sampling point according to the first loudness threshold.
[0014] In some embodiments, the processing method further includes: determining whether the current frame is a keyframe based on the loudness of the current frame; if the current frame is not a keyframe, determining the gain corresponding to all sampling points in the current frame as a preset gain value; if the current frame is a keyframe, determining the gain corresponding to the current sampling point of the current frame.
[0015] In some embodiments, determining whether a current frame is a keyframe based on the loudness of the current frame includes: determining whether the current frame is a keyframe based on a comparison between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold, wherein the second loudness threshold is calculated based on the average loudness of each frame in the audio.
[0016] In some embodiments, determining whether a current frame is a keyframe based on a comparison between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold includes: determining that the current frame is not a keyframe if the loudness of the current frame is less than the second loudness threshold or less than the third loudness threshold; and determining that the current frame is a keyframe if the loudness of the current frame is greater than or equal to the second loudness threshold and greater than or equal to the third loudness threshold.
[0017] In some embodiments, the loudness range of the audio is greater than or equal to a loudness range threshold, and the loudness of the audio is less than or equal to a fourth loudness threshold.
[0018] In some embodiments, the processing method further includes: performing loudness equalization processing on the current frame using FFMPEG when the loudness range of the audio is less than a loudness range threshold; and performing loudness equalization processing on the current frame using global linear gain when the loudness of the audio is greater than a fourth loudness threshold.
[0019] In some embodiments, estimating the future loudness peak based on the gain corresponding to the current sampling point of the current frame in the audio includes: estimating the loudness peak based on the gain corresponding to the current sampling point in the multi-channel fused audio.
[0020] According to some other embodiments of this disclosure, an audio processing apparatus is provided, comprising: an estimation unit, configured to estimate a loudness peak within a preset time length based on the gain corresponding to the current sampling point of the current frame in the audio; a determination unit, configured to determine the gain corresponding to the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point; and an equalization unit, configured to perform loudness equalization processing on the current frame using each sampling point of the current frame and its corresponding gain.
[0021] In some embodiments, the determining unit determines the gain corresponding to the next sampling point based on whether the loudness peak value is less than the peak threshold.
[0022] In some embodiments, when the loudness peak value is less than the peak threshold, the determining unit determines the gain corresponding to the next sampling point based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point; when the loudness peak value is greater than or equal to the peak threshold, the determining unit determines the gain corresponding to the next sampling point based on the difference between the loudness peak value and the target loudness.
[0023] In some embodiments, when the loudness peak value is less than the peak threshold, the determining unit determines the target gain of the next sampling point based on the difference between the loudness of the current sampling point and the target loudness, and the gain corresponding to the current sampling point. When the loudness peak value is greater than or equal to the peak threshold, the determining unit determines the target gain of the next sampling point based on the difference between the loudness peak value and the target loudness, and the gain corresponding to the current sampling point. The determining unit also determines the gain corresponding to the next sampling point based on the gain corresponding to the current sampling point and the target gain. The gain corresponding to the next sampling point is less than the target gain.
[0024] In some embodiments, when the loudness peak value is less than the peak value threshold, the determining unit determines whether to use a first convergence rate factor or a second convergence rate factor based on whether the difference between the loudness of the current sampling point and the target loudness is positive or negative. The first convergence rate factor and the second convergence rate factor are different. Based on the first convergence rate factor or the second convergence rate factor and the gain corresponding to the historical sampling points, the unit determines the gain corresponding to the next sampling point.
[0025] In some embodiments, the determining unit adjusts the loudness of the current sampling point using the gain corresponding to the current sampling point; and determines the gain corresponding to the next sampling point based on the difference between the adjusted loudness and the target loudness.
[0026] In some embodiments, the determining unit determines the candidate gain for the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point. If the candidate gain does not exceed the gain threshold, the candidate gain is determined as the gain corresponding to the next sampling point. If the candidate gain exceeds the gain threshold, the gain corresponding to the next sampling point is determined based on the gain threshold.
[0027] In some embodiments, the equalization unit adjusts the loudness of the current sampling point using the gain corresponding to the current sampling point. If the loudness of the audio does not exceed a first loudness threshold, the adjusted loudness is used as the output loudness of the current sampling point. If the loudness of the audio exceeds the first loudness threshold, the output loudness of the current sampling point is determined according to the first loudness threshold.
[0028] In some embodiments, the processing apparatus further includes: a determination unit, configured to determine whether the current frame is a key frame based on the loudness of the current frame; wherein, if the current frame is not a key frame, the determination unit determines the gain corresponding to all sampling points in the current frame as a preset gain value, and if the current frame is a key frame, determines the gain corresponding to the current sampling point of the current frame.
[0029] In some embodiments, the determining unit determines whether the current frame is a key frame based on the comparison result between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold. The second loudness threshold is calculated based on the average loudness of each frame in the audio.
[0030] In some embodiments, the determining unit determines that the current frame is not a key frame if the loudness of the current frame is less than a second loudness threshold or less than a third loudness threshold, and determines that the current frame is a key frame if the loudness of the current frame is greater than or equal to the second loudness threshold and greater than or equal to the third loudness threshold.
[0031] In some embodiments, the loudness range of the audio is greater than or equal to a loudness range threshold, and the loudness of the audio is less than or equal to a fourth loudness threshold.
[0032] In some embodiments, when the loudness range of the audio is less than the loudness range threshold, the equalization unit performs loudness equalization processing on the current frame using the FFMPEG method; when the loudness of the audio is greater than the fourth loudness threshold, it performs loudness equalization processing on the current frame using the global linear gain method.
[0033] In some embodiments, the estimation unit estimates the loudness peak based on the gain corresponding to the current sampling point in the multi-channel fused audio.
[0034] According to further embodiments of the present disclosure, an audio processing apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the audio processing method of any of the above embodiments based on instructions stored in the memory device.
[0035] According to further embodiments of the present disclosure, a non-volatile computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the audio processing method of any of the above embodiments.
[0036] In the above embodiments, the gain of each sampling point is determined based on the pre-estimated loudness peak value. In this way, loudness equalization processing based on the loudness peak value can solve the technical problem of audio clipping distortion and eliminate phenomena such as fluctuating loudness and over-amplification, thereby improving the effect of loudness equalization. Attached Figure Description
[0037] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.
[0038] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description:
[0039] Figure 1 Flowcharts illustrating some embodiments of the audio processing techniques of this disclosure;
[0040] Figure 2Flowcharts illustrating other embodiments of the audio processing techniques of this disclosure;
[0041] Figure 3 Flowcharts illustrating further embodiments of the audio processing techniques of this disclosure;
[0042] Figure 4 Block diagrams illustrating some embodiments of the audio processing apparatus of this disclosure;
[0043] Figure 5 Block diagrams illustrating other embodiments of the audio processing apparatus of this disclosure;
[0044] Figure 6 Block diagrams illustrating further embodiments of the audio processing apparatus of this disclosure are shown. Detailed Implementation
[0045] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0046] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0047] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0048] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0049] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0050] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0051] As mentioned earlier, the method of directly calculating the global loudness and then the global gain does not take into account the drastic loudness changes in each frame of a single audio file in reality. Calculating the global loudness is actually averaging the global loudness. Using the average value to amplify the audio is not favorable for extreme values and is prone to clipping distortion.
[0052] Directly using the maximum value to calculate the gain of audio also presents similar technical problems. If the extreme value in an audio clip is close to the full amplitude, but the loudness of other frames is extremely low, there is a possibility of further reducing the overall loudness of the audio. This can easily result in an overall low loudness value.
[0053] Furthermore, the loudness equalization scheme in FFMPEG is based on an algorithm developed according to the EBU R.128 standard. The entire EBU R.128 standard is based on the requirements and usage specifications of advertisements, music, and film and television works, and is not suitable for short video platforms.
[0054] The characteristics of works such as advertisements, music, and film and television are that the dubbing and music are all produced through rigorous post-production, with extremely high recording conditions and very strict post-production. Because there are no issues such as background noise, DC bias, etc., during the audio production process, noise and other problems are rare or even nonexistent.
[0055] However, those creating short videos are mostly ordinary users, and the types and quality of the uploaded videos are diverse. For example, there are personal music recordings, as well as works synthesized using software, some are candid snapshots of life, and others are introductions and explanations of things.
[0056] Therefore, the overall audio content is very diverse, including background music, noise, human voices, and other background noises such as insect chirping and car horns; moreover, there are many tools used for recording short videos, including mobile phones, video recorders, voice recorders, and audio pickup devices, with varying levels of equipment quality; it may also involve a wide variety of recording apps and audio production and editing software.
[0057] The EBU R.128 standard has three important metrics: loudness, dynamic range, and peak value. In order to make the overall loudness value of the original audio conform to the set metric value, the overall loudness range (i.e., dynamic range) and loudness magnitude of the original audio need to be rescaled.
[0058] The loudness equalization scheme in FFMPEG operates frame-by-frame, and the gain differences between frames are relatively large. This ultimately causes audio to fluctuate in volume, silence or noise to be excessively amplified, and background noise to be amplified, affecting sound quality and listening experience.
[0059] For example, the dynamic range of the audio itself may far exceed 7dB, while the loudness might only be -30dB, with a relatively large maximum value. The loudness equalization algorithm in FFMPEG requires significant gain boosting for low volumes and gain descent for originally louder areas to meet dynamic range requirements. If the gain changes between frames are too large, the final effect will severely damage the music, making it sound inconsistent in volume and significantly diminishing its rhythm and intonation.
[0060] To address the aforementioned technical issues, this disclosure proposes a loudness equalization technology solution suitable for short video platforms. While performing loudness equalization on audio, this disclosure reduces DC noise and limits noise amplification, mitigating the technical problems of unstable loudness caused by fluctuations in volume and noise resulting from methods such as FFMPEG. Simultaneously, it avoids technical problems such as clipping distortion and unintended loudness reduction, thereby improving the listening experience and enhancing the auditory perception.
[0061] For example, the technical solution of this disclosure can be implemented through the following embodiments.
[0062] Figure 1 Flowcharts illustrating some embodiments of the audio processing techniques of this disclosure are shown.
[0063] like Figure 1 As shown, in step 110, the loudness peak within a preset time length is estimated based on the gain corresponding to the current sampling point of the current frame in the audio.
[0064] In some embodiments, the loudness peak is estimated based on the gain corresponding to the current sampling point in the multi-channel fused audio.
[0065] For example, a high-pass filter is used to filter out DC noise and interference below 50Hz; a weighting method (such as K-weighting) is used to filter each channel; energy is calculated point by point for the filtered multi-channels; and the multi-channels are fused into a single-channel energy. This stage is the channel fusion and channel preprocessing stage (i.e., the preparation stage), after which the remaining steps of this disclosure can be used for the signal processing stage.
[0066] In some embodiments, the loudness of the current frame determines whether the current frame is a key frame; if the current frame is not a key frame, the gain corresponding to all sampling points in the current frame is determined as a preset gain value; if the current frame is a key frame, the gain corresponding to the current sampling point of the current frame is determined.
[0067] For example, based on the comparison between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold, it is determined whether the current frame is a key frame. The second loudness threshold is calculated based on the average loudness of each frame in the audio.
[0068] For example, if the loudness of the current frame is less than the second loudness threshold or less than the third loudness threshold, the current frame is determined not to be a key frame; if the loudness of the current frame is greater than or equal to the second loudness threshold and greater than or equal to the third loudness threshold, the current frame is determined to be a key frame.
[0069] In some embodiments, firstly, the processing of relevant parameters such as the pre-observed frame length, the processing single frame length, the target loudness value, the target peak value, the low noise or silence loudness threshold value, and the loudness silence segment fluctuation range value is set, and the loudness mean value, gain value, etc. are initialized; then, the processing single frame length and the pre-observed frame length are loaded, and the loudness of the current frame is calculated.
[0070] If the loudness is less than the third loudness threshold (which can be the difference between the loudness mean and the dynamic range) or less than the second loudness threshold, the target gain of the entire frame's sampling points is set to 1, and this target gain is used to smooth the historical gain in order to determine the gain of the entire frame's sampling points.
[0071] Otherwise, the current frame is identified as the key frame for which gain needs to be calculated point by point. The energy of the sample points in this frame is smoothed for loudness to obtain smoothed energy. The current gain is used to estimate the peak loudness of the sample points within the pre-observation length. If there are sample points that exceed the peak threshold, the loudness of the points exceeding the threshold is recursively calculated with the target loudness, and the gain value is adjusted. This stage is the gain calculation stage.
[0072] In step 120, the gain corresponding to the next sampling point of the current frame is determined based on the loudness peak value and the gain corresponding to the current sampling point.
[0073] In some embodiments, the gain corresponding to the next sampling point is determined based on whether the loudness peak value is less than the peak threshold.
[0074] In some embodiments, when the loudness peak value is less than the peak threshold, the gain corresponding to the next sampling point is determined based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point; when the loudness peak value is greater than or equal to the peak threshold, the gain corresponding to the next sampling point is determined based on the difference between the loudness peak value and the target loudness.
[0075] For example, if the loudness peak is less than the peak threshold, the target gain for the next sampling point is determined based on the difference between the loudness of the current sampling point and the target loudness, as well as the gain corresponding to the current sampling point; if the loudness peak is greater than or equal to the peak threshold, the target gain for the next sampling point is determined based on the difference between the loudness peak and the target loudness, as well as the gain corresponding to the current sampling point; the gain corresponding to the next sampling point is determined based on the gain corresponding to the current sampling point and the target gain, and the gain corresponding to the next sampling point is less than the target gain.
[0076] For example, when the loudness peak is less than the peak threshold, the first or second convergence rate factor is determined based on whether the difference between the loudness of the current sampling point and the target loudness is positive or negative. The first and second convergence rate factors are different. Based on the first or second convergence rate factor and the gain corresponding to the historical sampling points, the gain corresponding to the next sampling point is determined.
[0077] For example, if the estimated loudness peak within the pre-observation length does not exceed the target peak (i.e., the peak threshold), the gain calculation phase also begins. Gain calculation may include: first, setting different convergence rate factors for pull-up and push-down; calculating the gain on the energy-smoothed sampling points to obtain a new loudness; calculating the difference between the new loudness and the target loudness; selecting the appropriate convergence rate factor based on the direction of the difference; updating the historical gain based on the convergence rate factor and the difference; and selecting the updated value and the maximum gain value to determine the new gain.
[0078] In some embodiments, the loudness of the current sampling point is adjusted using the gain corresponding to the current sampling point; and the gain corresponding to the next sampling point is determined based on the difference between the adjusted loudness and the target loudness.
[0079] In some embodiments, a candidate gain for the next sampling point of the current frame is determined based on the loudness peak and the gain corresponding to the current sampling point; if the candidate gain does not exceed the gain threshold, the candidate gain is determined as the gain corresponding to the next sampling point; if the candidate gain exceeds the gain threshold, the gain corresponding to the next sampling point is determined based on the gain threshold.
[0080] In step 130, loudness equalization is performed on the current frame using each sampling point and its corresponding gain.
[0081] In some embodiments, the loudness of the current sampling point is adjusted using the gain corresponding to the current sampling point; if the loudness of the audio does not exceed a first loudness threshold, the adjusted loudness is used as the output loudness of the current sampling point; if the loudness of the audio exceeds the first loudness threshold, the output loudness of the current sampling point is determined according to the first loudness threshold.
[0082] For example, before outputting the loudness result, the adjusted loudness peak value is checked to ensure that the loudness peak value does not exceed the set value.
[0083] For example, repeat the above steps until all sampling points of the current frame have been processed; after processing the current frame, update the next frame. Update the next frame for the pre-observed frame until the calculation of all frames in the audio has been completed.
[0084] In some embodiments, the loudness range of the audio is greater than or equal to a loudness range threshold, and the loudness of the audio is less than or equal to a fourth loudness threshold. For example, when the loudness range of the audio is greater than or equal to a loudness range threshold, and the loudness of the audio is less than or equal to a fourth loudness threshold, the equalization method of the embodiments including steps 110 to 130 is performed.
[0085] In some embodiments, when the loudness range of the audio is less than the loudness range threshold, loudness equalization processing of the current frame is performed using the FFMPEG method; when the loudness of the audio is greater than the fourth loudness threshold, loudness equalization processing of the current frame is performed using the global linear gain method.
[0086] For example, using the EBU R.128 standard, loudness is calculated frame by frame for audio loudness, and the loudness range, loudness value, peak index, etc. are statistically analyzed. This process is the pre-statistical stage. Based on the results of the pre-statistical stage, different loudness equalization methods can be selected to complete the equalization operation. This process is the formal processing stage.
[0087] For example, the pre-statistical stage includes: using a high-pass filter to suppress low-frequency noise and DC components within 50Hz; calculating loudness range, loudness value, and peak index; and then proceeding to the formal processing stage.
[0088] For example, the formal processing stage includes: filtering each channel using the high-pass filter from the preprocessing stage; selecting an equalization method based on the calculated loudness range. If the loudness range is less than a loudness range threshold, the loudness equalization method in FFMPEG is selected; if the loudness range is not less than the loudness range threshold, the relationship between the loudness and a fourth loudness threshold is determined; if the loudness is greater than the fourth loudness threshold, the audio is reduced overall according to the difference gain (e.g., a global linear gain method); otherwise, the equalization method of the embodiment including steps 110-130 is used.
[0089] In some embodiments, in order to improve effectiveness and save computing resources, instead of performing pre-statistical calculations on the indicators, high-pass filtering is first applied to each channel, and then loudness equalization is performed using the equalization method of this disclosure.
[0090] The above embodiments can be applied to short video platforms, etc., and can eliminate the impact of DC noise on audio quality; avoid the shortcomings of linear gain in loudness equalization; alleviate the technical problems of open source algorithms such as fluctuating volume and over-amplification in the listening experience; and bring out the advantages of different algorithms by combining the pre-estimation stage and the formal processing stage, while avoiding the occurrence of bad cases.
[0091] The loudness equalization algorithm disclosed herein can also be used in direct processing methods without a pre-estimation stage to alleviate computational burden.
[0092] The technical solution disclosed herein takes into account the human ear's response to audio and uses weighting (such as K-weighting) to match digital loudness with human ear perception, thereby achieving the unity of digital and sensory perception.
[0093] For multi-channel audio, the same gain is used for each channel, and no changes are made to spatial and positional information such as sound image, so the spatial perception of the original audio can be completely preserved.
[0094] The technical solution disclosed herein enables automated processing, allowing for batch calculations and strategy selection without human intervention. It alleviates abrupt changes in audio loudness caused by varying production levels and recording quality on short video platforms, thereby improving the user experience, protecting users' hearing, and reducing auditory fatigue.
[0095] Figure 2 Flowcharts illustrating some other embodiments of the audio processing techniques of this disclosure are shown.
[0096] like Figure 2 As shown, in steps 210 and 220, each channel of the input raw audio file to be processed is subjected to high-pass filtering and "K" weighting.
[0097] In step 230, if the audio file is multi-channel audio, the multi-channel energy is fused according to different weights.
[0098] In some embodiments, if the overall processing flow adopts a two-stage model including a preprocessing and statistical stage, then it proceeds to the preprocessing and statistical stage; if the overall processing flow directly adopts a single-stage model, then it proceeds directly according to... Figure 1 The process of "loudness equalization algorithm disclosed herein" is processed.
[0099] For example, whether to use a single-stage or two-stage model can be manually specified in advance.
[0100] In step 240, the preprocessing statistics stage is entered to perform preliminary estimations. For example, the loudness range, loudness value, and peak value can be pre-calculated based on the EBU R.128 standard.
[0101] In steps 250 and 260, different loudness equalization methods are selected based on the pre-estimated results.
[0102] In some embodiments, if the loudness range is less than the loudness range threshold, the loudness equalization method in FFMPEG is selected; otherwise, if the loudness is greater than the fourth loudness threshold, the global linear gain method is used for loudness equalization; otherwise, the loudness equalization method in any of the above embodiments is used.
[0103] Figure 3 Flowcharts illustrating further embodiments of the audio processing techniques of this disclosure are shown.
[0104] like Figure 3 The diagram illustrates the flow of some embodiments of the loudness equalization method disclosed herein. For example, the input file for this method is a single-channel energy signal file after high-pass and "K" weighting and multi-channel fusion. The input file also includes preset values such as an initial smoothing factor, initial gain, loudness mean, silence and low-noise thresholds, speech pause range, frame length, and pre-observed future point length.
[0105] In step 310, the file to be processed is truncated at the frame level, and the loudness of the frame is calculated.
[0106] In step 320, the loudness is smoothed.
[0107] In step 330, it is determined whether the smoothed loudness is less than a second loudness threshold or a third loudness threshold. For example, the third loudness threshold includes the relative loudness value = the loudness mean - a preset range value.
[0108] In step 340, if the smoothed loudness is less than the second loudness threshold or less than the relative loudness value, the current frame is determined to be a non-critical frame, and the target gain of all sampling points in the entire frame is 1.
[0109] In step 350, otherwise, the energy of each sampling point in the current frame is smoothed.
[0110] In step 360, the loudness peak is predicted for future values using the current gain.
[0111] In step 370, it is determined whether the loudness peak value is less than the peak value threshold.
[0112] In step 375, if the loudness peak value is greater than the peak value threshold, the new gain is obtained by using the difference between the loudness peak value and the target loudness.
[0113] In step 378, otherwise, the new gain is obtained by using the difference between the loudness of the current sampling point and the target loudness.
[0114] In step 380, the target gain for the next sampling point is calculated by combining the above difference with the current gain.
[0115] In some embodiments, the target gain for the next sampling point is recursively calculated using historical gain. For example, pull-up and push-down utilize different convergence rate factors.
[0116] In step 385, the current gain is applied to each channel of the original audio signal.
[0117] In step 388, the new gain value is smoothed against the historical gain. For example, if the target gain for the next sampling point is 10 and the historical gain (such as the gain at the previous sampling point) is 2, the gain for the next sampling point can be smoothed to a value less than 10 and greater than 2 (such as 7).
[0118] In step 390, the amplitude of the processed current sampling point is limited using the first loudness threshold.
[0119] In step 395, a limit is set on the smoothed gain. For example, a gain threshold is used to limit the gain of the next sampling point. For instance, if the gain calculated in step 388 is 7, which exceeds the gain threshold of 5, then the gain of the next sampling point is limited to a value of 5 or less.
[0120] At this point, the adjustment of one sampling point is completed, and the gain value of the next sampling point is obtained. The process of processing and judging the next point begins. This process is repeated (steps 360 to 395) until the entire frame calculation is completed, and the next frame signal is input. The cycle of the next frame begins (steps 330 to 395).
[0121] In some embodiments, the multi-channel audio loudness equalization technique proposed in this disclosure integrates preprocessing and formal processing. The two-stage loudness equalization method, where preprocessing provides a basis for subsequent method selection, is applicable to both single-channel and multi-channel audio. For example, it may include: pre-processing the multi-channel signal using a preprocessing filter, fusing the multi-channel signal according to channel weights, calculating loudness range, loudness value, peak value, and other indicators based on the fused signal, and selecting a loudness equalization method based on the calculation results.
[0122] For example, the loudness equalization method disclosed herein, which is based on peak estimation and point-by-point calculation of single-frame loudness values, limits the gain and performs peak checks on the output results to ensure that the peak value does not exceed the set value.
[0123] For example, the preprocessing method disclosed herein includes: processing single-pass multi-channel or multi-channel signals using high-pass filtering and "K" weighting, and fusing them into a single-channel energy signal.
[0124] For example, the statistical method of this disclosure uses framing technology to calculate loudness range, loudness value, and peak value according to the EBU R.128 standard.
[0125] For example, the loudness equalization method selection of this disclosure includes: first, determining the equalization method based on the loudness range; if the original loudness range is within the set target, then the FFmpeg method is used; otherwise, the equalization method is determined based on the loudness value; if the loudness value is greater than the target value, then the global gain method is used to process the audio; otherwise, the loudness equalization algorithm proposed in this disclosure is used.
[0126] For example, the loudness equalization algorithm disclosed herein includes: when processing the current frame, evaluating the peak value of a signal with a certain pre-observation length in the future, and adjusting the current gain according to whether it will exceed the peak value limit.
[0127] For example, this loudness equalization algorithm is calculated based on filtering, K-weighting, and multi-channel fusion. This method can also be directly applied to single-channel loudness equalization. Furthermore, it allows direct signal processing without the need for pre-estimation.
[0128] For example, the loudness equalization algorithm disclosed herein includes: determining the frame length and the length of the pre-observed signal, calculating the loudness of the current frame, and calculating the average loudness of past frames. If the loudness value of the current frame is less than a set threshold, it is determined to be a silent frame or a low-noise frame; if the loudness value is less than a certain level of the average value, the frame is defined as a speech pause frame or a transition frame. In both cases, the target gain for the entire frame is set to 1.
[0129] For example, the loudness equalization algorithm disclosed herein includes: classifying frames that meet a set value as speech frames, and performing gain adjustment point by point within that frame. The sampling points and gain are calculated, and the adjusted loudness is compared with the target value to obtain the error; the gain is adjusted based on the error, and the new gain is smoothed against historical gains to obtain the gain value for the next sampling point.
[0130] For example, different gain limits can be imposed on pull-up and compression. The gain value for pull-up needs to be less than a set value; the gain value for compression should not be greater than the absolute value of the target loudness.
[0131] For example, the gain adjustment process of this disclosure includes using different convergence rates for boosting and compression. Simultaneously, the adjustment step size is proportional to the difference between the current loudness and the target value.
[0132] For example, before the loudness equalization algorithm of this disclosure is output, the adjustment value is subject to peak limit, and if the sampling point exceeds the peak limit, the maximum and minimum value limit is directly applied.
[0133] For example, the current gain value is used to estimate the peak value of the sampling points within the observation range. If the peak value is exceeded, the loudness at the location exceeding the peak value is compared with the set loudness to obtain the error, and the new gain is adjusted accordingly.
[0134] For example, the global linear gain method in this disclosure includes: the high-pass and "K" weighted signal is directly used to calculate the global loudness based on the EBU R.128 standard, the difference between the signal and the target loudness is obtained, and the difference is directly applied to the original audio.
[0135] Figure 4 Block diagrams illustrating some embodiments of the audio processing apparatus of this disclosure are shown.
[0136] like Figure 4 As shown, the audio processing device 4 includes: an estimation unit 41, used to estimate the loudness peak within a preset time length based on the gain corresponding to the current sampling point of the current frame in the audio; a determination unit 42, used to determine the gain corresponding to the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point; and an equalization unit 43, used to perform loudness equalization processing on the current frame using each sampling point of the current frame and its corresponding gain.
[0137] In some embodiments, the determining unit 42 determines the gain corresponding to the next sampling point based on whether the loudness peak value is less than the peak threshold.
[0138] In some embodiments, when the loudness peak value is less than the peak threshold, the determining unit 42 determines the gain corresponding to the next sampling point based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point; when the loudness peak value is greater than or equal to the peak threshold, the determining unit 42 determines the gain corresponding to the next sampling point based on the difference between the loudness peak value and the target loudness.
[0139] In some embodiments, when the loudness peak value is less than the peak threshold, the determining unit 42 determines the target gain of the next sampling point based on the difference between the loudness of the current sampling point and the target loudness, and the gain corresponding to the current sampling point. When the loudness peak value is greater than or equal to the peak threshold, the determining unit 42 determines the target gain of the next sampling point based on the difference between the loudness peak value and the target loudness, and the gain corresponding to the current sampling point. The determining unit 42 also determines the gain corresponding to the next sampling point based on the gain corresponding to the current sampling point and the target gain. The gain corresponding to the next sampling point is less than the target gain.
[0140] In some embodiments, when the loudness peak value is less than the peak value threshold, the determining unit 42 determines whether to use a first convergence rate factor or a second convergence rate factor based on whether the difference between the loudness of the current sampling point and the target loudness is positive or negative. The first convergence rate factor is different from the second convergence rate factor. Based on the first convergence rate factor or the second convergence rate factor and the gain corresponding to the historical sampling points, the gain corresponding to the next sampling point is determined.
[0141] In some embodiments, the determining unit 42 adjusts the loudness of the current sampling point using the gain corresponding to the current sampling point; and determines the gain corresponding to the next sampling point based on the difference between the adjusted loudness and the target loudness.
[0142] In some embodiments, the determining unit 42 determines the candidate gain for the next sampling point of the current frame based on the loudness peak and the gain corresponding to the current sampling point. If the candidate gain does not exceed the gain threshold, the candidate gain is determined as the gain corresponding to the next sampling point. If the candidate gain exceeds the gain threshold, the gain corresponding to the next sampling point is determined based on the gain threshold.
[0143] In some embodiments, the equalization unit 43 adjusts the loudness of the current sampling point using the gain corresponding to the current sampling point. If the loudness of the audio does not exceed the first loudness threshold, the adjusted loudness is used as the output loudness of the current sampling point. If the loudness of the audio exceeds the first loudness threshold, the output loudness of the current sampling point is determined according to the first loudness threshold.
[0144] In some embodiments, the processing device 4 further includes: a determination unit 44, configured to determine whether the current frame is a key frame based on the loudness of the current frame; wherein, if the current frame is not a key frame, the determination unit 42 determines the gain corresponding to all sampling points in the current frame as a preset gain value, and if the current frame is a key frame, determines the gain corresponding to the current sampling point of the current frame.
[0145] In some embodiments, the determining unit 44 determines whether the current frame is a key frame based on the comparison result between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold. The second loudness threshold is calculated based on the average loudness of each frame in the audio.
[0146] In some embodiments, the determining unit 44 determines that the current frame is not a key frame if the loudness of the current frame is less than the second loudness threshold or the loudness of the current frame is less than the third loudness threshold, and determines that the current frame is a key frame if the loudness of the current frame is greater than or equal to the second loudness threshold and the loudness of the current frame is greater than or equal to the third loudness threshold.
[0147] In some embodiments, the loudness range of the audio is greater than or equal to a loudness range threshold, and the loudness of the audio is less than or equal to a fourth loudness threshold.
[0148] In some embodiments, when the loudness range of the audio is less than the loudness range threshold, the equalization unit 43 performs loudness equalization processing on the current frame using the FFMPEG method, and when the loudness of the audio is greater than the fourth loudness threshold, it performs loudness equalization processing on the current frame using the global linear gain method.
[0149] In some embodiments, the estimation unit 41 estimates the loudness peak based on the gain corresponding to the current sampling point in the multi-channel fused audio.
[0150] Figure 5 Block diagrams illustrating other embodiments of the audio processing apparatus of this disclosure are shown.
[0151] like Figure 5 As shown, the audio processing apparatus 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. The processor 52 is configured to execute the audio processing method of any embodiment of this disclosure based on instructions stored in the memory 51.
[0152] The memory 51 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory stores, for example, the operating system, application programs, a boot loader, a database, and other programs.
[0153] Figure 6 Block diagrams illustrating further embodiments of the audio processing apparatus of this disclosure are shown.
[0154] like Figure 6 As shown, the audio processing apparatus 6 of this embodiment includes a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the audio processing method of any of the foregoing embodiments based on instructions stored in the memory 610.
[0155] The memory 610 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory may store, for example, the operating system, application programs, a boot loader, and other programs.
[0156] The audio processing device 6 may also include an input / output interface 630, a network interface 640, and a storage interface 650. These interfaces 630, 640, and 650, as well as the memory 610 and processor 620, can be connected, for example, via a bus 860. The input / output interface 630 provides a connection interface for input / output devices such as a monitor, mouse, keyboard, touchscreen, microphone, and speakers. The network interface 640 provides a connection interface for various networked devices. The storage interface 650 provides a connection interface for external storage devices such as SD cards and USB flash drives.
[0157] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable non-transitory storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc.
[0158] The audio processing method, audio processing apparatus, and non-volatile computer-readable storage medium according to this disclosure have been described in detail above. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions disclosed herein based on the above description.
[0159] The methods and systems of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0160] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. An audio processing method, comprising: Estimate the loudness peak within a preset time period based on the gain corresponding to the current sampling point of the current frame in the audio. Based on the loudness peak value and the gain corresponding to the current sampling point, determine the gain corresponding to the next sampling point of the current frame; Loudness equalization is performed on the current frame using each sampling point and its corresponding gain.
2. The processing method according to claim 1, wherein, The step of determining the gain corresponding to the next sampling point of the current frame based on the loudness peak value and the gain corresponding to the current sampling point includes: The gain corresponding to the next sampling point is determined based on whether the loudness peak value is less than the peak threshold.
3. The processing method according to claim 2, wherein, The step of determining the gain corresponding to the next sampling point based on whether the loudness peak value is less than the peak threshold includes: If the loudness peak value is less than the peak value threshold, the gain corresponding to the next sampling point is determined based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point. If the loudness peak value is greater than or equal to the peak value threshold, the gain corresponding to the next sampling point is determined based on the difference between the loudness peak value and the target loudness.
4. The processing method according to claim 3, wherein, Determining the gain corresponding to the next sampling point includes: If the loudness peak value is less than the peak value threshold, the target gain of the next sampling point is determined based on the difference between the loudness of the current sampling point and the target loudness, as well as the gain corresponding to the current sampling point. If the loudness peak value is greater than or equal to the peak value threshold, the target gain of the next sampling point is determined based on the difference between the loudness peak value and the target loudness, and the gain corresponding to the current sampling point. Based on the gain corresponding to the current sampling point and the target gain, the gain corresponding to the next sampling point is determined, wherein the gain corresponding to the next sampling point is less than the target gain.
5. The processing method according to claim 4, wherein, Determining the gain corresponding to the next sampling point based on the gain corresponding to the current sampling point and the target gain includes: If the loudness peak value is less than the peak value threshold, the first convergence rate factor or the second convergence rate factor is determined based on whether the difference between the loudness of the current sampling point and the target loudness is positive or negative. The first convergence rate factor and the second convergence rate factor are different. The gain corresponding to the next sampling point is determined based on the first convergence rate factor or the second convergence rate factor and the gain corresponding to the historical sampling points.
6. The processing method according to claim 3, wherein, The step of determining the gain corresponding to the next sampling point based on the difference between the loudness of the current sampling point and the target loudness of the current sampling point includes: The loudness of the current sampling point is adjusted using the gain corresponding to the current sampling point; The gain corresponding to the next sampling point is determined based on the difference between the adjusted loudness and the target loudness.
7. The processing method according to claim 1, wherein, Determining the gain corresponding to the next sampling point of the current frame includes: Based on the loudness peak value and the gain corresponding to the current sampling point, determine the candidate gain for the next sampling point of the current frame; If the candidate gain does not exceed the gain threshold, the candidate gain is determined as the gain corresponding to the next sampling point; If the candidate gain exceeds the gain threshold, the gain corresponding to the next sampling point is determined based on the gain threshold.
8. The processing method according to claim 1, wherein, The loudness equalization process for the current frame includes: The loudness of the current sampling point is adjusted using the gain corresponding to the current sampling point; If the loudness of the audio does not exceed the first loudness threshold, the adjusted loudness will be used as the output loudness of the current sampling point. If the loudness of the audio exceeds the first loudness threshold, the output loudness of the current sampling point is determined based on the first loudness threshold.
9. The processing method according to any one of claims 1-8, further comprising: Based on the loudness of the current frame, determine whether the current frame is a keyframe; If the current frame is not a keyframe, the gain corresponding to all sampling points in the current frame is determined as a preset gain value; If the current frame is a keyframe, determine the gain corresponding to the current sampling point of the current frame.
10. The processing method according to claim 9, wherein, The step of determining whether the current frame is a keyframe based on the loudness of the current frame includes: Based on the comparison result between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold, it is determined whether the current frame is a key frame. The second loudness threshold is calculated based on the average loudness of each frame in the audio.
11. The processing method according to claim 10, wherein, The step of determining whether the current frame is a keyframe based on the comparison result between the loudness of the current frame and a preset second loudness threshold and / or a calculated third loudness threshold includes: If the loudness of the current frame is less than the second loudness threshold, or if the loudness of the current frame is less than the third loudness threshold, then the current frame is determined not to be a key frame. If the loudness of the current frame is greater than or equal to the second loudness threshold, and the loudness of the current frame is greater than or equal to the third loudness threshold, the current frame is determined to be a key frame.
12. The processing method according to any one of claims 1-8, wherein, The loudness range of the audio is greater than or equal to the loudness range threshold, and the loudness of the audio is less than or equal to the fourth loudness threshold.
13. The processing method according to any one of claims 1-8, further comprising: If the loudness range of the audio is less than the loudness range threshold, loudness equalization processing is performed on the current frame using FFMPEG. If the loudness of the audio is greater than the fourth loudness threshold, loudness equalization is performed on the current frame using a global linear gain method.
14. The processing method according to any one of claims 1-8, wherein, The estimation of future loudness peaks based on the gain corresponding to the current sampling point of the current frame in the audio includes: The loudness peak is estimated based on the gain corresponding to the current sampling point in the multi-channel fused audio.
15. An audio processing apparatus, comprising: The estimation unit is used to estimate the loudness peak within a preset time length based on the gain corresponding to the current sampling point of the current frame in the audio. The determining unit is configured to determine the gain corresponding to the next sampling point of the current frame based on the loudness peak value and the gain corresponding to the current sampling point; An equalization unit is used to perform loudness equalization processing on the current frame using each sampling point of the current frame and its corresponding gain.
16. The processing apparatus according to claim 15, further comprising: The judgment unit is used to determine whether the current frame is a keyframe based on the loudness of the current frame. Wherein, if the current frame is not a keyframe, the determining unit determines the gain corresponding to all sampling points in the current frame as a preset gain value, and if the current frame is a keyframe, determines the gain corresponding to the current sampling point in the current frame.
17. An audio processing apparatus, comprising: Memory; and A processor coupled to the memory, the processor being configured to perform the audio processing method of any one of claims 1-14 based on instructions stored in the memory.
18. A non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio processing method according to any one of claims 1-14.
19. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-14.
Citation Information
Patent Citations
Method for controlling level of digital audio signal
CN105322904A
Method and apparatus for filtering an audio signal
US20150071463A1