Method and apparatus for normalizing an audio signal

The audio signal processing device normalizes loudness by integrating corrective audio processing steps to achieve consistent volume levels, addressing the 'loudness war' issue and user inconvenience.

JP7812141B2Active Publication Date: 2026-02-09GAUDI AUDIO LAB
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023531070
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-24
Filing Date
2021-11-24
Publication Date
2026-02-09
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The varying loudness levels in audio content due to differing international standards and the 'loudness war' create inconvenience for users who must repeatedly adjust volume, necessitating a method for consistent loudness normalization.

Method used

An audio signal processing device performs loudness normalization by receiving audio signals, integrating loudness information, and applying corrective audio processing steps such as dynamic range control and frequency domain changes to achieve a target volume.

Benefits of technology

This method ensures a consistent target volume across audio content, enhancing user convenience by minimizing volume adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007812141000016
    Figure 0007812141000016
  • Figure 0007812141000017
    Figure 0007812141000017
  • Figure 0007812141000018
    Figure 0007812141000018
Patent Text Reader

Abstract

A method for performing loudness normalization, the method being performed by an audio signal processing device, includes the steps of receiving an audio signal; receiving information about an integrated loudness of the audio signal; receiving information about a target loudness of the audio signal; correcting the integrated loudness based on one or more processing steps to obtain a corrected integrated loudness; and normalizing the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and apparatus for normalizing an audio signal. [Background technology]

[0002] As the method of providing audio to users has shifted from analog to digital, a wider range of volume has become possible. Additionally, the volume of audio signals is becoming more diverse depending on the content they represent. This is because the intended loudness may vary for each piece of audio content during the audio content production process. To address this issue, international standards organizations such as the International Telecommunication Union (ITU) and the European Broadcasting Union (EBU) have issued standards for audio loudness. However, the methods and standards for measuring loudness vary from country to country, making it difficult to apply the standards issued by these international standards organizations.

[0003] Content creators tend to create and provide users with content that is mixed with a relatively louder sound. This is due to the psychological acoustic property that an increase in the acoustic volume of an audio signal is perceived as an improvement in the sound quality of the audio signal. This has led to a competitive structure known as the "loudness war." This can result in loudness differences within a piece of content or between multiple pieces of content, which can be inconvenient for users, as they have to repeatedly adjust the volume of the device on which the content is being played. Therefore, a technology for normalizing the loudness of audio content is needed for the convenience of users of content playback devices. Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to provide a method for providing a constant target volume by loudness normalization. [Means for solving the problem]

[0005] This specification provides a method for performing loudness normalization.

[0006] Specifically, a method for performing loudness normalization, which is performed by an audio signal processing device, includes the steps of: receiving an audio signal; receiving information about integrated loudness of the audio signal; receiving information about target loudness of the audio signal; correcting the integrated loudness based on one or more audio processing steps to obtain a corrected integrated loudness, where the integrated loudness is corrected based on one or more audio processing steps, and the one or more audio processing steps include at least one of processing that changes a spectrum in a frequency domain of the audio signal and dynamic range control (DRC); and normalizing the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal.

[0007] In addition, in this specification, the method performed by the audio signal processing device is characterized in that it further includes a step of receiving order information indicating an order in which the one or more audio processing steps for correcting the integrated loudness are applied.

[0008] Also, in this specification, the method performed by the audio signal processing device further includes a step of receiving a bit flag indicating whether each of the one or more audio processing steps is activated, and the integrated loudness is corrected by the one or more audio processing steps activated based on the bit flag.

[0009] An audio signal processing device that performs loudness normalization includes: a receiving unit that receives an audio signal; and a processor that functionally controls the receiving unit, wherein the processor receives information about an integrated loudness of the audio signal, receives information about a target loudness of the audio signal, and corrects the integrated loudness based on one or more audio processing steps to obtain a corrected integrated loudness, the integrated loudness being corrected based on one or more audio processing steps, the one or more audio processing steps including at least one of processing that changes a spectrum in a frequency domain of the audio signal and dynamic range control (DRC), and further includes a step of normalizing the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal.

[0010] In the present specification, the processor may receive order information indicating an order in which the one or more audio processing steps for correcting the integrated loudness are applied.

[0011] Also, in this specification, the processor receives a bit flag indicating whether or not each of the one or more audio processing processes is activated, and the integrated loudness is corrected by the one or more audio processing processes that are activated based on the bit flag.

[0012] In addition, in this specification, the integrated loudness may be corrected by applying the one or more audio processing processes in an order determined based on the order information.

[0013] In this specification, the order information is set by the same flag value as the bit flag.

[0014] In addition, in this specification, the processing that changes the spectrum in the frequency domain of the audio signal is characterized by including at least one of an equalizer, processing related to user device characteristics, and processing related to the user's cognitive ability.

[0015] In addition, in this specification, the user device characteristics are a frequency band that the user device can output, and the user's cognitive ability is the user's sensitivity to the frequency band.

[0016] In addition, in this specification, any one of the one or more processing is characterized as being non-linear processing.

[0017] In this specification, the nonlinear processing is characterized by being the DRC.

[0018] In addition, in this specification, the normalization of the audio signal is performed based on parameters related to the user's surrounding environment, and the user's surrounding environment is at least one of the volume of noise in the location where the user is located and the frequency characteristics of the noise.

[0019] In addition, in this specification, the target loudness is set based on parameters related to the user's surrounding environment.

[0020] In this specification, the one or more audio processing units are characterized by being at least two. [Effects of the Invention]

[0021] The present invention has the advantage that it is possible to provide an efficient audio signal by using loudness normalization to unify the target volume. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a block diagram showing the operation of an audio signal processing device according to an embodiment of the present invention; [Figure 2] FIG. 2 illustrates a method for normalizing an audio signal according to one embodiment of the present invention. [Figure 3] FIG. 10 illustrates the syntax of metadata including loudness distribution information according to an embodiment of the present invention. [Figure 4] 3 is a flowchart illustrating a method for normalizing an audio signal according to an embodiment of the present invention. [Figure 5] 1 is a block diagram showing a configuration of an audio signal processing device according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0023] The terms used in this specification are currently commonly used and general terms that have been selected as much as possible while taking into consideration the functions of the present invention. However, these terms may change depending on the intentions of engineers in this field, customs, or the emergence of new technologies. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, the meanings thereof will be described in the relevant description of the invention. Therefore, it is clear that the terms used in this specification should be interpreted based on the substantive meanings of the terms and the overall content of this specification, rather than simply on the names of the terms.

[0024] FIG. 1 is a block diagram showing the operation of an audio signal processing device according to an embodiment of the present invention.

[0025] 1, an audio signal processing apparatus may receive an input audio signal. The audio signal processing apparatus may also receive metadata corresponding to the input audio signal and may further receive configuration information (config, state) for the audio signal processing apparatus. In this case, the metadata may refer to syntax components included in a bitstream received by the audio signal processing apparatus.

[0026] The audio signal processing device may correct the loudness of an input audio signal based on a received input audio signal, metadata, setting information, etc., and output the corrected input audio signal as an output audio signal. For example, the audio signal processing device may correct the loudness of the input audio signal and output a normalized audio signal.

[0027] Specifically, the audio signal processing device may include a parser and a decoder. The parser may calculate a deviation and a gain value for correcting an input audio signal using received metadata and setting information. The decoder may correct the input audio signal based on the input audio signal, the deviation, and the gain value, and output the corrected input audio signal. In other words, the decoder may perform a normalization process on the input audio signal based on the input audio signal, the deviation, and the gain value, and output the normalized audio signal as an output audio signal. The deviation may refer to the difference between the loudness of the input audio signal and the output audio signal.

[0028] Furthermore, the audio signal processing device may control the dynamic range of the output loudness level of the normalized audio signal before outputting the normalized audio signal. This is because sound quality distortion due to clipping may occur if the output loudness level for a specific frame of input content falls outside a preset dynamic range. The audio signal processing device may control the dynamic range of the output loudness level based on the preset dynamic range. For example, the audio signal processing device may control the dynamic range of the output loudness level using audio processing such as a limiter and dynamic range control (DRC).

[0029] In the following, a method for normalizing an input audio signal based on the input audio signal, metadata, and setting information received by an audio signal processing device will be specifically described. Also, in this specification, normalization of an audio signal may have the same meaning as loudness normalization.

[0030] The metadata may include information about integrated loudness and target loudness. The integrated loudness represents the average volume of an input audio signal, and the target loudness may refer to the loudness of an output audio signal, i.e., the target loudness that an audio signal processing device aims to output. The target loudness may be defined as a decibel scale with 0.1 dB resolution. Furthermore, loudness in this specification may be expressed in units such as LKFS (Loudness K-Weighted relative to Full Scale) or LUFS (Loudness Unit relative to Full Scale).

[0031] The setting information may include information about audio processing performed for normalizing the input audio signal. The audio processing performed for normalization may be device specific loudness control (DSLC), equal loudness equalizer (ELEQ), equalizer (EQ), or dynamic range control (DRC). That is, the audio signal processing device may normalize the input audio signal using at least one of DSLC, ELEQ, EQ, and DRC. The information about audio processing performed for normalization may include information about the number of audio processing methods used for normalization and an order in which the audio processing methods performed for normalization are applied. The information about the order may be a predetermined order of the audio processing methods applied for loudness normalization. The predetermined order may be the order of four methods: DSLC, ELEQ, EQ, and DRC. On the other hand, if the number of audio processing methods is less than four, the audio signal processing device may normalize the audio signal by taking into account order information about only the audio processing methods actually used. For example, if the order information indicates the order of DSLC -> ELEQ -> EQ -> DRC and only DSLC, EQ, and DRC are used for normalizing the audio signal, the audio signal device can normalize the audio signal based only on the order information of DSLC, EQ, and DRC in the order information. That is, even if the order information for ELEQ is included in the order information, the audio signal processing device can exclude it and normalize the audio signal based on the remaining order of DSLC, EQ, and DRC, i.e., the order of DSLC -> EQ -> DRC. In other words, the information indicating the number of audio processing operations used for normalizing the audio signal may be information indicating whether the audio processing operations are used.The audio signal processing device can know that audio processing for normalizing an audio signal is applied in the order in which it receives information indicating whether each processing is used, and can normalize the audio signal accordingly. The information indicating whether audio processing is used may be in the form of a bit flag. In this case, the order information may be indicated by the same value as the bit flag value indicating whether audio processing is used. Furthermore, the number of audio processing methods used for normalizing an audio signal may be at least two. In this specification, EQ and DRC may be processing performed inside a decoder of the input audio signal processing device or processing performed outside the decoder.

[0032] As used herein, EQ may refer to an effector or processing that changes the frequency characteristics of an input audio signal. That is, EQ processing can emphasize or attenuate a specific frequency band of the input audio signal. As used herein, dynamic range refers to the range between maximum and minimum values ​​of a measured quantity of an input audio signal and may be expressed in decibels (dB). As used herein, DRC may refer to processing that reduces the volume of relatively loud sounds and increases the volume of relatively quiet sounds in an input audio signal, thereby enabling an audio signal processing device to effectively output quiet sounds. As used herein, DSLC may refer to processing that processes an audio signal to reflect the output (reproduction) characteristics of a playback device (e.g., a user device). For example, DSLC may be processing that adjusts loudness by filtering out low-frequency signals that cannot be output (reproduced) by a playback device. Therefore, DSLC may be a method of adjusting the loudness of an audio signal using a bandpass filter (e.g., a high-pass filter) according to the output (reproduction) characteristics of a playback device, taking into account the output (reproduction) characteristics of the playback device (e.g., when low-frequency signal output is not possible) before applying k-weighting used in loudness measurement. ELEQ in this specification may refer to processing of an audio signal that reflects a user's cognitive characteristics depending on the output (volume) size of the audio signal. For example, ELEQ may be an EQ that compensates for changes in a user's sensitivity to high- or low-frequency ranges due to changes in playback level when the playback level is changed by adjusting the output (volume) of the audio signal. User sensitivity refers to the degree to which sounds of the same loudness are perceived differently depending on the frequency, and may be expressed as an equal-loudness contour. In other words, EQ, DSLC, and ELEQ may refer to methods of controlling the frequency spectrum of an input audio signal.

[0033] In addition, the audio signal processing device may receive (input) information related to the surrounding environment (i.e., listening environment) of a user (listener) listening to an output audio signal, and may normalize an input audio signal based on the information related to the user's surrounding environment. For example, the information related to the surrounding environment may refer to the noise level of the user's surrounding environment, frequency characteristics of the noise, and characteristics of the surrounding environment (e.g., indoors, outdoors, etc.). That is, when the noise level of the user's surrounding environment is high, the audio signal processing device may reduce the dynamic range and normalize the loudness of the input audio signal to output an audio signal suitable for the user. Specifically, the surrounding environment may refer to the user's surrounding environment for setting a target loudness or dynamic range for the user to ideally listen to the audio signal. For example, the information related to the surrounding environment may refer to the noise level of the user's surrounding environment and characteristics of the surrounding environment, and may be set as parameter values ​​for setting at least one of the target loudness and the dynamic range.

[0034] The audio processing described herein may modify the spectrum of an input audio signal in the frequency domain and modify the target loudness and dynamic range of the input audio signal in the time domain. In this case, the target loudness and dynamic range may be modified based on information related to a user's surrounding environment. As described above, the information related to a user's surrounding environment may be set using parameter values, and the audio signal processing device may normalize the input audio signal based on the parameter values. Specifically, the audio signal processing device may normalize the audio signal by modifying at least one of the target loudness and the dynamic range based on the parameter values.

[0035] In other words, the audio signal processing device can receive an input audio signal, information about the integrated loudness, and information about the target loudness, apply audio processing to correct the integrated loudness, and normalize the input audio signal based on the corrected integrated loudness and the target loudness, and then output the normalized audio signal.

[0036] FIG. 2 illustrates a method for normalizing an audio signal according to one embodiment of the present invention.

[0037] A method for normalizing an audio signal by the audio signal processing device will be specifically described with reference to FIG.

[0038] Some of the audio processing used for normalizing an audio signal may be non-linear. For example, DRC processing may be non-linear. Because non-linear processing affects the loudness of an output audio signal, the audio signal processing device may normalize the audio signal by taking into account a loudness deviation, which is a deviation between the loudness of an input audio signal and the loudness of an output audio signal, generated by the non-linear processing. However, unlike linear processing, non-linear processing processes an output signal non-linearly, making it difficult to predict the difference between the loudness of the input signal and the loudness of the output signal before the non-linear processing is applied. Therefore, when the audio signal processing device normalizes an audio signal, the order in which the audio processing is applied may be important in order to efficiently predict the loudness deviation generated by the non-linear processing.

[0039] The audio signal processing device may predict loudness deviation for four additional functions (e.g., EQ, DRC, DSLC, and ELEQ) related to the loudness of an audio signal and use the predicted loudness deviation for loudness normalization. The audio signal processing device may receive the frequency characteristics of the EQ and the gain characteristics of the DRC from an external device to predict the loudness deviation.

[0040] Referring to FIG. 2(a), the EQ and DRC for loudness deviation prediction may be applied before the Deviation Estimation Advanced Feature. Referring to FIG. 2(b), the EQ and DRC for loudness deviation prediction may be applied after the Deviation Estimation Advanced Feature. The Deviation Estimation Advanced Feature in FIG. 2 is audio processing used for normalizing the audio signal as described above, and may refer to Device Specific Loudness Control (DSLC), Equal Loudness Equalizer (ELEQ), Equalizer (EQ), and Dynamic Range Control (DRC). That is, referring to FIG. 2(a), the EQ and DRC may be applied for normalizing the audio signal before the remaining audio processing (i.e., DSLC, ELEQ), and referring to FIG. 2(b), the EQ and DRC may be applied for normalizing the audio signal after the remaining audio processing (i.e., DSLC, ELEQ).

[0041] As described above, when an audio signal processing device performs loudness normalization, nonlinear processing may be applied to the audio signal. Specifically, the above-mentioned DRC processing may be applied. Because nonlinear processing affects the loudness of an output audio signal, the audio signal processing device must correct the loudness of the audio signal by taking into account a loudness deviation, which is a deviation between the loudness of an input audio signal and the loudness of an output audio signal, caused by the nonlinear processing. However, unlike linear processing, nonlinear processing processes an output signal nonlinearly, making it difficult to predict the difference between the loudness of the input signal and the loudness of the output signal before the nonlinear processing is applied. Therefore, when an audio signal processing device must process an audio signal in real time, a method is needed to efficiently predict the loudness deviation caused by nonlinear processing. To address this issue, metadata including loudness distribution information of an audio signal included in content may be used.

[0042] FIG. 3 shows the syntax of metadata including loudness distribution information according to an embodiment of the present invention.

[0043] As described above, the metadata used by the audio signal processing apparatus may include information about loudness distribution. For convenience of description, the information about loudness distribution is referred to as loudness distribution information. In this case, the loudness distribution information may be a loudness histogram. Specifically, the loudness distribution information may be a normalized histogram. That is, the loudness distribution information may be a histogram configured with normalized ratios in which the sum of values ​​corresponding to each time interval is 1. In a specific embodiment, the metadata may include loudness distribution information indicating, for each of a plurality of stages divided according to loudness level, a ratio between the amount of audio signal corresponding to each stage among the audio signal and the total amount of the audio signal. In this case, the loudness may be measured within a predetermined time interval. For example, the metadata may indicate the ratio between the number of predetermined time intervals corresponding to each stage and the total number of time intervals. For convenience of description, the ratio between the amount of audio signal corresponding to each stage and the total amount of audio signal is referred to as an audio signal ratio. In addition, the predetermined time interval may be a frame of the audio signal. The loudness distribution information may include information indicating a reference loudness type. The loudness type may be classified according to the length of a time interval during which the loudness is measured. For example, the loudness type may indicate at least one of short-term loudness and momentary loudness. Specifically, the loudness distribution information may have a syntax as shown in FIG. 3.

[0044] In FIG. 3, "type" represents the loudness type indicated by the loudness distribution information. As described above, the loudness type may represent a type according to the length of the time interval in which the loudness of the loudness distribution information is measured. "bsMin" may represent the minimum loudness value coded in the loudness distribution information. "bsMax" may represent the maximum loudness value coded in the loudness distribution information. "bsStep" may represent the size of the loudness step used in the loudness distribution information. "numSteps" may represent the total number of steps included in the loudness distribution information. "bsRatio" may represent the ratio of the amount of audio signal corresponding to each step to the amount of the entire audio signal in the loudness distribution information. Specifically, "bsRatio" may represent a value representing the ratio between the value for each step of the loudness histogram and the sum of the values ​​for all steps. That is, "bsRatio" may represent the audio signal ratio described above. The loudness distribution range may be -90 to 0 LUFS / LKFS.

[0045] The audio signal ratios for each level included in the loudness distribution information may be encoded into a variable-length bitstream. This is because the audio signal ratios for each level may vary significantly. Therefore, when the audio signal ratios are encoded into a variable-length bitstream, the loudness distribution information may be encoded using a much smaller number of bits than when the audio signal ratios are encoded into a fixed-length bitstream. Specifically, audio signal ratios corresponding to each of a plurality of levels may be included in one bitstream. In this case, the bitstream may include an ending flag, which is repeatedly positioned a predetermined number of times and indicates whether it is the last bit of the bits indicating the audio signal ratios corresponding to each level. Specifically, the ending flag may be repeatedly positioned every 8 bits. Furthermore, when the value of the ending flag is a predetermined value, the bit immediately before the ending flag may be the last bit of the audio signal ratio. In this case, the predetermined value may be 1.

[0046] In this specific embodiment, the audio signal processing device can process the bit string including the audio signal ratio for each stage in units of 8 bits. The audio signal processing device acquires 7 bits of the 8 bits as part of the bits indicating the audio signal ratio and the last bit as an ending flag. If the value of the ending flag is a predetermined value, the audio signal processing device acquires the bit indicating the audio signal ratio by combining the previously acquired bits of the audio signal ratio. If the value of the ending flag is not a predetermined value, the audio signal processing device acquires the next 8 bits and repeats the previous operation. The audio signal device can acquire the audio signal ratio from the bit string including the audio signal component for each stage using the syntax of Table 1.

[0047] [Table 1]

[0048] The audio signal processing device can correct the loudness of the audio signal based on the loudness distribution information.

[0049] As described above, the audio signal processing device may perform nonlinear processing on an audio signal. At this time, the audio signal processing device may predict loudness deviations caused by the nonlinear processing based on loudness distribution information and information about nonlinear processing characteristics. At this time, the information about the nonlinear processing characteristics may include at least one of frequency characteristics of the nonlinear processing or gain characteristics of the nonlinear processing. For example, the information about the nonlinear processing characteristics may include frequency characteristics of an equalizer. Furthermore, the information about the nonlinear processing characteristics may include parameters related to DRC (e.g., gain characteristics of the DRC). At this time, the information about the nonlinear processing characteristics may be information received (input) from an external device into the audio signal processing device.

[0050] The audio signal processing device can correct the loudness of the audio signal based on the loudness deviation caused by the nonlinear processing. Specifically, the audio signal processing device can correct the loudness by the difference between the target loudness and a value obtained by adding the loudness of the audio signal to the loudness deviation caused by the nonlinear processing. The audio signal processing device can normalize the input audio signal based on the corrected loudness.

[0051] Depending on the embodiment, the audio signal processor may apply non-linear processing before loudness correction, or the audio signal processor may apply non-linear processing after loudness correction.

[0052] The audio signal processing device can obtain loudness deviation caused by the DRC based on a DRC mapping curve that maps the reference magnitude of the DRC input signal and the reference magnitude of the DRC output signal. The reference magnitude may be an average magnitude, a maximum magnitude (Peak value), or a loudness value calculated based on a certain section expressed in a linear magnitude or logarithmic scale level of the input signal. The integrated loudness of the DRC input audio signal input to the DRC is expressed as L. I and the integrated loudness of the input audio signal input to the audio signal processing device is L I,org Then, the loudness deviation ΔL caused by other audio processing before DRC is prev may be defined as follows: The DRC mapping curve may have values ​​ranging from -127 to 0 dB.

[0053] ΔL prev =L I -L I,org

[0054] The audio signal processing device is prevThe loudness distribution of the input audio signal obtained from the loudness distribution information can be adjusted to the range of the loudness distribution input to the DRC using ΔL. prev may be 0. Specifically, if the loudness distribution of the audio signal is -127 <input DB When defined in <0, the audio signal processor generates a loudness distribution, h DB [k(input DB )] range is -127+ΔL prev <input DB <ΔL prev Specifically, the audio signal processing device can adjust the range of the loudness distribution of the audio signal to the loudness distribution range of the audio signal input to the DRC by the following mathematical formula: That is, the audio signal processing device can perform DRC processing on the loudness distribution and generate a new loudness distribution to be input to the DRC.

[0055]

number

[0056] The audio signal processor is a h DB , prev From DRC mapping curve drc DB [K(input DB After DRC is applied based on the )], the round-to-border distribution of the audio signal, h DB,DRC can be initialized.

[0057]

number

[0058] The audio signal processing device is DB,DRC to DRC output audio signal loudness L DRC,out and from this value, ΔL, the loudness deviation caused by the DRC, is calculated. DRCSpecifically, the audio signal processing device can obtain the average value of the distribution components from which components below an absolute threshold have been removed, and the relative threshold J derived from this value, according to ITU-R Recommendation BS.1770-4. O The average of the above distributions is the loudness L of the DRC output audio signal. DRC,out For example, an audio signal processing device can calculate ΔL by the following mathematical formula: DRC can be obtained.

[0059]

number

[0060] In other words, the audio signal processing device can receive loudness distribution information and information on nonlinear processing characteristics (e.g., parameters related to DRC) from an external device. Then, the audio signal processing device can update (acquire new) the loudness distribution information based on the information on the nonlinear processing characteristics. The audio signal processing device can normalize the input audio signal based on the updated loudness distribution information, information on the integrated loudness, and information on the target loudness.

[0061] The audio signal processor applies a gain value (G loud ) to normalize the audio signal. The audio signal processing device can calculate a gain value to obtain an output audio signal that matches the target loudness. At this time, various methods for normalizing the audio signal based on the gain value will be described. The gain value is calculated based on the target loudness (L T ) and integrated loudness (L I ) may be calculated based on

[0062] i) The audio signal processing device calculates a target loudness (L T ) and integrated loudness (LI ) deviation can be compensated for. T L I If the gain is greater than 1, the audio signal gain value may be greater than 1, causing clipping, and a peak limiter may be applied to prevent this. The gain value may be calculated using the following mathematical formula:

[0063]

number

[0064] ii) The audio signal processor shall provide a limiter in the metadata containing the maximum target loudness L that can be provided without artifacts. QSHI The gain value can be calculated using the following mathematical formula: T and L QSHI In the following mathematical formula, min(x, y) is a function that returns the smaller value of x and y.

[0065]

number

[0066] iii) The audio signal processing device can calculate a gain value based on an offset value for normalizing the audio signal and a reference loudness value. Specifically, the audio signal processing device calculates a gain value based on an offset value for normalizing the audio signal (STLN offset ) and the reference loudness value (L ref ) based on the offset gain (G offset ) can be calculated. Then, the audio signal processing apparatus can correct the offset value for normalizing the audio signal and calculate a gain value based on the corrected offset value. A specific mathematical formula for calculating the gain value is as follows:

[0067]

number

number

[0068] iv) When the audio signal processing device does not use a limiter to normalize the audio signal, the sample peak value smpl ) the gain value can be calculated based on the following formula:

[0069]

number

[0070] v) If the audio signal processing device does not use a limiter to normalize the audio signal, the true peak value true ) the gain value can be calculated based on the true peak. True peak can mean the exact peak of an analog signal that cannot be measured by a sample peak meter. The specific mathematical formula for calculating the gain value is as follows:

[0071]

number

[0072] vi) The audio signal processing device may normalize the audio signal using only an attenuation gain. When only the attenuation gain is used, the gain value may not exceed 1. The specific mathematical formula for calculating the gain value is as follows:

[0073]

number

[0074] The audio signal processing device may normalize an audio signal by applying all or some of audio processing including EQ, DRC, DSLC, and ELEQ. As an example, the audio signal processing device may normalize an audio signal by applying audio processing in the order of DSLC->ELEQ. That is, the audio signal processing device may normalize an audio signal by applying DSLC filtering and ELEQ filtering.

[0075] DSLC filtering

[0076] The DSLC filter may be a filter for ensuring a signal dynamic range rather than a filter that reflects the frequency response of the user device. For example, since the speaker of a mobile device does not have good low-frequency characteristics, the DSLC filter may be a low-cut filter that removes components below 100 Hz. In this case, the DSLC filter may be a finite impulse response (FIR) filter or an infinite impulse response (IIR) filter.

[0077] The audio signal processing device is configured to process an input audio signal (x) when the DSLC filter is a finite impulse response type filter. DSLC [n]) and apply finite impulse response filtering to the filtered output signal (y DSLC The specific calculation method for the output signal is as follows: DSLC denotes the DSLC filter order for the user device, and b DSLC is a DSLC filter coefficient for the user device, which may be a numerator of the finite impulse response filter coefficient and may have a data type of float32.

[0078]

number

[0079] The audio signal processing device is configured to process an input audio signal (x) when the DSLC filter is an infinite impulse response type filter. DSLC [n]) and apply infinite impulse response filtering to the filtered output signal (y DSLC The specific calculation method for the output signal is as follows: DSLC denotes the DSLC filter order for the user device, and b DSLC is a DSLC filter coefficient for the user device, which may be the numerator of the infinite impulse response filter coefficient and may have a data type of float32. DSLC are DSLC filter coefficients for the user device, which may be the denominator of the infinite impulse response filter coefficients and may have a data type of float32.

[0080]

number

[0081] On the other hand, if the DSLC filter is not applied, the DSLC processing is bypassed and the output signal (y DSLC [n]) is the input audio signal (x DSLC [n]).

[0082] ELEQ filtering

[0083] The ELEQ filter adjusts the target loudness (L T ) and user volume (L vol ) and the ELEQ reference loudness (L ELEQ、ref ) signal is 0 dB, it may be a filter that compensates for the difference between the volume tones.

[0084] The audio signal processor uses the filter index (i ELEQ ) filter coefficients (y DSLC [n]) to the input signal (x ELEQ [n]) and perform finite impulse response filtering on the filtered output signal (y ELEQ The output signal (y [n]) can be output. ELEQ [n]) may be calculated as follows: ELEQ [i ELEQ ][k] is a value that is preset according to the frequency of the input audio signal, and the frequency may be 44100 Hz or 48000 Hz. ELEQ [n] is y DSLC May be the same as [n].

[0085]

number

[0086] When the ELEQ filter is not applied, the ELEQ processing is bypassed. ELEQ [n] is x ELEQ May be the same as [n].

[0087] The audio signal processing device calculates the gain value (G loud ) and the signal output after ELEQ filtering, y ELEQ Based on [n], a normalized audio signal (y[n]) can be output. In this case, the normalized audio signal may be calculated using the following mathematical formula:

[0088]

number

[0089] FIG. 4 is a flow chart illustrating a method for normalizing an audio signal according to one embodiment of the present invention.

[0090] Referring to FIG. 4, an audio signal processing device may receive an audio signal (S410). The audio signal processing device may receive information about integrated loudness of the audio signal (S420). The audio signal processing device may receive information about target loudness of the audio signal (S430). The audio signal processing device may correct the integrated loudness based on one or more audio processing processes to obtain a corrected integrated loudness (S440). The integrated loudness may be corrected based on one or more audio processing processes, and the one or more audio processing processes may include at least one of processing that changes a spectrum in a frequency domain of the audio signal and dynamic range control (DRC). The audio signal processing device may normalize the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal (S450).

[0091] The audio signal processing apparatus may receive order information indicating an order in which the one or more audio processing steps for correcting the overall loudness are applied, and the overall loudness may be corrected by applying the one or more audio processing steps in an order determined based on the order information.

[0092] The audio signal device may receive bit flags indicating whether the one or more audio processing steps are activated, and the integrated loudness may be corrected by the one or more audio processing steps activated based on the bit flags. The order information may be set by a flag value equal to the bit flags.

[0093] The processing for modifying the spectrum in the frequency domain of the audio signal may include at least one of an equalizer, processing related to user device characteristics, and processing related to the user's cognitive ability. In this case, the user device characteristics may refer to a frequency band that the user device can output. The user's cognitive ability may refer to the user's sensitivity to a frequency band. Any one of the one or more processing may be non-linear processing. Specifically, the non-linear processing may be the DRC. There may be at least two of the one or more audio processing used for normalizing the audio signal.

[0094] The normalization of the audio signal may be performed based on a parameter related to a user's surrounding environment, where the user's surrounding environment may be at least one of a noise level in a location where the user is located and a frequency characteristic of the noise, and the target loudness may be set based on the parameter related to the user's surrounding environment.

[0095] FIG. 5 is a block diagram showing the configuration of an audio signal processing device according to an embodiment of the present invention.

[0096] The audio signal processing device illustrated in Fig. 5 may be the audio signal processing device illustrated in Fig. 4. Specifically, the audio signal processing device may include a receiving unit that receives information related to an audio signal and a processor that functionally controls the receiving unit. In this case, the processor may perform the method for normalizing an audio signal illustrated in Figs. 1 to 4.

[0097] According to an embodiment, the audio signal processing device 1000 may include a receiving unit 1100, a processor 1200, an output unit 1300, and a storage unit 1400. However, not all of the components shown in FIG. 5 are essential components of the audio signal processing device. The audio signal processing device 1000 may further include components not shown in FIG. 5. Furthermore, at least some of the components of the audio signal processing device 100 shown in FIG. 5 may be omitted. For example, the audio signal processing device according to an embodiment may not include the receiving unit 1100 and the output unit 1300.

[0098] The receiving unit 1100 may receive an input audio signal input to the audio signal processing apparatus 1000. The receiving unit 1100 may receive an input audio signal to be normalized by the processor 1200. Specifically, the receiving unit 1100 may receive input content from an external server via a network. The receiving unit 1100 may also acquire an input audio signal from a storage medium. In this case, the audio signal may include at least one of an Ambisonic signal, an object signal, or a channel signal. The audio signal may be a single object signal or a mono signal. The audio signal may also be a multi-object or multi-channel signal. According to an embodiment, the receiving unit 1100 may include an input terminal for receiving an input audio signal transmitted via a wired connection. The receiving unit 1100 may also include a wireless receiving module for receiving an input audio signal transmitted via a wireless connection.

[0099] According to an embodiment, the audio signal processing apparatus 1000 may include a separate decoder. In this case, the receiving unit 1100 may receive an encoded bitstream corresponding to the input audio signal. The encoded bitstream may be decoded into input content by the decoder. The receiving unit 1100 may also receive metadata associated with the input audio signal.

[0100] According to an embodiment, the receiving unit 1100 may include a transceiver for transmitting and receiving data to and from an external device through a network. In this case, the data may include at least one of a bitstream of an input audio signal and metadata. The receiving unit 1100 may include a wired transceiver terminal for receiving data transmitted via a wired connection. Alternatively, the receiving unit 1100 may include a wireless transceiver module for receiving data transmitted wirelessly. In this case, the receiving unit 1100 may receive data transmitted wirelessly using a Bluetooth or Wi-Fi communication method. The receiving unit 1100 may also receive data transmitted based on a mobile communication standard such as LTE (Long Term Evolution) or LTE-advanced, but the present disclosure is not limited thereto. The receiving unit 1100 may receive various types of data transmitted based on various wired and wireless communication standards.

[0101] The processor 1200 may control the overall operation of the audio signal processing device 100. The processor 1200 may control each component of the audio signal processing device 100. The processor 1200 may perform calculations and processes of various data and signals. The processor 1200 may be embodied as hardware in the form of a semiconductor chip or an electronic circuit, or as software that controls the hardware. The processor 1200 may also be embodied in a form in which the hardware and the software are combined. For example, the processor 1200 may control the operations of the receiving unit 1100, the output unit 1300, and the storage unit 1400 by executing at least one program. The processor 1200 may also perform the operations described above with reference to FIGS. 1 to 4 by executing at least one program.

[0102] According to an embodiment, the processor 1200 may normalize the input audio signal. For example, the processor 1200 may normalize the input audio signal based on audio processing. In this case, the audio signal processing may include at least one of Device Specific Loudness Control (DSLC), Equal Loudness Equalizer (ELEQ), Equalizer (EQ), and Dynamic Range Control (DRC). In this case, any one of the audio processing may be non-linear processing. In addition, the processor 1200 may normalize the input audio signal by applying audio processing according to order information indicating the order in which the audio processing is applied. In addition, the processor 1200 may output the normalized audio signal. In this case, the processor 1200 may output the normalized audio signal from the output unit 1300, which will be described later.

[0103] The output unit 1300 may output a normalized audio signal. The output unit 1300 may output a normalized audio signal obtained by normalizing the input audio signal by the processor 1200. In this case, the output audio signal may include at least one of an Ambisonic signal, an object signal, or a channel signal. The output audio signal may be a multi-object or multi-channel signal. Alternatively, the output audio signal may include a two-channel output audio signal corresponding to each ear of a listener. The output audio signal may include a binaural two-channel output audio signal.

[0104] According to an embodiment, the output unit 1300 may include an output means for outputting output content. For example, the output unit 1300 may include an output terminal for outputting an output audio signal to an external device. In this case, the audio signal processing apparatus 100 may output the output audio signal to an external device connected to the output terminal. The output unit 1300 may include a wireless audio transmission module for outputting the output audio signal to an external device. In this case, the output unit 1300 may output the output audio signal to the external device using a wireless communication method such as Bluetooth or Wi-Fi.

[0105] The output unit 1300 may also include a speaker. In this case, the audio signal processing apparatus 100 can output an output audio signal from the speaker. The output unit 1300 may also include a converter (e.g., a digital-to-analog converter, DAC) that converts a digital audio signal into an analog audio signal. The output unit 1300 may also include a display means that outputs a video signal included in the output content.

[0106] The storage unit 1400 may store at least one of data or programs for processing and control of the processor 1200. For example, it may store various information for performing audio processing (e.g., ELEQ filter coefficients, etc.). The storage unit 1400 may also store results calculated by the processor 1200. For example, the storage unit 1400 may store a signal after DSLC filtering. The storage unit 1400 may also store data input to or output from the audio signal processing device 1000.

[0107] The storage unit 1400 may include at least one memory, which may include at least one type of storage medium selected from the group consisting of a flash memory type, a hard disk type, a multimedia card micro type, a card-type memory (e.g., SD or XD memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.

[0108] Some embodiments may be embodied in the form of a recording medium containing computer-executable instructions, such as program modules, that are executed by a computer. A computer-readable medium may be any available medium accessible by a computer, and may include both volatile and nonvolatile media, and both separate and non-separate media. A computer-readable medium may also include a computer storage medium. A computer storage medium may include any volatile and non-volatile, separate and non-separate media embodied in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data.

[0109] Although the present disclosure has been described above using specific examples, those skilled in the art with ordinary skill in the art to which the present disclosure pertains can make modifications and changes without departing from the spirit and scope of the present disclosure. Therefore, anything that can be easily inferred by a person in the art to which the present disclosure pertains from the detailed description and examples of the present disclosure is deemed to fall within the scope of the present disclosure. [Explanation of symbols]

[0110] 1100 Receiver 1200 processor 1300 Output Unit 1400 Preservation Department

Claims

1. 1. A method for loudness normalization, comprising: The method performed by the audio signal processing device comprises: receiving an audio signal; receiving information regarding the integrated loudness of the audio signal; receiving information regarding a target loudness of the audio signal; receiving order information indicating an order in which the plurality of audio processing steps are to be performed; correcting the integrated loudness using the plurality of audio processing methods in accordance with the order determined based on the order information to obtain a corrected integrated loudness, the plurality of audio processing methods including at least one of processing that changes a spectrum in a frequency domain of the audio signal and dynamic range control (DRC); normalizing the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal.

2. receiving a bit flag indicating whether each of the plurality of audio processings is activated; 2. The method of claim 1, wherein the integrated loudness is corrected by the plurality of audio processings activated based on the bit flags.

3. The method of claim 2 , wherein the order information is set by the same flag value as the bit flag.

4. 2. The method of claim 1, wherein the processing that changes the spectrum in the frequency domain of the audio signal includes at least one of an equalizer, processing related to user device characteristics, and processing related to the user's cognitive ability.

5. The user equipment characteristics are a frequency band that the user equipment can output, The method of claim 4, wherein the user's cognitive ability is the user's sensitivity to a frequency band.

6. 5. The method of claim 4, wherein at least one of the plurality of audio processing processes is non-linear processing.

7. 7. The method of claim 6, wherein the non-linear processing is the DRC.

8. The normalization of the audio signal is performed based on parameters related to the user's surroundings; The method of claim 1 , wherein the surrounding environment of the user is at least one of the volume of noise in a location where the user is located and frequency characteristics of the noise.

9. The method of claim 8 , wherein the target loudness is set based on parameters related to the user's surroundings.

10. 5. The method of claim 4, wherein the plurality of audio processing units is at least two.

11. An audio signal processing apparatus for performing loudness normalization comprises: a receiving unit for receiving an audio signal; a processor for functionally controlling the receiver; The processor: receiving information regarding the integrated loudness of the audio signal; receiving information regarding a target loudness of the audio signal; receiving order information indicating an order in which the plurality of audio processing steps are to be performed; correcting the integrated loudness using the plurality of audio processing methods according to the order determined based on the order information to obtain a corrected integrated loudness; The plurality of audio processing processes include at least one of processing that changes a spectrum in a frequency domain of the audio signal and dynamic range control (DRC); 2. An audio signal processing apparatus comprising: normalizing the audio signal based on the corrected integrated loudness and the target loudness to obtain a normalized audio signal.

12. the processor receives a bit flag indicating whether each of the plurality of audio processings is activated; The audio signal processing apparatus according to claim 11, wherein the integrated loudness is corrected by the plurality of audio processings activated based on the bit flag.

13. The audio signal processing apparatus according to claim 12, wherein the order information is set by the same flag value as the bit flag.

14. The processing for modifying the frequency domain spectrum of the audio signal includes: The audio signal processing device according to claim 11, further comprising at least one of an equalizer, processing related to user device characteristics, and processing related to user cognitive ability.

15. The user equipment characteristics are a frequency band that the user equipment can output, The audio signal processing apparatus according to claim 14, wherein the user's cognitive ability is the user's sensitivity to a frequency band.

16. 15. The audio signal processing apparatus according to claim 14, wherein at least one of the plurality of audio processings is non-linear processing.

Citation Information

Patent Citations

  • Sound recording device

    JP2014078298A

  • Program audio channel number conversion device, broadcast program receiver and program audio channel number conversion program

    JP2016208189A

  • Concept for combined dynamic range compression and induced clipping prevention for audio device

    JP2018151639A

  • Optimizing loudness and dynamic range across different playback devices

    JP2019037011A

  • Loudness control for user interactivity in audio coding systems

    JP2019124953A