Audio signal processing method and apparatus for controlling loudness level
The audio signal processing apparatus addresses the 'loudness war' by using QSHI metadata to normalize loudness levels, ensuring consistent and high-quality audio output across diverse content, enhancing user convenience and adherence to international standards.
Patent Information
- Application Number
- JP2024199667
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-07
- Filing Date
- 2024-11-15
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2040-03-12
AI Technical Summary
The challenge of varying audio loudness levels across different content types and platforms leads to user inconvenience, as users must repeatedly adjust volume settings due to the 'loudness war' caused by content producers aiming for louder sounds, making it difficult to apply international loudness standards consistently.
An audio signal processing apparatus that measures and adjusts loudness levels using Quality Secure Histogram Index (QSHI) metadata, ensuring no perceptual sound quality impairment, by generating and applying loudness metadata to normalize output levels.
Effectively normalizes loudness levels across content, providing a stable and convenient listening experience without degrading sound quality, aligning with user preferences and international standards.
Smart Images

Figure 0007768603000002 
Figure 0007768603000003 
Figure 0007768603000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio signal processing method and apparatus for effectively reproducing an audio signal, and more particularly to an audio signal processing method and apparatus for adjusting the loudness level at which an audio signal of content is output, and providing a user with an audio signal with a more immersive feeling. [Background technology]
[0002] As the method of providing audio to users has shifted from analog to digital, a wider range of volume has become possible. Furthermore, the volume of audio signals is becoming more diverse depending on the content associated with the audio signal. During the audio content production process, the intended loudness may be set differently for each piece of audio content. For this reason, international standards organizations such as the International Telecommunication Union (ITU) and the European Broadcasting Union (EBU) have issued standards for audio loudness. However, since loudness measurement methods and standards vary from country to country, it is difficult to apply the standards issued by international standards organizations.
[0003] Content producers tend to create and provide users with content that is mixed with a relatively louder sound. This is due to the psychological acoustic property that when the sound volume of an audio signal increases, the sound quality of the audio signal is perceived as having improved. This has led to a competitive structure known as the "loudness war." This can lead to loudness differences within content or between multiple pieces of content, which can cause inconvenience to users, who must repeatedly adjust the volume of the device on which the content is being played. Therefore, a technology for normalizing the loudness of audio content is desired for the convenience of users of content playback devices. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of one embodiment of the present invention is to efficiently adjust the output loudness level of content in an audio signal processing method for reproducing content including an audio signal. [Means for solving the problem]
[0005] According to an embodiment of the present invention, an audio signal processing apparatus includes a receiving unit that receives an input audio signal, a processor that generates loudness metadata corresponding to the input audio signal, and an output unit that transmits the loudness metadata generated by the processor. The processor measures the loudness of the input audio signal to obtain loudness information of the input audio signal, converts the loudness information to generate the loudness metadata, and transmits the generated loudness metadata from the output unit to an output device that outputs the input audio signal. The loudness information includes information indicating a Quality Secure Histogram Index (QSHI) of the input audio signal, and the QSHI indicates a threshold loudness level at which no perceptual sound quality impairment occurs.
[0006] The processor may obtain the QSHI based on a loudness histogram of the input audio signal.
[0007] The processor may obtain the loudness histogram based on a distribution of at least one short-term loudness level of the input audio signal, and may obtain the QSHI based on the loudness histogram. The short-term loudness level may be measured over a period shorter than the entire period of the input audio signal.
[0008] The loudness histogram may be a size histogram relating to peak values or root-mean-square (RMS) for each section of the input audio signal.
[0009] The processor can predict loudness parameters when the input audio signal is output according to a target loudness level based on a loudness histogram of the input audio signal, obtain a predicted loudness histogram of the input audio signal based on the predicted loudness parameters, and obtain the QSHI based on the predicted loudness predicted histogram.
[0010] The loudness information may include a cumulative loudness level of the input audio signal, and the QSHI may be greater than the cumulative loudness level of the input audio signal, and the cumulative loudness level may be a loudness level calculated based on loudness measurement values obtained from a setup point in time set in the audio signal processing device.
[0011] The QSHI may be a parameter that is corrected depending on whether or not post-processing is performed on the input audio signal in the output device.
[0012] The processor can set the QSHI so that the short-term loudness level of the entire section of the input audio signal output from the output device is equal to or lower than a previously set level.
[0013] According to another aspect of the present invention, an audio signal processing apparatus includes a processor that adjusts an output loudness level of an input audio signal. The processor may receive loudness metadata corresponding to the input audio signal, parse the loudness metadata to obtain loudness information for the input audio signal, determine a loudness gain for the input audio signal based on the loudness information and a target loudness level, and adjust the output loudness level of the input audio signal based on the loudness gain. The loudness information may include information indicating a Quality Secure Histogram Index (QSHI) of the input audio signal, and the QSHI may indicate a threshold loudness level at which no perceptual sound quality impairment occurs.
[0014] The processor may compare a target loudness level of the input audio signal with the QSHI and determine the loudness gain based on the comparison result.
[0015] The processor may determine the loudness gain based on the smaller of a target loudness level of the input audio signal and the QSHI.
[0016] The processor may receive a cumulative loudness level of the input audio signal, and determine the loudness gain based on the cumulative loudness level of the input audio signal, the QSHI, and the target loudness level. The cumulative loudness level may be a loudness level calculated based on loudness measurements obtained from a setup point in time set in a device that measures the loudness of the input audio signal.
[0017] The QSHI may be a loudness parameter calculated based on a loudness histogram of the input audio signal.
[0018] The loudness histogram may be a size histogram of short-term loudness levels over time of the input audio signal, and the short-term loudness levels may be measured over a period shorter than the entire period of the input audio signal.
[0019] The loudness histogram may be a size histogram relating to peak values or root-mean-square (RMS) for each section of the input audio signal.
[0020] The QSHI may be a parameter calculated based on a predicted loudness histogram predicted from the loudness histogram of the input audio signal, and the predicted loudness histogram may be a histogram generated based on loudness parameters predicted when the input audio signal is output according to the target loudness level.
[0021] The QSHI may be greater than the cumulative loudness level of the input audio signal, and the cumulative loudness level may be a loudness level calculated based on loudness measurements obtained from a setup point in time set in an apparatus that measures the loudness of the input audio signal.
[0022] The processor may generate an output audio signal by adjusting the output loudness level of the input audio signal by the loudness gain, and may output the output audio signal by applying a loudness limiter to limit the loudness level of the output audio signal.
[0023] The QSHI may be a loudness parameter determined based on the number of times a limiter is activated in the audio signal processing device.
[0024] The processor can perform post-processing on the input audio signal, receive post-processing information indicating characteristics of post-processing on the input audio signal, correct the acquired QSHI based on the post-processing information, and determine the loudness gain based on the corrected QSHI.
[0025] The processor can correct the QSHI based on the post-processing information and a previously stored function.
[0026] The processor may correct the QSHI based on the post-processing information and a pre-stored look-up table, which may include information regarding QSHI correction according to post-processing characteristics.
[0027] The information related to the QSHI correction may include information indicating a QSHI correction value according to characteristics of post-processing. The processor may obtain a QSHI correction value corresponding to post-processing of the input audio signal based on the already-stored lookup table, and may correct the QSHI by adding the QSHI correction value to the obtained QSHI.
[0028] The loudness gain may be a fixed gain having a fixed value over the entire duration of the input audio signal.
[0029] The loudness gain may be a time-varying gain over the time that the input audio signal is played.
[0030] The processor may generate an output audio signal by adjusting an output loudness level of the input audio signal by the loudness gain. The QSHI may be a parameter set so that the short-term loudness level of the entire section of the output audio signal is equal to or lower than a predetermined level. [Effects of the Invention]
[0031] The apparatus and method according to an embodiment of the present invention can effectively normalize the loudness level of an audio signal when playing content including the audio signal, and can provide a user with convenience for improving sound quality and adjusting the volume.
[0032] In particular, according to an embodiment of the present invention, it is possible to control a loudness level without impairing sound quality. In addition, an audio signal processing apparatus according to an embodiment of the present invention can provide output content having a more stable output loudness level by using loudness metadata. In addition, it is possible to perform loudness normalization that is close to the loudness actually perceived by a listener. [Brief explanation of the drawings]
[0033] [Figure 1] FIG. 2 is a diagram illustrating loudness levels that change over time while multiple contents are being played according to one embodiment of the present invention.
[0034] [Figure 2] 1 is a schematic diagram illustrating a system including a first audio signal processing device and a second audio signal processing device according to an embodiment of the present invention.
[0035] [Figure 3] 4 is a flowchart illustrating a method for adjusting the loudness level of an input audio signal according to one embodiment of the present invention.
[0036] [Figure 4] 1 is a block diagram specifically illustrating a method for extracting loudness information of an input audio signal by an audio signal processing apparatus according to an embodiment of the present invention;
[0037] [Figure 5]The frequency response of a first-order pre-filter as defined in ITU-R BS.1770-4 is shown.
[0038] [Figure 6] The frequency response of a second-order prefilter is shown.
[0039] [Figure 7] FIG. 2 illustrates a method for generating loudness metadata for an input audio signal by a server according to an embodiment of the present invention.
[0040] [Figure 8] FIG. 10 illustrates how a client outputs an input audio signal using loudness metadata according to an embodiment of the present invention.
[0041] [Figure 9] 10 is a diagram illustrating a histogram of short-term loudness magnitude of an input audio signal according to an embodiment of the present invention;
[0042] [Figure 10] 1 is a block diagram illustrating a system in which an audio signal processing device optimizes the loudness gain of an input audio signal taking into account a target loudness level and perceived sound quality degradation according to an embodiment of the present invention.
[0043] [Figure 11] FIG. 10 is a diagram showing the loudness levels of the input audio signal over time and a fixed gain for the target loudness level. [Figure 12] FIG. 10 is a diagram showing the loudness levels of the input audio signal over time and a fixed gain for the target loudness level.
[0044] [Figure 13] FIG. 2 is a schematic diagram illustrating how the output loudness level of an input audio signal is adjusted according to one embodiment of the present disclosure. [Figure 14]FIG. 2 is a schematic diagram illustrating how the output loudness level of an input audio signal is adjusted according to one embodiment of the present disclosure.
[0045] [Figure 15] 1 is a diagram illustrating a method in which an audio signal processing apparatus according to an embodiment of the present invention acquires loudness information of an input audio signal.
[0046] [Figure 16] 3 is a diagram illustrating a method for adjusting the output loudness level of an input audio signal by an audio signal processing apparatus according to an embodiment of the present invention;
[0047] [Figure 17] 3 is a diagram illustrating a method in which an audio signal processing apparatus according to an embodiment of the present invention adjusts the output loudness level of an input audio signal based on a target loudness range.
[0048] [Figure 18] 2 is a diagram illustrating a method for an audio signal processing device to measure the loudness of input content according to an embodiment of the present invention.
[0049] [Figure 19] 4 is a flowchart illustrating an operation of the audio signal processing device according to an embodiment of the present invention.
[0050] [Figure 20] FIG. 1 is a block diagram showing a configuration of an audio signal processing apparatus 2000 according to an embodiment of the present invention.
[0051] [Figure 21] 4 is a diagram illustrating peak values for each time period of an input audio signal according to an embodiment of the present invention.
[0052] [Figure 22] 1 is a diagram illustrating a method for adjusting the output loudness level of an input audio signal using smoothing in an audio signal processing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] Hereinafter, with reference to the accompanying drawings, embodiments of the present invention will be described in detail so as to be easily understood by those skilled in the art to which the present invention pertains. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. In order to clearly explain the present invention, parts in the drawings that are not relevant to the description will be omitted, and similar parts will be designated by similar reference numerals throughout the specification. Furthermore, when a part "includes" a certain element, this does not mean that it excludes other elements, but that it may further include other elements, unless otherwise specified.
[0054] The present disclosure relates to a method for an audio signal processing device to adjust an output loudness level of input content. In the present disclosure, the input content may be content including an audio signal. In the present disclosure, the input content may be referred to as an input audio signal. Furthermore, loudness may represent the loudness of a sound perceived by the ear. The loudness level may be a numerical value indicating loudness. For example, the loudness level may be expressed in units such as LKFS (Loudness K-Weighted relative to Full Scale) or LUFS (Loudness Unit relative to Full Scale). Furthermore, the loudness level may be expressed in units such as sones or phons.
[0055] The loudness of an audio signal will be described below with reference to FIG. 1. FIG. 1 is a diagram illustrating loudness levels that change over time while multiple contents are being played back according to an embodiment of the present invention. Referring to FIG. 1, time-varying average loudness, short-term loudness, and loudness dynamic range are shown. The average loudness level may be a single loudness value corresponding to one content. The average loudness level may differ for each content (content1, content2, content3). In FIG. 1, solid lines represent the average loudness level for each content (content1, content2, content3). The average loudness in FIG. 1 may represent integrated loudness. The aforementioned integrated loudness and short-term loudness may follow the definitions of loudness standards such as ITU-R BS.1770-4, EBU R 128, EBU TECH 3341, and EBU TECH 3342.
[0056] According to an embodiment, the short-term loudness level may be a loudness level measured over a period shorter than the entire period of the input audio signal. The short-term loudness level may be a loudness measurement value for a portion of content. In this case, the portion of content may be a portion included in one measurement window. The audio signal processing device may obtain multiple short-term loudness levels for one content. Furthermore, the average loudness level may be an average of the multiple short-term loudness levels.
[0057] In FIG. 1, each of the multiple contents to be played and converted has different loudness characteristics. For example, when different contents are converted in a video service platform, advertising content may be inserted between the converted contents. In this case, it may be difficult for an audio signal processing device to maintain a loudness level within a certain range. Also, there may be a large difference in loudness dynamic range between different contents. In such an environment, it may be difficult for an audio signal processing device to provide a loudness level within a range desired by a listener.
[0058] Specifically, when content is converted, a listener may first perceive a sudden change in the loudness level for a short period. This may require the listener to adjust the volume of the device that outputs the audio signal. Furthermore, the listener may need to adjust the volume again to set an appropriate gain according to the average loudness while the converted content is being played. For example, when the converted content is played at a volume adjusted based on the loudness of the initial period of the converted content, a situation may occur in which the loudness level suddenly increases or decreases depending on the content characteristics. If the sudden increase or decrease in the loudness level makes it difficult to understand the content, the listener may need to adjust the volume of the device that outputs the audio signal again.
[0059] For this reason, an audio signal processing device according to an embodiment of the present invention can control an output loudness level of input content to enhance listener convenience. Specifically, the audio signal processing device can adjust the loudness level based on a loudness gain of the input content. In this case, the audio signal processing device can use loudness metadata including loudness information of the input audio signal.
[0060] According to an embodiment of the present invention, the loudness level of input content generated according to different standards or without any specific standard may be normalized based on a target loudness level. Here, the target loudness level may be a loudness level to be output by an audio signal processing device. For example, the target loudness level may be set by a content creator of the input content. In this case, the audio signal processing device may receive information regarding the target loudness along with the input content. The target loudness level may also be set to different values depending on the genre of the input content. In this case, the audio signal processing device may determine the target loudness level based on the genre of the input content. The target loudness level may be set to a default value already stored in the audio signal processing device. In this case, the target loudness level may be set to a value unrelated to the input content or the genre of the input content. The audio signal processing device may adjust the output loudness level of the input content based on the target loudness level.
[0061] According to an embodiment, the audio signal processing device may obtain a loudness gain based on a relationship between the loudness level of the input content and the target loudness level, which may include a difference or ratio between the loudness level of the input content and the target loudness level.
[0062] For example, the audio signal processing device may acquire a loudness gain based on a relationship between a representative loudness level of the input content and a target loudness level. Here, the representative loudness level may be a loudness level that represents the loudness levels for the entire section of the input content. The audio signal processing device may receive the representative loudness level of the input content along with the input content. Alternatively, the audio signal processing device may acquire the representative loudness level based on loudness information analyzed from the input content. In this case, the audio signal processing device may acquire the loudness information based on loudness measurement values for the input content. In the present disclosure, the loudness information of the input audio signal may include loudness metadata converted into a metadata format.
[0063] The audio signal processing device may also adjust an output loudness level of the input content based on the loudness gain. Specifically, the audio signal processing device may apply the loudness gain to the input content to obtain an output audio signal with an adjusted loudness level.
[0064] An audio signal processing device according to an embodiment of the present invention can adjust an output loudness level of an input audio signal by using loudness metadata of the input audio signal, thereby controlling the loudness level of the input content without impairing the sound quality of the input audio signal included in the input content.
[0065] For example, a previously set target loudness level may be greater than a representative loudness level of an input audio signal. In this case, if the input audio signal is output at the previously set target loudness level, sound quality degradation may occur. Therefore, the audio signal processing device may obtain a loudness gain based on the loudness characteristics and the previously set target loudness. The audio signal processing device may obtain a loudness gain that does not cause sound quality degradation of the input audio signal based on the loudness characteristics. The audio signal processing device may adjust the output loudness level of the input audio signal based on the obtained loudness gain.
[0066] In this case, the audio signal processing device can acquire loudness information using loudness metadata of the input audio signal. Specifically, the audio signal processing device can receive loudness metadata of the input audio signal from a device external to the audio signal processing device. The external device can analyze loudness characteristics of the input audio signal and generate loudness metadata of the input audio signal based on the analyzed loudness characteristics. In addition, the external device can transmit the loudness metadata of the input audio signal to the audio signal processing device.
[0067] A method for adjusting the output loudness level of input content according to an embodiment of the present invention will be described below with reference to Fig. 2. Fig. 2 is a schematic diagram showing a system 200 including a first audio signal processing device 210 and a second audio signal processing device 220 according to an embodiment of the present invention. In Fig. 2, the first audio signal processing device 210 may be a server. In Fig. 2, the second audio signal processing device 220 may be a client device.
[0068] Although Figure 2 illustrates the sequence of operations for loudness normalization of input content as being performed by a system with a server-client architecture, the present disclosure is not limited thereto. For example, the sequence of operations illustrated in Figure 2 may be performed by a single audio signal processing device.
[0069] According to an embodiment of the present invention, the first audio signal processing device 210 may generate loudness metadata for an input audio signal. The first audio signal processing device 210 may transmit the generated loudness metadata to the second audio signal processing device 220 that is to output the input audio signal. The second audio signal processing device 220 may receive the loudness metadata from the first audio signal processing device 210. The second audio signal processing device 220 may also adjust the output loudness level of the input audio signal based on the received loudness metadata. Specifically, the second audio signal processing device 220 may determine a loudness gain to be applied to the input audio signal based on the loudness metadata. The second audio signal processing device 220 may also adjust the loudness level of the input audio signal based on the determined loudness gain.
[0070] Specifically, the first audio signal processing device 210 may receive input content. In the present disclosure, the input content may be an input audio signal composed of a plurality of frames. Next, the first audio signal processing device 210 may measure a loudness level of the input content. The first audio signal processing device 210 may obtain a loudness measurement value of the audio signal using a loudness filter based on an auditory scale. Specifically, the loudness filter may be at least one of an inverse filter of equal-loudness contours or a K-weighting filter that approximates the inverse filter.
[0071] For example, the first audio signal processing device 210 may apply a loudness filter to at least a portion of the previously received input content to obtain a loudness measurement value. Here, the portion may be a unit time used to obtain one loudness measurement value. The portion may include at least one frame. In the present disclosure, the unit time used to obtain one loudness measurement value may be referred to as a measurement window.
[0072] The first audio signal processing device 210 may acquire loudness measurement values for each measurement window for the input content. In this case, the acquired loudness measurement values may be instantaneous loudness levels or short-term loudness levels depending on the length of the measurement window. The instantaneous loudness level may be a loudness measurement value measured for a shorter time interval than the short-term loudness level. For example, the length of a measurement window used to acquire one instantaneous loudness level may be 400 milliseconds (ms). Also, the length of a measurement window used to acquire one short-term loudness level may be 3 seconds. However, the present disclosure is not limited thereto. The length of the measurement window for loudness analysis may vary depending on the input content. According to an embodiment, the length of the measurement window may be determined based on additional information of the input content. A method for determining the length of a measurement window by the audio signal processing device will be described below with reference to FIG. 18.
[0073] Next, the first audio signal processing device 210 may acquire loudness information of the input content based on loudness measurements for the input content. The loudness information may include at least one loudness measurement for the input content. The loudness information may also include information calculated based on the loudness measurements for the input content. The first audio signal processing device 210 may update the loudness information in real time. For example, the loudness information may include at least one of a cumulative loudness level, a short-term loudness level, and an instantaneous loudness level. The first audio signal processing device 210 may acquire a cumulative loudness level representing a plurality of loudness measurements accumulated from the start of loudness measurement for the input content to the current time.
[0074] In the present disclosure, the cumulative loudness level may represent a loudness level accumulated from a setup point set in a device for measuring a loudness level. According to an embodiment, the cumulative loudness level may be a loudness level calculated based on loudness measurements measured from a setup point set in the first audio signal processing device 210. For example, the cumulative loudness level may be an average loudness level calculated based on loudness measurements for each section acquired from the setup point. In this case, the loudness measurements for each section may represent either a short-term loudness level or an instantaneous loudness level.
[0075] According to one embodiment, the cumulative loudness level may be obtained based on an average of effective loudness measurements measured between the setup time and the current time, where the effective loudness measurement may be a loudness measurement that satisfies at least one criterion requirement among a plurality of loudness measurements measured between the setup time and the current time.
[0076] For example, the effective loudness measurement value may be a loudness measurement value having a loudness level equal to or greater than a specific level. First, the first audio signal processing device 210 may calculate a first average of loudness measurement values having a loudness level equal to or greater than a first critical value among the plurality of loudness measurement values. In this case, the first critical value may be a value set based on the minimum audible loudness. Next, the first audio signal processing device 210 may calculate a second average of loudness measurement values having a loudness level equal to or greater than a second critical value among the loudness measurement values used in calculating the first average. In this case, the second critical value may be a value obtained by subtracting a predetermined value from the first average. In addition, the first audio signal processing device 210 may use the second average as a cumulative loudness level of the input content. Meanwhile, the first audio signal processing device 210 may reset a setup point for the cumulative loudness level according to specific requirements.
[0077] Next, the first audio signal processing device 210 may generate loudness metadata based on the loudness information. For example, the first audio signal processing device 210 may remove unnecessary information from the loudness information and generate loudness metadata in a syntax form understandable by the second audio signal processing device 220. Furthermore, the first audio signal processing device 210 may generate loudness metadata including additional information related to the input audio signal. The additional information related to the input audio signal may include at least one of information indicating the length, genre, content provider, content creator, popularity, number of views, album, and channel of the input audio signal. In this way, the first audio signal processing device 210 allows another device that outputs the input audio signal to adjust the output loudness level of the input audio signal using the additional information.
[0078] For example, the input audio signal may be a sound source of the same content creator as the already-played audio signal. In this case, the input audio signal and the already-played audio signal may have similar sound characteristics, such as similar style / timbre. Therefore, a device that outputs the input audio signal (e.g., the second audio signal processing device 220) can determine the loudness gain of the input audio signal based on the target loudness level of the already-played audio signal. In this case, the second audio signal processing device 220 can use loudness metadata including additional information.
[0079] Next, the loudness metadata generated by the first audio signal processing device 210 may be stored in a metadata database (hereinafter, referred to as 'DB'). The first audio signal processing device 210 may receive a request for loudness metadata of an input audio signal from the second audio signal processing device 220. In this case, the first audio signal processing device 210 may transmit the loudness metadata of the input audio signal to the second audio signal processing device.
[0080] The second audio signal processing device 220 according to an embodiment of the present invention may acquire loudness information of an input audio signal from the first audio signal processing device 210. Specifically, the second audio signal processing device 220 may request loudness metadata of the input audio signal from the first audio signal processing device 210. Furthermore, the second audio signal processing device 220 may receive the loudness metadata of the input audio signal from the first audio signal processing device 210. The second audio signal processing device 220 may acquire the loudness information of the input audio signal based on the received loudness metadata.
[0081] The second audio signal processing device 220 may acquire a loudness gain to be applied to the input content based on the loudness information. Specifically, the second audio signal processing device 220 may acquire a loudness gain to be applied to a specific frame of the input content based on the loudness information and a target loudness level. According to an embodiment, the second audio signal processing device 220 may acquire a loudness gain to be applied to a specific frame of the input content. The loudness gain to be applied to each frame in a specific section of the input content may be dynamically adjusted over time. The loudness gain to be applied to each frame in sections other than the specific section may be a static gain that is not dynamically adjusted. Furthermore, the loudness gain in a specific section of the input content may be limited to a value within a specific range.
[0082] Next, the second audio signal processing device 220 may adjust the output loudness level of the input content based on the loudness gain. For example, the second audio signal processing device 220 may adjust the output loudness level by applying a loudness gain to the input content. According to an embodiment, the loudness gain may be applied to each frame constituting the input content. In this case, the second audio signal processing device 220 may adjust the output loudness level of the input content by applying the loudness gain to an audio signal corresponding to each frame. The second audio signal processing device 220 may obtain, from the input content, output content whose output loudness level has been adjusted according to the loudness gain. The second audio signal processing device 220 may also output the obtained output content. For example, the second audio signal processing device 220 may play the output content. Alternatively, the second audio signal processing device 220 may transmit the output content to a playback device via a wired / wireless interface.
[0083] Furthermore, the second audio signal processing device 220 can control the dynamic range of the adjusted output loudness level. If the output loudness level for a specific frame of the input content is outside the preset dynamic range, sound quality distortion due to clipping may occur. The second audio signal processing device 220 can control the dynamic range of the output loudness level based on the preset dynamic range. For example, the second audio signal processing device 220 can control the dynamic range of the output loudness level using processing such as a limiter and a dynamic range compressor (DRC).
[0084] 3 is a flowchart illustrating a method for adjusting the loudness level of an input audio signal according to an embodiment of the present invention. For convenience of explanation, a series of operations for adjusting the output loudness level of an input audio signal is described in FIG. 3 as being performed by a single audio signal processing device, but the present disclosure is not limited thereto. For example, some of the operations described in FIG. 3 may be performed by a server and others by a client.
[0085] 3, the audio signal processing device may perform a post-processing operation on an input audio signal. For example, the audio signal processing device may perform at least one of equalization and sound field mode on the input audio signal. In this case, the equalization and sound field mode performed by the audio signal processing device may be operations of a general media playback system.
[0086] In step S303, the audio signal processing device may extract loudness information of the input audio signal. According to an embodiment, when step S301 is performed, in step S303, the audio signal processing device may extract loudness information based on frequency characteristics of post-processing. The audio signal processing device may obtain band-specific loudness level information (weight of post-processing, w_Proc) that changes due to post-processing based on the frequency characteristics of post-processing. Furthermore, the audio signal processing device may extract loudness information using w_Proc.
[0087] For example, when the above-described equalization is performed on the input audio signal, w_Proc may include equalization curve information in the frequency domain. The audio signal processing device may extract loudness information of the input audio signal based on the equalization curve information. When the above-described sound field mode is applied to the input audio signal, w_Proc may include at least one of filter characteristic information and reverb information used in the sound field mode.
[0088] In another embodiment, the environment in which the input audio signal is output may have uneven frequency characteristics and a small response to low frequencies, such as a small speaker used in a mobile phone. In this case, w_Proc may include frequency characteristic information of the output environment. Finally, the audio signal processing device can adjust the output loudness level of the input audio signal based on w_Proc. This allows the audio signal processing device to provide output loudness level adjustment that reflects the characteristics of the device in which the input audio signal is output.
[0089] According to an embodiment of the present disclosure, the loudness information extracted in step S303 may include at least one of integrated loudness information (L_Integ), a quality secure histogram index (QSHI), and a predicted loudness change value (dL_Proc). Here, L_Integ may conform to the ITU-R BS.1770-4 standard. Furthermore, QSHI may represent a threshold loudness level at which no perceptual sound quality impairment occurs due to an output limiter. In the present disclosure, QSHI may include a maximum target loudness (Max_TL). The QSHI may be calculated based on an automatic algorithm or defined by a content creator. A specific method for obtaining the QSHI will be described later with reference to FIG. 4. Furthermore, dL_Proc may be a predicted loudness change value for the input audio signal after post-processing. The audio signal processing device can acquire dL_Proc based on post-processing information set by a user. The audio signal processing device can acquire dL_Proc based on at least one of frequency characteristics of the input audio signal and w_Proc.
[0090] In step S305, the audio signal processing device may determine a loudness gain G_target of the input audio signal. For example, the audio signal processing device may determine the loudness gain G_target based on a previously set target loudness level L_target and the loudness information extracted in step S303. In this case, the previously set target loudness level may be a value set by a user. In step S307, the audio signal processing device may apply a final loudness gain to the input audio signal post-processed in step S301 to output an output audio signal.
[0091] In this case, the output audio signal may be a signal that has passed through a limiter. For example, the audio signal processing device may generate a first output audio signal by applying a final loudness gain to the post-processed input audio signal. Also, the audio signal processing device may generate a second output audio signal by applying a limiter to the first output audio signal. Finally, the audio signal processing device may output the second output audio signal to which the limiter has been applied.
[0092] Hereinafter, a method for extracting loudness information by an audio signal processing device will be described in detail with reference to FIG. 4. FIG. 4 is a block diagram specifically illustrating a method for extracting loudness information of an input audio signal by an audio signal processing device according to an embodiment of the present invention. For convenience of explanation, FIG. 4 illustrates each unit / part performing a respective operation, but the present disclosure is not limited thereto. For example, the operation of each unit / part of the loudness information extraction unit 400 in FIG. 4 may be a series of operations performed by a processor included in the audio signal processing device.
[0093] 4, the loudness information extraction unit 400 may include a loudness measurement unit 401, a frequency-specific loudness analysis unit 402, a post-processing loudness prediction unit 403, and a QSHI extraction unit 404. The loudness information extraction unit 400 may perform the operation described in step S303 of FIG.
[0094] According to one embodiment, the loudness measurement unit 401 may acquire a loudness measurement value of the input audio signal. For example, the loudness measurement unit 401 may acquire at least one of a short-term loudness level and a cumulative loudness level of the input audio signal. Specifically, the loudness measurement unit 401 may acquire cumulative loudness information L_Integ and short-term loudness information L_ShortTerm from the input audio signal through a process such as that described in the ITU-R BS.1770-4 standard.
[0095] According to an embodiment, the frequency-specific loudness analysis unit 402 may acquire a frequency-specific loudness ratio (Multi-band Weight in loudness, WLoud_MB) of the entire input audio signal. For example, the frequency-specific loudness analysis unit 402 may acquire WLoud_MB by applying a K-weighting filter to the input audio signal. The frequency-specific loudness analysis unit 402 may calculate WLoud_MB by frequency-transforming the signal to which the K-weighting filter has been applied.
[0096] A specific method by which the frequency-specific loudness analysis unit 402 calculates WLoud_MB will be described below with reference to Equations 1 to 8.
[0097] [Number 1]
[0098] x_k = filter ( h_kweight, x_in )
[0099] Or,
[0100] x_k = filter ( h_pre2_kweight, filter ( h_pre1_kweight, x_in ) )
[0101] In Equation 1, x_k represents a signal obtained by applying a K-weighted filter to the input audio signal (x_in). In Equation 1, "filter(A,B)" represents an operation of filtering the input audio signal B with the filter coefficient A. In Equation 1, h_kweight may represent a single K-weighted filter. Also, h_pre2_kweight and h_pre1_kweight may represent a first-order pre-filter and a second-order pre-filter, respectively, as defined in ITU-R BS.1770-4. The frequency-specific loudness analysis unit 402 may filter the input audio signal with the K-weighted filter coefficients and apply them. FIG. 5 represents the frequency response of the first-order pre-filter defined in ITU-R BS.1770-4. Also, FIG. 6 represents the frequency response of the second-order pre-filter.
[0102] The frame-by-frame signal of signal x_k obtained from Equation 1 may be expressed as Equation 2. In Equation 2, x_frame[l] represents the l-th frame signal of signal x_k, where NF represents the frame length and NH represents the hop size.
[0103] [Number 2]
[0104] x_frame[l] = x_k[ ((l-1)*NH+1) : ((l-1)*NH+NF) ]
[0105] Referring to Equation 3, the frequency-specific loudness analysis unit 402 may obtain xw_frame[l][-] by windowing x_frame[l]. In this case, the frequency-specific loudness analysis unit 402 may obtain xw_frame[l][-] using a rectangular window function in which all coefficients of the window function are 1. Alternatively, the frequency-specific loudness analysis unit 402 may obtain xw_frame[l][-] using various window functions such as a Hamming window function or a Hanning window function. The window operation may be an operation for frequency analysis of the input audio signal. In Equation 3, wind[n] represents the n-th coefficient of the window function, and n may be the sample number of the window. For example, if the NF is 512, the value of n may be any one of 1 to 512.
[0106] [Number 3]
[0107] xw_frame[l][n] = x_frame[l][n] * wind[n] for n=1, 2, …, NF
[0108] Furthermore, the frequency-specific loudness analysis unit 402 can perform a Discrete Fourier Transform (DFT) on xw_frame[l][-]. The frequency domain signal (XW_frame[l]) obtained by performing a Discrete Fourier Transform on xw_frame[l][-] may be expressed as in Equation 4. In Equation 4, DFT{x} represents the Discrete Fourier Transform of the time domain signal 'x'.
[0109] [Number 4]
[0110] XW_frame[l] = DFT { xw_frame[l][1:NF]}
[0111] Next, referring to Equation 5, the frequency-specific loudness analysis unit 402 can obtain the power for each frequency bin of the converted frequency signal XW_frame[l]. In Equation 5, P_frame_bin[l][k] represents the power in the k-th frequency bin of the l-th frame. Also, conj(x) represents the conjugation function of 'x'.
[0112] [Number 5]
[0113] P_frame_bin[l][k] = XW_frame[l][k] * conj(XW_frame[l][k]) for k=1, 2, …, NF
[0114] Next, referring to Equation 6, the frequency-specific loudness analysis unit 402 may map P_frame_bin[l][k] to the previously set frequency bands to obtain the frequency band-specific power (P_frame_band[l][b]) of the l-th frame. In Equation 6, band[b] represents the index of the start frequency bin of the b-frequency band. That is, the frequency-specific loudness analysis unit 402 may obtain the frequency band-specific power by summing the frequency bin-specific powers from band[b] to band[b+1]-1. In Equation 6, sum_{y}(x) may represent the sum of the function 'x' index having index k as a factor. In this case, 'y' may represent the range of the index for this calculation.
[0115] [Number 6]
[0116] P_frame_band[l][b]
[0117] = sum_{k from band[b] to band[b+1]-1} (P_frame_bin[l][k])
[0118] Referring to Equation 7, the frequency-specific loudness analysis unit 402 can obtain the frequency band-specific power (P_band[b]) for the entire section of the input audio signal based on the frequency band-specific power (P_frame_band[l][b]) for the l-th frame. The frequency-specific loudness analysis unit 402 can obtain the frequency band-specific power (P_band[b]) for the entire section of the input audio signal by summing the frequency band-specific powers (P_frame_band[l][b]) obtained for each frame for the same frequency band. In Equation 7, NumberOfFrames represents the number of all frames. Furthermore, l, representing a frame index, is defined within the range from 1 to NumberOfFrames.
[0119] [Number 7]
[0120] P_band[b] = sum_{l from 1 to NumberOfFrames} (P_frame_band[l][b])
[0121] Next, referring to Equation 8, the frequency-specific loudness analysis unit 402 can obtain a frequency band-specific loudness ratio (WLoud_MB[b]) based on the frequency band-specific power (P_band[b]). Specifically, the frequency-specific loudness analysis unit 402 can normalize the specific frequency band-specific power (P_band[b]) based on the sum of all frequency band-specific powers. In Equation 8, NumberOfBands represents the total number of divided frequency bands. Furthermore, b, representing a band index, is defined within the range from 1 to NumberOfBands.
[0122] [Number 8]
[0123] WLoud_MB[b] = P_band[b] / [sum_{b from 1 to NumberOfBands} (P_band[b])]
[0124] WLoud_MB[b] calculated from Equation 8 represents the ratio of cumulative loudness levels for each frequency band of the input audio signal. For example, if the input audio signal is a two-band signal and the cumulative loudness level of the input audio signal is L_Integ=-20 LKFS, WLoud_MB
[10] =0.8 and WLoud_MB[1]=0.2. In this case, the loudness level for the first frequency band of the input audio signal may be predicted to be -20+10*log10(0.8)=-20.97 LKFS, and the loudness level for the second frequency band may be predicted to be -20+10*log10(0.2)=-26.99 LKFS.
[0125] In one embodiment, the post-processing loudness prediction unit 403 can obtain a loudness change prediction value based on at least one of the band-specific loudness level information (w_Proc) that changes due to post-processing and the frequency-specific loudness ratio (WLoud_MB) of the entire input audio signal.
[0126] In this case, the post-processing loudness prediction unit 403 can use the frequency-specific loudness ratio (WLoud_MB) of the entire input audio signal acquired from the frequency-specific loudness analysis unit 402. In addition, the band-specific loudness level information (w_Proc) that changes due to post-processing may be acquired based on the characteristics of post-processing for the input audio signal. The characteristics of post-processing for the input audio signal may be determined based on information input by the user.
[0127] Specifically, equalization set by a user may be applied to the input audio signal, and the frequency band gain of the equalization may be set to w_ProcBand_dB in decibel units for each of the NumberOfBands frequency bands, and the total gain of the equalization may be set to w_ProcGain_dB. In this case, the frequency-specific loudness analysis unit 402 may obtain the loudness ratio for each frequency band based on the frequency band gain (w_ProcBand_dB) and the total gain (w_ProcGain_dB). The calculation method used by the frequency-specific loudness analysis unit 402 to obtain the loudness ratio for each frequency band may be expressed as Equation 9.
[0128] [Number 9]
[0129] w_Proc[b] = 10^((w_ProcBand_dB[b] + 0.5*w_ProcGain_dB) / 10)
[0130] for 1= <b=<NumberOfBands
[0131] Furthermore, the method by which the post-processing loudness prediction unit 403 obtains the loudness change predicted value dL_Proc can be expressed as in Equation 10.
[0132] [Number 10]
[0133] dL_Proc = 10 * log10 ( sum_{b from 1 to NumberOfBands} (WLoud_MB[b] * w_Proc[b]) )
[0134] According to one embodiment, the QSHI extraction unit 404 can extract a quality assurance histogram index QSHI based on the short-term loudness information L_ShortTerm. As described above, the quality assurance histogram index (hereinafter, referred to as 'QSHI') may be a threshold loudness level at which no perceptual impairment of sound quality occurs. The QSHI extraction unit 404 can acquire the QSHI based on the short-term loudness information L_ShortTerm acquired from the loudness measurement unit 401.
[0135] For example, the QSHI extraction unit 404 may acquire the QSHI by analyzing the short-interval loudness information L_ShortTerm. In this case, the short-interval loudness information L_ShortTerm may include one or more short-interval loudness levels of the input audio signal. Specifically, the QSHI extraction unit 404 may acquire a short-interval loudness magnitude histogram of the input audio signal based on the one or more short-interval loudness levels. Furthermore, the QSHI extraction unit 404 may acquire the QSHI of the input audio signal based on the acquired short-interval loudness magnitude histogram.
[0136] Hereinafter, a specific method in which the QSHI extraction unit 404 extracts QSHI from the short-term loudness information L_ShortTerm of the input audio signal will be described with reference to Equations 11 and 12. In Equation 11, L_ShortTerm_Sorted represents information in which one or more short-term loudness levels included in the short-term loudness information L_ShortTerm of the input audio signal are sorted in descending order. For example, the QSHI extraction unit 404 can sort one or more short-term loudness levels in descending order.
[0137] [Number 11]
[0138] L_ShortTerm_Sorted = sort ( L_ShortTerm, 'descending' )
[0139] Furthermore, the QSHI extraction unit 404 may acquire a loudness level corresponding to a previously set index from among one or more short-term loudness levels of the input audio signal based on L_ShortTerm_Sorted. In Equation 12, EffectiveIndex may represent a previously set effective index. Specifically, the previously set effective index (EffectiveIndex) may indicate one or more short-term loudness levels of the input audio signal in a previously set order of magnitude. That is, the QSHI extraction unit 404 may acquire the short-term loudness level that is the EffectiveIndex-th largest from among one or more short-term loudness levels of the input audio signal. In this case, the short-term loudness level that is the EffectiveIndex-th largest from among one or more short-term loudness levels of the input audio signal may be referred to as the effective short-term loudness level (L_ShortTerm_Effective) of the input audio signal.
[0140] [Number 12]
[0141] L_ShortTerm_Effective = L_ShortTerm_Sorted[EffectiveIndex]
[0142] Next, the QSHI extraction unit 404 can obtain the QSHI based on at least one of the effective short-term loudness level (L_ShortTerm_Effective) and the cumulative loudness level of the input audio signal, and the QSHI may be greater than or equal to the cumulative loudness level.
[0143] Furthermore, the QSHI extraction unit 404 may acquire an effective short-term loudness level (L_ShortTerm_Effective_Shift) to be changed when the input audio signal is output according to a previously set target loudness level. Specifically, the QSHI extraction unit 404 may predict the short-term loudness information (L_ShortTerm_Shft) to be changed based on the short-term loudness information L_ShortTerm of the input audio signal. In this case, the short-term loudness information (L_ShortTerm_Shft) to be changed may include one or more short-term loudness levels to be changed when the input audio signal is output according to the previously set target loudness level. In this case, the QSHI extraction unit 404 may acquire the QSHI based on the acquired L_ShortTerm_Effective_Shift. For example, the QSHI may be a maximum allowable target loudness value when the L_ShortTerm_Effective_Shift[EffectiveIndex] short-term loudness level is limited to be equal to or less than a threshold value.
[0144] For example, the L_ShortTerm_Effective_Shift of the input audio signal may be used as a threshold (L_Threshold) for the short-term loudness level. The QSHI extraction unit 404 may correct the maximum allowable target loudness value based on the L_ShortTerm_Effective_Shift. The QSHI extraction unit 404 may use the corrected maximum allowable target loudness value as the QSHI value. Alternatively, the QSHI extraction unit 404 may select the larger value of the maximum allowable target loudness value corrected in the above manner and the cumulative loudness of the input audio signal as the QSHI value.
[0145] By using the above method, the audio signal processing device can effectively prevent degradation of the sound quality of the input audio signal due to the limiter, since the sound quality may be degraded by the limiter in the portion of the input audio signal where the volume is set relatively high.
[0146] According to one embodiment, QSHI may be a value set such that the number of short-duration loudness levels greater than a specific value among one or more short-duration loudness levels of an input audio signal is smaller than EffectiveIndex. In this case, EffectiveIndex may be a value determined based on characteristics of a limiter of the audio signal processing device. For example, EffectiveIndex may be changed depending on the degree of sound quality degradation caused by operation of the limiter. Furthermore, the short-duration loudness threshold (L_Threshold) may be a value determined based on characteristics of the limiter of the audio signal processing device. For example, the short-duration loudness threshold (L_Threshold) may be changed depending on the degree of sound quality degradation caused by operation of the limiter.
[0147] According to a specific embodiment, the input audio signal may have a relatively large dynamic range. For example, the cumulative loudness level of the input audio signal may be L_Integ=-24LKFS, and the effective short-term loudness level may be calculated as L_ShortTerm_Effective=-10LKFS. In this case, when EffectiveIndex=10 and the short-term loudness threshold=-7LKFS, the QSHI may be calculated as -21LKFS.
[0148] In the above-described embodiment, the QSHI of an input audio signal is extracted based on a histogram of short-term loudness magnitudes. However, the present disclosure is not limited to this. For example, the QSHI of an input audio signal may be defined as a value arbitrarily set by a producer of content including the input audio signal or an operator of a sound system that outputs the input audio signal. In addition to the short-term loudness level, the audio signal processing device may acquire the QSHI by performing a histogram analysis on at least one of the peak envelope and RMS of the input audio signal.
[0149] According to an embodiment, the QSHI of an input audio signal may change depending on changes in the short-term loudness magnitude histogram. For example, the aforementioned short-term loudness magnitude histogram may change depending on whether or not post-processing is performed, as determined by a user's input. In this case, the QSHI of the input audio signal may be changed to another value based on a pre-set table. Alternatively, the QSHI of the input audio signal may be changed to a value calculated based on the characteristics of the post-processing.
[0150] Furthermore, a method for an audio signal processing device according to an embodiment of the present disclosure to determine a loudness gain of an input audio signal based on the above-described loudness information will be described. Equation 13 represents the modified cumulative loudness level (L_IntegProc) of the input audio signal when a post-processing process is performed on the input audio signal. The audio signal processing device may obtain the modified cumulative loudness level (L_IntegProc) of the input audio signal based on the loudness change prediction value dL_Proc due to post-processing. Referring to Equation 13, the audio signal processing device may obtain the modified cumulative loudness level (L_IntegProc) by adding the loudness change prediction value dL_Proc due to post-processing to the cumulative loudness level of the input audio signal.
[0151] [Number 13]
[0152] L_IntegProc = L_Integ + dL_Proc
[0153] The audio signal processing device can calculate a loudness gain for adjusting the output loudness level based on the above-mentioned QSHI, the previously set target loudness level (L_Target), and the cumulative loudness level changed by post-processing.
[0154] In the above-described embodiment, the target loudness level (L_Target) may be a value set by a user. However, the present disclosure is not limited thereto. For example, the pre-set target loudness level (L_Target) may be a default value provided by a playback system that outputs the input audio signal. Alternatively, the pre-set target loudness level (L_Target) may be a value set based on the playback environment that outputs the input audio signal. The audio signal processing device may apply a loudness gain (G_Target) to a first intermediate audio signal that has been post-processed from the input audio signal. For practical convenience of implementation, the post-processing process may be performed after the loudness gain (G_Target) is applied to the input audio signal before post-processing. Furthermore, the audio signal processing device may pass the second intermediate audio signal to which the loudness gain (G_Target) has been applied through a limiter and output the second intermediate audio signal.
[0155] Meanwhile, multimedia streaming services are currently widely used in the media market. A system providing a multimedia streaming service may generally include a server that stores content to be streamed and a user device (i.e., a client). At this time, the multimedia streaming service may be provided on the client side in the form of in-application playback or web playback. Each of the server and the client may be an audio signal processing device that performs the operations described in this disclosure. In such a server-client architecture, the server may analyze input content and provide loudness information. The client may adjust the output loudness level of the input content based on the loudness information provided by the server. Specifically, the server may transmit loudness metadata including loudness information of an input audio signal to the client. The client may receive the loudness metadata of the input audio signal from the server. The client may also obtain a loudness gain to be applied to the input audio signal based on the loudness metadata of the input audio signal.
[0156] FIG. 7 illustrates a method for generating loudness metadata of an input audio signal by a server according to an embodiment of the present invention. The server according to an embodiment of the present invention may encode the input audio signal and generate and / or output an audio stream. The server according to an embodiment of the present invention may extract loudness information from the input audio signal. For example, the server of FIG. 7 may perform the operations described with reference to the loudness information extraction (step S303) of FIG. 3 and the operations described with reference to the loudness information extraction unit 400 of FIG. 4. The server may also generate loudness metadata including the extracted loudness information. The server may output the generated loudness metadata to an external device. For example, the server may transmit the generated loudness metadata to a client in the form of a metadata stream.
[0157] 8 is a diagram illustrating a method in which a client according to an embodiment of the present invention outputs an input audio signal using loudness metadata. The client according to an embodiment of the present invention may receive an audio stream. The client may also decode the received audio stream to obtain an input audio signal. The client may perform a post-processing process on the input audio signal. In this case, whether or not to perform the post-processing process and the characteristics thereof may be determined based on an input received from a user or a setting value already stored in the system.
[0158] A client according to an embodiment of the present invention may determine a loudness gain of an input audio signal based on loudness metadata of the input audio signal. For example, the client may receive loudness metadata in the form of a metadata stream. The client may acquire loudness information of the input audio signal by parsing the loudness metadata of the input audio signal. Specifically, the client may acquire at least one of WLoud_MB, L_Integ, and QSHI, as described above in FIGS. 3 and 4, from the loudness metadata of the input audio signal. The client may determine a loudness gain of the input audio signal based on the acquired loudness information. The client may adjust an output loudness level by applying a loudness gain to the input audio signal. The client may generate an output audio signal by applying a limiter to the intermediate audio signal whose output loudness level has been adjusted. The client may also output the output audio signal.
[0159] In one embodiment, the client of FIG. 8 can perform the operations described with reference to the post-processing (step S301), loudness gain determination (step S305), and loudness gain application (step S307) of FIG. 3, and the operations described with reference to the post-processing loudness prediction unit 403 of FIG. 4.
[0160] Meanwhile, music content may have various loudness levels depending on the era and / or genre. For example, the cumulative loudness level of classical music is relatively low to provide a wide dynamic range, while the cumulative loudness level of popular music from the 2000s is relatively high. Specifically, the cumulative loudness level of popular music from the 2000s may be -13 to -8 LKFS, and the cumulative loudness level of a quieter movement of classical music may be about -30 LKFS.
[0161] When determining the target loudness level, it is possible to utilize -23 to -24 LKFS defined in the broadcasting standard. However, this may not provide sufficient volume to overcome external noise in a noisy environment such as a subway. For this reason, an audio signal processing device according to an embodiment of the present invention can determine different target loudness levels depending on the playback environment. When the target loudness level of popular music from the 2000s is set to -10, the volume of the popular music from the 2000s may not change significantly. In contrast, when the target loudness level of music with a relatively low integrated loudness level, such as classical music or music from the 1970s and 1980s, is set to -10, the volume may change significantly.
[0162] 9 is a diagram illustrating a histogram of short-term loudness magnitudes of an input audio signal according to an embodiment of the present invention. In the embodiment illustrated in FIG. 9, the genre of the input audio signal may be classical. Also, in the embodiment illustrated in FIG. 9, the cumulative loudness of the input audio signal may be -21 LKFS. For example, the target loudness level of the input audio signal may be L_Target=-10 LKFS. In this case, the short-term loudness magnitude histogram shifts +11 LKFS to the right. At this time, a segment having a short-term loudness level greater than -7 LKFS occurs.
[0163] According to one embodiment, sound quality degradation due to the limiter may occur in sections having short-term loudness levels greater than -7LKFS. For this reason, an audio signal processing device according to one embodiment of the present invention may perform loudness normalization of an input audio signal based on QSHI as described above. In this case, although loudness normalization performance may be relatively reduced, a best-effort method may be used that performs the most aggressive adjustment within a range that prevents sound quality degradation.
[0164] According to an embodiment of the present invention, an audio signal processing device can use a loudness gain correction method that most closely approximates a target loudness level based on loudness information of an input audio signal, and can provide equalization without changing the loudness level by using the method.
[0165] Equalization refers to adjusting the frequency energy of an input audio signal to achieve a desired tone color. Depending on the degree of adjustment of the input audio signal, the overall energy may increase. In this case, the input audio signal may be clipped. Furthermore, a limiter may impair the sound quality compared to the input audio signal. Therefore, an audio signal processing device according to an embodiment of the present invention may set a previously set target loudness level (L_Target), cumulative loudness level (L_Integ), and QSHI to the same arbitrary value. In this case, the loudness gain (G_Target) of the input audio signal may be expressed as in Equation 14. That is, the audio signal processing device can obtain a linear loudness gain (G_Target) because the target loudness level (L_Target), cumulative loudness level (L_Integ), and QSHI cancel each other out.
[0166] [Number 14]
[0167] G_Target = power ( 10, -dL_Proc) / 20
[0168] The audio signal processing device may apply the loudness gain (G_Target) of Equation 14 to the input audio signal. The audio signal processing device may compensate for loudness changes due to post-processing and provide an output loudness level that is the same as the loudness level of the input audio signal. The audio signal processing device may compensate for loudness changes due to post-processing and maintain the loudness level of the input audio signal. The audio signal processing device may set the loudness level of the intermediate audio signal to be the same as the loudness level of the input audio signal using the predicted loudness change due to post-processing. In this case, the intermediate audio signal may be a signal post-processed from the input audio signal. This means that the audio signal processing device provides the intermediate audio signal with the same loudness level as the original input audio signal, even though its tone is changed compared to the input audio signal through the post-processing process. Meanwhile, the predicted loudness change due to post-processing may be obtained by the method described above with reference to FIGS. 3 and 4. The predicted loudness change due to post-processing may be obtained based on a WLoud_MB provided by analysis or a WLoud_MB based on content characteristics.
[0169] 10 is a block diagram showing a system in which an audio signal processing device optimizes the loudness gain of an input audio signal in consideration of a target loudness level and perceived sound quality degradation according to an embodiment of the present invention. The audio signal processing device can determine a target loudness gain that a dynamic processor can accept based on the target loudness level and loudness information of the input audio signal. Here, the dynamic processor can represent a processing process that clips a signal according to the loudness level, like the limiter or compressor described above. The loudness information of the input audio signal can include at least one of a cumulative loudness level, a short-term loudness level, an instantaneous loudness level, a sample peak, a true peak, a loudness range, and RMS (root-mean-square).
[0170] A specific example in which an audio signal processing device determines a loudness gain of an input audio signal will be described below. According to one example, the maximum target loudness level that can be set by a user may be −10 LKFS, and the cumulative loudness of the input audio signal may be −22 LKFS. Furthermore, the short-term loudness level corresponding to the tenth of the multiple short-term loudness levels of the input audio signal may be −18 LKFS. In this case, the tenth short-term loudness level may be a specific example of the effective short-term loudness level (L_ShortTerm_Effective) described with reference to the QSHI extraction unit 404 of FIG. 4 . That is, −18 LKFS may be used as an index for determining whether or not sound quality has deteriorated due to DRC. When the maximum target loudness level is −10 LKFS, the maximum amplification amount may be 12 LU (Loudness Unit). In this case, the audio signal processing device can acquire the QSHI based on the tenth short-term loudness level amplified by the maximum amplification amount.
[0171] The audio signal processing device may compare a pre-set target loudness level input by a user with the QSHI. The audio signal processing device may determine a loudness gain of the input audio signal based on the comparison result. For example, the audio signal processing device may determine a loudness gain of the input audio signal based on a relatively smaller value of the pre-set target loudness level or the QSHI. In the above-described embodiment, the short-term loudness level for determining the indicator for determining the presence or absence of DRC sound quality degradation is selected as the top 10 in descending order, but the present disclosure is not limited thereto. In addition to the short-term loudness level, the audio signal processing device may obtain the QSHI by performing a histogram analysis on at least one of the peak value and RMS of the signal.
[0172] 11 and 12 are diagrams illustrating the loudness levels of time-varying input audio signals and fixed gains for target loudness levels. FIG. 11 illustrates a fixed gain for adjusting the loudness level of a first input audio signal having a loudness distribution smaller than the target loudness level to the target loudness level. In this case, the first input audio signal may be clipped in sections greater than 0 dBFS, resulting in excessive timbre distortion. As such, there are limitations to using a loudness level adjustment method using a fixed gain to obtain a value close to the target loudness level. For this reason, the audio signal processing device may apply a gain smaller than the fixed gain value to sections (2) and (4) of the first input audio signal.
[0173] 12, the second input audio signal has a larger dynamic range than the first input audio signal of FIG 11. Therefore, when the audio signal processing device applies a fixed gain for a target loudness level to the second input audio signal, the loudness level may be relatively low in some sections. Therefore, the audio signal processing device may apply a gain larger than the fixed gain value to sections (1) and (3) of the second input audio signal.
[0174] According to a further embodiment, the audio signal processing device may apply a gain boost. For example, the audio signal processing device may acquire a target loudness range. The audio signal processing device may set an additional gain for each section of the input audio signal based on the acquired target loudness range. Specifically, the audio signal processing device may apply the set additional gain to a section having a loudness level outside the target loudness range among all time sections of the input audio signal.
[0175] As described above, an audio signal processing device according to an embodiment of the present disclosure may adjust an output loudness level of an input audio signal by applying a time-varying gain to the input audio signal. The audio signal processing device may adjust the output loudness level of the input audio signal based on loudness metadata of the input audio signal. In this case, the loudness metadata of the input audio signal may include information that changes over time. In order to apply a time-varying gain, the audio signal processing device may normalize the output loudness level of the input audio signal according to a target loudness level or a target loudness range by referring to the time-varying metadata. As a result, in the present disclosure, the audio signal processing device may solve the above-mentioned problem when applying a fixed gain to the input audio signal for loudness normalization.
[0176] 13 and 14 are schematic diagrams illustrating a method for adjusting the output loudness level of an input audio signal according to an embodiment of the present disclosure. FIG. 13 illustrates an embodiment in which loudness information of an input audio signal is extracted and the output loudness level of the input audio signal is adjusted within a single audio signal processing device. In this case, the audio signal processing device can measure the loudness level of the input audio signal. The audio signal processing device can obtain the loudness information of the input content as a loudness measurement value. A method for the audio signal processing device to measure the loudness level of an input audio signal in real time will be specifically described with reference to FIG. 19.
[0177] FIG. 14 shows the server-client structure described above with reference to FIGS. 7 and 8. First, the server can analyze an input audio signal to extract loudness information of the input audio signal. The server can also convert the loudness information of the input audio signal into a metadata format and generate loudness metadata. Next, the client can receive the input audio signal and receive the loudness metadata of the input audio signal separately from the input audio signal. The client can also parse the loudness metadata to obtain loudness information used to adjust the output loudness level of the input audio signal. The client can also obtain a loudness gain for the input audio signal based on the loudness information and a previously set target loudness level. The client can adjust the output loudness level of the input audio signal based on the loudness gain of the input audio signal.
[0178] 15 is a diagram illustrating a method for an audio signal processing device according to an embodiment of the present invention to acquire loudness information of an input audio signal. The audio signal processing device may acquire loudness information by analyzing the input audio signal. For example, the method of FIG. 15 may be performed by the server of FIG. 7 described above. The audio signal processing device may output the loudness information in the form of loudness metadata.
[0179] According to an embodiment, the loudness information may include static loudness metadata and dynamic loudness metadata. The static loudness metadata may include at least one static loudness parameter. For example, the static loudness metadata may include at least one of a cumulative loudness level of the input audio signal, a Max. Sample Peak (Max. Sample Peak), a Loudness Range (LRA), a Peak to Loudness Range (PLR), an Album Integrated Loudness (Album Integrated Loudness), a Relative Threshold (Relative Threshold), a Min. Momentary Loudness (Min. Momentary Loudness), a Max. Momentary Loudness (Max. Momentary Loudness), and a Sample Per Frame (Sample Per Frame).
[0180] The audio signal processing device can acquire static loudness metadata of an input audio signal. Specifically, the audio signal processing device can measure at least one of an instantaneous loudness level of the input audio signal and a short-term loudness level of the input audio signal using a loudness filter based on a hearing scale. The audio signal processing device can generate static loudness metadata including at least one static loudness parameter.
[0181] The dynamic loudness metadata may indicate loudness information that changes over time. The dynamic loudness metadata may include at least one dynamic loudness parameter. For example, the dynamic loudness metadata may include at least one of a time-dependent short-term loudness level and a peak value (Peak Envelope) of an input audio signal. A method for the audio signal processing device to acquire the peak value will be described in detail with reference to FIG. 21.
[0182] According to one embodiment, an audio signal processing device may acquire dynamic loudness metadata of an input audio signal. For example, the audio signal processing device may acquire short-term loudness measurements for a specific section of the input audio signal. The audio signal processing device may acquire peak values of the input audio signal for the section. The audio signal processing device may generate dynamic loudness metadata including at least one dynamic loudness parameter. The audio signal processing device may also correct time delays or advances of dynamic loudness parameters such as short-term loudness measurements and peak values. For example, the audio signal processing device may shift the dynamic loudness parameters. This will be described in detail with reference to FIG. 21.
[0183] The audio signal processing device can acquire short-term loudness levels for past sample values and subsequently input sample values based on a specific time point. As a result, the audio signal processing device can more stably control the loudness level in response to loudness changes in the input audio signal. For example, the audio signal processing device can shift a time reference value of an already acquired dynamic loudness parameter to acquire short-term loudness levels for past sample values and subsequently input sample values. Furthermore, the audio signal processing device can acquire short-term loudness levels for past sample values and subsequently input sample values using a buffer. In this case, the audio signal processing device can set a sufficient look-ahead time.
[0184] FIG. 16 is a diagram illustrating a method for adjusting an output loudness level of an input audio signal by an audio signal processing device according to an embodiment of the present invention. The audio signal processing device may obtain a loudness gain of an input audio signal based on a target loudness level and loudness metadata of the input audio signal. Specifically, the audio signal processing device may calculate a gain parameter based on the target loudness level and the static loudness metadata. The audio signal processing device may obtain a loudness gain to be applied to a particular frame of the input audio signal based on the calculated gain parameter and the dynamic loudness metadata. For example, the audio signal processing device may parse the dynamic loudness metadata to obtain at least one of a short-term loudness level and a peak value corresponding to the frame. The audio signal processing device may obtain a loudness gain to be applied to the frame based on at least one of a short-term loudness level and a peak value corresponding to the frame. Specifically, the audio signal processing device may obtain a loudness gain to be applied to the frame based on the calculated gain parameter and the short-term loudness level corresponding to the frame. In this case, the loudness gain to be applied to the frame may be limited so as to prevent clipping due to the loudness level within the frame. The audio signal processing device may correct a loudness gain applied to a frame based on a peak value so as to prevent clipping due to loudness level within the frame. The audio signal processing device may generate an intermediate audio signal by applying a final loudness gain to an input audio signal. The audio signal processing device may also generate an output audio signal by applying a limiter to the intermediate audio signal. The audio signal processing device may output the output audio signal. According to a further embodiment, when a difference in frame-by-frame loudness gain between adjacent frames is equal to or greater than a predetermined magnitude, the audio signal processing device may correct the frame-by-frame loudness gain.At this time, the audio signal processing device may adjust the loudness gain so that it changes gradually using a smoothing method. As a result, the audio signal processing device may prevent tone distortion due to changes in the loudness gain for each frame and volume pumping due to sudden large level changes. The method by which the audio signal processing device smooths the loudness gain will be described in detail with reference to FIG. 22.
[0185] 17 is a diagram illustrating a method in which an audio signal processing device according to an embodiment of the present invention adjusts an output loudness level of an input audio signal based on a target loudness range. The audio signal processing device may further consider the target loudness range in the process of calculating the gain parameters of FIG. 16. As described with reference to FIG. 12, the target loudness range may be narrower than the dynamic range of the input audio signal. Depending on the environment, when listening to video / audio at a low volume or when listening to music in a noisy environment such as a subway or road, it is necessary to reduce the dynamic range of the input audio signal before reproduction.
[0186] As a result, the audio signal processing device can calculate a gain parameter of the input audio signal based on a target loudness range of the input audio signal. In this case, the gain parameter may include a gain ratio used for loudness compression. The audio signal processing device can apply an additional boost gain to frames having a short-term loudness smaller than a predetermined level among a plurality of frames included in the input audio signal based on the gain ratio. The audio signal processing device can apply an additional cut gain to frames having a short-term loudness larger than a predetermined level among a plurality of frames included in the input audio signal based on the gain ratio. As a result, the audio signal processing device can adjust the output loudness level of the entire section of the input audio signal to approximate the target loudness level.
[0187] According to an additional embodiment, the audio signal processing device may perform loudness normalization for each time interval based on loudness parameters measured differently for each time interval. Specifically, the audio signal processing device may determine a loudness gain (G_loud) for each time interval of the input audio signal based on a target loudness level (L_T), an accumulated loudness level (L_I), a short-term loudness level (L_S), a relative threshold (L_Rel), a noise floor level (L_Noise), and a peak value (P). Here, L_Rel may be a value obtained by adding a preset value to an average of dynamic loudness parameters effective over the entire interval of the input audio signal. In this case, the preset value may be −20LU. In addition, the dynamic loudness parameter may be an instantaneous loudness level or a short-term loudness level.
[0188] For example, L_Rel may be a value calculated based on an average of short-term loudness levels, among the short-term loudness levels for each section of the input audio signal, that are at least greater than the effective loudness level. L_Rel may be a value calculated based on an average of instantaneous loudness levels, among the instantaneous loudness levels for each section of the input audio signal, that are at least greater than the effective loudness level. Here, the effective loudness level may be a value set based on a loudness level that is difficult to perceive auditorily. The effective loudness level may be a value set based on the loudness level of an audio signal with almost no sound. For example, the effective loudness level may be a value set based on -70LKFS.
[0189] Furthermore, L_Noise may be a value calculated based on at least one of the loudness level of an interval in which there is almost no sound in the input audio signal or the loudness level of an interval in the input audio signal corresponding to a very low level of background noise.
[0190] According to one embodiment, L_T, L_I, L_S, L_Rel, L_Noise, and P can be obtained from the loudness metadata described above. The time interval may include a frame. In the above embodiment, the short-term loudness level (L_S) may be replaced with a loudness representative value representing a specific time interval. For example, the short-term loudness level (L_S) may be replaced with an instantaneous loudness level of an input audio signal. A method in which the audio signal processing device obtains the loudness gain (G_loud) for each time interval based on L_T, L_I, L_S, L_Rel, L_Noise, and P can be expressed as Equation 16 below.
[0191] [Number 16]
[0192]
number
[0193] In Equation 16, r_1 and r_2 may represent loudness compression ratios for controlling the dynamic range of the output audio signal with respect to the input audio signal. r_1 may be a loudness compression ratio used to obtain a loudness gain for a section in which the input loudness level of the input audio signal is smaller than at least the cumulative loudness level. r_1 may be set based on at least one of LRA, PLR, or maximum instantaneous loudness value, which indicate the loudness range of the input audio signal. r_1 may be an arbitrary constant between 0 and 1. r_2 may be a compression ratio used to obtain a loudness gain for a section in which the input loudness level of the input audio signal is smaller than the cumulative loudness level and the input loudness level is smaller than L_Rel. In this case, r_2 may be set to a value at least smaller than r_1 to minimize boosting of noise components. The audio signal processing device may smooth G_loud[n] and apply it to the input audio signal. Furthermore, clippingThreshold may represent a maximum allowable sample peak value. ClippingThreshold may be a value set based on at least one of the QSHI, maximum true peak, and maximum sample peak value. For example, clippingThreshold may be the same value as QSHI. Alternatively, clippingThreshold may be a value arbitrarily set in the audio signal processing device or audio delivery system.
[0194] Hereinafter, a method for an audio signal processing device according to an embodiment of the present invention to acquire loudness measurement values will be described in detail with reference to FIG. 18. FIG. 18 is a diagram illustrating a method for an audio signal processing device to measure loudness of input content according to an embodiment of the present invention. According to an embodiment, the audio signal processing device may measure the loudness of input content based on the above-described measurement window. Furthermore, the audio signal processing device may acquire loudness measurement values for each measurement window of the input content. The audio signal processing device may acquire loudness information based on the loudness measurement values for each measurement window.
[0195] In the embodiment of FIG. 18 , the audio signal processing device may acquire a measurement value for each measurement window based on the length of the measurement window 801. In this case, the length of the measurement window 801 may be a default value already stored in the audio signal processing device. According to an embodiment of the present invention, the length of the measurement window 801 may vary depending on the input content. For example, the audio signal processing device may acquire the length of the measurement window corresponding to the input content based on additional information of the input content. In the embodiment of FIG. 18 , the length of the measurement window corresponding to the input content may be 400 ms. The audio signal processing device may acquire a loudness measurement value corresponding to a specific 400 ms long section in the entire section of the input content.
[0196] According to an embodiment, the length of the measurement window may be obtained based on the additional information. For example, the length of the measurement window may be obtained based on the loudness range of the input content. Here, the loudness range may be a value representing the loudness level distribution for the entire section of the content. The loudness range may be expressed using a unit indicating a relative measurement quantity, such as LU. The audio signal processing device may obtain information regarding the loudness range of the input content from the additional information. Then, the audio signal processing device may determine the length of the measurement window based on the loudness range of the input content. In this case, the length of the measurement window of the input content may be set to a value shorter than the length of the measurement window of another content having a loudness range wider than the loudness range of the input content. For example, if the loudness range of a first input content is wider than the loudness range of a second input content, the length of the measurement window for the first input content may be longer than the length of the measurement window for the second input content.
[0197] Furthermore, the audio signal processing device may acquire loudness measurement values for each measurement window according to a measurement period for acquiring measurement values for the input content. In the present disclosure, the measurement period may represent a temporal distance over which the measurement window moves. Referring to FIG. 18 , a first measurement value 802 may be a loudness measurement value corresponding to a section (300 ms to 700 ms) based on the time point at which the input content starts to be played. Furthermore, a second measurement value 803 may be a loudness measurement value corresponding to a section (400 ms to 800 ms) based on the time point at which the input content starts to be played. If the time length from the time point at which the input content starts to be played to the current time is shorter than the length of the measurement window, the audio signal processing device may acquire loudness measurement values in the nearest measurement period after the current time. In this case, the audio signal processing device may acquire loudness measurement values corresponding to a section shorter than the length of the measurement window.
[0198] Specifically, the audio signal processing device can determine a measurement period based on the additional information. For example, the measurement period may be determined based on the length of the input content. For example, if the length of the second input content is longer than the length of the first input content, the measurement period for the first input content may be shorter than the measurement period for the second input content. The audio signal processing device can also acquire loudness measurement values for each measurement window based on the determined measurement period. In the example of FIG. 18, the measurement period may be 100 ms. The audio signal processing device can move the measurement window every 100 ms and acquire loudness measurement values for each measurement window. The audio signal processing device can also acquire the above-mentioned loudness information based on the multiple loudness measurement values measured in FIG. 18.
[0199] 19 is a flowchart showing the operation of an audio signal processing device according to an embodiment of the present invention. The audio signal processing device according to an embodiment of the present invention may receive an input audio signal (step S1901). At this time, the input audio signal may include the input content described in FIG. 2. Next, the audio signal processing device may receive loudness metadata corresponding to the input audio signal (step S1902).
[0200] Next, the audio signal processing device may parse the loudness metadata to acquire loudness information of the input audio signal (step S1903). According to an embodiment of the present invention, the loudness information may include at least one of information indicating a cumulative loudness level of the input audio signal, at least one short-term loudness level, a Quality Secure Histogram Index (QSHI), a dynamic range of the input audio signal, frequency-specific loudness energy, frequency-specific loudness ratio, and peak envelope. The methods by which the audio signal processing device acquires each piece of information included in the loudness information may be applied to the embodiments described with reference to FIGS. 2 to 18.
[0201] The QSHI may indicate a threshold loudness level at which no perceptual sound quality impairment occurs. The QSHI may be obtained by the above-described step S303 of FIG. 3, the QSHI extraction unit 404 of FIG. 4, and the embodiment described in FIG. 10. For example, the QSHI may be a loudness parameter calculated based on a loudness histogram of the input audio signal. In this case, the loudness histogram may be a size histogram of the short-term loudness level of the input audio signal over time. Alternatively, the loudness histogram may be a size histogram related to the peak value or root-mean-square (RMS) of the input audio signal over each section. The QSHI may be greater than the cumulative loudness level of the input audio signal.
[0202] According to one embodiment, the QSHI may be a parameter calculated based on a predicted loudness histogram predicted from the loudness histogram of the input audio signal, where the predicted loudness histogram may be a histogram generated based on loudness parameters predicted when the input audio signal is output according to a target loudness level.
[0203] According to an embodiment, the QSHI may be determined based on the number of times a limiter in the audio signal processing device is driven. In this case, the audio signal processing device may apply a loudness limiter that limits the loudness level of an output audio signal to the output audio signal and output the output audio signal. In this case, the output audio signal may be a signal in which the output loudness level of an input audio signal is adjusted by a loudness gain. The QSHI may be a parameter set so that the short-term loudness level of the entire section of the output audio signal is equal to or lower than a predetermined level.
[0204] Next, the audio signal processing apparatus may obtain a loudness gain of the input audio signal based on the loudness information and the target loudness level (S1904). According to one embodiment, the loudness gain of the input audio signal may be a fixed gain having a fixed value throughout the entire duration of the input audio signal. According to another embodiment, the loudness gain of the input audio signal may be a gain that varies over time while the input audio signal is being played back.
[0205] According to one embodiment of the present invention, an audio signal processing apparatus may receive a cumulative loudness of an input audio signal, and may determine a loudness gain based on the cumulative loudness of the input audio signal, a QSHI, and the target loudness level.
[0206] According to one embodiment, the audio signal processing device may compare a target loudness level of an input audio signal with the QSHI. The audio signal processing device may then determine a loudness gain based on the comparison result. The audio signal processing device may determine the loudness gain based on the smaller value of the target loudness level of the input audio signal and the QSHI. The specific embodiment described with reference to FIG. 10 may be applied to this.
[0207] According to one embodiment, an audio signal processing device may obtain a loudness gain of an input audio signal based on a QSHI corrected from the QSHI of the input audio signal. For example, the audio signal processing device may perform post-processing on the input audio signal. In this case, the audio signal processing device may receive post-processing information indicating characteristics of the post-processing on the input audio signal. The audio signal processing device may also correct an already-obtained QSHI based on the post-processing information. According to one embodiment, the audio signal processing device may correct an already-obtained QSHI based on the post-processing information and a previously-stored function. The audio signal processing device may correct an already-obtained QSHI based on the post-processing information and a previously-stored look-up table. In this case, the previously-stored look-up table may be a table containing information regarding QSHI correction according to characteristics of the post-processing. Furthermore, the information regarding QSHI correction may include information indicating a QSHI correction value according to the characteristics of the post-processing. The audio signal processing device may obtain a QSHI correction value corresponding to the post-processing on the input audio signal based on the previously-stored look-up table. The audio signal processing device can correct the acquired QSHI by adding a QSHI correction value to the acquired QSHI. The audio signal processing device can determine the loudness gain of the input audio signal based on the QSHI corrected by the above-mentioned method.
[0208] According to an embodiment, the audio signal processing device may determine a loudness gain of the input audio signal based on frequency-specific loudness energy and post-processing information indicating characteristics of post-processing on the input audio signal. The audio signal processing device may determine a loudness gain of the input audio signal based on band-specific loudness levels that change due to post-processing.
[0209] According to an embodiment, an audio signal processing device may obtain a band-specific loudness level that changes due to post-processing based on frequency-specific loudness energy and post-processing information indicating characteristics of post-processing for an input audio signal. The audio signal processing device may obtain a band-specific loudness level that changes due to post-processing based on frequency-specific loudness ratios and the post-processing information for the input audio signal. The band-specific loudness level that changes due to post-processing may be calculated based on an inner product of frequency-specific loudness ratios of the input audio signal. The band-specific loudness level that changes due to post-processing may also be a parameter obtained based on a perceptual loudness characteristic. The audio signal processing device may obtain a band-specific loudness level that changes due to post-processing of the input audio signal based on a loudness filter based on an auditory scale. Specifically, the loudness filter may be at least one of an inverse filter of equal-loudness contours or a K-weighting filter that approximates the inverse filter. When the loudness level of a specific frame among a plurality of frames included in the input audio signal is less than or equal to the relative threshold, the audio signal processing device may not calculate the band-specific loudness level that is changed by post-processing corresponding to the frame. As another example, the band-specific loudness level that is changed by post-processing of the input audio signal may be a parameter set based on at least one of the genre of the input audio signal and a user input.
[0210] The frequency-specific loudness ratio and / or frequency-specific loudness energy of the input audio signal may be values calculated based on loudness measurement values for the input audio signal. The frequency-specific loudness ratio may be a parameter obtained based on a perceptual loudness characteristic. The audio signal processing device may obtain the frequency-specific loudness ratio of the input audio signal based on a loudness filter based on a hearing scale. Specifically, the loudness filter may be at least one of an inverse filter of equal-loudness contours or a K-weighting filter that approximates the inverse filter. If the loudness level of a specific frame among multiple frames included in the input audio signal is less than or equal to a relative threshold, the audio signal processing device may not calculate the frequency-specific loudness ratio corresponding to the frame. The frequency-specific loudness ratio may be obtained by the embodiment described with reference to the frequency-specific loudness analysis unit 402 of FIG. 4. As another example, the frequency-specific loudness ratio of the input audio signal may be a parameter set based on at least one of the genre of the input audio signal and a user input.
[0211] The audio signal processing device may acquire post-processing information for an input audio signal based on a user input. The user input may be an input related to the input audio signal. The user may be a user who uses the audio signal processing device. The post-processing information may include at least one of information indicating an output characteristic of the audio signal processing device, a genre of the input audio signal, a post-processing mode based on the user input, an equalization type, reverberation, and room compensation. The method for determining a loudness gain of an input audio signal based on a band-specific loudness level that changes through post-processing by the audio signal processing device may be the same as the embodiment described in step S303 of FIG. 3.
[0212] According to an embodiment, the audio signal processing device may determine a loudness gain of an input audio signal based on a loudness change prediction value. The loudness change prediction value may be a prediction value for a loudness change of the input audio signal due to post-processing. The audio signal processing device may acquire the loudness change prediction value based on post-processing information set by a user. The audio signal processing device may acquire the loudness change prediction value based on at least one of frequency characteristics of the input audio signal and a band-specific loudness level changed by post-processing. The loudness change prediction value may be calculated based on an inner product of frequency-specific loudness ratios of the input audio signal. The loudness change prediction value may be a parameter acquired based on a perceptual loudness characteristic. The audio signal processing device may acquire the loudness change prediction value of the input audio signal based on a loudness filter based on an auditory scale. Specifically, the loudness filter may be at least one of an inverse filter of equal-loudness contours or a K-weighting filter that approximates the inverse filter. When the loudness level of a specific frame among a plurality of frames included in the input audio signal is smaller than or equal to the relative threshold, the audio signal processing apparatus may not calculate a loudness change prediction value corresponding to the specific frame. The method of the audio signal processing apparatus to obtain the loudness change prediction value may be the same as the embodiment described with reference to the frequency-specific loudness analysis unit 402 and the post-processing loudness prediction unit 403 in FIG. 4 .
[0213] According to an embodiment of the present invention, an audio signal processing apparatus may determine a loudness gain of an input audio signal based on frame-by-frame loudness information of the input audio signal. The audio signal processing apparatus may obtain a frame-by-frame loudness gain of the input audio signal based on the frame-by-frame loudness information of the input audio signal. The loudness gain of the input audio signal may be a gain that changes over time during playback of the input audio signal. According to an embodiment, the audio signal processing apparatus may receive loudness metadata including frame-by-frame loudness information of the input audio signal. The audio signal processing apparatus may parse the loudness metadata to obtain the frame-by-frame loudness information of the input audio signal. The frame-by-frame loudness information may include a dynamic loudness parameter. According to an embodiment, the frame-by-frame loudness information may include information indicating a frame-by-frame peak value. The frame-by-frame peak value may be obtained based on a maximum absolute value of an audio signal included in frames of a predetermined length.
[0214] According to one embodiment, an audio signal processing device may determine a loudness gain for each frame of an input audio signal based on a peak value for each frame of the input audio signal. The audio signal processing device may determine the loudness gain for each frame of the input audio signal based on a target loudness level and a peak value for each frame of the input audio signal. For example, the audio signal processing device may set the loudness gain for each frame based on the target loudness level so as not to exceed the peak value for each frame. Furthermore, the audio signal processing device may adjust the output loudness level of a corresponding frame of the input audio signal based on the loudness gain for each frame. The embodiment described above with reference to FIG. 17 may be applied to a method in which the audio signal processing device determines the loudness gain based on the loudness information for each frame.
[0215] Next, the audio signal processing device may adjust an output loudness level of the input audio signal based on the loudness gain (S1905). According to one embodiment, the audio signal processing device may generate an output audio signal by adjusting the output loudness level of the input audio signal. In this case, the audio signal processing device may use the determined loudness gain. According to one embodiment, the audio signal processing device may apply a loudness limiter to the generated output audio signal and output it.
[0216] According to a further embodiment of the present invention, the audio signal processing device may adjust an output loudness level of an input audio signal based on a section loudness gain for a portion of the entire section of the input audio signal. According to one embodiment, the audio signal processing device may obtain a loudness gain corresponding to a specific section of the input audio signal based on a loudness parameter corresponding to the specific section. For example, the loudness parameter corresponding to the specific section of the input audio signal may include at least one representative value for the specific section. In this case, the representative value may include at least one of a maximum absolute value of the loudness level of the input audio signal corresponding to the specific section and a short-section loudness level.
[0217] According to an embodiment, an audio signal processing device may determine a loudness gain for each time interval of an input audio signal based on a target loudness level, a cumulative loudness level, and an input loudness level. In this case, the input loudness level may be a loudness level representative of a specific interval. For example, the input loudness level may be a short-term loudness level. The audio signal processing device may compare at least two of the target loudness level, the cumulative loudness level, the input loudness level, a relative threshold, a noise floor level, and a peak value with each other. Furthermore, the audio signal processing device may determine a loudness gain for each time interval of the input audio signal based on a comparison result.
[0218] For example, the audio signal processing device may compare a target loudness level with a cumulative loudness level. The audio signal processing device may compare an input loudness level with a cumulative loudness level. If the target loudness level is smaller than the cumulative loudness level and the input loudness level is larger than the cumulative loudness level, the audio signal processing device may apply a first segment-specific loudness gain to the input audio signal for the segment.
[0219] As another example, when the target loudness level is greater than the cumulative loudness level, the input loudness level is less than the cumulative loudness level, and the input loudness level is greater than a relative threshold, the audio signal processing device can apply a second section-specific loudness gain to the input audio signal of the section.
[0220] As yet another example, when the target loudness level is greater than the cumulative loudness level, the input loudness level is less than the cumulative loudness level, the input loudness level is less than the relative threshold, and the input loudness level is greater than the noise floor level, the audio signal processing device can apply a third section-specific loudness gain to the input audio signal of the section.
[0221] In yet another embodiment, when the target loudness level is greater than the cumulative loudness level, the input loudness level is less than the cumulative loudness level, the input loudness level is less than the relative threshold, and the input loudness level is less than the noise floor level, the audio signal processing device may apply a fourth sectional loudness gain to the input audio signal of the corresponding section. In this case, the fourth sectional loudness gain may be the loudness gain of a frame prior to the corresponding frame. For example, when the target loudness level is greater than the cumulative loudness level, the input loudness level corresponding to the Nth frame is less than the cumulative loudness level, the input loudness level corresponding to the Nth frame is less than the relative threshold, and the input loudness level corresponding to the Nth frame is less than the noise floor level, the audio signal processing device may use the loudness gain corresponding to the (N-1)th frame as the loudness gain corresponding to the Nth frame.
[0222] According to another embodiment, the fourth section loudness gain may represent a fixed gain applied to the entire input audio signal. Also, the first section loudness gain, the second section loudness gain, and the third section loudness gain may be gains corrected in individual ways based on the fourth section loudness gain. Also, the first section loudness gain, the second section loudness gain, and the third section loudness gain may be gains having individual values.
[0223] According to an embodiment, the loudness representative value of an Nth section of an input audio signal may be a representative value corresponding to a section adjacent to the Nth section of the input audio signal. For example, the loudness representative value of an Nth specific section of the input audio signal may be a representative value corresponding to an (N+L)th or NLth section. In this case, L may be an index value corresponding to a section smaller than the time section for obtaining the representative value. For example, the time section for obtaining the representative value may be 3 seconds. Furthermore, the audio signal processing device may obtain a representative value of a specific section of the input audio signal based on a time-delayed input audio signal. In this case, the audio signal processing device may time-delay the input audio signal based on a preset delay time and obtain at least one loudness measurement value used for obtaining the representative value.
[0224] According to one embodiment, an audio signal processing device may obtain a fixed loudness gain to be applied to the entire input audio signal. In this case, the audio signal processing device may correct the fixed loudness gain based on a loudness parameter corresponding to a specific section of the input audio signal. The audio signal processing device may also adjust the output loudness level of the input audio signal for that section based on the corrected gain. The embodiment described above with reference to FIG. 17 may be applied to a method in which the input audio signal processing device adjusts the output loudness level of the input audio signal based on a section loudness gain for a portion of the entire section of the input audio signal.
[0225] FIG. 20 is a block diagram showing a configuration of an audio signal processing device 2000 according to an embodiment of the present invention. According to an embodiment, the audio signal processing device 2000 may include a receiving unit 2100, a processor 2200, and an output unit 2300. However, not all of the components shown in FIG. 10 are necessarily essential components of the audio signal processing device. The audio signal processing device 2000 may further include components not shown in FIG. 20. For example, the audio signal processing device according to an embodiment may further include a storage unit (not shown). Note that at least some of the components of the audio signal processing device 2000 shown in FIG. 20 may be omitted. For example, the audio signal processing device according to an embodiment may not include at least one of the receiving unit 2100 and the output unit 2300.
[0226] The receiving unit 2100 may receive input content input to the audio signal processing device 2000. The receiving unit 2100 may receive input content whose output loudness level is to be adjusted by the processor 2200. As described above, the input content may include an audio signal. In this case, the audio signal may include at least one of an Ambisonic signal, an object signal, or a channel signal. The audio signal may be a single object signal or a mono signal. The audio signal may be a multi-object or multi-channel signal. According to an embodiment, the receiving unit 2100 may include an input terminal for receiving input content transmitted via a wire. The receiving unit 2100 may also include a wireless receiving module for receiving input content transmitted via a wireless network.
[0227] According to an embodiment, the audio signal processing apparatus 2000 may include a separate decoder. In this case, the receiving unit 2100 may receive an encoded bitstream of the input content. The encoded bitstream may be decoded as the input content by the decoder. The receiving unit 2100 may also receive additional information related to the input content.
[0228] According to an embodiment, the receiving unit 2100 may include a transceiver for transmitting and receiving data to and from an external device via a network. The data may include at least one of a bitstream of input content and additional information. The receiving unit 2100 may include a wired transceiver terminal for receiving data transmitted via a wired connection. The receiving unit 2100 may also include a wireless transceiver module for receiving data transmitted wirelessly. In this case, the receiving unit 2100 may receive data transmitted wirelessly using a Bluetooth or Wi-Fi communication method. The receiving unit 2100 may also receive data transmitted according to a mobile communication standard such as LTE (Long Term Evolution) or LTE-Advanced, but the present disclosure is not limited thereto. The receiving unit 2100 may receive various types of data transmitted according to various wired and wireless communication standards.
[0229] The processor 2200 may control the overall operation of the audio signal processing device 2000. The processor 2200 may control each component of the audio signal processing device 2000. The processor 2200 may perform calculations and processes of various data and signals. The processor 2200 may be implemented by hardware in the form of a semiconductor chip or an electronic circuit, or by software that controls the hardware. The processor 2200 may also be implemented by a combination of hardware and software. For example, the processor 2200 may control the operations of the receiving unit 2100 and the output unit 2300 by executing at least one program. The processor 2200 may also execute at least one program to perform the operations described with reference to FIGS. 1 to 19.
[0230] According to one embodiment, the processor 2200 may adjust the output loudness level of the input content. For example, the processor 2200 may adjust the output loudness level of the input content based on a loudness gain. The loudness information may be loudness characteristics of the input content analyzed from the input content. In this case, the loudness gain may be obtained based on the loudness information. The processor 2200 may also output output content whose output loudness level has been adjusted from the input content. In this case, the processor 2200 may output the output content from the output unit 2300, which will be described later.
[0231] The output unit 2300 may output output content. The output unit 2300 may output output content whose output loudness level has been adjusted from the input content by the processor 2200. Here, the output content may include an output audio signal. In this case, the output audio signal may include at least one of an Ambisonic signal, an object signal, or a channel signal. The output audio signal may be a multi-object or multi-channel signal. Furthermore, the output audio signal may include a two-channel output audio signal corresponding to each ear of a listener. The output audio signal may include a binaural two-channel output audio signal. The output unit 2300 may output an audio headphone signal whose output loudness level has been adjusted by the processor 2200.
[0232] According to an embodiment, the output unit 2300 may include an output means for outputting output content. For example, the output unit 2300 may include an output terminal for outputting an output audio signal to an external device. In this case, the audio signal processing device 2000 may output the output audio signal to an external device connected to the output terminal. The output unit 2300 may include a wireless audio transmission module for outputting the output audio signal to an external device. In this case, the output unit 2300 may output the output audio signal to the external device using a wireless communication method such as Bluetooth or Wi-Fi.
[0233] The output unit 2300 may also include a speaker. In this case, the audio signal processing device 2000 can output an output audio signal from the speaker. The output unit 2300 may also include a converter (e.g., a digital-to-analog converter, DAC) that converts a digital audio signal into an analog audio signal. The output unit 2300 may also include a display means that outputs a video signal included in the output content.
[0234] As described above, the audio signal processing apparatus 2000 may further include a storage unit (not shown). The storage unit may store at least one of data or programs for processing and control of the processor 2200. The storage unit may also store loudness information. The storage unit may store loudness information extracted from received loudness metadata. The storage unit may store a received target loudness level. Alternatively, the storage unit may store a loudness measurement value acquired by the processor 2200. The storage unit may also store a result of a calculation performed by the processor 2200. For example, the storage unit may store a loudness gain determined based on the loudness information. The storage unit may also store data input to or output from the audio signal processing apparatus 2000.
[0235] The storage unit may include at least one memory, which may include at least one type of storage medium selected from the group consisting of flash memory, hard disk, micro multimedia card, card-type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk.
[0236] FIG. 21 is a diagram illustrating peak values for each time interval of an input audio signal according to an embodiment of the present invention. In the embodiment of FIG. 21, the peak values for each time interval may be values obtained based on loudness measurements measured from the input audio signal. In FIG. 21, values indicated by solid lines represent loudness measurements for each time interval of the input audio signal. Also, values indicated by a first dashed line (-*-) represent representative values for each time interval of the loudness measurements for each time interval of the input audio signal. The audio signal processing device may obtain peak values for each time interval based on the representative values for each time interval. In this case, since the representative values are calculated based on values input to an input buffer of the loudness measurement device, errors may occur if the representative values are calculated based on the actual input audio signal.
[0237] In FIG. 21, the values indicated by the second dashed line (-△-) may be representative values for each time interval obtained by a time delay of about 15 ms. The audio signal processing device may obtain the representative values for each time interval by applying a time delay to the input audio signal. As a result, the audio signal processing device may correct the obtained peak value so that it can more accurately correspond to loudness changes in the input audio signal. In this case, the delay duration used for the time delay may be set based on the length of the measurement frame of the input audio signal. The time delay correction method for peak values described in FIG. 21 may also be applied to other dynamic loudness parameters described in FIG. 15. For example, the audio signal processing device may obtain a short-term loudness level using a time delay.
[0238] 22 is a diagram illustrating a method for adjusting an output loudness level of an input audio signal using smoothing in an audio signal processing device according to an embodiment of the present invention. According to an embodiment of the present invention, the audio signal processing device can adjust the output loudness level of the input audio signal using smoothing so that the loudness gain changes smoothly. In this case, since smoothing is performed based on the loudness measurement value of the input audio signal (causal processing), it may be difficult for the audio signal processing device to correctly provide parameters required in a corresponding frame for actual loudness changes.
[0239] Therefore, the audio signal processing device can perform a smoothing operation on the loudness gain of the input audio signal using the loudness parameter obtained by the time delay, which may be a parameter obtained by the method described above in FIG.
[0240] In FIG. 22, values represented by solid lines may represent loudness gains for individual frames of an input audio signal. Here, values represented by solid lines may represent loudness gains to which no smoothing has been applied. Furthermore, values represented by the third dashed line (--) and the fourth dashed line (-·-) may represent loudness gains to which smoothing has been applied from the frame loudness gains. Here, each of the frame loudness gains represented by the third dashed line (--) may represent a first frame loudness gain (smoothing from shifted input) obtained based on a measurement value to which a time delay has been applied. Meanwhile, each of the frame loudness gains represented by the fourth dashed line (-·-) may represent a second frame loudness gain (smoothing from org.input) obtained based on a measurement value to which a time delay has not been applied.
[0241] 22, the loudness gain for each second frame may change more similarly to the loudness level of the input audio signal than the loudness gain for each first frame. Referring to the section of frame indexes 110 to 130 on the horizontal axis of FIG. 22, the loudness gain for each frame to which smoothing of the input audio signal is not applied suddenly decreases. In this section, the loudness gain for each first frame gradually decreases compared to the loudness gain for each second frame. The loudness gain for each second frame also decreases suddenly compared to the loudness gain for each first frame. Furthermore, the loudness gain for each first frame starts decreasing a certain number of frames earlier than the loudness gain for each second frame. Thus, the audio signal processing device can prevent a listener from perceiving a sudden change in loudness by using the loudness gain for each first frame acquired based on a measurement value to which a time delay is applied.
[0242] According to one embodiment of the present invention, an audio signal processing device may apply a loudness gain determined for each section to an input audio signal in order to process the characteristics of the input audio signal according to a target loudness level. In this case, an excessive loudness gain value may be applied to a specific section. This may result in clipping greater than 0 dBFS or a loudness greater than a predefined threshold value. For this reason, the audio signal processing device may apply a limiter to an output audio signal. Thus, the audio signal processing device may apply a limiter to a section in which the loudness level of an output audio signal, obtained by adjusting the output loudness level of the input audio signal, is greater than a previously set loudness level.
[0243] In this case, the limiter may process the output audio signal in real time or in a time sequence (causal processing) according to limiter parameters associated with the limiter. When an audio signal processing device uses a limiter, the audio signal processing device may generate unintended timbre distortion. As described above, the audio signal processing device may adjust the output loudness level of an input audio signal using a loudness gain determined for each section. In this case, the loudness gain determined for each section may be a gain that takes into account peak values for each section. Based on the peak values for each section, the audio signal processing device may predict clipping that will occur in the corresponding section or the occurrence of a section having a level exceeding a target loudness level. Furthermore, the audio signal processing device may determine a loudness gain for each section of the input audio signal based on the prediction. That is, the audio signal processing device may conversely correct the loudness gain based on the prediction. As a result, the audio signal processing device may prevent timbre distortion of the output audio signal caused by the limiter.
[0244] Some embodiments may be embodied in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. A computer-readable medium may be any available medium accessible by a computer, and may include both volatile and nonvolatile media, and both separate and non-separate media. A computer-readable medium may also include a computer storage medium. A computer storage medium may include both volatile and non-volatile, separate and non-separate media embodied in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data.
[0245] Although the present disclosure has been described above using specific examples, those skilled in the art who have ordinary skill in the art of the present disclosure can make modifications and changes without departing from the spirit and scope of the present disclosure. That is, although the present disclosure has described an example of adjusting the loudness level of an audio signal, the present disclosure can be similarly applied and extended to various multimedia signals, including video signals, in addition to audio signals. Therefore, anything that can be easily inferred by a person skilled in the art in the art of the present disclosure from the detailed description and examples of the present disclosure is construed as falling within the scope of the present disclosure. [Explanation of symbols]
[0246] 801 measurement window 2000 Audio Signal Processing Device 2100 Receiver 2200 processor 2300 Output Unit
Claims
1. 1. An audio signal processing device for controlling a loudness level, the audio signal processing device comprising: a receiving unit for receiving an input audio signal; a processor for generating loudness metadata corresponding to the input audio signal; an output unit that transmits the generated loudness metadata to the processor; The processor: measuring the loudness of the input audio signal to obtain loudness information of the input audio signal; transforming the loudness information to generate the loudness metadata; outputting the generated loudness metadata to an output device for outputting the input audio signal via the output unit, the loudness information including information representing a loudness ratio for each frequency of the input audio signal; the loudness ratios for each frequency of the input audio signal are used by an audio signal processing device that controls the loudness level of the input audio signal using the loudness metadata to obtain a loudness difference modified by post-processing; Audio signal processing device.
2. The audio signal processing device of claim 1 , wherein the post-processing includes at least one of equalization, reverberation, and spatial compensation.
3. The post-treatment is applying an output characteristic of an audio signal processing device using the loudness metadata to control the loudness level of the input audio signal.
2. The audio signal processing apparatus according to claim 1.
4. 2. The audio signal processing device according to claim 1, wherein the audio signal processing device that controls the loudness level of the input audio signal using the loudness metadata acquires the loudness difference changed by the post-processing based on information representing the loudness level for each band changed by the post-processing.
5. 5. The audio signal processing device according to claim 4, wherein the audio signal processing device that controls the loudness level of the input audio signal using the loudness metadata acquires the loudness difference changed by the post-processing based on an inner product of the loudness level for each band changed by the post-processing and the loudness ratio of each frequency of the input audio signal.
6. 6. The audio signal processing device of claim 5, wherein the loudness difference modified by the post-processing is a parameter obtained by the audio signal processing device using the loudness metadata to control the loudness level of the input audio signal based on perceptual loudness characteristics.
7. 7. The audio signal processing device of claim 6, wherein the audio signal processing device that controls the loudness level of the input audio signal using the loudness metadata obtains the loudness difference modified by the post-processing based on a K-weighted filter.
8. 1. An audio signal processing device for controlling a loudness level, the audio signal processing device comprising: a processor for adjusting an output loudness level of an input audio signal, the processor comprising: receiving loudness metadata corresponding to the input audio signal; parsing the loudness metadata to obtain loudness information of the input audio signal, the loudness information including information representing loudness ratios per frequency of the input audio signal; obtaining a post-processed modified loudness difference based on the loudness ratio for each frequency of the input audio signal; adding the post-processed modified loudness difference to a cumulative loudness level to obtain a modified cumulative loudness level; determining a loudness gain of the input audio signal based on a relationship between the modified cumulative loudness level and a target loudness level, the relationship comprising a difference or ratio between the modified cumulative loudness level and the target loudness level; adjusting an output loudness level of the input audio signal based on the loudness gain. Audio signal processing device.
9. The audio signal processing apparatus of claim 8 , wherein the processor outputs the generated output audio signal by applying a loudness limiter to the input audio signal.
10. The audio signal processing device of claim 8 , wherein the post-processing includes at least one of equalization, reverberation, and spatial compensation.
11. The post-treatment is applying an output characteristic of an audio signal processing device using the loudness metadata to control the loudness level of the input audio signal. The audio signal processing device according to claim 8 .
12. the processor obtains the loudness difference through the post-processing based on information representing the loudness level for each band changed by the post-processing. The audio signal processing device according to claim 8 .
13. The audio signal processing device according to claim 12 , wherein the processor obtains the loudness difference due to the post-processing based on an inner product of the loudness level for each band changed by the post-processing and a loudness ratio of each frequency of the input audio signal.
14. The audio signal processing apparatus according to claim 13 , wherein the loudness difference due to the post-processing is a parameter obtained based on a perceptual loudness characteristic.
15. The audio signal processing device of claim 14 , wherein the processor obtains the post-processed loudness difference based on a K-weighted filter.
16. 1. A method for generating loudness metadata for an input audio signal by an audio signal processing apparatus, comprising: receiving an input audio signal; measuring the loudness of the input audio signal to obtain loudness information of the input audio signal; transforming the loudness information to generate the loudness metadata; outputting the generated loudness metadata to an output device for outputting the input audio signal via an output unit, the loudness information including information representing a loudness ratio for each frequency of the input audio signal; the loudness ratios for each frequency of the input audio signal are used by an audio signal processing device that controls the loudness level of the input audio signal using the loudness metadata to obtain a loudness difference modified by post-processing; method.
17. 1. A method for adjusting an output loudness level of an input audio signal by an audio signal processing apparatus, comprising: receiving loudness metadata corresponding to the input audio signal; parsing the loudness metadata to obtain loudness information of the input audio signal, the loudness information including information representing loudness ratios per frequency of the input audio signal; obtaining a post-processed modified loudness difference based on the loudness ratio for each frequency of the input audio signal; adding the post-processed modified loudness difference to a cumulative loudness level to obtain a modified cumulative loudness level; determining a loudness gain of the input audio signal based on a relationship between the modified cumulative loudness level and a target loudness level, the relationship comprising a difference or ratio between the modified cumulative loudness level and the target loudness level; adjusting an output loudness level of the input audio signal based on the loudness gain. method.
Citation Information
Patent Citations
Loudness range control system, transmitting device, receiving device, transmitting program and receiving program
JP2013157659A
Loudness level and range processing
US9565508B1
Gain control apparatus and gain control method, and voice output apparatus
WO2010131470A1