Audio signal processing device and program
The audio signal processing device controls audio signal levels through a rendering and gain application system to prevent distortion in object-based audio systems, addressing the challenge of diverse playback environments and viewer preferences.
Patent Information
- Application Number
- JP2024141183
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Existing audio signal processing systems for object-based audio fail to prevent audio distortion in diverse playback environments due to insufficient dynamic range control, especially when multiple audio signals are mixed, and the dynamic range of audio objects is difficult to determine in advance.
An audio signal processing device that includes a rendering unit, gain calculation unit, and gain application unit to generate and apply gains to audio signals based on acoustic metadata and user settings, ensuring the signals do not exceed a threshold, thereby preventing audio distortion across various playback environments.
The device effectively limits audio signal levels to prevent distortion regardless of viewer preferences or playback equipment capabilities, ensuring high-quality audio reproduction.
Smart Images

Figure 2026037866000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio signal processing device and a program. [Background technology]
[0002] In object-based audio, an "object" refers to the audio material that makes up a program's audio, such as background sounds and dialogue. Object-based audio program audio signals are created for each audio object, and are recorded or transmitted along with audio metadata that describes how the audio objects are combined to generate the program audio, as well as the playback position and volume of each audio object. When played back at home, a playback device called a renderer generates (renders) the actual playback signal based on the audio metadata, playback environment, and viewer preferences. In other words, object-based audio playback signals vary depending on the playback environment. In recent years, there has been a global movement to introduce object-based audio into broadcasting. For example, the International Telecommunication Union (ITU-R) has defined the international audio metadata standard, the Audio Definition Model (ADM) (see, for example, Non-Patent Document 1), and a production renderer compatible with ADM, commonly known as the ADM renderer (see, for example, Non-Patent Document 2). Other standardization organizations have also standardized audio coding methods compatible with object-based audio, such as MPEG-H 3D Audio and AC-4.
[0003] On the other hand, audio levels for programs currently produced for broadcast are typically controlled very strictly to prevent distortion. In the current audio system (channel-based audio), the audio signals produced at the broadcasting station and those played back at home are essentially the same, meaning that if there is no audio distortion at the broadcasting station, there will be no audio distortion at home. Therefore, multiple distortion prevention measures are implemented to prevent audio distortion during program production at the broadcasting station. There are several stages in program production where audio distortion is likely to occur, but the most stringent level control is required at the digital audio signal processing stage, where audio distortion will definitely occur if the audio level exceeds a certain value. For example, in an audio console that generates program audio by mixing multiple digital audio signals and processing them to change the timbre, it is common to insert an audio effector called a compressor or limiter, which controls the dynamic range (the ratio between the maximum and minimum values of the signal), in the channel that actually mixes the audio signals (known as the master bus, for example), to adjust the audio signal so that it does not exceed a certain level.
[0004] Digital audio signals are expressed as amplitudes by quantizing the range of -1 to +1 as antilogarithms according to the bit depth. For example, if quantization is at 24 bits, the range of -1 to +1 is expressed as 24 bits. 24When mixing multiple signals or increasing the audio level, if the absolute value of the amplitude exceeds 1, the amplitude cannot be represented digitally, resulting in audio distortion. To address this issue, professional equipment such as audio consoles often convert digital audio signals to 32-bit or 64-bit floating-point values in their signal processing sections. This allows the amplitude to be properly represented digitally, even if the absolute value exceeds 1 as a result of signal processing, and audio distortion does not occur at this stage. However, typical digital audio interfaces can only handle audio signals represented as 16-bit or 24-bit integer values. When outputting digital audio signals to a device that has undergone signal processing, audio distortion occurs when the floating-point audio signal is converted to an integer value if the amplitude exceeds 1. When a limiter is inserted into the master bus of an audio console, even if the amplitude exceeds 1 after mixing multiple audio signals, audio distortion can be avoided by lowering the audio level so that the amplitude is below 1 before outputting the signal from the audio console. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] “Audio Definition Model”, Recommendation ITU-R BS.2076-02, 10 / 2019 [Non-patent document 2] “Audio Definition Model renderer for advanced sound systems”, Recommendation ITU-R BS.2127-0,06 / 2019 Summary of the Invention [Problem to be solved by the invention]
[0006] Because object-based audio playback signals are generated (rendered) in playback environments such as homes based on audio object-specific audio signals—that is, audio signals before they are mixed on the master bus—applying limiters or other controls on the master bus during program production, as is the case with current methods, cannot prevent audio distortion in playback environments (the master bus signal is not broadcast or distributed). Furthermore, even if limiters or other controls are applied on an audio object-by-audio object basis, the possibility cannot be denied that the audio level may exceed the level at which audio distortion occurs due to the rising audio level caused by mixing multiple audio signals. Furthermore, considering that object-based audio allows viewers to change rendering settings according to their preferences on the playback side, that audio metadata contains multiple program configurations, and that the same audio object may be reused multiple times between the main audio and one or more secondary audio tracks, it is difficult to determine the extent to which the dynamic range of the audio signal for each audio object should be controlled (audio level limitation) in advance of program production.
[0007] Audio distortion can be prevented if the signal processing device (renderer) that generates the playback signal in the playback environment supports floating-point signal processing and has a built-in limiter function. However, with the diversification of program playback environments and equipment in recent years, not all playback equipment has sufficient signal processing capabilities, so in order to prevent audio distortion in any playback environment, it is necessary to control the dynamic range in the production environment.
[0008] The object of the present invention, which has been made in consideration of the above circumstances, is to provide an audio signal processing device and program that can limit the level of an audio signal in advance so that audio distortion does not occur regardless of how rendering is performed on the playback side according to the viewer's preferences. [Means for solving the problem]
[0009] The gist of the present invention for solving the above problems is as follows.
[0010] (1) An audio signal processing device for multi-channel digital audio signals, comprising: a rendering unit that generates main and secondary audio signals consisting of a main audio signal and one or more secondary audio signals based on acoustic metadata and the digital audio signal; a gain calculation unit that calculates gains for the main and secondary audio signals to attenuate them so that they do not exceed a threshold value for audio samples over the entire time; and a gain application unit that applies the gains to the digital audio signals on an audio object basis and outputs the signals.
[0011] (2) The audio signal processing device described in (1), wherein the rendering unit downmixes the digital audio signal to a speaker arrangement with the smallest number of speakers in the playback environment and applies the maximum gain that can be set in the playback environment.
[0012] (3) The audio signal processing device according to (1) or (2), wherein the rendering unit has a plurality of renderers, each of which generates the main audio signal or the secondary audio signal.
[0013] (4) An audio signal processing device according to any one of (1) to (3), wherein the gain calculation unit calculates the minimum gain for the main and secondary audio signals for audio objects commonly used between the main and secondary audio signals, calculates a gain for the main audio signal for audio objects other than the commonly used audio objects used in generating the main audio signal, and calculates a gain for the secondary audio signal for audio objects other than the commonly used audio objects used in generating the secondary audio signal.
[0014] (5) The audio signal processing device according to any one of (1) to (3), wherein the gain calculation unit calculates a gain for each audio object of the main and secondary audio signals to attenuate the gain so that the threshold value is not exceeded in audio samples over the entire time period, and then calculates the minimum gain for an object commonly used between the main and secondary audio signals for that audio object, calculates a gain for the main audio signal for an audio object other than the commonly used audio object used in generating the main audio signal, and calculates a gain for the audio object of the secondary audio signal for an audio object other than the commonly used audio object used in generating the secondary audio signal.
[0015] (6) A program for causing a computer to function as the audio signal processing device according to any one of (1) to (5). [Effects of the Invention]
[0016] According to the present invention, in an object-based audio program, it is possible to limit the level of the audio signal in advance so that audio distortion does not occur regardless of how rendering is performed on the playback side according to the viewer's preferences. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a block diagram illustrating an example of the configuration of an audio signal processing device according to an embodiment. [Figure 2] 10 is a flowchart illustrating an example of a processing procedure of an audio signal processing device according to an embodiment. [Figure 3] FIG. 2 is a block diagram illustrating a detailed configuration example of an audio signal processing device according to an embodiment. [Figure 4] 1 is a block diagram showing an example of the configuration of a DRC gain calculation unit in an audio signal processing device according to an embodiment. FIG. [Figure 5] 10A and 10B are diagrams illustrating processing by a static characteristic value calculation unit in the audio signal processing device according to an embodiment. [Figure 6] 5A and 5B are diagrams illustrating processing by a gain smoothing unit in the audio signal processing device according to an embodiment. [Figure 7] FIG. 1 is a diagram illustrating an example in which an audio signal processing device according to an embodiment is implemented in a digital audio console. [Figure 8] FIG. 10 is a diagram illustrating an example of gain when DRC is performed on only one of two audio objects. [Figure 9] FIG. 3 is a diagram illustrating processing by a rendering unit in the audio signal processing device according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating processing by a rendering unit in the audio signal processing device according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0019] 1 is a block diagram showing an example of the configuration of an audio signal processing device 1 according to this embodiment. The audio signal processing device 1 shown in FIG. 1 includes a rendering setting unit 10, a rendering unit 20, a gain calculation unit 30, and a gain application unit 40.
[0020] The audio signal processing device 1 is a device that prevents signal distortion during playback by controlling the dynamic range in the production of object-based audio programs. The audio signal processing device 1 distributes the input digital audio signal in audio object units and inputs it to the rendering unit 20 and the gain application unit 40.
[0021] Prior to signal processing, the rendering setting unit 10 performs rendering settings for the rendering unit 20 based on the audio metadata and user settings input to the rendering setting unit 10 .
[0022] The rendering unit 20 generates (renders) a main audio signal and one or more secondary audio signals (hereinafter collectively referred to as "main and secondary audio signals") based on the input acoustic metadata and digital audio signal. The rendering unit 20 then outputs the generated main and secondary audio signals to the gain calculation unit 30.
[0023] The gain calculation unit 30 calculates the dynamic range control (DRC) gain required to prevent audio distortion. The DRC gain is a gain that attenuates the main and secondary audio signals so that the audio samples do not exceed a threshold value over the entire time period.
[0024] The gain application unit 40 applies the gain calculated by the gain calculation unit 30 to the digital audio signal input to the audio signal processing device 1 on an audio object basis, and outputs the result to the outside of the audio signal processing device 1.
[0025] 2 is a flowchart showing an example of a processing procedure of an audio signal processing device according to one embodiment of the present invention. Each step of the processing will be described below with reference to FIG.
[0026] In step S101, the rendering setting unit 10 analyzes the audio metadata and user settings input to the rendering setting unit 10. In the analysis, in addition to the analysis required for normal object-based audio rendering, information for audio signal processing is also acquired. Regarding the analysis required for normal rendering, depending on the usage status of the audio signal processing device 1, information required by the production renderer described in Non-Patent Document 2 may be acquired in accordance with the renderer (input by user settings) expected to be used in the playback environment, or information required by the MPEG-H 3D Audio or AC-4 renderer may be acquired.
[0027] The information for audio signal processing includes the number of all objects included in the audio metadata, the total number of main audio and secondary audio included in the audio metadata, the speaker arrangement expected in the playback environment, and information on whether the user is permitted to control the gain of each object. The audio metadata may be the ADM described above, or may be metadata such as MPEG-H 3D Audio or AC-4.
[0028] In step S102, the rendering setting unit 10 sets the renderer of the rendering unit 20 based on the results of analyzing the acoustic metadata and the user settings.
[0029] In step S103, the rendering unit 20 performs at least one type of rendering in real time. The rendering unit 20 may have multiple renderers so that multiple types of rendering can be performed in parallel at the same time.
[0030] In the rendering settings, the renderer expected to be used in the playback environment is set so that the audio signal level is maximized based on the speaker arrangement of the expected playback environment. For example, when playing multi-channel audio in stereo, the fewer channels are downmixed, the more audio is distorted. Also, the greater the gain of the object on the playback side, the more audio is distorted. By reproducing rendering for the environment most susceptible to audio distortion, the rendering unit 20 prevents audio distortion in any playback environment.
[0031] Specifically, the rendering setting unit 10 performs rendering settings assuming that the speaker arrangement is converted to a speaker arrangement with the minimum number of speakers in the playback environment, and that the viewer's preferences in the playback environment are manipulated to increase the object gain to its maximum value (the upper limit can be specified in the audio metadata). That is, the rendering unit 20 downmixes the digital audio signal to a speaker arrangement with the minimum number of speakers in the playback environment and applies the maximum gain that can be changed in the playback environment. Typically, a speaker arrangement with the minimum number of speakers is stereo. When playing outside the home, such as for public viewing, the speaker arrangement with the minimum number of speakers may not be stereo but may be 5.1ch or the like. Furthermore, with extremely inexpensive speakers available on the market, the left and right channels of a stereo signal may be mixed to produce mono.
[0032] However, for some renderers, such as those in which the renderer for the playback environment is MPEG-H 3D Audio, corrections are made based on the correlation of audio signals mixed when the speaker layout is changed so that the audio level does not change even when the speaker layout is changed in the renderer to one that differs from that in the production environment. In such cases, the rendering unit 20 also sets the speaker layout that is the same as that at the time of production, rather than the speaker layout that minimizes the number of speakers.
[0033] Furthermore, with regard to the gain of an object, some renderers, such as those using MPEG-H 3D Audio in the playback environment, may apply corrections so that the loudness of the total audio signal does not change when the viewer operates the gain of a specific object. In such cases, the rendering unit 20 also uses the default gain described in the audio metadata rather than changing the gain of the object to its maximum.
[0034] Whether or not correction is to be applied when speaker placement is changed and when the viewer operates the gain is set by a user input to the audio signal processing device 1. Other possible user inputs include information required for general digital audio signal processing, such as the buffer size of the device, the sampling frequency of the digital audio signal, and the bit depth.
[0035] After the renderers are set up and ready, real-time DRC processing (audio signal processing) can begin. In real-time processing, the rendering unit 20 first converts the audio signal into a floating-point data type, and then a number of renderers equal to the total number of main audio signals and secondary audio signals to be generated simultaneously perform rendering.
[0036] In step S104, the gain calculation unit 30 calculates the DRC gain required to prevent audio distortion. The process of the gain calculation unit 30 will be described in detail later.
[0037] In step S105, the gain application unit 40 applies the gain calculated by the gain calculation unit 30 to the digital audio signal input to the audio signal processing device 1 on an audio object basis.
[0038] In step S106, the audio signal processing device 1 determines whether or not an instruction to end the real-time audio signal processing has been received, and repeats the processes from step S103 to step S105 until an instruction to end the processing is received.
[0039] 3 is a block diagram showing a detailed configuration example of the rendering unit 20, the gain calculation unit 30, and the gain application unit 40. The rendering unit 20 has M renderers. The gain calculation unit 30 has M dB conversion units 31 (31-1 to 31-M), M DRC gain calculation units 32 (32-1 to 32-M), N gain determination units 33 (33-1 to 33-N), and N true value conversion units 34 (34-1 to 34-N). Here, M is the total number of main audio signals and secondary audio signals to be generated, and N is the number of objects. For example, when background sound objects, dialogue (Japanese) objects, and dialogue (English) objects are input to the audio signal processing device 1, and the main audio signal is reproduced by rendering the background sound objects and the dialogue (Japanese) objects, and the secondary audio signal is reproduced by rendering the background sound objects and the dialogue (English) objects, N=3 and M=2.
[0040] The dB conversion unit 31 converts each rendered audio signal acquired from each renderer of the rendering unit 20 into a decibel value, and outputs the converted signal to the DRC gain calculation unit 32 .
[0041] Fig. 4 is a block diagram showing an example configuration of the DRC gain calculation unit 32. The DRC gain calculation unit 32 shown in Fig. 4 includes a static characteristic value calculation unit 321, a difference calculation unit 322, a gain smoothing unit 323, and a gain addition unit 324. In the DRC, parameters that are set in advance by the user include a threshold, a soft knee width, an attack time, a release time, and a make-up gain.
[0042] 5 is a diagram illustrating the processing of the static characteristic value calculation unit 321. The static characteristic value calculation unit 321 outputs (output audio) without changing the audio level of the input audio signal (input audio) up to a certain value (Threshold). For input audio that exceeds the Threshold, the static characteristic value calculation unit 321 lowers the audio level so that the output audio does not exceed the Threshold. Furthermore, to avoid sudden changes in values around the Threshold, the static characteristic value calculation unit 321 also smooths changes in input and output according to the Soft Knee width ratio.
[0043] The difference calculation unit 322 calculates the difference between the input sound and the output sound of the static characteristic value calculation unit 321 in order to calculate an adjustment value (gain) for the sound level.
[0044] FIG. 6 is a diagram illustrating the processing of the gain smoothing unit 323. The gain smoothing unit 323 smooths the gain in the time direction. The gain smoothing unit 323 does not make a sudden change from when the input signal exceeds the threshold and it becomes necessary to suppress the audio level until the gain actually reaches that level, but rather changes it over a certain amount of time. This time is set as the attack time. The gain smoothing unit 323 also does not make a sudden change from when the input signal falls below the threshold and it is no longer necessary to use gain to suppress the audio level until it returns to the original gain, or until it reaches a large gain (since the gain is basically negative, a larger gain is closer to 0 dB = no adjustment), but rather changes it over a certain amount of time. This time is set as the release time.
[0045] Since DRC is generally based on processing to suppress audio levels, the overall audio level may seem low if the threshold is set low so that DRC is applied immediately. To prevent this, the gain adder 324 increases the overall calculated and smoothed gain by the value of the make-up gain. However, when the threshold is set to a high value (close to 0 dB), such as in a limiter, the make-up gain is often not used and is set to 0 dB.
[0046] The processing of the DRC gain calculation unit 32 described above is the same as that of existing DRC calculations. Any DRC and algorithm can be used as long as it can suppress the main and secondary audio signals after rendering to an audio level that does not cause audio distortion. Therefore, a limiter or compressor calculation method different from the above may be used to achieve the quality or sound quality characteristics desired for the audio signal processing device 1.
[0047] Referring back to Figure 3, after the DRC gain calculation unit 32 calculates the post-DRC gain (DRC gain) for each audio signal, the gain determination unit 33 finds the minimum DRC gain for each audio object and determines the gain to be applied to that audio object. In object-based audio, the same audio object is often reused between different main and secondary audio signals, such as background sound objects. Therefore, in order to prevent audio distortion in any of the main and secondary audio signals, it is desirable to select the DRC gain of the reused common audio object (common object) so that the controlled gain value is minimized (the signal level of the processed audio object is minimized) based on the audio signal rendered so that the audio signal level of the main and secondary audio signals is the highest. For audio objects other than common objects, the DRC gain of the main and secondary audio signals is used as the gain of that audio object as is.
[0048] That is, the gain determination unit 33 determines the smallest gain (gain with the largest attenuation) for the main and secondary audio signals as the gain to be applied to a common object used in common between the main and secondary audio signals. The gain determination unit 33 also determines the gain for the main audio signal as the gain to be applied to an audio object other than the common object that is used to generate the main audio signal. The gain determination unit 33 also determines the gain for a certain secondary audio signal as the gain to be applied to an audio object other than the common audio object that is used to generate the secondary audio signal.
[0049] The true value conversion unit 34 converts the gain determined by the gain determination unit 33 from decibels back to a true value, and outputs the converted value to the gain application unit 40 .
[0050] The gain application unit 40 applies the gain calculated by the gain calculation unit 30 to the original digital audio signal input to the audio signal processing device 1 and outputs the result. That is, the gain application unit 40 applies the smallest gain for the main and secondary audio signals to the common object. Furthermore, the gain application unit 40 applies the gain for the main audio signal to audio objects other than the common object used in generating the main audio signal, and applies the gain for a certain secondary audio signal to audio objects other than the common object used in generating the certain secondary audio signal.
[0051] For example, the rendering unit 20 generates a main audio signal using a common object and a dialogue (Japanese) audio object, and generates a secondary audio signal using the common object and a dialogue (English) audio object. The gain calculated by the DRC gain calculation unit 32 for the main audio signal is g1, the gain calculated for the secondary audio signal is g2, and g1 is smaller than g2 (the amount of attenuation is greater). In this case, the gain application unit 40 applies a gain g', which is a true value of g1, to the common object and the dialogue (Japanese) audio object. Lin-1 For the dialogue (English) audio object, the gain g' is applied, which is the true value of g2. Lin-2 applies.
[0052] In the series of processes from gain calculation in the gain calculation unit 30 to gain application in the gain application unit 40, it is assumed that the gain is calculated for each audio sample, as in a general limiter or compressor, but the same gain may be applied to a certain number of consecutive samples to reduce the amount of processing when developing inexpensive equipment, etc. Furthermore, when an audio object is made up of multiple audio channels, different gains may be applied to each channel of the audio object, or the same gain may be applied to all channels, depending on the gain calculation method used.
[0053] Furthermore, when simultaneously generating multiple main and secondary audio signals, the rendering unit 20 may share processing related to common objects rather than performing rendering processes completely in parallel. This eliminates the need for the rendering unit 20 to perform the same processing in each rendering of the main and secondary audio signals when generating multiple main and secondary audio signals, which is expected to have the effect of reducing the amount of processing.
[0054] 7, the audio signal processing device 1 can be implemented as a DRC processing unit of a digital audio console 100. The audio signal processing device 1 can also be implemented as one function of a digital audio signal output board of a computer.
[0055] (Variation) Although the above-described DRC gain calculation unit 32 calculates a single DRC gain for each of the main and secondary audio signals, the DRC gain calculation unit 32 may calculate a DRC gain for each object constituting the main and secondary audio signals. As a specific example, a case will be described in which a background sound object, a dialogue (Japanese) object, and a dialogue (English) object are input to the audio signal processing device 1, the main audio signal is played back by rendering the background sound object and the dialogue (Japanese) object, and the secondary audio signal is played back by rendering the background sound object and the dialogue (English) object.
[0056] In this case, the audio level obtained by adding a background sound object to which a DRC gain for background sound has been applied and a dialogue (Japanese) object to which a DRC gain for dialogue (Japanese) has been applied must be the same as the audio level of a main audio signal to which a single DRC gain for the main audio has been applied. Similarly, for the secondary audio, a DRC gain for background sound and a DRC gain for dialogue (English) are calculated. The gain determination unit 33 compares the DRC gain for background sound calculated from the main audio signal with the DRC gain for background sound calculated from the secondary audio signal to find the minimum value.
[0057] Figure 8 shows an example of calculating the gain when performing DRC on only one of two audio objects (background sound and dialogue in this example). The DRC gain (G A ,G B [dB]) is the level difference (Diff [dB]) between the single DRC gain (G [dB]) calculated by the DRC gain calculation unit 32 and the level difference (Diff _AB The level difference of the audio object B relative to the audio object A (Diff _BA When expressed in terms of dB, it takes the same form as equation (1), and equations (1) and (2) are essentially the same.
number
number
[0058] Two DRC gains G A ,G B Contribution rate CR of how much the signal contributes to the overall attenuation of the audio signal A , C.R. B [-] is the audio level L of each audio object A , L B [dB], and the audio level L when the two audio objects are combined to form the program audio. M The DRC gain is expressed by equations (3) and (4) using [dB], and the relationship is expressed by equation (5). When calculating the DRC gain for each object, the contribution rate of the DRC gain may be set to any value. A ,CR B )=(1,0), the case where the level of the audio object A is higher is shown by a solid line, and the case where the level of the audio object A is lower is shown by a dashed line. A ,CR B ) as in equation (6), G A =G BThis corresponds to the above example where a single DRC gain is calculated for each of the main and secondary audio signals.
number
number
number
number
[0059] That is, the gain calculation unit 30 calculates a gain for the main and secondary audio signals, which is used to attenuate the audio samples over the entire time so that the gain does not exceed a threshold, based on the contribution of the attenuation amount due to the gain to the attenuation amount of the entire audio signal, for each audio object in the main and secondary audio signals. Then, for an audio object used in common between the main and secondary audio signals, the gain calculation unit 30 calculates the minimum gain for that audio object in the main and secondary audio signals. Furthermore, for an audio object other than the commonly used audio object used in generating the main audio signal, the gain calculation unit 30 calculates a gain for that audio object in the main audio signal. Furthermore, for an audio object other than the commonly used audio object used in generating the secondary audio signal, the gain calculation unit 30 calculates a gain for that audio object in the secondary audio signal.
[0060] When applying DRC gain only to audio objects with high audio levels, the DRC gain for those objects is kept relatively small compared to when applying DRC gain only to audio objects with low audio levels, minimizing the impact on program audio. On the other hand, when applying DRC gain only to audio objects with low audio levels, audio distortion can be suppressed without changing audio objects with high audio levels that are likely to be more prominent in the program. Furthermore, when using the DRC gain of only one of the audio objects in this way, the audio object to which the DRC gain is applied may be selected based on the content of the object, such as background sound or dialogue, rather than the audio level.
[0061] (First embodiment) As a first embodiment of the present invention, we will explain an example in which the audio metadata is ADM, the algorithm of the rendering unit 20 in the playback environment is assumed to be an ADM renderer, and multiple audio signals are generated based on user input without any correction for speaker placement changes or gain operations performed by the viewer, but the processing related to common objects that are reused is standardized. Below, the alphabetical characters enclosed in " " are XML descriptors written in the ADM.
[0062] First, the rendering setting unit 10 analyzes the acoustic metadata (ADM) and user settings input to the audio signal processing device 1. In addition to the analysis required for rendering by the ADM renderer described in Non-Patent Document 2, the rendering setting unit 10 also acquires information for the audio signal processing of the present invention.
[0063] The rendering setting unit 10 acquires from the ADM a list of "audioProgramme" indicating the main audio and secondary audio, and "audioObject" indicating the audio objects, and specifies the number of "audioProgramme" (the total number M of main audio signals and secondary audio signals) and the number of "audioObject" (the number N of objects). In addition, the number of samples S of the processing buffer (input / output buffer), the sampling frequency Fs, etc. are acquired from user input.
[0064] Next, the rendering setting unit 10 obtains the speaker layout assumed in the playback environment, i.e., the playback speaker layout to be used in the calculation. If speaker layout conversion correction is not performed, the smallest assumed speaker layout is obtained. Specifically, if "authoringInformation," which is information about the production environment, is present as a child element of the first "audioProgramme" in the ADM, the "audioPackFormat" (audio format information) with the smallest number of channels is obtained from the speaker layouts described in "audioProgramme / authoringInformation / referenceLayout / audioPackFormatIDRef" and "audioProgramme / authoringInformation / renderer / audioPackFormatIDRef." The number of channels in each "audioPackFormat" is equal to the number of child elements "audioChannelFormatIDRef" (audio channel information). If the ADM "audioProgramme / authoringInformation" does not exist, the "audioPackFormatID" is set to "AP_00010002", i.e., stereo. The audio formats indicated by each ADM ID are specified in Recommendation ITU-R BS.2094 "Common definitions for the Audio Definition Model."
[0065] The number of channels is identified from the acquired playback speaker layout "audioPackFormat." Also, a channel weighting coefficient array (array of the number of channels in the speaker layout before conversion x the number of channels in the speaker layout after conversion) used in speaker layout conversion for each object, which is determined for each renderer (and can be stored by the audio signal processing device 1), is identified. However, this does not include audio objects with "typeDefinition"="object", i.e., audio objects whose playback position may change dynamically within the content.
[0066] Next, the rendering setting unit 10 specifies the gain that can be added for each object under the control of the viewer. If no correction is performed, and if each "audioObject" has an upper limit value for the gain related to interaction restrictions, "audioObject / audioObjectInteraction / gainInteract / gainInteractionRange / bound"="max", this is specified as the gain that can be added; if not, the true value 1.0 (0 dB) is used.
[0067] The above operations determine the values of each variable required by the rendering unit 20 and the size of the array used for signal processing. By creating a cell array for storing the signal being processed based on the determined size, it is possible to calculate the memory size and processing load required for real-time DRC processing (audio signal processing) before processing begins.
[0068] Up to this point, the setting and preparation of the rendering unit 20 is completed, and then DRC processing (audio signal processing) is started in real time in response to a user instruction or the like.
[0069] Fig. 9 shows the rendering process of the rendering unit 20 when the processes for converting a common object to the minimum speaker arrangement possible on the playback side and applying the maximum gain that the viewer can set for the audio object are standardized for the common object. In this example, the common object is a background sound object. When the processing load calculated in advance is high, the processing load can be reduced by standardizing the processes as shown in Fig. 9.
[0070] The input multi-channel audio signal is divided into a specific windowed time width for each "audioObject". The size of the divided data is the number of channels of the audio object x the windowed time width (number of samples), and can be specified in advance based on settings such as the buffer size.
[0071] Each audio object is converted to the specified speaker layout after the speaker layout conversion and multiplied by the maximum gain value that can be set by the viewer. Then, for each "audioProgramme", the audio signals of the audio objects that make up the main audio signal or secondary audio signal are added together. However, for audio objects with "typeDefinition"="object", the signals are distributed to the speaker layout expected in the playback environment according to their playback positions "audioChannelFormat" / "audioBlockFormat" / "position" and the 3D panning method of the rendering unit 20. The subsequent application of the maximum gain value that can be set by the viewer and signal addition are the same as for objects other than those with "typeDefinition"="object". The processing from the rendering unit 20 onwards is as described above.
[0072] (Second embodiment) As a second embodiment of the present invention, we will explain an example in which the audio metadata is ADM, the algorithm of the rendering unit 20 in the playback environment is assumed to be MPEG-H 3D Audio, and the user inputs speaker placement transformations and corrections when the viewer performs gain operations, generating multiple audio signals, but not sharing the processing related to common objects that are reused.
[0073] For speaker layout conversion information, the same speaker layout as that used during content creation (when the object-based audio program audio (digital audio signal and audio metadata) was created) is obtained. In other words, downmixing is not performed. Specifically, if "authoringInformation" exists as a child element of the first "audioProgramme" in the ADM, the "audioPackFormat" referenced by "audioProgramme / authoringInformation / referenceLayout / audioPackFormatIDRef" is obtained. If the ADM does not have "audioProgramme / authoringInformation," the "audioPack" referenced by the "audioObject" with the largest number of channels among all the "audioObjects" contained in the ADM is obtained. The number of channels in an "audioObject" is equal to the number of "audioTrackUIDs" (unique IDs of audio tracks in the audio metadata) that are child elements of "audioObject".
[0074] Furthermore, for gains that can be added under viewer control, if no correction is performed, a true value of 1.0 (0 dB) is used, meaning that no substantial gain is added.
[0075] 10 shows the rendering process of the rendering unit 20 of the second embodiment. The processes other than those described above are the same as those of the first embodiment, and therefore will not be described again.
[0076] As described above, the audio signal processing device 1 generates main and secondary audio signals consisting of a main audio signal and one or more secondary audio signals based on acoustic metadata and a digital audio signal, calculates gains for attenuating the main and secondary audio signals so that they do not exceed a threshold value for all audio samples over the entire time period, and applies these gains to the digital audio signals before outputting them. Therefore, according to the present invention, it is possible to control the dynamic range of the audio signal on an audio object-by-audio object basis so that audio distortion does not occur regardless of how rendering is performed on the playback side according to the viewer's preferences.
[0077] (program) A computer capable of executing program instructions can also be used to function as the above-described audio signal processing device 1. Here, the computer may be a general-purpose computer, a special-purpose computer, a workstation, a PC (Personal Computer), an electronic notepad, etc. The program instructions may be program code, code segments, etc. for performing the necessary tasks.
[0078] The computer includes a processor, a storage unit, an input unit, and an output unit. The processor may be a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), SoC (System on a Chip), etc., and may be composed of multiple processors of the same or different types. The processor reads and executes programs from the storage unit to control the above components and perform various arithmetic processing. Note that at least a portion of these processing contents may be realized by hardware. The input unit is an input interface that accepts user input operations and acquires information based on the user operations, and may be a pointing device, keyboard, microphone, etc. The output unit is an output interface that outputs information, and may be a display, speaker, etc.
[0079] The program may be recorded on a computer-readable recording medium. Using such a recording medium, the program can be installed on a computer. Here, the recording medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a CD-ROM, a DVD-ROM, or a USB (Universal Serial Bus) memory. Furthermore, the program may be downloaded from an external device via a network.
[0080] The audio signal processing device 1 may be configured with one or more semiconductor chips. The semiconductor chip may include a CPU that executes a program that describes the processing contents for realizing each function of the modulation device 1.
[0081] Although the above-described embodiments have been described as typical examples, it will be apparent to those skilled in the art that many modifications and substitutions can be made within the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited by the above-described embodiments, and various modifications or alterations can be made without departing from the scope of the claims. For example, it is possible to integrate multiple building blocks shown in the block diagrams of the embodiments, or to divide one building block. [Explanation of symbols]
[0082] 1 Audio signal processing device 10 Rendering Settings 20 Rendering Section 30 Gain calculation section 31 dB converter 32 DRC gain calculation section 33 Gain determination unit 34 True value conversion section 34 Value conversion part 40 Gain application section 100 Digital Audio Console 321 Static characteristic value calculation unit 322 Difference calculation part 323 Gain smoothing section 324 Gain Adder
Claims
1. An audio signal processing device for multi-channel digital audio signals, comprising: a rendering unit that generates a main / secondary audio signal consisting of a main audio signal and one or more secondary audio signals based on the acoustic metadata and the digital audio signal; a gain calculation unit that calculates a gain for attenuating the main and secondary audio signals so that the gain does not exceed a threshold value for audio samples over the entire time period; a gain application unit that applies the gain to the digital audio signal on an audio object basis and outputs the result; An audio signal processing device comprising:
2. The audio signal processing device according to claim 1 , wherein the rendering unit downmixes the digital audio signal to a speaker arrangement with a minimum number of speakers in a playback environment, and applies a maximum gain that can be set in the playback environment.
3. The audio signal processing device according to claim 1 , wherein the rendering unit includes a plurality of renderers, each of which generates the main audio signal or the secondary audio signal.
4. the gain calculation unit calculates a minimum gain for the main and secondary audio signals for an audio object commonly used between the main and secondary audio signals; Calculating a gain relative to the main audio signal for audio objects other than the commonly used audio object used in generating the main audio signal; The audio signal processing device according to claim 1 , further comprising: calculating a gain for the secondary audio signal for an audio object other than the commonly used audio object that is used in generating the secondary audio signal.
5. the gain calculation unit calculates a gain for each audio object of the main and secondary audio signals to attenuate the main and secondary audio signals so that the gain does not exceed a threshold value for all audio samples over the entire time period, and then calculating a minimum gain for an object commonly used between the main and secondary audio signals for the audio object; For audio objects other than the commonly used audio object used in generating the main audio signal, a gain of the main audio signal for the audio object is calculated; The audio signal processing device according to claim 1 , further comprising: calculating a gain of the secondary audio signal for an audio object other than the commonly used audio object that is used in generating the secondary audio signal.
6. A program for causing a computer to function as the audio signal processing device according to claim 1.
Citation Information
Patent Citations
ITRBS.2076-02,10/2019
ITRBS.2127-0,06/2019