Method for adjusting the dynamic range of an input audio signal, audio signal processing device, and storage medium

JP7909561B2Active Publication Date: 2026-08-21DOLBY LABORATORIES LICENSING CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024062458
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-10-12
Filing Date
2024-04-09
Publication Date
2026-08-21
Estimated Expiration
2033-05-02

Smart Images

  • Figure 0007909561000009
    Figure 0007909561000009
  • Figure 0007909561000010
    Figure 0007909561000010
  • Figure 0007909561000011
    Figure 0007909561000011
Patent Text Reader

Abstract

To provide a method, an audio signal processor, and a storage medium that enable delivery of audiovisual media in a bandwidth-saving manner.SOLUTION: In a single-mode decoding system 51, a bitstream P includes an encoded core signal Y∼, a multichannel coding parameter α, a preprocessing dynamic range control (DRC) parameter DRC2, and a compensated postprocessing DRC parameter DRC3, and a demultiplexer 70 placed at the input extracts these quantities. A core signal decoder 71 receives an encoded core signal Y∼ and outputs m-channel core signals Y (1≤m≤n). A DRC processor 74 outputs an intermediate signal YC, which is input to a parametric synthesis stage 72. The parametric synthesis stage 72 forms an n-channel linear combination of the m channels in the intermediate signal, and outputs a reconstructed n-channel audio signal X.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The inventions disclosed herein relate, in general terms, to audiovisual media distribution. Specifically, to an adaptive distribution format that enables seamless mode transitions between both high-bitrate and low-bitrate modes during decoding. The invention further relates to methods and devices for encoding and decoding signals using the distribution format. [Background technology]

[0002] Parametric stereo and multichannel coding methods are known to be scalable and efficient in terms of listening quality. This makes them particularly attractive for low-bitrate applications. However, if the bitrate limitations are temporary (e.g., network jitter, additional fluctuations), the full benefit of available network resources can be obtained by using adaptive distribution formats that employ relatively high bitrates under normal conditions and lower bitrates when network performance is insufficient.

[0003] Existing adaptive distribution formats and associated decoding techniques have been improved in terms of bandwidth efficiency, computational efficiency, error recovery, and algorithmic latency. Furthermore, in audiovisual media distribution, improvements have been made in how easily bitrate switching events are noticeable to those enjoying the decoded media. However, as legacy decoders are expected to continue to be used in parallel with newer dedicated equipment, there are limits to these potential improvements insofar as backward compatibility must be maintained.

[0004] Dynamic range control (DRC) methods, which ensure a more consistent dynamic range during the playback of audiovisual signals, are well known in the art. For an overview, please refer to Non-Patent Document 1 and the references cited therein. This method allows the receiver to adjust the dynamic range of the audiovisual signal to suit relatively unsophisticated playback equipment, while the signal itself is broadcast in full dynamic range for the benefit of more sophisticated equipment. A simple implementation of DRC uses a metadata field that encodes the gain factor in the range of 0 to 1. The decoder can choose whether or not to apply it.

[0005] Using known DRC techniques, encoded audiovisual signals can be transmitted with metadata that allows the user to compress or boost the playback dynamic range to suit their preferences, or to manually adjust the dynamic range to the available playback equipment. However, known DRC techniques are incompatible with adaptive bitrate coding techniques, and switching between the two bitrates sometimes results in dynamic range mismatches, especially with legacy equipment. This invention solves this problem. [Cross-references to related applications] This application claims priority to U.S. Provisional Patent Application No. 61 / 649,036 filed on 18 May 2012, U.S. Provisional Patent Application No. 61 / 664,507 filed on 25 July 2012, and U.S. Provisional Patent Application No. 61 / 713,005 filed on 12 October 2012, the above documents are incorporated herein by reference in their entirety. [Prior art documents] [Patent Documents]

[0006] [Non-Patent Document 1] T. Carroll and J. Riedmiller, “Audio for Digital Television”, published as chapter 5.18 of EA Williams et al. (eds.), NAB Engineering Handbook, 10th ed. (2007), Academic Press [Brief explanation of the drawing]

[0007] Embodiments of the present invention will be described with reference to the attached drawings. [Figure 1] This is a block diagram showing an audio encoding system according to one embodiment of the present invention. [Figure 2] Block diagram of an audio decoding system according to one embodiment of the present invention. [Figure 3] This is a block diagram showing an audio encoding system according to one embodiment of the present invention. [Figure 4] Block diagram of an audio decoding system according to one embodiment of the present invention. [Figure 5] This figure shows a part of the parametric analysis stage in an audio coding system. [Figure 6] Block diagram of an audio decoding system according to one embodiment of the present invention. [Figure 7] This is a block diagram showing an audio encoding system according to one embodiment of the present invention. [Figure 8] This figure shows the calculation of compensated post-processing DRC parameters based on pre-processing and post-processing DRC parameters that refer to time blocks of the same length. [Figure 9] This figure shows the calculation of compensated post-processing DRC parameters based on pre-processing and post-processing DRC parameters that reference time blocks of different lengths. [Figure 10] This is a block diagram showing an audio encoding system according to one embodiment of the present invention. [Figure 11]A diagram showing a part of the parametric synthesis stage in an audio decoding system. [Figure 12] A diagram showing a part of the parametric synthesis stage in an audio decoding system. [Figure 13] A block diagram showing an audio decoding system according to an embodiment of the present invention. All drawings are schematic diagrams, showing only the necessary parts for explaining the present invention, and other parts are omitted or only suggested. Unless otherwise specified, the same reference numerals indicate the same parts even in different drawings.

Embodiments for Carrying Out the Invention

[0008] I. Overview Here, an "audio signal" refers to a pure audio signal or the audio part of an audiovisual signal or a multimedia signal.

[0009] According to an embodiment of the present invention, a method and a device for enabling the distribution of audiovisual media in a bandwidth-saving manner are proposed. In particular, according to one embodiment, a coding format for audiovisual media distribution is proposed, which can output an audio part having a consistent conversation level for legacy receivers and more recent devices. In particular, according to one embodiment, a coding format having an adaptive bit rate is proposed. Here, the switching between two bit rates does not need to be accompanied by a sharp change in the conversation level. Otherwise, this will result in perceptible artifacts in the audio signal or the audio part of the signal during playback.

[0010] [[ID=*23]] According to an embodiment of the present invention, an encoding method, an encoder, a decoding method, a decoder, a computer program product, and a media coding format having the features described in the independent claims are provided.

[0011] It should be noted that there seems to be a formatting issue with the original text where line breaks are used in an inconsistent way within some paragraphs. I've tried to maintain the overall structure as accurately as possible during translation. Also, the tag seems to be a bit misaligned in the original text, but I've left it as is in the translation. If this is a formatting error in the source, it might need to be corrected in the original for better readability.According to an embodiment of the present invention, a decoding system is provided for reconstructing an n-channel audio signal X based on a bitstream P. The decoding system is operable at least in a parametric coding mode and includes the following: · Receiving the bitstream and a coded core signal Y ~ (Note: "~" is above "Y") and one or more multi-channel coding parameters, which are collectively represented by α; · A core signal decoder that receives the coded core signal and outputs an m-channel core signal, where 1 ≦ m < n; · A parametric synthesis stage that receives the core signal and the multi-channel coding parameters and outputs the n-channel signal by forming a linear combination of the channels of the core signal using a gain that depends on the multi-channel coding parameters.

[0012] In this first embodiment, the bitstream further includes one or more pre-processing DRC parameters DRC2, which quantitatively characterize the dynamic range limitation performed by the encoder that generated the bitstream. The decoding system is operable to cancel the encoder-side dynamic range limitation based on the pre-processing DRC parameters. Preferably, the signal is partitioned into time blocks, and the pre-processing DRC parameter DRC2 is determined by the resolution of one time block of the signal. Therefore, each value of the parameter DRC2 is applied to at least one time block, and it is possible to associate each time block with a value specific to that time block. Without departing from the scope of the present invention, the value of the parameter DRC2 may be constant over a plurality of consecutive blocks. For example, the value of the parameter DRC2 is updated only once per time frame, and thus includes a plurality of time blocks in which the parameter DRC2 is constant.

[0013] The advantage of the first embodiment is that the pre-processing DRC parameter DRC2 provides the decoding system with the option to recover the audio signal to its original dynamic range at a time when the encoder, for whatever reason, has limited (or compressed) the dynamic range. Recovery cancels the dynamic range limitation, i.e., expands (or boosts) the dynamic range. One reason for limiting the dynamic range in an encoder is to avoid clipping. Whether or not recovery is performed depends on factors such as manually input user input, automatically detected playback device characteristics, a target DRC level obtained from an external source, or other factors. The target DRC level represents a portion of the original post-processing dynamic range control (quantified by the post-processing DRC parameter DRC1) applied to the decoding system. It is represented by a parameter f∈[0,1] that modifies the magnitude of the applied DRC from DRC1 to f×DRC1 (logarithmic units).

[0014] In a simple implementation, the DRC2 parameter is encoded in the form of a broad-spectrum (i.e., broadband) gain factor expressed in logarithmic form as a positive dB value that quantifies the relative amplitude reduction already experienced by the signal. Therefore, if DRC2 = x > 0, the relative amplitude change on the encoder side is

number

[0015] Actual cancellation is performed completely or partially according to the target DRC level and the input DRC level (or decoder input DRC level), that is, the DRC level that the n-channel audio signal would have after reconstruction without dynamic range compression or dynamic range boosting. The input DRC level is the original dynamic range reduced by the amount corresponding to the preprocessing DRC parameter DRC2. The target DRC level may be the original dynamic range reduced by the amount corresponding to the product of the parameter f and the postprocessing DRC parameter DRC1, that is, f×DRC1 (in logarithmic units). In the above simple implementation, the condition f×DRC1 < DRC2 indicates partial cancellation, that is, cancellation only by the amount corresponding to DRC2 - f×DRC1 instead of DRC2. For example, when the target DRC level corresponds to the input DRC level (for example, the dynamic range of the audio signal originally encoded by an encoder that generates a bitstream), this can be expressed as f = 0, but in that case, complete cancellation by the amount of DRC2 is required. When the target DRC level is smaller than the input DRC level, as in the case of 0 < f < 1 and f×DRC1 < DRC2, it is sufficient to partially cancel the dynamic range limitation. When the target DRC level is larger than the input DRC level, that is, f×DRC1 > DRC2, the specified DRC level can be achieved by further dynamic range compression in the decoder, that is, dynamic range compression by the amount corresponding to f×DRC1 - DRC2. In this case, it is not necessary to cancel the preprocessing DRC first. Finally, when the target DRC level is the full DRC quantified by DRC1, as expressed by f = 1, whether to perform partial cancellation or further compression of the encoder-side dynamic range limitation depends on whether DRC1 < DRC2 or DRC1 > DRC2.

[0016] In a second embodiment, a method for reconstructing an n-channel audio signal X based on a bitstream is provided. According to this method, an encoded core signal Y ~(Translator's note: "~" is above "Y") The following action is triggered upon receiving a bitstream containing one or more multi-channel coding parameters α and pre-processing DRC parameters DRC2: • Decodes the above-mentioned encoded core signal into an m-channel core signal, 1 < m <nであるステップと、 - A step of performing parametric synthesis to reconstruct the n-channel signal based on the core signal and the multi-channel coding parameters.

[0017] In the second embodiment, decoding includes canceling the encoder-side dynamic range limitation based on the parameter DRC2.

[0018] The first and second embodiments are functionally similar and generally share the same advantages.

[0019] In a further development of the first embodiment, the decoding system also receives one or more compensated post-processing DRC parameters DRC3 as part of the bitstream while the system is in parametric coding mode. This quantifies the DRC applied by the decoder. The application of DRC depends on manual user input, automatically detected playback device characteristics, etc. Therefore, the DRC applied by the decoder may be fully, partially, or not applied at all. Generally speaking, the pre-processing DRC parameter DRC2 is useful for boosting the dynamic range relative to the input DRC level, while the compensated post-processing DRC parameter DRC3 is useful for adjusting the dynamic range from the input DRC level, including range compression. The DRC3 parameter can be expressed in logarithmic form as a positive or negative dB value. Thus, DRC3 = y > 0, the relative amplitude change on the decoder side is

number

[0020] In yet another development of the above, the decoding system includes a DRC processor capable of operating to cancel encoder-side dynamic range compression based on parameter DRC2. Optionally, the DRC processor is capable of operating to cancel a portion of the dynamic range compression applied on the encoder side, represented by parameter f above.

[0021] In yet another development, the decoding system further includes a DRC preprocessor that controls the DRC processor and core signal decoder to achieve a target DRC level. Therefore, the DRC preprocessor determines whether the target DRC level (e.g., f × DRC1) is greater than or less than the input DRC level. The input DRC level is the dynamic range of the originally encoded audio signal, reduced by the encoder-side DRC quantified by the preprocessing DRC parameter DRC2. Based on this determination, if the decoded audio signal needs to be boosted, the DRC preprocessor instructs the DRC processor to (i) partially or completely cancel the encoder-side dynamic range limitation. If the decoded audio signal needs to be compressed (e.g., f × DRC1 > DRC2), the DRC preprocessor instructs the DRC processor to (ii) partially or completely perform the applicable decoder-side DRC quantified by parameter DRC3.

number

[0022] In one embodiment, the decoding system further encodes n-channel signals X in discrete decoding mode. ~The audio signal is reconstructed based on a bitstream containing (Translator's note: "~" is above "X"). Thus, this embodiment is a dual-mode or multiple-mode decoding system. From an adaptive coding perspective, discrete coding modes represent high bitrate modes, while parametric coding modes generally correspond to low bitrate modes.

[0023] In one embodiment, the decoding system is of a dual-mode type, i.e., it operates in either parametric coding mode or discrete coding mode. The decoding system can apply a decoder-side DRC in each of these modes. In discrete coding mode, the decoding system uses a post-processing DRC parameter DRC1 as a guide for the DRC. However, in parametric coding mode, the n-channel audio signal is generated based on a core signal that is potentially determined with respect to the encoder-side dynamic range limit, at least for some time blocks. To account for the dynamic range changes already made (i.e., dynamic range limits in some time blocks), the decoding system uses a compensated post-processing DRC parameter DRC3 as a guide for the DRC. Both parameters DRC1 and DRC3 are derived from the bitstream, but under normal system operation, only one of the parameter types, and not both, can be derived in a time block. Including both parameters DRC1 and DRC3 would transmit redundant information if parameter DRC2 were present. The decoding system of this embodiment uses parameter DRC2 to scale parameter DRC1 to parameter DRC3, or to scale parameter DRC3 to DRC1. For example, the decoding system includes a DRC down compensator that receives parameters DRC2 and DRC3 and outputs a recovered post-processing DRC parameter applied by the decoder system. The recovered post-processing DRC parameter is compared (on the same scale) with the post-processing DRC parameter DRC1. In other words, the decoder-side DRC represented by the recovered DRC parameter is quantitatively equivalent to the coupling of the encoder-side dynamic range limit of the core signal and the decoder-side DRC represented by the compensated post-processing DRC parameter. In the simple implementation described above, the relationship between each DRC parameter is as follows: the recovered DRC parameter is obtained as DRC2 + DRC3, which is equal to DRC1.

[0024] In a second aspect of the present invention, one embodiment provides an encoding system that encodes an n-channel audio signal X partitioned into time blocks as a bit stream P. The encoding system · receives the n-channel signal and, based thereon, outputs an m-channel core signal Y and one or more multichannel coding parameters α in a parametric coding mode of the encoding system, 1 < a parametric analysis stage where m < n, and · receives the core signal and outputs an encoded core signal Y ~ (Note: "~" is above "Y").

[0025] In the coding system, the parametric analysis stage performs time-segment-based adaptive dynamic range limiting and outputs a preprocessing DRC parameter DRC2 that quantifies the applied dynamic range limit. A time segment can be a single time block or a series of consecutive time blocks (e.g., a time frame containing six time blocks). The coding system is configured to transmit the preprocessing DRC parameter DRC2 with the bitstream (preferably as part of it, though not necessarily required). By transmitting the preprocessing DRC parameter DRC2, the coding system enables the decoding system to receive the bitstream and cancel the dynamic range limit imposed on the core signal by the parametric analysis stage. When dynamic range limiting is performed on a time-block basis, the parameter DRC2 has the resolution of a time block. Alternatively, when dynamic range limiting is performed on a frame basis, the parameter DRC2 has the resolution of a frame. In other words, each time block is associated with or refers to a value of parameter DRC2, but this value can be updated on a frame-based or block-based basis. Furthermore, dynamic range limiting in the parametric analysis stage can be performed either directly on the core signal (for example, by applying dynamic range limiting to the core signal) or indirectly (for example, by applying dynamic range limiting to the signal from which the core signal is desired).

[0026] In a further development of the above-described embodiment, the coding system can operate in parametric coding mode and discrete coding mode. To enable DRC on the decoder side, the encoder is configured to require one or more preprocessing steps that quantify the decoder-side DRC to be applied. Parameter DRC1 operates in discrete coding mode. However, in parametric coding mode, parameter DRC1 is compensated to take into account the dynamic range limit already performed in the parametric analysis stage. The output of this compression process includes a compensated post-processing DRC parameter DRC3. The guiding principle of this compensation process is that the decoder-side DRC represented by the post-processing DRC parameter is quantitatively equivalent to the combination of the dynamic range limit applied by the parametric analysis stage (quantified by parameter DRC2) and the decoder-side DRC (quantified by the compensated post-processing DRC parameter DRC3). Preferably, the three parameter types are all expressed on compatible scales, for example, using corresponding linear or logarithmic units. In the simple implementation described above, the relationships between each DRC parameter (still on a logarithmic scale) are as follows: the compensated post-processed DRC parameters are obtained as DRC1-DRC2.

[0027] In yet another embodiment of the second aspect, the encoding method includes the following steps: Step 1: Receive an n-channel audio signal X partitioned into time blocks; The steps include: generating an m-channel core signal Y and one or more multi-channel coding parameters α, while simultaneously performing dynamic range limiting on a time-block basis and generating one or more preprocessed DRC parameters DRC2 that quantify the applied dynamic range limit; - A step of outputting a bitstream P that includes the core signal, multi-channel coding parameters, and pre-processing DRC parameters DRC2.

[0028] In yet another embodiment, a computer program product is provided that includes a computer-readable medium having computer-executable instructions for performing the decoding or encoding method according to the above embodiment. The computer program product can be executed on a general-purpose computer. The general-purpose computer does not necessarily include dedicated hardware components.

[0029] In yet another embodiment, the present invention provides a data structure for storing or transmitting an audio signal. This structure includes an m-channel core signal Y, one or more mixing parameters α, and one or more preprocessing DRC parameters DRC2 that quantify encoder-side dynamic range limitations. This structure can be decoded by a linear combination of n channels of downmix signal channels (and optionally, a linear combination of channels of uncorrelated signals) (the one or more mixing parameters control the gain of at least the linear combination), and by canceling the encoder-side dynamic range limitations. Specifically, the present invention provides a computer-readable medium for storing information configured by the above data structure. In the above data structure, the preprocessing DRC parameter DRC2 is encoded as a 3-bit field representing the exponent and a 4-bit field representing the mantissa. During decoding, the exponent and mantissa are combined to form a scalar value corresponding to the gain value. Alternatively, the preprocessing DRC parameter DRC2 is encoded as a 2-bit field representing the exponent and a 5-bit field representing the mantissa.

[0030] Further embodiments are provided in the dependent claims. It should be noted that the present invention relates to all combinations of features, even if they are described in different claims. II. Embodiment: Encoding side Figure 1a shows a dual-mode coding system 1 according to one embodiment in a generalized block diagram format. An n-channel audio signal X is supplied to the upper (which is active in at least one discrete coding mode of the coding system 1) and the lower (which is active in at least one parametric coding mode of the system 1), respectively.

[0031] At the top, a discrete-mode DRC analyzer 10 is typically placed in parallel with an encoder 11, both receiving an audio signal X as input. Based on this signal, the encoder 11 outputs an encoded n-channel signal X~ (Note: "~" is above "X"). In response, the DRC analyzer 10 outputs one or more post-processing DRC parameters DRC1 that quantify the applicable decoder-side DRC. The parallel outputs from both units 10 and 11 are collected by a discrete-mode multiplexer 12. The discrete-mode multiplexer 12 outputs a bitstream P.

[0032] The lower part of the symbolization system 1 has a parametric analysis stage 22 arranged in parallel with the parametric mode DRC analyzer 21, and the parametric analysis stage 22 receives an n-channel audio signal X. Based on the n-channel audio signal X, the parametric analysis stage 22 outputs one or more multi-channel coding parameters (collectively denoted as α) and an m-channel (1 ≦ m < n) core signal Y. The m-channel core signal Y is then processed by the core signal encoder 23. The core signal encoder 23 outputs an encoded core signal Y~ (note: the "~" is above "Y") based on the m-channel core signal Y. As indicated by the notation "g↓", the parametric analysis stage 22 performs dynamic range limitation in a time block as necessary. Conditions that may control when to apply dynamic range limitation are the "non-clip condition" or the "in-range condition", where in time segments with large core signal amplitudes, the signal is processed so that it falls within a defined range. This condition is implemented based on a time frame consisting of one time block or multiple time blocks. Preferably, this condition is implemented by using broad-spectrum gain reduction rather than simply truncating peak values or using a similar approach. As is well known in the art, when limitation is required only for a plurality of time blocks, there is a technique for rendering a temporary dynamic range limitation operation less noticeable, such as gradually applying and / or removing the limitation. In particular, the system 1 may include a feedback loop (not shown) configured to smooth the DRC parameters. For example, the current parameter value to be output is obtained as the sum of a part 0 < a < 1 of the parameter value of the previous segment and a part (1 - a) of the parameter value that is the result of implementing the "non-clip condition" in the current segment. The processed DRC parameter DRC1 and the pre-processed DRC parameter DRC2 are, of course, smoothed independently, and different constants a are used.

[0033] Figure 5 shows a possible implementation of the parametric analysis stage 22, which includes a preprocessor 527 and a parametric analysis processor 528. The preprocessor 527 is responsible for performing dynamic range limiting on the n-channel signal X, thereby outputting a dynamic range limited n-channel signal Xc. The signal Xc is fed to the parametric analysis processor 528. The preprocessor 527 also outputs block or frame values ​​of the pre-processing DRC parameter DRC2. The parameter DRC2, along with the multi-channel coding parameter α and the m-channel core signal Y from the parametric analysis processor 528, is included in the output from the parametric analysis stage 22.

[0034] Referring again to Figure 1a, the discrete-mode DRC analyzer 10 functions similarly to the parametric-mode DRC analyzer 21 in that it outputs one or more post-processing DRC parameters DRC1 that define the decoder used. However, the parameter DRC1 supplied by the parametric-mode DRC analyzer 21 is not included in the bitstream of the parametric coding mode, but is compensated to take into account the dynamic range limiting performed by the parametric analysis stage 22. For this purpose, the DRC up-compensator 24 receives the post-processing DRC parameter DRC1 and the pre-processing DRC parameter DRC2. For each time block, the DRC up-compensator 24 determines the value of one or more compensated post-post-processing DRC parameters DRC3. This is so that the combined action of the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 is quantitatively equal to the DRC quantified by the post-processing DRC parameter DRC1. In other words, the DRC up-compensator 24 is configured to reduce the post-processing DRC parameters output by the DRC analyzer 21 by the amount (if any) already done by the parametric analysis stage 22. What should be included in the bitstream is the compensated post-processing DRC parameter DRC3. Continuing to refer to the bottom of System 1, the parameter-mode multiplexer 25 collects the compensated post-processing DRC parameter DRC3, the pre-processing DRC parameter DRC2, the multi-channel coding parameter α, and the coded core signal Y~ (Translator's note: "~" is above "Y", and so on) and forms the bitstream P based on them. In possible implementations, the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 are coded in logarithmic form as dB values ​​that affect amplitude upscaling or downscaling on the decoder side. The sign of the compensated post-processing DRC parameter DRC3 can be arbitrary. However, the pre-processing DRC parameter DRC2 is the result of implementing conditions such as "non-clipping conditions" and is always expressed in non-negative dB.

[0035] In common to both the upper and lower parts of the encoding system 1, a selector 26 (symbolizing any signal selection means implemented in hardware or software) determines, depending on the actual coding mode, whether the bitstream from the upper or lower part of the encoding system 1 constitutes the final output from the encoding system 1. Similarly, the input side of the system 1 is provided with a switch (not shown in Figure 1a) that directs the audio signal X to either the upper or lower part of the system 1. The input switch can be activated in correspondence with the output switch 26.

[0036] Referring to Figure 1a and the diagrams described below, bitstream P can be encoded in a format compliant with Dolby Digital Plus (DD+ or E-AC-3, Enhanced AC-3). The bitstream includes at least two metadata fields, dynrng and compr. According to one DD+ specification, dynrng has a resolution of one time block, while compr has a resolution of one frame. One frame has four to six time blocks. Regarding the importance of these metadata fields, the post-processing DRC parameter DRC1 defined above corresponds to either dynrng or compr, functioning in a way that ensures the monophonic downmix does not exceed a certain peak level, depending, for example, whether "heavy compression" is activated. In a normal environment, both the dynrng and compr fields are transmitted, and it is the decoder's problem to decide which to use. Thus, the post-processing DRC parameter DRC1 has either block resolution or frame resolution, but can be transmitted in the legacy part of the format and understood by legacy decoders. However, the preprocessing DRC parameter DRC2 has no corresponding value in the DD+ format and is preferably encoded in a new metadata field. It is recalled that the preprocessing DRC parameter DRC2 concerns the dynrng and / or compr portion that ensures the signal is not clipped when the signal is downmixed from 5.1 format (n=6) to stereo format (m=2). The guaranteed postprocessing DRC parameter DRC3 is the result after compensating the dynrng or compr value by inferring the clipping prevention quantified by the preprocessing DRC parameter DRC2; therefore, it is transmitted in the dynrng or compr field in the DD+ bitstream.

[0037] The new metadata field for the preprocessing DRC parameter DRC2 has 7 bits (xxyyyyy), where the bit at position x represents an integer in the range [0,3] and the bit at position y represents an integer in the range [0,31]. The preprocessing DRC parameter DRC2 is the gain factor.

number

[0038] Another metadata parameter in the DD+ format is dialnorm, which is the (potentially time-averaged) loudness level of the content. In embodiments, the target output reference level LT is a setting in the decoder configuration, which may be controlled by the user. To achieve the target output reference level LT, the decoding system applies static attenuation quantified by the difference dialnorm-LT. To determine the total attenuation to apply, the decoding system increases this difference by an additional attenuation defined by the (uncompensated) post-processing DRC parameter DRC1, or the compensated post-processing DRC parameter DRC3, or the target DRC, which is expressed as a part of the post-processing DRC parameters f × DRC1. This results in the respective dialnorm- T +DRC1 or dialnorm-L T +DRC3 or dialnorm- T +f × DRC1 is obtained. If one of these three linear combinations is positive, a non-zero amount of total attenuation is applied to the decoding system; if it is negative, the signal is effectively boosted.

[0039] FIG. 7 shows an encoding system 701 that functions similarly to the encoding system 1 shown in FIG. 1a according to yet another embodiment. Using similar reference symbols and notations consistent with FIG. 1a for signals, it is believed that a detailed description of the operating principle of the encoding system 701 is not necessary. However, one important difference is that one DRC analyzer 721 fulfills the tasks of both the discrete mode DRC analyzer 10 and the parametric mode DRC analyzer 21 of FIG. 1a. For this purpose, the DRC analyzer 721 receives the n-channel audio signal X encoded by the encoding system 701. The DRC analyzer 721 supplies the post-processing DRC parameter DRC1 generated based on the n-channel audio signal X to both the discrete mode multiplexer 712 and the DRC up-compensator 724. The DRC up-compensator 724 is functionally equivalent to the DRC up-compensator 24 of the encoding system 1 in FIG. 1a.

[0040] FIG. 3 shows an encoding system 301. It is relatively simpler than that shown in FIG. 1a in that it does not generate post-processing DRC parameters as an output. Therefore, a decoder that receives the bitstream P generated by the encoding system 301 does not necessarily have to perform dynamic range compression. However, such a decoder can cancel the dynamic range limitation applied by the encoding system 301. Typically, this is to increase the dynamic range of a time block in which the n-channel audio signal X contains relatively high amplitude peaks.

[0041] In FIG. 3, the upper part of the encoding system 301 is active at least in the discrete coding mode of the encoding system 301, and the n-channel signal X encoded based on the n-channel signal X encoded by the system 301 ~It does not need to include anything other than the encoder 311 configured to provide (Translator's note: "~" is above "X"). The lower part corresponds to the discrete coding mode and includes fewer components than the analog portion of the coding system shown in Figure 1a, namely, only a parameter analysis stage 322 that outputs preprocessing DRC parameters DRC2, multichannel coding parameters α and m-channel core signals Y based on n-channel audio signals X. The core signal Y is encoded by the core signal encoder 323 (core signal Y ~ After being processed by (translator's note: "~" is above "Y"), the output set from the parametric analysis stage 322 is synthesized by the parametric mode multiplexer 325 into bitstream P. A selector 326 located downstream of both the upper and lower parts of the coding system 301 is responsible for outputting the bitstream generated by either the upper or lower part, depending on the current coding mode of the coding system 301.

[0042] Encoding system 1001 shown in FIG. 10 shows further simplification. This encoding system 1001 is configured to process an n-channel audio signal X in a format suitable for storage and transmission without further encoding operations. Therefore, in the discrete coding mode, as indicated by the position of selector 1026 shown in FIG. 10, the audio signal X is output from the encoding system 1001 without further processing. In the parametric coding mode, parametric analysis stage 1022 analyzes the n-channel audio signal X and outputs preprocessing DRC parameter DRC2, multi-channel coding parameter α, and m-channel core signal Y. As described above, parametric analysis stage 1022 is configured to act on the n-channel audio signal even when the n-channel audio signal is in a format suitable for transmission and storage. In the encoding system 1001 of FIG. 10, the core signal Y is in a format that can be transmitted or stored, and in the parametric coding mode, this signal is combined by parametric mode multiplexer 1025 together with multi-channel coding parameter α and parameter DRC2 to form a bitstream and is output from the encoding system 1001.

[0043] FIG. 1b shows a single-mode encoding system according to one embodiment. An n-channel audio signal X is provided to a DRC analyzer 21 and a parametric analysis stage 22 (both are arranged in parallel). The parametric analysis stage 22 outputs one or more multi-channel coding parameters (collectively denoted as α) and an m-channel (1≦m<n) core signal Y based on the n-channel audio signal X. The m-channel core signal Y is then processed by a core signal encoder 23. The core signal encoder 23, based on the m-channel core signal Y, encodes the core signal Y ~(Translator's note: "~" is above "Y") is output. Parametric analysis stage 22 applies dynamic range limits to the time blocks where this is required. DRC up compensator 24 receives the post-processing DRC parameter DRC1 and the pre-processing DRC parameter DRC2. For each time block (in this example, the resolution at which the value of post-processing DRC parameter DRC1 is generated is one time block), DRC up compensator 24 determines one or more compensated post-processing DRC parameter DRC3 values. This is such that the combined operation of the compensated post-processing DRC parameter DRC3 and the pre-processing DRC parameter DRC2 is quantitatively equal to the DRC quantified by post-processing DRC parameter DRC1.

[0044] Figure 8 provides a more detailed view of the possible functions of the DRC up-compensators 24, 724 in Figures 1 and 7. Each DRC up-compensator 24, 724 is configured to generate a compensated post-processing DRC parameter DRC3 based on a pre-processing DRC parameter DRC2 and a post-processing DRC parameter DRC1. Each bar represents a time frame of the signal. Each frame is associated with the values ​​of the pre-processing DRC parameter DRC2 and the post-processing DRC parameter DRC1. In Figures 8 and 9, these are in dBFS units with a negative sign. As the legend indicates, the solid line represents the post-processing DRC parameter DRC1, while the other two DRC parameter values ​​correspond to different shading patterns. Each value of the compensated post-processing DRC parameter DRC3 is generated on the condition that the combined operation of the pre-processing DRC parameter DRC2 and the compensated post-processing DRC parameter DRC3 is quantitatively equal to the decoder-side DRC represented by the post-processing DRC parameter DRC1. Figures 8 and 9 are simplified and do not faithfully represent the effect of DRC by the specific approach (i.e., the Carroll and Riedmiller paper mentioned above) as a scalar, i.e., a linear quantity. Figures 8 and 9 would show a complete picture of the simplified embodiment described above, where the DRC parameters are encoded as scalars.

[0045] Figure 8 shows a state where the post-processing DRC parameter DRC1 is constant within each time frame, as described above, similar to the compr parameter in the DD+ format. However, this is not always the case. For example, legacy-type DRC analyzers are configured to analyze segments with a fixed number of p1 time blocks, where p1 is equal to 4, 6, 8, 16, 24, 32, 64, or other integers, which is generally significantly less than the number of time blocks in the entire program (e.g., songs, tracks, episodes of a radio program). This number p1 may or may not match the number of frames p2 between each update of the pre-processing DRC parameter. Figure 8 shows a specific case where p1=6 and p2=6. Preferably, the post-processing DRC parameter DRC1 is small enough to be re-evaluated at least once per second of the audio signal X, more preferably tens or hundreds of times per second of the audio signal X.

[0046] Figure 9 shows the case where p1=1, similar to the dynrng parameter in DD+ format. However, the dynamic range limit of the parametric analysis stage 22,722 is performed based on p2=6 time blocks at a time, so that a new value of the preprocessing DRC parameter DRC2 is generated every 6 time blocks. Each thin bar represents one time block. The up-compensator 24,724 is configured to determine each value of the compensated post-processing DRC parameter DRC3 such that the decoder-side DRC represented by the post-processing DRC parameter DRC1 is quantitatively equal to the combination of the dynamic range limit applied by each parameter analysis stage 22,722 across each time block and the decoder-side DRC quantified by the compensated post-processing DRC parameter DRC3. III. Embodiment: Decoder side Figure 2a shows a single-mode decoding system 51 that reconstructs an n-channel audio signal based on a bitstream P. The bitstream P is the encoded core signal Y. ~(Translator's note: "~" is above "Y"), the multi-channel coding parameter α, the pre-processing DRC parameter DRC2, and the compensated post-processing DRC parameter DRC3 are extracted from the bitstream by a demultiplexer 70 located at the input of the decoding system 51. The core signal decoder 71 encodes the core signal Y ~ (Translator's note: "~" is above "Y") is received and, based on this, outputs an m-channel core signal Y (1 ≤ m ≤ n). With respect to decoding, the core signal decoder 71 further performs DRC quantified by the decoded post-processing DRC parameter DRC3. The core signal decoder 71 operates to produce a full DRC represented by the compensated post-processing DRC parameter DRC3 or a portion thereof. This decision may be manually controlled by the user or based on detection of the characteristics of the playback device. Downstream of the core signal decoder 71 is a DRC processor 74 which recovers the dynamic range of the core signal by canceling the dynamic range limitation imposed on the encoder side, which is quantified by the pre-processing DRC parameter DRC2, as indicated by the g↑ indicator. The DRC processor 74 outputs an intermediate signal YC, which is the same as the core signal Y except for its dynamic range, and is input to the parametric synthesis stage 72. The parametric synthesis stage 72 forms an n-channel linear combination of m channels in the intermediate signal YC (the applied gain is controlled by the multi-channel coding parameter α) and outputs a reconstructed n-channel audio signal X. The linear combination of the parametric synthesis stage 72 further includes a decorrelated signal obtained from the intermediate signal YC or the core signal Y. The decorrelated signal is further subjected to nonlinear processing such as artifact attenuation. The decorrelated signal may be generated by a core signal modification unit or a decorrelator (not shown). In the simple embodiment outlined above, the cancellation of the dynamic range limitation imposed on the encoder side in the DRC processor 74 means scaling the signal over a wide spectral range by a coefficient corresponding to the reciprocal of the parameter DRC2 that quantifies the preprocessing dynamic range limitation.

[0047] Figure 2b shows the decoding system 51, which is a slightly more advanced version of the decoding system in Figure 2a. This decoding system 51 is equipped with a DRC preprocessor 77, which adjusts the DRC-related operation of the core signal decoder 71 and the DRC processor 74, respectively. On the one hand, the core signal decoder 71 can compress the dynamic range of the signal to a limit determined by the compensated post-processing DRC parameter DRC3, or it can operate to compress the dynamic range. On the other hand, the DRC processor 74 can operate to fully increase the dynamic range to the pre-encoded level, or partially. With this setting, it is generally possible to achieve a certain target DRC level by activating the DRC processing of only one of the core signal decoder 71 or the DRC processor 74. If the compensated post-processing DRC parameter DRC3 indicates dynamic range compression, operating both units simultaneously suggests that there is some degree of mutual counter-action (mutual cancellation), which may have a negative impact on the quality of the output.

[0048] The DRC preprocessor 77 receives both a preprocessing DRC parameter DRC2 and a compensated postprocessing DRC parameter DRC3. The DRC preprocessor 77 also has access to a predetermined or variable (e.g., user-specified) DRC target level, represented by a parameter f (e.g., f × DRC1) and the input DRC level of the signal corresponding to the original dynamic range obtained from DRC2. Based on a comparison of the two DRC levels, the DRC preprocessor 77 determines whether the DRC target level can be achieved by dynamic range compression in the core signal decoder 71 or by dynamic range boosting in the DRC processor 74. For this purpose, the DRC preprocessor 77 receives the decoded control signal k 71 , k 74 These outputs are supplied to the core signal decoder 71 and the DRC processor 74, respectively.

[0049] Control signal k supplied from the DRC preprocessor 77 to the core signal decoder 71 and the DRC processor 74. 71、 k 74 The behavior of the first control signal k is explained here. 71 This controls how much of the decoder-side DRC, quantified by the compensated post-processing DRC parameter DRC3, is applied by the core signal decoder 71. In the simple embodiment described above, the resulting relative gain change is a coefficient

number

number

[0050] f × DRC1 = k 74 ×DRC2+k 71 ×DRC3 Here, f∈[0,1] is predetermined, and DRC2 > 0 and DRC1 = DRC2 + DRC3 (logarithmic scale). From the above, it can be seen that DRC1 and DRC3 can be positive or negative. As stated above, if the operation of the core signal decoder 71 is range compacting (DRC3 = y > 0), it is desirable to avoid operating both the core signal decoder 71 and the DRC processor 74. This is because k 71 = 0 or k 74 We will solve the above equation when the value is 0.

[0051] A further possible representation is a loudness-dependent gain coefficient, which may be on a logarithmic scale. For example, a set of gain coefficients is transmitted along with the dialogue level. The first gain coefficient is applied to time segments louder than the dialogue level, while the second gain coefficient is applied to quiet time segments. This allows for dynamic range compression and expansion, as the first and second gain coefficients can be assigned mutually dependent values.

[0052] Figure 2c shows a dual-mode decoding system 51, which is configured to receive a bitstream P containing audio signals that are either parametrically encoded or directly encoded. In the parametric mode of the decoding system 51, the downstream upper part of the parametric mode demultiplexer 70 is active and provides channel audio signals X, similar to the function of the system shown in Figure 2a. In discrete mode, the bitstream P is encoded into n channel signals X. ~ (Translator's note: "~" is above "X") and one or more post-processing DRC parameters DRC1 are extracted and supplied to the discrete-mode demultiplexer 60. The input and output selectors 52, 82 of the decoding system 51 (symbolizing signal selection means implemented in any hardware or software) operate according to the current mode. The selectors may operate together so that they are always in the upper or lower position. In discrete mode, the encoded n-channel signal X ~ (Translator's note: "~" is above "X") is processed by decoder 61. Decoder 61 can perform DRC using the post-processing DRC parameter DRC1. Dialogue-level consistency between discrete coding mode and parametric coding mode is ensured by the decoding system 51 being configured to use a compensated post-processing DRC parameter DRC3 instead of the (uncompensated) post-processing DRC parameter DRC1 in parametric mode. The relationship between parameters DRC1 and DRC3 is described above.

[0053] Figure 4 is a generalized block diagram showing a simplified decoding system 451. This simplified decoding system 451 does not have the capability to perform post-processing DRC. However, the decoding system 451 in Figure 4 is operable to cancel the dynamic range limitation applied on the encoder side, which is quantified by the pre-processing DRC parameter DRC2. More precisely, the parametric synthesis stage 472 is configured to completely or partially cancel this dynamic range limitation, as indicated by the symbol g↑.

[0054] Figures 11 and 12 show two possible implementations of the parametric synthesis stage 472 shown in Figure 4. Similar implementations are also useful for the type of coding system shown in Figure 13, which will be discussed further later. In the first possible implementation, as shown in Figure 11, the preconditioner 1174 performs dynamic range limiting cancellation on the m-channel core signal Y, thereby obtaining the m-channel intermediate signal YC. The intermediate signal YC is processed in the parametric synthesis processor 1175, which forms a linear combination of the channels in the intermediate signal YC (and possibly additional uncorrelated signals). Here, the gain applied to the linear combination is controllable by the multi-channel coding parameter α, which is also supplied to the parametric synthesis processor 1175.

[0055] The second implementation shown in Figure 12 represents an alternative. In the second implementation, parametric synthesis is performed as a processing step before dynamic range limit cancellation. This is evident from the fact that the parametric synthesis processor 1275 is located upstream of the post-conditioner 1276. The post-conditioner 1276 is responsible for canceling the encoder-side dynamic range limit, which is quantified by the pre-processing DRC parameter DRC2. Therefore, the signal supplied from the parametric synthesis processor 1275 to the post-conditioner 1276 relates to the n-channel signal XC with limited dynamic range.

[0056] Figure 13 shows a decoding system 1351 according to yet another embodiment. The decoder-side DRC is influenced by a DRC processor 1383 located downstream of both the discrete-mode and parametric-mode portions of system 1351. As with the decoding systems described with reference to Figures 2a, 2b, 2c, and 4, this decoding system 1351 can also cancel the dynamic range limit applied to the encoder side, which is quantified by the post-processing DRC parameter DRC2. The DRC processor 1383 operates in both the discrete coding mode, where the (uncompensated) post-processing DRC parameter DRC1 is included in the received bitstream P, and the parametric coding mode, where the compensated post-processing DRC parameter DRC3 is received. It should be noted that decoding system 1351 differs from system 51 shown in Figure 2b in that the post-processing DRC acts on the n-channel output signal, i.e., downstream of the parametric synthesis stage 1372. In system 51 in Figure 2b, the corresponding operation occurs in the core signal decoder 71.

[0057] The DRC processor 1383 receives the target DRC level from the user, memory, hardware diagnostics performed on the playback device, or other external or internal data sources. For example, the target DRC level f represents the portion of the full post-processing DRC that the user wants the decoding system 1351 to act upon. As can be seen from the diagram, the structure of the decoding system 1351 has the advantage that only the DRC processor 1383 is required to consider the value of parameter f. This makes the implementation of partial DRC convenient. For this purpose, it is configured to convert the compensated post-processing DRC parameter DRC3 to the scale of the (uncompensated) post-processing DRC parameter DRC1. In fact, the n-channel audio signal X output from the parametric synthesis stage 1372 performs encoder-side dynamic range limitation cancellation. Therefore, applying DRC based on the compensated post-processing DRC parameter DRC3 involves under-range compression. To prevent this scenario, the DRC down compensator 1373 recovers the compensated post-processing DRC parameter DRC3 based on the pre-processing DRC parameter DRC2, thereby obtaining the recovered post-processing DRC parameter in parametric coding mode, which is supplied to the DRC processor 1383. As mentioned above, the decoder-side DRC represented by the recovered DRC parameter is quantitatively equal to the combination of the encoder-side dynamic range limit already imposed on the core signal and the decoder-side DRC represented by the compensated post-processing DRC parameter DRC3, as suggested by Figures 8 and 9.

[0058] In another embodiment, the decoding system 1351 is implemented without the discrete-mode demultiplexer 1360 and decoder 1361. The DRC parameter selectors 1381 and 1382 in Figure 13 are replaced by connections between the DRC processor 1383 and the DRC down compensator 1373 (which receives the recovered post-processed DRC parameters) and the parametric synthesis stage 1372 (which supplies the n-channel audio signal X), respectively. This alternative embodiment is simplified in that it operates in single-parametric decoding mode. Furthermore, the implementation is simpler because a legacy-type DRC processor 1383, which is not necessarily configured to handle compensated post-processed DRC parameters, can be used.

[0059] Figure 6 shows the legacy decoding system 651 that decodes the received bitstream P into an m-channel audio signal. In parametric coding mode, the upper part downstream of the parameter-mode demultiplexer 670 is active, and the encoded m-channel core signal Y ~ (Translator's note: "~" is above "Y") and outputs the compensated post-processing DRC parameter DRC3. Encoded m-channel core signal Y ~ (Translator's note: "~" is above "Y") is decoded into an m-channel core signal Y by the first decoder 671. In discrete coding mode, the output audio signal is generated by the lower part located downstream of the discrete-mode demultiplexer 660. The discrete-mode demultiplexer 660 decodes the bitstream P into an encoded n-channel signal X ~ (Translator's note: "~" is above "X") and extract the (uncompensated) post-processing DRC parameter DRC1. Encoded n-channel signal X ~The signal (Translator's note: "~" is above "X") is decoded by the second decoder 661 and downmixed to an m-channel signal Y in the downmix stage 662. Both this signal Y and the signal Y described for the parametric mode are supplied to the DRC processor 683, which is common to both modes. In the parametric mode, the quantitative characteristics of the DRC processor 683 are controlled by the compensated post-processing DRC parameter DRC3. In the discrete mode, these characteristics are controlled by the (uncompensated) post-processing DRC parameter DRC1. In this way, consistency in the dialogue level of the m-channel audio signal output from the decoding system 651 can be maintained. It should be noted that this decoding system 651 is of the legacy type because it treats compensated and uncompensated post-processing DRC parameters the same unless they are the same.

[0060] IV. Reference symbols in drawings

[0061] [Table 1-1] [Table 1-2] V. Equivalents, extensions, modifications, etc. Further embodiments of the present invention will become apparent to those skilled in the art from the above description. While this specification and drawings disclose embodiments and examples, the present invention is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the invention as defined in the appended claims. Reference numerals appearing in the claims should not be considered limiting.

[0062] The systems and methods disclosed herein can be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks between functional units as referred to above does not necessarily correspond to the division to physical units.

[0063] Conversely, one physical component has multiple functions, and one task is executed by multiple physical components in cooperation. Some or all of the components can be implemented as software executed by a digital signal processor or a microprocessor, or can be implemented as hardware or an application-specific integrated circuit. Such software can be distributed on a computer-readable medium. The computer-readable medium includes a computer storage medium (i.e., a non-transitory medium) and a communication medium (i.e., a transitory medium). As is well known to those skilled in the art, the term computer storage medium includes any method or technology for storing information such as computer-readable instructions, data structures, program modules, and other data, including volatile and non-volatile, removable and non-removable media. The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory, and other memory technologies, CD-ROM, digital versatile disk (DVD), and other optical disk storage media, magnetic cassettes, magnetic tapes, magnetic disk storage, and other magnetic storage devices, or any other medium that can be used to store the desired information. Furthermore, as is well known to those skilled in the art, the communication medium generally embodies data in a modulated data signal such as computer-readable instructions, data structures, program modules, and other carrier waves or other transmission mechanisms, and includes any information distribution medium. Appendices are made to the embodiments. (Appendix 1) A decoding system configured to reconstruct an n-channel audio signal based on a bitstream, receiving the bitstream and, based thereon, outputting an encoded core signal and multi-channel coding parameters in a parametric coding mode of the system, a parametric mode demultiplexer; receiving the encoded core signal and, based thereon, outputting an m-channel core signal, where 1 ≦ m < n, a core signal decoder; The system includes a parametric synthesis stage that receives the core signal and the multi-channel coding parameters and outputs the n-channel signal based on them, The parameter-mode demultiplexer is further configured to output preprocessing dynamic range control (DRC) parameters that quantify the encoder-side dynamic range limit of the core signal based on the bitstream. The decoding system is operable to cancel the encoder-side dynamic range limitation based on the preprocessing DRC parameters. Decryption system. (Note 2) The parametric mode demultiplexer is further configured to output a compensated post-processing DRC parameter that quantifies the decoder-side DRC applied in the parametric coding mode of the system, based on the bitstream. The decoding system is 1) Within or downstream of the parametric synthesis stage, 2) Within the core signal decoder, On the other hand, it is operable to apply the decoder-side DRC, The decryption system described in Appendix 1. (Note 3) Furthermore, the system has a DRC processor that can operate to cancel or part of the encoder-side dynamic range limitation and output a compensated core signal. The core signal decoder is operable to apply the decoder-side DRC or a portion thereof. The decryption system described in Appendix 2. (Note 4) The DRC preprocessor is further coupled to the core signal decoder and the DRC processor in a communicative manner, and the DRC preprocessor receives the target DRC level, the preprocessing DRC parameters, and the compensated postprocessing DRC parameters. -When the target DRC level corresponds to a dynamic range boost with respect to the decoder input DRC level of the core signal, the DRC processor is instructed to cancel the encoder-side dynamic range limit or a portion thereof based on the target DRC level. -When the target DRC level corresponds to dynamic range compression with respect to the decoder input DRC level of the core signal, the core signal decoder is instructed to apply the decoder-side DRC or a portion thereof based on the target DRC level. The DRC preprocessor determines a portion of the above according to the target DRC level. The decryption system described in Appendix 3. (Note 5) The parametric mode demultiplexer is further configured to output compensated post-processing DRC parameters in the parametric coding mode of the system based on the bitstream. The aforementioned system further, A DRC down compensator receives the compensated post-processing DRC parameters and the pre-processing DRC parameters, and based on them outputs a recovered post-processing DRC parameter that quantifies the decoder-side DRC to be applied. The system includes, in the parametric coding mode, a DRC processor configured to apply DRC to the n-channel audio signal according to the recovered post-processing DRC parameters, The decoder-side DRC represented by the recovered DRC parameters is quantitatively equivalent to the combination of the encoder-side dynamic range limit of the core signal and the decoder-side DRC represented by the compensated post-processing DRC parameters. A decryption system as described in any one of the appendices 1 through 4. (Note 6) A discrete-mode demultiplexer that receives the bitstream and outputs, based thereon, an encoded n-channel signal and a post-processed DRC parameter that quantifies the applied decoder-side DRC in the discrete coding mode of the system, The system includes a decoder that receives an encoded n-channel signal contained in the bitstream and outputs the n-channel audio signal in the discrete coding mode of the system based on that signal. The DRC processor is further configured to apply DRC to the n-channel audio signal in the discrete coding mode of the system, according to the post-processing DRC parameters. The decryption system described in Appendix 5. (Note 7) The parametric synthesis stage is A preconditioner that receives the core signal and the preprocessing DRC parameters and outputs a dynamic range compensated core signal obtained by canceling the encoder-side dynamic range limit, A parametric synthesis processor that receives the dynamic range compensated core signal and the multi-channel coding parameters and outputs the n-channel signal based thereon, The decryption system described in Appendix 5 or 6. (Note 8) The parametric synthesis stage is A parametric synthesis processor that receives the core signal and the multi-channel coding parameters and outputs an intermediate signal based thereon, The system includes a post-conditioner that receives the intermediate signal and the pre-processed DRC parameters, and outputs an n-channel signal obtained by canceling the encoder-side dynamic range limit. The decryption system described in Appendix 5 or 6. (Note 9) The parametric mode demultiplexer is further configured to read each value of the preprocessing DRC parameter as a 2-bit field representing the exponent and a 5-bit field representing the mantissa. A decryption system as described in any one of the appendices 1 through 4. (Note 10) A method for reconstructing an n-channel audio signal based on a bitstream, Depending on the bitstream, which includes an encoded code signal, multi-channel coding parameters, and pre-processed dynamic range control (DRC) parameters that quantify the encoder-side dynamic range limit of the core signal, a-1) Decode the above-mentioned encoded core signal into an m-channel core signal, < m <nであるステップと、 a-2) The step of performing parametric synthesis and reconstructing the n-channel signal based on the core signal and the multi-channel coding parameters, The method further comprises the step of canceling the encoder-side dynamic range limit based on the preprocessing DRC parameters. method. (Note 11) Depending on whether the bitstream includes an encoded core signal, multi-channel coding parameters, and pre-processing DRC parameters, and further includes compensated post-processing DRC parameters that quantify the applicable decoder-side DRC, Steps a-1 and a-2, a-3) A step of canceling the encoder-side dynamic range limit or a part thereof based on the preprocessing DRC parameters, and a-4) The step of performing at least one of the steps of applying the decoder-side DRC or a portion thereof in accordance with the compensated post-processing DRC parameters, The method described in Appendix 10. (Note 12) By performing steps a-1 and a-2, the steps corresponding to the above case are performed, The process involves receiving a target DRC level, comparing it to the decoder input DRC level, and determining whether the target DRC level corresponds to dynamic range boosting or dynamic range compression. Based on the above comparison, a-3) A step of canceling the encoder-side dynamic range limit or a part thereof based on the preprocessing DRC parameters, and a-4) Performing one of the selected steps of applying the decoder-side DRC or a part thereof according to the compensated post-processing DRC parameter; The method according to appended claim 11, comprising: (Appended claim 13) The bitstream further includes a post-processing DRC parameter obtained by quantifying the decoder-side DRC to be applied; The method further includes applying DRC to the n-channel signal according to the post-processing DRC parameter. When the bitstream includes a pre-processing DRC parameter and the post-processing DRC parameter in the bitstream is a compensated post-processing DRC parameter, using a recovered post-processing DRC parameter instead of the compensated post-processing DRC parameter; The recovered post-processing DRC parameter is obtained based on the compensated post-processing DRC parameter and the pre-processing DRC parameter, and the decoder-side DRC represented by the recovered DRC parameter is numerically equivalent to the combination of the encoder-side dynamic range limitation of the core signal and the decoder-side DRC represented by the post-processing DRC parameter; The method according to any one of appended claims 10 to 12. (Appended claim 14) In response to the bitstream including an encoded n-channel signal, b) further comprising reconstructing the n-channel signal by decoding the encoded n-channel signal; The method according to appended claim 13. (Appended claim 15) A decoding system configured to encode an n-channel audio signal partitioned into time blocks as a bitstream, receiving the n-channel signal and outputting, based thereon, an m-channel core signal and multi-channel coding parameters in a parametric coding mode of the encoding system, 1 < where m < n, a parametric analysis stage; receiving the core signal and having a core signal encoder that outputs an encoded core signal based thereon; The parametric analysis stage further performs time-segment-based adaptive dynamic range limiting and outputs preprocessed dynamic range control (DRC) parameters that quantify the applied dynamic range limiting. The system further includes a parametric mode multiplexer capable of operating to form a bitstream output from the system, which includes at least the coding core signal, the multi-channel coding parameters, and the pre-processing DRC parameters, in the parametric coding mode of the system. Decryption system. (Note 16) At least one DRC analyzer that receives the n-channel audio signal and outputs post-processed DRC parameters that quantify the applicable decoder-side DRC based thereon, A DRC up compensator receives the post-processing DRC parameter and the pre-processing DRC parameter, and outputs a compensated post-processing DRC parameter that quantifies the applicable decoder-side DRC based on them, wherein the compensated post-processing DRC parameter is included in the bitstream in the parametric coding mode. The decoder-side DRC represented by the post-processing DRC parameters is quantitatively equivalent to the combination of the dynamic range limit applied by the parametric analysis stage and the decoder-side DRC quantified by the compensated post-processing DRC parameters. The encoding system described in Appendix 15. (Note 17) The above at least one DRC analyzer is a first number p1 > It is configured to calculate the value of the post-processing DRC parameter based on a single segment including a time lock, The parametric analysis stage is performed by the second number p2 > It is configured to calculate the value of the preprocessing DRC parameter based on a single segment containing one time block, The first number is less than or equal to the second number, i.e., p1 < p2 is The encoding system described in Appendix 16. (Note 18) An encoder that receives the n-channel signal and outputs an encoded n-channel signal that forms part of the bitstream output from the system in the discrete coding mode of the system, The system has a discrete-mode multiplexer capable of operating to form a bitstream output from the system in the discrete coding mode of the system, wherein the bitstream includes at least the coded n-channel signal and the post-processing DRC parameters. The encoding system described in Appendix 16 or 17. (Note 19) The system has two DRC analyzers, which are functionally equivalent, namely a discrete-mode DRC analyzer and a parametric-mode DRC analyzer. The encoding system described in any one of the appendices 15 to 18. (Note 20) The system further comprises a discrete-mode multiplexer capable of receiving the post-processing DRC parameters and the encoded n-channel signal and operating to form a bitstream to be output from the system in discrete coding mode. The encoding system described in any one of the appendices 15 to 19. (Note 21) The parametric analysis stage is A preprocessor that receives the aforementioned n-channel signal and outputs a dynamic range-limited n-channel signal and DRC parameters, A parametric analysis processor that receives the aforementioned dynamic range-limited n-channel signal and outputs the aforementioned m-channel signal and multi-channel coding parameters based on it, The encoding system described in any one of the appendices 15 to 20. (Note 22) The parametric mode demultiplexer is configured to include each value of the preprocessing DRC parameter as a 2-bit field representing the exponent and a 5-bit field representing the mantissa. The encoding system described in any one of the items 15 to 21 of the appendix. (Note 23) A method for encoding an n-channel audio signal partitioned into time blocks, The above method generates an m-channel core signal and multi-channel coding parameters. < m <nであるステップを有し、The generation step includes the step of performing a time-block-based dynamic range limit and the step of generating preprocessed dynamic range control (DRC) parameters that quantify the applied dynamic range limit. The method further comprises the step of transmitting the preprocessing DRC parameters simultaneously with the core signal and the multi-channel coding parameters. method. (Appendix 24) A computer program product including a computer-readable medium having computer-executable instructions that perform the methods described in any one of the appendices 10 to 14 and 23. (Note 25) A system, method, or computer program product described in any one of Notes 1 through 24, wherein n=6 and m=2.

Claims

1. A method for adjusting the dynamic range of an audio signal, performed by an audio signal processing device, Receiving a bitstream containing an encoded audio signal and encoder-generated dynamic range control (DRC) metadata, wherein the encoder-generated DRC metadata includes a plurality of sets of DRC gains, the plurality of sets of DRC gains including a first set of DRC gains representing a first portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal, and a second set of DRC gains representing a second portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal. Decoding the encoded audio signal to obtain the audio signal, Downmixing the aforementioned audio signal, This includes adjusting the dynamic range of the audio signal by applying a second set of the DRC gains to the audio signal after the downmix in order to apply the total DRC gains to be applied to the audio signal. method.

2. An audio signal processing device for adjusting the dynamic range of an audio signal, wherein the audio signal processing device has one or more processors, and the one or more processors are The encoder receives a bitstream containing an encoded audio signal and encoder-generated dynamic range control (DRC) metadata, wherein the encoder-generated DRC metadata includes a plurality of sets of DRC gains, the plurality of sets of DRC gains including a first set of DRC gains representing a first portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal, and a second set of DRC gains representing a second portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal. The encoded audio signal is decoded to obtain the audio signal, The aforementioned audio signal is downmixed, To apply the total DRC gain to the audio signal, the dynamic range of the audio signal is adjusted by applying a second set of the DRC gain to the audio signal after the downmix. device.

3. A non-temporary computer-readable storage medium containing software instructions, wherein, when executed by an audio signal processing device, the software instructions cause the audio signal processing device to perform a method for adjusting the dynamic range of an audio signal, and the method is Receiving a bitstream containing an encoded audio signal and encoder-generated dynamic range control (DRC) metadata, wherein the encoder-generated DRC metadata includes a plurality of sets of DRC gains, the plurality of sets of DRC gains including a first set of DRC gains representing a first portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal, and a second set of DRC gains representing a second portion of the total DRC gain to be applied to the audio signal to adjust the dynamic range of the audio signal. Decoding the encoded audio signal to obtain the audio signal, Downmixing the aforementioned audio signal, This includes adjusting the dynamic range of the audio signal by applying a second set of the DRC gains to the audio signal after the downmix in order to apply the total DRC gains to be applied to the audio signal, Non-temporary computer-readable storage medium.

Citation Information

Patent Citations

  • Encoding / decoding device and method

    JP2009526259A

  • A system for maintaining reversible dynamic range control information related to a parametric audio coder.

    JP2015517688A

  • Audio signal processing device and audio signal processing method

    WO2012026092A1