Audio object processing
The method corrects object reconstruction information to align levels in object-based audio coding, addressing level errors and ensuring consistent audio quality across diverse playback configurations by applying correction gains in the bitstream.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-09
- Publication Date
- 2026-03-17
AI Technical Summary
Existing object-based audio coding methods, such as Joint Object Coding (JOC), fail to guarantee that the covariance matrix of reconstructed objects matches the original, leading to level errors like loss or build-up, especially when rendered on different playback configurations.
A method for modifying object reconstruction information by analyzing the rendered presentations of original and processed objects, applying correction gains to align the levels, and including these gains in the bitstream for decoding, ensuring accurate rendering across various playback configurations.
Reduces level errors by aligning the rendering of reconstructed objects with the original, providing consistent audio quality across different playback systems like 5.1.2, 7.1.4, and 9.1.6 configurations.
Smart Images

Figure 0007831905000031 
Figure 0007831905000032 
Figure 0007831905000033
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims priority to the following priority application: U.S. Provisional Application No. 63 / 153,719, filed on 25 February 2021 (reference: D21011USP1) (incorporated herein by reference).
[0002] [Technical field] This disclosure relates to the processing of audio objects, and more particularly to the encoding and decoding of audio objects. [Background technology]
[0003] Object-based representation of immersive audio content is a powerful technique that combines intuitive content creation with optimal playback across a wide range of playback configurations using appropriate rendering systems. Object-based audio is a key component of systems like Dolby Atmos. Audio objects include the actual audio signal and associated metadata such as the object's location. To deliver object-based audio to consumer entertainment devices, an efficient representation is required that enables broadcast, streaming, download, or similar transmission scenarios. For this purpose, various object processing techniques, such as spatial coding and object coding, are employed.
[0004] One specific coding technique is the Joint Object Coding (JOC) method, as discussed in H. Purnhagen, T. Hirvonen, L. Villemoes, J. Samuelsson, J. Klejsa, “Immersive Audio Delivery Using Joint Object Coding”, in AES 140th Convention, Paris, FR, May 2016. An example of this is the Dolby Digital Plus (DD+) JOC system in “Backwards-compatible object audio carriage using Enhanced AC-3”, ETSI TS 103 420 V1.1.1 (2016-07). As discussed in J. Breebaart, G. Cengarle, L. Lu, T. Mateos, H. Purnhagen, N. Tsingos, “Spatial Coding of Complex Object-Based Program Material,” J. Audio Eng. Soc., vol. 67, no. 7 / 8, pp. 486-497, July 2019, joint object coding can be used in conjunction with spatial coding as a preprocessor to reduce the number of objects that need to be transmitted.
[0005] In a JOC encoder, objects are rendered to a downmix signal, e.g., a 5.1 surround representation, and JOC parameters are calculated to enable a JOC decoder to reconstruct the objects from the downmix signal. The JOC encoder sends the downmix signal, JOC parameters, and object metadata to the JOC decoder. Typically, object-based content contains more objects than the number of downmix signals, thus enabling more efficient transmission. Furthermore, the downmix signal itself can be efficiently transmitted using perceptual audio coding systems such as DD+. Typically, JOC parameters control how objects are reconstructed as a linear combination of the downmix signal, and JOC parameters vary over time and frequency and are transmitted per time / frequency (T / F) tile. A common initial technique for calculating JOC parameters for a given object within a given T / F tile is to achieve the best approximation in terms of least mean squared error (MMSE). However, if an exact reconstruction is not possible, the approximation error means that the reconstructed object will have a lower level (measured as energy or variance). To achieve a more perceptually appropriate approximation, it is advantageous to boost (i.e., gain) the reconstructed object so that it has the same level (i.e., energy) as the original object, and this boost can be achieved by appropriately changing the JOC parameters.
[0006] However, this method does not guarantee that the complete covariance matrix of the reconstructed objects will match the covariance matrix of the original objects. Only the diagonal elements of the covariance matrix (i.e., object energies) are guaranteed to be correctly restored. Often, an increase in correlation between the reconstructed objects can be observed, resulting in a level build-up effect when the reconstructed objects are rendered, for example, for playback by a 7.1.4 loudspeaker system. This build-up can be observed when compared to the rendering of the original objects and may manifest, for example, as an increase in the perceived loudness of the objects in the affected content. [Overview of the project]
[0007] The object of the present invention is to improve the processing of audio objects, which includes avoiding level errors such as level loss and level build-up in object coding.
[0008] According to a first aspect of the present invention, this and other objectives are achieved by a method for modifying object reconstruction information, the method comprising the steps of: obtaining a set of N spatial audio objects, each spatial audio object comprising an audio signal and spatial metadata; obtaining an audio presentation representing the N spatial audio objects; obtaining object reconstruction information configured to reconstruct the N spatial audio objects from the audio presentation; applying the reconstruction information to the audio presentation to form a set of N reconstructed spatial audio objects; rendering the N spatial audio objects using a first rendering configuration to obtain a first rendered presentation; rendering the N reconstructed spatial audio objects to obtain a second rendered presentation; and modifying the reconstruction information based on the difference between the first rendered presentation and the second rendered presentation to thereby form modified reconstruction information.
[0009] By analyzing (comparing) the rendered presentations of the original object and the processed object, the reconstruction information can be corrected, thereby making the rendering of the reconstructed object more closely match the rendering of the original object.
[0010] In some embodiments, the method according to the first aspect is used for audio object coding. In this case, the audio presentation is a set of M audio signals encoded into a set of encoded audio signals, and the encoded audio signals and the modified reconstruction information are combined into a bitstream for transmission. In a more specific example, the M audio signals represent the downmix of the audio signals of N spatial audio objects, the object reconstruction information is a set of reconstruction parameters configured to reconstruct N spatial audio objects from the M audio signals, and the modified reconstruction information is a set of modified reconstruction parameters.
[0011] In these embodiments, the decoding process may remain unchanged, but will use the modified reconstruction information transmitted in the bitstream. This reduces, for example, the level errors that occur when unmodified reconstruction parameters are used on the decoder side.
[0012] The method further includes using a second rendering configuration to render N spatial audio objects to generate a third rendered presentation, rendering N reconstructed spatial audio objects to generate a fourth rendered presentation, determining a second set of object-specific modification gains associated with the second rendering configuration, and including in the encoded bitstream 1) both the first and second sets of object-specific modification gains and 2) one of the ratios of the first set of object-specific modification gains to the second set.
[0013] In this approach, the encoded bitstream includes information that enables a decoder on the receiving side to obtain the modified reconstructed objects associated with one of a plurality of rendering configurations, for example, 5.1.2 or 7.1.4.
[0014] According to a second aspect of the present invention, this and other objectives are achieved by a method for decoding spatial audio objects in a bitstream, the method comprising the steps of decoding a bitstream to obtain a set of M audio channels and a set of reconstruction parameters configured to reconstruct a set of N spatial audio objects from the M audio signals, wherein the reconstruction parameters include a set of reconstruction parameters associated with a first rendering configuration and a modified gain associated with a second rendering configuration. The method further includes the steps of determining a playback rendering configuration, applying a modified gain to the reconstruction parameters to obtain alternative reconstruction parameters in response to determining the playback rendering configuration, and applying the alternative reconstruction parameters to the M audio signals to obtain a set of N reconstructed spatial audio objects.
[0015] For example, if it is determined that the playback rendering configuration corresponds to a second rendering configuration, a modification gain can be applied so that the alternative reconstruction parameters are associated with the second rendering configuration.
[0016] In one example, the modification gains include a first set of object-specific modification gains associated with a first rendering configuration and a second set of object-specific modification gains associated with a second rendering configuration, and the step of applying the modification gains to the reconfiguration parameter includes the step of applying the first set of modification gains to remove the association of the reconfiguration parameter with the first rendering configuration and the step of applying the second set of modification gains to associate the reconfiguration parameter with the second rendering configuration.
[0017] In another example, the modification gains include a set of ratios h(n) / h2(n) between a first object-specific modification gain h(n) associated with the first rendering configuration and a second object-specific modification gain h2(n) associated with the second rendering configuration.
[0018] A further aspect of the present invention relates to an encoder, which includes: a downmix renderer configured to receive a set of N spatial audio objects and generate a set of M audio signals representing the N spatial audio objects; an object encoder for obtaining object reconstruction information configured to reconstruct N spatial audio objects from the M audio signals; an object decoder for applying the reconstruction information to the M audio signals to form a set of N reconstructed spatial audio objects; a renderer configured to render the N spatial audio objects using a first rendering configuration to obtain a first rendered presentation and to render the N reconstructed spatial audio objects to obtain a second rendered presentation; a modifier for modifying the reconstruction information based on the difference between the first rendered presentation and the second rendered presentation, thereby forming modified reconstruction information; an encoder configured to encode the M audio signals into a set of encoded audio signals; and a multiplexer for combining the encoded audio signals and the modified reconstruction information into a bitstream for transmission.
[0019] A further aspect of the present invention relates to a decoder configured to reconstruct a set of M audio channels and a set of N spatial audio objects from the M audio signals using a reconstruction parameter c mod The decoder includes a set of (n,m) and a set of reconstruction parameters, the set of reconstruction parameters being associated with a first rendering configuration, and a set of modification gains associated with a second rendering configuration, for decoding a bitstream. The decoder applies the modification gains to the reconstruction parameters c in response to the determined re-rendering configuration. mod Apply (n,m) to the alternative reconstruction parameter c mod2An alternative unit configured to obtain (n,m), and an alternative reconfiguration parameter c mod2 Includes an object decoder for applying (n,m) to M audio signals to obtain a set of N reconstructed spatial audio objects.
[0020] A further embodiment includes a computer program product which includes a portion of computer program code configured to perform the methods according to the first and second embodiments when executed on a computer processor. [Brief explanation of the drawing]
[0021] The present invention will be described in more detail with reference to the accompanying drawings illustrating currently preferred embodiments of the invention. [Figure 1] The first implementation form of the present invention is shown. [Figure 2a] This shows an encoding system including further implementations of the present invention. [Figure 2b] This document shows a decoding system including further implementations of the present invention. [Figure 3A] This is a flowchart of the encoding process according to one implementation of the present invention. [Figure 3B] This is a flowchart of the decoding process according to one implementation of the present invention. [Figure 4a] This document describes an encoding system including yet another implementation of the present invention. [Figure 4b] This document shows a decoding system including yet another implementation of the present invention. [Figure 5a] This document describes an encoding system including yet another implementation of the present invention. [Figure 5b] This document shows a decoding system including yet another implementation of the present invention. [Modes for carrying out the invention]
[0022] Although not explicitly mentioned in the following description, those skilled in the art will understand that all signals are typically divided into time (frames) and frequency (bandwidth), and therefore processing is performed in time-frequency tiles. For the sake of simplicity of notation, time and frequency dependencies are excluded from the description.
[0023] Furthermore, in the following disclosures, “object,” “audio object,” or “spatial audio object” should be understood to include an audio signal and associated metadata, including spatial rendering information. overview Front mounting
[0024] A rendering configuration is a set of rules that, given metadata about a spatial audio object, such as object position, yield a rendering gain g(k,n) that describes how much an object signal S(n) contributes to a rendering signal L(k). The set of rendering signals L(k), k=1,...,K is called the rendered representation of the set of objects S(n), n=1,...,N, or simply the rendition of the set of objects. The rendition of the original set of objects S(n), n=1,...,N is called the original rendition, and the rendition of the processed set of objects is called the processed rendition. Similarly, the rendition of the modified (level-aligned) set of objects is called the modified rendition.
[0025] The calculation of the original renditions L(k), k=1,...,K can be expressed based on the following equation.
number
number
number
number
number
[0026] The goal of level alignment is to compute the modified object such that, given the original object and the modified object, the rendered representation calculated from the modified object (modified rendition) exhibits a rendering signal level as close as possible to the level of the rendered representation from the original object (original rendition).
[0027] To enable level alignment while preserving the object's properties as much as possible, a modification gain h(n) is applied to the object. Modified object S M (n) is,
number
number
[0028] The following presents a method for calculating the modified gain h(n). The energy of the signal and the cross-correlation between signals are calculated as part of these methods. The energy of an object is calculated based on
Number
Number
[0029] First, the MSE method that minimizes the M mean square error
Number
Number
[0030] A modified MMSE method that avoids the latter phenomenon predicts the target L(k) as f(k)L PThis can be obtained by substituting (k), where f(k) is the rendering signal alignment gain aimed at obtaining the desired output level. Gain distribution method
[0031] Alternatively, the signal energy of the original rendition ||L(k)|| 2 and the signal energy ||L of the processed rendition P (k)|| 2 These are each calculated, and the rendering signal alignment gain f(k) is calculated based on the following equation.
number
[0032] From the rendering signal alignment gain, the object correction gain can be calculated based on the following formula:
number
[0033] In other words, the modified gain h(n) is calculated as a weighted sum of the alignment gains f(k), where the sum of the weights over all k for any given n is 1. This can be described as the distribution of the alignment gains according to the weights (which are determined from the rendering gains) for obtaining the modified gains. If the processed objects are uncorrelated, these gains are exactly the same as those obtained by the modified MMSE method described in the previous section.
[0034] An alternative example for calculating the corrected gain is given by the following equation:
number
[0035] The deviation of the rendering signal k, i.e., f(k) ≠ 1, affects the object in proportion to the object's contribution to that rendering signal. Furthermore, in all of these equations, the desired effect ||L| is obtained when the object is not rendered to two or more rendering signals, i.e., when at most one of the rendering gains g(k,n), k=1,...,K is non-zero for each n=1,...,N. p (k)|| 2 =||L p (k)|| 2 This will achieve...
number
[0036] For example, adjust the corrected gain.
number
[0037] Modified Rendition Energy || L M (k)|| 2 These are monitored, and their energy ||L(k)|| 2If not close enough, the overall gain g is the same for all objects so that the total energy of the modified rendition is equal to the total energy of the original rendition. overall There may be advantages to the second processing step, where this can be applied. Specifically,
number
number
number
number
number
[0038] In many cases, the threshold is the energy ||L(k)|| of the original rendering signal. 2 It is a function of, for example, the following:
number
[0039] In the above monitoring and threshold calculation of the energy of the modified rendition, the energy of the processed rendition ||Lp(k)|| 2 The original rendition energy ||L(k)|| 2It can be used instead. It may seem pointless, but the gain distribution method can, for some set of objects, yield a modified rendering signal energy that deviates from the original rendering signal energy, rather than the processed rendering signal energy. Recursive gain distribution
[0040] In some use cases, it may be beneficial to perform the above process recursively. Energy ||L of the modified rendition M (k)|| 2 These quantities can be fed back into a recursive process in which they are calculated based on the following:
number
number
[0041] In situations where audio objects are encoded to be included in a bitstream, the encoder may calculate a correction gain and transmit it to the decoder, where playback rendering takes place.
[0042] In one example, the original object is a set of downmix signals Y(m) and reconstruction parameters.
number
number
number
number
[0043] The so-called nominal rendering configuration used for level analysis and level correction may differ from the playback rendering configuration. For example, the playback rendering configuration on the decoder side may not be known at the time of encoding.
[0044] In many practical cases, the methods presented herein are robust to differences in rendering configurations for the rendering configurations actually relevant (e.g., 5.1.2, 5.1.4, 7.1.4, 9.1.6). By calculating the corrected gain using the nominal rendering configuration of 7.1.4, robust level adjustments are also provided for the rendering configurations of 5.1.2, 5.1.4, and 9.1.6.
[0045] It may be beneficial to calculate the corrected gain for several nominal rendering configurations.
number
[0046] For example, when J=4, these rendering configurations can be, for example, 5.1.2, 5.1.4, 7.1.4, and 9.1.6, where h1(n), n=1,...,N are the modification gains associated with the rendering configuration 5.1.2, h2(n), n=1,...,N are the modification gains associated with 5.1.4, and so on. A common set of modification gains h(n), n=1,...,N can be calculated by combining these sets of gains. This combination can be calculated, for example, as a weighted sum.
number
[0047] If there is a mismatch between the nominal rendering configuration and the replay rendering configuration, and the averaging method does not work, the corrective gain can be stored / sent along with the processed objects or reconstruction parameters. If the replay rendering configuration matches one of the stored nominal configurations, the corresponding corrective gain can be applied "just in time". If there is still a mismatch, the "closest" nominal configuration can be used, or the averaging of the nominal configurations can be used. Practical implementation forms
[0048] Figure 1 shows that N processed (e.g., spatially encoded or decoded and reconstructed) objects S take a set of N* original objects S(n*) as input. P The audio system 100 includes an object processor 101 that produces a set of (n) as output.
[0049] Using object metadata (not shown separately), N* original objects S(n*) and N processed objects S P (n) can be rendered to a nominal replay configuration (e.g., 7.1.4) by two renderers 102 and 103, resulting in rendered representations L(k) and L P(k) is obtained. By analyzing and comparing the levels of both rendered representations in the level analyzer 104, the processed object S P (n) is taken as input, and the modified object S M It is possible to extract information to control the object modifier 105 which generates (n) as output. The renderer 106 renders the modified object and renders the presentation L M (k) is provided. The goal of object modification is the modified object S M (n) Rendered representation L M Object S(n) is introduced and processed by object processor 101, bringing (k) closer to the rendered representation L(k) of the original object S(n). P (n) Rendered representation L P The goal is to mitigate all errors, such as level errors, observed for (k).
[0050] If the object processor is a spatial coder, fewer objects will be processed (N*>N). In a typical spatial coding process, 128 audio objects are clustered into 20 audio objects (N*=128, N=20).
[0051] The object processor 101 in Figure 1 may be a combination of encoder and decoder that occur in the codec process. In this case, N* = N. Figures 2a and 2b show how the principles of the present invention may be implemented in an exemplary encoding and decoding (codec) process 200. The codec may be based on, for example, a Dolby Digital Plus (DD+) codec with Joint Object Coding (JOC). It may also be based on an AC-4 codec with Advanced Joint Object Coding (A-JOC), in which case contributions from an uncorrelated version of the downmix signal are also taken into account. The A-JOC encoder may alternatively use a downmix generated by a spatial coder instead of a downmix renderer.
[0052] The encoder side 201 (Figure 2a) includes a downmix renderer 202, a downmix encoder 203, an object encoder 204, and a multiplexer 205. In one example, blocks 202, 203, 204, and 205 are substantially equivalent to their corresponding blocks in the DD+JOC encoder.
[0053] In the illustrated example, encoder 201 further comprises object decoder 206 (e.g., JOC decoder) and two renderers 207, 208. The object decoder processes the object S P To generate (n), object reconstruction parameters c(n,m) from object encoder 204 are used to decode the downmix Y(m) from downmix renderer 202. Renderers 207 and 208 are configured to decode the original object S(n) and the processed object S, respectively. P (n) is received and the selected playback rendering configuration, for example, configuration 7.1.4, is used to render the first rendered presentation L(k) and the second rendered presentation L PThe system is configured to use object metadata (not shown separately) to provide (k). The selected rendering configuration is called the “nominal” rendering configuration. The level analyzer 209 is configured to use the rendering L(k) and L(k) rendered from each renderer 207, 208. P The parameter modifier 210 is configured to receive (k) and provide a set of parameters h(n) (one parameter for each object) that represent the difference between the two rendered presentations. The parameter modifier 210 is configured to receive the parameter h(n) and perform a modification of the reconstruction parameter c(n,m). The modified reconstruction parameter is c mod It is called (n,m).
[0054] The decoder side 211 (Figure 2b) includes a demultiplexer 212, a downmix decoder 213, and an object decoder 214. In one example, blocks 212, 213, and 214 are substantially equivalent to their corresponding blocks in a DD+ JOC decoder. The output from the decoder side 211 is provided to the replay renderer 221.
[0055] During use, referring to Figure 3, the original set of objects S(n) is first rendered in the downmix renderer 202 to generate the downmix signal Y(m) (step S1). In a typical encoder, a 5.1 configuration is used for downmixing, and the downmix rendering uses object metadata (not shown). Both the original objects S(n) and the downmix signal Y(m) are used by the object encoder 204 to calculate the reconstruction parameter c(n,m) (step S2). The downmix signal is also encoded by the downmix encoder 203 (step S3).
[0056] In parallel with step S3, the object decoder 206 takes the downmix signal Y(m) as input and processes (i.e., reconstructs) the object S P(n) is generated (step S4). Then, the original object S(n) and the processed object S P Both (n) are rendered (step S5), and the first rendered representation L(k) and the second rendered representation L P (k) is obtained for each. Next, both rendered representations are analyzed (step S6) and a set of parameters h(n) called object modification gains is calculated. In step S7, the parameter modifier 210 applies the object modification gains h(n) to the reconstruction parameter c(n,m) to obtain the modified reconstruction parameter c mod Generate (n,m).
[0057] In step S8, the encoded downmix is converted in the multiplexer to the modified reconstruction parameter c mod (n,m) and object metadata (not shown) are combined to form the final bitstream. This bitstream is then sent to the decoder 211 (step S9).
[0058] On the decoder side, the bitstream is demultiplexed by the demultiplexer 212 (step S11), decoded by the downmix decoder 213 to obtain the downmix signal Y(m) (step S12). These downmix signals Y(m) are modified with the reconstruction parameter c mod Using (n,m), the object S is processed by the object decoder 214 and modified. M (n) is generated (step S13).
[0059] Finally, the modified object S M (n) is a representation L for a desired playback configuration (e.g., 7.1.4 loudspeaker playback) in a playback renderer 221 that uses object metadata (not shown) transmitted in a bitstream. M (k) is rendered (step S14).
[0060] Referring to Figures 4a and 4b, the encoding side (Figure 4a) also includes a spatial coder 231 configured to perform reduction (clustering) of the original set of N* audio objects. In a typical example, 128 original audio objects are spatially coded into 20 objects before being provided to the object encoder process. In the illustrated case, as an alternative to the process in Figures 2a and 2b, the original audio objects S(n*) (e.g., 128 objects) are used by the renderer 207 to obtain a first rendition L(k).
[0061] Figures 5a and 5b illustrate yet another implementation of the present invention, in which multiple sets of object-specific modification gains h1(n) and h2(n) are determined, and a set of modification parameters based on these sets of modification gains is made available to the decoder. In the illustrated example, only two sets of object-specific modification gains exist, but of course, any number may exist.
[0062] In this implementation, the renderers 307 and 308 on the encoder side 301 (Figure 5a) are configured to perform multiple renditions associated with multiple rendering configurations. In the illustrated case, two renditions are provided. These may be associated, for example, with configurations 7.1.4 and 9.1.6. The level analyzer 309 performs a level analysis on each pair of renditions, resulting in two sets of object-specific modification gains h1(n) and h2(n). One of the gain sets is used by a parameter modifier to modify the reconstruction parameter c(n,m). In addition to the encoded downmix Y(m) and the modified reconstruction parameter, the multiplexer 205 is also provided with a set of modified parameters based on the two sets of modification gains h1(n) and h2(n), so these modified parameters are also included in the bitstream.
[0063] Decoder 311 (Figure 5b) contains elements similar to those of decoder 211 in Figures 2b and 4b. These elements are given the same reference numbers (212, 213, 214, 221) in Figure 5b. Decoder 311 also includes an alternate block 312 configured to apply modified parameters to the original reconstruction parameters in order to obtain an alternate set of modified reconstruction parameters. This alternate set of modified reconstruction parameters may correspond to a second rendering configuration. The operation of alternate block 312 is optional and controlled by appropriate logic. For example, the activation of alternate block 312 may be based on a configuration decision of the replay renderer 221.
[0064] In the first example shown in Figure 5b, the modification parameters include two sets of object-specific modification gains, h1(n) and h2(n). In this case, the alternative block 312 includes the following two units: 1) An undo unit 313 configured to apply a first set of gains h1(n) (inverse) to return the reconfiguration parameters to their original "unmodified" state, and 2) A gain applying unit 314 is configured to apply a second set of gains h2(n) to the “unmodified” reconstruction parameters in order to obtain an alternative set of modified reconstruction parameters corresponding to the second rendering configuration.
[0065] It is clear that the implementation shown in Figure 5B provides three different object decoding options. 1) Modified reconstruction parameter c mod (n,m) is used to provide a modified and reconfigured object for improved rendering by the first rendering configuration. 2) Provide a modified and reconstructed object for improved rendering by a second rendering configuration, using alternative modified reconstruction parameters. 3) Use the “unmodified” reconfiguration parameter to provide an object that has been reconfigured without modification.
[0066] In another example, the modification parameters include the ratio h2(n) / h1(n) of a second set of object-specific modification gains h2(n) to the first set h1(n). In this case, on the decoder side, these ratios can be applied to the modified reconstruction parameters corresponding to the first rendering configuration to achieve a conversion to alternative modified reconstruction parameters corresponding to the second rendering configuration.
[0067] In this case, the following two alternative decoding options are available on the decoder side: 1) Modified reconstruction parameter c mod (n,m) is used to provide a modified and reconfigured object for improved rendering by the first rendering configuration. 2) Provide a modified and reconstructed object for improved rendering by the second rendering configuration, using alternative modified reconstruction parameters.
[0068] However, a special case in this particular example is that a second set of modification gains h2(n) can be set to correspond to the unity gain, i.e., the unmodified reconstruction parameter. In other words, the modified parameter in the bitstream becomes 1 / h1(n). On the decoder side, applying these gains cancels out the modification gains h1(n), thus providing the original "unmodified" reconstruction parameter.
[0069] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on a digital signal processor or microprocessor. Other components may be implemented as hardware and / or as application-specific integrated circuits. Signals encountered in the methods and systems described herein may be stored on a medium such as random-access memory or optical storage media. They may be transmitted over a network such as a wireless network, satellite network, wireless network, or wireline network, such as the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer equipment used to store and / or render audio signals.
[0070] Unless otherwise specified, as is evident from the following description, throughout this disclosure, any use of terms such as “process,” “calculate,” “calculate,” “determine,” and “analyze” is understood to refer to the actions and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or convert data, which is represented as a physical quantity such as an electronic quantity, into other data, which is similarly represented as a physical quantity.
[0071] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together with a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various embodiments of the invention. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly described in each claim. Rather, as reflected in the following claims, the embodiments of the invention consist of fewer features than all of the features of a single, aforementioned disclosed embodiment. Thus, the claims following the embodiments for carrying out the invention are expressly incorporated into the embodiments for carrying out the invention, and each claim stands independently as a separate embodiment of the invention. Furthermore, some embodiments described herein include some features included in other embodiments, but not others, and as will be understood by those skilled in the art, combinations of features of different embodiments are within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0072] Furthermore, some of the embodiments described herein are described as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing the function. Thus, a processor having the instructions necessary to perform such a method or element of a method forms a means for performing the method or element of a method. Note that when a method includes several elements, for example, several steps, the order of such elements is not indicated unless specifically stated. Furthermore, the elements of the embodiments of the apparatus described herein are examples of means for performing the function performed by the elements for the purpose of carrying out the present invention. Numerous specific details are described in the description provided herein. However, it is understood that embodiments of the present invention can be carried out without these specific details. In other instances, well-known methods, structures and techniques are not described in detail so as not to obscure the understanding of this description.
[0073] Therefore, although specific embodiments of the present invention have been described, those skilled in the art will recognize that other and further modifications can be made thereto without departing from the spirit of the invention, and it is intended that all such changes and modifications be claimed to fall within the scope of the invention. For example, other object encoding / decoding techniques may be implemented.
[0074] The present invention includes exemplary embodiments (EEE) listed below.
[0075] EEE1. A method for aligning the levels of the original rendition and the processed rendition, Steps to receive the original set of objects, A step of receiving a set of processed objects, A step of receiving a rendering configuration, wherein the rendering configuration describes a mapping from a set of original objects to a set of original rendering signals, and the rendering configuration also describes a mapping from a set of processed objects to a set of processed rendering signals. The steps include: correcting the processed audio object set to align the level of the processed rendering signal set to the level of the original rendering signal set; and A method that includes this.
[0076] EEE2. A step of calculating the level of the original rendering signal set, The steps include: calculating the level of the processed rendering signal set and The method described in EEE1, further including the method described in EEE1.
[0077] EEE3. The step of rendering the original set of objects to the original set of rendering signals, The steps include rendering the processed set of objects into a set of processed rendering signals, The steps include measuring the level of the original rendering signal set, The steps include measuring the level of a set of processed rendering signals and The method described in EEE1, further including the method described in EEE1.
[0078] Aligning EEE4 levels is For each object, calculate the object modification gain and apply the object modification gain to that object. The method described in EEE1, including the method described in EEE1.
[0079] EEE5. A method for aligning the levels of a rendering signal, Steps to receive the original set of objects, A step of receiving a set of processed objects, A step of receiving a rendering configuration, wherein the rendering configuration describes a mapping from a set of original objects to a set of original rendering signals, and the rendering configuration also describes a mapping from a set of processed objects to a set of processed rendering signals. The steps include calculating the optimal set of object correction gains and A method that includes this.
[0080] EEE6. A method for aligning the levels of a rendering signal, Steps to receive the original set of objects, A step of receiving a set of processed objects, A step of receiving a rendering configuration, wherein the rendering configuration describes a mapping from a set of original objects to a set of original rendering signals, and the rendering configuration further describes a mapping from a set of processed objects to a set of processed rendering signals. The steps include calculating the levels of the original rendering signal set, The steps include calculating the level of the processed rendering signal set, The steps include calculating a set of rendering signal correction gains, Distribution of the rendering signal alignment gain set to the object modification gain set and A method that includes this.
[0081] EEE7. The mapping of the set of rendering signal alignment gains to the set of object modification gains is: Step 1: Calculate the modification gain for each object as a weighted sum of the rendering signal alignment gains. The method described in EEE6, including the method described in EEE6.
[0082] EEE8. The weights in a weighted sum are functions of the rendering gain, as described in EEE7.
[0083] EEE9. The method of EEE6, wherein a modification gain is applied to the processed object to obtain a modified object.
[0084] EEE10. The step of rendering the modified object to a set of modified rendering signals, The steps include: calculating the total correction level of the corrected rendering signal, The steps include: calculating the total reference level of the set of reference rendering signals, The steps include calculating the total correction gain from the total correction level and the total reference level, and The method described in EEE9, which further includes the method described in EEE9.
[0085] EEE11. Replace the processed object with the corrected object and repeat the procedure. The method described in EEE9, which further includes the method described in EEE9.
[0086] EEE12. Object modification gain is applied to at least one set of audio object reconstruction parameters, e.g., a set of JOC parameters, according to any of the methods described in EEE4 to 11.
[0087] EEE13. The object correction gain is calculated in the encoder. The object modification gain is applied in the encoder to at least one set of audio object reconstruction parameters, for example, a set of JOC parameters, to obtain modified JOC parameters. The modified audio object reconstruction parameters replace at least one set of audio object reconstruction parameters in the encoder bitstream. The method described in any of EEE4 to 11.
[0088] EEE14. Multiple sets of object modification gains are calculated for multiple rendering configurations. The total set of object modification gains is calculated by combining multiple sets of object modification gains. The method described in any of EEE4 to 13.
[0089] EEE15. The method described in EEE14, wherein the combination is performed by a weighted average of a set of object modification gains.
[0090] EEE16. Multiple sets of object modification gains are calculated for multiple rendering configurations. Multiple sets of object modification gains are stored along with the processed object. The best match set of object modification gains is applied before playback rendering. The method described in any of EEE4 to 15.
[0091] EEE17. A method for decoding an encoded audio bitstream, A step of decoding an encoded audio bitstream in order to obtain multiple decoded audio signals, wherein the multiple decoded audio signals include a multi-channel downmix of multiple audio object signals. A step of extracting multiple sets of audio object reconstruction parameters from an encoded audio bitstream, wherein each set of audio object reconstruction parameters corresponds to a different channel configuration, and The steps to determine the playback rendering configuration, The steps include determining a set of audio object reconstruction parameters from multiple sets of audio object reconstruction parameters based on the determined playback rendering configuration, To obtain the reconstruction of multiple audio object signals, the steps include applying a determined set of audio object reconstruction parameters to multiple decoded audio signals. A method that includes this.
[0092] EEE18. The determined set of audio object reconstruction parameters is the set of audio object reconstruction parameters corresponding to the determined playback rendering configuration, as described in EEE17.
[0093] EEE19. If none of the sets of audio object reconstruction parameters correspond to a channel configuration that matches the determined playback rendering configuration, the set of determined audio object reconstruction parameters corresponds to the channel configuration that is closest to the determined playback rendering configuration, as described in EEE17.
[0094] EEE20. If none of the sets of audio object reconstruction parameters match the determined playback rendering configuration, the determined set of audio object reconstruction parameters corresponds to the average of the sets of audio object reconstruction parameters, as described in EEE17.
[0095] EEE21. The average is a weighted average, as described in EEE20.
[0096] EEE22. A method according to any one of EEE17 to 21, further comprising the steps of extracting object metadata from an encoded bitstream and rendering a reconstruction of multiple audio object signals in response to the object metadata into a determined playback rendering configuration.
[0097] EEE23. A method for decoding an encoded audio bitstream, A step of decoding an encoded audio bitstream in order to obtain multiple decoded audio signals, wherein the multiple decoded audio signals include a multi-channel downmix of multiple audio object signals. The steps include extracting a set of audio object reconstruction parameters from an encoded audio bitstream, To obtain the reconstruction of multiple audio object signals, the steps include applying a set of audio object reconstruction parameters to multiple decoded audio signals. Includes, Multiple reconstruction parameters were calculated according to the EEE13 method. method.
[0098] EEE24. The method according to EEE23, further comprising the steps of extracting object metadata from an encoded bitstream and rendering a reconstruction of multiple audio object signals into a playback rendering configuration in response to the object metadata.
Claims
1. A method for modifying object reconstruction information, A step of obtaining a set of N spatial audio objects, where each spatial audio object includes an audio signal and spatial metadata, The steps include obtaining an audio presentation representing the N spatial audio objects, The steps include obtaining object reconstruction information configured to reconstruct the N spatial audio objects from the audio presentation, The steps include applying the object reconstruction information to the audio presentation to form a set of N reconstructed spatial audio objects, The steps include rendering the N spatial audio objects using a first rendering configuration to obtain a first rendered presentation, and rendering the N reconstructed spatial audio objects to obtain a second rendered presentation, The steps include modifying the object reconstruction information based on the difference between the first rendered presentation and the second rendered presentation, thereby forming the modified reconstruction information, The aforementioned audio presentation is a set of M audio signals, The steps include encoding the M audio signals to form a set of encoded audio signals, The further step includes combining the encoded audio signal and the modified reconstruction information into a bitstream for transmission. The M audio signals represent the downmix of the audio signals of the N spatial audio objects, the object reconstruction information is a set of reconstruction parameters c(n, m) configured to reconstruct the N spatial audio objects from the M audio signals, and the modified reconstruction information is a set of modified reconstruction parameters c mod (n, m). The modification step comprises determining a set of object-specific modification gains h1(n) associated with the first rendering configuration, wherein the object-specific modification gains h1(n) are applied to a set of reconstruction parameters c(n, m), in a method.
2. The method according to claim 1, wherein the set of N spatial audio objects is obtained by spatially coding a set of L spatial audio objects, where L > N, and the first rendered presentation is obtained by rendering the L spatial audio objects.
3. The object-specific modification gain h 1 (n) is, Determining a first level of the first rendered presentation, Determining the second level of the second rendered presentation, Calculating a set of level alignment gains based on the difference between the first level and the second level, The object-specific correction gain h is a linear combination of the level alignment gains. 1 Forming (n) and The method according to claim 1, as determined by...
4. Each object's unique modification gain h 1 The method according to claim 3, further comprising the step of calculating (n) as a weighted sum of the level alignment gains, wherein the weights in the weighted sum are optionally a function of the rendering gains used to generate the first rendered presentation and the second rendered presentation.
5. The steps include rendering the N spatial audio objects using a second rendering configuration to generate a third rendered presentation, and rendering the N reconstructed spatial audio objects to generate a fourth rendered presentation, A second set of object-specific modification gains h associated with the second rendering configuration. 2 The step of determining (n), During the encoding bitstream, 1) A first set of object-specific modification gains h 1 (n) and the second set h 2 Both of (n), and 2) The ratio h of the second set to the first set of object-specific correction gains 2 (n) / h 1 (n) A step that includes one of the following The method according to any one of claims 3 to 4, further comprising:
6. A decoding method for decoding spatial audio objects in a bitstream, Decode the aforementioned bitstream, A set of M audio signals, A reconstruction parameter c configured to reconstruct a set of N spatial audio objects from the M audio signals mod A set of (n, m), wherein the reconstruction parameter is a set of reconstruction parameters associated with a first rendering configuration, and The change parameters associated with the second rendering configuration and Steps to obtain, The steps to determine the playback rendering configuration, In response to determining the aforementioned playback rendering configuration, the modified parameter is set to the reconstruction parameter c mod Apply to (n, m) to the alternative reconstruction parameter c mod2 Steps to obtain (n, m), The alternative reconstruction parameter c mod2 The steps include applying (n, m) to the M audio signals to obtain a set of N reconstructed spatial audio objects, and Decryption methods including [specific methods].
7. The aforementioned playback rendering configuration is determined to correspond to the second rendering configuration, and the alternative reconstruction parameter c mod2 The decoding method according to claim 6, wherein the modification parameters are applied such that (n, m) is associated with the second rendering configuration.
8. The alternative reconstruction parameter c mod2 (n, m) is the reconstruction parameter c mod A set of (n, m) and the reconfiguration parameter c after applying the modified parameters. mod The decoding method according to claim 6, wherein the modification parameters are partially applied to correspond to a weighted average with a set of (n, m).
9. The aforementioned modification parameter is a second object-specific modification gain h associated with the second rendering configuration. 2 (n) and the first object-specific modification gain h associated with the first rendering configuration 1 (n) ratio h 2 (n) / h 1 A decoding method according to any one of claims 6 to 8, comprising a set of (n).
10. The modification parameter is a first set of object-specific modification gains associated with the first rendering configuration h 1 (n) and a second set h of object-specific modification gains associated with the second rendering configuration. 2 (n) and The step of applying the modified parameters to the reconfiguration parameters is: The steps include applying a first set of object-specific modification gains to remove the association of the reconstruction parameters with the first rendering configuration, The steps of relating the reconstruction parameters to the second rendering configuration by applying a second set of object-specific modification gains, and The decoding method according to any one of claims 6 to 8.
11. It is a decoder, A set of M audio signals Reconstruction parameter c configured to reconstruct a set of N spatial audio objects from the M audio signals. mod A set of (n, m), wherein the reconstruction parameters are associated with a first rendering configuration, and Modified gain associated with the second rendering configuration and A decoder for decoding a bitstream containing, In response to the determined playback rendering configuration, the modified gain is adjusted to the reconstruction parameter c mod Apply to (n, m) to the alternative reconstruction parameter c mod2 An alternative unit configured to obtain (n, m), The alternative reconstruction parameter c mod2 An object decoder for applying (n, m) to the M audio signals to obtain a set of N reconstructed spatial audio objects, A decoder that includes this.
12. A computer program that causes a computer to perform the method described in any one of claims 1 to 5.
13. A computer program that causes a computer to perform the method described in any one of claims 6 to 10.
Citation Information
Patent Citations
Apparatus and method for constructing a multi-channel output signal or for generating a downmix signal
US20050157883A1
Scalable downmix design for object-based surround codec with cluster analysis by synthesis
US20140023197A1
Apparatus and method for spatial audio object coding employing hidden objects for signal mixture manipulation
US20150348559A1
Reconstruction of Audio Scenes from a Downmix
US20190311724A1
Exploiting metadata redundancy in immersive audio metadata
WO2015150480A1